A water faucet start-stop control method and system based on somatosensory interaction

CN122176424BActive Publication Date: 2026-08-21HESHAN GAOYANG SANITARY WARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610624141.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-08-21
Estimated Expiration
2046-05-08

AI Technical Summary

Technical Problem

水龙头近场区域同时存在用户手部、水流、金属反光、水渍反光和泡沫遮挡,普通视觉摄像头采集的图像容易出现亮斑、深度缺失和轮廓断裂,导致手部目标检测不稳定;现有手势识别方法多关注完整手势轨迹或二维图像特征,难以在打泡沫、冲洗和短时遮挡过程中持续获得可靠的手腕关键点三维空间坐标;当泡沫或水流遮挡导致手腕关键点短时丢失时,传统感应控制通常直接判定用户离开并关闭电磁阀,容易出现误关水、频繁启停和状态跳变;同时,多人靠近水龙头或手部交叉时,现有方法缺少基于出水口正下方基准点的目标手腕锁定机制,容易将非目标手部误作为控制对象;此外,现有控制逻辑一般只依据单帧感应结果或简单手势结果切换状态,缺少对遮挡持续时间、轨迹运动参数和上一控制状态的联合判定,难以区分正常出水状态、打泡沫悬浮状态和真实离开状态,影响水龙头启停控制的稳定性和使用体验

Benefits of technology

本发明通过偏振去反光处理和深度置信度约束,对水龙头近场区域内的金属反光像素、水渍反光像素和无效深度位置进行识别、替换与填补,解决普通视觉采集在水流反光和深度缺失下图像不稳定的问题,提高带有深度信息的图像帧质量;通过抗遮挡定制化YOLOv8-pose模型删除全身骨骼关键点输出、保留手腕关键点检测头,并结合深度坐标回归分支、遮挡置信度输出分支和泡沫遮挡区域权重,解决泡沫遮挡下手腕关键点检测易丢失、二维手势信息不足的问题,提高了手腕关键点三维空间坐标的连续性和可信度;通过自适应轨迹补全、死区阈值滤除搓洗抖动位移以及Z轴相对位移带符号累积,解决泡沫遮挡、搓洗抖动和真实靠近或远离动作难以区分的问题,使方向累积值能够更准确反映用户启停意图;通过半马尔可夫持续时间约束结合遮挡观测参数、轨迹运动参数和上一控制状态,区分正常出水状态、打泡沫悬浮状态和真实离开状态,并通过暂停记忆态屏蔽打泡沫期间的误关闭控制意图,解决传统感应水龙头频繁误关水和状态跳变的问题;通过目标手腕关键点锁定和超时熔断,降低多人干扰下的误触发风险。其重要意义在于,使水龙头启停控制从单帧感应式切换提升为面向真实洗手过程的连续状态理解控制,提高无接触用水的稳定性、安全性和交互自然性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176424B_ABST
    Figure CN122176424B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on somatosensory interaction's faucet start-stop control method and system, including the following steps: obtaining three-dimensional video stream, and generating image frame with depth information by polarization anti-reflection processing;Image frame is input into anti-shielding customized YOLOv8-pose model, and the three-dimensional space coordinates of wrist key point and shielding confidence are generated;According to shielding confidence, adaptive trajectory completion is executed, and continuous hand space trajectory sequence is generated;Z-axis coordinate sequence is intercepted and is filtered out by washing and shaking, and direction cumulative value is generated;Based on semi-Markov duration constraint, shielding state retention determination is carried out;According to direction cumulative value and shielding state retention determination result, control intention is generated, and solenoid valve start-stop is controlled;In multi-wrist scene, target wrist key point is locked, and through timeout fuse reset, off state is closed.The application improves the stability, anti-interference and use safety of somatosensory start-stop control of faucet in complex water use scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent bathroom control technology, and in particular to a faucet start / stop control method and system based on motion-sensing interaction. Background Technology

[0002] With the increasing demand for intelligent public sanitation fixtures and contactless water supply, motion-sensing interactive control technology for faucet operation has received widespread attention. Existing smart faucets mainly rely on infrared sensors, visual camera gesture recognition, or simple distance detection to control the solenoid valve's operation, switching the water flow status when the user approaches, leaves, or makes specific gestures. However, in actual handwashing scenarios, the following problems commonly exist: The near-field area of ​​a faucet is simultaneously obscured by the user's hand, water flow, metallic reflections, water stain reflections, and foam. Images captured by ordinary vision cameras are prone to bright spots, lack of depth, and broken contours, leading to unstable hand target detection. Existing gesture recognition methods often focus on complete gesture trajectories or 2D image features, making it difficult to continuously obtain reliable 3D spatial coordinates of wrist key points during foaming, rinsing, and short-term obstruction. When foam or water flow obstructs the wrist key point and causes a temporary loss, traditional sensor control usually directly determines that the user has left and closes the solenoid valve, which is prone to accidental water shut-off, frequent start-stop, and state jumps. At the same time, when multiple people approach the faucet or their hands are crossed, existing methods lack a target wrist locking mechanism based on a reference point directly below the water outlet, which can easily mistake non-target hands for control objects. In addition, existing control logic generally only switches states based on single-frame sensing results or simple gesture results, lacking joint determination of obstruction duration, trajectory motion parameters, and the previous control state, making it difficult to distinguish between normal water flow, foaming and floating states, and actual departure states, affecting the stability of faucet start-stop control and user experience.

[0003] Therefore, how to provide a faucet start-stop control method and system based on motion-sensing interaction is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] One objective of this invention is to propose a faucet start-stop control method and system based on motion-sensing interaction. This invention fully utilizes three-dimensional visual perception, anti-occlusion customized YOLOv8-pose model, trajectory completion algorithm, and finite state machine control technology. It describes in detail the process of realizing intelligent faucet start-stop control in scenarios with foam occlusion, water flow reflection, and interference from multiple people's hands. It has the advantages of high recognition stability, low false trigger rate, good water-saving effect, and high safety in use.

[0005] A faucet start / stop control method based on motion-sensing interaction according to an embodiment of the present invention includes the following steps: Step 1: Acquire a 3D video stream containing the user's hand and water flow, and generate image frames with depth information through polarization de-reflection processing; Step 2: Input the image frame with depth information into the anti-occlusion customized YOLOv8-pose model. The anti-occlusion customized YOLOv8-pose model deletes the output of the whole body skeleton key points, retains the wrist key point detection head, configures the depth coordinate regression branch and the occlusion confidence output branch, and adds the foam occlusion region weight to the training loss to generate the three-dimensional spatial coordinates and occlusion confidence of the wrist key points. Step 3: Perform adaptive trajectory completion based on occlusion confidence. When the occlusion confidence is lower than the preset safety threshold, generate virtual coordinates and a continuous hand spatial trajectory sequence based on the historical valid 3D spatial coordinate sequence. Step 4: Extract the Z-axis coordinate sequence from the continuous hand spatial trajectory sequence, filter out the rubbing and shaking displacement according to the dead zone threshold, and perform signed accumulation on the filtered Z-axis relative displacement to generate directional accumulation value; Step 5: Determine the occlusion state based on the semi-Markov duration constraint, and generate the normal water discharge state, foam suspension state, or real departure state according to the occlusion observation parameters, trajectory motion parameters and the previous control state. Step 6: Generate control intent based on the direction accumulation value and the occlusion state determination result, and input the control intent into a finite state machine containing the closed state, water outlet state, and pause memory state to control the start and stop of the solenoid valve; Step 7: When multiple wrist key points are detected, the target wrist key point is locked based on the Euclidean distance from each wrist key point to the reference point directly below the outlet. When the duration of the pause memory state exceeds the preset safe duration, a timeout circuit is triggered, the integral cache is cleared, and the finite state machine is reset to the off state.

[0006] Optionally, step one specifically includes: Acquire the original color image, depth channel image, and depth confidence image of the near-field area of ​​the faucet, and generate a 3D video stream by binding them according to the frame sequence number; A near-field effective region is established based on the reference point directly below the outlet. Orthogonal polarization sampling is performed on the original color image within the near-field effective region to obtain the first polarization intensity image and the second polarization intensity image. Compare the brightness difference at the same pixel location in the first polarization intensity image and the second polarization intensity image, and combine the pixel locations in the depth confidence image that are below a preset depth confidence threshold to generate a reflective pixel mask and an invalid depth marker; The reflective pixels in the original color image are replaced with the median brightness of the neighborhood based on the reflective pixel mask, and the median depth of the neighborhood is filled in the depth channel image based on the invalid depth marker. The original color image after brightness replacement, the depth channel image after depth filling, and the corresponding frame number are re-bound to generate an image frame with depth information.

[0007] Optionally, the anti-occlusion customized YOLOv8-pose model includes a dual-modal input layer, a feature extraction and fusion layer, and a customized decoupled detection head: The bimodal input layer splits the image frame with depth information into an RGB image tensor and a depth map tensor, writes depth pixels and invalid depth markers according to the pixel mapping relationship, and normalizes them to generate bimodal input data; The feature extraction and fusion layer inputs the RGB image tensor into the CSPDarknet backbone network to extract multi-scale visual features, and then inputs the multi-scale visual features into the feature fusion neck network to generate hand color texture response, foam brightness response and wrist contour response; The depth map tensor is input into the depth parallel branch to generate wrist depth boundary response, water flow depth fracture response and invalid depth region response. The responses at the same pixel position are then concatenated by channel dimension and dimensionality reduced by one-to-one convolution to generate a fused feature map. A custom decoupling detection head is used to remove the output of the full-body skeletal key points in the YOLOv8-pose model, while retaining the wrist key point detection head. The wrist key point 2D pixel coordinates are generated through the wrist 2D localization branch. The customized decoupled detection head extracts the depth features in the neighborhood of the two-dimensional pixel coordinates of the wrist key point through the depth coordinate regression branch, removes invalid depth markers and depth values ​​that cross the water flow depth fracture area, and generates the physical depth value of the wrist key point. The two-dimensional pixel coordinates of the wrist key points and the physical depth values ​​of the wrist key points are combined to form the three-dimensional spatial coordinates of the wrist key points. During the training phase, a binary mask of the foam region is generated based on the annotation of the foam-covered pixels. The loss of the key point coordinates corresponding to the foam-covered pixels is multiplied by a preset foam weighting coefficient, and the edge loss corresponding to the wrist contour boundary position is multiplied by a preset edge enhancement coefficient to form a training loss with added weights for the foam-occluded region. The occlusion confidence output branch reads the mean gray-level gradient, the proportion of invalid depth pixels, and the proportion of foam-covered pixels in the candidate area of ​​the wrist, and generates the occlusion confidence through fully connected mapping and Sigmoid activation function.

[0008] Optionally, step three specifically includes: Read the three-dimensional spatial coordinates and occlusion confidence of wrist key points, and mark the three-dimensional spatial coordinates of wrist key points with occlusion confidence not lower than the preset safety threshold and not in an invalid depth position as valid three-dimensional spatial coordinates; The effective three-dimensional spatial coordinates, occlusion confidence and frame number are stored according to the frame sequence number to form a historical trajectory cache queue, and the jump coordinates of adjacent frames whose spatial displacement exceeds the preset single frame displacement limit are deleted. When the occlusion confidence of the current frame is lower than the preset safety threshold, the most recent consecutive valid three-dimensional spatial coordinates in the historical trajectory cache queue are read, the X-axis displacement, Y-axis displacement and Z-axis displacement are calculated, and the median displacement is selected as the complete displacement. The last valid 3D spatial coordinates in the historical trajectory cache queue are superimposed to complete the displacement, generating the predicted spatial position of the current frame, and written as virtual coordinates to the completion mark and the corresponding frame number; When the number of consecutive frames generating virtual coordinates does not exceed the preset maximum number of completion frames, the virtual coordinates are written to the historical trajectory cache queue; when the number of consecutive frames generating virtual coordinates exceeds the preset maximum number of completion frames, completion is stopped and a trajectory interruption flag is output. When a subsequent frame re-obtains valid 3D spatial coordinates and the spatial position difference between the frame and the previous virtual coordinates does not exceed the preset reconnection distance threshold, the re-obtained valid 3D spatial coordinates will be added to the historical trajectory cache queue. The effective three-dimensional spatial coordinates and virtual coordinates are arranged according to the frame number and then subjected to median smoothing to generate a continuous hand spatial trajectory sequence.

[0009] Optionally, step four specifically includes: Read the three-dimensional spatial coordinates of wrist key points in a continuous hand spatial trajectory sequence, and extract the Z-axis coordinate sequence according to the frame number; Set up an integral sliding window, calculate the difference in Z-axis coordinates between two adjacent frames frame by frame, and generate the relative displacement of the Z-axis. The dead zone threshold is generated based on the sorting results of the absolute values ​​of the relative displacements along the Z-axis in the static rubbing calibration samples. The Z-axis relative displacements with absolute values ​​less than the dead zone threshold within the integral sliding window are marked as rubbing and shaking displacements and written to the zero displacement marker. The Z-axis relative displacements with absolute values ​​not less than the dead zone threshold are retained to generate the filtered Z-axis relative displacements. The filtered Z-axis relative displacement is accumulated with a sign according to the frame number. The Z-axis relative displacement away from the outlet is written into the positive accumulation term, and the Z-axis relative displacement closer to the outlet is written into the negative accumulation term. The positive and negative accumulation terms are superimposed to generate the directional accumulation value. Write the filtered Z-axis relative displacement, zero displacement marker, positive cumulative term, negative cumulative term, and direction cumulative value into the integration buffer.

[0010] Optionally, step five specifically includes: Read the occlusion confidence and effective 3D spatial coordinate output results of the current frame. When the occlusion confidence is lower than the preset safety threshold or the current frame does not output effective 3D spatial coordinates, write the occlusion observation flag and update the continuous occlusion duration frame count. Otherwise, write the no-occlusion observation flag and clear the continuous occlusion duration frame count. Within the current decision window, read the X-axis coordinates, Y-axis coordinates, and Z-axis coordinates from the continuous hand spatial trajectory sequence, and generate the mean planar displacement, Z-axis direction continuity, and trajectory completion percentage. Normal water outflow state, foaming and suspension state, and actual departure state are selected as candidate occlusion states, and state transition admission conditions, preset minimum duration frame number, preset maximum duration frame number, duration frame number center value, and duration frame number fluctuation value are configured for each candidate occlusion state. For each candidate occlusion state, a semi-Markov duration constraint calculation is performed. When the number of consecutive occlusion duration frames exceeds the corresponding preset minimum duration frame number and preset maximum duration frame number limit range, the corresponding duration weight is reset to zero. When the number of consecutive occlusions is within the corresponding limit, the duration weight is generated based on the difference between the number of consecutive occlusions and the center value of the number of consecutive frames, and the fluctuation value of the number of consecutive frames. Based on the occlusion observation markers, unocclusion observation markers, previous control state, mean plane displacement, Z-axis continuity, and trajectory completion ratio, the state matching weight is accumulated to generate a matching score for each candidate occlusion state. The matching score of the candidate occlusion state that passes the state transition admission condition is multiplied by the duration weight to generate a state score after semi-Markov constraint. The candidate occlusion state with the highest state score is selected to generate a normal water discharge state, a foaming and floating state, or a real departure state.

[0011] Optionally, step six specifically includes: Read the direction accumulation value and the occlusion state retention determination result, compare the direction accumulation value with the preset positive pause threshold and the preset negative recovery threshold, and write the occlusion state retention determination result into the current state determination bit; When the cumulative value of the direction is not less than the preset positive pause threshold and the current state determination bit is the normal water discharge state, a pause control intention is generated. When the cumulative value of the direction is not greater than the preset negative recovery threshold and the current state determination bit is normal water outflow state or foam suspension state, a water outflow control intention is generated. When the current state determination bit indicates a true departure state, a shutdown control intent is generated; When the current state determination bit is the foaming suspension state and the direction accumulation value has not reached the preset negative recovery threshold, a pause hold control intention is generated; When the finite state machine is in the paused memory state and the occlusion state holds the decision result as the bubble-floating state, the control intention to close is masked and a pause hold control intention is generated. The control intention input is a finite state machine containing the off state, water outlet state and pause memory state. When the pause control intention is received in the water outlet state, it switches to the pause memory state and outputs the off drive signal. When the water outlet control intention is received in the pause memory state, it switches to the water outlet state and outputs the on drive signal. When the off control intention is received, it switches to the off state and outputs the off drive signal. After the finite state machine completes the state transition, the accumulated direction values ​​that have participated in the generation of control intent in the current decision window are cleared, the previous control state corresponding to the pause memory state is retained, and the state transition result is written into the occlusion state retention decision input of the next frame.

[0012] Optionally, step seven specifically includes: Read the 3D spatial coordinates of all wrist key points output by the anti-occlusion customized YOLOv8-pose model in the current frame, and remove wrist key points whose occlusion confidence is lower than the preset safety threshold; Read the spatial coordinates of the reference point directly below the outlet, calculate the X-axis coordinate difference, Y-axis coordinate difference, and Z-axis coordinate difference between the three-dimensional spatial coordinates of each wrist key point and the reference point directly below the outlet, and generate the corresponding Euclidean distance; The remaining wrist keypoints are sorted in ascending order of Euclidean distance. The wrist keypoint with the smallest Euclidean distance is written into the target wrist keypoint marker. Only the target wrist keypoint is input into the historical trajectory cache queue, direction accumulation value calculation, and occlusion status preservation determination. When the target wrist keypoints switch in several consecutive frames and the difference in Euclidean distance between the new target wrist keypoint and the previous target wrist keypoint does not exceed the preset switching tolerance, the previous target wrist keypoint mark is maintained. Read the duration of the finite state machine in the paused memory state. When the paused memory state duration exceeds the preset safe duration, trigger the timeout circuit breaker, clear the integral cache, historical trajectory cache queue, completion flag and target wrist key point flag, and reset the finite state machine to the closed state.

[0013] According to an embodiment of the present invention, a faucet start / stop control system based on motion-sensing interaction includes the following modules: The 3D vision access module is used to access a 3D video stream containing the user's hand and water flow and perform polarization de-reflection processing. The wrist key point parsing module, connected to the 3D vision access module, is used to run an anti-occlusion customized YOLOv8-pose model and output the 3D spatial coordinates and occlusion confidence of all wrist key points. The target locking module, connected to the wrist key point parsing module, is used to generate target wrist key point markers based on the Euclidean distance from each wrist key point to the reference point directly below the outlet. The trajectory continuity maintenance module, connected to the target locking module, is used to generate a continuous hand spatial trajectory sequence based on the occlusion confidence of the target's wrist key points and the three-dimensional spatial coordinates of the wrist key points. The directional intent parsing module, connected to the trajectory continuity maintenance module, is used to generate directional cumulative values ​​based on a continuous sequence of hand spatial trajectories. The state maintenance determination module is connected to the wrist key point parsing module, trajectory continuity maintenance module and finite state control module. It is used to generate occlusion state maintenance determination results based on occlusion confidence, continuous hand spatial trajectory sequence and the previous control state of the finite state machine. The finite state control module, connected to the direction intent parsing module, the state holding determination module, and the target locking module, is used to generate control intent based on the direction accumulation value, the occlusion state holding determination result, and the target wrist constraint signal. It switches between the closed state, the water outlet state, and the pause memory state to control the start and stop of the solenoid valve, and triggers the timeout fuse after the pause memory state expires.

[0014] The beneficial effects of this invention are: This invention addresses the instability of images captured by conventional visual acquisition under conditions of water flow reflection and depth deficiency by using polarization de-reflection processing and depth confidence constraints to identify, replace, and fill in metallic reflective pixels, water stain reflective pixels, and invalid depth locations in the near-field region of a faucet. This improves the quality of image frames containing depth information. Furthermore, by using an anti-occlusion customized YOLOv8-pose model to remove full-body skeletal keypoint outputs while retaining the wrist keypoint detection head, and combining depth coordinate regression branches, occlusion confidence output branches, and foam occlusion region weights, this invention solves the problems of easy loss of wrist keypoint detection and insufficient 2D gesture information under foam occlusion, improving the continuity of the 3D spatial coordinates of wrist keypoints. This system improves reliability and accuracy. Through adaptive trajectory completion, dead-zone threshold filtering to eliminate rubbing and shaking displacement, and signed accumulation of Z-axis relative displacement, it addresses the difficulty in distinguishing between foam occlusion, rubbing and shaking, and actual approaching or moving away movements, enabling the directional accumulation value to more accurately reflect the user's start / stop intentions. By combining semi-Markov duration constraints with occlusion observation parameters, trajectory motion parameters, and the previous control state, it differentiates between normal water output, foaming and suspension, and actual leaving states. Furthermore, by using pause memory to shield against erroneous shut-off intentions during foaming, it solves the problems of frequent accidental shut-offs and state jumps in traditional sensor faucets. Target wrist key point locking and timeout circuit breaking reduce the risk of accidental triggering under multi-user interference. Its significance lies in elevating faucet start / stop control from single-frame sensor switching to continuous state understanding control oriented towards the actual handwashing process, improving the stability, safety, and naturalness of contactless water use. Attached Figure Description

[0015] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a faucet start / stop control method based on motion-sensing interaction proposed in this invention; Figure 2 This is a framework diagram of the anti-occlusion customized YOLOv8-pose model in the faucet start / stop control method based on motion-sensing interaction proposed in this invention. Figure 3 This is a schematic diagram of the overall framework of a faucet start / stop control system based on motion-sensing interaction proposed in this invention. Detailed Implementation

[0016] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0017] refer to Figures 1-2 A method for controlling the start and stop of a faucet based on motion-sensing interaction includes the following steps: Step 1: Acquire a 3D video stream containing the user's hand and water flow, and generate image frames with depth information through polarization de-reflection processing; Step 2: Input the image frame with depth information into the anti-occlusion customized YOLOv8-pose model. The anti-occlusion customized YOLOv8-pose model removes the full-body skeleton keypoint output, retains the wrist keypoint detection head, configures the depth coordinate regression branch and the occlusion confidence output branch, and adds the foam occlusion region weight to the training loss to generate the wrist keypoint 3D spatial coordinates and occlusion confidence. Step 3: Perform adaptive trajectory completion based on occlusion confidence. When the occlusion confidence is lower than the preset safety threshold, generate virtual coordinates and a continuous hand spatial trajectory sequence based on the historical valid 3D spatial coordinate sequence. Step 4: Extract the Z-axis coordinate sequence from the continuous hand spatial trajectory sequence, filter out the rubbing and shaking displacement according to the dead zone threshold, and perform signed accumulation on the filtered Z-axis relative displacement to generate directional accumulation value; Step 5: Determine the occlusion state based on the semi-Markov duration constraint, and generate the normal water discharge state, foam suspension state, or real departure state according to the occlusion observation parameters, trajectory motion parameters and the previous control state. Step 6: Generate control intent based on the direction accumulation value and the occlusion state determination result, and input the control intent into a finite state machine containing the closed state, water outlet state, and pause memory state to control the start and stop of the solenoid valve; Step 7: When multiple wrist key points are detected, the target wrist key point is locked based on the Euclidean distance from each wrist key point to the reference point directly below the outlet. When the duration of the pause memory state exceeds the preset safe duration, a timeout circuit is triggered, the integral cache is cleared, and the finite state machine is reset to the off state.

[0018] In this embodiment, step one specifically includes: Acquire the original color image, depth channel image, and depth confidence image of the near-field area of ​​the faucet. The near-field area of ​​the faucet is the image acquisition area corresponding to the preset handwashing space below the water outlet. The original color image is the image that records the red channel value, green channel value, and blue channel value. The depth channel image is the image that records the distance value from the pixel to the acquisition end. The depth confidence image is the image that records the confidence level of the depth value of each pixel in the depth channel image. The original color image, depth channel image, and depth confidence image acquired at the same time are bound according to the frame number to generate a 3D video stream; A near-field effective area is established based on the reference point directly below the water outlet. The reference point directly below the water outlet is the spatial reference point after the center of the water outlet is projected vertically onto the handwashing space. The near-field effective area is the pixel area whose depth value is within the preset handwashing distance and is distributed around the reference point directly below the water outlet. Orthogonal polarization sampling is performed on the original color image within the effective near-field region to obtain the first polarization intensity image and the second polarization intensity image. Orthogonal polarization sampling is a sampling process that records the brightness value of the same image frame under two mutually perpendicular polarization directions. Compare the brightness difference at the same pixel location in the first polarization intensity image and the second polarization intensity image, and write invalid depth markers at pixel locations below the preset depth confidence threshold in the depth confidence image to mark metallic reflective pixels and water stain reflective pixels, thereby generating a reflective pixel mask; The brightness of reflective pixels in the original color image is replaced based on the reflective pixel mask. The brightness replacement is to read the median brightness of the pixels in the neighborhood of the reflective pixel that are not marked by the reflective pixel mask, and write the median brightness of the pixels to the corresponding reflective pixel position. Depth filling is performed on the positions corresponding to invalid depth markers. Depth filling involves reading the valid depth values ​​at the boundaries of the same connected region and writing them into the corresponding invalid depth positions according to the median of the neighborhood. The original color image after brightness replacement, the depth channel image after depth filling, and the corresponding frame number are re-bound to generate an image frame with depth information. In one specific implementation, based on the installation height from the faucet outlet to the bottom of the washbasin in the public restroom of the office building and the statistical range of the distance between the adult's hand and the water outlet when washing hands normally, the preset handwashing distance range is 120mm to 420mm, and the preset depth confidence threshold is 0.55. In this embodiment, the anti-occlusion customized YOLOv8-pose model includes a dual-modal input layer, a feature extraction and fusion layer, and a customized decoupled detection head: The bimodal input layer receives image frames with depth information and splits them into an RGB image tensor and a TOF depth map tensor. The RGB image tensor is a three-dimensional array converted from the original color image, with each pixel position recording the red, green, and blue channel values. The TOF depth map tensor is a depth array obtained by time-of-flight ranging, with each pixel position recording the distance from the corresponding object surface to the acquisition end. Based on the pixel mapping relationship between the RGB image tensor and the TOF depth map tensor, depth pixels in the TOF depth map tensor are written to the corresponding pixel positions in the RGB image tensor. Pixels without matching depth values ​​are marked as invalid depth. The RGB pixel values ​​and depth values ​​are mapped to a unified numerical range through normalization processing to generate bimodal input data. The feature extraction and fusion layer inputs the RGB image tensor into the CSPDarknet backbone network to extract multi-scale visual features, and then inputs these multi-scale visual features into the feature fusion neck network. Through convolution, batch normalization, and SiLU activation, shallow spatial detail features and deep semantic features are fused to generate hand color texture response, foam brightness response, and wrist contour response. The TOF depth map tensor is input into the depth parallel branch, and the wrist depth boundary response, water flow depth fracture response, and invalid depth region response are extracted through convolution mapping. Channel dimensions are then stitched together, and then dimensionality reduction is performed by one-to-one convolution to generate a fused feature map. A customized decoupled detection head replaces the multi-keypoint output structure of the human body in the YOLOv8-pose model with a single wrist center point output structure. The full-body skeletal keypoint outputs corresponding to the head, shoulders, elbows, hips, knees, and ankles are deleted, and only the wrist keypoint detection head is retained. The fused feature map is input into the wrist 2D localization branch, and a wrist center response map is generated through heatmap regression. The pixel region with the largest wrist center response value is selected as the wrist candidate region. The position of the wrist center point in the wrist candidate region is corrected by offset prediction, and the 2D pixel coordinates of the wrist keypoint are generated. A customized decoupled detection head inputs the fused feature map into the depth coordinate regression branch. It extracts depth features within the neighborhood of the wrist keypoint's two-dimensional pixel coordinates through convolutional and fully connected mappings. Invalid depth values ​​are removed, and the remaining TOF depth values ​​are sorted by magnitude and the median depth value is selected. Depth values ​​crossing water flow depth fracture regions are also removed based on the wrist depth boundary response, generating the physical depth value of the wrist keypoint. The two-dimensional pixel coordinates and physical depth value of the wrist keypoint are then combined to form the three-dimensional spatial coordinates of the wrist keypoint. During the training phase, a binary mask for the foam region is generated based on the annotations of foam-covered pixels in the training samples. The binary mask for the foam region is a label map in which the foam-covered pixels take the first label value and the non-foam pixels take the second label value. The overlap result between the wrist prediction position and the binary mask for the foam region is read. The loss of the key point coordinates corresponding to the foam-covered pixels is multiplied by a preset foam weighting coefficient, and the edge loss corresponding to the wrist contour boundary position is multiplied by a preset edge enhancement coefficient to form a training loss with added weights for the foam occlusion region. The occlusion confidence output branch reads the mean gray-level gradient, the proportion of invalid depth pixels, and the proportion of foam-covered pixels in the wrist candidate region. After passing the mean gray-level gradient, the proportion of invalid depth pixels, and the proportion of foam-covered pixels through a fully connected mapping, the data is input into the Sigmoid activation function to generate the occlusion confidence. In one specific implementation, the preset foam weighting coefficient is set to 0.45, and the preset edge enhancement coefficient is set to 1.8. This is based on the calibration results of the training set containing foam-occluded samples and unoccluded samples. When foam-covered pixels still participate in training with ordinary keypoint coordinate loss, wrist keypoints are prone to shifting to the foam bright spot area. Therefore, the keypoint coordinate loss corresponding to foam-covered pixels is weighted down, while the edge loss at the wrist contour boundary position is enhanced to improve the stability of the wrist contour response under foam occlusion. The preset safety threshold is set to 0.62, based on the ROC curve boundary point of the occlusion confidence output branch on foam-occluded, water flow-occluded, and unoccluded samples, so that the occlusion misclassification rate and false negative rate remain at a low level. In a specific implementation, the anti-occlusion customized YOLOv8-pose model retains the basic framework of the original YOLOv8-pose model for target detection, feature extraction and key point localization. It extracts multi-scale visual features of the image through the backbone network, and uses the feature fusion neck network to enhance the expression of the hand region at different scales. Finally, it outputs the key point position through the detection head. Compared to the original YOLOv8-pose model, this implementation removes the full-body skeletal keypoint outputs corresponding to the head, shoulders, elbows, hips, knees, and ankles, retains only the wrist keypoint detection head, and adds TOF depth map tensor input, depth coordinate regression branch, and occlusion confidence output branch. At the same time, foam occlusion region weights are added to the training loss, so that the model is transformed from ordinary human pose estimation into a three-dimensional analytical model of wrist keypoints for near-field handwashing scenarios with a faucet. Through the above improvements, the model can jointly identify the wrist position using RGB image tensors and TOF depth map tensors. Even under the interference of foam occlusion, water flow depth breaks, and metal reflection, it can still output the three-dimensional spatial coordinates and occlusion confidence of wrist key points. This solves the problems of false detection, missed detection, and lack of depth judgment of key points in the original YOLOv8-pose model in close-range handwashing scenarios, and provides stable input for subsequent adaptive trajectory completion, direction accumulation value calculation, and occlusion state maintenance judgment.

[0019] In this embodiment, step three specifically includes: Read the three-dimensional spatial coordinates and occlusion confidence of wrist key points, and mark the three-dimensional spatial coordinates of wrist key points with occlusion confidence not lower than the preset safety threshold and not in an invalid depth position as valid three-dimensional spatial coordinates; Valid 3D spatial coordinates, occlusion confidence and frame number are stored in chronological order to form a historical trajectory cache queue. The valid 3D spatial coordinates in the historical trajectory cache queue are skipped and removed. Coordinates whose spatial displacement in adjacent frames exceeds the preset single frame displacement limit are marked as skipped coordinates and deleted from the historical trajectory cache queue. When the occlusion confidence of the current frame is lower than the preset safety threshold, the most recent consecutive valid three-dimensional spatial coordinates in the historical trajectory cache queue are read, and the X-axis displacement, Y-axis displacement and Z-axis displacement between adjacent valid three-dimensional spatial coordinates are calculated respectively. The median displacement is selected from the most recent several sets of displacements as the complete displacement. The last valid 3D spatial coordinates of the last frame in the historical trajectory cache queue are superimposed with the X-axis displacement, Y-axis displacement and Z-axis displacement in the completion displacement to generate the predicted spatial position of the current frame and use it as virtual coordinates. The virtual coordinates are the alternative wrist spatial position data generated when the wrist key point is occluded by foam or the depth is missing for a short time, and the completion mark and the corresponding frame number are written into it. When the number of consecutive frames generating virtual coordinates does not exceed the preset maximum number of completion frames, the virtual coordinates are written to the historical trajectory cache queue; when the number of consecutive frames generating virtual coordinates exceeds the preset maximum number of completion frames, completion is stopped and a trajectory interruption flag is output. When a subsequent frame re-acquires valid 3D spatial coordinates, the spatial position difference between the re-acquired valid 3D spatial coordinates and the previous virtual coordinates is compared. If the spatial position difference does not exceed the preset reconnection distance threshold, the re-acquired valid 3D spatial coordinates are added to the historical trajectory cache queue. The effective three-dimensional spatial coordinates and virtual coordinates are arranged according to the frame number, and the small jumps between adjacent coordinates in a single frame are smoothed by median smoothing to generate a continuous hand spatial trajectory sequence. In one specific implementation, the preset single-frame displacement upper limit is set to 45mm, based on the statistical value of wrist displacement in a single frame of a user's normal handwashing action under 30fps acquisition conditions. Coordinates exceeding this value are mostly caused by false detections due to reflection or drift due to foam occlusion. The preset maximum number of completed frames is set to 12 frames, corresponding to a short occlusion time of approximately 0.4s, based on the duration distribution of short-term loss of key wrist points during foaming and water flow occlusion tests. The preset reconnection distance threshold is set to 65mm, based on the statistical results of the reconnection error between virtual coordinates and the re-obtained valid three-dimensional spatial coordinates. The trajectory can remain continuous when it is below this threshold.

[0020] In this embodiment, step four specifically includes: Read the three-dimensional spatial coordinates of wrist key points in a continuous hand spatial trajectory sequence, and extract the Z-axis coordinates of each frame according to the frame number to generate a Z-axis coordinate sequence. The Z-axis coordinate is the physical depth value of the wrist key point along the depth direction in the camera coordinate system. Set up an integral sliding window, which is a calculation window that extracts several consecutive frames of Z-axis coordinates in chronological order, and reads the Z-axis coordinates of two adjacent frames one by one within the integral sliding window. The difference between the Z-axis coordinates of the next frame and the Z-axis coordinates of the previous frame is calculated to generate the relative Z-axis displacement. The relative Z-axis displacement is the amount of displacement change of the wrist key point along the depth direction between adjacent frames. The dead zone threshold is generated based on the static rubbing calibration samples. The static rubbing calibration samples are the Z-axis relative displacement samples collected when the user holds the handwashing position below the water outlet and performs the rubbing action. The dead zone threshold is the upper limit value of the jitter selected after sorting the absolute values ​​of the Z-axis relative displacement samples. The Z-axis relative displacements with absolute values ​​less than the dead zone threshold within the integral sliding window are marked as rubbing and shaking displacements, and the rubbing and shaking displacements are written into the zero displacement marker. The Z-axis relative displacements with absolute values ​​not less than the dead zone threshold are retained to generate the filtered Z-axis relative displacements. The filtered Z-axis relative displacement is accumulated with a sign according to the frame number. The Z-axis relative displacement away from the outlet is written into the positive accumulation item, and the Z-axis relative displacement closer to the outlet is written into the negative accumulation item. The positive and negative accumulation items are superimposed according to their signs to generate the directional accumulation value. Write the filtered Z-axis relative displacement, zero displacement marker, positive cumulative term, negative cumulative term, and direction cumulative value within the integral sliding window into the integral buffer; In one specific implementation, the integral sliding window takes 18 frames, corresponding to a continuous action judgment time of approximately 0.6 seconds under 30fps acquisition; the dead zone threshold is set to 7mm, based on the 95th percentile of the absolute value of the relative displacement of the Z-axis in the static rubbing calibration sample; the preset positive pause threshold is set to 55mm, and the preset negative recovery threshold is set to -45mm, based on the cumulative distribution of the relative displacement of the Z-axis when the user moves away from the water outlet to foam and then moves back to the water outlet to rinse, so that ordinary rubbing and shaking will not trigger the pause control intention or the water output control intention.

[0021] In this embodiment, step five specifically includes: Read the occlusion confidence and effective 3D spatial coordinates output results of the current frame. When the occlusion confidence is lower than the preset safety threshold or the current frame does not output effective 3D spatial coordinates, write an occlusion observation flag and set the continuous occlusion duration frame number to the previous frame's continuous occlusion duration frame number plus one. When the occlusion confidence is not lower than the preset safety threshold and the current frame outputs effective 3D spatial coordinates, write an unoccluded observation flag and clear the continuous occlusion duration frame number to zero. Using the most recent consecutive frames as the current decision window, read the X-axis coordinates, Y-axis coordinates, and Z-axis coordinates in the continuous hand spatial trajectory sequence. Add the square of the difference between the X-axis coordinates and the square of the difference between the Y-axis coordinates of adjacent frames and take the square root to generate the inter-frame displacement in the plane. Then, average all the inter-frame displacements in the current decision window to generate the mean value of the inter-frame displacement. Calculate the Z-axis coordinate difference between adjacent frames within the current decision window, generate the relative Z-axis displacement, count the number of frames with the same Z-axis relative displacement direction, and divide the number of frames with the same direction by the total number of valid adjacent frames to generate the Z-axis direction continuity. Count the number of virtual coordinates within the current judgment window, divide the number of virtual coordinates by the total number of coordinates within the current judgment window, and generate the trajectory completion percentage; Normal water outflow state, foaming and suspension state, and actual departure state are selected as candidate occlusion states, and a preset minimum duration frame count, a preset maximum duration frame count, a duration frame count center value, and a duration frame count fluctuation value are configured for each candidate occlusion state. The state transition admission conditions for candidate occlusion states are set according to the previous control state of the finite state machine in the previous frame. When the previous control state is the off state, it is allowed to enter the normal water discharge state or the real departure state. When the previous control state is the water discharge state, it is allowed to enter the normal water discharge state, the foaming suspension state, or the real departure state. When the previous control state is the pause memory state, it is allowed to enter the foaming suspension state, the normal water discharge state, or the real departure state. For each candidate occlusion state, a semi-Markov duration constraint calculation is performed. When the number of consecutive occlusion frames is lower than the preset minimum duration frame or higher than the preset maximum duration frame of the corresponding candidate occlusion state, the duration weight of the corresponding candidate occlusion state is reset to zero. When the continuous occlusion duration is between the corresponding preset minimum duration and the preset maximum duration, the absolute value of the difference between the continuous occlusion duration and the center value of the corresponding duration is taken, and the duration weight is generated based on the ratio of the absolute value to the duration fluctuation value. The smaller the ratio, the larger the duration weight, and the larger the ratio, the smaller the duration weight. State matching weights are configured for occlusion observation markers, unocclusion observation markers, previous control state, mean plane displacement, Z-axis continuity, and trajectory completion ratio. The state matching weights that meet the corresponding candidate occlusion state conditions are accumulated to generate a matching score for each candidate occlusion state. Unobstructed observation markers and effective three-dimensional spatial coordinates correspond to the normal water discharge state. Obstructed observation markers, previous control state being water discharge state or paused memory state, average planar displacement within a preset stable range, Z-axis continuity reaching a preset direction continuity threshold and trajectory completion ratio not exceeding a preset completion ratio correspond to the foaming suspension state. Continuous obstruction lasting for a number of frames reaching a preset departure judgment frame number, the presence of a trajectory interruption marker in the current judgment window, or the average planar displacement exceeding a preset departure displacement threshold correspond to the actual departure state. The matching score of each candidate occlusion state that passes the state transition admission condition is multiplied by the corresponding duration weight to generate a state score after semi-Markov constraint. The candidate occlusion state with the highest state score is selected to generate a normal water outflow state, a foaming and floating state, or a real departure state. In one specific implementation, the preset stability range is 0mm to 12mm, the preset directional continuity threshold is 0.65, the preset completion ratio is 0.5, the preset departure judgment frame number is 15 frames, and the preset departure displacement threshold is 80mm. The above values ​​are based on the statistical results of samples where the wrist plane displacement is small, the Z-axis direction change is continuous, and the virtual coordinate ratio is not too high during the foaming suspension state. When the continuous occlusion duration reaches 15 frames or the average plane displacement exceeds 80mm, the probability of the user's hand leaving the water outlet area increases significantly, thus corresponding to the real departure state.

[0022] In this embodiment, step six specifically includes: Read the direction accumulation value and the occlusion status retention determination result, compare the direction accumulation value with the preset positive pause threshold and the preset negative recovery threshold, and write the normal water outlet state, foaming suspension state or real departure state in the occlusion status retention determination result into the current status determination bit. When the cumulative directional value is not less than the preset positive pause threshold and the current state determination bit is in normal water discharge state, a pause control intention is generated; when the cumulative directional value is not greater than the preset negative recovery threshold and the current state determination bit is in normal water discharge state or foaming suspension state, a recovery water discharge control intention is generated; when the current state determination bit is in true departure state, a shut-off control intention is generated; when the current state determination bit is in foaming suspension state and the cumulative directional value has not reached the preset negative recovery threshold, a pause hold control intention is generated. When the finite state machine is in the paused memory state and the occlusion state holds the decision result as the bubble-floating state, the control intention to close is masked and a pause hold control intention is generated. The closed state, the water outlet state, and the pause memory state are used as state nodes of the finite state machine. The finite state machine is a state control structure that switches between different state nodes according to the control intention and outputs the solenoid valve drive signal. The closed state is the state node in which the solenoid valve is in the closed drive state, the water outlet state is the state node in which the solenoid valve is in the open drive state, and the pause memory state is the state node in which the solenoid valve is in the closed drive state and retains the previous water outlet recovery conditions. When the finite state machine is in the water outlet state and receives the pause control intention, it switches the finite state machine to the pause memory state and outputs a shut-off drive signal to the solenoid valve; when the finite state machine is in the pause memory state and receives the resume water outlet control intention, it switches the finite state machine to the water outlet state and outputs an open drive signal to the solenoid valve. When the finite state machine is in the paused memory state and receives the pause hold control intention, the finite state machine remains in the paused memory state and continues to output the shut-off drive signal to the solenoid valve; when the finite state machine is in the shut-off state or the water outlet state and receives the shut-off control intention, the finite state machine is switched to the shut-off state and the shut-off drive signal is output to the solenoid valve. After the finite state machine completes the state transition, the accumulated direction values ​​that have participated in the generation of control intent in the current decision window are cleared, the previous control state corresponding to the pause memory state is retained, and the state transition result is written into the occlusion state retention decision input of the next frame.

[0023] In this embodiment, step seven specifically includes: Read the 3D spatial coordinates of all wrist key points output by the anti-occlusion customized YOLOv8-pose model in the current frame, and remove wrist key points whose occlusion confidence is lower than the preset safety threshold; Read the spatial coordinates of the reference point directly below the outlet, calculate the X-axis coordinate difference, Y-axis coordinate difference, and Z-axis coordinate difference between the three-dimensional spatial coordinates of each wrist key point and the reference point directly below the outlet, and take the square root of the sum of the squares of the three coordinate differences to generate the corresponding Euclidean distance; The remaining wrist keypoints are sorted in ascending order of Euclidean distance. The wrist keypoint with the smallest Euclidean distance is written into the target wrist keypoint marker. Only the target wrist keypoint is input into the historical trajectory cache queue, direction accumulation value calculation, and occlusion status preservation determination. When the target wrist keypoints switch in several consecutive frames, compare the Euclidean distance difference between the new target wrist keypoint and the previous target wrist keypoint. If the Euclidean distance difference does not exceed the preset switching tolerance, the previous target wrist keypoint mark is retained. Read the duration of the finite state machine in the paused memory state. When the paused memory state duration exceeds the preset safe duration, trigger the timeout circuit breaker, clear the integral cache, historical trajectory cache queue, completion flag and target wrist key point flag, and reset the finite state machine to the closed state. In one specific implementation, the preset switching tolerance is 35mm, based on the statistical results of the difference in Euclidean distance between the target wrist key point and the non-target wrist key point when two people are close to the faucet, in order to avoid frequent switching of the target wrist key point when the two hands are close to each other; the preset safe duration is 45s, based on the common duration of a single foaming pause in public restrooms and water-saving safety requirements. When the pause memory state duration exceeds 45s, a timeout circuit is triggered, clearing the integral cache, historical trajectory cache queue, completion flag and target wrist key point flag, and resetting the finite state machine to the closed state.

[0024] refer to Figure 3 A faucet start / stop control system based on motion-sensing interaction includes the following modules: The 3D vision access module is used to access a 3D video stream containing the user's hand and water flow and perform polarization de-reflection processing. The wrist key point parsing module, connected to the 3D vision access module, is used to run an anti-occlusion customized YOLOv8-pose model and output the 3D spatial coordinates and occlusion confidence of all wrist key points. The target locking module, connected to the wrist key point parsing module, is used to generate target wrist key point markers based on the Euclidean distance from each wrist key point to the reference point directly below the outlet. The trajectory continuity maintenance module, connected to the target locking module, is used to generate a continuous hand spatial trajectory sequence based on the occlusion confidence of the target's wrist key points and the three-dimensional spatial coordinates of the wrist key points. The directional intent parsing module, connected to the trajectory continuity maintenance module, is used to generate directional cumulative values ​​based on a continuous sequence of hand spatial trajectories. The state maintenance determination module is connected to the wrist key point parsing module, trajectory continuity maintenance module and finite state control module. It is used to generate occlusion state maintenance determination results based on occlusion confidence, continuous hand spatial trajectory sequence and the previous control state of the finite state machine. The finite state control module, connected to the direction intent parsing module, the state holding determination module, and the target locking module, is used to generate control intent based on the direction accumulation value, the occlusion state holding determination result, and the target wrist constraint signal. It switches between the closed state, the water outlet state, and the pause memory state to control the start and stop of the solenoid valve, and triggers the timeout fuse after the pause memory state expires.

[0025] Example 1: To verify the feasibility of this invention in practice, it was applied to a scenario of controlling the start and stop of sensor-operated faucets in a public restroom in an office building. The restroom had six faucets, each equipped with an RGB camera, a TOF depth acquisition unit, and a polarization sampling component. The acquisition area covered a near-field effective area approximately 120mm to 420mm below the water outlet, with an image acquisition frame rate of 30fps. In the test, 30 users were selected, each performing actions including normal rinsing, pausing while applying foam, rinsing again, actually leaving, and two people approaching each other for interference, resulting in a total of 900 complete handwashing processes. The preset depth confidence threshold is 0.55, the preset safety threshold is 0.62, the preset single-frame displacement limit is 45mm, the preset maximum number of completed frames is 12, the preset reconnection distance threshold is 65mm, and the integral sliding window is 18 frames. The dead zone threshold is obtained from the 95th percentile of the absolute value of the relative displacement of the Z-axis in the static rubbing calibration sample, and is 7mm in this embodiment. The preset positive pause threshold is 55mm, the preset negative recovery threshold is -45mm, the preset switching tolerance is 35mm, and the preset safety duration is 45s.

[0026] In the comparison methods, the infrared sensing method uses conventional infrared distance detection to control the start and stop of the solenoid valve; the ordinary vision method uses the original YOLOv8-pose to identify key points of the hand and combines it with a simple finite state machine to control the water output; the unmodified model method retains the basic structure of YOLOv8-pose, but does not add TOF depth map tensor, foam occlusion region weight, occlusion confidence output branch and semi-Markov duration constraint; the present invention uses polarization de-reflection processing, anti-occlusion customized YOLOv8-pose model, adaptive trajectory completion, direction accumulation value, occlusion state retention judgment and pause memory state linkage control.

[0027] Table 1 Comparison of the Implementation Effects of Faucet Start-Stop Control

[0028] As shown in Table 1, the infrared sensing method has the lowest response delay, but it only determines the presence of objects in the near field and cannot distinguish between normal water flow, foaming and suspension, and actual departure. Therefore, it is easy to mistakenly assume that the user has left during foaming, resulting in a false shut-off rate of 18.6%. The ordinary visual method can identify the hand area, but the detection rate of key points on the wrist decreases when there is water reflection, foam coverage, or hand crossing. The trajectory continuity rate is only 72.5% during foam obstruction, leading to a large number of frequent start-stop cycles.

[0029] The unimproved model method shows improvement over ordinary vision methods, indicating that the visual keypoint framework of YOLOv8-pose has a certain adaptability to the faucet scene; however, due to the lack of depth coordinate regression branch, occlusion confidence output branch, and foam occlusion region weights, keypoint drift still occurs in foam-covered and water flow depth break areas, with an average 3D spatial coordinate error of 23.5mm, making it difficult to stably support the judgment of directional cumulative values. This invention improves the wrist keypoint detection rate to 97.8% and reduces the average 3D spatial coordinate error to 11.6mm by fusing RGB image tensors and TOF depth map tensors; through adaptive trajectory completion, the trajectory continuity rate reaches 94.7% during foam occlusion; and by filtering out rubbing and shaking displacement and Z-axis relative displacement with sign accumulation through dead zone thresholding, the situation where ordinary rubbing actions are misjudged as start and stop actions is reduced.

[0030] In practical use, when a user moves their hand away from the water outlet and lathers their hands, this invention generates a pause control intention based on the direction accumulation value, causing the finite state machine to switch from the water outlet state to the pause memory state. The solenoid valve closes but retains the recovery conditions. When foam obstruction causes a brief loss of the wrist's key point, the obstruction confidence triggers adaptive trajectory completion. The semi-Markov duration constraint, combined with the trajectory completion ratio, Z-axis direction continuity, and the previous control state, determines the scene as a foam-laden suspension state, rather than directly determining it as a real departure state. When the user moves their hand back to the water outlet, the direction accumulation value reaches a preset negative recovery threshold, the finite state machine switches to the water outlet state, and water flow resumes. Therefore, this invention does not simply rely on a single frame of image or a single distance value, but rather comprehensively judges the continuous actions and the duration of obstruction during the handwashing process.

[0031] As can be seen from the embodiments, the beneficial effects of the present invention are as follows: improving the quality of near-field images and depth data through polarization de-reflection processing; improving the stability of the three-dimensional spatial coordinates of wrist key points under foam, water flow and reflective conditions through anti-occlusion customized YOLOv8-pose model; reducing the false shut-off rate through adaptive trajectory completion and occlusion state retention determination; reducing false triggering due to interference from multiple people through target wrist key point locking; and balancing ease of use and safety through pause memory state and timeout circuit breaker, thereby making the faucet start-stop control more consistent with the real handwashing process.

[0032] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for controlling the start and stop of a faucet based on motion-sensing interaction, characterized in that, Includes the following steps: Step 1: Acquire a 3D video stream containing the user's hand and water flow, and generate image frames with depth information through polarization de-reflection processing; Step 2: Input the image frame with depth information into the anti-occlusion customized YOLOv8-pose model. The anti-occlusion customized YOLOv8-pose model deletes the output of the whole body skeleton key points, retains the wrist key point detection head, configures the depth coordinate regression branch and the occlusion confidence output branch, and adds the foam occlusion region weight to the training loss to generate the three-dimensional spatial coordinates and occlusion confidence of the wrist key points. Step 3: Perform adaptive trajectory completion based on occlusion confidence. When the occlusion confidence is lower than the preset safety threshold, generate virtual coordinates and a continuous hand spatial trajectory sequence based on the historical valid 3D spatial coordinate sequence. Step 4: Extract the Z-axis coordinate sequence from the continuous hand spatial trajectory sequence, filter out the rubbing and shaking displacement according to the dead zone threshold, and perform signed accumulation on the filtered Z-axis relative displacement to generate directional accumulation value; Step 5: Determine the occlusion state based on the semi-Markov duration constraint, and generate the normal water discharge state, foam suspension state, or real departure state according to the occlusion observation parameters, trajectory motion parameters, and the previous control state. Step 6: Generate control intent based on the direction accumulation value and the occlusion state determination result, and input the control intent into a finite state machine containing the closed state, water outlet state, and pause memory state to control the start and stop of the solenoid valve; Step 7: When multiple wrist key points are detected, the target wrist key point is locked based on the Euclidean distance from each wrist key point to the reference point directly below the outlet. When the duration of the pause memory state exceeds the preset safe duration, a timeout circuit is triggered, the integral cache is cleared, and the finite state machine is reset to the off state.

2. The faucet start / stop control method based on motion-sensing interaction according to claim 1, characterized in that, Step one specifically includes: Acquire the original color image, depth channel image, and depth confidence image of the near-field area of ​​the faucet, and generate a 3D video stream by binding them according to the frame sequence number; A near-field effective region is established based on the reference point directly below the outlet. Orthogonal polarization sampling is performed on the original color image within the near-field effective region to obtain the first polarization intensity image and the second polarization intensity image. Compare the brightness difference at the same pixel location in the first polarization intensity image and the second polarization intensity image, and combine the pixel locations in the depth confidence image that are below a preset depth confidence threshold to generate a reflective pixel mask and an invalid depth marker; The reflective pixels in the original color image are replaced with the median brightness of the neighborhood based on the reflective pixel mask, and the median depth of the neighborhood is filled in the depth channel image based on the invalid depth marker. The original color image after brightness replacement, the depth channel image after depth filling, and the corresponding frame number are re-bound to generate an image frame with depth information.

3. The faucet start / stop control method based on motion-sensing interaction according to claim 1, characterized in that, The anti-occlusion customized YOLOv8-pose model includes a dual-modal input layer, a feature extraction and fusion layer, and a customized decoupled detection head: The bimodal input layer splits the image frame with depth information into an RGB image tensor and a depth map tensor, writes depth pixels and invalid depth markers according to the pixel mapping relationship, and normalizes them to generate bimodal input data; The feature extraction and fusion layer inputs the RGB image tensor into the CSPDarknet backbone network to extract multi-scale visual features, and then inputs the multi-scale visual features into the feature fusion neck network to generate hand color texture response, foam brightness response and wrist contour response; The depth map tensor is input into the depth parallel branch to generate wrist depth boundary response, water flow depth fracture response and invalid depth region response. The responses at the same pixel position are then concatenated by channel dimension and dimensionality reduced by one-to-one convolution to generate a fused feature map. A custom decoupling detection head is used to remove the output of the full-body skeletal key points in the YOLOv8-pose model, while retaining the wrist key point detection head. The wrist key point 2D pixel coordinates are generated through the wrist 2D localization branch. The customized decoupled detection head extracts the depth features in the neighborhood of the two-dimensional pixel coordinates of the wrist key point through the depth coordinate regression branch, removes invalid depth markers and depth values ​​that cross the water flow depth fracture area, and generates the physical depth value of the wrist key point. The two-dimensional pixel coordinates of the wrist key points and the physical depth values ​​of the wrist key points are combined to form the three-dimensional spatial coordinates of the wrist key points. During the training phase, a binary mask of the foam region is generated based on the annotation of the foam-covered pixels. The loss of the key point coordinates corresponding to the foam-covered pixels is multiplied by a preset foam weighting coefficient, and the edge loss corresponding to the wrist contour boundary position is multiplied by a preset edge enhancement coefficient to form a training loss with added weights for the foam-occluded region. The occlusion confidence output branch reads the mean gray-level gradient, the proportion of invalid depth pixels, and the proportion of foam-covered pixels in the candidate area of ​​the wrist, and generates the occlusion confidence through fully connected mapping and Sigmoid activation function.

4. The faucet start / stop control method based on motion-sensing interaction according to claim 1, characterized in that, Step three specifically includes: Read the three-dimensional spatial coordinates and occlusion confidence of wrist key points, and mark the three-dimensional spatial coordinates of wrist key points with occlusion confidence not lower than the preset safety threshold and not in an invalid depth position as valid three-dimensional spatial coordinates; The effective three-dimensional spatial coordinates, occlusion confidence and frame number are stored according to the frame sequence number to form a historical trajectory cache queue, and the jump coordinates of adjacent frames whose spatial displacement exceeds the preset single frame displacement limit are deleted. When the occlusion confidence of the current frame is lower than the preset safety threshold, the most recent consecutive valid three-dimensional spatial coordinates in the historical trajectory cache queue are read, the X-axis displacement, Y-axis displacement and Z-axis displacement are calculated, and the median displacement is selected as the complete displacement. The last valid 3D spatial coordinates in the historical trajectory cache queue are superimposed to complete the displacement, generating the predicted spatial position of the current frame, and written as virtual coordinates to the completion mark and the corresponding frame number; When the number of consecutive frames generating virtual coordinates does not exceed the preset maximum number of completion frames, the virtual coordinates are written to the historical trajectory cache queue; when the number of consecutive frames generating virtual coordinates exceeds the preset maximum number of completion frames, completion is stopped and a trajectory interruption flag is output. When a subsequent frame re-obtains valid 3D spatial coordinates and the spatial position difference between the frame and the previous virtual coordinates does not exceed the preset reconnection distance threshold, the re-obtained valid 3D spatial coordinates will be added to the historical trajectory cache queue. The effective three-dimensional spatial coordinates and virtual coordinates are arranged according to the frame number and then subjected to median smoothing to generate a continuous hand spatial trajectory sequence.

5. The faucet start / stop control method based on motion-sensing interaction according to claim 1, characterized in that, Step four specifically includes: Read the three-dimensional spatial coordinates of wrist key points in a continuous hand spatial trajectory sequence, and extract the Z-axis coordinate sequence according to the frame number; Set up an integral sliding window, calculate the difference in Z-axis coordinates between two adjacent frames frame by frame, and generate the relative displacement of the Z-axis. The dead zone threshold is generated based on the sorting results of the absolute values ​​of the relative displacements along the Z-axis in the static rubbing calibration samples. The Z-axis relative displacements with absolute values ​​less than the dead zone threshold within the integral sliding window are marked as rubbing and shaking displacements and written to the zero displacement marker. The Z-axis relative displacements with absolute values ​​not less than the dead zone threshold are retained to generate the filtered Z-axis relative displacements. The filtered Z-axis relative displacement is accumulated with a sign according to the frame number. The Z-axis relative displacement away from the outlet is written into the positive accumulation term, and the Z-axis relative displacement closer to the outlet is written into the negative accumulation term. The positive and negative accumulation terms are superimposed to generate the directional accumulation value. Write the filtered Z-axis relative displacement, zero displacement marker, positive cumulative term, negative cumulative term, and direction cumulative value into the integration buffer.

6. The faucet start / stop control method based on motion-sensing interaction according to claim 1, characterized in that, Step five specifically includes: Read the occlusion confidence and effective 3D spatial coordinate output results of the current frame. When the occlusion confidence is lower than the preset safety threshold or the current frame does not output effective 3D spatial coordinates, write the occlusion observation flag and update the continuous occlusion duration frame count. Otherwise, write the no-occlusion observation flag and clear the continuous occlusion duration frame count. Within the current decision window, read the X-axis coordinates, Y-axis coordinates, and Z-axis coordinates from the continuous hand spatial trajectory sequence, and generate the mean planar displacement, Z-axis direction continuity, and trajectory completion percentage. Normal water outflow state, foaming and suspension state, and actual departure state are selected as candidate occlusion states, and state transition admission conditions, preset minimum duration frame number, preset maximum duration frame number, duration frame number center value, and duration frame number fluctuation value are configured for each candidate occlusion state. For each candidate occlusion state, a semi-Markov duration constraint calculation is performed. When the number of consecutive occlusion duration frames exceeds the corresponding preset minimum duration frame number and preset maximum duration frame number limit range, the corresponding duration weight is reset to zero. When the number of consecutive occlusions is within the corresponding limit, the duration weight is generated based on the difference between the number of consecutive occlusions and the center value of the number of consecutive frames, and the fluctuation value of the number of consecutive frames. Based on the occlusion observation markers, unocclusion observation markers, previous control state, mean plane displacement, Z-axis continuity, and trajectory completion ratio, the state matching weight is accumulated to generate a matching score for each candidate occlusion state. The matching score of the candidate occlusion state that passes the state transition admission condition is multiplied by the duration weight to generate a state score after semi-Markov constraint. The candidate occlusion state with the highest state score is selected to generate a normal water discharge state, a foaming and floating state, or a real departure state.

7. The faucet start / stop control method based on motion-sensing interaction according to claim 1, characterized in that, Step six specifically includes: Read the direction accumulation value and the occlusion state retention determination result, compare the direction accumulation value with the preset positive pause threshold and the preset negative recovery threshold, and write the occlusion state retention determination result into the current state determination bit; When the cumulative value of the direction is not less than the preset positive pause threshold and the current state determination bit is the normal water discharge state, a pause control intention is generated. When the cumulative value of the direction is not greater than the preset negative recovery threshold and the current state determination bit is normal water outflow state or foam suspension state, a water outflow control intention is generated. When the current state determination bit indicates a true departure state, a shutdown control intent is generated; When the current state determination bit is the foaming suspension state and the direction accumulation value has not reached the preset negative recovery threshold, a pause hold control intention is generated; When the finite state machine is in the paused memory state and the occlusion state holds the decision result as the bubble-floating state, the control intention to close is masked and a pause hold control intention is generated. The control intention input is a finite state machine containing the off state, water outlet state and pause memory state. When the pause control intention is received in the water outlet state, it switches to the pause memory state and outputs the off drive signal. When the water outlet control intention is received in the pause memory state, it switches to the water outlet state and outputs the on drive signal. When the off control intention is received, it switches to the off state and outputs the off drive signal. After the finite state machine completes the state transition, the accumulated direction values ​​that have participated in the generation of control intent in the current decision window are cleared, the previous control state corresponding to the pause memory state is retained, and the state transition result is written into the occlusion state retention decision input of the next frame.

8. The faucet start / stop control method based on motion-sensing interaction according to claim 1, characterized in that, Step seven specifically includes: Read the 3D spatial coordinates of all wrist key points output by the anti-occlusion customized YOLOv8-pose model in the current frame, and remove wrist key points whose occlusion confidence is lower than the preset safety threshold; Read the spatial coordinates of the reference point directly below the outlet, calculate the X-axis coordinate difference, Y-axis coordinate difference, and Z-axis coordinate difference between the three-dimensional spatial coordinates of each wrist key point and the reference point directly below the outlet, and generate the corresponding Euclidean distance; The remaining wrist keypoints are sorted in ascending order of Euclidean distance. The wrist keypoint with the smallest Euclidean distance is written into the target wrist keypoint marker. Only the target wrist keypoint is input into the historical trajectory cache queue, direction accumulation value calculation, and occlusion status preservation determination. When the target wrist keypoints switch in several consecutive frames and the difference in Euclidean distance between the new target wrist keypoint and the previous target wrist keypoint does not exceed the preset switching tolerance, the previous target wrist keypoint mark is maintained. Read the duration of the finite state machine in the paused memory state. When the paused memory state duration exceeds the preset safe duration, trigger the timeout circuit breaker, clear the integral cache, historical trajectory cache queue, completion flag and target wrist key point flag, and reset the finite state machine to the closed state.

9. A faucet start / stop control system based on motion-sensing interaction, applied to the faucet start / stop control method based on motion-sensing interaction as described in any one of claims 1-8, characterized in that, Includes the following modules: The 3D vision access module is used to access a 3D video stream containing the user's hand and water flow and perform polarization de-reflection processing. The wrist key point parsing module, connected to the 3D vision access module, is used to run an anti-occlusion customized YOLOv8-pose model and output the 3D spatial coordinates and occlusion confidence of all wrist key points. The target locking module, connected to the wrist key point parsing module, is used to generate target wrist key point markers based on the Euclidean distance from each wrist key point to the reference point directly below the outlet. The trajectory continuity maintenance module, connected to the target locking module, is used to generate a continuous hand spatial trajectory sequence based on the occlusion confidence of the target's wrist key points and the three-dimensional spatial coordinates of the wrist key points. The directional intent parsing module, connected to the trajectory continuity maintenance module, is used to generate directional cumulative values ​​based on a continuous sequence of hand spatial trajectories. The state maintenance determination module is connected to the wrist key point parsing module, trajectory continuity maintenance module and finite state control module. It is used to generate occlusion state maintenance determination results based on occlusion confidence, continuous hand spatial trajectory sequence and the previous control state of the finite state machine. The finite state control module, connected to the direction intent parsing module, the state holding determination module, and the target locking module, is used to generate control intent based on the direction accumulation value, the occlusion state holding determination result, and the target wrist constraint signal. It switches between the closed state, the water outlet state, and the pause memory state to control the start and stop of the solenoid valve, and triggers the timeout fuse after the pause memory state expires.

Citation Information

Patent Citations

  • Intelligent faucet

    CN120907000A

  • Foam hand washing machine

    CN213551466U