Nursing bed with video chat function and communication system thereof

By designing a folding bracket and lifting module to install a smart screen on the nursing bed, and combining it with a facial tracking and voice control module, automatic video calls for users with special conditions are realized. This solves the problems of facial tracking and voice recognition in low light environments and background noise, and improves the naturalness of video calls and the clarity of voice.

CN121868053APending Publication Date: 2026-04-17GUANGZHOU LIJIE MEDICAL EQUIP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-13
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing smart nursing beds cannot automatically adjust the smart screen to enable video chat, especially for users with special conditions such as hemiplegia, dystonia, fractures, or joint ankylosis, who find it difficult to operate the smart screen for video calls on their own.

Method used

A nursing bed with video chat function was designed. A smart screen is installed through a folding bracket and lifting module. Combined with a face tracking module and a voice control module, it can automatically track the user's face, prioritize emergency signals, adapt to low light environment and background noise, and adopt a shoulder and face correlation tracking mode to improve the accuracy of voice recognition.

Benefits of technology

It enables automatic video calls for users with special medical conditions, solves the problems of difficult facial tracking in low light environments and low speech recognition accuracy in background noise, and improves the naturalness of video calls and speech clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121868053A_ABST
    Figure CN121868053A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of nursing beds, and provides a nursing bed with a video chat function and a communication system thereof, the nursing bed comprises a bed body, an intelligent screen, a folding support, a lifting module and a control electric box, the intelligent screen is movably mounted on a headboard or a guardrail of the bed body through the folding support and the lifting module; wherein a call program is built in the intelligent screen; a main control module is arranged in the control electric box; wherein the main control module starts a communication system of the intelligent screen through a first signal, and the first signal comprises a user body abnormity signal, a voice communication request signal and a one-key communication signal installed on a headboard; the intelligent screen is integrated with a camera, a microphone and a touch display unit; and the main control module is connected with the intelligent screen, the driving mechanism of the folding bracket and the driving mechanism of the lifting module so as to control the spatial position adjustment of the intelligent screen and the starting and operation of a video call function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart bed technology, and in particular to a nursing bed with video chat function and its communication system. Background Technology

[0002] Intelligent nursing beds are a representative product that integrates an aging society with smart healthcare, and they have shown a rapid development trend in recent years.

[0003] Existing smart beds are equipped with smart screens, whose main function is to enable video entertainment through user control. Some nursing bed manufacturers have also proposed adding call functionality to the smart screens of nursing beds to assist users in video chatting.

[0004] However, for users of nursing beds who suffer from hemiplegia, dystonia, fractures, or joint ankylosis (folded-over person), these users often cannot operate the smart screen themselves. Therefore, the smart bed needs to have certain mechanical control functions to minimize the number of times the user needs to adjust their position in order to enable video calls.

[0005] Existing smart beds do not yet have the above functions. Summary of the Invention

[0006] This application proposes a nursing bed with video chat function and its communication system, which is used to automatically track the user's face to achieve tracking-based automatic communication.

[0007] To achieve the above objectives, this application provides the following technical solution: Firstly, this application proposes a nursing bed with video chat functionality, comprising a bed frame, a smart screen, a folding support, a lifting module, and a control box: The smart screen is movably installed on the headboard or guardrail of the bed frame via a folding bracket and a lifting module; the smart screen has a built-in call program. The control box contains a main control module; the main control module activates the smart screen's call system via a first signal, which includes a user's abnormal physical signal, a voice call request signal, and a one-button call signal installed on the headboard. The smart screen integrates a camera, microphone, and touch display unit; The main control module is connected to the drive mechanism of the smart screen, the folding bracket, and the lifting module, respectively, to control the spatial position adjustment of the smart screen and the start and operation of the video call function.

[0008] Secondly, this application proposes a communication system applicable to the aforementioned nursing bed with video chat functionality, the system comprising: Call response module: used to receive the first signal and initiate the call program on the smart screen; Face tracking module: When the call is started, it drives the folding bracket and lifting module to control the smart screen to keep in sync with the user's face based on the user's real-time facial position; Voice control module: Used to set the sound reception range of the smart screen and the user's voiceprint according to the synchronization trajectory, and to synchronously collect the voice of the call from the location of the user's face.

[0009] In conjunction with the second aspect, receiving the first signal further includes: When the first signal includes both a user's physical abnormality signal and a voice call request signal, the response priority of the physical abnormality signal is higher than that of the voice call request signal, and a first call instruction is generated. The first call instruction is used to initiate a video call with the nurses' station and delay the processing of voice call requests until the abnormal signal is cleared.

[0010] In conjunction with the second aspect, generating the first call instruction further includes: The pressure change rate of the back pressure sensor array, the heart rate value of the heart rate sensor, and the pressure drop amplitude of the bed exit detection sensor were collected. A signal is considered a genuine anomaly when at least two sensor data points meet a threshold condition. If only one condition is met or the data fluctuation duration does not reach the preset time, it is determined to be a false trigger, a suspected abnormal log is generated and stored locally, and when the number of false triggers reaches the preset number within the specified time, the smart screen automatically pops up a sensor calibration prompt and sends a calibration request to the administrator terminal.

[0011] In conjunction with the second aspect, when the call program is started, it also includes: Detect real-time light intensity and determine whether the real-time light intensity is lower than the preset light intensity value; Among them, when the real-time light intensity is lower than the preset light intensity value, the built-in infrared fill light of the smart screen frame is automatically turned on, and the facial feature point extraction algorithm is switched to infrared thermal imaging and contour priority mode. The three-dimensional facial contour data is obtained through the TOF depth sensor and matched with skeletal feature points to reduce the feature point matching threshold. When the area covered by facial occlusion exceeds a preset ratio, shoulder and face correlation tracking is activated. The face position is predicted by the preset coordinate mapping relationship between the center point of the shoulder contour and the center point of the face, so that the tracking interruption time is controlled within a preset time range.

[0012] In conjunction with the second aspect, when the call program is started, it also includes: The system determines whether the pressure drop detected by the hip pressure sensor exceeds a preset ratio and the duration reaches a preset value, and whether the infrared ranging sensor does not detect the presence of a human body within the preset range, or whether the user manually triggers the privacy mode icon through the smart screen touch unit and passes the verification. Disconnect the power supply to the smart screen's camera, control the folding bracket to rotate the smart screen to face the foot of the bed, switch the microphone array to silent monitoring mode, and when the user returns to the bed, the system needs to be reactivated via voice command or one-click call button.

[0013] In conjunction with the second aspect, the drive folding bracket and lifting module control the smart screen to move in a synchronized trajectory with the user's face, and also includes: The user priority sorting is preset and stored in the facial feature template library. When the facial tracking module detects both the user and the medical staff at the same time, if the facial feature matching degree of the medical staff reaches the preset value and continues to appear in the field of view of the smart screen for a preset time, the dual-target tracking mode is activated. When the dual-target tracking mode is executed, the smart screen displays a main screen and a picture-in-picture. The folding bracket and lifting module respond to changes in the user's position first. Changes in the position of medical staff only trigger horizontal rotation fine-tuning. When medical staff leave or the facial feature matching degree is lower than the preset value, the single-target tracking mode is restored within a specified time and the picture-in-picture is automatically turned off.

[0014] In conjunction with the second aspect, the synchronous acquisition of call audio based on the user's facial location includes: The sound source azimuth and distance are calculated by a time delay estimation algorithm and a sound source heat map is generated. When the sound source corresponding to the user's voiceprint is detected to be within the preset distance and azimuth range, the first-level beamforming is activated. If other interfering sound sources are present, gain attenuation is applied to the direction of the interfering sound sources. When the signal-to-noise ratio of the user's voiceprint signal is lower than the preset value, secondary enhancement is triggered. The gain of the weak signal area is increased by dynamic range compression and howling suppression is started. Howling suppression is stopped when the speech clarity meets the preset standard.

[0015] In conjunction with the second aspect, the synchronous acquisition of call audio based on the location of the user's face also includes: Predefined core voice commands and their corresponding dialect versions and speech rate models; When a user issues a voice command, the frequency band signal is extracted by FFT filtering and power frequency noise and high frequency noise are removed. The real-time voice features are compared with the pre-stored template by a dynamic template matching algorithm. Among them, the instruction is considered valid when the similarity reaches a preset value and the instruction duration is within the specified range; When the similarity is within a preset range, a list of command candidates is displayed on the smart screen for the user to confirm. When the number of consecutive command recognition failures reaches a preset value, the system automatically switches to touch-assisted mode and pops up a command shortcut button at the bottom of the screen.

[0016] In conjunction with the second aspect, the synchronous acquisition of call audio based on the location of the user's face also includes: Collect the user's voice recordings and convert them into text. Based on the call text, generate call captions on the smart screen.

[0017] The beneficial effects of this invention are as follows: This application primarily addresses the challenges faced by users with specific medical conditions who have difficulty operating smart screens for video calls, including difficulties with facial tracking in low-light environments and low speech recognition accuracy in noisy environments, particularly in hospitals where the environment is noisy and patient rooms are often poorly lit. When the face is occluded, the system automatically switches from a conventional facial tracking mode to a shoulder-face correlation tracking mode. It predicts the facial position using a preset coordinate mapping relationship between the center point of the shoulder contour and the center point of the face, thus resolving the tracking interruption problem caused by facial occlusion. Through time delay estimation algorithms and beamforming, the system calculates the azimuth and distance of the sound source, and dynamically adjusts the directivity of the microphone array using a delay-summation beamforming algorithm, addressing the low speech recognition accuracy in noisy environments. Furthermore, the speech localization in speech recognition and the visual localization in the face occlusion situation complement each other; facial tracking improves the patient's speech reception angle, and speech localization improves the efficiency of facial tracking.

[0018] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0019] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0021] In the attached diagram: Figure 1 This is a schematic diagram illustrating the structure of a nursing bed with video chat functionality according to an embodiment of the present invention. Figure 2 This is a diagram illustrating the composition of a communication system according to an embodiment of the present invention; Figure 3 This is a simulation diagram of a video chat scenario in an embodiment of the present invention. Detailed Implementation

[0022] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention. Example 1:

[0023] See Figure 1 and Figure 3 This application proposes a nursing bed with video chat function, including a bed frame, a smart screen, a folding support, a lifting module and a control box.

[0024] In this application, the basic load-bearing component of the nursing bed, used for the patient to lie down, is the physical carrier of the system. The core hardware for the smart screen's video chat function integrates display, input / output, and communication modules. The folding bracket is a mechanical structure connecting the smart screen to the bed frame, used to achieve multi-angle folding and storage of the smart screen (folded to fit the bed frame or unfolded to a usable angle). The lifting module is a mechanical device for adjusting the height of the smart screen, adapting to the viewing needs of users of different heights or in lying / sitting positions. The control box is the electrical control center, with built-in circuitry and control modules, coordinating the operation of all components.

[0025] The smart screen is movably mounted on the headboard or guardrail of the bed via a folding bracket and a lifting module. The smart screen has a built-in calling app. The folding bracket and lifting module allow for flexible installation and adjustment of its position. The built-in calling app is the core software, supporting video chat functionality. Its placement on the headboard or guardrail is close to the user's head for easy operation and viewing, without taking up sleeping space. The smart screen's built-in calling app comes pre-installed with video calling software, allowing users to initiate and receive video calls without additional equipment.

[0026] The control box contains a main control module; the main control module activates the smart screen's call system via a first signal, which includes a user's abnormal physical signal, a voice call request signal, and a one-button call signal installed on the headboard. The main control module acts as the brain, receiving external signals and triggering the communication system to enable automatic / manual call triggering in multiple scenarios. The main control module is the core control unit, responsible for signal processing, logical judgment, and command issuance. The first signal is the set of input conditions that trigger a call, covering three scenarios: Abnormal user physical signals: possibly from sensors integrated into the bed (pressure sensor detects falls and heart rate sensor detects abnormalities), enabling emergency calls. The voice call request signal is triggered by the user via voice command (calling a nurse), requiring the voice recognition module. The one-button call signal is a physical button located on the headboard, allowing users (especially those with limited mobility) to quickly initiate a call.

[0027] The smart screen integrates a camera, microphone, and touch display unit. It achieves video chat input and output functions through hardware integration, eliminating the need for external devices. The camera captures the user's video feed for the other party to view in real time. The microphone captures the user's voice signal, enabling two-way audio communication. The touch display unit combines display (outputting video / user interface) and input (touch operations, such as dialing and adjusting volume) functions.

[0028] The main control module is connected to the drive mechanisms of the smart screen, the folding bracket, and the lifting module to control the spatial position adjustment of the smart screen and the activation and operation of the video call function. Through the electrical connection of the main control module, centralized control of mechanical adjustment (position) and communication functions (video chat) is achieved, ensuring the coordinated operation of all components. The drive mechanism (motor, hydraulic device) is the power component of the folding bracket and lifting module. After receiving commands from the main control module, it drives the mechanical structure to move, adjusting the height and angle of the smart screen. Electrical connection is used: electrical signals are transmitted through wires or a wireless module, enabling the main control module to control the smart screen (initiating calls, transmitting audio and video signals) and the drive mechanism (controlling position adjustment) in real time. Example 2:

[0029] See Figure 2 and Figure 3 This application proposes a communication system applicable to the aforementioned nursing bed with video chat functionality. The system uses a modular design to separate functions, corresponding to signal processing, mechanical control, and audio optimization.

[0030] Call Response Module: This module receives the initial signal and initiates the smart screen's call process. It acts as an intermediate layer for signal reception and execution, connecting the underlying hardware signals (one-button call, sensors) with the upper-layer call software to ensure reliable triggering logic (preventing accidental touches and prioritizing signals). The initial signal primarily includes abnormal user bodily signals (e.g., falling out of bed detected by sensors, abnormal heart rate), voice call request signals (e.g., the user saying "call a doctor"), and one-button call signals (physical buttons on the bedside table), covering both proactive calls and reactive emergency triggering scenarios.

[0031] Face tracking module: When the call is started, it drives the folding bracket and lifting module to control the smart screen to keep in sync with the user's face based on the user's real-time facial position; This application combines computer vision technology with mechanical control to control the dynamic adjustment of a smart screen. Facial images are captured by a camera, and algorithms locate facial coordinates. The drive mechanisms (such as stepper motors) of the folding bracket and lifting module are then controlled to adjust the screen's angle and height in real time, ensuring the face remains centered on the screen. During real-time facial positioning, images are captured by the camera integrated into the smart screen. AI algorithms, specifically OpenCV face detection and deep learning keypoint localization, calculate the face's coordinates (X / Y / Z axis positions) in three-dimensional space. The folding bracket and lifting module are then driven, generating specific mechanical control commands based on the facial coordinates to determine the bracket's rotation angle and the module's lifting height. During operation, the main control module sends these commands to the drive mechanisms, such as motors and hydraulic devices, for physical position adjustment. The smart screen's movement path dynamically mirrors the face's movement path; the screen rotates synchronously when the user turns their head and rises when they sit up. This method solves the problem of viewing angle shift in traditional fixed screens, improving the naturalness of video call interaction.

[0032] The audio reception control module primarily uses a synchronized trajectory to set the reception range of the smart screen and the user's voiceprint (i.e., the patient's voiceprint). During facial tracking, it simultaneously collects the voice recordings of the user's face. By optimizing audio acquisition based on facial location information, it achieves spatial localization and biometric recognition of the patient's face, resulting in directional audio reception and noise reduction enhancement, thus improving voice call clarity. Simultaneously, the reception range is set according to the synchronized trajectory. This means that the beamforming direction of the microphone array is controlled by the real-time facial coordinates of the facial tracking module, locking the focus of the reception and the user's facial position. When the user's face moves to the left, the microphone array's pickup direction shifts to the left to reduce background noise from non-target directions. The user's voiceprint is obtained through pre-stored user voiceprint features. Voice samples collected during registration are used for identity verification and noise reduction, amplifying only the voice signal matching the voiceprint and filtering out environmental noise or other people's conversations. During user registration, a privacy authorization function is set. This application's smart bed collects necessary privacy data related to the user's smart bed functions. During the process of simultaneously capturing the voice of the face, not only is the audio capture and image tracking synchronized in time and space, ensuring that the facial position in the image is consistent with the position of the sound source, but also the problems of off-screen sound or sound delay are avoided, thus enhancing the immersive experience of the call.

[0033] In the process of facial tracking, this application employs a deep learning model for facial tracking. This model is based on an improved MobileNetV3-Small backbone network, and its structure includes: one input layer for extracting facial features, 15 convolutional layers for multi-scale facial feature fusion, three Squeeze-and-Excitation modules for facial heatmap tracking, and two fully connected layers for outputting the specific tracking results. The facial tracking uses a dual-path network structure: one path uses MobileNetV3 to extract global features, and the other path uses a lightweight U-Net to extract local details. The features from both paths are fused at 1 / 4 resolution. The network input is a 160×120×3 RGB image, and the output is the coordinates of 68 facial key points and their corresponding heatmaps. Training employs a multi-task loss function, including coordinate regression loss, heatmap loss, and occlusion perception loss, with a weight ratio of 0.6:0.3:0.1.

[0034] In actual implementation, upon initial startup of the smart bed, a concise privacy policy is displayed on the smart screen, including the scope, purpose, and storage period of data collection. Users must explicitly authorize the data via touch or voice confirmation. Authorization information is encrypted and stored in a secure chip, including the scope of authorization, timestamp, and digital signature. Facial data processing employs three levels of privacy protection: Level 1 involves short-term caching of the original image in memory; Level 2 involves irreversible transformation of the extracted facial feature vectors; and Level 3 involves adding differential privacy noise to the transmitted data, with a privacy budget ε=0.5. The system automatically clears local raw data every 24 hours, retaining only irreversible feature summaries for algorithm optimization. Users can disable all AI functions at any time with a single physical button press 'Privacy Protection,' at which point the system immediately deletes the biometric data from memory and switches to basic call mode. Example 3:

[0035] This application addresses the issue of multiple signal conflicts when receiving the first signal, specifically the simultaneous occurrence of multiple trigger signals.

[0036] When the first signal includes both a user's physical abnormality signal and a voice call request signal, the response priority of the physical abnormality signal is higher than that of the voice call request signal, and a first call instruction is generated. This application resolves signal conflicts through a priority sorting mechanism—when two types of input are received simultaneously, namely, emergency signals and proactive requests, emergency scenarios are responded to first, ensuring that user safety needs take precedence over regular communication needs.

[0037] It is understandable that when a user experiences abnormal bodily signals, it indicates that the patient, i.e., the user, is in a dangerous state and requires immediate intervention. In practice, emergency signal types include: sensors detecting a fall from bed, sudden changes in heart rate, and prolonged inactivity. It is also understood that this application provides dedicated instructions for emergency scenarios, triggering the system to execute preset emergency response procedures, distinct from regular voice communication instructions. For example: calling family members.

[0038] In this application, the first call command is used to initiate a video call with the nurse station and delay processing voice call requests until the abnormal signal is cleared. During execution, the first call command directly connects to the responsible medical party (nurse station) and suspends non-emergency requests to ensure priority handling of emergency events while preventing signal loss. The first call command is used for targeted calls, which make emergency calls through a pre-stored contact list. The pre-stored contact list is stored in the main control module or the cloud. The first call command can directly dial the nurse station terminal or directly connect to the nursing station computer and nurse wristband, skipping the conventional dialing process and shortening the response time.

[0039] In one embodiment, when there are multiple call requests, the voice call requests are stored in a cache queue instead of being directly rejected or discarded, and the abnormal signal is cleared. At this time, if the nurse manually resets the system after confirming the user's safety, or if the sensor detects that the user's status has returned to normal, the processing of the original voice request is automatically triggered again.

[0040] Once the abnormal signal is cleared, the conditions for triggering the delay to end are established. These conditions include manual confirmation and automatic recovery.

[0041] Manual confirmation involves a nurse verifying the user's safety via video call and then sending a command to remove the abnormality through the terminal at the nurse's station. Automatic recovery means that if abnormal bodily signals disappear during continuous sensor monitoring, i.e., the heart rate returns to normal and the risk of falling out of bed is eliminated, and the patient gives other normal instructions, the abnormality is automatically determined to be resolved. Example 4:

[0042] In generating the first call command, this application also constructs an input matrix for judging abnormal signals by collecting data from multi-dimensional physiological and behavioral sensors, thereby avoiding false triggering by a single sensor and improving the reliability of emergency signals.

[0043] The system collects data on the pressure change rate of the back pressure sensor array, the heart rate value from the heart rate sensor, and the pressure drop amplitude from the bed exit detection sensor. The back pressure sensor array, a distributed pressure sensor (thin-film pressure sensor) installed inside the mattress, detects changes in pressure distribution (changes in pressure areas when the user turns over, gets out of bed, or falls out of bed). The pressure change rate is the speed at which the pressure value changes per unit time (the rate at which pressure drops from a normal value to 0 when falling out of bed), used to distinguish between slow bed exit (voluntary getting out of bed) and rapid bed exit (accidental event). The heart rate sensor, a physiological monitoring module (PPG photoplethysmography sensor) integrated into the mattress or wristband, collects heart rate data in real time; abnormal values ​​may alert the user to a physiological crisis. The bed exit detection sensor, typically a pressure sensor along the edge of the bed or mattress, monitors whether the user has completely left the bed; its pressure drop amplitude refers to the amount of pressure decrease at the moment the user leaves the bed.

[0044] A genuine abnormal signal is determined when at least two sensor data points meet a threshold condition. Meeting at least two threshold conditions is based on a sensor cross-validation scheme, which reduces false triggers caused by single sensor failures or interference (such as occasional fluctuations in heart rate sensors or external impacts to pressure sensors) by ensuring consistency across independent data sources.

[0045] If only one condition is met or the data fluctuation duration does not reach the preset time, it is judged as a false trigger, a suspected anomaly log is generated and stored locally, and when the number of false triggers accumulates to the preset number within a specified time, the smart screen automatically pops up a sensor calibration prompt and sends a calibration request to the administrator terminal. In order to eliminate instantaneous interference or single faults, and to record suspected events for traceability and analysis, a time window is set to filter short-term fluctuations (instantaneous pressure changes when the user turns over, or a brief increase in heart rate due to coughing). Only when abnormal data lasts for more than a preset time (abnormal heart rate lasting more than 5 seconds) is it included in the judgment, reducing false judgments. Detailed information of false triggers (time, sensor type, abnormal value, duration) is recorded and stored in the local storage module of the control box, and is not uploaded to the cloud to protect user privacy, while also providing maintenance personnel with information to troubleshoot sensor faults (frequent false alarms of a certain sensor may be due to hardware aging). Example 5:

[0046] This application also needs to ensure that facial tracking fails when the call process is initiated.

[0047] Detect real-time light intensity and determine whether the real-time light intensity is lower than the preset light intensity value; Real-time light intensity represents the current ambient light level and is collected by an ambient light sensor (photoresistor or CMOS sensor) integrated into the smart screen.

[0048] The preset light intensity value represents the preset light threshold. When the light intensity is below this value, it is determined to be a low light environment, triggering supplemental lighting and algorithm switching.

[0049] Among them, when the real-time light intensity is lower than the preset light intensity value, the built-in infrared fill light of the smart screen frame is automatically turned on, and the facial feature point extraction algorithm is switched to infrared thermal imaging and contour priority mode. The three-dimensional facial contour data is obtained through the TOF depth sensor and matched with skeletal feature points to reduce the feature point matching threshold. In one embodiment, in low-light environments, the problem of blurred facial features caused by insufficient visible light is solved by using hardware supplemental lighting, algorithm optimization, and 3D data enhancement, ensuring that the facial tracking module can still work stably.

[0050] The infrared fill light is a near-infrared LED light source built into the bezel of the smart screen. It is used to provide active illumination without affecting the human eye, enabling the camera to capture clear infrared facial images, and avoiding the glare or sleep disturbance caused by visible light fill light.

[0051] Infrared thermal imaging and contour-priority mode: Infrared thermal imaging uses the temperature distribution on the face to generate a thermal image. At this time, the skin temperature is higher than that of the environment, highlighting the facial area, which is not affected by the intensity of visible light.

[0052] The contour-first mode algorithm extracts detailed feature points, specifically extracting the texture of the eyes and nose, and then switches to edge contour extraction to determine the edge feature regions such as the outer contour of the face and the jawline, reducing the dependence on the clarity of local details.

[0053] Time-of-flight (TOF) depth sensors emit infrared signals and calculate the time difference between the emission and reflection of light back to the sensor to construct a 3D point cloud model of the face. This allows them to acquire X / Y / Z axis coordinate data and determine facial features such as nose bridge height and facial contours.

[0054] Matching skeletal feature points is based on 3D contour data to locate rigid skeletal feature points such as the skull and mandible, rather than occluded skin features, in order to improve tracking stability. In terms of actual technical effect, even when the user is wearing a mask, tracking can still be performed through the mandibular contour.

[0055] Lowering the feature point matching threshold, in specific implementation, means that under normal lighting conditions, 80% of feature points need to be successfully matched to determine that tracking is effective, and in low light conditions, it is reduced to 50%. At this point, more missing feature points can be received, preventing some features from becoming blurred and causing tracking interruption.

[0056] In actual implementation, if the area covered by the facial occlusion exceeds the preset ratio, shoulder and face correlation tracking is activated. The facial position is predicted by the preset coordinate mapping relationship between the center point of the shoulder contour and the center point of the face, so as to achieve the effect of controlling the tracking interruption time within the preset time range.

[0057] This application addresses scenarios involving partial facial occlusion, such as wearing masks or covering the face with blankets. It infers facial position through a body part association model to achieve continuous tracking and prevent video loss or smart screen malfunction due to occlusion. If the facial occlusion area exceeds a preset proportion, it indicates the user is covering half their face with their hand or wearing an oxygen mask. Shoulder-face association tracking is an alternative tracking scheme based on the correlation of human body structure. The shoulder contour is relatively stable and not easily completely obscured; the face position is inferred from the shoulder position. Tracking time is controlled within a preset range to ensure smooth smart screen adjustments. Even when the user covers their face, the screen can still roughly locate the face based on shoulder movement, avoiding severe image jitter.

[0058] In practical implementation, the preset ratio represents the percentage threshold of the face covered by occlusions relative to the entire facial detection area. When the detected facial occlusion ratio exceeds the corresponding threshold, the system automatically switches from the regular face tracking mode to the shoulder-face correlation tracking mode. The typical ratio is 50%, preferably 45%–55%. In actual implementation, adaptive adjustments are needed based on environmental conditions; for example, the threshold is 60% in low-light environments and reduced to 40% in bright light environments. Users can also adjust this parameter in the settings interface based on their personal habits, with the adjustment range limited to 30%–70%. For certain specific conditions, such as Parkinson's disease patients, a personalized preset threshold of up to 75% can be used to accommodate their involuntary facial movement characteristics. Example 6:

[0059] This application covers proactive privacy needs when a user leaves their device without turning it off, thus protecting user privacy in terms of privacy protection.

[0060] The system determines whether a sudden drop in pressure detected by the hip pressure sensor exceeds a preset percentage and lasts for a preset duration, and also determines whether the infrared ranging sensor does not detect a human presence within a preset range, or whether the user manually triggers the privacy mode icon via the smart screen touch unit and passes verification. The hip pressure sensor, a pressure sensing device installed in the hip area of ​​the mattress, detects whether the user has left the bed. The determination of a sudden drop in pressure is based on whether the user has left the bed. A preset duration ensures that the user has indeed left the bed by eliminating interference from brief posture adjustments. The infrared ranging sensor, installed on the side of the bed or on the smart screen, detects the presence of the user or doctor within a preset range, thus determining whether the user has left the bed and no one is nearby.

[0061] Disconnect the power supply to the smart screen's camera, control the folding bracket to rotate the smart screen to face the foot of the bed, switch the microphone array to silent monitoring mode, and when the user returns to the bed, the system needs to be reactivated via voice command or one-click call button.

[0062] Disconnecting the power supply to the smart screen's camera completely disables it at the hardware level, which is safer than software shutdown and prevents malicious programs from bypassing software permissions to launch the camera. Directly cutting off the camera module's power line (controlled by a relay switch via the main control module) ensures the camera has no operating current and cannot capture images or videos, eliminating the risk of privacy leaks (hackers remotely waking up the camera) at the source. Users lying on nursing beds generally require even greater privacy protection. In practical implementation, this application allows the folding bracket to rotate the smart screen towards the foot of the bed, using a mechanical structure to physically block the view. Even if the camera is not completely powered off (e.g., due to hardware failure), adjusting the screen's orientation shifts the lens away from the user area, providing both software and hardware privacy protection. The foot of the bed is usually a wall or open area; therefore, after rotation, the camera's field of view cannot cover the bed frame and bedside table, such as during user dressing or care scenarios, physically isolating the shooting range. Simultaneously, the microphone array switches to silent monitoring mode, protecting voice privacy while retaining necessary wake-up functionality, thus balancing privacy and convenience. When muted monitoring is enabled, call audio recording must cease, and the microphone array must not transmit audio to the nurses' station or other terminals to prevent ambient sound leakage, such as private conversations between the user and their family. Local recognition of specific wake words is retained. Local recognition includes system activation and call initiation; no audio data is uploaded, and the system recovery process is triggered only upon detection of the wake word. Upon the user's return, active control capabilities are triggered to resume the call. Example 7:

[0063] This application drives the folding bracket and lifting module to control the smart screen to be in a synchronized trajectory with the user's face. Through user identification and priority preset, it realizes dynamic switching of multi-target tracking, solves the core problem of which target the smart screen should prioritize tracking when the user (patient) and medical staff appear in the field of vision at the same time in medical scenarios, and ensures that the main user experience is not disturbed.

[0064] User priority sorting is pre-configured, with built-in identity priority rules stored in the main control module's configuration file, defining resource allocation logic. Facial feature data of known individuals is pre-stored, such as facial key points and contour features of users and frequently used medical staff, for rapid identity matching. Briefly appearing individuals (passing nurses) are filtered out, typically for 3 seconds, ensuring that medical staff actually need to interact to trigger the dual-target mode.

[0065] The user priority sorting is preset and stored in the facial feature template library. When the facial tracking module detects both the user and the medical staff at the same time, if the facial feature matching degree of the medical staff reaches the preset value and continues to appear in the field of view of the smart screen for a preset time, the dual-target tracking mode is activated. When the dual-target tracking mode is executed, the smart screen displays a main screen and a picture-in-picture. The folding bracket and lifting module respond to changes in the user's position first. Changes in the position of medical staff only trigger horizontal rotation fine-tuning. When medical staff leave or the facial feature matching degree is lower than the preset value, the single-target tracking mode is restored within a specified time and the picture-in-picture is automatically turned off.

[0066] In one embodiment, this application employs a screen space allocation method for multi-target visualization. In screen space allocation mode, the main screen interacts with the user, while the picture-in-picture (Picture-in-Picture) provides auxiliary display for medical personnel, preventing information loss due to single-screen switching. The main screen displays the user's face, i.e., the tracked object, ensuring the remote end can clearly observe the patient's state, specifically including facial expressions and lip movements. The Picture-in-Picture displays the medical personnel's face for two-way communication; for example, when medical personnel instruct the patient on medication, the patient can see the medical personnel's gestures. In this application, a preset time represents the maximum allowable time interval between when the face is completely obscured and when the user's face is repositioned through shoulder-face correlation tracking, used to maintain tracking continuity, typically 0.6 seconds. In actual implementation, a three-level setting is used, for example: when facial occlusion exceeds a preset proportion, shoulder contour detection is immediately initiated (Level 1, time ≤ 0.2 seconds), simultaneously activating a preset coordinate mapping model (Level 2, time ≤ 0.15 seconds), and finally completing facial position prediction (Level 3, time ≤ 0.25 seconds).

[0067] In one embodiment, mechanical motion priority allocation ensures the accuracy of primary target tracking, avoiding frequent adjustments to the support due to the movement of multiple targets (the smart screen will not rotate violently and affect the patient's view when medical staff move around). Responding to changes in the user's position, the user's (patient's) head movement triggers full-dimensional adjustments to the folding support and lifting module (X-axis horizontal rotation, Y-axis pitch angle, Z-axis height adjustment), ensuring the main image remains centered and at a suitable distance (e.g., the support rises and tilts forward when the patient sits up). When medical staff move, only small-angle adjustments are made via the horizontal rotation motor of the smart screen base (rather than the lifting or pitch mechanism), maintaining the picture-in-picture position stable and avoiding patient viewpoint shifts or device noise interference caused by large-scale rotation.

[0068] In one embodiment, system resources are released through target disappearance detection and timeout recovery mechanisms to avoid meaningless multi-target tracking; for example, after medical staff leave, the empty picture-in-picture window affects the user experience. In this application, after medical staff leave, the infrared ranging sensor or facial tracking module determines that the medical staff is no longer in the field of view through the video feed. When the facial feature matching degree is lower than a preset value, indicating that the medical staff's face is obscured (e.g., turning around or wearing a mask resulting in a matching degree <50%), it is determined that effective tracking has been lost. The single-target mode is restored within a specified time, typically with a delay of 3-5 seconds, to avoid frequent mode switching due to brief obstructions (e.g., when a medical staff member bends down to pick something up). After the timeout, the picture-in-picture automatically closes, the main screen fills the entire screen, and the bracket only tracks the user. Example 8:

[0069] This application simultaneously collects the voice recordings of the call based on the location of the user's face, including: The sound source azimuth and distance are calculated by a time delay estimation algorithm and a sound source heat map is generated. When the sound source corresponding to the user's voiceprint is detected to be within the preset distance and azimuth range, the first-level beamforming is activated. In actual implementation, this application uses spatial acoustic localization technology based on microphone arrays to infer the location of the sound source by the time difference of sound received by multiple microphones, and generates a visual heat map by combining the intensity distribution for directional enhancement.

[0070] The time delay estimation algorithm characterizes the time difference between different microphones receiving the same sound signal in a microphone array. It calculates the distance difference from the sound source to each microphone using the formula: distance difference = speed of sound × time delay difference. Combined with the array's geometric position (microphone spacing), it solves for the three-dimensional coordinates (azimuth and distance) of the sound source. The sound source azimuth represents the horizontal angle of the sound source relative to the smart screen and is used for directional beamforming.

[0071] Because the distance represents the straight-line distance from the sound source to the smart screen, it affects the gain adjustment strategy (reducing gain in the near field to avoid overload, and increasing gain in the far field to enhance the signal). The sound source heatmap uses light and dark colors (red-yellow-blue) to represent the sound intensity of sound sources at different locations, intuitively displaying the sound field distribution (red areas represent strong sound sources, and blue represents background noise), helping the system to quickly locate the main sound source (the user).

[0072] In this application, a time delay estimation algorithm is used to estimate the measurement results of the time difference of arrival of the same sound source signal received by the microphone array between different microphones. First-level beamforming combines identity recognition and spatial positioning for dual screening, triggering basic directional sound pickup to ensure that only the target sound source (user) within the effective area is amplified, reducing invalid processing. Beamforming is used to amplify the sound source signal in a specific direction and suppress noise in other directions by applying specific time delays and weighting coefficients to the received signals of each channel of the microphone array.

[0073] In this application, a 4-channel circular microphone array with a diameter of 10cm is used for rapid localization of the main sound source. The audio signal is input at a sampling rate of 16kHz, and after being windowed by a 32ms Hamming window, it undergoes a 256-point STFT transformation. The GCC-PHAT algorithm calculates the time delay values ​​of the 6 microphone pairs to construct the sound source azimuth probability distribution. The beamforming module dynamically adjusts the phase delay of the 4 microphones according to this distribution to achieve directional reception with a 60° main lobe width. The system updates the beam pointing every 200ms to ensure tracking of moving speakers.

[0074] If other interfering sound sources are present, gain attenuation is applied to the direction of the interfering sound sources. When the signal-to-noise ratio of the user's voiceprint signal is lower than the preset value, secondary enhancement is triggered. The gain of the weak signal area is increased by dynamic range compression and howling suppression is started. Howling suppression is stopped when the speech clarity meets the preset standard.

[0075] In one embodiment, directional noise reduction suppresses interference from specific directions, preserving the target sound source while reducing multi-source aliasing and improving speech clarity. The direction of the interfering sound source is identified using a sound source heatmap to determine high-intensity sound sources other than the user. Gain attenuation applies a -10dB to -20dB attenuation to the sound signal in the direction of interference, reducing the intensity of the interference signal and minimizing masking of the target speech. Secondary enhancement, targeting weak signal scenarios, improves the clarity of the target speech through deep signal processing. Feedback suppression addresses weak signal distortion and audio loop feedback issues through dynamic range adjustment and feedback suppression, ensuring that the speech signal is not distorted or harsh during amplification. Feedback suppression introduces some signal distortion; therefore, processing ceases once the speech quality meets the standard, balancing noise reduction and sound quality preservation. Example 9:

[0076] This application simultaneously collects the voice recordings of the user's face location during the call, and also uses core voice commands and multi-scenario adaptation methods to ensure accurate command recognition.

[0077] Predefined core voice commands and their corresponding dialect versions and speech rate models; Core voice commands are key operational commands, and predefined definitions ensure a concise command set that covers high-frequency needs. Dialect versions are designed for elderly users in medical scenarios, with pre-stored dialect voice templates and dialect acoustic models optimizing recognition accuracy. A speech rate model is used to handle different speech rates; pre-stored speech rate templates prevent dialect comprehension errors caused by excessively fast or slow speech rates, and a duration normalization algorithm matches command features at different speech rates to prevent template matching failures due to speech rate differences.

[0078] When a user issues a voice command, this application extracts the frequency band signal through FFT filtering and removes power frequency noise and high frequency noise, and then uses a dynamic template matching algorithm to compare the real-time voice features with the pre-stored template. In practical implementation, this application uses spectrum analysis and targeted noise reduction to purify the speech signal, eliminate typical noise interference in the medical environment, and output the denoised target speech signal.

[0079] In this application, FFT filtering converts a time-domain speech signal into a frequency-domain spectrum to identify the energy distribution of different frequency components.

[0080] In this application, power frequency noise refers to power line interference generated by medical equipment, and the signal in this frequency band is attenuated in the frequency domain by a notch filter.

[0081] In this application, high-frequency noise refers to high-frequency noise generated by electronic components. A low-pass filter is used to retain the human voice frequency band and filter out high-frequency interference.

[0082] In this application, comparing real-time speech features with pre-stored templates improves matching flexibility through time-varying feature adaptation, addressing the issue that static templates cannot handle variations in speech duration and intonation (e.g., signal length differences caused by varying speech speed when a user says "call for a nurse"). Real-time speech features are extracted from the filtered speech signal, focusing on key features (Mel-frequency cepstral coefficients, short-time energy, zero-crossing rate), reflecting the spectral envelope and dynamic changes of the speech. Dynamic template matching differs from static templates (fixed length, fixed feature sequence) in that it allows for alignment of the time axis between real-time speech and pre-stored templates using time warping algorithms. Weights are dynamically adjusted during the matching process to enhance anti-interference capabilities.

[0083] Among them, the instruction is considered valid when the similarity reaches a preset value and the instruction duration is within the specified range; When the similarity is within a preset range, a list of candidate instructions is displayed on the smart screen for user confirmation. This application uses human-computer interaction to correct fuzzy matching scenarios, giving users the right to choose while ensuring recognition efficiency, and avoiding misjudgment or missed judgment due to similarity being close to the threshold.

[0084] When the number of consecutive command recognition failures reaches a preset value, the system automatically switches to touch-assisted mode and displays a command shortcut button at the bottom of the screen. This application ensures the availability of core functions through a multimodal degradation mechanism, seamlessly switching to touch input when voice interaction fails, preventing users from being stuck due to inability to operate (especially suitable for medical emergency scenarios). Example 10:

[0085] This application also includes simultaneously collecting the call audio from the user's facial location, and further includes: Collect the user's voice recordings and convert them into text. Based on the call text, generate call captions on the smart screen.

[0086] In one embodiment, the user's voice is collected, processed by an acoustic front-end, and output as a clean voice data stream. A deep learning model maps the voice data to text, and combined with predefined dialect versions and speech rate models, improves recognition accuracy in medical scenarios. This application converts the text output by ASR into visual subtitles, solving communication barriers caused by hearing impairment, noisy environments, or unclear speech. Punctuation is automatically added based on speech pauses and semantic models to avoid reading difficulties caused by long sentences without pauses. ASR recognition errors are corrected by combining context and a medical terminology database. After receiving the text data, the smart screen's display driver module generates an image layer through a subtitle rendering engine and plays the corresponding subtitles.

[0087] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A nursing bed with video chat function, comprising a bed frame, a smart screen, a folding support, a lifting module, and a control box, characterized in that: The smart screen is movably installed on the headboard or guardrail of the bed frame via a folding bracket and a lifting module; the smart screen has a built-in call program. The control box contains a main control module; the main control module activates the smart screen's call system via a first signal, which includes a user's abnormal physical signal, a voice call request signal, and a one-button call signal installed on the headboard. The smart screen integrates a camera, microphone, and touch display unit; The main control module is connected to the drive mechanism of the smart screen, the folding bracket, and the lifting module, respectively, to control the spatial position adjustment of the smart screen and the start and operation of the video call function.

2. A communication system applicable to the nursing bed with video chat function as described in claim 1, characterized in that, The system includes: Call response module: used to receive the first signal and initiate the call program on the smart screen; Face tracking module: When the call is started, it drives the folding bracket and lifting module to control the smart screen to keep in sync with the user's face based on the user's real-time facial position; Voice control module: Used to set the sound reception range of the smart screen and the user's voiceprint according to the synchronization trajectory, and to synchronously collect the voice of the call from the location of the user's face.

3. A communication system as described in claim 2, characterized in that, The receiving of the first signal further includes: When the first signal includes both a user's physical abnormality signal and a voice call request signal, the response priority of the physical abnormality signal is higher than that of the voice call request signal, and a first call instruction is generated. The first call instruction is used to initiate a video call with the nurses' station and delay the processing of voice call requests until the abnormal signal is cleared.

4. A communication system as described in claim 3, characterized in that, The generation of the first call instruction also includes: The pressure change rate of the back pressure sensor array, the heart rate value of the heart rate sensor, and the pressure drop amplitude of the bed exit detection sensor were collected. A signal is considered a genuine anomaly when at least two sensor data points meet a threshold condition. If only one condition is met or the data fluctuation duration does not reach the preset time, it is determined to be a false trigger, a suspected abnormal log is generated and stored locally, and when the number of false triggers reaches the preset number within the specified time, the smart screen automatically pops up a sensor calibration prompt and sends a calibration request to the administrator terminal.

5. A communication system as described in claim 2, characterized in that, When the call program is started, it also includes: Detect real-time light intensity and determine whether the real-time light intensity is lower than the preset light intensity value; Among them, when the real-time light intensity is lower than the preset light intensity value, the built-in infrared fill light of the smart screen frame is automatically turned on, and the facial feature point extraction algorithm is switched to infrared thermal imaging and contour priority mode. The three-dimensional facial contour data is obtained through the TOF depth sensor and matched with skeletal feature points to reduce the feature point matching threshold. When the area covered by facial occlusion exceeds a preset ratio, shoulder and face correlation tracking is activated. The face position is predicted by the preset coordinate mapping relationship between the center point of the shoulder contour and the center point of the face, so that the tracking interruption time is controlled within a preset time range.

6. A communication system as described in claim 2, characterized in that, When the call program is started, it also includes: The system determines whether the pressure drop detected by the hip pressure sensor exceeds a preset ratio and the duration reaches a preset value, and whether the infrared ranging sensor does not detect the presence of a human body within the preset range, or whether the user manually triggers the privacy mode icon through the smart screen touch unit and passes the verification. Disconnect the power supply to the smart screen's camera, control the folding bracket to rotate the smart screen to face the foot of the bed, switch the microphone array to silent monitoring mode, and when the user returns to the bed, the system needs to be reactivated via voice command or one-click call button.

7. A communication system as described in claim 2, characterized in that, The drive folding bracket and lifting module control the smart screen to move in sync with the user's face, and also include: The user priority sorting is preset and stored in the facial feature template library. When the facial tracking module detects both the user and the medical staff at the same time, if the facial feature matching degree of the medical staff reaches the preset value and continues to appear in the field of view of the smart screen for a preset time, the dual-target tracking mode is activated. When the dual-target tracking mode is executed, the smart screen displays a main screen and a picture-in-picture. The folding bracket and lifting module respond to changes in the user's position first. Changes in the position of medical staff only trigger horizontal rotation fine-tuning. When medical staff leave or the facial feature matching degree is lower than the preset value, the single-target tracking mode is restored within a specified time and the picture-in-picture is automatically turned off.

8. A communication system as described in claim 2, characterized in that, The synchronous acquisition of call audio based on the user's facial location includes: The sound source azimuth and distance are calculated by a time delay estimation algorithm and a sound source heat map is generated. When the sound source corresponding to the user's voiceprint is detected to be within the preset distance and azimuth range, the first-level beamforming is activated. If other interfering sound sources are present, gain attenuation is applied to the direction of the interfering sound sources. When the signal-to-noise ratio of the user's voiceprint signal is lower than the preset value, secondary enhancement is triggered. The gain of the weak signal area is increased by dynamic range compression and howling suppression is started. Howling suppression is stopped when the speech clarity meets the preset standard.

9. A communication system as described in claim 2, characterized in that, The method of synchronously collecting the call audio from the location of the user's face also includes: Predefined core voice commands and their corresponding dialect versions and speech rate models; When a user issues a voice command, the frequency band signal is extracted by FFT filtering and power frequency noise and high frequency noise are removed. The real-time voice features are compared with the pre-stored template by a dynamic template matching algorithm. Among them, the instruction is considered valid when the similarity reaches a preset value and the instruction duration is within the specified range; When the similarity is within a preset range, a list of command candidates is displayed on the smart screen for the user to confirm. When the number of consecutive command recognition failures reaches a preset value, the system automatically switches to touch-assisted mode and pops up a command shortcut button at the bottom of the screen.

10. A communication system as described in claim 2, characterized in that, The method of synchronously collecting the call audio from the location of the user's face also includes: Collect the user's voice recordings and convert them into text. Based on the call text, generate call captions on the smart screen.