Intelligent health physiotherapy entertainment interaction device, multi-source audio and video bidirectional collaborative interaction system and control method
By using multimodal sensors and adaptive weight fusion algorithms, bidirectional rhythm synchronization between the device and the presentation end is achieved, solving the problems of single sensing mode and low synchronization accuracy in existing technologies. It provides low latency, high precision rhythm synchronization and personalized adaptation, meeting the health needs of multi-source audio and video linkage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-06
- Publication Date
- 2026-05-05
AI Technical Summary
Existing physiotherapy equipment and fitness devices suffer from problems such as single sensing modes, low synchronization accuracy, lack of multi-source audio and video adaptation logic, single rhythm interaction dimension, and inability to achieve closed-loop interaction between devices and audio and video. They cannot meet users' immersive health needs for multi-sensor/sensorless adaptation, bidirectional rhythm synchronization, and multi-source audio and video linkage.
By employing multimodal sensors combined with an adaptive weight fusion algorithm, bidirectional rhythm synchronization between the device and the presentation end is achieved. Through a multimodal data acquisition unit, an action execution drive core module, a microcontroller, and a wireless communication module, it supports multi-dimensional action execution and audio-visual content scheduling. It also features intelligent resource scheduling and adaptive switching mechanisms, enabling bidirectional independent or collaborative interaction between the device and the presentation end.
It achieves low-latency, high-precision rhythm synchronization, supports multimodal sensing and sensorless full-solution adaptation, has continuous learning and personalized adaptation capabilities, meets the collaborative needs of refined health care, physiotherapy and entertainment scenarios, and provides a customized experience for each individual.
Smart Images

Figure CN121983268A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent health devices and human-computer interaction, and specifically relates to an intelligent health physiotherapy entertainment interactive device, a multi-source audio-video two-way collaborative interaction system, and a control method. Background Art
[0002] The existing interactive solutions for physiotherapy devices, fitness devices, physiological dysfunction trainers, postpartum rehabilitation trainers, dance training, and sports training are all one-way interaction architectures, which cannot meet the immersive health needs of users for multi-sensor / non-sensor adaptation, two-way rhythm synchronization, and multi-source audio-video linkage. Specifically, the following technical problems exist: the sensing scheme is single and the functions are fragmented, only supporting a single sensing mode, without integrating multi-sensing combinations of single acceleration, single gyroscope, single pressure sensor, acceleration + gyroscope + pressure sensor, and full scheme adaptation of non-sensor mechanical detection; and the sensing chip is only used for collecting the actions of the device itself, unable to drive the device to move or drive audio-video content; the rhythm interaction dimension is single, without two-way synchronization ability, and some technologies can only achieve simple sound effects driven by device actions, but without the ability to reverse-control device actions with audio-video, unable to form a device-audio-video closed-loop interaction; the types of audio-video-driven actions are limited, with strong sensor dependence, only supporting the basic vibration of the device driven by audio-video, unable to achieve regular synchronization of multi-dimensional actions such as stretching, rotation, slapping, dot vibration, and swinging; and it needs to rely on sensors to complete the linkage, and the audio-video cannot accurately drive the device actions in the non-sensor scenario, with insufficient technical versatility; the multi-source audio-video adaptation logic is missing and the robustness is insufficient. The existing AI soothing audio-video in the cloud interaction scheme has no rhythm binding with the device actions, and no rhythm scheduling mechanism for multi-source audio-video is established, resulting in problems such as strong cloud dependence, interruption of mode switching, and inability to adapt to private personalized health physiotherapy entertainment scenarios; the synchronization accuracy is low and the rhythm matching degree is insufficient. The existing technology is a simple trigger response, without an accurate mapping model for action-audio-video rhythm, with a rhythm synchronization delay > 200 ms and a matching error > 15%, unable to meet the rhythm coordination requirements of fine health physiotherapy entertainment scenarios.
[0003] In order to solve the deficiencies of the existing technology, people have carried out long-term explorations and proposed various solutions. For example, a Chinese patent document discloses a portable integrated intelligent interactive fitness system and method [201910647389.1], which includes: a portable integrated intelligent fitness apparatus, a data acquisition sensing module, a mobile terminal, an app program, an external audio-video device, and a wireless communication module.
[0004] The above solution solves the problem of single sensing mode to a certain extent, but there are still many deficiencies in this solution, such as low synchronization accuracy and lack of a scheduling mechanism for multi-source audio-video. Summary of the Invention
[0005] The purpose of this invention is to address the above-mentioned problems by providing a multimodal sensing intelligent health therapy and entertainment interactive device and a multi-source audio and video bidirectional collaborative interactive system that can realize the scheduling of multi-source audio and video.
[0006] Another objective of this invention is to address the aforementioned problems by providing a highly interactive intelligent health and wellness therapy entertainment device and a multi-source audio and video bidirectional collaborative interactive control method.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: an intelligent health and wellness entertainment interactive device and a multi-source audio and video two-way collaborative interactive system, including a device end, a presentation end connected to the device end, and a two-way rhythm synchronization engine integrated between the device end and the presentation end, the two-way rhythm synchronization engine being used to achieve the following interactions: Positive interaction: Transforms the motion characteristics of the device into control parameters to drive the presentation end to play rhythmically synchronized audio and video content; Reverse interaction: Converting the audio and video features of the presentation end into control parameters to drive the device end to perform rhythmic synchronization physical actions; Forward and reverse interactions can run independently or simultaneously.
[0008] In the aforementioned intelligent health and wellness entertainment interactive device and multi-source audio and video two-way collaborative interactive system, the device includes a multimodal data acquisition unit for collecting motion data, an action execution drive core module for performing multi-dimensional actions, a microcontroller, and a wireless communication module.
[0009] In the aforementioned intelligent health and wellness entertainment interactive device and multi-source audio-visual two-way collaborative interactive system, the multimodal data acquisition unit includes one or more of an accelerometer, gyroscope, pressure sensor, blood oxygen sensor, and heart rate sensor, and uses an adaptive weighted fusion algorithm to process the data. The fusion weights of acceleration, gyroscope, and pressure sensing data are dynamically allocated based on the motion state coefficient. In mobile scenarios, the weight of the gyroscope is increased; in fixed vibration scenarios, the weight of acceleration is increased; and in pressing scenarios, the weight of pressure sensing is increased.
[0010] In the aforementioned intelligent health and wellness entertainment interactive device and multi-source audio-visual bidirectional collaborative interactive system, the multimodal data acquisition unit is a sensorless mechanical detection solution, including: Hall effect sensors are used to detect changes in the displacement of moving parts. And / or a motor parameter detection module, used to collect the back electromotive force or phase current of the motor; By using displacement changes or motor parameters, the simulated motion acceleration, velocity, or rhythm characteristics are calculated through mathematical mapping relationships.
[0011] In the aforementioned intelligent health and wellness entertainment interactive device and multi-source audio and video bidirectional collaborative interactive system, the presentation end includes a multi-source audio and video scheduling module for scheduling audio and video content, a motion-audio-video mapping engine for realizing bidirectional feature conversion, and an audio and video perception and control conversion module for extracting features from audio and video.
[0012] In the aforementioned intelligent health and wellness entertainment interactive device and multi-source audio and video bidirectional collaborative interactive system, the multi-source audio and video scheduling module has an intelligent seamless switching function: When at least one of the following conditions is detected: network interruption, AI generation latency exceeding a threshold such as 500ms, action repetition exceeding a threshold such as 500ms, or device battery level below a threshold such as 20%, the scheduling source will be automatically switched from the current resource pool to a resource pool with lower latency or more locality, with a switching latency of ≤50ms.
[0013] A smart health and wellness therapy entertainment interactive device and a multi-source audio-visual two-way collaborative interactive control method, employing the aforementioned smart health and wellness therapy entertainment interactive device and multi-source audio-visual two-way collaborative interactive system, includes the following steps: S1: System initialization, connection establishment; S2: Receives the user's selected interaction mode command; S3: Execute the selected interaction mode: If it is in positive mode, then the positive interaction process will be executed; If it is in reverse mode, then the reverse interaction process will be executed; In bidirectional mode, forward and reverse interaction processes are executed concurrently.
[0014] In the aforementioned intelligent health and wellness entertainment interactive device and multi-source audio-visual two-way collaborative interactive control method, the control process for forward interaction includes: S41: Acquire raw motion data from the device. S42: Process the raw motion data and extract standardized motion rhythm features; S43: Transmit the motion rhythm features to the rendering end; S44: Based on action rhythm features and real-time scene information, schedule matching audio and video content from the AI-generated resource pool, pre-generated resource pool, or preset resource pool; S45: Based on the preset mapping rules, convert the motion rhythm features into audio and video control parameters and play the corresponding content; S46: Monitor and adjust synchronization errors.
[0015] In the aforementioned intelligent health and wellness entertainment interactive device and multi-source audio-visual two-way collaborative interactive control method, the control process for reverse interaction includes: S51: Load audio and video sources on the rendering end; S52: Extract visual motion features and / or audio rhythm features from audio and video sources through the machine vision submodule and / or sound recognition submodule; S53: Transform features into standardized device control commands; S54: Transmit device control commands to the device. S55: The device parses and executes the device control commands, driving the action execution module to complete the corresponding actions.
[0016] The aforementioned intelligent health and wellness entertainment interactive device and multi-source audio-visual bidirectional collaborative interactive control method also includes user preference learning and system optimization steps, specifically: Data collection dimensions: Real-time recording of core user behavior data in various scenarios, including selection of positive / reverse / bidirectional interaction modes; selection of audio and video style preferences, such as natural sound effects / white noise selection in sleep aid scenarios and level type preferences in gamified scenarios; mapping rule adjustment records, such as manually modifying the matching relationship between vibration frequency and audio and video rhythm; and also including action intensity tolerance range, training duration preferences, synchronization error correction operations, etc. Model Training and Optimization: Based on reinforcement learning algorithms, a user preference model is constructed, using collected behavioral data as training samples, and three core strategies are continuously optimized: Action-content matching strategy: accurately match user action characteristics with preferred audio and video types, such as identifying user preference for strong rhythmic actions corresponding to exciting sound effects and automatically strengthening this mapping relationship; Resource scheduling strategy: Prioritize scheduling resource pools frequently selected by users. For example, if users frequently use preset resources when the network is down, optimize the system to directly switch to the preset resource pool by default when the battery is low. Mapping parameters: dynamically adjust rhythm matching threshold, synchronization delay compensation parameters, etc., such as when the user manually increases the audio and video playback speed multiple times, automatically optimize the mapping coefficient of motion frequency-playback speed; Iterative update mechanism: Model iteration is triggered after every 10 valid interactions. The personalization accuracy gradually improves with the number of uses, eventually stabilizing at over 90%, achieving an adaptive experience that becomes more and more in line with user habits the more it is used.
[0017] Compared with existing technologies, the advantages of this invention are: 1. It pioneered a two-way closed-loop interactive architecture, breaking through the limitations of traditional one-way interaction, and realizing the two-way independent / cooperative operation of device-driven audio and video and audio and video-driven devices, forming a complete interactive link of acquisition-scheduling-mapping-control-feedback, solving the core pain point that existing technologies cannot form a closed-loop collaboration; 2. Covers all solutions for multimodal sensing and sensorless applications, supporting sensorless solutions such as single accelerometer, single gyroscope, pressure sensor, multi-sensor fusion, and Hall element / motor parameter detection. It can achieve accurate interaction without relying on a single sensor, and is suitable for all scenarios such as sleep aid, deep relaxation, professional physiotherapy assistance, postpartum recovery, functional disorder training, fitness, dance, children's movement correction, and personalized entertainment, significantly improving system robustness. 3. The core engine achieves industry-leading low-latency, high-precision rhythm synchronization, with forward interaction synchronization latency ≤100ms, reverse interaction response latency ≤80ms, and rhythm matching degree ≥90%, far exceeding the performance indicators of existing technologies and meeting the collaborative needs of refined health and wellness therapy scenarios. 4. The intelligent scheduling mechanism ensures the continuity of interaction in complex environments. When network interruption, AI generation delay > 500ms, action repetition ≥ 80%, or device battery < 20% is detected, the resource pool is automatically and seamlessly switched according to the path of AI generation → pre-generation → preset, with a switching delay ≤ 50ms, to avoid interaction interruption. 5. Possesses continuous learning and personalized adaptation capabilities. By capturing user behavior preferences through reinforcement learning algorithms, it dynamically optimizes action-content matching, resource scheduling, and mapping parameters, gradually improving the adaptation accuracy to over 90%, providing users with a customized experience tailored to each individual. 6. Expanding the boundaries of multi-dimensional motion and audio-visual linkage, the device can achieve regular synchronization of multi-dimensional motions such as vibration, extension, rotation, swinging, and point vibration. Audio and video support AI real-time generation, pre-generation processing, and local preset resource calling, which not only meets the needs of personalized creation, but also ensures the professionalism of precise physiotherapy rhythm, taking into account both entertainment and functionality. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the system structure of the present invention; Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation
[0019] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0020] like Figure 1-2 As shown, an intelligent health and wellness entertainment interactive device and a multi-source audio-visual two-way collaborative interactive system include a device end connected to a presentation end. A two-way rhythm synchronization engine is integrated between the device end and the presentation end. The hub of this two-way rhythm synchronization engine system includes a motion feature extraction unit, a motion mapping model unit, and an adaptive synchronization controller, specifically implementing the following interactions: Positive interaction: Transforms the motion characteristics of the device into control parameters to drive the presentation end to play rhythmically synchronized audio and video content; Reverse interaction: Converting the audio and video features of the presentation end into control parameters to drive the device end to perform rhythmic synchronization physical actions; Forward and reverse interactions can operate independently or simultaneously, ultimately forming a complete closed loop of acquisition, scheduling, mapping, reverse control, and feedback.
[0021] Specifically, the device includes a multimodal data acquisition unit for collecting motion data, an action execution driver core module for performing multidimensional actions, a microcontroller and wireless communication module, and a power supply module.
[0022] In detail, the multimodal data acquisition unit includes accelerometers and gyroscopes, as well as sensing elements such as pressure sensors and Hall elements. It can also use an adaptive weight fusion algorithm to process data: dynamically allocate the fusion weights of acceleration, gyroscope and pressure sensor data according to the motion state coefficients, increase the weight of gyroscopes in moving scenarios, increase the weight of acceleration in fixed vibration scenarios, and increase the weight of pressure in pressure scenarios.
[0023] The motion execution drive core module receives control commands and can achieve frequency adjustment of multiple levels of vibration / swaying / rotation, with an adjustment frequency of 0.1-5Hz, multi-angle attitude adjustment within 0-180 degrees, and precise actions such as multi-stroke extension and retraction control; the microcontroller can dynamically adjust the acquisition frequency within the range of 1Hz-200Hz and execute control commands to form a two-way linkage.
[0024] Furthermore, the multimodal data acquisition unit is a sensorless mechanical detection solution, including: a Hall element for detecting displacement changes of moving parts at any time; and / or a motor parameter detection module for acquiring the back electromotive force or phase current of the motor; and through displacement changes or motor parameters or other parameters that approximate the intensity of mechanical motion, such as voltage, power, and speed, the simulated motion acceleration, velocity, or rhythm characteristics are calculated by mathematical mapping relationships.
[0025] Furthermore, the presentation end includes a multi-source audio and video scheduling module for scheduling audio and video content, a motion-audio-video mapping engine for bidirectional feature conversion, and an audio and video perception and control conversion module for extracting features from audio and video. It also includes a data parsing module, a communication receiving / sending module, and a storage module. The storage module supports storage in multiple types of audio and video resource pools. The audio and video perception and control conversion module, as the core of the interaction, includes an audio and video feature extraction and parsing module for analyzing the rhythm, tempo, and intensity curves of audio and video, and an audio and video generation / scheduling module for generating or calling audio and video content that conforms to the rhythm.
[0026] In addition, the multi-source audio and video scheduling module has an intelligent seamless switching function: when at least one of the following conditions is detected, such as network interruption, AI generation latency exceeding the threshold, action repetition exceeding the threshold, or device battery level below the threshold, the scheduling source is automatically switched from the current resource pool to a resource pool with lower latency or more locality, with a switching latency of ≤50ms.
[0027] Example 1 This embodiment describes an intelligent multi-frequency relaxation device and its interaction method for sleep aid and deep relaxation. The device's multimodal data acquisition unit employs a sensorless mechanical detection scheme, specifically including a high-precision Hall element for detecting the linear displacement of the internal transmission mechanism with an accuracy of ±1mm; and a motor parameter detection circuit for acquiring the real-time phase current of the drive motor. The microcontroller calculates the extension / retraction stroke and speed using the Hall element signal, and inversely calculates the output torque and vibration intensity using the motor phase current signal. Through mathematical mapping, it calculates the simulated acceleration, velocity, and motion rhythm characteristics. The motion execution drive core module uses a precision micro-motor and transmission mechanism, enabling multi-level vibration, adjustable frequency within 0.1-5Hz, multi-angle oscillation within 0-60°, and short-stroke extension / retraction within 0-20mm. The microcontroller and wireless communication module use a low-power Bluetooth BLE chip for data processing, command parsing, and communication.
[0028] The presentation end is a mobile application containing a multi-source audio and video scheduling module. Its resource pool is divided into: an AI-generated resource pool, which connects to a cloud-based video generation model; a pre-generated resource pool, which stores various standardized audio and video clips with rhythmic tags, such as ocean waves, rain sounds, and forest sounds; and a pre-set resource pool, which stores local basic white noise and natural scenery animations.
[0029] The algorithm logic of the bidirectional rhythm synchronization engine is distributed between the device-side microcontroller and the presentation-side App, working collaboratively through a bidirectional communication link. During sensorless mechanical detection, the Hall element detects the relative position change Δd of the magnet. The microcontroller calculates the instantaneous velocity v=Δd / Δt based on the unit time Δt, and further calculates the acceleration a=Δv / Δt. The motor parameter detection module collects the phase current I, and based on the motor torque constant Kt, maps the torque τ=Kt*I into a simulated intensity signal of the action. Combined with the motor speed n obtained through back electromotive force or encoder, the rhythm characteristics of the action are jointly determined. Finally, the calculated velocity, acceleration, and intensity signals are normalized according to a preset range to form standardized JSON format feature data such as {freq:1.0,intensity:0.8}.
[0030] In the positive interaction process, the user selects the positive sleep aid mode and manually sets the device to 1Hz low-frequency vibration combined with a 15° gentle sway. The device's multimodal data acquisition unit is activated, and the Hall element and motor detection circuit begin to work. The microcontroller calculates standardized motion rhythm features in real time using a preset mathematical model based on the collected displacement changes and motor parameters. This feature is transmitted to the presentation app via BLE. The presentation app's multi-source audio and video scheduling module receives the feature and, when the network is good, prioritizes calling the AI-generated resource pool. The motion-audio-video mapping engine converts the motion features into a structured prompt: 1Hz low-frequency vibration and swaying, intensity 0.8, generating a matching sleep aid wave scene video and surround sound. The cloud model is then generated and returned as an audio and video stream.
[0031] In this positive interaction, the device-side microcontroller performs mean filtering on the calculated acceleration sequence A=[a1,a2,…an], where n is the number of sampling points. A sliding window mean filtering method is used, where N=5~10 and is dynamically adjusted according to the sampling frequency. For 10Hz sampling, N=5, and the filtered data a'i=(a_{i-4}+a_{i-3}+…+a_i) / 5. Boundary points are padded with default zeros or other fixed values to eliminate noise and extract stable vibration frequencies. Before transmitting standardized features, the device can use Huffman coding to compress the data, achieving a coding efficiency ≥85%, to ensure a transmission latency ≤100ms under Bluetooth BLE bandwidth.
[0032] Furthermore, when the rendering device detects a network interruption or an AI generation latency greater than 500ms, the scene switching control unit immediately takes action. The switching path is not a simple jump, but follows a degradation path of AI generation → pre-generation → preset, ensuring that the latency of each switch is ≤50ms. In low-power scenarios with reverse interaction, if switching to the pre-generated resource pool, the system uses a cosine similarity matching algorithm to match the BPM features of the current audio with the tags of segments in the resource library, quickly calling up the most suitable white noise segment.
[0033] In the reverse-interaction sleep aid mode, users set a target soothing rhythm of 1-2Hz in the app and select to activate the professional sleep aid mode. Based on this setting, the system plays corresponding sleep aid audio and video content, such as content from local storage, the cloud, or generated content. The sound recognition submodule of the audio-visual perception and control conversion module on the presentation end is activated, performing real-time spectrum analysis and beat tracking to extract the core rhythmic features of the audio. This feature is sent to the motion-audio-video mapping engine, which converts it into device control parameters according to mapping rules. These control parameters are then sent to the device via BLE. The device's microcontroller parses the instructions and drives the core module to begin gentle vibrations at a frequency of 1-2Hz, executing an intelligent degradation switching mechanism: when the app detects trigger conditions such as device battery level below 20%, network interruption, or cloud processing latency >500ms, the scene switching control unit is immediately activated. The switching follows a degradation path of AI-generated / streaming media, locally pre-generated resource pool, and device-preset basic resources. During this process, the system matches the spectral characteristics of the current audio, such as the main frequency distribution and energy envelope, and, in conjunction with the feature tags of audio segments in the local resource library, quickly calls up the local audio segment with the closest sonic quality for seamless transition, such as basic white noise, ensuring that the overall audio switching latency is ≤50ms. After switching, the system can simultaneously shut down unnecessary communications such as high-power communication modules running to acquire streaming media or perform cloud processing, thereby significantly reducing the overall power consumption of the device and ensuring that the sleep aid process continues uninterrupted.
[0034] This embodiment reduces hardware costs and failure rates through a sensorless solution. Combining forward and reverse interaction modes, the system fully supports two independent operating modes: device-driven content and content-driven device. The integrated intelligent resource scheduling and degradation switching mechanism effectively improves the system's robustness under complex operating conditions and the consistency of the user experience.
[0035] Example 2 This embodiment describes a professional auxiliary device for health training or physiotherapy teaching. The device's multimodal data acquisition unit employs a fusion scheme of accelerometer and gyroscope sensors. It integrates a six-axis IMU (three-axis accelerometer + three-axis gyroscope) with a sampling frequency of 30Hz. The microcontroller runs an adaptive weighted fusion algorithm, calculating the motion state coefficient S. When the device is actively oscillated, the gyroscope weight is increased to 80%, and when the handheld device is stationary and only the motor vibrates, the accelerometer weight is increased to 70%. This dynamic allocation of fusion weights ensures an attitude angle detection error of ≤±3°. The motion execution drive core module includes an eccentric wheel motor for vibration and a servo motor for rotation, enabling precise angle positioning from 0-180° and variable frequency vibration from 0.1-5Hz.
[0036] The presentation platform is a tablet computer software for physiotherapy teaching. The machine vision submodule of the audio-visual perception and control conversion module has been trained to recognize key joint angles and movement rhythms in the physiotherapist's demonstration videos. The pre-generated resource pool of the multi-source audio-visual scheduling module stores encouraging animation and sound effect clips that are linked to states such as whether the movement has reached the target or needs adjustment.
[0037] The raw acceleration data is mean-filtered to obtain a_x', a_y', a_z'; the raw gyroscope data is Kalman-filtered to eliminate drift, yielding angular velocities ω_x, ω_y, ω_z, which are then integrated to obtain attitude angles θ_g, φ_g, ψ_g. An adaptive weight fusion algorithm, as the core algorithm, calculates the motion state coefficient S in real time: S = max(Δa_x', Δa_y', Δa_z') / max(Δω_x, Δω_y, Δω_z). When the device is actively swayed (i.e., in a moving scenario where S < 1), the system automatically assigns a gyroscope weight of ω_g = 80% and an acceleration weight of ω_a = 20%. The fused attitude angle calculation formula is: θ = ω_a * θ_a + ω_g * θ_g, where θ_a is the angle calculated by the accelerometer, taking the roll angle as an example. Through this fusion, the attitude detection error is ≤ ±3°, providing a data foundation for precise guidance.
[0038] During reverse interaction, the user selects the reverse teaching mode, which plays a standard physiotherapy operation video. The machine vision submodule on the presentation side analyzes the video, extracting key features of the demonstrated movements, such as the wrist should externally rotate 45° at the 2-second mark and maintain a uniform up-and-down rhythm of 1Hz throughout. These features are then converted into device control commands by the motion-audio-video mapping engine. The commands are sent to the device, which begins vibrating at a 1Hz frequency, guiding the user's wrist to rotate to the 45° position. This process eliminates the need for the user to manually find the angle by holding the device; the device actively executes the standard movements to guide the user and achieve motion correction.
[0039] During positive interaction, while the device guides the user's actions, the device's multimodal data acquisition unit continuously works, using a fusion algorithm to accurately collect the user's actual hand movement data. It then extracts the actual angle, frequency, and other rhythmic features of the user's actions and uploads them to the presentation end in real time. The presentation end compares the user's action features with the standard features of the instructional video in real time, such as cosine similarity matching. When the system calculates that the synchronization degree between the user's actions and the standard actions is ≥80%, the multi-source audio-visual scheduling module immediately schedules a soothing green plant growth animation and encouraging sound effects from the pre-generated resource pool to play, providing the user with positive visual and auditory feedback. If the synchronization degree is below the threshold, prompting content can be scheduled. The bidirectional rhythmic synchronization engine ensures the rhythmic correlation between the guiding actions, the user's actual actions, and the feedback audio-visual content.
[0040] This embodiment provides a reliable data foundation for professional scenarios through a high-precision multi-sensor fusion scheme. Reverse interaction enables precise machine-led instruction, while forward interaction provides real-time motivational feedback. The combination of the two forms a complete two-way closed loop of teaching-following-assessment-motivation, significantly improving the professionalism, accuracy, and enjoyment of physiotherapy training, and achieving a movement correction accuracy rate of ≥85%.
[0041] Example 3 This embodiment presents a highly customizable, privacy-focused entertainment and relaxation device. It emphasizes personalized generation and privacy, with the device's multimodal data acquisition unit employing a single accelerometer solution, resulting in a simple structure and low power consumption. Furthermore, a triaxial accelerometer collects device motion data on demand at a frequency of 5-50Hz. The motion execution drive core module supports combinations of various complex motion trajectories.
[0042] The presentation end is dedicated client software on a personal computer, capable of running the core module offline. The AI generation resource pool of the multi-source audio and video scheduling module integrates a locally deployed lightweight AIGC model, which can run even without internet access. The motion-audio-video mapping engine contains custom-level mapping rule units that can incorporate complex motion trajectory features, such as trajectory curvature and acceleration change rate, into the generated cues.
[0043] During positive interaction, users select positive creation mode and operate the device with specific rhythms and trajectories, such as quickly drawing circles followed by slow, wavy movements. The device's accelerometer collects data, which, after mean filtering and feature extraction, yields complex motion rhythm features including frequency, intensity, and trajectory characteristics. These features are transmitted to the rendering end, triggering the custom-level mapping rule unit of the motion-audio-visual mapping engine. This transforms features such as arc-shaped, slow-swinging trajectories and quick circles into creative prompts: a meandering stream animation followed by fireworks effects, accompanied by corresponding rhythmic electronic sound effects.
[0044] When engaging in reverse interaction, users select the reverse resonance mode and upload a private audio or video clip containing no publicly available rhythmic information. The audio-visual perception and control conversion module on the display end analyzes the uploaded content, including sound recognition analysis of audio BPM and spectrum, and machine vision analysis of screen flickering or motion rhythm, extracting its inherent rhythmic features. These features are mapped to device control parameters, driving the device to rhythmically match the content with perfect precision. Users then no longer need to manually adjust the device to match the content's rhythm; the system automatically achieves deep synchronization, providing a highly private and convenient immersive experience.
[0045] During the aforementioned two-way interaction, the system records user preferences. For example, if a user always likes to map a certain swing to a certain visual style, the system will optimize the future action-content matching strategy through the background algorithm, so that the accuracy of personalized adaptation will gradually improve with the use time.
[0046] This embodiment utilizes customized mapping and AI generation capabilities to transform user actions into unique multimedia content, meeting high-end personalized needs. Simultaneously, reverse interaction allows any private media content to be dynamically controlled by the device. Both bidirectional interactions can run independently locally, maximizing privacy and security, and achieving a personalized experience where content and device deeply resonate in private settings.
[0047] Example 4 This embodiment describes a professional auxiliary device for postpartum recovery exercises: a smart Kegel ball. The device includes a multimodal data acquisition unit, which can incorporate a pressure sensor with a measurement range of 0-20N and an accuracy of ±0.1N, used to collect pressure data when the user squeezes the ball. It can also integrate a three-axis or six-axis gyroscope sensor with a sampling frequency of 5-50Hz to help capture changes in the ball's posture.
[0048] The motion execution drive core module adopts a miniature silent motor and eccentric wheel structure, supporting multiple vibration levels from 0.2-3Hz and preset waveform motion, such as gradually increasing and decreasing sine waves and pulsed square waves.
[0049] The control module includes a low-power MCU responsible for data processing and instruction parsing; and a wireless communication module, Bluetooth BLE 5.2, to achieve low-latency communication with the presentation device, with a latency of ≤80ms.
[0050] The presentation platform is an app on a personal mobile phone, capable of running its core modules offline. It includes: 1. The video scheduling module's resource pool contains three types of Kegel course resources: the AI-generated resource pool generates personalized course videos based on user stress data and training progress, such as animations of strengthening training progress for the weak side of the pelvic floor muscles; the pre-generated resource pool stores standardized Kegel training videos, such as 4 sets of 10-repetition contraction-relaxation cycles, as well as guiding audio, such as voice prompts for contracting for 3 seconds and relaxing for 5 seconds; and the pre-set resource pool locally stores low-latency training beat audio and video, milestone animations, etc.
[0051] 2. The audio perception and control conversion module extracts the rhythmic features of the course audio and video, such as the training beat BPM=60 and the contraction phase duration of 3 seconds, and converts them into equipment control parameters.
[0052] 3. The data recording module records the user's stress peak, duration, and action consistency for each training session, and optimizes the course matching strategy through reinforcement learning.
[0053] When engaging in positive interaction with sensors, users drive the course progress by squeezing actions; that is, users apply pressure to drive the Kegel ball, and the ball drives the audio and video content of the APP. For example, when a user activates the self-training mode, they insert a Kegel ball into their body and begin squeezing. A pressure sensor collects real-time squeezing pressure data F(t), and a gyroscope sensor, if present, assists in determining the completeness of the squeezing action, such as whether it is accompanied by ball displacement, hip lifting, or body rotation. The device's MCU processes the pressure data, using mean filtering with a window size of N=5 to remove interference, and extracts the pressure peak F_max, contraction duration T_s (the duration when pressure ≥ 1N), and relaxation duration T_r, generating standardized motion features, such as F_max: 8N, T_s: 3.2s, T_r: 4.8s. After the features are transmitted to the app via Bluetooth, the multi-source audio and video scheduling module matches audio and video content based on the features: during pressure changes, AI can be invoked to generate a real-time scaling animation of the Kegel ball inside the body matching the user's motion features, achieving a WYSIWYG effect; simultaneously, it can also generate progress videos representing training effectiveness, showcasing training progress and completion. All audio and video preferences, including characters, voices, and styles, can be generated according to the user's preset timing or pre-generated for use. If the user selects a standard training course, two-way interaction is also possible. That is, the user's motion characteristics are compared with the motion characteristics retrieved from the standard course, and audio and video prompts are generated to adjust the synchronization error, providing feedback to the user. If the deviation between the motion frequency and the video rhythm is >10%, the APP will accelerate the contraction speed through voice prompts, generate an animation of the difference between the current pressure curve and the target pressure curve in the video, and at the same time, the device's motor will start a slight vibration to remind and assist in calibrating the rhythm.
[0054] When engaging in positive interaction with sensors, the pressure value generated by the user's squeezing action is directly converted into device control commands, such as when the device itself has a motor. Specifically, during the contraction phase, progressively stronger vibrations are initiated, with the intensity increasing from 0 to 40, and a maximum intensity of 100, corresponding to F_max=20N, to guide the user to squeeze synchronously; during the relaxation phase, progressively weaker vibrations are initiated, with the intensity decreasing from 40 to 0, to prompt the user to relax.
[0055] When performing positive interaction without sensors, users can operate the device via the control keys or infrared remote control. For example, they can select a preset waveform mode and switch to pulse training. In this mode, the motor moves in a square wave pattern: 1 second of vibration, 2 seconds of stop, with a frequency of 0.33Hz. The device's motor moves in this preset square wave pattern, and the characteristic data of this square wave motion is transmitted to the APP. The APP's multi-source audio and video scheduling module calls the beat audio and video from the preset resource pool to synchronize the audio and video with the rhythm of the motor's motion. A beeping sound is played when the motor vibrates and silences when it stops. The APP interface simultaneously displays the motion waveform and training duration to help users master the rhythm. Alternatively, if computing power and network allow, AI can be used to generate rhythmically synchronized videos to enhance the fun, alleviate discomfort, and promote the release of endorphins.
[0056] During reverse interaction, the system guides user training through pre-set courses, where the courses drive the ball, and the ball drives the audio and video. For example, a user selects a 42-day postpartum pelvic floor muscle repair course from a pre-generated resource pool or imports it from the internet. This course includes voice guidance and animated demonstrations such as contracting for 3 seconds and relaxing for 5 seconds. The audio and video perception and control conversion module retrieves course features from the resource pool or extracts them from the audio and video: contraction phase trigger signal T=3s, relaxation phase trigger signal T=5s, and medium intensity level, corresponding to a pressure target of 8N. These features are converted into device control commands, such as {mode: sine wave vibration, max_intensity: 40, contraction_time: 3s, relax_time: 5s}, and transmitted to the Kegel ball. The device's motor responds to the commands: during the contraction phase, it initiates gradually increasing vibration, with the intensity rising from 0 to a peak of 40 within 3 seconds according to a sine wave, guiding the user to squeeze synchronously; during the relaxation phase, it initiates gradually decreasing vibration, with the intensity dropping from the peak of 40 to 0 within 5 seconds according to a sine wave, prompting the user to relax. Simultaneously, if the user selects two-way interaction, the pressure sensor collects the squeezing pressure data F(t) in real time, and the gyroscope sensor, if available, assists in judging the completeness of the squeezing action, such as whether it is accompanied by ball displacement, hip lifting, or body rotation, generating standardized movement characteristics. These characteristics are compared with the movement characteristics retrieved from the standard course, and audio and video prompts for synchronization error adjustment are generated and fed back to the user. If the peak pressure of the action does not reach the target (<8N), the APP provides a voice prompt indicating insufficient compression exercise, urging the user to squeeze harder. The video generates an animation of the difference between the current pressure curve and the target pressure curve, and the device's motor starts slight vibration to remind and assist in calibrating the rhythm.
[0057] The system performs two-way closed-loop optimization: the APP records the user's stress achievement rate and movement synchronization degree for each training session, and adjusts the intensity target of the next course through reinforcement learning. For example, if the achievement rate is ≥90%, the stress target will be increased to 9N.
[0058] This embodiment provides a systematic approach to postpartum recovery by utilizing pressure sensors and / or three-axis / six-axis gyroscopes. The pressure sensors accurately capture the contraction strength of the pelvic floor muscles, addressing the limitation of traditional Kegel balls in quantifying training effectiveness. Two-way interaction enhances both scientific rigor and engagement: positive feedback allows users to visually observe training progress and motivational animations, while reverse guidance reduces operational difficulty and increases enjoyment. This two-way interaction effectively assesses training results and promotes scientific rehabilitation, making it suitable for postpartum users who are physically weak and lack training experience. Furthermore, multiple modes adapt to different stages. For example, a preset waveform mode is suitable for initial self-adaptation, a course-driven mode is suitable for intermediate systematic training, and an AI-generated mode is suitable for personalized reinforcement, covering the entire postpartum recovery cycle.
[0059] Example 5 This embodiment is an interactive system for a health ring and AI glasses for gamified fitness. The smart health ring on the device side includes a multi-mode multimodal data acquisition unit, which integrates a high-precision blood oxygen sensor with a measurement range of 70%-100% and an accuracy of ±1%; a heart rate sensor with a sampling frequency of 50Hz, a measurement range of 30-240 beats / minute, and an accuracy of ±2 beats / minute; and a triaxial accelerometer with a sampling frequency of 20Hz and a measurement range of ±8g.
[0060] The control and communication module includes a low-power MCU responsible for data fusion processing and instruction parsing; and a Bluetooth BLE 5.3 module to achieve low-latency data transmission with the AI glasses, with a latency of ≤50ms.
[0061] The auxiliary module includes a built-in rechargeable battery with a battery life of ≥24 hours and supports magnetic fast charging; as well as a waterproof and dustproof structure with an IP67 rating, suitable for sports scenarios such as running and fitness.
[0062] The display device is an AI smart glasses, whose core display unit uses a Micro OLED retina microdisplay with a resolution of 1920×1080 and a refresh rate of 60Hz. It supports semi-transparent AR display and does not obstruct the field of vision during movement.
[0063] The multi-source audio and video scheduling module's resource pool includes three types of gamified content: an AI-generated resource pool that integrates lightweight text-based images and game scene generation models, capable of generating level scenes, monster images, and reward animations in real time based on heart rate and movement posture; a pre-generated resource pool that stores standardized sports level templates, such as four types of levels: Novice Village → Forest Secret Realm → Volcano Challenge → Peak of the Summit, each containing three mini-tasks, as well as indicator animation components, such as heart rate waveforms, blood oxygen progress bars, and step counters; a pre-set resource pool that locally stores low-latency game sound effects, such as level completion prompts and energy replenishment sound effects; and basic UI components, such as health bars and level progress bars, providing fallback content when the network is disconnected or power consumption is low.
[0064] The audio-visual perception and control conversion module receives physiological indicators and motion data transmitted by the ring in real time and converts them into game control parameters; the built-in bone conduction speaker plays game sound effects and voice prompts in sync.
[0065] During positive interactions, users drive the ring through human movement or physiological indicators, which in turn drives the AI glasses to present a gamified experience, i.e., indicators drive level progress.
[0066] For example, a user wearing a health ring and AI glasses starts a challenge-based exercise mode, selecting an exercise type such as running or indoor fitness, with the initial level being the default "beginner village." The device's data acquisition unit activates: a blood oxygen sensor monitors real-time blood oxygen saturation (SpO2), a heart rate sensor collects heart rate (HR), and an accelerometer captures movement posture, such as cadence, arm swing amplitude, and exercise intensity. Data processing and feature extraction are then performed: the MCU fuses the raw data, using a sliding window mean filter with a window size of N=10 to remove motion interference, calculating the heart rate variability (HRV), average cadence (S_f) in running scenarios, and movement repetition frequency (M_f) in fitness scenarios, generating standardized feature data, such as {HR: 120, SpO2: 98%, S_f: 170 steps / minute, intensity: 0.7, posture: standard}. Feature transmission and gamification mapping are then performed: the feature data is transmitted to the AI glasses via Bluetooth, and the multi-source audio and video scheduling module, based on the current level status, calls upon the corresponding resource pool. For example, in a running scenario: when the cadence is ≥160 steps / minute, the AI generates a resource pool to create game visuals such as character sprinting and wind speed effects, and simultaneously displays a heart rate waveform animation; in a fitness scenario: when the repetition frequency of the movement reaches the target, such as 30 squats per minute and blood oxygen ≥95%, the pre-generated resource pool triggers a level completion animation, the AI generates an exclusive level-passing badge animation, unlocks the next level, and the bone conduction speaker plays a congratulatory sound indicating that the level has been passed and energy replenishment has been obtained; if the heart rate remains higher than the target range, such as the target heart rate of 110-130 beats / minute for more than 3 minutes, a low battery warning and prompt sound are triggered, and slow-motion game visuals, such as Tai Chi, are generated, and the background music rhythm is adjusted simultaneously to guide the user to adjust the exercise intensity from an auditory, visual, and subconscious level.
[0067] During reverse interaction, motion adjustments are driven by game tasks. Specifically, the AI glasses' game tasks drive the ring, which then provides motion guidance. For example, when a user enters a forest adventure level, the AI glasses display the task: maintain a cadence of 220-250 steps / minute and a heart rate below 130 beats / minute for 3 minutes, and defeat 3 speed monsters ahead. The audio-visual perception and control conversion module extracts the level task characteristics: target cadence of 235±15 steps / minute, target heart rate of 125 beats / minute, and task duration of 3 minutes, and converts these into device control parameters and guidance commands. Command transmission and motion guidance are then performed: the AI glasses send parameters to the ring via Bluetooth. The ring's MCU parses the parameters and uses an accelerometer to compare the user's current cadence with the target value in real time. If the cadence is below 220, the ring's built-in micro-vibration motor activates a 1Hz pulse vibration to remind the user to accelerate; if the cadence is above 250, the vibration frequency increases to 2Hz to remind the user to decelerate. Real-time feedback and level adjustments are also provided: the AI glasses continuously receive heart rate and cadence data transmitted from the ring, and the game screen adjusts synchronously. When the user's step frequency reaches the target, the character automatically launches an attack skill to defeat the monster; when the heart rate stabilizes within the target range, the character gains a shield buff; if the user completes the task of defeating 3 monsters within 3 minutes, the AI generates a resource pool, generates a level-clearing animation in real time, unlocks the Forest Secret Realm, reveals a treasure chest reward, and unlocks the next level, the Volcano Challenge; if the task is not completed, the game screen displays "Task Failure," and the user can choose to retry or reduce the difficulty.
[0068] The system performs two-way closed-loop optimization: it records the user's completion rate of each exercise, preferred exercise type such as running / fitness, heart rate comfort zone, endurance intensity, and the relationship between heart rate and running speed. Through reinforcement learning, it optimizes the system to achieve personalized adaptation between the user and the level. The AI can also customize exclusive level scenarios based on the user's body type and athletic ability. For example, if the user is good at long-distance running, an endurance challenge level will be generated; if the user is good at explosive sports, a sprint challenge level will be generated.
[0069] This embodiment utilizes a smart health ring and AI glasses solution to create a two-way interactive gamified fitness or exercise scenario. It offers several key features: Visualized exercise data: Transforming monotonous indicators like blood oxygen, heart rate, and cadence into dynamic animations, allowing users to intuitively view their status in real time and avoid aimless exercise; Gamified incentives to improve endurance: Breaking down long workouts into shorter goals through level progression, task challenges, and instant rewards reduces boredom and increases commitment; Two-way interaction to adapt to personalized needs: Supporting both physiological indicators driving game progress and game tasks guiding exercise adjustments, balancing autonomy and scientific rigor, and adapting to users of different fitness levels; Strong scenario adaptability: AR display does not obstruct the view, and the lightweight ring design does not interfere with exercise, covering various sports scenarios such as running, indoor fitness, outdoor hiking, and health and wellness.
[0070] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
[0071] Although this document uses terms such as "device side" and "presentation side" frequently, the possibility of using other terms is not excluded. These terms are used merely for the convenience of describing and explaining the essence of this invention; interpreting them as any additional limitation would contradict the spirit of this invention.
Claims
1. A smart health and wellness therapy entertainment interactive device and a multi-source audio and video two-way collaborative interactive system, comprising a device end connected to a presentation end, characterized in that, A bidirectional rhythm synchronization engine is integrated between the device and the presentation end, and the bidirectional rhythm synchronization engine is used to achieve the following interaction: Positive interaction: Converting the motion characteristics of the device into control parameters to drive the presentation terminal to play rhythmically synchronized audio and video content; Reverse interaction: Converting the audio and video features of the presentation end into control parameters to drive the device end to perform rhythm synchronization physical actions; The positive and negative interactions can operate independently or simultaneously.
2. The intelligent health and wellness therapy entertainment interactive device and multi-source audio-visual two-way collaborative interactive system according to claim 1, characterized in that, The device includes a microcontroller, a wireless communication module, a multimodal data acquisition unit for collecting motion data, and an action execution driver core module for performing multidimensional actions.
3. The intelligent health and wellness therapy entertainment interactive device and multi-source audio-visual two-way collaborative interactive system according to claim 2, characterized in that, The multimodal data acquisition unit includes one or more combinations of an accelerometer, a gyroscope, a pressure sensor, a blood oxygen sensor, and a heart rate sensor.
4. The intelligent health and wellness entertainment interactive device and multi-source audio-visual two-way collaborative interactive system according to claim 3, characterized in that, The multimodal data acquisition unit uses an adaptive weighted fusion algorithm to process motion data. The fusion weights of acceleration, gyroscope, and pressure sensor data are dynamically allocated based on the motion state coefficient. In mobile scenarios, the weight of the gyroscope is increased; in fixed vibration scenarios, the weight of acceleration is increased; and in pressing scenarios, the weight of pressure sensitivity is increased.
5. The intelligent health and wellness therapy entertainment interactive device and multi-source audio-visual two-way collaborative interactive system according to claim 2, characterized in that, The multimodal data acquisition unit is a sensorless mechanical detection solution, including: Hall effect sensors are used to detect displacement changes of moving parts at any given moment. And / or a motor parameter detection module, used to collect the back electromotive force or phase current of the motor or other parameters that approximate the intensity of mechanical motion; The simulated motion acceleration, velocity, or rhythm characteristics are calculated using the displacement changes or motor parameters through mathematical mapping relationships.
6. The intelligent health and wellness therapy entertainment interactive device and multi-source audio-visual two-way collaborative interactive system according to claim 1, characterized in that, The presentation end includes a multi-source audio and video scheduling module for scheduling audio and video content, a motion-audio-video mapping engine for realizing bidirectional feature conversion, and an audio and video perception and control conversion module for extracting features from audio and video.
7. The intelligent health and wellness entertainment interactive device and multi-source audio-visual two-way collaborative interactive system according to claim 6, characterized in that, The multi-source audio and video scheduling module has an intelligent seamless switching function: When at least one of the following conditions is detected: network interruption, AI generation latency exceeding a threshold, action repetition exceeding a threshold, or device battery level below a threshold, the scheduling source will be automatically switched from the current resource pool to a resource pool with lower latency or more locality.
8. A smart health and wellness therapy entertainment interactive device and a multi-source audio-visual bidirectional collaborative interactive control method, comprising the smart health and wellness therapy entertainment interactive device and the multi-source audio-visual bidirectional collaborative interactive system described in any one of claims 1-7, characterized in that, Includes the following steps: S1: System initialization, connection establishment; S2: Receives the user's selected interaction mode command; S3: Execute the selected interaction mode: If it is in positive mode, then the positive interaction process will be executed; If it is in reverse mode, then the reverse interaction process will be executed; In bidirectional mode, forward and reverse interaction processes are executed concurrently.
9. The intelligent health and wellness therapy entertainment interactive device and multi-source audio and video bidirectional collaborative interactive control method according to claim 8, characterized in that, The control flow for the positive interaction includes: S41: Acquire raw motion data from the device. S42: Process the raw motion data and extract standardized motion rhythm features; S43: Transmit the motion rhythm features to the rendering end; S44: Based on action rhythm features and real-time scene information, schedule matching audio and video content from the AI-generated resource pool, pre-generated resource pool, or preset resource pool; S45: Based on the preset mapping rules, convert the motion rhythm features into audio and video control parameters and play the corresponding content; S46: Monitor and adjust synchronization errors.
10. The intelligent health and wellness entertainment interactive device and multi-source audio-visual bidirectional collaborative interactive control method according to claim 8, characterized in that, The control flow of the reverse interaction includes: S51: Load audio and video sources on the rendering end; S52: Extract visual motion features and / or audio rhythm features from audio and video sources through the machine vision submodule and / or sound recognition submodule; S53: Transform features into standardized device control commands; S54: Transmit device control commands to the device. S55: The device parses and executes the device control commands, driving the action execution module to complete the corresponding actions.
Citation Information
Patent Citations
Portable comprehensive intelligent interactive fitness system and method
CN110957018A