An automated tremor detection process control system and method

CN122531610APending Publication Date: 2026-08-07BEIJING JINGMIAN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING JINGMIAN TECHNOLOGY CO LTD
Filing Date
2026-05-14
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005](一)发明目的:为解决上述现有技术中存在的问题,本发明的目的是提供一种自动化震颤检测流程控制系统及方法,解决现有技术中震颤检测流程自动化程度不足、硬件缺乏自适应调节能力、遮挡处理机制不完善等问题,实现从检前准备到检后收尾的全流程无人干预自动化运行,提升检测效率和数据采集精准性

Benefits of technology

[0047](三)有益效果:本发明提供一种自动化震颤检测流程控制系统及方法,首先基于有限状态机的流程控制模块实现从检前准备到检后收尾的全流程自动化运行,大幅压缩单患者检测时长,减少医护人员人工操作。其次通过身高识别模块与设备控制模块的联动,根据患者身高自动调节动作捕捉相机高度角度及座椅位置,实现站姿与坐姿检测场景的无缝切换,无采集盲区;再者通过遮挡监测模块与引导提示模块的闭环联动,实时识别关键部位遮挡并自动暂停计时、标记无效数据、触发调整提示,遮挡解除后仅重测被遮挡片段,同一片段遮挡超次时自动向医生端告警并启动双向语音指导,保证数据有效性且无需人工后期筛选。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531610A_ABST
    Figure CN122531610A_ABST
Patent Text Reader

Abstract

An automatic tremor detection process control system and method, a doctor end enters patient basic information and monitors the running state of a detection room end in real time, receives a standardized detection report generated by a cloud end management platform, the detection room end is respectively connected with the doctor end and the cloud end management platform, tremor detection whole process state management and automatic operation are realized based on a scene state machine, and multi-modal data of a patient are synchronously collected and preprocessed, and effective multi-modal data are output; the cloud end management platform processes the effective multi-modal data, extracts tremor characteristics according to a clinical scale, and generates a standardized detection report. The process control module based on the finite state machine realizes automatic operation of the whole process from pre-examination preparation to post-examination ending, real-time identification of key part shielding, re-measurement of the shielded segment only after the shielding is removed, and guarantee of data effectiveness without manual post-screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of movement disorder detection technology, and in particular to a quantitative detection system and method for Parkinson's disease and essential tremor based on computer vision, multi-sensor fusion and automated control. Background Technology

[0002] With the aging population, the incidence of movement disorders such as Parkinson's Disease (PD) and Essential Tremor (ET) is increasing year by year. Clinical diagnosis and condition assessment mainly rely on doctors observing patients perform designated movements and making subjective scores based on scales such as the Fahn-Tolosa-Marin Tremor Scale and the Tetras Scale. This assessment method has inherent limitations: first, it relies on the doctor's experience, and different doctors may give different scores, making it highly subjective; second, it is difficult to quantify the subtle characteristics of tremor; third, it cannot simultaneously analyze the correlation between tremors in multiple sites such as the head, upper limbs, and lower limbs; and fourth, the testing process is time-consuming and requires multiple follow-up visits to compare treatment effects.

[0003] Existing tremor detection solutions based on sensors or vision only achieve partial automation, lacking hardware adaptive adjustment, full-process state machine control, and real-time occlusion closed-loop processing, thus failing to meet the clinical requirements for standardized, efficient, and unmanned detection.

[0004] Therefore, there is an urgent need for a system or method that can automate the entire tremor detection process, improve detection efficiency and data acquisition accuracy, and solve hardware adaptive adjustment and real-time occlusion handling. Summary of the Invention

[0005] (I) Purpose of the invention: In order to solve the problems existing in the prior art, the purpose of the present invention is to provide an automated tremor detection process control system and method, which solves the problems of insufficient automation of the tremor detection process, lack of adaptive adjustment capability of hardware, and imperfect occlusion handling mechanism in the prior art, so as to realize the fully automated operation without human intervention from pre-inspection preparation to post-inspection completion, thereby improving detection efficiency and data acquisition accuracy.

[0006] (II) Technical Solution: In order to solve the above technical problems, this technical solution provides an automated tremor detection process control system, including a doctor's end, a testing room end, and a cloud management platform;

[0007] The doctor's terminal is used to input basic patient information and output it to the testing room terminal, monitor the operating status of the testing room terminal in real time, and receive and display standardized test reports generated by the cloud management platform.

[0008] The testing room is connected to both the doctor's terminal and the cloud management platform. Based on the patient's basic information entered by the doctor's terminal, it implements full-process state management and automated operation of tremor detection using a finite state machine. It also collects and preprocesses the patient's multimodal data and outputs effective multimodal data.

[0009] The detection chamber includes a height recognition module, an equipment control module, an occlusion monitoring module, and a guidance and prompting module. The height recognition module acquires patient images and calculates patient height data. The equipment control module adaptively adjusts the height and angle of the motion capture camera and the position of the electric seat according to the height data to adapt to the current detection scene. The occlusion monitoring module outputs the visibility confidence score of each key point based on the skeleton key point detection network and graph neural network. When the visibility confidence score is lower than the preset visibility threshold and the duration exceeds the preset time threshold, and the key point belongs to the key area of ​​the current detection scene, it is considered effective occlusion. The timing and data acquisition of the current scene are paused, and the guidance and prompting module issues an adjustment prompt. After the occlusion is removed, the detection resumes and only the occluded segment is re-acquired.

[0010] The cloud management platform processes the effective multimodal data and extracts tremor features based on clinical scales to generate a standardized test report.

[0011] Preferably, the doctor's terminal includes a patient information entry module, a testing process control module, and a testing report management module;

[0012] The patient information entry module provides a graphical entry interface for receiving basic patient information entered by medical staff.

[0013] The detection process control module includes a voice communication submodule and an abnormal state intervention submodule. It establishes a communication connection with the detection room through the hospital intranet or wireless network to obtain the real-time operation status information of the detection room and remotely issue control commands.

[0014] The test report management module receives standardized test reports pushed by the cloud management platform in real time and presents the information in the standardized test reports in a graphical form to the doctor's end.

[0015] Preferably, the graphical information in the standardized test report for doctors includes a patient information section, a test overview section, a tremor index panel, a grading result display area, and a historical comparison curve area.

[0016] Preferably, the testing room includes an embedded server, a visual acquisition device, a wearable sensing device, environmental auxiliary equipment, and standardized props; the embedded server in the testing room is equipped with a process control module, which coordinates the collaborative work of each module based on a finite state machine to realize the management and transition control of the state nodes of the entire vibration detection process.

[0017] Preferably, the state nodes maintained by the finite state machine include standby state, pre-inspection preparation state, equipment calibration state, scene detection state, and post-inspection closing state;

[0018] Upon receiving the start-up command and basic patient information from the doctor's terminal, the system transitions from the standby state to the pre-test preparation state. In the pre-test preparation state, hardware self-tests and initial parameter adjustments are performed. After all hardware devices report a ready status, the system transitions to the device calibration state. In the device calibration state, multi-hardware device time synchronization and initial data quality verification are performed. After successful calibration, the system transitions to the scene detection state. In the scene detection state, based on the scale type selected by the doctor, the system loads the corresponding scene combination list from a preset scene library, performs hardware adaptive adjustments according to the scene type, guides patient actions, and collects multimodal data. If effective occlusion is detected during the test, the current scene timing and data collection are paused, an adjustment prompt trigger command is sent to the guidance prompt module, and the subsequent scene type conversion is determined based on the duration of occlusion. This continues until all scenes have been tested. In the post-test closing state, all data collection channels are stopped, valid data is summarized, and the system returns to the standby state.

[0019] Preferably, when the patient stands at the preset guide sign position, the height recognition module triggers the motion capture camera installed on the top of the detection room to acquire the patient's image; the built-in human posture estimation model locates the key points on the top of the patient's head and the key points on the bottom of the feet, and the pixel-to-actual distance mapping relationship is obtained by combining the calibration parameters of the motion capture camera, and the vertical distance from the top of the patient's head to the bottom of the feet is calculated as the patient's height data; and multi-frame image fusion and depth information correction are used to eliminate perspective errors and output the patient's height data.

[0020] Preferably, the occlusion monitoring module includes a skeleton key point detection network, a key point topology construction submodule, and a graph neural network occlusion determination submodule;

[0021] The skeleton key point detection network receives real-time video frames captured by various motion capture cameras and performs preprocessing to extract the two-dimensional skeleton key point coordinates and detection confidence of the target key points; it selects or weights and fuses the multi-view detection results to form multi-frame spatiotemporal skeleton feature data.

[0022] The key point topology construction submodule constructs a spatial topology map of target key points within the same frame based on the connection relationship of the human upper limb skeleton, and establishes a temporal connection between the same key points in consecutive frames to form a spatiotemporal skeleton map.

[0023] The graph neural network occlusion determination submodule takes the spatiotemporal skeleton graph as input, uses key point coordinates, detection confidence, and motion change in adjacent frames as node features, aggregates spatial and temporal neighborhood information, and outputs the visibility confidence score of each key point; when the visibility confidence score of a key region is lower than a preset visibility threshold and the duration exceeds a preset time threshold, it is considered effective occlusion.

[0024] Preferably, the skeleton keypoint detection network filters target keypoints related to tremor detection from the human keypoint set according to the current detection scene. For each video frame, it outputs the detection results of each target keypoint, including horizontal coordinates, vertical coordinates, and detection confidence. For a video segment containing T frames, the detection results of V target keypoints are the original keypoint feature tensor. The data includes horizontal coordinates, vertical coordinates, detection confidence, and motion increments between adjacent frames. For video streams captured by multiple motion capture cameras, based on the motion capture camera calibration parameters, key point detection confidence, and historical trajectory continuity, similar key points detected from different perspectives are associated as candidate point sets for the same physical joint. Outliers are removed, and the remaining multi-view detection results for the same target key point are weighted and fused to obtain spatiotemporal skeleton feature data.

[0025] The graph neural network occlusion determination submodule receives the spatiotemporal skeleton feature data of the target key points as input. Based on the connection relationship of the human skeleton, it constructs a spatial topology graph of the target key points in the same frame, and establishes temporal connections between the same target key points in consecutive frames to form a spatiotemporal skeleton graph. It aggregates the spatial neighborhood features and temporal neighborhood features of the target key points respectively, performs weighted fusion on the spatial neighborhood features and temporal neighborhood features, and outputs the visibility confidence score of the target key points. It also determines the occluded segment based on the visibility confidence change trend within a continuous time window.

[0026] Within a sliding time window, the visibility confidence of a target key point is statistically analyzed. When the visibility confidence of a target key point is lower than a preset visibility threshold and the duration exceeds a preset time threshold, it is determined to be a suspected occlusion. When the adjacent visible joints of the target key point remain stable in visibility, and the target key point experiences trajectory interruption, coordinate drift, or a sudden drop in confidence, and the target key point belongs to a key area of ​​the current detection scene, causing at least one of the following—tremor frequency, amplitude, direction, or action completion—to be unreliably extracted, the time period corresponding to the suspected occlusion is determined to be a valid occlusion segment.

[0027] Preferably, for the k-th target key point in the view of the i-th motion capture camera, the final weight of its weighted fusion is determined by the detection confidence of the key point, the distance attenuation coefficient of the key point from the edge of the field of view, and the angle between the observation line of sight and the normal of the limb surface.

[0028] The detection confidence is obtained directly from the output layer of the skeleton key point detection network;

[0029] The distance attenuation coefficient from the keypoint to the edge of the field of view is calculated using a Gaussian function based on the Euclidean distance from the keypoint pixel coordinates to the optical center of the image; the image resolution of the i-th motion capture camera is W×H, and the coordinates of the image optical center are... Key points Euclidean distance to the center of the image Distance attenuation coefficient Where σ is the attenuation control constant;

[0030] The angle between the observation line of sight and the normal to the limb surface is calculated based on the inner product relationship of spatial geometric vectors, and the world coordinates of the i-th motion capture camera are: The key point's estimated location in the previous frame's 3D space is... Then the observation line vector Based on the connection relationship between this key point and its adjacent key points in three-dimensional space, a limb skeleton vector centered on this key point is obtained. Calculate the spatial angle between the observation line-of-sight vector and the limb skeleton vector. : Normal angle coefficient ;

[0031] Calculate the initial joint weights of the keypoint from this perspective. Normalize the initial weights of all M valid motion capture camera views to obtain the final fusion weights for that keypoint. The weighted least squares method is used to output the fused spatiotemporal skeleton feature data.

[0032] Preferably, when in the first When a frame detects effective occlusion of the target joint, it backtracks along the time axis, extracts a continuous high-confidence history window of length W, extracts the displacement sequence s(t) of the target key point from the high-confidence history window, maps the time-domain signal to the frequency domain, and extracts the dominant frequency. Peak amplitude and initial phase Construct a periodic prior signal P(t): The periodic prior signal is extended to the same length as the current segment to be compensated, and then converted into high-dimensional prior embedding features through a linear mapping layer. .

[0033] Preferably, the multimodal patient data collected in the testing room includes: high frame rate video data collected by multiple motion capture cameras, acceleration and angular velocity data collected by smart sensing gloves and inertial measurement units, interactive data collected by standardized props, and physiological data collected by sensors and microphones. The testing room verifies the collected data, removes data segments marked as invalid due to occlusion or with abnormal data quality, and uploads the remaining valid data to the cloud management platform after timestamp alignment, encapsulation, and encryption.

[0034] Preferably, the detection chamber further includes an exogenous environmental parameter intervention module, a tremor signal determination module, and a tremor signal decoupling module;

[0035] When the patient's current tremor state stabilizes, the exogenous environmental parameter intervention module applies an intervention stimulus to the patient and records the baseline timestamp of the intervention stimulus taking effect.

[0036] The tremor signal determination module uses the reference timestamp as the absolute zero point to simultaneously acquire the autonomic nervous stress response time series and the kinematic agitation response time series, calculates the difference between the peak values ​​of the autonomic nervous stress response time series and the kinematic agitation response time series, and determines the source of tremor energy during that period.

[0037] The tremor signal decoupling module uses the autonomic nerve stress response time series as a priori guide vector to strip off the components that are synchronized with the prior guide vector and output a pure organic pathological tremor signal.

[0038] Preferably, the cloud management platform includes a patient information management module, a data acquisition and processing module, and a data analysis and reporting module;

[0039] The patient information management module centrally stores basic patient information, historical test records and diagnostic conclusions, and supports multi-center data sharing, access control and data de-identification export.

[0040] The data acquisition and processing module receives multimodal data uploaded from each testing room, performs secondary verification and format standardization, and stores the data in a distributed time-series database after sensor noise removal and video transcoding preprocessing.

[0041] The data analysis and reporting module has a built-in multi-scale grading algorithm library, which is used to automatically extract tremor frequency, amplitude, direction and movement completion features, generate local tremor grading scores for each part and a total comprehensive grading score for the whole body, combine the patient's historical data to draw tremor parameter time-series change curves, and output standardized test reports.

[0042] An automated vibration detection process control method, applied to the aforementioned automated vibration detection process control system, includes the following steps:

[0043] Enter the patient's basic information;

[0044] Based on the patient's basic information, a finite state machine is used to realize the state management and automated operation of the entire tremor detection process. The patient's image is collected and the patient's height data is calculated. Based on the height data, the height and angle of the motion capture camera and the position of the electric seat are adaptively adjusted to adapt to the current detection scene. The occlusion monitoring module outputs the visibility confidence score of each key point based on the skeleton key point detection network and graph neural network. When the visibility confidence score is lower than the preset visibility threshold and the duration exceeds the preset time threshold, and the key point belongs to the key area of ​​the current detection scene, it is considered effective occlusion. The timing of the current scene and data acquisition are paused, and the guidance prompt module issues an adjustment prompt. After the occlusion is removed, the detection is resumed and only the occluded segment is re-acquired. The patient's multimodal data is collected synchronously and preprocessed to output effective multimodal data.

[0045] The effective multimodal data is processed, and tremor features are extracted based on clinical scales to generate a standardized test report;

[0046] Display standardized testing reports.

[0047] (III) Beneficial Effects: This invention provides an automated tremor detection process control system and method. First, based on a finite state machine-based process control module, it achieves fully automated operation from pre-examination preparation to post-examination completion, significantly reducing the testing time for a single patient and minimizing manual operations by medical staff. Second, through the linkage between the height recognition module and the equipment control module, the height and angle of the motion capture camera and the seat position are automatically adjusted according to the patient's height, achieving seamless switching between standing and sitting posture detection scenarios with no blind spots. Third, through the closed-loop linkage between the occlusion monitoring module and the guidance prompt module, key occlusions are identified in real time, and the timing is automatically paused, invalid data is marked, and adjustment prompts are triggered. After the occlusion is removed, only the occluded segment is retested. If the same segment is occluded more than once, an alarm is automatically sent to the doctor's end and two-way voice guidance is initiated, ensuring data validity and eliminating the need for manual post-screening. Attached Figure Description

[0048] Figure 1 This is a system structure diagram of an automated vibration detection process control system according to the present invention;

[0049] Figure 2 This is a network architecture diagram between the cloud management platform of this invention and the testing room and doctor's terminals;

[0050] Figure 3 This is a flowchart of the steps of an automated vibration detection process control method of the present invention. Detailed Implementation

[0051] The present invention will be further described in detail below with reference to preferred embodiments. More details are set forth in the following description in order to provide a full understanding of the present invention. However, the present invention can obviously be implemented in many other ways different from those described herein. Those skilled in the art can make similar extensions and derivations based on actual application situations without departing from the spirit of the present invention. Therefore, the scope of protection of the present invention should not be limited by the content of this specific embodiment.

[0052] The accompanying drawings are schematic diagrams of embodiments of the present invention. It should be noted that these drawings are for illustrative purposes only and are not drawn to scale, and should not be construed as limiting the actual scope of protection of the present invention.

[0053] An automated vibration detection process control system, such as Figure 1 As shown, the system comprises interconnected doctor's end, testing room end, and cloud management platform, which work together to automate the entire tremor detection process. The doctor's end is used to input basic patient information and output it to the testing room end, monitor the operating status of the testing room end in real time, and receive and display standardized test reports generated by the cloud management platform, providing human-computer interaction and diagnostic decision support for medical staff. The testing room end, based on the basic patient information input by the doctor's end, implements full-process status management and automated operation of tremor detection using a scenario state machine, and simultaneously collects and preprocesses multimodal patient data, outputting valid multimodal data to the cloud management platform. The cloud management platform processes the valid multimodal data, extracts tremor features based on clinical scales, and generates standardized test reports, providing support for clinical diagnosis, efficacy evaluation, and research data accumulation.

[0054] Specifically, the doctor's terminal is deployed on a computer terminal in the examination room or nurse's station, serving as the entry point for medical staff to interact with the system. The doctor's terminal specifically includes a patient information entry module, a testing process control module, and a testing report management module.

[0055] The patient information entry module provides a graphical entry interface for receiving basic patient information entered by medical staff. The basic patient information includes the patient's name, age, height, weight, and medical history.

[0056] The graphical input interface refers to a display screen panel on the doctor's end that presents multiple interactive input controls. These controls include at least one of the following: text input box, number input box, drop-down selection box, date picker, radio button, and checkbox. The graphical input interface responds to the input operations of medical staff, retrieves the data field values ​​of the patient's basic information entered or selected in each input control, and performs data validation on the data field values ​​of the patient's basic information. The validated data field values ​​of the patient's basic information are then aggregated into a structured patient basic information dataset.

[0057] The input controls establish a bidirectional mapping with the data field values ​​of the patient's basic information through the program interface. When the information entered in the input control changes, the data field values ​​of the patient's basic information are updated synchronously. When the data field values ​​of the patient's basic information change due to a call, the content displayed in the corresponding input control is also refreshed synchronously.

[0058] The data validation is performed locally on the graphical input interface, including checking whether required fields are empty, whether input values ​​conform to preset data types, whether numeric fields are within preset reasonable ranges, whether specific field formats conform to regular expression rules, whether string lengths exceed limits, checking logical consistency between multiple fields, and filtering or escaping dangerous characters in the input content. Only when all data validations are successful can a structured patient basic information dataset be output. If any data validation fails, the graphical input interface displays an error message in real time and prevents form submission.

[0059] The patient information entry module is also used to encrypt the entered information and synchronize the encrypted patient basic information to the testing room terminal via the hospital intranet or wireless network, as the basic data for configuring the testing process parameters at the testing room terminal.

[0060] Preferably, when synchronizing patient basic information to the testing room, an encrypted transmission channel is established using TLS 1.2 or higher, so that the patient basic information is transmitted in encrypted form over the network; when storing patient basic information in a local database or cloud management platform, the AES algorithm is used to perform field-level encryption storage on sensitive fields, and the encrypted data is stored in the database instead of plaintext data, and the encryption key is stored and managed separately through a key management service.

[0061] The testing process control module includes a voice communication submodule and an abnormal status intervention submodule. The testing process control module establishes a communication connection with the core control system at the testing room end via the hospital's intranet or wireless network, acquiring real-time operating status information of the testing room and remotely issuing control commands.

[0062] When a preset type of anomaly occurs during the detection process at the detection chamber end and the automatic handling mechanism cannot resolve it, the detection chamber end pushes an anomaly notification message to the anomaly state intervention submodule. Upon receiving the anomaly notification message, the anomaly state intervention submodule generates a pop-up reminder message on the display screen panel. The pop-up reminder message includes a description of the anomaly type, the time of the anomaly occurrence and the corresponding detection scenario identifier, and suggested handling measures.

[0063] When the pop-up notification is triggered, the voice communication submodule automatically activates the microphone and speaker on the doctor's end and establishes a real-time audio stream transmission link with the built-in or external voice playback and acquisition device in the testing room. Medical staff can issue voice guidance commands in real time through the microphone on the doctor's end. These commands are transmitted over the network and then played back to the patient through the speaker in the testing room. Simultaneously, the patient's voice collected in the testing room can also be transmitted back to the doctor's end in real time via the audio stream transmission link, enabling two-way communication between medical staff and patients, facilitating timely understanding of the patient's condition and providing personalized guidance.

[0064] Preferably, the testing process control module sends status query requests to the testing room terminal at preset time intervals, or passively receives updates of the testing room's operating status information actively pushed by the testing room terminal via a WebSocket long connection. The testing process control module displays the testing room's operating status information in the form of a graphical dashboard on the doctor's terminal display screen panel, allowing medical staff to clearly understand the real-time operation of the testing room.

[0065] The operational status information of the testing room includes the current testing scene identifier, the remaining execution time of the scene, the working status of the hardware equipment, the occlusion status indicator, and the data acquisition quality identifier.

[0066] The current detection scene identifier indicates the name of the currently executing detection scene, including a standing still scene, a standing postural tremor scene, and a sitting action tremor scene. The remaining execution time of the scene displays the difference between the current execution time and the preset total time, presented as a countdown. The hardware device working status displays the online and operational status of each hardware device, including online / offline, normal / abnormal, in motion / in position, etc. The hardware device includes multiple motion capture cameras, an electric seat, smart sensing gloves, an IMU sensor module, a guidance display screen, and standardized props. The occlusion status indicator marks the occlusion status and occluded area in a prominent color or icon on the interface when effective occlusion occurs in the current detection scene. The data acquisition quality identifier indicates whether the packet loss rate, signal-to-noise ratio, and other quality indicators of the current data acquisition from each channel are within the normal range.

[0067] The graphical interface of the detection process control submodule is equipped with multiple interactive control controls, which are used to receive control commands from medical staff and send the control commands to the detection room to execute corresponding operations.

[0068] The control controls include a start detection control, a pause detection control, a terminate detection control, and a skip current scene control.

[0069] The "Start Detection" control responds to a click by medical staff, sending a command to the testing room to start the testing process, triggering the testing room to enter the pre-test preparation state from the standby state and begin executing the automated testing process. The "Pause Detection" control responds to a click by medical staff, sending a command to pause the testing process, pausing the currently executing testing scenario, halting timing and data acquisition; clicking the pause control again resumes the testing process. The "Terminate Detection" control responds to a click by medical staff, sending a command to terminate the testing process, forcibly ending the current testing process and jumping to the post-test closing state to perform data aggregation and saving operations. The "Skip Current Scenario" control optionally provides the function of skipping the current scenario. When medical staff determine that the current testing scenario is not suitable for the patient, they can use the "Skip Current Scenario" control to skip the current testing scenario and directly proceed to the next testing scenario.

[0070] The detection process control module also logs key events throughout the entire detection process. The log entries include the detection start time, start and end times for each scenario, received anomaly notifications, the types and durations of control operations performed by medical personnel, and the on / off times of voice communication. These log entries are stored locally on the doctor's end or synchronized to the cloud management platform for subsequent process auditing and quality traceability.

[0071] The test report management module is deployed on the doctor's computer terminal and establishes a communication connection with the cloud management platform through the hospital intranet or a dedicated data interface. It is used to receive, store, retrieve, and visualize standardized test reports generated by the cloud management platform.

[0072] The test report management module is equipped with a data receiving interface for receiving standardized test reports pushed by the cloud management platform in real time or near real time, and presenting the information in the standardized test reports in a graphical form on the display screen panel of the doctor's end. Corresponding to the standardized test report, it includes a patient information column, a test overview column, a tremor index panel, a grading result display area, and a historical comparison curve area.

[0073] The standardized test report supports printing or archiving in PDF, CSV, and Excel formats. The standardized test report can be de-identified of patient information according to the medical staff's choice to meet privacy protection requirements during research data sharing or external consultations.

[0074] The patient information section displays basic identification information such as the patient's name, age, and medical record number. The test overview section displays summary information such as the test date, scale used, total test duration, and data completeness. The tremor index panel displays the core quantitative tremor indicators of this test in the form of a numerical dashboard or bar chart, with normal reference ranges for comparison. The grading results display area displays the tremor grading score for each location and the overall score in the form of a table or radar chart.

[0075] The historical comparison curve area plots the time-series curves of tremor characteristic parameters for the same patient across multiple tests in the form of a line graph, visually presenting the trend of disease progression or treatment effectiveness. Specifically, all historical test report data for the patient under the same assessment scale are obtained from the local cache or by requesting the cloud management platform. The core tremor characteristic parameters for each test are extracted, including at least tremor frequency and tremor amplitude, and optionally also the graded scores for each body part or the overall total score. A line graph or smooth curve is plotted with the test date as the horizontal axis and the tremor characteristic parameter values ​​as the vertical axis; multiple curves are overlaid on the same coordinate system, using different colors or line types to distinguish different tremor characteristic parameters or different body parts; the specific values ​​for each test are marked on the curves, and important clinical intervention nodes are marked next to the curves.

[0076] The standardized test report provides report retrieval, allowing medical staff to search historical test reports based on various search criteria and their combinations. These search criteria include searching by patient information, by date range, by scale type, by testing laboratory or doctor, and combinations of these criteria.

[0077] The patient-information search retrieves a list of all historical test reports for a patient by inputting a unique identifier such as the patient's name, medical record number, or ID card number. The date-range search filters for all test reports generated within a specified time period by selecting a start and end date. The scale-type search filters for test reports assessed using the corresponding scale by selecting a specific scale name. The laboratory or physician-based search filters for test reports handled by the corresponding laboratory or physician by selecting the laboratory number or responsible medical staff identifier.

[0078] The search results are presented in list format, with each record displaying key fields such as patient name, test date, scale type, and overall score. Healthcare professionals can click on any record in the list to load and display the complete standardized test report for that test.

[0079] The testing room includes an embedded server, visual acquisition equipment, wearable sensing devices, environmental auxiliary equipment, and standardized props. The embedded server, deployed within a cabinet in the testing room, operates the entire testing room to detect tremors in patients. The visual acquisition equipment includes multiple motion capture cameras and electrically operated lifting and rotating supports for each camera. The motion capture cameras are distributed throughout the testing room, including at least a frontal motion capture camera, left and right side motion capture cameras, an overhead motion capture camera, and left and right lower side motion capture cameras, used to acquire high frame rate video streams. The electrically operated lifting and rotating supports support electric adjustment of the height and angle of the motion capture cameras. The wearable sensing device includes bilateral smart sensing gloves, a headband with an IMU (In-Mechanical Unit) sensor, and bilateral leg IMU straps. The smart sensing gloves are worn on the patient's hands and contain built-in IMU sensors, fingertip pressure sensors, heart rate sensors, and microphones to collect hand tremor data. The headband with the IMU sensor is worn on the patient's head and contains built-in IMU sensors to collect head tremor data. The leg IMU straps are worn on both legs and contain built-in IMU sensors to collect leg tremor data. The environmental support equipment includes an electric seat, a floor rail for moving the electric seat, a guidance display screen, and floor guidance markers. The electric seat is mounted on the floor rail and supports horizontal movement along the rail, seat height adjustment, and automatic armrest deployment and locking. The guidance display screen is installed within the patient's field of vision in the testing room and has built-in or external speakers for playing graphic, animated, and voice guidance prompts. The floor guidance markers are placed on the floor of the testing room to indicate the patient's standing position. The standardized props include, but are not limited to, e-ink drawing boards, weighted water cups, chopstick clips, prop positioning slots, and LED indicator lights, used to collect tremor data from patients from multiple angles.

[0080] An embedded server within the testing chamber houses a process control module, serving as the core control hub of the automated vibration detection process control system. This embedded server, housed in a cabinet within the testing space, connects to various hardware devices deployed within the testing chamber via multiple communication interfaces, performing functions such as fully automated process control, multimodal data acquisition and preprocessing, and data uploading.

[0081] The process control module is the core scheduling unit of the testing chamber. Based on a finite state machine, it coordinates the collaborative work of various modules to manage and control the state nodes of the entire vibration detection process. The state nodes maintained by the finite state machine include standby state, pre-test preparation state, equipment calibration state, scene detection state, and post-test closing state. The transition trigger delay between state nodes in the process control module does not exceed 3 seconds, ensuring seamless connection between standing and sitting posture detection.

[0082] The standby state refers to the system being idle and waiting for the detection task to start. Upon receiving the start detection command and basic patient information from the doctor's end, the system transitions from the standby state to the pre-detection preparation state. The pre-detection preparation state performs hardware self-tests and initial parameter adjustments. After all hardware devices report a ready status, the system transitions to the device calibration state. The device calibration state performs time synchronization of multiple hardware devices and initial data quality verification. After successful calibration, the system transitions to the scene detection state. The scene detection state loads the corresponding scene combination list from a preset scene library based on the scale type selected by the doctor's end. Hardware adaptive adjustment is performed according to the scene type to guide patient actions and collect multimodal data. If effective occlusion is detected during the detection process, the current scene timing and data collection are paused, an adjustment prompt trigger command is sent to the guidance prompt module, and the subsequent scene type conversion is determined based on the duration of occlusion; this continues until all scenes have been detected. The post-detection closing state stops all data acquisition channels, summarizes valid data, calls the data verification module for integrity verification, releases hardware resources, and returns to the standby state after completion.

[0083] Specifically, the testing chamber includes a system self-test module and a time synchronization module. Before testing begins, the system self-test module automatically verifies the communication connectivity of each hardware device, checks the quality of sensor data, and confirms that the actuator is in its initial position.

[0084] The testing process is executed automatically at preset intervals before and during the testing process, checking the working status of each hardware device deployed in the testing room, including device connection status detection, sensor data quality detection, and actuator position verification. When an anomaly is detected, an alarm message containing the abnormal device identifier, anomaly type, and suggested troubleshooting measures is generated and pushed to the testing process control module on the doctor's end for display. The testing results are also written to the testing log file for subsequent maintenance and traceability.

[0085] The device connection status detection uses polling or a heartbeat mechanism to check the communication link between each hardware device and the detection chamber, including the motion capture camera, electric seat, smart sensing gloves, IMU sensor module, guide display screen, and standardized props. The sensor data quality detection collects the output data of each sensor in a static state, calculates the packet loss rate, signal-to-noise ratio, and baseline drift, and determines whether it is within a preset normal range. The actuator position verification reads the current position encoder value of the motion capture camera's electric bracket and the limit switch status of the electric seat, compares it with the preset initial position, and confirms that the actuator has been reset.

[0086] The time synchronization module establishes and maintains a unified time reference among all data acquisition devices before the detection begins, ensuring the timestamp alignment accuracy of multimodal data.

[0087] A master clock signal is broadcast to all motion capture cameras via Precision Time Protocol (PTP). The detection chamber acts as the PTP master clock source, and each motion capture camera acts as a PTP slave clock node, periodically exchanging timestamp messages to compensate for network transmission latency, ultimately achieving a time synchronization error of less than or equal to 1 millisecond among the motion capture cameras. Time synchronization beacon frames are sent to the smart sensing gloves, head IMU headband, and leg IMU straps via Bluetooth Low Energy (BLE) broadcast channel. The time synchronization beacon frame contains the current system timestamp from the detection chamber. Upon receiving the beacon frame, each wearable device associates its local sampling clock with the received reference timestamp and includes both its local timestamp and the synchronized global timestamp in each subsequent IMU data packet it reports.

[0088] The detection chamber also includes a height recognition module, a scene recognition module, an anomaly handling module, and a data acquisition module.

[0089] Once the patient stands at the pre-set guide marker position on the ground, the height recognition module triggers a motion capture camera mounted on the ceiling of the testing room to capture a top-down image of the patient. The height recognition module incorporates a pre-trained human pose estimation model, which employs a keypoint detection network based on the YOLO-pose architecture. This model includes a pre-processing layer, a backbone feature extraction network, a feature pyramid fusion network, a keypoint detection head, and a keypoint post-processing layer, connected sequentially. The human pose estimation model infers from the captured top-down image, outputting the distance values ​​of each pixel in the patient's head region relative to the motion capture camera, and the distance values ​​of each pixel in the foot region relative to the motion capture camera. Combining this with the motion capture camera calibration parameters, the pixel-to-actual distance mapping is obtained, and the vertical distance from the patient's head to the ground is calculated, which is the patient's height data.

[0090] Specifically, after the motion capture camera at the top of the testing room captures a top-down image of the patient, the image is input into the image preprocessing layer. The image preprocessing layer performs image cropping, size normalization, distortion correction, and pixel normalization on the top-down image to obtain a standardized image tensor. The backbone feature extraction network performs convolutional feature extraction on the standardized image tensor and outputs intermediate feature maps at different scales. The feature pyramid fusion network fuses the intermediate feature maps at different scales to generate a fused feature map containing semantic information of the human body contour, the top of the head region, and the bottom of the feet region. The keypoint detection head outputs a set of human keypoints based on the fused feature map. Each keypoint in the set includes horizontal coordinates, vertical coordinates, and a confidence score. The keypoint post-processing layer filters the set of human keypoints according to the confidence score and performs temporal smoothing on the detection results of multiple consecutive frames to determine the top of the head keypoints and the bottom of the feet keypoints. The height recognition module calculates the vertical distance between the key points on the top of the head and the key points on the bottom of the feet in the actual spatial coordinate system, based on the image coordinates, combined with the calibration parameters of the motion capture camera on the top of the detection chamber and depth information, to obtain the patient's height data. Multi-frame image fusion and depth information correction are used to eliminate perspective errors, outputting patient height data with an error of no more than 2 cm.

[0091] The scene recognition module identifies the patient's positional changes in real time during the testing process. By receiving real-time video streams from the overhead motion capture camera installed on the top of the testing room, it analyzes the patient's positional status frame by frame, providing a trigger basis for the automatic switching between standing and sitting posture testing scenarios.

[0092] The scene recognition module incorporates a human posture analysis algorithm. It determines the patient's position by analyzing changes in key geometric features of the patient's body contour in video frames captured by the motion capture camera. Specifically, it extracts the vertical span of the patient's body contour (the pixel height from the top of the head to the bottom of the feet) in the image; it also extracts the height coordinates of the patient's head region in the image coordinate system. When the vertical span of the patient's body contour abruptly changes from a larger value when standing to a smaller value when sitting, and the height coordinates of the head region decrease significantly, it is determined that the patient has completed the transition from standing to sitting. Conversely, when the vertical span of the body contour increases from a smaller value when sitting to a larger value when standing, and the height coordinates of the head region increase significantly, it is determined that the patient has completed the transition from sitting to standing. Simultaneously, corresponding adjustments are made to various hardware devices, including the height of the motion capture camera and the position of the electric seat.

[0093] Once a body position switching event is determined, the scene recognition module generates a body position switching event signal, which includes the switching direction, and sends the body position switching event signal to the process control module. This triggers the process control module to switch the state machine from the scene execution state to the scene switching state, thereby initiating the corresponding hardware adaptive adjustment process.

[0094] During the detection process, standardized action guidance and abnormal status prompts are provided to the patient, and occlusion behavior is monitored in real time. This invention incorporates a cascaded model of a YOLO-based skeleton recognition network and a graph neural network, including a skeleton keypoint detection network, a keypoint topology construction submodule, and a graph neural network occlusion determination submodule. The skeleton keypoint detection network is based on the YOLO-pose architecture, receiving real-time video frames from various motion capture cameras. After preprocessing the video frames, it extracts the two-dimensional skeleton keypoint coordinates and detection confidence of the target keypoints. It selects or weights and fuses the multi-view detection results to form multi-frame spatiotemporal skeleton feature data. The keypoint topology construction submodule constructs a spatial topology map of target keypoints within the same frame based on the connection relationship of the human upper limb skeleton, and establishes temporal connections between the same keypoints in consecutive frames to form a spatiotemporal skeleton map. The graph neural network occlusion determination submodule uses the spatiotemporal skeleton map as input, takes the keypoint coordinates, detection confidence, and motion change in adjacent frames as node features, aggregates spatial and temporal neighborhood information through a message passing mechanism, and outputs the visibility confidence score of each keypoint. When the visibility confidence score of a key area is lower than a preset visibility threshold and the duration exceeds a preset time threshold, it is considered a valid occlusion. Upon determining valid occlusion, the timer is immediately paused, the occluded period data is marked as invalid, and a guidance prompt is triggered to instruct the patient to adjust their posture. After the occlusion is removed, only the occluded action segment is re-sampled; the entire scene does not need to be repeated. If occlusion occurs more than once, an alarm is automatically sent to the doctor's end, and two-way voice intervention is initiated.

[0095] Specifically, when extracting the coordinates of two-dimensional key points of the patient's hand and upper limb skeleton, the skeleton key point detection network filters target key points related to tremor detection from the human key point set according to the current detection scenario. The target key points include at least the shoulder joint, elbow joint, wrist joint, palm key points, and fingertip key points. For each video frame, the detection results of each target key point are output, including the horizontal coordinate, vertical coordinate, and detection confidence. For a video segment containing T frames, the detection results of V target key points are organized into an original key point feature tensor. The feature channels of the original keypoint feature tensor include horizontal coordinates, vertical coordinates, detection confidence, and motion increments of adjacent frames. For video streams captured by multiple motion capture cameras, based on the motion capture camera calibration parameters, keypoint detection confidence, and historical trajectory continuity, similar keypoints detected from different perspectives are associated as candidate point sets for the same physical joint. Outliers with low detection confidence are eliminated, and the remaining multi-view detection results for the same target keypoint are weighted and fused to form spatiotemporal skeleton feature data.

[0096] In cases of high-frequency tremor or rapid hand movements by the patient, conventional multi-motion capture camera simple averaging fusion methods are highly susceptible to interference from edge distortion and blind spots, leading to coordinate drift in the extracted tremor spatial trajectory and severely affecting the accuracy of tremor amplitude assessment. Therefore, for the k-th target keypoint from the i-th motion capture camera's viewpoint, its final weighted fusion weight... The key point detection confidence level, the distance attenuation coefficient from the key point to the edge of the field of view, and the angle between the observation line of sight and the normal to the limb surface are all determined. In the case of instantaneous self-occlusion caused by rapid hand movement, the motion continuity of adjacent frames is compared, and the view detection results with displacement abrupt changes exceeding a preset displacement threshold are directly discarded. The fused spatiotemporal skeleton feature data is then output to ensure that the coordinates of the target key points remain stable after fusion.

[0097] Furthermore, the detection confidence is directly obtained from the output layer of the skeleton keypoint detection network. The skeleton keypoint detection network regresses the two-dimensional keypoint coordinates. Simultaneously, the product of the category probability corresponding to the key point and the IoU of the bounding box / key point localization is output as the detection confidence of the key point. The detection confidence score range is normalized to [0,1]. The higher the value, the more reliable the observation from that perspective.

[0098] The distance attenuation coefficient of the keypoint from the edge of the field of view is calculated using a Gaussian function based on the Euclidean distance from the keypoint pixel coordinates to the optical center of the image. This is used to suppress the positioning error caused by edge distortion of the motion capture camera lens, ensuring that keypoints located in the center of the image receive higher weights, while the weights of keypoints near the edges are smoothly attenuated, thereby effectively eliminating the influence of perspective error on the measurement of shakiness amplitude. Let the image resolution of the i-th motion capture camera be W×H, and the coordinates of the image optical center be... ,in , Calculate key points Euclidean distance to the center of the image The maximum safe radius of the image diagonal is Construct a distance decay coefficient based on a Gaussian function. Where σ is the attenuation control constant, .

[0099] The angle between the observation line of sight and the normal to the limb surface is calculated based on the inner product relationship of spatial geometric vectors, and is used to solve self-occlusion and feature compression caused by an excessively small angle between the observation line of sight and the limb surface. Let the world coordinates of the i-th motion capture camera be... The key point's estimated location in the previous frame's 3D space is... Then the observation line vector Simultaneously, based on the connection relationship between this key point and its adjacent key points in three-dimensional space, a limb skeleton vector centered on this key point is constructed. Calculate the spatial angle between the observation line-of-sight vector and the limb skeleton vector. : Define the normal angle coefficient When the line of sight is perpendicular to the surface of the limb, Take the maximum value of 1; when the line of sight is parallel to the surface of the limb, Approaching 0, thus effectively reducing the fusion weight of non-orthogonal perspectives.

[0100] Calculate the initial joint weights of the keypoint from this perspective. To address instantaneous self-occlusion caused by rapid hand movements, a temporal motion continuity constraint is introduced. The 2D pixel displacement ΔS of the keypoint between the current and previous frames is calculated from the i-th viewpoint. If ΔS exceeds a preset physiological motion threshold, indicating a feature point jump or flying point due to occlusion, the initial joint weight is forcibly reset to zero. =0, directly remove the abnormal observation.

[0101] Normalize the initial weights of all M valid motion capture camera views to obtain the final fusion weights for that keypoint. The projection matrices of each motion capture camera are used, combined with the normalized final weights. By using the weighted least squares method, the high-precision, distortion-resistant true coordinates of the key point in three-dimensional space are obtained, and the fused spatiotemporal skeleton feature data is output.

[0102] The attenuation of the distance from key points to the edge of the field of view and the angle between the observation line of sight and the normal to the limb surface dynamically reduce the weight of distortion edges and non-orthogonal viewpoints. At the same time, the elimination of displacement abrupt changes solves the trajectory tearing caused by instantaneous self-occlusion, enabling the extraction of truly clinically diagnostic, distortion-resistant micro-tremor displacement features in complex motion detection.

[0103] For each key point, confidence information that changes over time is retained for subsequent dynamic occlusion identification. Unlike prediction using only coordinate values, this invention uses confidence as an important basis for occlusion determination and information fusion, enabling the subsequent compensation process to distinguish between true and unreliable observations.

[0104] Secondly, the graph neural network occlusion determination submodule receives the spatiotemporal skeleton feature data of the target key points as input. Based on the connection relationship of the human upper limb skeleton, it constructs a spatial topology graph of target key points such as shoulders, elbows, wrists, palms, and fingertips within the same frame, and establishes temporal connections between the same target key points in consecutive frames to form a spatiotemporal skeleton graph. The nodes in the spatiotemporal skeleton graph correspond to hand and upper limb key points, and the node features include key point coordinates, detection confidence, displacement increment between adjacent frames, velocity features, acceleration features, and historical visibility status. The edges in the spatiotemporal skeleton graph include spatial edges and temporal edges. The spatial edges represent the connection relationship between adjacent skeletal key points within the same video frame, and the temporal edges represent the connection relationship between the same key point in adjacent video frames.

[0105] The graph neural network occlusion determination submodule takes a spatiotemporal skeleton graph as input and a target keypoint as the objective. It collects feature vectors corresponding to each neighboring node in the target's spatial neighborhood, as well as the relative displacement and angle between the neighboring nodes and the target node. A multilayer perceptron is used to learn the spatial influence weights of each neighboring node on the target node, and then the spatial neighborhood features of the target keypoint are obtained through mean aggregation. These spatial neighborhood features include coordinate changes of neighboring keypoints, detection confidence, relative distance, relative angle, and joint motion consistency features.

[0106] Centered on the current video frame, a time window encompassing several frames before and after is extracted to construct a local trajectory feature sequence for the target keypoint. Using the target node's features in the current frame as the query vector, the association weights between each historical frame and the current frame within the time window are automatically learned. Where low-confidence observations exist in historical frames, their corresponding association weights are reduced. Finally, a weighted sum is obtained to obtain the temporal neighborhood features. These temporal neighborhood features include the target keypoint's trajectory in consecutive frames, confidence change sequence, displacement change, velocity change, and acceleration change.

[0107] After obtaining the spatial and temporal neighborhood features, the graph neural network occlusion determination submodule concatenates the spatial and temporal neighborhood features, calculates a gating coefficient with a value between 0 and 1 using a fully connected layer and a sigmoid activation function. The spatial neighborhood features are then weighted using this gating coefficient, and the temporal neighborhood features are weighted using the remainder after subtracting the gating coefficient from 1, outputting the visibility confidence score of the target keypoint.

[0108] When the visibility confidence of a target key point is lower than a preset visibility threshold in consecutive frames and the duration exceeds a preset time threshold, or when it decreases significantly relative to its historical average, the time period corresponding to the key point is determined to be an occluded segment.

[0109] After identifying occluded segments, a dynamic mask matrix M is generated. Each element in the matrix corresponds to the available state of a keypoint at a given time. If the keypoint is observed validly in the frame, the corresponding position is assigned a value of 1. If the keypoint is occluded or in a low-confidence state, the corresponding position is assigned a value of 0. By multiplying the mask matrix element-wise with the original feature tensor, the masked input feature tensor can be obtained. : , where ⊙ represents element-wise product, which allows the subsequent reconstructed network to explicitly perceive which inputs come from reliable observations and which locations need to rely on historical priors and information from adjacent joints to complete compensation.

[0110] A spatiotemporal graph structure G, incorporating both spatial and temporal connections, is further constructed. Nodes in the graph correspond to human keypoints, spatial edges correspond to the connections between adjacent joints in the human skeletal topology, and temporal edges correspond to the connections between the same keypoint in adjacent frames. This structural representation allows for the simultaneous encoding of human structural constraints and trajectory evolution constraints within the same model. If a hand keypoint is occluded, the system can still jointly infer the occluded segment using the motion changes of nearby visible joints such as the elbow and shoulder, as well as the state information of the keypoint at previous and subsequent time points, and determine whether the time segment corresponding to the keypoint is a valid occluded segment.

[0111] The effective occlusion segment is determined based on the visibility confidence of key points, duration, spatial neighborhood consistency, and the configuration of key parts in the current detection scene. Specifically, the visibility confidence of target key points is statistically analyzed within a sliding time window. When the visibility confidence of a target key point is lower than a preset visibility threshold and the duration exceeds a preset time threshold, it is determined to be a suspected occlusion. Furthermore, when the adjacent visible joints of the target key point maintain stable visibility, and the target key point experiences trajectory interruption, coordinate drift, or a sudden drop in confidence, and the target key point belongs to a key area of ​​the current detection scene, causing at least one of the following—tremor frequency, amplitude, direction, or motion completion—to be unreliably extracted, the time period corresponding to the suspected occlusion is determined to be an effective occlusion segment.

[0112] Both the visibility threshold and the time threshold can be adaptively set according to the device type, imaging quality, key point category, and historical statistical distribution.

[0113] Once a valid occlusion is determined, individualized tremor periodic priors are extracted using valid observation information from the same object, joint, and time period prior to the occlusion. Furthermore, tremor periodic priors for the patient and joint in the current detection scenario are extracted using valid observation data prior to the occlusion, which assists in trajectory compensation and validity assessment of the occluded segment. The process control module triggers closed-loop control actions based on the valid occlusion determination result, including pausing timing, prompting the patient to adjust their posture, resuming detection after the occlusion is removed, and only re-collecting the corresponding segment to reconstruct the occluded segment.

[0114] Utilizing effective observation data before occlusion for auxiliary analysis improves the system's tolerance to temporary data loss caused by occlusion. Specifically, when the system... When a frame detects occlusion of the target joint, it traces backward along the time axis and extracts a continuous high-confidence historical window of length W. This window is preferably selected from a period of continuous unocclusion or stable high confidence before the occlusion occurred, used to characterize the true tremor rhythm of the target joint during the current detection process.

[0115] In one embodiment, the system extracts the displacement sequence s(t) of the target keypoint in a certain direction from the historical window. This direction can be the horizontal axis, vertical axis, depth axis, or a synthetic displacement modulus. Subsequently, frequency domain analysis is performed on the displacement sequence, which can be achieved using Fast Fourier Transform or other spectral estimation methods to map the time-domain signal to the frequency domain. Extracting the main frequency Peak amplitude and initial phase By using parameters such as these, we can obtain the individualized tremor periodicity prior within the target tremor frequency band.

[0116] The spectral peaks are searched within a target frequency band of 3Hz to 12Hz to obtain the dominant tremor frequency that best matches the current patient. Compared to fixed templates or empirical frequencies, this invention extracts the rhythmic information that has actually occurred in the current patient during the current time period, and is therefore more suitable as the driving basis for occlusion compensation.

[0117] After obtaining parameters such as the dominant frequency, peak amplitude, and initial phase, the corresponding periodic prior signal P(t) is constructed: The periodic prior signal is used to describe the underlying oscillation pattern that the target joint may continue to exhibit after occlusion occurs. To facilitate fusion with network features, the periodic prior signal can be extended to the same length as the current segment to be compensated and converted into high-dimensional prior embedding features through a linear mapping layer or multilayer perceptron. In this way, the subsequent trajectory reconstruction module not only receives the incomplete observations after occlusion, but also receives explicit priors reflecting the current individual rhythmic characteristics. Reconstructing the occluded segments is no longer simply filling in coordinate gaps, but rather preserving the original tremor rhythm as much as possible while maintaining a reasonable structure.

[0118] When the occlusion detection module receives an occlusion trigger signal, the guidance prompt module interrupts the playback of the current normal guidance content and prioritizes sending occlusion adjustment prompts, such as displaying the text "Hands detected to be occluded, please move your hands into the motion capture camera's field of view" on the screen, and simultaneously playing the corresponding voice prompt.

[0119] When the number of effective occlusions within the same detection scene segment exceeds a preset occlusion threshold during the detection process, the anomaly handling module determines it as an occlusion over-limit anomaly. At this time, an anomaly notification message is pushed to the doctor's terminal, which includes an anomaly type description, occurrence time, corresponding scene identifier, and suggested handling measures. Simultaneously, a two-way voice communication channel is activated between the doctor's terminal and the detection room, allowing medical staff to remotely issue posture adjustment guidance instructions to the patient. The current detection process is paused according to a preset configuration, awaiting a continue or termination instruction from the doctor's terminal.

[0120] When the number of effective occlusions within the same detection scene segment during the detection process is less than the preset occlusion threshold, the data acquisition module synchronously collects the patient's multimodal tremor data during the detection process, providing a data basis for subsequent tremor feature analysis and scale scoring.

[0121] The data acquisition module includes a video data acquisition submodule, an IMU data acquisition submodule, a prop data acquisition submodule, and a physiological data acquisition submodule.

[0122] The video data acquisition submodule establishes a data stream transmission channel with multiple motion capture cameras distributed at preset locations within the detection space via a gigabit Ethernet interface, based on the GigE Vision protocol. The multiple motion capture cameras output raw video frame data at a preset high frame rate, with each frame marked with a precise hardware timestamp at the motion capture camera hardware level. After receiving the video stream, the video data acquisition submodule segments the video stream according to the current scene identifier provided by the process control module, performs real-time compression using a video encoding format, and stores it to the local disk array. The naming rules for the stored video files include the patient identifier, detection date, scene number, and motion capture camera number.

[0123] The IMU data acquisition submodule establishes a data connection with the wearable device via the BLE GATT protocol. Each wearable device reports IMU data packets at a preset sampling frequency. Each data packet contains a unique device identifier, three-axis acceleration data, three-axis angular velocity data, optional magnetometer data, and a synchronized global timestamp. In addition, the data packets reported by the smart sensing glove also include pressure sensor data and heart rate sensor data from each fingertip. The IMU data acquisition submodule writes the received data into a time-series data buffer in timestamp order.

[0124] The prop data acquisition submodule establishes a data connection with the standardized props configured in the testing room via a Bluetooth serial port or a dedicated wireless receiver, and receives interactive data generated during the patient's operation of the standardized props. The types of data collected include the patient's finger sliding trajectory coordinate sequence and timestamp reported by the e-ink screen drawing board, weight change data and tilt angle data reported by the weighted water cup with built-in weighing sensor, and clamping force data reported by the chopsticks clamp with built-in bending sensor.

[0125] The physiological data acquisition submodule uses the photoplethysmography (PPG) sensor built into the smart sensing glove to collect the patient's heart rate and heart rate variability parameters in real time during the testing process; it also collects the patient's respiratory sounds and speech signals through the sensors in the smart sensing glove or microphones deployed in the testing room. The physiological data acquisition submodule stores the collected physiological data in timestamp order for subsequent auxiliary assessment of mental state.

[0126] The multimodal patient data collected in the testing room includes high frame rate video data collected by multiple motion capture cameras, acceleration and angular velocity data collected by smart sensing gloves and inertial measurement units, interactive data collected by standardized props, and physiological data collected by sensors and microphones.

[0127] The testing chamber also includes a data verification module and a data caching module.

[0128] After data acquisition is completed, the data verification module performs real-time or near-real-time verification of the integrity and validity of the data acquired from each channel. It removes data segments marked as invalid due to occlusion or with abnormal data quality, and uploads the remaining valid data to the cloud management platform after timestamp alignment and encapsulation encryption to ensure reliable data quality, including packet loss rate, video frame continuity, timestamp monotonicity, and abnormal data value detection.

[0129] The packet loss rate is the ratio of the number of data packets that each data channel should receive to the number of data packets actually received within a statistical unit of time. When the packet loss rate of a certain channel exceeds a preset threshold, the data of that channel is determined to be incomplete, triggering a quality anomaly flag. The video frame continuity refers to checking whether the timestamp interval between consecutive video frames is stable near a preset frame interval, detecting whether there are frame drops, frame duplication, or frame out-of-order phenomena. The timestamp monotonicity verifies whether the timestamps of each data frame within the same data channel are strictly increasing, preventing timestamp rollback or data out-of-order due to abnormal device clocks. Threshold judgments are performed on the acceleration, angular velocity, and other data output by the IMU sensor; if the values ​​exceed the limits of human physiological movement, they are marked as abnormal values.

[0130] Data segments that pass verification are allowed to proceed to the subsequent synchronization and upload process; data segments that fail verification are marked as invalid, and are removed by the data synchronization module after detection, and the exception handling module is notified to trigger the corresponding retest or alarm process.

[0131] During data acquisition and uploading, the data caching module provides local temporary storage capabilities, writing the raw data streams output by each data acquisition module into a circular buffer in memory in real time. This provides low-latency data temporary storage for the data verification and data synchronization modules to read from nearby locations. Simultaneously with writing to the circular buffer in memory, data is asynchronously written to a persistent disk cache to prevent data loss due to unexpected system power outages or process crashes.

[0132] After the data synchronization module confirms that the data has been successfully uploaded to the cloud management platform, it automatically clears the corresponding local cache data and frees up storage space for subsequent testing tasks.

[0133] The cloud management platform is deployed in the hospital's data center or a medical cloud server, such as Figure 2 As shown, a communication connection is established with the testing room and the doctor's end through the hospital intranet or the Internet to centrally manage patient information, process multimodal testing data, and generate standardized test reports.

[0134] The cloud management platform includes a patient information management module, a data acquisition and processing module, and a data analysis and reporting module.

[0135] The patient information management module is used to centrally store and manage a patient information database containing all patient records. This database stores basic patient information, historical test records, and diagnostic conclusions. Basic patient information includes patient name, age, gender, height, weight, and a summary of medical history. Historical test records include the date of each test, the type of scale used, the testing room number, the test duration, and the data storage path. When patient data needs to be used for scientific research statistical analysis or cross-institutional sharing, the module, according to preset anonymization rules, masks, generalizes, or replaces personally identifiable fields such as patient name, ID number, and contact information, generating an anonymized dataset and exporting it in a standardized data format.

[0136] The patient information management module supports patient data sharing between multiple testing centers or hospital campuses. Through a preset access control list, tiered permissions are set for medical staff in different departments and hospital campuses, including levels such as view-only, editable, and exportable. After logging into the doctor's terminal with authorized credentials, doctors can only access patient files within their authorized scope.

[0137] The data acquisition and processing module is equipped with a data receiving interface, which receives composite data packets of multimodal data uploaded by the data synchronization module at the testing room end via an SSL / TLS encrypted channel. The composite data packet includes basic patient information, valid video clip files, IMU time-series data files, prop interaction data files, physiological data files, and occlusion event log files.

[0138] Upon receiving the composite data packet, the data acquisition and processing module performs a secondary verification on the uploaded data. This verification sequentially checks the correspondence between the patient's basic information and existing patient records in the system to ensure correct data attribution; it also verifies whether the format of each data file conforms to preset standards; and it verifies the integrity of the composite data packet to ensure no data corruption or truncation occurred during the upload process. If the secondary verification fails, an error code is returned to the testing room, requesting a re-upload of the data.

[0139] If the secondary verification passes, the data acquisition and processing module performs format standardization processing on the data from each modality, converting the heterogeneous data formats from different testing chambers into a unified standard data format within the system. This standardization processing includes sensor noise removal, video transcoding, data slicing, and indexing.

[0140] The sensor noise removal refers to using low-pass or band-pass filters to remove high-frequency noise and baseline drift from IMU acceleration and angular velocity data; the video transcoding refers to uniformly transcoding the original video files into a preset encoding format and resolution suitable for cloud storage and network access; the data slicing and indexing refers to slicing the time-series data collected continuously over a long period of time according to the detection scenario and establishing a time index to facilitate rapid retrieval by subsequent analysis modules.

[0141] After preprocessing, the data acquisition and processing module stores the processed data in a distributed time-series database and creates a globally unique detection data identifier for each detection data, which is then associated with the corresponding patient information database.

[0142] The data analysis and reporting module has a built-in library of multi-scale grading algorithms for various clinical scales.

[0143] The multi-scale grading algorithm library supports clinical scales including, but not limited to, the Fahn-Tolosa-Marin Tremor Scale (FTM), the Unified Parkinson's Disease Rating Scale (UPDRS) tremor section, and the Tetras Tremor Scale. The scoring rules for each scale are encoded as callable algorithm functions, which define the mapping relationship between the input parameters and output scores for tremor ratings of each body part.

[0144] The data analysis and reporting module automatically extracts tremor feature parameters from the preprocessed multimodal data. These tremor feature parameters include tremor frequency, tremor amplitude, tremor direction, and movement completion.

[0145] The tremor frequency refers to the peak frequency of the power spectral density identified by performing a Fast Fourier Transform on the IMU acceleration or angular velocity signal. The tremor amplitude refers to the root mean square value or peak-to-peak value of the acceleration or angular velocity signal in the time domain. The tremor direction refers to the spatial principal direction vector of the tremor motion calculated based on the triaxial acceleration components. The action completion rate refers to quantifying the quality of action completion by comparing the deviation between the trajectory of the patient's actual action in the video and the trajectory of a standard action template.

[0146] Based on the extracted tremor feature parameters, the corresponding clinical scale algorithm function is called to score the local tremor in each body part; and based on the local tremor score, the total score of the whole body comprehensive grading is calculated according to the scale's preset summarization rules.

[0147] The data analysis and reporting module also includes historical test data of the same patient under the same clinical scale system. Tremor characteristic parameters from each test are extracted, and a time-series curve of tremor parameter changes is plotted with the test date on the horizontal axis and the parameter value on the vertical axis. By comparing the trends of the curves from different tests, the module can help assess disease progression or the effectiveness of treatment interventions.

[0148] The analysis results are packaged into a standardized test report. The standardized test report includes basic patient information and test summary, numerical tables and graphical displays of tremor characteristic parameters for each body part, local tremor grading scores for each body part and total scores for the overall grading, historical trend curves of tremor characteristic parameters, summaries of occlusion events during the test, and data quality descriptions.

[0149] The standardized test report is stored in PDF format and structured data package on the cloud management platform, and is pushed to the test report management module on the doctor's end through the data interface of the cloud management platform for medical staff to view.

[0150] An automated vibration detection process control method is applied to the aforementioned automated vibration detection process control system, such as... Figure 3 As shown, the specific steps include:

[0151] Enter the patient's basic information;

[0152] Based on the patient's basic information, the system realizes full-process state management and automated operation of tremor detection using a scenario state machine, and simultaneously collects and preprocesses the patient's multimodal data to output effective multimodal data.

[0153] The effective multimodal data is processed, and tremor features are extracted based on clinical scales to generate a standardized test report;

[0154] Display standardized testing reports.

[0155] Example 1: The entire process of tremor detection for a Parkinson's disease patient, Mr. Zhang, is used as an example.

[0156] Patient basic information: Mr. Zhang, male, 65 years old, height 170cm, weight 70kg, chief complaint of resting tremor in both hands for six months, preliminary clinical diagnosis of Parkinson's disease pending evaluation. The doctor selected the Fahn-Tolosa-Marin tremor scale for quantitative assessment.

[0157] Doctors log into the system via their computer terminals and access the graphical user interface for patient information entry. The interface displays text input boxes, numeric input boxes, drop-down selection boxes, and a date picker. The doctor enters "Zhang Mou" in the name text box, "65" in the age numeric box, "170" in the height field, and "70" in the weight field. The doctor selects "Male" from the gender drop-down selection box and the date of birth from the date picker. In the medical history field, the doctor enters "Holding tremor of both hands for six months, Parkinson's disease pending evaluation," and selects "Fahn-Tolosa-Marin Tremor Scale" from the scale type drop-down box. During the entry process, the front-end validation program verifies each field in real time. After all required fields are entered and validated, the doctor clicks the "Confirm Submission" button. The patient basic information entry module aggregates the input data from each control into a structured dataset in JSON format, encrypts it using the TLS 1.3 protocol, and synchronizes it to the testing room via the hospital intranet. The doctor then clicks the "Start Testing" control, and the testing command is sent synchronously.

[0158] After receiving the start command and patient information, the process control module in the embedded server at the testing room transitions from standby to pre-test preparation state.

[0159] First, the system self-test module runs automatically. A heartbeat detection mechanism verifies the communication links of each hardware device, including the front, left, right, overhead binocular, lower left, and lower right motion capture cameras, all of which return to online status. The electric seat controller returns to standby status via the RS485 bus. The dual smart sensor gloves, head IMU headband, and dual leg IMU straps report sufficient power and normal sensor operation via BLE protocol. Data collected by each sensor in a static state shows a packet loss rate of 0.02%, and the signal-to-noise ratio is within the preset normal range. The encoder readings of the current position of each motion capture camera's electric bracket match the preset initial position, and the electric seat is in the guide rail storage position. The self-test passes, and the log is completed.

[0160] The nurse guided patient Zhang into the testing room and fitted him with smart sensor gloves, an IMU headband, and IMU straps on both legs. All devices completed their power-on self-tests and reported normal status.

[0161] The guidance display screen plays a voice prompt: "Please stand at the blue marker on the ground." After the patient stands in position, the height recognition module triggers the overhead binocular motion capture camera to capture a top-down image. The left and right eyes of the binocular motion capture camera simultaneously image, obtaining depth information of each pixel in the patient's head area through the principle of parallax. The built-in YOLO-pose keypoint detection network infers from the image, locating the two-dimensional coordinates of keypoints on the top of the head and the bottom of the feet. Combining the focal length, baseline distance, and installation height of the binocular motion capture camera, the vertical distance from the top of the head to the bottom of the feet is calculated through the pixel-to-actual distance mapping relationship. The system continuously acquires 5 frames of images, and the results of the 5 calculations are processed with mean filtering. The final output of the patient's height is 170.2cm, with an error within ±2cm.

[0162] After receiving the height data, the device control module calculates the optimal position parameters for each motion capture camera. The height of the front motion capture camera is approximately 145cm (170.2 × 0.85); the height of the left and right motion capture cameras is approximately 119cm (170.2 × 0.70); and the height of the bottom motion capture camera is approximately 77cm (170.2 × 0.45). The device control module then sends command messages containing the target height and angle values ​​to the motorized support controllers of each motion capture camera sequentially via TCP protocol. The motorized support motors of each motion capture camera start, and upon reaching the target position, they provide a completion response. Once all the motion capture camera motorized supports are adjusted, the system displays a "ready" status.

[0163] The process control module transitions the state machine from the pre-inspection preparation state to the equipment calibration state. The testing chamber acts as the PTP master clock source, broadcasting PTP synchronization messages to the six motion capture cameras. Each motion capture camera, acting as a slave clock node, periodically exchanges timestamp messages. After network latency compensation, the clock alignment error of each motion capture camera is 0.8ms, meeting the requirement of no more than 1ms. Simultaneously, the time synchronization module sends time synchronization beacon frames to the smart sensing glove, head IMU headband, and bilateral leg IMU straps via the BLE broadcast channel. Upon receiving these frames, each device associates its local sampling clock with the reference timestamp, and each subsequent reported data packet contains both the local and global timestamps. After synchronization, the data verification module performs initial quality checks on each channel. A brief period of static state data acquisition is performed on each channel to confirm that the packet loss rate and signal-to-noise ratio are within the preset normal range. Calibration is successful, and the state machine transitions to the scene detection state.

[0164] The process control module loads the corresponding scenario combination list from the preset scenario library based on the Fahn-Tolosa-Marin tremor scale selected by the doctor, and arranges them in the optimization order as follows: Scenario 1: Standing still, Scenario 2: Standing postural tremor (arms outstretched), Scenario 3: Standing action tremor (finger-to-nose test), Scenario 4: Sitting still, Scenario 5: Sitting postural tremor (arms outstretched), Scenario 6: Sitting action tremor (cup grip).

[0165] The first scenario is a standing posture. The motion capture camera was already adjusted during the pre-test preparation stage and requires no further adjustment. After the hardware is ready, the scenario execution state is entered. The guidance prompt module sends a standing still motion demonstration animation and text instructions to the guidance display screen via the WebSocket protocol: "Please maintain a standing posture, look straight ahead, let your arms hang naturally, and remain still for 30 seconds," and simultaneously plays a pre-recorded voice command. The video data acquisition submodule receives video streams output from 6 motion capture cameras at 120fps via the GigE Vision protocol. Each frame carries a hardware timestamp, is segmented according to the scene identifier "P001_20250424_S01," and is compressed and stored on the local disk array using H.265 encoding. The IMU data acquisition submodule receives IMU data packets reported at a frequency of 100Hz from the bilateral smart sensor gloves, the head IMU headband, and the bilateral leg IMU straps via the BLE GATT protocol. These data packets include three-axis acceleration, three-axis angular velocity, and global timestamps. The gloves also additionally report fingertip pressure and heart rate data. The physiological data acquisition submodule recorded the patient's heart rate as 78 bpm and respiratory rate as 16 breaths / minute. The occlusion monitoring module continuously analyzed 6 video streams. The YOLO skeleton recognition network extracted the coordinates of 21 key points on the patient's hand and upper limb frame by frame, and the GNN graph neural network output the confidence score of each key point. Throughout Scenario 1, the confidence score of each key point was above 92%, and no occlusion events occurred. After 30 seconds, Scenario 1 was completed.

[0166] The patient transitions from a standing to a sitting position. The device control module sends commands to the electric seat controller via the Modbus RTU protocol through the RS485 bus: the seat moves from its wall-mounted storage position along the guide rail to 50cm directly behind the standing marker; the lifting column automatically adjusts the seat height to 45cm based on the patient's height of 170cm; and the armrest motor drives the armrests to unfold and lock. The entire process takes 2.8 seconds, not exceeding 3 seconds. The guidance display screen prompts: "Please take a seat." After the patient sits down, the scene recognition module analyzes the real-time video stream from the overhead motion capture camera. It detects a sudden change in the patient's body contour longitudinal span from approximately 680 pixels in the standing position to approximately 420 pixels in the sitting position, and a decrease in the head height coordinate from approximately 180 pixels to approximately 105 pixels, thus determining that the position switch is complete. The device control module then executes a parameter switch for the motion capture cameras: the height of the front motion capture camera decreases to 130cm, the height of the left and right side motion capture cameras decreases to 105cm, and the angle of the overhead motion capture camera is slightly adjusted to cover the detection platform. The switch is completed within 0.9 seconds. Once the hardware is ready, the guide display plays a demonstration of a seated, static posture, and data acquisition proceeds normally, similar to scenario 1.

[0167] During the execution of Scenario 6, when patient Zhang was grasping a water cup and raising his arm, his left hand involuntarily shifted to the left side of his body, causing some key points of his left hand to be obscured by the left side of his torso.

[0168] The occlusion detection module detected that the confidence level of the key points of the left index and middle fingers dropped to 78% at time t1 and remained so until time t3, exceeding the preset time threshold, and was therefore determined to be valid occlusion. This triggered a closed-loop processing flow: the process control module switched the state machine from the scene execution state to the occlusion processing state, pausing the scene 6 timer; each data acquisition module marked the data within the time window from t1-0.5s to t3+0.5s as "occlusion invalid"; the guidance prompt module interrupted the normal guidance screen, displayed the text "Left hand detected as occluded, please keep both hands in front of your body within the motion capture camera's field of view," and simultaneously played a voice prompt.

[0169] After hearing the prompt, the patient adjusted their arm position, and their left hand re-entered the motion capture camera's field of view. The occlusion monitoring module detected that the confidence level of the left hand's key points had recovered to 93%. The occlusion was immediately removed, and the process control module switched to a guided prompt state. After confirming that the posture had returned to normal, it returned to the scene execution state. The system only retested the occluded motion segment; the remaining valid data collected in Scene 6 were retained. The remaining time period of Scene 6 was occluded and completed successfully.

[0170] After all six scenarios have been executed and the scenario list is empty, the process control module transitions the state machine to the post-inspection closing state. All data acquisition modules stop all data acquisition channels. The data synchronization module, based on the marker information generated during the detection process by the occlusion monitoring module and the data verification module, removes invalid occlusion data segments from t1-0.5s to t3+0.5s in scenario 6; it aligns the global timestamps of valid data from each channel, generating a composite data packet containing six valid video files, IMU timing data files, prop interaction data files, physiological data files, and occlusion event logs. The composite data packet is encrypted using SSL / TLS and uploaded to the cloud management platform via the hospital's intranet. The device control module sends a reset command: the electric seat returns to its wall-mounted storage position along the guide rail, and the electric brackets of each motion capture camera reset to their initial positions. Hardware resources are released.

[0171] The anomaly handling module summarizes the anomaly logs for this test: 0 instances of occlusion exceeding the limit, 0 instances of device communication timeout, and 0 instances of data quality anomalies. The logs are stored locally. The process control module returns the state machine to standby mode, awaiting the next test task.

[0172] After receiving the encrypted composite data packet, the data acquisition and processing module of the cloud management platform verifies that the patient's name and ID number match the existing records in the database, that the formats of each data file conform to preset standards, and that the data packet integrity check passes. Subsequently, a 0.5-20Hz bandpass filter is used to remove high-frequency noise and baseline drift from the IMU acceleration data; the video is uniformly transcoded to H.264 encoding at 1080p resolution; and the time-series data is sliced ​​into six scenes and a time index is created. The processed data is stored in a distributed time-series database and linked to the patient Zhang's records.

[0173] The data analysis and reporting module automatically extracted tremor characteristic parameters, including: Scene 1 (standing still): left hand tremor frequency 5.2Hz, amplitude 2.8mm; right hand tremor frequency 5.1Hz, amplitude 3.1mm; Scene 4 (sitting still): left hand tremor frequency 5.3Hz, amplitude 2.6mm; right hand tremor frequency 5.0Hz, amplitude 2.9mm. The FTM scale algorithm function was used for grading and scoring. Both hands were rated as level 2 for resting tremors, level 2 for postural tremors, and level 1 for action tremors. The total score was 12 points. Combined with Zhang's historical test data from 3 months prior, the FTM total score was 9 points. A time-series curve of tremor frequency and amplitude was plotted, showing a slight upward trend for both parameters.

[0174] Once the standardized test report is generated, it is stored in PDF format and as a structured data package, and then pushed to the doctor's end through the data interface.

[0175] After receiving the report, the doctor's report management module displays the report content on the screen. The doctor can view the patient information section, the test overview section, the tremor index panel, the grading result display area, and the historical comparison curve area. Based on the report and the clinical manifestations, the doctor adjusts the medication treatment plan for Zhang and archives the test report.

[0176] Example 2: In real clinical diagnosis, after entering the testing room, patients may experience greater tremor amplitude than usual due to emotional factors such as tension, anxiety, and unfamiliar environment. It is impossible to determine whether the currently detected tremor is a genuine organic pathological tremor or a pseudo-tremor temporarily aggravated by tension. In this example, the testing room further integrates an exogenous environmental parameter intervention module, a tremor signal determination module, and a tremor signal decoupling module. During the execution of the testing scenario, when the patient's current tremor state is determined to be stable, the exogenous environmental parameter intervention module applies an intervention stimulus variable of preset intensity to the patient or the testing environment and records the baseline timestamp of the intervention stimulus variable taking effect, which serves as the time calibration starting point for inducing and extracting non-pathological stress response sequences. The tremor signal determination module uses the reference timestamp of the intervention stimulus variable taking effect as the absolute zero point, simultaneously acquiring two time series: physiological stress flow and kinematic catastrophic flow. It calculates the difference between the peak values ​​of the autonomic nervous stress response time series and the kinematic catastrophic response time series, and outputs the tremor energy source within that time period based on the joint determination results of multiple preset conditions. The tremor signal decoupling module divides the total tremor signal into organic pathological tremor components, stress-induced tremor components, and environmental noise. Using the extracted autonomic nervous stress response time series as a priori guiding vector, it establishes a correlation mapping between the total tremor signal and the priori guiding vector. A constrained independent component analysis algorithm is used to forcibly remove the pseudo-tremor components synchronized with the priori guiding vector, outputting a pure organic pathological tremor signal. This provides more precise medication guidance and psychological intervention basis for clinical practice.

[0177] Taking Mr. Zhang, a Parkinson's disease patient in Example 1, as an example, during routine static posture detection, the data analysis module of the cloud management platform analyzes the triaxial acceleration signal of the right wrist collected by the IMU sensor in real time. A local sampling window with a fixed time span of 2 seconds is set, and the sampling slides across the continuous data stream in 0.5-second increments, calculating the root mean square of the synthetic acceleration magnitude within the current window as the local tremor amplitude. Specifically, for the triaxial acceleration within each window... , , Calculate the resultant acceleration The root mean square value of the composite acceleration within this window is used as the average flutter amplitude of that window. Simultaneously record the peak amplitude of the acceleration signal within this window. For each consecutive three-window group, the average flutter amplitude is calculated. The mean and standard deviation are used to calculate the coefficient of variation (CV) = (standard deviation / mean) × 100%. When the coefficient of variation of the tremor amplitude in three consecutive windows is lower than the preset variability threshold, it is determined that the patient's current tremor state has reached a stable baseline and exogenous stimulation can be triggered.

[0178] The variation threshold is determined based on clinical pre-experiment data. The coefficient of variation for spontaneous physiological tremor in healthy adults at rest is typically 8%–15%, while it can reach 3%–6% in Parkinson's disease patients during the effective period of the medication. In this embodiment, the variation threshold is set to 5%.

[0179] The exogenous environmental parameter intervention module pre-stores various intervention stimulus variables. Based on the patient's medical history or prior questionnaire, the doctor pre-enters the patient's mental stress sensitivity and selects an appropriate intervention stimulus type from the exogenous environmental parameter intervention module. These intervention stimulus types include, but are not limited to, auditory intervention stimuli, environmental visual intervention stimuli, and tactile cognitive composite stimuli. Auditory intervention stimuli are achieved by guiding the built-in speaker on the display screen to play sudden low-frequency white noise, or by issuing cognitive load commands such as reciting numbers backwards through a voice system. Environmental visual intervention stimuli are achieved by adjusting the color temperature of the testing room lights, or by guiding the screen to flash visual stimulus patterns at specific frequencies. Tactile cognitive composite stimuli are achieved by providing vibration prompts through smart gloves and simultaneously issuing voice commands for mental arithmetic tasks.

[0180] Once the patient's tremor has reached a stable baseline, the process control module triggers the exogenous environmental parameter intervention module, injecting an intervention stimulus variable of preset intensity into the current detection environment. Simultaneously, the absolute time of this instruction's activation is recorded by the real-time task scheduler, serving as the baseline timestamp in the multimodal data stream. This baseline timestamp is broadcast to all data acquisition channels via the time synchronization module, serving as the time zero point for all subsequent data processing.

[0181] In this embodiment, the patient reported being particularly sensitive to sudden sounds, therefore auditory intervention stimulation was chosen: [the stimulation was administered at specific times]. When the intervention stimulus command is triggered, the built-in speaker on the guidance display screen at the detection chamber suddenly plays a low-frequency white noise with a frequency of 150Hz, an intensity of 65dB, and a duration of 0.5 seconds. At the same time, the voice system issues the intervention stimulus command: Please count backwards: 100, 99, 98...

[0182] After the auditory intervention stimulus is injected, the photoplethysmography (PPG) sensor built into the smart sensing glove continuously collects the vascular volume pulse wave signal at the patient's fingertip at a sampling rate of 256Hz. The system uses... Using time zero, the vascular volume pulse wave signal was bandpass filtered from 0.5Hz to 15Hz and motion artifacts suppressed before extracting the inter-heartbeat sequence. Based on the inter-heartbeat sequence, heart rate variability indices were calculated in real time, including low-frequency power, high-frequency power, and the low-frequency / high-frequency ratio. An increased low-frequency / high-frequency ratio is a typical marker of sympathetic nerve activation and stress response.

[0183] The system uses a base timestamp Starting from the baseline, physiological data within 15 seconds is extracted, and an autonomic nervous system stress response time series S(t) is generated in real time. The time point at which the low-frequency / high-frequency ratio jumps from the baseline value to the peak value is recorded as . This time point reflects the delay in the patient's autonomic nervous system's response to interventional stimuli.

[0184] Simultaneously, visual acquisition equipment and inertial measurement units synchronously extract the tremor response after auditory intervention stimulation injection. Multiple motion capture cameras reconstruct the 3D spatial coordinates of key points on the right wrist at 120fps, calculate the 3D displacement between adjacent frames, and extract the displacement envelope using the sliding root mean square method. The triaxial acceleration is bandpass filtered and the synthesized acceleration magnitude is calculated; the acceleration envelope is also extracted using the sliding root mean square method. Both are fused in real-time to obtain the kinematic catastrophic response time series E(t). The system detects the time point where the kinematic catastrophic response time series shows a significant increase and reaches its peak, denoted as t0. This time point is typically 100 to 300 milliseconds later than the peak of the autonomic nervous system stress response, reflecting the inherent delay in the transmission of neural signals from the autonomic nervous system to skeletal muscle.

[0185] The tremor signal determination module dynamically aligns the physiological stress flow and kinematic dysregulation flow time series. When 100ms ≤ Within a reasonable neural conduction window of ≤300ms, if the correlation coefficient between the rising slope of the autonomic stress response time series and the rising slope of the kinematic agitation response time series on the time axis is greater than 0.7, and if the frequency band with a coherence coefficient greater than 0.6 in the increase of tremor energy accounts for more than 50%, then the surge in tremor energy during this period is determined to originate from the patient's emotional stress response, rather than their inherent organic pathological tremor, and the autonomic stress response time series is used as the prior guiding vector. Otherwise, if any condition is not met, the surge in tremor energy is considered to originate from organic pathological tremor, the process is terminated, and no prior guiding vector is generated.

[0186] The total tremor signal acquired at the detection chamber is the magnitude of the acceleration measured by the IMU sensor on the right wrist. This total tremor signal includes components of organic pathological tremor, stress-induced tremor, and environmental noise. The time-varying intensity of the stress-induced tremor component is highly correlated with physiological stress flow.

[0187] The autonomic nervous system stress response time series is used as a priori guiding vector and injected into the tremor signal decoupling module along with the total tremor signal. Specifically, a constrained independent component analysis framework is adopted, adding a constraint term to the objective function to force the maximization of mutual information between one output component and the kinematic stress response time series. During the separation process, the tremor signal decoupling module forcibly classifies the component in the total tremor signal that is highly synchronized with the physiological stress flow rhythm and has a highly similar waveform morphology as stress-induced tremor and separates it from the total tremor signal. The remaining signal component is the pure organic pathological tremor. When the correlation coefficient between the stress-induced tremor and the autonomic nervous system stress response time series is greater than or equal to 0.6, the separation is considered successful; when the correlation coefficient is less than 0.6, the separation is considered unsuccessful, and the total tremor signal is treated as a pathological signal.

[0188] The data analysis and reporting module of the cloud management platform calculates the percentage of the energy of the stress-induced tremor component relative to the energy of the original total tremor signal based on the stripped organic pathological tremor signal, defining it as the stress susceptibility index; and recalculates the tremor score of each body part based on the organic pathological tremor signal, and calls the clinical scale algorithm to generate the organic tremor baseline score.

[0189] The baseline score for organic tremor reflects the true severity of the patient's lesions and is used to guide drug dosage adjustments.

[0190] This invention provides an automated tremor detection process control system and method. First, based on a finite state machine-based process control module, it achieves fully automated operation from pre-detection preparation to post-detection completion, significantly reducing the detection time per patient and minimizing manual operations by medical staff. Second, through the linkage between the height recognition module and the equipment control module, the height and angle of the motion capture camera and the seat position are automatically adjusted according to the patient's height, achieving seamless switching between standing and sitting detection scenarios with no blind spots. Third, through the closed-loop linkage between the occlusion monitoring module and the guidance prompt module, key occlusions are identified in real time, and the timing is automatically paused, invalid data is marked, and adjustment prompts are triggered. After the occlusion is removed, only the occluded segment is retested. If the same segment is occluded more than once, an alarm is automatically sent to the doctor's end and two-way voice guidance is initiated, ensuring data validity and eliminating the need for manual post-screening.

[0191] In summary, the automated tremor detection process control system and method provided by this invention solves the problems of insufficient process automation, lack of adaptive hardware adjustment, lack of occlusion handling mechanism, poor patient experience and weak data quality control in the prior art, and provides objective, comprehensive and reliable quantitative data support for the clinical diagnosis, disease monitoring and efficacy evaluation of Parkinson's disease and essential tremor.

[0192] The above description illustrates preferred embodiments of the present invention and helps those skilled in the art to more fully understand the technical solution of the present invention. However, these embodiments are merely illustrative and should not be construed as limiting the specific implementation of the present invention to these embodiments. For those skilled in the art, several simple deductions and modifications can be made without departing from the inventive concept, and all such modifications should be considered within the protection scope of the present invention.

Claims

1. An automated vibration detection process control system, characterized in that, This includes the doctor's end, the testing room end, and the cloud management platform; The doctor's terminal is used to input basic patient information and output it to the testing room terminal, monitor the operating status of the testing room terminal in real time, and receive and display standardized test reports generated by the cloud management platform. The testing room is connected to both the doctor's terminal and the cloud management platform. Based on the patient's basic information entered by the doctor's terminal, it implements full-process state management and automated operation of tremor detection using a finite state machine. It also collects and preprocesses the patient's multimodal data and outputs effective multimodal data. The detection chamber includes a height recognition module, an equipment control module, an occlusion monitoring module, and a guidance and prompting module. The height recognition module acquires patient images and calculates patient height data. The equipment control module adaptively adjusts the height and angle of the motion capture camera and the position of the electric seat according to the height data to adapt to the current detection scene. The occlusion monitoring module outputs the visibility confidence score of each key point based on the skeleton key point detection network and graph neural network. When the visibility confidence score is lower than the preset visibility threshold and the duration exceeds the preset time threshold, and the key point belongs to the key area of ​​the current detection scene, it is considered effective occlusion. The timing and data acquisition of the current scene are paused, and the guidance and prompting module issues an adjustment prompt. After the occlusion is removed, the detection resumes and only the occluded segment is re-acquired. The cloud management platform processes the effective multimodal data and extracts tremor features based on clinical scales to generate a standardized test report.

2. The automated vibration detection process control system according to claim 1, characterized in that, The doctor's terminal includes a patient information entry module, a testing process control module, and a testing report management module; The patient information entry module provides a graphical entry interface for receiving basic patient information entered by medical staff. The detection process control module includes a voice communication submodule and an abnormal state intervention submodule. It establishes a communication connection with the detection room through the hospital intranet or wireless network to obtain the real-time operation status information of the detection room and remotely issue control commands. The test report management module receives standardized test reports pushed by the cloud management platform in real time and presents the information in the standardized test reports in a graphical form to the doctor's end.

3. The automated vibration detection process control system according to claim 2, characterized in that, The standardized test report for doctors includes graphical information such as patient information, test overview, tremor index panel, grading results display area, and historical comparison curve area.

4. The automated vibration detection process control system according to claim 1, characterized in that, The testing room includes an embedded server, visual acquisition equipment, wearable sensing equipment, environmental auxiliary equipment, and standardized props. The embedded server in the testing room is equipped with a process control module, which coordinates the collaborative work of each module based on a finite state machine to realize the management and transition control of the state nodes of the entire tremor detection process.

5. An automated vibration detection process control system according to claim 4, characterized in that, The state nodes maintained by the finite state machine include standby state, pre-inspection preparation state, equipment calibration state, scene detection state, and post-inspection closing state. Upon receiving the start-up command and basic patient information from the doctor's terminal, the system transitions from the standby state to the pre-test preparation state. In the pre-test preparation state, hardware self-tests and initial parameter adjustments are performed. After all hardware devices report a ready status, the system transitions to the device calibration state. In the device calibration state, multi-hardware device time synchronization and initial data quality verification are performed. After successful calibration, the system transitions to the scene detection state. In the scene detection state, based on the scale type selected by the doctor, the system loads the corresponding scene combination list from a preset scene library, performs hardware adaptive adjustments according to the scene type, guides patient actions, and collects multimodal data. If effective occlusion is detected during the test, the current scene timing and data collection are paused, an adjustment prompt trigger command is sent to the guidance prompt module, and the subsequent scene type conversion is determined based on the duration of occlusion. This continues until all scenes have been tested. In the post-test closing state, all data collection channels are stopped, valid data is summarized, and the system returns to the standby state.

6. The automated vibration detection process control system according to claim 1, characterized in that, Once the patient stands at the preset guide sign position, the height recognition module triggers the motion capture camera installed on the top of the testing room to capture the patient's image; The built-in human pose estimation model locates the key points on the top of the patient's head and the key points on the bottom of the feet. Combined with the calibration parameters of the motion capture camera, the pixel-to-actual distance mapping relationship is obtained, and the vertical distance from the top of the patient's head to the bottom of the feet is calculated as the patient's height data. It also uses multi-frame image fusion and depth information correction to eliminate perspective errors and outputs the patient's height data.

7. The automated vibration detection process control system according to claim 1, characterized in that, The occlusion monitoring module includes a skeleton key point detection network, a key point topology construction submodule, and a graph neural network occlusion determination submodule. The skeleton key point detection network receives real-time video frames captured by various motion capture cameras and performs preprocessing to extract the two-dimensional skeleton key point coordinates and detection confidence of the target key points; it selects or weights and fuses the multi-view detection results to form multi-frame spatiotemporal skeleton feature data. The key point topology construction submodule constructs a spatial topology map of target key points within the same frame based on the connection relationship of the human upper limb skeleton, and establishes a temporal connection between the same key points in consecutive frames to form a spatiotemporal skeleton map. The graph neural network occlusion determination submodule takes the spatiotemporal skeleton graph as input, uses key point coordinates, detection confidence, and motion change in adjacent frames as node features, aggregates spatial and temporal neighborhood information, and outputs the visibility confidence score of each key point; when the visibility confidence score of a key region is lower than a preset visibility threshold and the duration exceeds a preset time threshold, it is considered effective occlusion.

8. An automated vibration detection process control system according to claim 7, characterized in that, The skeleton keypoint detection network filters target keypoints related to tremor detection from the human keypoint set according to the current detection scene. For each video frame, it outputs the detection results of each target keypoint, including horizontal coordinates, vertical coordinates, and detection confidence. For a video segment containing T frames, the detection results of V target keypoints are the original keypoint feature tensor. The data includes horizontal coordinates, vertical coordinates, detection confidence, and motion increments between adjacent frames. For video streams captured by multiple motion capture cameras, based on the motion capture camera calibration parameters, key point detection confidence, and historical trajectory continuity, similar key points detected from different perspectives are associated as candidate point sets for the same physical joint. Outliers are removed, and the remaining multi-view detection results for the same target key point are weighted and fused to obtain spatiotemporal skeleton feature data. The graph neural network occlusion determination submodule receives the spatiotemporal skeleton feature data of the target key points as input. Based on the connection relationship of the human skeleton, it constructs a spatial topology graph of the target key points in the same frame, and establishes temporal connections between the same target key points in consecutive frames to form a spatiotemporal skeleton graph. It aggregates the spatial neighborhood features and temporal neighborhood features of the target key points respectively, performs weighted fusion on the spatial neighborhood features and temporal neighborhood features, and outputs the visibility confidence score of the target key points. It also determines the occluded segment based on the visibility confidence change trend within a continuous time window. Within a sliding time window, the visibility confidence of target key points is statistically analyzed. When the visibility confidence of a target key point is lower than a preset visibility threshold and the duration exceeds a preset time threshold, it is judged as a suspected occlusion. When the adjacent visible joints of the target key point remain stable and the target key point experiences trajectory interruption, coordinate drift, or sudden drop in confidence, and the target key point belongs to the key area of ​​the current detection scene, resulting in at least one of the tremor frequency, amplitude, direction, or action completion being unable to be reliably extracted, the time period corresponding to the suspected occlusion is determined as a valid occlusion segment.

9. An automated vibration detection process control system according to claim 8, characterized in that, For the k-th target keypoint from the viewpoint of the i-th motion capture camera, the final weight of its weighted fusion is determined by the detection confidence of the keypoint, the distance attenuation coefficient of the keypoint from the edge of the field of view, and the angle between the observation line of sight and the normal of the limb surface. The detection confidence is obtained directly from the output layer of the skeleton key point detection network; The distance attenuation coefficient from the keypoint to the edge of the field of view is calculated using a Gaussian function based on the Euclidean distance from the keypoint pixel coordinates to the optical center of the image; the image resolution of the i-th motion capture camera is W×H, and the coordinates of the image optical center are... Key points Euclidean distance to the center of the image Distance attenuation coefficient Where σ is the attenuation control constant; The angle between the observation line of sight and the normal to the limb surface is calculated based on the inner product relationship of spatial geometric vectors, and the world coordinates of the i-th motion capture camera are: The key point's estimated location in the previous frame's 3D space is... Then the observation line vector Based on the connection relationship between this key point and its adjacent key points in three-dimensional space, a limb skeleton vector centered on this key point is obtained. ; Calculate the spatial angle between the observation line vector and the limb skeleton vector. : Normal angle coefficient ; Calculate the initial joint weights of the keypoint from this perspective. Normalize the initial weights of all M valid motion capture camera views to obtain the final fusion weights for that keypoint. The weighted least squares method is used to output the fused spatiotemporal skeleton feature data.

10. An automated vibration detection process control system according to claim 8, characterized in that, When in the When a frame detects effective occlusion of the target joint, it backtracks along the time axis, extracts a continuous high-confidence history window of length W, extracts the displacement sequence s(t) of the target key point from the high-confidence history window, maps the time-domain signal to the frequency domain, and extracts the dominant frequency. Peak amplitude and initial phase Construct a periodic prior signal P(t): The periodic prior signal is extended to the same length as the current segment to be compensated, and then converted into high-dimensional prior embedding features through a linear mapping layer. .

11. An automated vibration detection process control system according to claim 1, characterized in that, The multimodal patient data collected in the testing room includes: high frame rate video data collected by multiple motion capture cameras, acceleration and angular velocity data collected by smart sensing gloves and inertial measurement units, interactive data collected by standardized props, and physiological data collected by sensors and microphones. The testing room verifies the collected data, removes data segments marked as invalid due to occlusion or with abnormal data quality, and uploads the remaining valid data to the cloud management platform after timestamping, encapsulation, and encryption.

12. The automated vibration detection process control system according to claim 1, characterized in that, The detection chamber also includes an exogenous environmental parameter intervention module, a tremor signal determination module, and a tremor signal decoupling module; When the patient's current tremor state stabilizes, the exogenous environmental parameter intervention module applies an intervention stimulus to the patient and records the baseline timestamp of the intervention stimulus taking effect. The tremor signal determination module uses the reference timestamp as the absolute zero point to simultaneously acquire the autonomic nervous system stress response time series and the kinematic agitation response time series, calculates the difference between the peak values ​​of the autonomic nervous system stress response time series and the kinematic agitation response time series, and determines the source of tremor energy within the time period. The tremor signal decoupling module uses the autonomic nerve stress response time series as a priori guide vector to strip off the components that are synchronized with the prior guide vector and output a pure organic pathological tremor signal.

13. The automated vibration detection process control system according to claim 1, characterized in that, The cloud management platform includes a patient information management module, a data acquisition and processing module, and a data analysis and reporting module; The patient information management module centrally stores basic patient information, historical test records and diagnostic conclusions, and supports multi-center data sharing, access control and data de-identification export. The data acquisition and processing module receives multimodal data uploaded from each testing room, performs secondary verification and format standardization, and stores the data in a distributed time-series database after sensor noise removal and video transcoding preprocessing. The data analysis and reporting module has a built-in multi-scale grading algorithm library, which is used to automatically extract tremor frequency, amplitude, direction and movement completion features, generate local tremor grading scores for each part and a total comprehensive grading score for the whole body, combine the patient's historical data to draw tremor parameter time-series change curves, and output standardized test reports.

14. An automated vibration detection process control method, applied to the aforementioned automated vibration detection process control system, characterized in that, Including the following steps: Enter the patient's basic information; Based on the patient's basic information, a finite state machine is used to realize the state management and automated operation of the entire tremor detection process. The patient's image is acquired and the patient's height data is calculated. Based on the height data, the height and angle of the motion capture camera and the position of the electric seat are adaptively adjusted to adapt to the current detection scene. The visibility confidence score of each key point is output based on the skeleton key point detection network and graph neural network. When the visibility confidence score is lower than the preset visibility threshold and the duration exceeds the preset time threshold, and the key point belongs to the key area of ​​the current detection scene, it is considered effective occlusion. The timing of the current scene and data acquisition are paused, and an adjustment prompt is issued. After the occlusion is removed, the detection is resumed and only the occluded segment is re-acquired. The patient's multimodal data is collected synchronously and preprocessed to output effective multimodal data. The effective multimodal data is processed, and tremor features are extracted based on clinical scales to generate a standardized test report; Display standardized testing reports.