Intelligent visual training collaborative diagnosis and treatment method and platform based on multi-modal perception and 5g communication

The intelligent vision training system, which utilizes multimodal perception and 5G communication, collects and analyzes user behavior data in real time to generate personalized training tasks. This solves the problems of insufficient data collection and lack of remote collaboration in traditional training systems, enabling refined and real-time collaborative diagnosis and treatment during the training process.

CN122117237APending Publication Date: 2026-05-29GUANGDONG NO 2 PROVINCIAL PEOPLES HOSPITAL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG NO 2 PROVINCIAL PEOPLES HOSPITAL
Filing Date
2026-02-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing behavioral vision training systems lack the ability to collect multimodal behavioral data, making it impossible to quantify the training process. This results in training plans that are difficult to personalize, training effects that are untraceable, and a lack of remote doctor intervention and dynamic adjustment of training tasks, leading to insufficient matching between the training pace and the user's capabilities.

Method used

By using multimodal perception and 5G communication, the system collects the user's eye movement trajectory, spatial movement path, and response time in real time, generates training data packets, and sends them to a remote doctor's terminal via the 5G network. Combined with the doctor's intentions, the system generates task control vectors, dynamically matches training tasks, and executes them by the robot, thus realizing personalized and remote collaboration in the training process.

Benefits of technology

It achieves refined identification, personalized scheduling, and real-time collaboration in the training process, improving the controllability and continuity of training results. It is suitable for multimodal intelligent collaborative training in hospitals, rehabilitation centers, and home settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117237A_ABST
    Figure CN122117237A_ABST
Patent Text Reader

Abstract

The application provides an intelligent visual training collaborative diagnosis and treatment method and platform based on multi-modal perception and 5G communication, which comprises the following steps: collecting the eye movement trajectory, spatial action path and user response time of a user in real time, calculating the fixation stability index based on the eye movement trajectory data, and generating a training data package; dividing the training task type and extracting the corresponding behavior performance index of each stage to calculate the perception stage performance score and the cognitive stage performance score; sending the user state snapshot to the remote doctor terminal through the 5G communication network, combining the doctor's regulation intention, and generating a task control vector; according to the training state vector and the task control vector, dynamically matching and combining the task parameters suitable for the current ability of the user and the regulation intention of the doctor from a preset task component library, and analyzing the task parameters into operation instruction streams by a robot guide module. The application can be applied to a multi-modal intelligent collaborative training system suitable for hospital, rehabilitation center and family scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent vision training and collaborative diagnosis and treatment, and particularly relates to an intelligent vision training and collaborative diagnosis and treatment method and platform based on multimodal perception and 5G communication. Background Technology

[0002] Behavioral visual training relies on neuroplasticity mechanisms to improve the visual system's processing capabilities in real-world task scenarios through repeated training in gaze control, pattern recognition, dynamic tracking, spatial localization, and motor response. It has been widely used in the rehabilitation of amblyopia, binocular vision abnormalities, visual cognitive impairments, and learning-related visual problems. However, traditional training methods heavily rely on real-time guidance and manual judgment from therapists. Training tools are mostly based on fixed cards, static exercises, or simple gamified software, making it difficult to reflect the dynamic performance of users in the "seeing, understanding, and moving" stages, and also unable to quantify the location of obstacles at each stage of training. Existing technologies generally lack the ability to collect multimodal behavioral data such as gaze trajectories, movement paths, and reaction times throughout the entire process, making it difficult for training systems to form a unified behavioral structure, resulting in a lack of personalized training plans and untraceable training effects.

[0003] Meanwhile, because behavioral vision training typically requires continuous training over periods of weeks to months, users lack professional guidance when training at home. Existing remote training systems only support uploading training results or simple video feedback, lacking a structured understanding of the real training process and thus failing to provide doctors with effective diagnostic information. Furthermore, current training platforms generally use static task libraries, unable to dynamically generate adaptive training tasks based on user performance at different stages, nor can they leverage telemedicine resources to collaboratively adjust training strategies. This results in a significant disconnect between home training and clinical guidance, making it difficult to accurately match training pace, task difficulty, and user capabilities over the long term.

[0004] Therefore, there is an urgent need for an intelligent behavioral vision training system that can collect multimodal behavioral data in real time during training, accurately characterize the differences in user performance at different stages of the visual perception and cognition-action chain, enable structured intervention by remote doctors through high-speed communication networks, and automatically generate training tasks corresponding to user capabilities for collaborative execution by robots, in order to solve the systemic deficiencies of existing solutions in data collection, state recognition, dynamic task control, and remote collaboration. Summary of the Invention

[0005] The purpose of this invention is to propose an intelligent vision training and collaborative diagnosis and treatment method and platform based on multimodal perception and 5G communication to solve the above-mentioned problems.

[0006] To achieve the above objectives, a first aspect of the present invention provides an intelligent vision training-based collaborative diagnosis and treatment method based on multimodal perception and 5G communication, the method comprising the following steps: Step S1: During the user's visual training task, the user's eye movement trajectory, spatial movement path and user response time are collected in real time, and the gaze stability index is calculated based on the eye movement trajectory data. The task parameters, gaze stability index, eye movement trajectory, spatial movement path and user response time are packaged to generate a training data package. Step S2: Based on the training data package, the training task type is divided into perception stage, cognition stage and execution stage, and the corresponding behavioral performance indicators of each stage are extracted to calculate the performance score of perception stage and the performance score of cognition stage, and then combined to generate training state vector. Step S3: Send a snapshot of the user's state, including the training state vector and training data packet, to the remote doctor's terminal via the 5G communication network, and generate a task control vector by combining the doctor's control intention input based on the visual interface; wherein, the adjustment intensity factor of the task control vector is adaptively adjusted according to the difference between the user's current state and the target state; Step S4: Based on the training state vector and task control vector, dynamically match and combine task parameters from the preset task component library to generate task parameters that are adapted to the user's current ability and the doctor's control intention. The robot guidance module then parses the task parameters into an operation instruction stream to guide the user to complete the training task, while collecting a new round of behavioral data.

[0007] Furthermore, the user response time is the difference between the time when the target is first displayed and the time when the user first makes a valid response. The gaze stability index is generated by calculating the Euclidean norm based on eye movement trajectory data and the average coordinate position of all gaze points during the task; the Euclidean norm is used to measure the deviation of each gaze point from the average position.

[0008] Furthermore, the behavioral performance indicators in the perception stage are eye movement trajectory and gaze stability indicators; the behavioral performance indicators in the cognition stage are user response time and action path start-point offset; and the behavioral performance indicators in the execution stage are path offset sequences. The offset of the starting point of the action path is obtained by comparing the initial segment of the spatial action path with the optimal starting action position set in the task parameters. The path offset sequence is used to evaluate whether the action execution deviates from the preset route.

[0009] Furthermore, the performance score in the perception stage is generated by combining the gaze stability index, the visual complexity of the task setting, and the variance of target changes in the task; the performance score in the cognition stage is generated by combining the user response time, the offset of the starting point of the action path, and the coefficient of the task's requirement for reaction speed. The visual complexity of the task is obtained by linearly transforming the number of interfering objects and the target's change rate; the variance of the target's change in the task is generated by analyzing the target's motion amplitude and speed changes.

[0010] Furthermore, when combining the doctor's control intentions input based on the visual interface, the doctor's intention commands are mapped into three control dimensions: Visual stimulus complexity factor, path structure adjustment factor, and cue rhythm adjustment factor; The visual stimulus complexity factor corresponds to the image recognition difficulty in the task, the path structure adjustment factor is used to control the number of targets and path length in the spatial task, and the prompt rhythm adjustment factor is used to control the trigger frequency of robot voice or large screen prompts.

[0011] Furthermore, the task control vector is calculated and generated based on the visual stimulus complexity factor, path structure adjustment factor, cue rhythm adjustment factor, and adjustment intensity factor. The task control vector is mapped to the doctor's control intention. Figure 1 In this way, visual stimulus complexity factor, path structure adjustment factor, and cue rhythm adjustment factor are generated as the final task control vector based on the task control vector.

[0012] Furthermore, the adjustment intensity factor is generated based on the performance scores of the perception stage, the performance scores of the cognition stage, the doctor's expected target score for the user, and the variance of the structural complexity index of the current task. The adjustment intensity factor is used to ensure that the target strategy input by the doctor is not out of sync with the user's current ability, thus preventing training blockage or efficiency decline. The variance of the structural complexity index of the current task is used to suppress the problem of control changes being too rapid in complex tasks.

[0013] Furthermore, the preset task component library includes three types of task elements: visual target units, spatial path units, and prompting interaction units.

[0014] Furthermore, the process of dynamically matching and combining task parameters from a preset task component library to generate task parameters that adapt to the user's current capabilities and the doctor's control intentions, and then having the robot guidance module parse these task parameters into an operation instruction stream, specifically involves: A fit scoring function is designed to calculate a fit score to assess the degree of matching between the current state and the task element; wherein, the fit scoring function is composed of the performance scores of the perception stage and the cognitive stage, as well as the structural repeatability measure; the task element is used to penalize the task element with too high repetition in the past 3 rounds of the task; The task components with the highest adaptation scores are selected to generate task parameters. After the robot guidance module is started, the received task parameters are parsed into an operation instruction stream, and user behavior data is collected synchronously from multiple channels during the execution process.

[0015] In a second aspect, the present invention provides an intelligent vision training and collaborative diagnosis platform based on multimodal perception and 5G communication, the platform comprising: The multimodal behavior acquisition module is used to collect the user's eye movement trajectory, spatial movement path and user response time in real time during the user's visual training task, and calculate the gaze stability index based on the eye movement trajectory data. The task parameters, gaze stability index, eye movement trajectory, spatial movement path and user response time are packaged to generate a training data package. The cognitive path modeling module is used to divide the training task type into a perception stage, a cognition stage and an execution stage based on the training data package, and extract the corresponding behavioral performance indicators for each stage to calculate the performance scores of the perception stage and the cognition stage, and combine them to generate a training state vector. The 5G collaborative diagnosis and treatment module is used to send a snapshot of the user's state, including the training state vector and training data packet, to a remote doctor's terminal via a 5G communication network, and generate a task control vector by combining the doctor's control intentions input based on a visual interface; wherein, the adjustment intensity factor of the task control vector is adaptively adjusted according to the difference between the user's current state and the target state. The task generation and execution control module is used to dynamically match and combine task parameters from a preset task component library based on the training state vector and task control vector to generate task parameters that are adapted to the user's current ability and the doctor's control intention. The robot guidance module parses the task parameters into an operation instruction stream to guide the user to complete the training task, while collecting a new round of behavioral data.

[0016] The beneficial technical effects of the present invention are at least as follows: This invention proposes an intelligent vision training method and platform based on multimodal behavior acquisition, cognitive path modeling, 5G collaborative diagnosis and treatment, and robot training execution. By acquiring the user's gaze trajectory, action path, and reaction latency throughout the training process, structured training data is constructed. Based on this, a cognitive path model is established that can distinguish the differences between the three stages of visual perception, cognitive judgment, and action execution. A training state vector reflecting the user's true ability state is generated. The state vector and key behavioral data are synchronized to a remote doctor using high-speed, low-latency 5G communication. The doctor inputs the training strategy intent, which is then mapped into a quantified control vector by the system, thereby enabling remote control of task difficulty, visual stimulus complexity, spatial path structure, and robot prompting rhythm. This invention further proposes a task generation mechanism jointly driven by state and control vectors. Through structured combination of task component libraries, training tasks matching the user's current ability are dynamically generated. The robot execution control module guides the user through voice, action, light, or path guidance, forming a closed-loop chain from "data acquisition—state recognition—remote diagnosis and treatment—task generation—training execution." This invention breaks through the technical bottlenecks of traditional behavioral vision training, which relies on human experience, has fixed tasks, lacks process data, and cannot be remotely intervened. It achieves refined recognition during the training phase, personalized scheduling of training tasks, and real-time training-diagnosis collaboration, making the training process closer to real-life scenarios and significantly improving the controllability and continuity of training effects. It is applicable to multimodal intelligent collaborative training systems in hospitals, rehabilitation centers, and home settings. Attached Figure Description

[0017] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0018] Figure 1 This is a flowchart of the intelligent vision training and collaborative diagnosis method based on multimodal perception and 5G communication of the present invention. Detailed Implementation

[0019] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0020] like Figure 1 As shown in the embodiment of the present invention, the intelligent vision training and collaborative diagnosis method based on multimodal perception and 5G communication provides the following method: Step S1: During the user's visual training task, the user's eye movement trajectory, spatial movement path and user response time are collected in real time, and the gaze stability index is calculated based on the eye movement trajectory data. The task parameters, gaze stability index, eye movement trajectory, spatial movement path and user response time are packaged to generate a training data package. Step S2: Based on the training data package, the training task type is divided into perception stage, cognition stage and execution stage, and the corresponding behavioral performance indicators of each stage are extracted to calculate the performance score of perception stage and the performance score of cognition stage, and then combined to generate training state vector. Step S3: Send a snapshot of the user's state, including the training state vector and training data packet, to the remote doctor's terminal via the 5G communication network, and generate a task control vector by combining the doctor's control intention input based on the visual interface; wherein, the adjustment intensity factor of the task control vector is adaptively adjusted according to the difference between the user's current state and the target state; Step S4: Based on the training state vector and task control vector, dynamically match and combine task parameters from the preset task component library to generate task parameters that are adapted to the user's current ability and the doctor's control intention. The robot guidance module then parses the task parameters into an operation instruction stream to guide the user to complete the training task, while collecting a new round of behavioral data.

[0021] Specifically, step S1 includes: The goal of this step is to collect and encode the core behavioral performance information of users during the behavioral vision training process in a structured manner into data packets, providing basic data support for the subsequent cognitive path modeling module. The system then uses the configuration parameters of the training task... The data dimensions and sampling parameters to be observed in this round of training are set, and the data is collected and stored in real time through multi-source sensing devices during the task execution.

[0022] Before training begins, the system receives data from the task control module. This parameter describes the basic type of the training task (e.g., gaze tracking, pattern recognition, spatial walking), execution method (e.g., gesture response, gait triggering), number of targets, stimulus velocity, and complexity of distractors. For example, in a typical "visual target tracking and spatial response" training exercise, The system instructs the user to continuously track a target with changing color under dynamic image interference, and then quickly move a distance towards the target after it comes to a standstill to complete the task. Under this configuration, the system needs to simultaneously collect the following three types of key behavioral data.

[0023] First, eye-tracking data The data acquisition is performed by an infrared eye-tracking module located below the display screen of the training terminal. This module uses binocular stereo pupil tracking technology, combined with screen calibration points to generate a sequence of two-dimensional coordinate points, which are sampled at a frequency of 60 Hz during task execution, with the sampling period equal to the task duration. The data format is a sequence. Each item represents a screen coordinate point. It is used to reconstruct the changing trend of the user's gaze point during the switching of visual targets.

[0024] Secondly, spatial motion paths This is primarily achieved through the coordinated operation of binocular stereo vision cameras deployed at the four corners of the training space and inertial motion tracking ankle bracelets worn on the user's lower limbs. The vision system identifies the user's position and synchronizes with the inertial sensors for calibration, forming a sequence of the user's spatial movement paths indexed by timestamps. Each frame records the user's spatial coordinates and relative velocity This data is primarily used to assess a user's ability to effectively translate visual information into spatial positioning and motor execution capabilities.

[0025] The third data point is user response time. This is generated by an internal system trigger mechanism. In each training round, the system automatically records the time point when the target first appears. and the time point of the user's first effective response. If a user clicks, confirms with voice, or raises their hand or moves their foot to reach a trigger threshold, the system calculates: ; This indicates the reaction time for this task. Indicates response time; This indicates the time when the system detected the user's first completed response; This indicates the start time of the task stimulus set by the system.

[0026] Furthermore, during the data acquisition process, the system tracks the gaze trajectory. Calculate gaze stability index This is used to reflect whether the user frequently deviates from the target or experiences gaze skipping during fixation. The specific formula is as follows: ; in, Indicates the first The gaze coordinates of the frame; This represents the average coordinate position of all gaze points during the task; Indicates the total number of frames sampled during the task; This represents the Euclidean norm, used to measure the deviation of the gaze point from its average position in each frame.

[0027] Here's a specific operational procedure: In a training task primarily involving image recognition and spatial movement, the user needs to identify three targets of different colors against a dynamic image background and quickly move them to the target locations that match the system's voice prompts. After the task starts, the system immediately activates the eye-tracking and spatial motion acquisition modules, and... The system samples at a pre-defined pace (e.g., a 10-second task duration). After the task, the system saves approximately 600 frames of eye-tracking data, 45 spatial location nodes, and records a response time of 2.3 seconds. The system then automatically calculates... And match these data with the task number Package them together into a single data structure The training data package Includes five fields: Task Parameters Eye movement trajectory gaze stability index Spatial path Response latency .

[0028] Specifically, step S2 includes: This step is used to process the training data package output in the previous step. Based on this, a cognitive path model is constructed that can realistically reflect the differences in user performance in the three stages of "perception-cognition-execution" during behavioral vision training, and the training state vector is output. The core innovation of this step lies in introducing a task-dependent stage-weighted mechanism, which couples user performance with task structure for evaluation, so that the final output can more accurately reflect the obstacle locations in real training.

[0029] The input is the training data packet. This includes task parameters. Eye movement trajectory gaze stability Spatial motion path and response time These data reflect the user's visual concentration, motion planning ability, and reaction speed during training, and are essential foundational information for building a path model. This step first considers the task parameters... Three time boundaries are defined to ensure that the behavioral process is divided into corresponding cognitive stages. For example, in the task of "performing spatial actions after dynamic visual tracking," the period from the first appearance of the target to the user's gaze stabilizing is the perception stage, the period from gaze stabilization to the start of action planning is the cognitive stage, and the period from action initiation to task completion is the execution stage. The stage boundaries are automatically set by the system when the task starts, ensuring that different tasks have a consistent structural definition in the overall model.

[0030] Furthermore, after dividing the training into stages, this step extracts the performance metrics most relevant to behavioral visual training for each stage. The perception stage primarily relies on gaze trajectories. With gaze stability The cognitive stage uses user response time to determine whether the user can maintain a stable gaze in a distracting background. Offset from the start of the action path The latter, through comparison The initial segment and The optimal starting position is obtained from the set position; during the execution phase, the path offset sequence is used. This step assesses whether the action execution deviates from the preset route. Since the goal of this step is to generate a vectorized overall performance, the stage results will not be directly output, but will be used to construct a comprehensive performance score.

[0031] In the scoring phase, this step introduces an innovative regularization term that incorporates the characteristics of behavioral visual training. This ensures that the formula reflects not only the user's physical performance but also the "difficulty contribution" within the task structure, thereby enhancing the scoring's sensitivity to real-world training conditions. Perception Phase Performance Scoring Adopt the following form: ; in, The fixation stability metric defined in the previous step; The visual complexity set for the task is obtained by linearly transforming the number of distractions and the rate of change of the target. The variance of the target's changes during the task is generated by analyzing the target's motion amplitude and velocity changes. The scaling factor set during task definition is used to adjust the impact of target perturbations on the evaluation. The formula introduces... As the denominator, it makes the high-interference task in the same This avoids excessively punishing users, thereby improving the model's adaptability to different task scenarios.

[0032] Furthermore, cognitive stage performance scores Then take into account the user's response delay Offset from the start of the action path In conjunction with the task's requirements for decision-making speed, it is defined as follows: ; in The response time for the user's first valid action; This is the starting offset of the action path; The coefficient representing the task's required reaction speed is given by... The task rhythm field in the data is converted to obtain the result. , , Weighting coefficients are used when defining tasks to balance the importance of velocity and spatial bias in different tasks. (Introduction) The significance lies in distinguishing between "tasks that allow for slow responses" and "tasks that require fast responses," which is one of the key characteristics of behavioral vision training. It truly reflects whether the user can keep up with the pace of the task, rather than absolute speed.

[0033] Finally, the system combines the scores from the two stages into a training state vector: ; This vector is directly used in the next step's training task generation module as a crucial criterion for adjusting task difficulty, stimulus frequency, and prompting methods. Because the fundamental goal of behavioral vision training is to improve the "see-understand-move" ability chain, this vector structure comprehensively describes the user's behavioral performance at key stages. The output of this step is the training state vector. This includes the user's visual adaptation to task stimuli ( ) and the ability to recognize action conversion ( ).

[0034] Specifically, step S3 includes: This step aims to establish a remote collaborative diagnosis and treatment structure based on 5G communication mechanisms, enabling doctors to formulate task strategy control instructions based on user training performance under non-real-time manual control conditions, and to utilize structured parameters. Implement task adjustments on the intelligent platform.

[0035] The input to this step is the training state vector output from the previous step. These represent the user's performance scores in the perception and cognition phases of the current training task, respectively, and also include the current task structure parameters. And auxiliary state data such as gaze stability Reaction time and path offset index These data all include the structured training data output in steps one and two. The above data is structured and packaged to form a user state snapshot. It is pushed to the remote doctor's terminal via a secure compression protocol through a 5G communication module.

[0036] Furthermore, the doctor's terminal system received Subsequently, a multimodal visualization interface is generated locally, including a gaze trajectory playback map, a path distribution map, and a phased score line graph, to assist doctors in making judgments based on the training process. Unlike static report-based decision-making, this step allows doctors to input control objectives in the form of intent parameters, such as "increase the user's focus task duration" or "reduce path complexity," which the system automatically converts into executable control command vectors. This vector will participate in the task structure adjustment in the subsequent task generation module.

[0037] Furthermore, to ensure the coupling between control commands and user states, this step designs a state-aware adjustment intensity factor. Furthermore, task-dependent regularization is introduced to ensure that the target strategy input by the doctor is not out of sync with the user's current capabilities, thus preventing training blockage or efficiency decline. Adjustment intensity factor The definition is as follows: ; in These are the user's current perception and cognitive scores, respectively. Rate the goals that doctors expect users to achieve; Adjust the weights for deviations in each dimension; The variance is a structural complexity index for the current task, used to suppress the problem of control changes being too rapid in complex tasks and improve the robustness of regulation.

[0038] Furthermore, in generating the final control vector At that time, the system maps the doctor's intended instructions into three adjustment dimensions: visual stimulus complexity factor. Path structure adjustment factor With prompt rhythm adjustment factor .in This corresponds to the difficulty of image recognition in the task (such as the number of images, the rate of change, and the proportion of target occlusion). Controlling the number of targets and path length in space missions. Control the trigger frequency of robot voice or large screen prompts. Final control vector. The definition is as follows: ; in To adjust the global scaling weights of the intensity factor for the system. Introducing... This design ensures that the doctor's instructions are not directly and forcefully applied when the user's current state fluctuates greatly, but are automatically weakened or delayed until the user's state stabilizes. This design reflects the deep coupling between this step and the cognitive path modeling stage.

[0039] Taking a specific scenario as an example, if a user fails in two consecutive rounds of training... Once the value stabilized below 0.6, the doctor wanted to adjust the task to focus more on image focusing rather than spatial movement. Therefore, they selected "Reduce visual distractions and increase image exposure time" in the system interface, and the system automatically mapped it to... And according to the current task and Automatic calculation The value is 0.35, and the final control command is... This command will control the generation of task templates in the next stage. The control command is cached locally by the system and written to the task queue for subsequent training template selection and parameter setting. The task queue supports traceability and strategy version recording, ensuring that the doctor's strategy input is interpretable and rollbackable. 5G communication ensures... The transmission latency is no more than 50 milliseconds, ensuring that instructions can take effect during training intervals and maintaining the system's operational rhythm without interruption. The output of this step is the task control vector. Through with The coupling and differential adjustment mechanism enables remote doctors to not only make personalized judgments based on user performance, but also provide structural interventions at the task generation level, achieving true "collaborative behavioral visual intelligence training".

[0040] Specifically, step S4 includes: This step involves combining the training state vector generated in the previous stage with the task control vector remotely transmitted by the doctor via 5G. Combined as input, dynamically generate training task parameters They also collaborate with a robot guidance system to complete task execution.

[0041] Furthermore, the system first receives the state vector. ,in Eye-track stability from the user's previous training The normalized analysis results reflect its visual focusing ability; Originating from response latency and path offset The combined score reflects the matching between cognitive judgment and spatial execution. Simultaneously, the system receives the final task control vector. ,in As a visual complexity modulator, for example, doctors may want to enhance visual focusing training by increasing the number of targets or shortening the duration of stimuli; This indicates adjustments to the complexity of the execution space, such as the density of path points or the number of orientation changes; This refers to cues and rhythm factors, such as adjusting the delay of voice cues or the dwell time of graphic cues. These control factors are selected and generated by the doctor through an interactive control panel, and then adjusted by the platform based on the intensity of the differences in the patient's condition. The final instruction is formed after adaptive fine-tuning.

[0042] The task generation module combines and matches the above inputs in the local task component library. The task component library includes three types of task elements: visual target units (such as graphic recognition, color interference blocks, etc.), spatial path units (such as spatial pattern navigation, direction recognition, etc.), and prompting interaction units (voice prompt templates, action demonstration templates, etc.). For example, the visual target unit stores 100 preset graphics, including standard geometric figures, children's cognitive graphics, and motion target videos, while the spatial path unit contains commonly used path shapes (such as "L", "Z", and "U" shapes) and their parameter configurations, and the prompting unit contains audio synthesis templates and action guidance videos.

[0043] Furthermore, the system uses a fit scoring function. Calculate the degree of matching between the current state and the task elements: ; in, This indicates the priority weight of perception and cognition indicators in the current training phase. It is a structural repeatability metric used to penalize task elements that have too much repetition with the past 3 rounds of the task. This is the regularization term strength coefficient. The above scoring mechanism reflects both state adaptation and mitigates the risk of task template convergence. The system selects the task elements with the highest scores to generate structured task parameters. It includes fields such as visual element type, stimulus duration, graphic interference density, path point sequence, start and end angle change rate, and maximum allowable path deviation.

[0044] For example, in a dynamic vision training session for a child with amblyopia, if Doctor settings The system will prioritize selecting circular patches with a moderate amount of dynamic perturbation within the visual target and set the stimulus duration to 2 seconds. If simultaneously... and The system will select a straight path task with a length not exceeding 3 meters and no turns, and limit the spatial offset tolerance to 20 centimeters. The system will then set... The robot issues a voice prompt every 4 seconds and indicates the target location through directional gestures.

[0045] Furthermore, after the robot guidance module is activated, it will receive the task parameters. The commands are interpreted as an operational instruction stream. When performing static visual recognition tasks, the robot uses its voice module to announce "Please look at the red target," simultaneously highlighting the target's location on the screen and playing a 3-second countdown voice message before the target disappears. In spatial pathfinding tasks, the robot generates a trajectory based on path points, adjusts its pace using its built-in stepping control system, and provides directional hand gestures at each corner. Furthermore, the robot projects its path onto the ground using a laser-assisted system, accompanied by the prompt "Please walk towards the target along the direction of the light," enhancing visual-motor guidance.

[0046] User behavior data is collected simultaneously from multiple sources during execution. Gaze points are captured by a binocular infrared tracker at the screen edge with a 60Hz sampling rate; motion trajectories are tracked in real-time by a ceiling-mounted camera and corrected in real-time by a wearable inertial measurement unit (IMU), with path offset calculated to an accuracy of 5 centimeters. Voice responses are triggered by a microphone array, and keyword recognition is used to determine if the user is responding correctly.

[0047] This invention also provides an intelligent vision training and collaborative diagnosis platform based on multimodal perception and 5G communication, the platform comprising: The multimodal behavior acquisition module is used to collect the user's eye movement trajectory, spatial movement path and user response time in real time during the user's visual training task, and calculate the gaze stability index based on the eye movement trajectory data. The task parameters, gaze stability index, eye movement trajectory, spatial movement path and user response time are packaged to generate a training data package. The cognitive path modeling module is used to divide the training task type into a perception stage, a cognition stage and an execution stage based on the training data package, and extract the corresponding behavioral performance indicators for each stage to calculate the performance scores of the perception stage and the cognition stage, and combine them to generate a training state vector. The 5G collaborative diagnosis and treatment module is used to send a snapshot of the user's state, including the training state vector and training data packet, to a remote doctor's terminal via a 5G communication network, and generate a task control vector by combining the doctor's control intentions input based on a visual interface; wherein, the adjustment intensity factor of the task control vector is adaptively adjusted according to the difference between the user's current state and the target state. The task generation and execution control module is used to dynamically match and combine task parameters from a preset task component library based on the training state vector and task control vector to generate task parameters that are adapted to the user's current ability and the doctor's control intention. The robot guidance module parses the task parameters into an operation instruction stream to guide the user to complete the training task, while collecting a new round of behavioral data.

[0048] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for intelligent visual training and collaborative diagnosis based on multimodal perception and 5G communication, characterized in that, The method includes: Step S1: During the user's visual training task, the user's eye movement trajectory, spatial movement path and user response time are collected in real time, and the gaze stability index is calculated based on the eye movement trajectory data. The task parameters, gaze stability index, eye movement trajectory, spatial movement path and user response time are packaged to generate a training data package. Step S2: Based on the training data package, the training task type is divided into perception stage, cognition stage and execution stage, and the corresponding behavioral performance indicators of each stage are extracted to calculate the performance score of perception stage and the performance score of cognition stage, and then combined to generate training state vector. Step S3: Send a snapshot of the user's state, including the training state vector and training data packet, to the remote doctor's terminal via the 5G communication network, and generate a task control vector by combining the doctor's control intention input based on the visual interface; wherein, the adjustment intensity factor of the task control vector is adaptively adjusted according to the difference between the user's current state and the target state; Step S4: Based on the training state vector and task control vector, dynamically match and combine task parameters from the preset task component library to generate task parameters that are adapted to the user's current ability and the doctor's control intention. The robot guidance module then parses the task parameters into an operation instruction stream to guide the user to complete the training task, while collecting a new round of behavioral data.

2. The intelligent visual training and collaborative diagnosis method based on multimodal perception and 5G communication according to claim 1, characterized in that, The user response time is the difference between the time when the target is first displayed and the time when the user first makes a valid response. The gaze stability index is generated by calculating the Euclidean norm based on eye movement trajectory data and the average coordinate position of all gaze points during the task; the Euclidean norm is used to measure the deviation of each gaze point from the average position.

3. The intelligent visual training and collaborative diagnosis method based on multimodal perception and 5G communication according to claim 1, characterized in that, The behavioral performance indicators for the perception stage are eye movement trajectory and gaze stability; the behavioral performance indicators for the cognition stage are user response time and action path origin offset; and the behavioral performance indicator for the execution stage is path offset sequence. The offset of the starting point of the action path is obtained by comparing the initial segment of the spatial action path with the optimal starting action position set in the task parameters. The path offset sequence is used to evaluate whether the action execution deviates from the preset route.

4. The intelligent visual training and collaborative diagnosis method based on multimodal perception and 5G communication according to claim 3, characterized in that, The performance score for the perception stage is calculated by combining the gaze stability index, the visual complexity of the task setting, and the variance of the target changes in the task; the performance score for the cognition stage is calculated by combining the user response time, the offset of the starting point of the action path, and the coefficient of the task's requirement for reaction speed. The visual complexity of the task is obtained by linearly transforming the number of interfering objects and the target's change rate; the variance of the target's change in the task is generated by analyzing the target's motion amplitude and speed changes.

5. The intelligent visual training and collaborative diagnosis method based on multimodal perception and 5G communication according to claim 1, characterized in that, When combining the doctor's control intentions input based on the visual interface, the doctor's intention commands are mapped into three control dimensions: Visual stimulus complexity factor, path structure adjustment factor, and cue rhythm adjustment factor; The visual stimulus complexity factor corresponds to the image recognition difficulty in the task, the path structure adjustment factor is used to control the number of targets and path length in the spatial task, and the prompt rhythm adjustment factor is used to control the trigger frequency of robot voice or large screen prompts.

6. The intelligent visual training and collaborative diagnosis method based on multimodal perception and 5G communication according to claim 5, characterized in that, The task control vector is calculated and generated based on the visual stimulus complexity factor, path structure adjustment factor, cue rhythm adjustment factor, and adjustment intensity factor. The task control vector is mapped to the same dimension as the doctor's regulatory intention, and a visual stimulus complexity factor, path structure adjustment factor, and cue rhythm adjustment factor based on the task control vector are generated as the final task control vector.

7. The intelligent visual training and collaborative diagnosis method based on multimodal perception and 5G communication according to claim 1, characterized in that, The adjustment intensity factor is generated based on the performance scores of the perception stage, the performance scores of the cognition stage, the doctor's expected target score for the user, and the variance of the structural complexity index of the current task. The adjustment intensity factor is used to ensure that the target strategy input by the doctor is not out of sync with the user's current ability, thus preventing training blockage or efficiency decline. The variance of the structural complexity index of the current task is used to suppress the problem of control changes being too rapid in complex tasks.

8. The intelligent visual training and collaborative diagnosis method based on multimodal perception and 5G communication according to claim 1, characterized in that, The preset task component library includes three types of task elements: visual target units, spatial path units, and prompting interaction units.

9. The intelligent visual training and collaborative diagnosis method based on multimodal perception and 5G communication according to claim 8, characterized in that, The process involves dynamically matching and combining task parameters from a pre-defined task component library to generate task parameters that are adapted to the user's current capabilities and the doctor's control intentions. The robot guidance module then parses these task parameters into an operation command stream, specifically: A fit scoring function is designed to calculate a fit score to assess the degree of matching between the current state and the task element; wherein, the fit scoring function is composed of the performance scores of the perception stage and the cognitive stage, as well as the structural repeatability measure; the task element is used to penalize the task element with too high repetition in the past 3 rounds of the task; The task components with the highest adaptation scores are selected to generate task parameters. After the robot guidance module is started, the received task parameters are parsed into an operation instruction stream, and user behavior data is collected synchronously from multiple channels during the execution process.

10. An intelligent vision training and collaborative diagnosis platform based on multimodal perception and 5G communication, characterized in that: The platform includes: The multimodal behavior acquisition module is used to collect the user's eye movement trajectory, spatial movement path and user response time in real time during the user's visual training task, and calculate the gaze stability index based on the eye movement trajectory data. The task parameters, gaze stability index, eye movement trajectory, spatial movement path and user response time are packaged to generate a training data package. The cognitive path modeling module is used to divide the training task type into a perception stage, a cognition stage and an execution stage based on the training data package, and extract the corresponding behavioral performance indicators for each stage to calculate the performance scores of the perception stage and the cognition stage, and combine them to generate a training state vector. The 5G collaborative diagnosis and treatment module is used to send a snapshot of the user's state, including the training state vector and training data packet, to a remote doctor's terminal via a 5G communication network, and generate a task control vector by combining the doctor's control intentions input based on a visual interface; wherein, the adjustment intensity factor of the task control vector is adaptively adjusted according to the difference between the user's current state and the target state. The task generation and execution control module is used to dynamically match and combine task parameters from a preset task component library based on the training state vector and task control vector to generate task parameters that are adapted to the user's current ability and the doctor's control intention. The robot guidance module parses the task parameters into an operation instruction stream to guide the user to complete the training task, while collecting a new round of behavioral data.