Millimeter wave radar rehabilitation exercise recognition method and system based on knowledge distillation

Through knowledge distillation-based millimeter wave radar technology, combined with image and point cloud data, the problem of insufficient effect of existing rehabilitation motion recognition methods in complex environments is solved, efficient and real-time rehabilitation motion recognition is achieved, and lightweight deployment is carried out to edge devices.

CN120164259APending Publication Date: 2025-06-17NANHUA UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510280691.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing rehabilitation exercise recognition methods are not effective in scenarios such as insufficient light, overlapping multiple objects, and complex backgrounds. Visual recognition of the human skeleton consumes a lot of computing resources and is difficult to deploy to edge devices.

Method used

The millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation is adopted. By acquiring image data sets and point cloud data sets, training is used for online bidirectional distillation, parameters are updated alternately, and parameters of the teacher model are finally frozen, and the student model is guided to perform offline distillation and optimization, and deployed in edge computing equipment.

Benefits of technology

It improves the efficiency of human skeleton recognition, improves the real-time and efficiency of rehabilitation exercise recognition, solves the problems of poor environmental adaptability and high computing resource consumption, and realizes lightweight deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164259A_ABST
    Figure CN120164259A_ABST
Patent Text Reader

Abstract

The invention provides a millimeter wave radar rehabilitation exercise recognition method and system based on knowledge distillation. The method comprises the steps that an image data set and a point cloud data set are acquired; training the image data set by using a teacher model A; training the point cloud data set by using a teacher model B; performing online bidirectional distillation on the teacher model A and the teacher model B, alternately updating parameters of the teacher model A and the teacher model B, and outputting the teacher model B; freezing parameters of the teacher model B, starting off-line distillation, guiding student model training, and optimizing the student model; and acquiring patient information, inputting the patient information into the student model, and outputting a millimeter wave radar rehabilitation exercise recognition result. According to the method, a knowledge distillation method is combined, a high-precision and complex neural network model is used as a teacher model, a lightweight student model is guided to be trained, and finally the lightweight student model is deployed in an edge computing device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rehabilitation exercises, in particular to a millimeter-wave radar rehabilitation exercise recognition method and system based on knowledge distillation. Background Art

[0002] Currently, in the field of rehabilitation exercise recognition, the combination of sensors and computer vision technology is an emerging topic. Compared with traditional rehabilitation exercise recognition methods, inertial measurement unit (IMU) sensors and computer vision (CV) have higher accuracy, objectivity, and personalization, can provide real-time feedback and adjustment, and can better improve the training effect of patients. At the same time, remote monitoring and intelligent management also make the rehabilitation process more efficient and convenient. However, in some cases, such as scenes with insufficient light, multi-object overlap, complex background, etc., or due to users' concerns about privacy, the applicability of computer vision is greatly reduced, and visual recognition of the human skeleton consumes a large amount of computer resources and is difficult to be lightweight deployed to edge devices.

[0003] Traditional rehabilitation exercise recognition methods mostly rely on doctors' observations and manual measurements, emphasizing qualitative evaluation and manual records. These methods can help understand the overall movement ability of patients, but they usually lack accurate quantitative data and may be affected by the experience and subjective factors of operators. With the development of technology, more modern methods based on sensors and intelligent devices have gradually replaced these traditional means, providing more accurate and real-time motion analysis and feedback.

[0004] Current rehabilitation exercise recognition methods include computer vision, wearable sensors, inertial motion capture, remote monitoring and evaluation, etc. Although these technologies are effective to a certain extent, they generally have some problems, including being sensitive to environmental conditions, methods based on optical sensors being difficult to achieve ideal results in strong light or poor lighting environments, the frequent wearing of wearable devices being relatively inconvenient, possibly infringing on privacy, and there still being problems with the real-time performance of vision technology, etc.

[0005] The invention patent application with application publication number CN117669693A discloses a knowledge distillation method and system based on a multi-teacher multimodal model. The present invention performs multimodal knowledge distillation to a student model through a plurality of teacher models. These teacher models have different architectures, initializations, training data or tasks. This diversity helps to extract knowledge from different angles and types, thereby improving the robustness of the student model and the ability to understand images, texts and graphic multimodalities, and improving the accuracy of image recognition, the accuracy of text understanding, and the recall rate and accuracy of multimodal retrieval. The disadvantages of this method are that the joint training of multiple teacher models requires a large amount of computing resources, and the consistency of the structure of the student model and the teacher model may result in too many parameters, which cannot meet the needs of a lightweight student model; the performance of the multimodal fusion device is highly dependent on the output quality of the encoder, and the design is complex, which may offset the lightweight advantage brought by distillation; the joint distillation of multiple teachers involves more hyperparameters, requires complex parameter adjustment strategies and experimental verification, and the optimization difficulty is significantly increased. Summary of the invention

[0006] In order to solve the above technical problems, the present invention proposes a millimeter-wave radar rehabilitation motion recognition method and system based on knowledge distillation. Combined with the knowledge distillation method, a high-precision and complex neural network model is used as a teacher model to guide the training of a lightweight student model, and finally the lightweight student model is deployed in the edge computing device.

[0007] The first object of the present invention is to provide a millimeter wave radar rehabilitation motion recognition method based on knowledge distillation, comprising acquiring an image data set and a point cloud data set, characterized in that it also includes the following steps:

[0008] Step 1: Use teacher model A to train the image dataset;

[0009] Step 2: Use teacher model B to train the point cloud dataset;

[0010] Step 3: The teacher model A and the teacher model B are bidirectionally distilled online, the parameters of the teacher model A and the teacher model B are updated alternately, and the teacher model B is output;

[0011] Step 4: Freeze the parameters of the teacher model B, start offline distillation, guide the student model training, and optimize the student model;

[0012] Step 5: Obtain patient information, input it into the student model, and output the millimeter-wave radar rehabilitation movement recognition result.

[0013] Preferably, the online bidirectional distillation means that the teacher model A and the teacher model B are teacher-student models of each other, and the teacher model A and the teacher model B are bidirectionally semantically decoupled.

[0014] Preferably, in any of the above solutions, the online two-way distillation includes the following sub-steps:

[0015] Step 31: Build a symmetric network structure, allowing the teacher model A and the teacher model B to be teachers and students of each other. At time step t, model A guides model B, and at time step t+1, model B guides model A;

[0016] Step 32: Design a shared dynamic feature cache pool to store the latest N batches of cross-modal feature pairs in real time;

[0017] Step 33: The point cloud semantics extracted by the teacher model B generate pseudo-image labels through a cross-modal classifier, and the image features of the teacher model A and the pseudo-image labels are used to calculate the contrast loss L cont and the SSIM loss SSIM(x,y), forcing the teacher model A to understand the point cloud semantics;

[0018] Step 34: When the teacher model B is the teacher of the teacher model A, use a point cloud network for high-order semantic extraction and map the high-order semantics to the feature space;

[0019] Step 35: The gating network generates two-way distillation weights based on the modal quality assessment of the current batch of data;

[0020] Step 36: The teacher model A is updated based on the contrast loss provided by the teacher model B, and the teacher model B is updated based on the contrast loss L cont provided by the teacher model A and the SSIM loss SSIM(x,y);

[0021] Step 37: After the model is updated, output the teacher model B.

[0022] Preferably, in any of the above solutions, the contrast loss L cont is calculated as

[0023]

[0024] where f i image is the feature vector extracted by the teacher model A for the i-th sample, is the pseudo-label feature generated by the teacher model B, P(i) is the set of pseudo-label features belonging to the same category as the i-th sample, A(i) is all the pseudo-label features in the current batch, τ is the temperature coefficient, and p is the sample belonging to the same category as the current sample i.

[0025] Preferably, in any of the above solutions, the SSIM loss SSIM(x,y) is calculated as

[0026]

[0027] Among them, μ x and μ y are the local means of x and y, and are the local variances of x and y, σ xy is the covariance between x and y, is the local variance of the image x, is the local variance of the pseudo-image label y, C1 and C2 are constant terms, x is the image data, and y is the pseudo-image label.

[0028] Preferably, in any of the above solutions, the bidirectional distillation weight includes the distillation intensity ω A->B for controlling the distillation from teacher model A to teacher model B and the distillation intensity ω B->A .

[0029] Preferably, in any of the above solutions, the distillation intensity ω A->B for controlling the distillation from teacher model A to teacher model B is calculated as:

[0030] ω A->B =σ(f gate (I qualitu ,P density ))

[0031] Among them, ω A->B represents that model A guides model B for distillation, σ is the activation function, f gate is the gating network, I qualitu is the image quality evaluation index, which is evaluated according to the image sharpness, and P density is the point cloud quality evaluation index.

[0032] Preferably, in any of the above solutions, the distillation intensity ω B->A for controlling the distillation from teacher model B to teacher model A is calculated as:

[0033] ω B->A =σ(f gate (I qualitu ,P density ))

[0034] Among them, ω B->A represents that model B guides model A for distillation.

[0035] Preferably, in any of the above solutions, in step 4, teacher model B only serves as a teacher model to provide soft labels for the student model.

[0036] Preferably, in any of the above solutions, the offline distillation includes the following sub-steps:

[0037] Step 41: Use the frozen teacher model B to perform inference on the training set, generate soft labels and intermediate features for each sample;

[0038] Step 42: Design the total loss function;

[0039] Step 43: Use the total loss function to train the student model;

[0040] Step 44: Use the strong supervision of the teacher model B to guide the continuous optimization of the student model.

[0041] Preferably, in any of the above solutions, the total loss function includes a task loss and a distillation loss.

[0042] Preferably, in any of the above solutions, the calculation formula of the task loss is

[0043] L task =-Σy true log y pred

[0044] where L tack is the cross-entropy loss used for the task loss, L tack is the cross-entropy loss used for the task loss, y true is the probability distribution predicted by model B for this category, y pred is the predicted category probability of the student model.

[0045] Preferably, in any of the above solutions, the distillation loss includes a soft target loss and an intermediate feature distillation loss.

[0046] Preferably, in any of the above solutions, the calculation formula of the soft target loss is

[0047]

[0048] where Z student is the output of the student model, Z student is the output of the teacher model B, T is the temperature hyperparameter, and KL is a difference metric for measuring the probability distribution between the teacher model B and the student model.

[0049] Preferably, in any of the above solutions, the calculation formula of the intermediate feature distillation loss is

[0050]

[0051] where l is the feature layer, is the intermediate feature of the student model at the l-th layer, It is a feature adapter that maps the features of the teacher model to the feature space of the student model. is the intermediate feature of the teacher model at the l-th layer.

[0052] In any of the above solutions, preferably, the calculation formula of the total loss function is

[0053] L total = αL tack + βL soft + γL feat

[0054] where α, β, and γ are hyperparameters.

[0055] The second object of the present invention is to provide a millimeter-wave radar rehabilitation motion recognition system based on knowledge distillation, including a data acquisition unit and a server. It is characterized in that it further includes an edge computing device unit.

[0056] The edge computing device unit is used for model training and motion recognition based on online bidirectional distillation.

[0057] The system uses the method described in the first object for millimeter-wave radar rehabilitation motion recognition based on knowledge distillation.

[0058] Preferably, the data acquisition unit includes a millimeter-wave radar and a device with a mature bone tracking algorithm.

[0059] In any of the above solutions, preferably, the motion recognition includes: calculating the Euclidean distance error, the mean absolute error (MAE), and the joint angle error, and assigning weights to the three errors to obtain a compliance score, and then normalizing the score.

[0060] In any of the above solutions, preferably, the server is used to store data.

[0061] The present invention proposes a millimeter-wave radar rehabilitation motion recognition method and system based on knowledge distillation, which improves the efficiency of human skeleton recognition and enhances the real-time performance and efficiency of rehabilitation motion recognition. Description of the Drawings

[0062] Figure 1 is a flowchart of a preferred embodiment of the millimeter-wave radar rehabilitation motion recognition method according to the present invention.

[0063] Figure 2 is a schematic diagram of the composition of a preferred embodiment of the millimeter-wave radar rehabilitation motion recognition system according to the present invention.

[0064] Figure 3Schematic diagram of the architecture of a preferred embodiment of a millimeter-wave radar rehabilitation motion recognition system based on knowledge distillation according to the present invention.

[0065] Figure 4 Schematic diagram of the server-side process of a preferred embodiment of a millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation according to the present invention.

[0066] Figure 5 Schematic diagram of the client-side process of a preferred embodiment of a millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation according to the present invention.

[0067] Figure 6 Flowchart of an embodiment of a rehabilitation motion of a millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation according to the present invention.

[0068] Figure 7 Schematic diagram of the principle of an embodiment of knowledge distillation of a millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation according to the present invention.

[0069] Figure 8 Flowchart of an embodiment of knowledge distillation of a millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation according to the present invention.

[0070] Figure 9 Flowchart of an embodiment of error analysis of a millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation according to the present invention.

[0071] Figure 10 Schematic diagram of 19 skeletal joints of the human body in a preferred embodiment of a millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation according to the present invention.

[0072] Figure 11 Flowchart of an embodiment of model training of a millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation according to the present invention. Detailed implementation manners

[0073] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0074] Embodiment 1

[0075] As Figure 1 shown, a millimeter-wave radar rehabilitation motion recognition method based on knowledge distillation performs step 100 to obtain an image data set and a point cloud data set.

[0076] Perform step 110 to train the image data set using teacher model A.

[0077] Perform step 120 to train the point cloud data set using teacher model B.

[0078] Execute step 130, where the teacher model A and the teacher model B perform online bidirectional distillation, alternately update the parameters of the teacher model A and the teacher model B, and output the teacher model B. The online bidirectional distillation means that the teacher model A and the teacher model B are each other's teacher-student models, and perform bidirectional semantic decoupling on the teacher model A and the teacher model B.

[0079] The online bidirectional distillation includes the following sub-steps:

[0080] Execute step 131 to build a symmetric network structure, allowing the teacher model A and the teacher model B to be each other's teacher and student. At time step t, model A guides model B, and at time step t + 1, model B guides model A.

[0081] Execute step 132 to design a shared dynamic feature cache pool to store the latest N batches of cross-modal feature pairs in real time.

[0082] Execute step 133, where the point cloud semantics extracted by the teacher model B generate pseudo-image labels through a cross-modal classifier, and the image features of the teacher model A and the pseudo-image labels are used to calculate the contrast loss L cont and the SSIM loss SSIM(x, y), forcing the teacher model A to understand the point cloud semantics.

[0083] The contrast loss L cont has the following calculation formula

[0084]

[0085] where f i image is the feature vector extracted by the teacher model A for the i-th sample, is the pseudo-label feature generated by the teacher model B, P(i) is the set of pseudo-label features belonging to the same category as the i-th sample, A(i) is all the pseudo-label features in the current batch, τ is the temperature coefficient, and p is the sample belonging to the same category as the current sample i.

[0086] The calculation formula of the SSIM loss SSIM(x, y) is

[0087]

[0088] where μ x and μ y are the local means of x and y, and are the local variances of x and y, σ xy is the covariance of x and y, is the local variance of the image x, The local variance of the pseudo-image label y, C1 and C2 are constant terms, x is the image data, and y is the pseudo-image label.

[0089] Execute step 134. When the teacher model B is the teacher of the teacher model A, use the point cloud network for high-order semantic extraction and map the high-order semantics to the feature space.

[0090] Execute step 135. The gating network generates bidirectional distillation weights based on the modal quality assessment of the current batch of data. The bidirectional distillation weights include the distillation intensity ω that controls the teacher model A to the teacher model B A->B and the distillation intensity ω that controls the teacher model B to the teacher model A B->A .

[0091] The distillation intensity ω that controls the teacher model A to the teacher model B A->B The calculation formula is:

[0092] ω A->B =σ(f gate (I qualitu ,P density ))

[0093] Among them, ω A->B represents that model A guides model B for distillation, σ is the activation function, f gate is the gating network, I qualitu is the image quality assessment index, which is evaluated according to the image clarity, and P density is the point cloud quality assessment index, and the value range is [0,1]. The closer it is to 1, the higher the point cloud density.

[0094] The distillation intensity ω that controls the teacher model B to the teacher model A B->A The calculation formula is:

[0095] ω B->A =σ(f gate (I qualitu ,P density ))

[0096] Among them, ω B->A represents that model B guides model A for distillation.

[0097] Execute step 136. The teacher model A is updated based on the contrastive loss provided by the teacher model B, and the teacher model B is updated based on the contrastive loss L cont provided by the teacher model A and the SSIM loss SSIM(x,y).

[0098] Execute step 137. After the model is updated, output the teacher model B.

[0099] Execute step 140, freeze the parameters of the teacher model B, start offline distillation, guide the training of the student model, optimize the student model, and the teacher model B only serves as a teacher model to provide soft labels for the student model.

[0100] The offline distillation includes the following sub-steps:

[0101] Execute step 141, use the frozen teacher model B to perform inference on the training set, generate soft labels and intermediate features for each sample.

[0102] Execute step 142, design the total loss function, and the total loss function includes a task loss and a distillation loss.

[0103] The calculation formula for the task loss is

[0104] L task =-∑y true log y pred

[0105] where L tack is the cross-entropy loss used for the task loss, y true is the probability distribution predicted by model B for this category, and y pred is the predicted category probability of the student model.

[0106] The distillation loss includes a soft target loss and an intermediate feature distillation loss, and the calculation formula for the soft target loss is

[0107]

[0108] where Z student is the output of the student model, Z student is the output of the teacher model B, T is the temperature hyperparameter, and KL is a difference metric for measuring the probability distribution between the teacher model B and the student model.

[0109] The calculation formula for the intermediate feature distillation loss is

[0110]

[0111] where l is the feature layer, is the intermediate feature of the student model at the l-th layer, is the feature adapter that maps the features of the teacher model to the feature space of the student model, is the intermediate feature of the teacher model at the l-th layer.

[0112] The calculation formula for the total loss function is

[0113] L total =αL tack +βLsoft +γL feat

[0114] Among them, α, β, and γ are hyperparameters.

[0115] Execute step 143 to train the student model using the total loss function.

[0116] Execute step 144 to guide the continuous optimization of the student model using the strong supervision of the teacher model B.

[0117] Execute step 150 to obtain patient information, input it into the student model, and output the recognition result of the millimeter-wave radar rehabilitation movement.

[0118] Embodiment 2

[0119] The object of the present invention is to implement a convenient and home-use action recognition system for assisting rehabilitation training, using millimeter-wave signals for recognition to solve the problem of low environmental adaptability of computer vision technology in the field of recognizing rehabilitation actions, and to protect the necessary privacy of customers. The rehabilitation action recognition principle of the present invention is that the millimeter-wave radar emits high-frequency electromagnetic waves through the transmitting antenna, and then the receiving antenna captures the electromagnetic waves reflected by various parts of the human body. Through a series of digital signal processing, the signal is converted into a sparse point cloud in three-dimensional space. These point clouds are handed over to a pre-trained neural network model to obtain the point cloud information about the human skeleton. According to the human kinematics model, the predicted points are connected to form a complete skeleton. The edge computing device will match the rehabilitation action that the current user needs to perform with the human point cloud information of the user's current action, analyze and judge whether the action of the user's current rehabilitation movement is correct. If it is incorrect, it will make a prompt for the limb position that does not conform particularly according to the point cloud information of the user's current action.

[0120] The system structure schematic diagram of the present invention is as Figure 2As shown in the figure. This system mainly consists of the following parts: millimeter-wave radar, edge computing device, and server. Inside the server, there are a neural network model, a system controller, and a voice announcer. These parts are connected in sequence. The millimeter-wave radar and the edge computing device are connected through a data transmission interface, and the server and the edge computing device use the HTTP communication protocol for information interaction through a Wi-Fi module. Among them, the millimeter-wave radar is responsible for collecting all data and the task of collecting real-time human body point cloud information of users. The edge computing device is the core part of the whole system, responsible for identifying the human body point cloud collected by the millimeter-wave radar through the neural network model deployed inside, and finally obtaining all the skeleton key point information. The built-in voice announcer (if nvidia jetson is selected as the edge device, the voice announcer module needs to be connected through an external interface) can guide the user's rehabilitation training process. The system controller is the operation part of this system, responsible for providing an interface that allows administrators or users to set and adjust this system and perform operations such as selecting rehabilitation training models. The Wi-Fi module enables the entire edge computing device to be connected to the local area network through a router, establishing a connection channel for the physical layer and data link layer for communication with the server, and completing functions such as requesting a specified model from the server or uploading a new model.

[0121] The present invention designs a millimeter-wave radar rehabilitation motion recognition system based on knowledge distillation. The system consists of a server side and multiple client sides, and they exchange data and models through a communication protocol to achieve synchronous update of the global model and the local model. The system architecture of the present invention is as Figure 3 shown in the figure. In the initial stage of the system, the administrator logs in to the edge computing device with permissions and deposits several initial models that have learned the knowledge of complex and highly accurate teacher models through knowledge distillation into the server as rehabilitation motion models that users can choose to use. Although there are not many widely known types of rehabilitation motions, different types of rehabilitation motions target different user groups. For example, users who have undergone knee replacement surgery need joint range of motion training, while users who have suffered a stroke need more balance and coordination training. The initial models deposited in the server need to meet the needs of different users' rehabilitation motion focuses, which can also further lightweight the neural network and ensure the high real-time performance and high accuracy of the edge computing device. The present invention uses a strategy of administrator-executed update. Users only need to request the corresponding model file from the server, which is convenient for the update of new types of rehabilitation actions and can greatly improve the user experience.

[0122] The subsequent stage of the system mainly involves information transmission between the millimeter-wave radar and the edge computing device. The scenario is that the user starts using the rehabilitation system. After logging into the edge computing device with the corresponding user ID, the user selects a specified model (the model file can be stored as a local offline file or a new model can be requested from the server) and starts the rehabilitation exercise. The voice broadcaster in the edge computing device prompts the user about the current rehabilitation action required. At this time, the millimeter-wave radar is responsible for collecting the five-dimensional human body point cloud data of the current user. These data are transmitted to the edge computing device and, through the processing of the neural network, the rehabilitation exercise performed by the current user is obtained. It is then determined whether the user's action is in line. If it is in line, the voice broadcaster starts to prompt the next action. If it is not in line, the user is prompted to correct it, and the edge computing device gives prompts for the corresponding action until the user completes it correctly. When all the actions of the rehabilitation exercise are completed, the current rehabilitation exercise ends, and the edge computing device uploads the rehabilitation data of this time to the server.

[0123] After the initial several model files are stored in the system of the present invention, it is equivalent to version 1.0 of the system. The server-side flowchart is as Figure 4 shown. First, the server is responsible for five operations: querying rehabilitation data, sending the specified model to the edge, receiving the model transmitted from the edge device, entering a new user ID, and entering an administrator ID. According to the requirements of the client, it sends or waits for upload. Entering a new user ID and entering a new administrator ID are similar, but when entering a new administrator, the key of other administrators already stored in the database needs to be input.

[0124] The working process of the client of the present invention is as Figure 5 shown. Four main functions are executed on the edge computing device of the client: logging into the user mode, logging into the developer mode, entering the ID, and real-time collection. Among them, the real-time collection function is an operation that can be performed by both users and administrators. The administrator can use real-time collection to enable the millimeter-wave radar to collect data. In a suitable environment, if an external device for collecting tag data (such as Kinect V2) is used to cooperate, the human body rehabilitation exercise data can be collected synchronously. After the collection is completed, the data collected by the millimeter-wave radar is automatically saved as a local file, and the edge computing device waits for the tag data collected by the external device to be transmitted. After receiving it, both files will be stored in the local dataset folder. These two dataset files can be directly selected when the edge computing device subsequently creates a new model. The administrator can click on "developer mode" and enter the key to enter this mode. After that, the administrator can choose to update an existing model or create a new model. Updating an existing model and creating a new model are similar. Subsequently, both the millimeter-wave dataset and the tag dataset are received. These two datasets can either directly use the local dataset stored by the real-time collection function or be transmitted from the outside through the edge computing device interface. After the two datasets are received, they wait for neural network training. After the training is completed, they are saved as model files and uploaded to the server using the communication protocol.

[0125] The conventional status system defaults to the user mode. The user starts using the rehabilitation training guidance function by "inputting the user ID". If the user selects to query the rehabilitation data, the edge computing device will send a corresponding request to the server. After receiving the response, the user can query the completion status of the rehabilitation actions during the historical rehabilitation training. For example, if the number of times a certain rehabilitation action is guided by the voice broadcaster during each rehabilitation training is significantly higher than that of other rehabilitation actions, it indicates that the rehabilitation condition of the corresponding joint of the user is not good. It will be recommended to increase the frequency of the rehabilitation actions for the corresponding part in the subsequent rehabilitation training. When officially using, the user first selects the applicable rehabilitation model. The user can import the offline model file from the local or import a new model from the server and wait for the server to send it down. Then click "Start rehabilitation training", and the user can perform the corresponding rehabilitation actions according to the prompts of the voice broadcaster. When all the rehabilitation actions are completed, the rehabilitation training ends, and the updated user rehabilitation data is uploaded to the server. If a new user starts using for the first time, click "Enter the number ID", select to enter the administrator ID or the user ID, and wait for the server response. The remaining part is shown in the server flowchart as Figure 4 shown below.

[0126] When the client of the present invention executes the rehabilitation training, it will identify the user's real-time rehabilitation actions. The flowchart is as Figure 6 shown below. When identifying the rehabilitation movement, first, the user performs the corresponding rehabilitation actions according to the instructions of the current voice broadcaster. At this time, the millimeter-wave radar starts to collect the human body point cloud information. After neural network prediction, error analysis is carried out. If the result of the error calculation of the user's performed rehabilitation action does not meet the standard, it is regarded as not performing the correct action, and the user is prompted again to perform the corresponding rehabilitation actions according to the current instructions until the user's current rehabilitation action meets the standard. If the recognition result of the user's rehabilitation action meets the standard, it will be judged whether the current rehabilitation action has been completely completed. If it has not been completely completed, it will return to receive the next instruction and repeat the previous process. When all the actions are completed, the entire rehabilitation movement process ends.

[0127] Embodiment III

[0128] As Figure 7As shown, the neural network model deployed inside the edge computing device in the client of the present invention is trained based on the knowledge distillation method. Since tasks such as human skeleton recognition usually require high computing resources, knowledge distillation is used to compress the model, making the model lightweight and suitable for edge deployment. The entire training process includes four stages: data preparation, teacher model training, knowledge distillation strategy, and student model training optimization. Data preparation is the collection of the dataset. As mentioned before, the dataset can be collected in real time using a device that uses millimeter radar to assist in collecting labeled data, or imported from an external interface and handed over to the neural network for training. Before training the teacher model, a suitable teacher model needs to be selected, which should have high-precision characteristics. Usually, a network with a large number of parameters (such as ResNet-152, HRNet-W48) is selected. Its output not only contains the coordinates of the skeleton key points, but also provides spatial semantic information through the intermediate layer feature maps. The Pretrain-Finetune method is used to accelerate convergence, and the cosine annealing learning rate scheduler is used to optimize the training stability. After training is completed, the parameters of the teacher model are frozen, and the soft labels of its output layer and the feature maps of the intermediate layer are recorded as the supervision signals for knowledge distillation.

[0129] The knowledge distillation strategy adopts offline knowledge distillation to improve the performance of the student model through static knowledge transfer (the parameters of the teacher model are fixed). A multi-level knowledge transfer mechanism needs to be designed, including output layer knowledge transfer, intermediate layer knowledge transfer, and multi-task joint optimization, so that the student model can learn the confidence distribution and localization accuracy of the teacher model for fuzzy samples. For the training optimization of the student model, first, a suitable lightweight backbone network (such as MobileNetV3) needs to be selected. First, use soft label distillation to train the student model, and then gradually introduce intermediate layer feature constraints to avoid overfitting too early. Dropout and LabelSmoothing are used to prevent the student model from relying too much on the teacher's knowledge. In terms of hardware adaptation, continuous operations such as Conv-BN-ReLU can be combined into a single computing unit, which can further reduce the memory access overhead during inference.

[0130] such as Figure 8As shown in the figure, the present invention uses multi-layer teacher collaborative distillation. The teacher model A trains the image dataset, and the teacher model B trains the point cloud dataset. The teacher model A and the teacher model B perform online bidirectional distillation, alternately updating the parameters of model A and model B. Finally, when the parameter update of model B is completed, it serves as the teacher of the student model to guide the update of the parameters of the student model. Since both the teacher model B and the student model train on point cloud data, the student model can directly learn in the feature space of the point cloud. After the online distillation stage, only the output of teacher B needs to be concerned about, which avoids the need to consider the feedback of multiple teacher models simultaneously during the training of the student model, thus reducing the computational and design complexity during the training process.

[0131] Overall, progressive distillation is adopted. In the online stage, bidirectional distillation between model A and model B focuses on feature space alignment. In the offline stage, model B guides the training of the student model for multi-level (feature and output) knowledge transfer. In the online stage, the parameters of model A and B are updated according to the modality alignment loss and the supervision loss. In the offline stage, the student model is continuously optimized according to the feature imitation loss, soft label distillation, and task loss.

[0132] Different from traditional bidirectional distillation methods and multi-teacher multi-modal distillation, the hierarchical distillation of the present invention is divided into two stages. The first stage is online bidirectional distillation, which performs bidirectional semantic decoupling on the image model and the point cloud model. Models A and B are teachers and students of each other, enabling the point cloud model B to more fully integrate the knowledge of the image model A, laying the foundation for guiding the point cloud student model next. The second stage is offline distillation, which only uses model B to guide the training of the student model.

[0133] The specific solution is as follows:

[0134] Step 1: Obtain the image dataset and the point cloud dataset.

[0135] Step 2: The teacher model A trains the image dataset, and the teacher model B trains the point cloud dataset (pre-training starts simultaneously).

[0136] Step 3: The image data and the point cloud data need to be modality-aligned. The traditional method based on projection often ignores the decoupling relationship between semantics and geometric structures across modalities. Here, the online bidirectional distillation is a cross-modal bidirectional semantic decoupling distillation method. First, a symmetric network structure is constructed, allowing teacher model A and teacher model B to be teachers and students of each other. At time step t, model A guides model B, and at time step t+1, model B guides model A. A shared dynamic feature cache pool is designed to store the latest N batches of cross-modal feature pairs (image + point cloud) in real time to alleviate the temporal mismatch problem in online learning. The point cloud semantics extracted by teacher model B generate pseudo-image labels through a cross-modal classifier. The image features of model A and the pseudo-image labels are used to calculate the contrast loss (5) and the SSIM loss (6), forcing the image model to understand the point cloud semantics. Since the image model needs complex geometric calculations to generate 3D pseudo-point clouds through the inverse projection method and is prone to introducing noise, when B is the teacher of model A, a point cloud network (such as PointNet++) can be used for high-order semantic extraction and mapping the high-order semantics to the feature space. After the feature alignment is prepared, the gating network generates bidirectional distillation weights (7) based on the modality quality assessment of the current batch of data (such as image clarity, point cloud density). The image model A is updated based on the contrast loss provided by the point cloud model B. The point cloud model is updated based on the contrast loss and the SSIM loss provided by the image model. Finally, the model update is completed, and model B is output.

[0137] Step 4: Freeze the parameters of model B. At this time, model B only serves as a teacher model to provide soft labels for the student model. Start offline distillation. Use the frozen model B to perform inference on the training set, generate the soft labels and intermediate features of each sample, and then design the loss function. The total loss includes the task loss (8) and the distillation losses (9, 10). Since model B and the student model belong to the same point cloud feature space, the entire training process can use the strong supervision of teacher model B and directly use the total loss (11) to train from scratch. Model B guides the student model to continuously optimize during this process.

[0138]

[0139] where, f i image is the feature vector extracted by the image model A for the i-th sample, is the pseudo-label feature generated by the point cloud model B; P(i) is the set of pseudo-label features belonging to the same category as the i-th sample; A(i) is all the pseudo-label features in the current batch; τ is the temperature coefficient; N is the batch size,

[0140]

[0141] where, μx and μ y are the local means (luminance contrast) of images x and y, and is the local variance (contrast contrast) of images x and y, σ xy is the covariance of images x and y (structural contrast C1 and C2 are constant terms. In this method, x and y are the image data and the pseudo-image labels respectively.

[0142] ω A->B = σ(f gate (I qualitu , P density )) (7)

[0143] The dynamic weight control network, ω A->B controls the distillation intensity from the image model to the point cloud model, ω B->A controls the distillation intensity from the point cloud model to the image model.

[0144] L task = -∑y true log y pred (8)

[0145] The cross-entropy loss used for the task loss.

[0146]

[0147] The soft target loss, where Z student represents the output of the student model, Z student Similarly, divide by the temperature hyperparameter and obtain the output probability through the σ function. In the loss function of knowledge distillation, the student model is generally optimized by comparing the differences between the soft target probability distributions of the student model and the teacher model, and this difference is measured using the KL divergence.

[0148]

[0149] The intermediate feature distillation loss.

[0150] L total = αL tack + βL soft + γL feat (11)

[0151] The total loss, and the three hyperparameters α, β, and γ are set according to the task performance.

[0152] The present invention utilizes millimeter-wave signals as the basis for identifying human movements in rehabilitation exercises. Compared with the mainstream visual signal-based recognition of rehabilitation exercises, it is not affected by environmental factors such as light intensity. Compared with wearable sensor solutions, it can achieve non-contact human skeleton recognition and can protect user privacy, making it more suitable for home rehabilitation training conditions. Combining the advantages of the knowledge distillation method, high timeliness, high accuracy, and lightweight deployment of neural networks on edge devices are achieved.

[0153] Through multi-layer teacher distillation, the image modality that can provide semantic priors is trained by Teacher A. The first stage is the online two-way distillation between Teacher B and Teacher A, and Teacher B can learn the knowledge of Teacher A; the second stage is that Teacher B guides the training of the student model. Since both Teacher B and the student model are trained on point cloud data, the student model can directly learn in the point cloud feature space without considering the feedback of multiple teachers in multiple modalities, greatly saving computational resources.

[0154] Embodiment III

[0155] As Figure 9 shown, to measure the similarity between the user's current movement and the action point cloud information in the label data, a common method is to measure the matching degree between the two by calculating the error. The goal of error measurement is to measure the difference between the user's current movement and the label data (standard movement). The present invention identifies rehabilitation actions based on 19 skeleton joint points of the human body as Figure 10 shown, and evaluates the standard of the current movement according to the error magnitude. Common error evaluation methods include: Euclidean distance error, mean absolute error (MAE), and joint angle error. After rigorous calculation, weights are assigned to the three errors to obtain a compliance score, and then the score is normalized. If the final score is higher than 0.7, it is evaluated as "the current movement meets the standard".

[0156] Euclidean error

[0157]

[0158] Among them, x j,user is the x coordinate of the user's current j-th movement, x j,label is the x coordinate label of the current j-th movement, y j,user is the y coordinate of the user's current j-th movement, y j,label is the y coordinate label of the current j-th movement, z j,user is the z coordinate of the current j-th movement, z j,label is the z coordinate label of the current j-th movement, and M is the total number of movements.

[0159] Mean absolute error

[0160]

[0161] Angle Error

[0162] Error3=Angle Error=|θ user -θ label | (3)

[0163] where θ user is the current joint angle of the user, and θ label is the joint angle of the label.

[0164] Conformance

[0165] G=ω1Error1+ω2Error2+ω3Error3 (4)

[0166] where ω1 is the weight of the Euclidean error, ω2 is the weight of the mean absolute error, and ω3 is the weight of the angle error. Among the three weights, ω1 and ω2 are equally important, and ω3 is less important.

[0167] Example 4

[0168] As Figure 11 shown, first is data preparation, including an image dataset and a point cloud dataset. After that, preprocessing operations such as cleaning and segmentation are performed on the dataset. Then, two teacher models are trained. After the training of the teacher models is completed, the parameters of the teacher models will be frozen. At the same time, when the student model is trained, the parameters are updated. Using the preprocessed data, the soft labels and feature maps obtained from the training of the teacher models, the distillation loss is obtained. The parameters of the student model are updated according to the distillation loss. After the parameters of the student model are updated, this model will be deployed on the edge device. Through operations such as pruning optimization, the student model is further lightweighted. Then, the validation dataset is used to calculate the accuracy metrics using real-scene data on the edge device. Considering the measured inference latency (FPS), memory occupancy, and average power consumption comprehensively, it is decided whether the model meets the standard. If it has met the standard, the model can be uploaded to the server, and the whole process ends. If it does not meet the standard, some steps or model data are fine-tuned manually.

[0169] Example 5

[0170] A millimeter-wave radar rehabilitation motion recognition system based on knowledge distillation, including a data acquisition unit 200, a server 210, and an edge computing device unit 220,

[0171] The data acquisition unit 200 includes a millimeter-wave radar and a device with a mature bone tracking algorithm.

[0172] The server 210 is used to store data.

[0173] The edge computing device unit 220 is used for model training and motion recognition based on online two-way distillation. The motion recognition includes: calculating the Euclidean distance error, the mean absolute error (MAE), and the joint angle error, assigning weights to the three errors, obtaining a compliance score, and then normalizing the score.

[0174] To better understand the present invention, the above has been described in detail in conjunction with specific embodiments of the present invention, but it is not a limitation of the present invention. Any simple modification made to the above embodiments based on the technical essence of the present invention still belongs to the scope of the technical solution of the present invention. Each embodiment in this specification focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

Claims

1. A method for recognizing rehabilitation motion using millimeter wave radar based on knowledge distillation, comprising acquiring an image dataset and a point cloud dataset, characterized in that: The following steps are also included: Step 1: Use teacher model A to train the image dataset; Step 2: Use teacher model B to train the point cloud dataset; Step 3: The teacher model A and the teacher model B are bidirectionally distilled online, the parameters of the teacher model A and the teacher model B are updated alternately, and the teacher model B is output; Step 4: Freeze the parameters of the teacher model B, start offline distillation, guide the student model training, and optimize the student model; Step 5: Obtain patient information, input it into the student model, and output the millimeter-wave radar rehabilitation movement recognition result.

2. The method for recognizing rehabilitation movements using millimeter wave radar based on knowledge distillation as claimed in claim 1, characterized in that: The online bidirectional distillation means that the teacher model A and the teacher model B are teacher-student models of each other, and bidirectional semantic decoupling of the teacher model A and the teacher model B is performed.

3. The method for recognizing rehabilitation movements using millimeter wave radar based on knowledge distillation as claimed in claim 2, characterized in that: The online bidirectional distillation comprises the following sub-steps: Step 31: Build a symmetric network structure, allowing the teacher model A and the teacher model B to be teachers and students to each other. At time step t, model A guides model B, and at time step t+1, model B guides model A. Step 32: Design a shared dynamic feature cache pool to store the latest N batches of cross-modal feature pairs in real time; Step 33: The point cloud semantics extracted by the teacher model B generates a pseudo image label through a cross-modal classifier, and the image features of the teacher model A and the pseudo image label calculate the contrast loss L cont With SSIM loss SSIM(x,y), the teacher model A is forced to understand the semantics of the point cloud; Step 34: When the teacher model B serves as the teacher of the teacher model A, a point cloud network is used to extract high-order semantics and map the high-order semantics to a feature space; Step 35: The gating network generates bidirectional distillation weights based on the modal quality assessment of the current batch of data; Step 36: The teacher model A is updated based on the contrast loss provided by the teacher model B, and the teacher model B is updated based on the contrast loss L provided by the teacher model A. cont And the SSIM loss SSIM(x,y) is updated; Step 37: After the model is updated, the teacher model B is output.

4. The method for recognizing rehabilitation movements using millimeter wave radar based on knowledge distillation as claimed in claim 3, characterized in that: The contrast loss L cont The calculation formula is in, is the feature vector extracted by the teacher model A for the i-th sample, is the pseudo-label feature generated by the teacher model B, P(i) is the set of pseudo-label features belonging to the same category as the i-th sample, A(i) is all the pseudo-label features in the current batch, τ is the temperature coefficient, and p is the sample belonging to the same category as the current sample i.

5. The method for recognizing rehabilitation movements using millimeter wave radar based on knowledge distillation as claimed in claim 4, characterized in that: The calculation formula of the SSIM loss SSIM(x,y) is: Among them, μ x and μ y is the local mean of x and y, and is the local variance of x and y, σ xy is the covariance of x and y, is the local variance of image x, is the local variance of the pseudo image label y, C1 and C2 are constant terms, x is the image data, and y is the pseudo image label.

6. The method for recognizing rehabilitation movements using millimeter wave radar based on knowledge distillation as claimed in claim 5, characterized in that: The bidirectional distillation weight includes controlling the distillation strength ω from teacher model A to teacher model B A->B and controls the distillation strength ω of teacher model B to teacher model A B->A .

7. The method for recognizing rehabilitation movements using millimeter wave radar based on knowledge distillation as claimed in claim 6, characterized in that: The control of the distillation strength ω from teacher model A to teacher model B A->B The calculation formula is: oh A->B =σ(f gate (I qualitu ,P density )) Among them, ω A->B It means that model A guides model B to perform distillation, σ is the activation function, and f gate For the gated network, I qualitu is an image quality evaluation index, which is evaluated based on image clarity. density It is a point cloud quality assessment indicator.

8. The method for recognizing rehabilitation movements using millimeter wave radar based on knowledge distillation as claimed in claim 7, characterized in that: The control of the distillation strength ω of the teacher model B to the teacher model A B->A The calculation formula is: oh B->A =σ(f gate (I qualitu ,P density )) Among them, ω B->A It means that model B guides model A to perform distillation.

9. The method for recognizing rehabilitation movements using millimeter wave radar based on knowledge distillation as claimed in claim 8, characterized in that: The off-line distillation comprises the following sub-steps: Step 41: Use the frozen teacher model B to infer the training set and generate soft labels and intermediate features for each sample; Step 42: Design a total loss function; Step 43: train the student model using the total loss function; Step 44: Use the strong supervision of the teacher model B to guide the student model to continuously optimize.

10. A millimeter wave radar rehabilitation movement recognition system based on knowledge distillation, comprising a data acquisition unit and a server, characterized in that: It also includes an edge computing device unit, wherein the edge computing device unit is used for model training and motion recognition based on online bidirectional distillation, The system uses the method as claimed in claim 1 to perform millimeter wave radar rehabilitation motion recognition based on knowledge distillation.

Citation Information

Patent Citations

  • Knowledge distillation method and system based on multi-teacher multi-modal model

    CN117669693A