Robot form recognition method and device based on knowledge distillation

Through knowledge distillation, the pre-trained knowledge of large models is migrated to small models, which solves the shortcomings in recognition accuracy and computational complexity of small models, and realizes real-time recognition effect on edge devices.

CN120164197APending Publication Date: 2025-06-17LONGWOOD VALLEY MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510228968.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Small models are difficult to take into account the recognition accuracy and computational complexity of large models, especially on resource-constrained edge devices, which are difficult to run in real time.

Method used

The pre-trained large models are migrated to small models through knowledge distillation to form a migration model, and the model is used to identify robotic morphology in the video stream on edge devices.

Benefits of technology

The recognition accuracy of small models on specific tasks is achieved close to that of large models, while significantly reducing the computational complexity, which is suitable for real-time operation on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164197A_ABST
    Figure CN120164197A_ABST
Patent Text Reader

Abstract

The invention provides a robot form recognition method and device based on knowledge distillation. The method comprises the steps that a pre-training large model and a to-be-trained model are acquired; the pre-trained large model is a large-scale language model of a Transform architecture; based on a knowledge distillation mode, migrating the pre-trained large model to the to-be-trained model to obtain a migration model; and identifying a robot form in the video stream based on the migration model. Through combination of knowledge distillation and transfer learning, pre-training knowledge of the large model is transferred to the small model, so that accurate identification of the robot form is realized through the small model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of medical image recognition. Specifically, it relates to a robot morphology recognition method and device based on knowledge distillation. Background Art

[0002] In existing robot morphology recognition methods, large models have high computational complexity and are difficult to run in real time on resource-constrained edge devices. Small models, due to their small number of parameters, are difficult to directly learn effective features from large-scale data, resulting in low recognition accuracy. Summary of the Invention

[0003] The problem solved by this application is that small models are difficult to combine the advantages of large models.

[0004] To solve the above problems, the first aspect of this application provides a robot morphology recognition method based on knowledge distillation, including:

[0005] Obtain a pre-trained large model and a model to be trained; the pre-trained large model is a large-scale language model with a Transformer architecture;

[0006] Based on the knowledge distillation method, transfer the pre-trained large model to the model to be trained to obtain a transferred model;

[0007] Based on the transferred model, recognize the morphology of the robot in the video stream.

[0008] The second aspect of this application provides a robot morphology recognition device based on knowledge distillation, which includes:

[0009] A model acquisition module, which is used to obtain a pre-trained large model and a model to be trained; the pre-trained large model is a large-scale language model with a Transformer architecture;

[0010] A knowledge distillation module, which is used to transfer the pre-trained large model to the model to be trained based on the knowledge distillation method to obtain a transferred model;

[0011] A morphology recognition module, which is used to recognize the morphology of the robot in the video stream based on the transferred model.

[0012] The third aspect of this application provides an electronic device, which includes: a memory and a processor;

[0013] The memory is used to store programs;

[0014] The processor is coupled to the memory and is used to execute the program for:

[0015] Obtain a pre-trained large model and a model to be trained; the pre-trained large model is a large-scale language model with a Transformer architecture;

[0016] Based on the knowledge distillation method, transfer the pre-trained large model to the model to be trained to obtain a transferred model;

[0017] Based on the transferred model, identify the robot form in the video stream.

[0018] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the above-mentioned robot form recognition method based on knowledge distillation.

[0019] In the present application, by combining knowledge distillation and transfer learning, transfer the pre-trained knowledge of the large model to the small model, so as to accurately identify the robot form through the small model.

[0020] In this way, the small model can achieve recognition accuracy close to that of the large model on specific tasks, while significantly reducing the computational complexity, making it suitable for real-time operation on edge devices. Description of the Drawings

[0021] Figure 1 It is a flowchart of the robot form recognition method based on knowledge distillation according to an embodiment of the present application;

[0022] Figure 2 It is a structural block diagram of the robot form recognition device based on knowledge distillation according to an embodiment of the present application;

[0023] Figure 3 It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed Embodiments

[0024] To make the above objects, features, and advantages of the present application more obvious and understandable, the following detailed description of the specific embodiments of the present application is provided in conjunction with the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0025] It should be noted that unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meaning understood by those skilled in the art to which the present application belongs.

[0026] To address the above problems, the present application provides a new robot form recognition solution based on knowledge distillation, which can transfer the pre-trained knowledge of the large model to the small model by combining knowledge distillation and transfer learning, and solve the problem that it is difficult for the small model to take into account the advantages of the large model.

[0027] The embodiment of the present application provides a robot form recognition method based on knowledge distillation. The specific solution of this method is as follows Figure 1 As shown, this method can be executed by a robot form recognition device based on knowledge distillation. This robot form recognition device based on knowledge distillation can be integrated into electronic devices such as computers, servers, computers, server clusters, and data centers. Combining Figure 1 As shown, it is a flowchart of a robot form recognition method based on knowledge distillation according to an embodiment of the present application; wherein, the robot form recognition method based on knowledge distillation includes:

[0028] S101, obtain a pre-trained large model and a model to be trained; the pre-trained large model is a large-scale language model with a Transformer architecture;

[0029] In the present application, the pre-trained large model has strong feature extraction capabilities and can capture complex spatio-temporal features. It has a large number of model parameters and high computational complexity, and is suitable as a teacher model.

[0030] In the present application, a lightweight model is selected as the model to be trained. The model to be trained has fewer parameters and low computational complexity, and is suitable for deployment on resource-constrained devices.

[0031] S102, based on the knowledge distillation method, transfer the pre-trained large model to the model to be trained to obtain a transferred model;

[0032] In the present application, the pre-trained large model is used to generate soft labels (i.e., probability distributions); the model to be trained is trained to make its output distribution close to the soft labels of the large model.

[0033] S103, based on the transferred model, recognize the form of the robot in the video stream.

[0034] In the present application, the transferred model is deployed to edge devices (such as Jetson Nano, NVIDIA Xavier) or cloud servers. Model quantization (such as INT8 quantization) and pruning techniques are used to further compress the model. Frames are read from a camera or video file, the frame resolution is adjusted (such as 224x224), and the pixel values are normalized (such as scaling the pixel values to [0,1]). The frame sequence is input into the transferred model, and the robot form recognition result is output.

[0035] In the present application, by combining knowledge distillation and transfer learning, the pre-trained knowledge of the large model is transferred to the small model, so as to accurately recognize the form of the robot through the small model.

[0036] In this way, the small model can achieve recognition accuracy close to that of the large model on specific tasks, while significantly reducing the computational complexity, and is suitable for real-time operation on edge devices.

[0037] In this application, by combining knowledge distillation and transfer learning, the performance of lightweight models in robot form recognition tasks is significantly improved, while the computational complexity is reduced, making it suitable for deployment on resource-constrained devices.

[0038] In one implementation, the S102, based on the knowledge distillation method, transfers the pre-trained large model to the model to be trained to obtain a transferred model, including:

[0039] Obtain a video dataset;

[0040] Fine-tune the pre-trained large model based on the video dataset;

[0041] Construct the model to be trained;

[0042] Based on soft label distillation, train the model to be trained through the fine-tuned pre-trained large model to obtain a transferred model.

[0043] In this application, the dataset source: public video datasets (such as Kinetics, UCF101, Something-Something). Robot-related video data (such as robot movement, form changes).

[0044] Dataset requirements: The video resolution is at least 224x224. The video length is moderate (such as 5-10 seconds). Data diversity (different scenarios, lighting, backgrounds, robot forms).

[0045] Dataset division: Divide the dataset into a training set, a validation set, and a test set (such as 8:1:1)

[0046] In this application, the goal of fine-tuning the pre-trained large model based on the video dataset: making the pre-trained large model adapt to the robot form recognition task.

[0047] Fine-tuning method: Freeze some layers of the large model (such as the shallow convolutional layer), and only fine-tune the top layer. Use a smaller learning rate (such as 1e-4) to avoid overfitting.

[0048] In one implementation, before fine-tuning the pre-trained large model, an adaptive adjustment module is added before the input layer of the pre-trained large model to adjust the input image frames.

[0049] The processing process of the adaptive adjustment module includes:

[0050] Divide the image frame into blocks to obtain independent blocks;

[0051] For each independent block, obtain the first neighborhood block and the second neighborhood block with different spacings;

[0052] Generate a first feature block based on the independent block and the first neighborhood blocks;

[0053] Generate a second feature block based on the independent block and the second neighborhood blocks;

[0054] Perform feature compression on the first feature block and the second feature block to obtain a compressed block;

[0055] Traverse all the independent blocks and generate an adjusted image based on the obtained compressed block.

[0056] In this application, partitioning the image frame into blocks means dividing the image frame into corresponding image blocks through a checkerboard; among them, the image block can be at the pixel level (that is, each pixel is an image block), or at other levels, and the specific partitioning depends on the actual processing situation.

[0057] In this application, the image is divided into blocks of the same size using a sliding window or a fixed step size.

[0058] It should be noted here that if the image frame is a two-dimensional image, it is directly divided into a checkerboard, and each grid is an image block; if the image frame is a three-dimensional image, a surface is selected for checkerboard partitioning, and each grid is a strip-shaped grid with a lot of depth (the depth is the depth of the three-dimensional image), and this strip-shaped grid is an image block.

[0059] Preferably, in this application, each image block has 100 - 1000 pixels, so as to perform more feature calculations between local regions on the basis of ensuring the generation accuracy and reducing the calculation amount.

[0060] In this application, an image block is selected as the independent block, and the adjacent image blocks above, below, left, and right of the independent block are the first neighborhood blocks; the image blocks separated by one grid above, below, left, and right of the independent block are the second neighborhood blocks. The distances between the first neighborhood blocks and the second neighborhood blocks and the independent block are different.

[0061] In this application, neighborhood information is extracted for each independent block to capture the local structure.

[0062] In this application, to generate the first feature block, a local feature representation is generated using the independent block and its first neighborhood blocks, and specifically, it can be: performing convolutional layer and attention layer processing on the independent block and the first neighborhood blocks to obtain the first feature block.

[0063] In this application, the specific structures and specific parameters of the convolutional layer and the attention layer can be obtained according to the training data or determined according to the actual situation.

[0064] It should be noted that in this application, there are four first neighborhood blocks and multiple first feature blocks.

[0065] In this application, the independent block and the first neighborhood block are processed through a convolutional layer and an attention layer to obtain the first feature block. The specific process is as follows: The independent block and four neighborhood blocks are concatenated together to form a multi-channel input, and the convolutional layer is used to extract features from the concatenated blocks; The self-attention mechanism or the channel attention mechanism is used to enhance important features, calculate the attention weights, and weight the output of the convolutional layer to enhance important features; The output of the attention layer is split into multiple feature blocks, and each feature block corresponds to the processing results of the independent block and at least one neighborhood block.

[0066] In this application, the second feature block is generated to generate a more extensive local feature representation using the independent block and its second neighborhood block. The specific generation process is the same as that of the first feature block, except that the parameters of the convolutional layer and the attention layer are different.

[0067] In this application, the generated feature blocks are compressed into a more compact representation to reduce the computational amount and retain key information. Pooling operations (such as max pooling or average pooling) or fully connected layers are used for feature compression.

[0068] In this way, through compression, multiple first feature blocks and second feature blocks are compressed into a compressed block, which corresponds to the independent block in size and position and is used to replace the independent block. All image blocks are replaced by compressed blocks to obtain an adjusted image.

[0069] In this application, each image block of the image frame is traversed in a traversal manner to obtain the corresponding compressed block.

[0070] In this application, for the image blocks / independent blocks near the edge, their first neighborhood blocks and second neighborhood blocks are incomplete. At this time, the first neighborhood blocks and second neighborhood blocks in the relative positions are copied for completion. For example, if the first neighborhood block above the independent block does not exist, the first neighborhood block below is copied and used as the block above.

[0071] In this application, through completion, the processing accuracy of the edge image blocks is greatly improved.

[0072] In this application, the adaptive adjustment module is used to capture the similarity relationship between local regions, thereby enhancing the feature representation. For images with rich textures or complex structures, high-quality images can be generated.

[0073] In this application, due to the defects of the pre-trained large model itself, if directly fine-tuned, the corresponding purpose may not be achieved. By adding an adaptive adjustment module, the pre-trained large model can adjust the parameters of the adaptive adjustment module during fine-tuning, thereby greatly increasing the amplitude of fine-tuning and achieving a better fine-tuning effect.

[0074] In one implementation, based on soft label distillation, training the model to be trained with a fine-tuned pre-trained large model to obtain a transfer model, including:

[0075] Obtain training data, where the training data is a part of the video dataset;

[0076] Input the training data into the fine-tuned pre-trained large model to obtain corresponding soft labels;

[0077] Input the training data into the model to be trained to obtain prediction data;

[0078] Based on the prediction data and the soft labels, calculate the overall loss of the model to be trained;

[0079] Adjust the model to be trained based on the overall loss until the loss converges.

[0080] In this application, randomly select a part from the video dataset as training data; the training data should contain diverse video segments, covering different scenarios, lighting conditions, and robot morphologies.

[0081] In this application, input the training data (video segments) into the fine-tuned pre-trained large model, use the temperature parameter to soften the probability distribution, and the soft labels are the probability distributions output by the large model, which contain more information (such as the relationships between classes).

[0082] In this application, input the same training data into the model to be trained, and the model to be trained generates an output (probability distribution).

[0083] In this application, after obtaining the overall loss, calculate the gradient of the total loss and update the parameters of the model to be trained. Among them:

[0084] Optimizer: Use Adam or SGD optimizer.

[0085] Learning rate: Set the initial learning rate to 1e-3 and dynamically adjust it using a learning rate scheduler (such as Cosine Annealing).

[0086] Training termination: Stop training when the loss converges (such as the validation set loss no longer decreases) or reaches the maximum number of training epochs.

[0087] Model saving: Save the transfer model with the best performance.

[0088] In one implementation, based on the prediction data and the soft labels, calculating the overall loss of the model to be trained includes:

[0089] Based on the KL divergence, calculate the difference loss between the prediction data and the soft labels;

[0090] Calculate the true loss between the predicted data and the annotations of the training data based on the cross-entropy loss;

[0091] Calculate the overall loss of the model to be trained based on the difference loss and the true loss.

[0092] In this application, the definition of KL divergence: KL divergence (Kullback-Leibler Divergence) is used to measure the difference between two probability distributions.

[0093] In this application, when calculating the KL divergence, a temperature parameter is introduced to soften the probability distribution; the KL divergence is used to calculate the difference loss between the predicted data and the soft labels.

[0094] In this application, the definition of cross-entropy loss: Cross-entropy loss is used to measure the difference between the predicted data and the true labels. The cross-entropy loss is used to calculate the true loss between the predicted data and the true labels.

[0095] In this application, the overall loss is the weighted sum of the difference loss and the true loss.

[0096] In this application, weight adjustment is also included: at the beginning of training (the first preset number of training rounds), a relatively high weight for the difference loss (such as 0.9) can be set to enable the student model to learn more knowledge from the teacher model; at the end of training (the last preset number of training rounds), the weight of the difference loss can be gradually reduced (such as 0.5) to enable the student model to pay more attention to the true labels.

[0097] In this application, by transferring the knowledge of the large model to the lightweight model through soft label distillation, the performance of the lightweight model in the robot form recognition task is significantly improved, while the computational complexity is reduced, making it suitable for deployment on resource-constrained devices.

[0098] In one implementation, after the S102, based on the knowledge distillation method, transferring the pre-trained large model to the model to be trained to obtain a transferred model, further includes:

[0099] Fine-tune the transferred model.

[0100] In this application, by fine-tuning the transferred model, the performance of the model in the robot form recognition task is further improved.

[0101] In one implementation, the fine-tuning of the transferred model includes:

[0102] Construct a fine-tuning dataset, where the fine-tuning dataset includes sample video frames with robot form annotations;

[0103] Freeze the shallow convolutional layers of the transferred model;

[0104] Train the frozen transfer model based on the fine-tuning dataset until the loss converges.

[0105] In this application, by fine-tuning the transfer model, the performance of the model in the robot form recognition task is further improved, and overfitting is avoided at the same time.

[0106] In this application, select sample video frames from the video dataset to ensure data diversity; the sample video frames should have robot form annotations (such as bounding box, key points, class labels).

[0107] This application also includes data preprocessing: adjust the video frame resolution (such as 224x224). Normalize the video frames (such as scale the pixel values to [0,1]). Perform data augmentation (such as random cropping, rotation, flipping)

[0108] In this application, the purpose of freezing the shallow convolutional layers of the transfer model: retain the general features learned by the transfer model during the knowledge distillation process and avoid overfitting; only fine-tune the top layer to make the model adapt to specific tasks

[0109] Freezing method: traverse the shallow convolutional layers (the first preset number of layers) of the transfer model and set their parameters to be non-trainable.

[0110] In this application, the specific process of training the frozen transfer model based on the fine-tuning dataset until the loss converges:

[0111] Input the sample video frames in the fine-tuning dataset into the transfer model to obtain the output probability distribution; dynamically apply data augmentation techniques (such as random cropping, rotation, flipping) during the training process; calculate the cross-entropy loss; calculate the gradient of the loss with respect to the model parameters; update the parameters of the unfrozen layers of the transfer model; repeat the above steps until the loss converges or the maximum number of training epochs is reached. Convergence condition: when the validation set loss no longer decreases or reaches a preset threshold, stop training; save the fine-tuning model with the best performance.

[0112] In this application, optimizer selection: use Adam or SGD optimizer.

[0113] Learning rate setting: Set the initial learning rate to a small value (such as 1e-4). Use a learning rate scheduler (such as ReduceLROnPlateau) to dynamically adjust the learning rate.

[0114] In this application, a fine-tuning dataset is constructed: sample video frames are selected from the video dataset and preprocessed and data-augmented. The shallow convolutional layers are frozen: the general features of the transfer model are retained to avoid overfitting. Fine-tuning training: the transfer model is trained using the fine-tuning dataset. The cross-entropy loss is calculated (the distillation loss can also be combined / continue to use the soft labels of the teacher model for weight calculation), and the parameters of the unfrozen layers are updated through backpropagation. The learning rate is dynamically adjusted using a learning rate scheduler. Training termination: when the loss converges, the fine-tuning model with the best performance is saved.

[0115] In one implementation, S103, based on the transfer model, identifying the robot form in the video stream includes:

[0116] Deploy the transfer model to a computer device or a cloud server;

[0117] Compress the transfer model using model quantization or pruning techniques;

[0118] Obtain the video stream to be identified;

[0119] Input the video stream to be identified into the transfer model to obtain the identified robot form.

[0120] In this application, deployment platform selection: Computer device: such as edge devices (Jetson Nano, NVIDIA Xavier) or local servers. Cloud server: such as AWS, Google Cloud, Azure.

[0121] In this application, deployment method: Convert the transfer model into a format suitable for deployment (such as ONNX, TensorRT). Use an inference engine (such as TensorRT, OpenVINO) to accelerate model inference.

[0122] In this application, environment configuration: Install necessary dependency libraries (such as PyTorch, TensorRT). Configure hardware acceleration (such as CUDA, Tensor Cores).

[0123] In this application, model quantization method: Convert the model weights and activation values from floating-point numbers (FP32) to 8-bit integers (INT8).

[0124] Quantization steps: Use a calibration dataset to determine the quantization parameters, convert the model into a quantized version, and ensure that the accuracy loss of the quantized model is within an acceptable range.

[0125] In this application, model pruning: Remove entire convolutional kernels or channels or remove individual weights.

[0126] Pruning steps: Evaluate the importance of the model weights, remove the unimportant weights, fine-tune the pruned model, and restore the performance.

[0127] In this application, the preprocessed video frames are input into the migration model, and the migration model outputs the probability distribution of the robot form categories for each frame; the output result is smoothed (such as moving average), and the category with the highest probability is selected as the final recognition result.

[0128] In this application, by combining knowledge distillation and transfer learning, the pre-trained knowledge of the large model is transferred to the small model, and its performance in the robot form recognition task is optimized.

[0129] An embodiment of this application provides a robot form recognition device based on knowledge distillation, which is used to execute a robot form recognition method based on knowledge distillation described above in this application. The following is a detailed description of the robot form recognition device based on knowledge distillation.

[0130] As Figure 2 shown, the robot form recognition device based on knowledge distillation includes:

[0131] A model acquisition module 101, which is used to acquire a pre-trained large model and a model to be trained; the pre-trained large model is a large-scale language model with a Transformer architecture;

[0132] A knowledge distillation module 102, which is used to transfer the pre-trained large model to the model to be trained based on the knowledge distillation method to obtain a migration model;

[0133] A form recognition module 103, which is used to recognize the robot form in the video stream based on the migration model.

[0134] In one implementation, the knowledge distillation module 102 is further used for:

[0135] Obtain a video data set; fine-tune the pre-trained large model based on the video data set; construct a model to be trained; and train the model to be trained through the fine-tuned pre-trained large model based on soft label distillation to obtain a migration model.

[0136] In one implementation, the knowledge distillation module 102 is further used for:

[0137] Obtain training data, where the training data is a part of the video data set; input the training data into the fine-tuned pre-trained large model to obtain corresponding soft labels; input the training data into the model to be trained to obtain prediction data; calculate the overall loss of the model to be trained based on the prediction data and the soft labels; and adjust the model to be trained based on the overall loss until the loss converges.

[0138] In one implementation, the knowledge distillation module 102 is further used for:

[0139] Based on the KL divergence, calculate the difference loss between the predicted data and the soft labels; based on the cross-entropy loss, calculate the true loss between the predicted data and the annotation of the training data; based on the difference loss and the true loss, calculate the overall loss of the model to be trained.

[0140] In one implementation, the knowledge distillation module 102 is further configured to:

[0141] Fine-tune the transfer model.

[0142] In one implementation, the knowledge distillation module 102 is further configured to:

[0143] Construct a fine-tuning data set, where the fine-tuning data set includes sample video frames with robot form annotations; freeze the shallow convolutional layers of the transfer model; train the frozen transfer model based on the fine-tuning data set until the loss converges.

[0144] In one implementation, the form recognition module 103 is further configured to:

[0145] Deploy the transfer model to a computer device or a cloud server; compress the transfer model using model quantization or pruning techniques; obtain the video stream to be recognized; input the video stream to be recognized into the transfer model to obtain the recognized robot form.

[0146] A robot form recognition device based on knowledge distillation provided by the above embodiments of the present application and a robot form recognition method based on knowledge distillation provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0147] The internal functions and structures of a robot form recognition device based on knowledge distillation are described above. As Figure 3 shown, in practice, the robot form recognition device based on knowledge distillation can be implemented as an electronic device, including: a memory 301 and a processor 303.

[0148] The memory 301 can be configured to store programs.

[0149] In addition, the memory 301 can also be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.

[0150] The memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0151] The processor 303, coupled to the memory 301, is configured to execute a program in the memory 301 for:

[0152] Obtain a pre-trained large model and a model to be trained; the pre-trained large model is a large-scale language model with a Transformer architecture;

[0153] Based on the knowledge distillation method, transfer the pre-trained large model to the model to be trained to obtain a transferred model;

[0154] Based on the transferred model, identify the robot form in the video stream.

[0155] In one embodiment, the processor 303 is further configured to:

[0156] Obtain a video data set; fine-tune the pre-trained large model based on the video data set; construct a model to be trained; based on soft label distillation, train the model to be trained through the fine-tuned pre-trained large model to obtain a transferred model.

[0157] In one embodiment, the processor 303 is further configured to:

[0158] Obtain training data, where the training data is a part of the video data set; input the training data into the fine-tuned pre-trained large model to obtain corresponding soft labels; input the training data into the model to be trained to obtain prediction data; based on the prediction data and the soft labels, calculate the overall loss of the model to be trained; adjust the model to be trained based on the overall loss until the loss converges.

[0159] In one embodiment, the processor 303 is further configured to:

[0160] Based on the KL divergence, calculate the difference loss between the prediction data and the soft labels; based on the cross-entropy loss, calculate the true loss between the prediction data and the annotation of the training data; based on the difference loss and the true loss, calculate the overall loss of the model to be trained.

[0161] In one embodiment, the processor 303 is further configured to:

[0162] Fine-tune the transferred model.

[0163] In one embodiment, the processor 303 is further configured to:

[0164] Construct a fine-tuning data set, where the fine-tuning data set includes sample video frames with robot form annotations; freeze the shallow convolutional layers of the transfer model; train the frozen transfer model based on the fine-tuning data set until the loss converges.

[0165] In one implementation, the processor 303 is further configured to:

[0166] Deploy the transfer model to a computer device or a cloud server; compress the transfer model using model quantization or pruning techniques; obtain a video stream to be recognized; input the video stream to be recognized into the transfer model to obtain the recognized robot form.

[0167] In this application, Figure 3 only some components are schematically shown, which does not mean that the electronic device only includes Figure 3 the components shown.

[0168] The electronic device provided in this embodiment and a method for recognizing robot form based on knowledge distillation provided in an embodiment of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0169] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CDROM, optical storage, etc.) containing computer-usable program code.

[0170] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0171] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0173] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0174] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read only memory (ROM) or flash memory (FlashRAM). Memory is an example of computer-readable media.

[0175] This application also provides a computer-readable storage medium corresponding to a method for robot form recognition based on knowledge distillation provided in the foregoing embodiments. A computer program (i.e., a program product) is stored thereon. When the computer program is run by a processor, it will execute a method for robot form recognition based on knowledge distillation provided in any of the foregoing embodiments.

[0176] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CDROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0177] The computer-readable storage medium provided by the above embodiments of the present application and a method for robot form recognition based on knowledge distillation provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0178] It should be noted that a large number of specific details are set forth in the specification provided herein. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0179] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0180] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A robot morphology recognition method based on knowledge distillation, characterized in that: include: Obtain a pre-trained large model and a model to be trained; the pre-trained large model is a large-scale language model of the Transformer architecture; Based on the knowledge distillation method, the pre-trained large model is migrated to the model to be trained to obtain the migration model; Recognize robot morphology in video streams based on transfer models.

2. The robot morphology recognition method based on knowledge distillation according to claim 1 is characterized in that: The method based on knowledge distillation is to migrate the pre-trained large model to the model to be trained to obtain a migration model, including: Get the video dataset; Fine-tune the pre-trained large model based on the video dataset; Build the model to be trained; Based on soft label distillation, the model to be trained is trained by the fine-tuned pre-trained large model to obtain the transfer model.

3. The robot morphology recognition method based on knowledge distillation according to claim 2 is characterized in that: Based on the soft label distillation, the model to be trained is trained by using the fine-tuned pre-trained large model to obtain the migration model, including: Acquire training data, where the training data is a part of the video data set; Input the training data into the fine-tuned pre-trained large model to obtain corresponding soft labels; Inputting the training data into the model to be trained to obtain prediction data; Calculate the overall loss of the model to be trained based on the predicted data and the soft labels; The model to be trained is adjusted based on the overall loss until the loss converges.

4. The robot morphology recognition method based on knowledge distillation according to claim 3 is characterized in that: The calculating the overall loss of the model to be trained based on the predicted data and the soft label includes: Based on the KL divergence, the difference loss between the predicted data and the soft label is calculated; Based on the cross entropy loss, calculate the true loss between the predicted data and the annotations of the training data; Based on the difference loss and the true loss, calculate the overall loss of the model to be trained.

5. The robot morphology recognition method based on knowledge distillation according to any one of claim 14, characterized in that: The method of migrating the pre-trained large model to the model to be trained based on the knowledge distillation method, and obtaining the migrated model, further includes: The migration model is fine-tuned.

6. The robot morphology recognition method based on knowledge distillation according to claim 5 is characterized in that: The fine-tuning of the migration model includes: Constructing a fine-tuning dataset, the fine-tuning dataset comprising sample video frames, the sample video frames having robot morphology annotations; Freezing the shallow convolutional layers of the migration model; The frozen migration model is trained based on the fine-tuning dataset until the loss converges.

7. The robot morphology recognition method based on knowledge distillation according to any one of claim 14, characterized in that: The method of identifying the robot form in the video stream based on the migration model includes: Deploy the migration model to a computer device or a cloud server; Compressing the migration model using model quantization or pruning techniques; Get the video stream to be identified; The video stream to be identified is input into the migration model to obtain the identified robot form.

8. A robot morphology recognition device based on knowledge distillation, characterized in that: include: A model acquisition module, which is used to acquire a pre-trained large model and a model to be trained; the pre-trained large model is a large-scale language model of the Transformer architecture; The knowledge distillation module is used to migrate the pre-trained large model to the model to be trained based on the knowledge distillation method to obtain a migration model; The morphology recognition module is used to identify the robot morphology in the video stream based on the migration model.

9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor, coupled to the memory, is configured to execute the program to: Obtain a pre-trained large model and a model to be trained; the pre-trained large model is a large-scale language model of the Transformer architecture; Based on the knowledge distillation method, the pre-trained large model is migrated to the model to be trained to obtain the migration model; Recognize robot morphology in video streams based on transfer models.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by the processor to implement a robot morphology recognition method based on knowledge distillation as described in any one of claim 17.