Arithmetic device and method for training artificial intelligence model
Through the knowledge distillation training method based on pre-trained models, combined with RT-DETR and CLIP models, high-precision and instant interactive behavior detection of artificial intelligence models in on-board systems and industrial monitoring systems is achieved, solving the problem of insufficient generalization in the existing technology, and improving the adaptability and detection efficiency of the model.
Patent Information
- Application Number
- CN202510398397.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively improve the generalization and instant detection capabilities of artificial intelligence models in on-board systems and industrial monitoring systems, especially interactive behavior recognition under multimodal data.
The knowledge distillation training method based on pretrained models is adopted, combined with the object detector with RT-DETR structure and the text encoder of the CLIP model, through the end-to-end training method, the knowledge of the pretrained model is migrated to the multi-layer perceptron, realizing high-precision and instant detection of object detection and interactive behavior decoder.
It improves the generalization and instant detection capabilities of artificial intelligence models, reduces dependence on pre-trained models, and improves flexibility and detection accuracy on the hardware platform.
Smart Images

Figure CN120338041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the application of artificial intelligence (AI), and particularly to an arithmetic device and method for improving the generalization of an artificial intelligence model (including an object detector, an interaction behavior decoder, and a multi-layer perceptron connecting the object detector and the interaction behavior decoder) by using knowledge distillation training based on a pre-trained model. Background Art
[0002] An in-vehicle system can be regarded as a computer system on a vehicle. With the development of vehicle networks, many new applications have been continuously added to the in-vehicle system, such as entertainment applications, navigation applications, communication applications, and so on. In addition, as the driver's requirement for driving safety is getting higher and higher, the current in-vehicle system also includes an Advanced Driver Assistance System (ADAS), which can continuously detect driving information inside and outside the vehicle for the driver and issue warnings or prompts immediately. For example, the advanced driver assistance system can include a cockpit monitoring system (including a driver monitoring system (DMS) and an occupant monitoring system (OMS)). An in-vehicle monitoring camera can provide real-time images of the driver / occupant. Therefore, the cockpit monitoring system can perform driving behavior analysis and passenger behavior monitoring through the image output of the in-vehicle monitoring camera. In addition to in-vehicle systems, other applications with monitoring requirements (such as industrial applications) can also achieve human behavior analysis and monitoring through the images of monitoring cameras. In recent years, driven by deep learning, artificial intelligence has achieved excellent performance in many fields. How to apply artificial intelligence technology to monitoring systems to greatly improve the functions and performance of existing monitoring systems has become an important issue. Summary of the Invention
[0003] Therefore, one of the objectives of the present invention is to propose an arithmetic device and method for improving the generalization of an artificial intelligence model (including an object detector, an interaction behavior decoder, and a multi-layer perceptron connecting the object detector and the interaction behavior decoder) by using knowledge distillation training based on a pre-trained model.
[0004] In an embodiment of the present invention, an arithmetic device for training an artificial intelligence model is disclosed. The arithmetic device includes a storage device and a processor. The storage device is used to store a program code, where the program code includes the artificial intelligence model, and the artificial intelligence model includes an object detector, an interaction behavior decoder, and a multi-layer perceptron connecting the object detector and the interaction behavior decoder. The processor is used to load and execute the program code, where the artificial intelligence model performs the following training operations: in a first stage, the object detector is first trained; and in a second stage, a part of the object detector is initialized using the parameters obtained in the first stage, and the artificial intelligence model is trained, where the training of the artificial intelligence model includes performing a knowledge distillation training using a pre-trained model to transfer the knowledge of the pre-trained model to the output of the multi-layer perceptron.
[0005] In an embodiment of the present invention, a method for training an artificial intelligence model is disclosed. The method includes: in a first stage, an object detector in the artificial intelligence model is first trained, where the artificial intelligence model includes the object detector, an interaction behavior decoder, and a multi-layer perceptron connecting the object detector and the interaction behavior decoder; and in a second stage, a part of the object detector is initialized using the parameters obtained in the first stage, and the artificial intelligence model is trained, where the training of the artificial intelligence model includes performing a knowledge distillation training using a pre-trained model to transfer the knowledge of the pre-trained model to the output of the multi-layer perceptron.
[0006] The training device and method of the artificial intelligence model of the present invention provide an end-to-end interaction behavior detection method, which integrates rich prior semantic knowledge of a multi-modal model (such as the CLIP model) to guide the interaction behavior decoder to produce more robust results. In addition, it no longer depends on the CLIP model during the subsequent inference process, so the deployment and application will be more flexible. Moreover, the object detector in the artificial intelligence model of the present invention is based on the RT-DETR structure, so it can meet the requirements of real-time detection while maintaining high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 It is a schematic diagram of an arithmetic device for training an artificial intelligence model according to an embodiment of the present invention.
[0008] Figure 2 It is a schematic diagram of the training process of an artificial intelligence model according to an embodiment of the present invention.
[0009] Figure 3 It is a schematic diagram of the architecture of an object detector based on RT-DETR.
[0010] Among them, the reference numerals are explained as follows: 100: computing device, 102: processor, 104: storage device, 106: artificial intelligence model, 108: object detector, 110: multilayer perceptron, 112: interaction behavior decoder, 202: convolutional neural network, 204: transformer encoder, 206: detection decoder, 208, 212: feed-forward neural network, 210: text encoder, PROG: program code, MD_PRE: pre-trained model, IMG_1~IMG_N: images, TXT_1~TXT_N: texts. Detailed implementation manners
[0011] Figure 1 It is a schematic diagram of a computing device for training an artificial intelligence model according to an embodiment of the present invention. The computing device 100 includes a processor 102 and a storage device 104. The storage device 104 stores a program code PROG, where the program code PROG includes an artificial intelligence model 106. In this embodiment, the artificial intelligence model 106 is an end-to-end human-object-interaction detection model, including an object detector 108, an interaction decoder 112, and a multilayer perception (MLP) 110 connecting the object detector 108 and the interaction decoder 100. The processor 102 loads and executes the program code PROG, and performs training operations on the artificial intelligence model 106. The training process of the artificial intelligence model 106 can be divided into two stages. In the first stage, the object detector 108 is trained first. In the second stage after the first stage, a part of the object detector 108 is initialized with the parameters obtained in the first stage, and the artificial intelligence model 106 is trained. The training of the artificial intelligence model 106 in the second stage includes performing a knowledge distillation training using a pre-trained model MD_PRE to transfer the knowledge of the pre-trained model MD_PRE to the output of the multilayer perceptron 110 (that is, transferring the knowledge learned by a large and complex model to a small and simple model through knowledge distillation). Note that the pre-trained model MD_PRE is only used in the training stage of the artificial intelligence model 106. After training to a certain extent, the trained artificial intelligence model 106 is deployed to a hardware platform, and the hardware platform is responsible for processing a large amount of unlabeled data according to the trained artificial intelligence model 106 in the inference stage. Since the artificial intelligence model 106 deployed to the hardware platform no longer depends on the pre-trained model MD_PRE, the artificial intelligence model 106 can adapt to more hardware platforms.
[0012] In some embodiments of the present invention, the object detector 108 may be an object detector based on a real-time detection transformer (hereinafter simply referred to as "RT-DETR"), and in the second stage, a part of the object detector 108 that initializes using the parameters obtained in the first stage is a detection decoder in the object detector based on the real-time detection transformer. Additionally, the pre-trained model MD_PRE may be a multimodal model based on contrastive learning, such as a Contrastive Language-Image Pre-training (hereinafter simply referred to as "CLIP") model. Therefore, the output of the detection decoder in the object detector based on RT-DETR can be used as the input to the multi-layer perceptron 110, and knowledge distillation training can utilize the output of a text encoder of the CLIP model.
[0013] To facilitate the description of the technical features of the present invention, the object detector based on RT-DETR and the text encoder of the CLIP model are used as examples hereinafter. However, the present invention is not limited thereto. In fact, the object detector 108 may also be an object detector using other methods, and the pre-trained model MD_PRE may also be a large model using other methods. These design variations all fall within the scope of the present invention. The training operations of the artificial intelligence model 106 disclosed in the present invention will be further described below in conjunction with the accompanying drawings.
[0014] Figure 2Schematic diagram of the training process of the artificial intelligence model 106 according to an embodiment of the present invention. The entire training process is divided into two stages. The first stage is to train the detection decoder 206 in the object detector based on RT-DETR, mainly to extract the human box, object box, object class, and interactive confidence score. The second stage is to use the parameters of the first stage to initialize the detection decoder 206. At the same time, the output generated by the text encoder 210 of the pre-trained CLIP model is used for knowledge distillation to extract rich semantic features, and then the sequence features output by the detection decoder 206 are fused as the target query, and finally the interactive behavior classification result is output through the interactive behavior decoder 112. The overall DETR structure adopts RT-DETR. Since RT-DETR introduces and draws on many optimization methods for DETR in recent years, it can achieve the benchmark of an instant detector while maintaining high precision.
[0015] The first stage completes the training of object detection (especially the detection decoder 206) through an object detector based on RT-DETR, so that the second stage can use the parameters obtained in the first stage to initialize the detection decoder 206. As Figure 2 shown in the lower part of, after the input images IMG_1~IMG_N are preprocessed, the feature maps {s3, s4, s5} are extracted by the Convolutional Neural Network (hereinafter simply referred to as "CNN") 202 and input into the transformer encoder 204. Please refer to Figure 3 , Figure 3Schematic diagram of the architecture of an object detector based on RT-DETR. The transformer encoder 204 of the object detector based on RT-DETR is a hybrid encoder, which can be composed of Attention-based Intrascale Feature Interaction (AIFI) and CNN-based Cross-scale Feature Fusion (CCFF). In order to reduce the computational load, only the standard encoding operation is performed on the last layer feature map s5, while the other feature maps s3 and s4 only participate in the feature fusion calculation in CCFF. The sequence feature Xd output by this hybrid encoder is input to the detection decoder 206, and another input Qd of the detection decoder 206 is the joint query of human-object, which encodes the feature information of the human-object pair. In addition, the sequence feature Xd can go through the IoU (intersection over union)-Aware Query Selection step to select the target queries with high classification scores and high IoU scores to initialize Qd. In the training stage, in order to solve the ambiguity of Hungarian matching (that is, the discreteness of the matching of the Hungarian Algorithm and the randomness of model training) resulting in unstable matching and slow convergence, additional queries can be introduced to learn the way of adding noise to the gt (ground truth), just like the de-noising method in DN-DETR (DeNoising-DETR). The detection decoder 206 adopts deformable self-attention, which can greatly reduce the computational load. In addition, the sequence feature output by the detection decoder 206 passes through a Feedforward Neural Network (hereinafter simply referred to as "FNN") 208 (which serves as a prediction head) to generate the final human bounding box, object bounding box, object category, and interaction behavior confidence scores.
[0016] The loss is calculated through one-to-one Hungarian matching. The calculation of the regression loss can be expressed as box_loss = L1_loss + GIoU_loss. The box classification loss uses the cross-entropy loss box_cls_loss, and the loss for whether there is an interaction behavior, interact_loss, uses BCE_loss. Therefore, the calculation of the final loss in the first stage can be expressed as step1_loss = box_loss + box_cls_loss + interact_loss. After the first stage of training, the detection decoder 206 can output the {human box, object box} pairs where an interaction behavior occurs.
[0017] In the second stage, knowledge distillation training is carried out using the text encoder 210 of the pre-trained CLIP model, and the final interaction behavior classification is completed. Interaction behavior recognition is a recognition task of a {human, object, interaction behavior} triple, which is more complex than general detection tasks. Traditional methods would adopt a two-stage approach for training, that is, first train a detector model to complete the object detection task, and then train the interaction behavior recognition model when the human position and object position are known. This method requires two models in actual deployment, which is usually not friendly to platforms with limited resources. Although this method is based on the object detection in the first stage, it only uses it for initialization, and ultimately the entire detection process is carried out end-to-end (that is, the artificial intelligence model 106 of the present invention is an end-to-end model). As Figure 2 shown in the upper part of, this method connects an interaction behavior decoder 112 behind the detection decoder 206 of the object detector based on RT-DETR, uses the human-object pairing coding information output by the detection decoder 206 as the benchmark target query, and at the same time, through the method of knowledge distillation, integrates the rich prior text semantic knowledge of the CLIP model to decode the interaction behavior occurring in the human-object pairing.
[0018] First, for the input images IMG_1 to IMG_N, image-text annotations are made to generate the corresponding texts TXT_1 to TXT_N for the images. After preprocessing by the tokenizer, the serialized text features Fi are generated through the text encoder 210 of the pre-trained CLIP model. Since the CLIP model is built on the training of a large number of images and has a strong zero-shot generalization ability for object classification and behavior recognition, at this time, Fi already has the ability to express rich high-level semantic features. Combining the text semantic features of Fi will be very beneficial to the subsequent classification and recognition of interaction behaviors.
[0019] Secondly, the parameters of the detection decoder 206 obtained in the first stage are used to initialize the detection decoder 206 in the second stage. The detection decoder 206 outputs the serialized feature Qd_out (which contains position information). Qd_out generates the feature Qi through the multi-layer perceptron 110. Then, the L1 loss (L1loss) is imposed on Fi and Qi for supervised knowledge distillation training to transfer the prior semantic knowledge in Fi to Qi (that is, to make Qi as close as possible to Fi). Then, the features of Qi and Qd_out are fused to generate Qa. In other words, through knowledge distillation training, the present invention can transfer the knowledge of the pre-trained CLIP model (especially the text feature Fi generated by the text encoder 210) to the output Qa of the multi-layer perceptron 110. The fusion process can be expressed as: Qi = MLP(Qd_out) and Qa = MLP(concat(Qd_out, Qi)). In addition, the calculation of the distillation loss can be expressed as At this time, Qa not only has the human-object pairing information but also absorbs the text prior knowledge of the CLIP model, providing key guiding information for the subsequent classification and recognition of interaction behaviors.
[0020] Next, the interaction behavior decoder 112 uses Qa (which contains action information) as the target query and Xd (which contains image information) as the input, and finally outputs the sequence feature Qa_out (which contains interaction behavior information). Similarly, the deformable attention is also used as the interaction attention in the interaction behavior decoder 112 for calculation.
[0021] Finally, Qa_out passes through the feed-forward neural network 212 to generate the prediction result of the final interaction behavior class. During training, the corresponding loss action_cls_loss uses the cross-entropy loss. The calculation of the second-stage loss can be expressed as step2_loss = distillation_loss + action_cls_loss. In addition, the calculation of the overall loss can be expressed as total_loss = step1_loss + step2_loss.
[0022] In summary, the training device and method of the artificial intelligence model of the present invention provide an end-to-end interaction behavior detection method, which integrates the rich prior semantic knowledge of the multi-modal model (such as the CLIP model) to guide the interaction behavior decoder to produce more robust results. In addition, it does not depend on the CLIP model during the subsequent inference process, so the deployment and application will be more flexible. Moreover, the object detector in the artificial intelligence model of the present invention is based on the RT-DETR structure, so it can meet the requirements of real-time detection while maintaining high accuracy.
[0023] The trained artificial intelligence model provided by the present invention can be deployed to the hardware platform of the vehicle-mounted system. For example, the trained artificial intelligence model can be applied to the cockpit monitoring system (including the driving monitoring system and the passenger monitoring system) for driving behavior analysis and passenger behavior monitoring, and can also be applied to the advanced driver assistance system to detect the driving information inside and outside the vehicle and issue warnings or prompts immediately. In addition to the vehicle-mounted system, the trained artificial intelligence model provided by the present invention can also be deployed to any hardware platform providing monitoring functions. For example, the trained artificial intelligence model can be applied to industrial security monitoring to provide human behavior retrieval and identification and tracking of dangerous actions.
[0024] The above are only the preferred embodiments of the present invention, but they are not intended to limit the scope of the present invention. Any person familiar with this technology can make further improvements and changes on the basis without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the scope defined by the claims of this application.
Claims
1. An arithmetic device for training an artificial intelligence model, characterized in that, Comprising: A storage device for storing program code, where the program code includes the artificial intelligence model, and the artificial intelligence model includes an object detector, an interaction behavior decoder, and a multi-layer perceptron connecting the object detector and the interaction behavior decoder; And A processor for loading and executing the program code, where the artificial intelligence model performs the following training operations: In the first stage, the object detector is trained first; And In the second stage, a part of the object detector is initialized with the parameters obtained in the first stage, and the artificial intelligence model is trained, where the training of the artificial intelligence model includes using a pre-trained model to perform knowledge distillation training to transfer the knowledge of the pre-trained model to the output of the multi-layer perceptron.
2. The computing device according to claim 1, wherein the object detector is an object detector based on an instant detection transformer.
3. The computing device according to claim 2, wherein a part of the object detector is the detection decoder of the object detector based on the instant detection transformer.
4. The computing device according to claim 3, wherein the pre-trained model is a multi-modal model based on contrastive learning.
5. The computing device according to claim 4, wherein the multi-modal model based on contrastive learning is a contrastive language-image pre-training model.
6. The computing device according to claim 5, wherein the output of the detection decoder is used as the input of the multi-layer perceptron, and the knowledge distillation training uses the output of the text encoder of the contrastive language-image pre-training model.
7. A method for training an artificial intelligence model, characterized in that, Comprising: In the first stage, the object detector in the artificial intelligence model is trained first, where the artificial intelligence model includes the object detector, the interaction behavior decoder, and the multi-layer perceptron connecting the object detector and the interaction behavior decoder; And In the second stage, a part of the object detector is initialized with the parameters obtained in the first stage, and the artificial intelligence model is trained, where the training of the artificial intelligence model includes using a pre-trained model to perform knowledge distillation training to transfer the knowledge of the pre-trained model to the output of the multi-layer perceptron.
8. The method according to claim 7, wherein the object detector is an object detector based on an instant detection transformer.
9. The method according to claim 8, wherein a part of the object detector is the detection decoder of the object detector based on the instant detection transformer.
10. The method according to claim 9, wherein the pre-trained model is a multi-modal model based on contrastive learning.
11. The method according to claim 10, wherein the multi-modal model based on contrastive learning is a contrastive language-image pre-training model.
12. The method according to claim 11, wherein the output of the detection decoder is used as the input of the multi-layer perceptron, and the knowledge distillation training uses the output of the text encoder of the contrastive language-image pre-training model.