Model training method, target detection method and device
By using multimodal data to train the object detection model and using mutual distillation and self-distillation losses for model training, the personnel detection and tracking problems of single-modal data limitation are solved, achieving higher detection accuracy and model robustness.
Patent Information
- Application Number
- CN202111537999.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-15
AI Technical Summary
The prior art is limited by single-modal data in personnel detection and tracking, resulting in limited detection accuracy and model robustness.
Multimodal data is used to train the object detection model, and the model is trained through mutual distillation and self-distillation losses, and multiple object detection models are fused to obtain the final object detection model.
Improve feature quality, optimize training process, enhance the robustness of the model, thereby improving the accuracy of personnel detection and tracking.
Smart Images

Figure CN114492563B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a model training method, a target detection method and a device thereof. Background Art
[0002] With the development of the economy and the acceleration of urbanization, the size of urban population continues to expand. How to manage people in public places in an orderly manner has gradually become a difficult problem. Intelligent tracking of people is undoubtedly one of the most effective ways.
[0003] The purpose of the personnel detection task is to determine the location and size of the personnel, so as to be used for subsequent target trajectory analysis. At present, hardware technology and deep learning technology are constantly innovating, and the application scenarios of personnel tracking are also constantly expanding. In the military field, it can be applied to battlefield detection, in the security field, it can be applied to suspect positioning, and in the field of urban transportation, it can be applied to automatic driving. At the same time, it plays a very important role in regulating traffic, reducing vehicle accidents, and improving vehicle flow efficiency. Pedestrian tracking has therefore become a research hotspot. A practical personnel tracking system must have the feature of full-time detection, but most of the current visual systems are based on visible light.
[0004] Using only single-modal data, such as visible light data, will only collect limited information in complex scenarios. Since single-modal data can only provide limited information, the accuracy of person detection and tracking is limited, which also has a certain impact on the training of related models and reduces the robustness of the models. Summary of the invention
[0005] The present application provides a model training method, a target detection method and a device thereof.
[0006] In order to solve the above technical problems, the first technical solution provided by the present application is: to provide a model training method, the model training method comprising:
[0007] Acquire training images of different modalities, wherein the training images of different modalities are obtained based on different imaging principles;
[0008] Inputting the training images of different modalities into different target detection models to obtain output results of each target detection model, wherein the number of the target detection models is consistent with the number of modalities;
[0009] Based on an output result of an object detection model and output results of other object detection models, obtain the mutual distillation loss of the object detection model;
[0010] Performing model training according to the mutual distillation loss of the target detection model;
[0011] Based on multiple trained target detection models, the final target detection model is obtained.
[0012] Wherein, the model training method further includes:
[0013] Based on the output result of an object detection model, obtain the self-distillation loss of the object detection model;
[0014] The model training is performed according to the mutual distillation loss and self-distillation loss of the target detection model.
[0015] The step of obtaining the self-distillation loss of a target detection model based on an output result of the target detection model includes:
[0016] Obtaining a current output result of the target detection model for the training image based on current model parameters, and a historical output result of the target detection model for the same training image based on historical model parameters;
[0017] Based on the current output result and the historical output results, a self-distillation loss of the target detection model is generated.
[0018] The historical output result is the average value of the training output results of the target detection model before the current training time.
[0019] Wherein, the target detection model includes a feature extractor and a target detector;
[0020] The model training method further includes:
[0021] Inputting a training image of a first modality into a first object detection model, obtaining a first feature map extracted by a first feature extractor from the training image of the first modality, and a first object detection result of a first object detector based on the first feature map;
[0022] Obtaining a self-distillation loss of the first feature extractor based on the first feature map, and obtaining a detector loss of the first object detector based on the first object detection result;
[0023] Inputting the training image of the second modality into the second object detection model, obtaining a second feature map extracted by a second feature extractor from the training image of the second modality, and a second object detection result of a second object detector based on the second feature map;
[0024] Obtaining a self-distillation loss of the second feature extractor based on the second feature map, and obtaining a detector loss of the second object detector based on the second object detection result;
[0025] Determining a mutual distillation loss of the first feature extractor based on the first feature map and the other second feature map;
[0026] determining a mutual distillation loss of the second feature extractor based on the second feature map and other first feature maps;
[0027] Performing model training on the first target detection model according to the self-distillation loss, the mutual distillation loss, and the detector loss of the first target detection model;
[0028] The second target detection model is trained according to the self-distillation loss, the mutual distillation loss, and the detector loss of the second target detection model.
[0029] The first feature extractor and the second feature extractor have the same structure, and each includes a plurality of feature extraction blocks;
[0030] The determining the mutual distillation loss of the first feature extractor based on the first feature map and the other second feature map includes:
[0031] Obtaining a first feature subgraph output by each feature extraction block in the first feature extractor;
[0032] Obtaining a second feature subgraph output by each feature extraction block in the second feature extractor;
[0033] Determine the mutual distillation loss of the feature extraction block at the same position according to the first feature subgraph and the second feature subgraph output by the feature extraction block at the same position;
[0034] Based on the mutual distillation losses of all feature extraction blocks in the first feature extractor, the mutual distillation loss of the first feature extractor is determined.
[0035] Wherein, the target detection model includes a feature extractor and a target detector;
[0036] The trained target detection models are fused to obtain the final target detection model, including:
[0037] The outputs of feature extractors of several object detection models are stacked according to channels and input into a unified object detector to obtain the final object detection model.
[0038] The training images of several modalities are input into a final object detection model to train the object detector in the final object detection model.
[0039] Wherein, after fusing the trained target detection models to obtain the final target detection model, the method further includes:
[0040] Constructing a positive sample library, wherein the positive sample library includes a number of positive samples whose detection confidence is greater than a first preset threshold;
[0041] Inputting the test image into the final object detection model to output a detection result of the test image, wherein the detection result includes a confidence level;
[0042] Obtaining the maximum correlation between the test image and the positive samples in the positive sample library;
[0043] Calculating a correlation confidence of the test image based on the confidence and the maximum correlation;
[0044] Determining whether the relevant confidence of the test image is greater than or equal to a second preset threshold;
[0045] If so, determining that the detection result of the test image is a positive detection;
[0046] If not, it is determined that the detection result of the test image is a false detection.
[0047] In order to solve the above technical problems, the second technical solution provided by the present application is: to provide a target detection method, the target detection method comprising:
[0048] Acquire the image to be detected;
[0049] Inputting the image to be detected into a pre-trained target detection model, wherein the target detection model is trained by the above-mentioned model training method, and the modality of the image to be detected is any modality among the different modalities;
[0050] The target in the image to be detected is detected based on the target frame output by the target detection model.
[0051] The detecting of the target in the image to be detected based on the target frame output by the target detection model includes:
[0052] Identify a current target in a current image to be detected based on a target frame output by the target detection model;
[0053] Identify other targets in the area where the current target is located in the previous frame of the image to be detected;
[0054] Calculating the correlation between the other targets and the current target;
[0055] Determining whether the correlation is greater than or equal to a third preset threshold;
[0056] If yes, setting the target identifier of the current target according to the target identifiers of the other targets;
[0057] If not, reallocate the target ID.
[0058] In order to solve the above technical problems, the third technical solution provided by the present application is: to provide a model training device, the model training device includes an acquisition module, an input module, a distillation module and a training module; wherein,
[0059] The acquisition module is used to acquire training images of several modalities;
[0060] The input module is used to input training images of different modalities into different target detection models to obtain output results of each target detection model, wherein the number of the target detection models is consistent with the number of modalities;
[0061] The distillation module is used to obtain the mutual distillation loss of the target detection model based on the output result of one target detection model and the output results of other target detection models;
[0062] The training module is used to perform model training according to the mutual distillation loss of the target detection model; and obtain a final target detection model based on multiple trained target detection models.
[0063] To solve the above technical problems, the fourth technical solution provided in the present application is: to provide a terminal device, the terminal device includes a processor and a memory connected to the processor, wherein the memory stores program instructions; the processor is used to execute the program instructions stored in the memory to implement the above-mentioned model training method and / or target detection method.
[0064] In order to solve the above technical problems, the fifth technical solution provided in the present application is: providing a computer-readable storage medium, wherein the storage medium stores program instructions, and when the program instructions are executed, the above-mentioned model training method and / or target detection method are implemented.
[0065] In the model training method provided by the present application, a terminal device obtains training images of different modalities, wherein the training images of different modalities are obtained based on different imaging principles; the training images of different modalities are input into different target detection models to obtain the output results of each target detection model, wherein the number of target detection models is consistent with the number of modalities; based on the output results of one target detection model and the output results of other target detection models, the mutual distillation loss of the target detection model is obtained; the model is trained according to the mutual distillation loss of the target detection model; based on multiple trained target detection models, the final target detection model is obtained. The target detection method of the present application introduces multimodal data and uses mutual distillation between modalities for model training, which can effectively improve feature quality, optimize the training process, and improve model robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0067] Figure 1 It is a flowchart of an embodiment of a model training method provided by the present application;
[0068] Figure 2 It is a flowchart of a multi-task oriented real-time tracking system provided by the present application;
[0069] Figure 3 It is a flowchart of another embodiment of the model training method provided by the present application;
[0070] Figure 4 It is a schematic diagram of the framework of the target detection model provided in this application;
[0071] Figure 5 It is a flowchart of an embodiment of a target detection method provided by the present application;
[0072] Figure 6 yes Figure 5 A schematic diagram of a specific process of step S33 in the target detection method shown;
[0073] Figure 7 It is a structural diagram of an embodiment of a terminal device provided by the present application;
[0074] Figure 8 is a structural schematic diagram of another embodiment of a terminal device provided by the present application;
[0075] Fig. 9 It is a schematic diagram of the structure of the computer-readable storage medium provided in this application. DETAILED DESCRIPTION
[0076] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0077] The present application is described in detail below with reference to the accompanying drawings and embodiments.
[0078] See also Figure 1 and Figure 2 , Figure 1is a flow chart of an embodiment of a model training method provided by the present application, Figure 2 It is a flowchart of a multi-task oriented real-time tracking system provided by the present application. Among them, the model training method described in the embodiment of the present application is applied to a terminal device, wherein the terminal device of the present application can be a server, or a system composed of a server and a local terminal cooperating with each other. Accordingly, the various parts included in the terminal device, such as various units, sub-units, modules, and sub-modules can all be set in the server, or can be set in the server and the local terminal respectively.
[0079] Furthermore, the above-mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules for providing distributed servers, or it can be implemented as a single software or software module, which is not specifically limited here. In some possible implementations, the target detection method of the embodiment of the present application can be implemented by a processor calling a computer-readable instruction stored in a memory.
[0080] like Figure 1 As shown, the specific steps of the model training method in the embodiment of the present application are as follows:
[0081] Step S11: acquiring training images of different modalities, wherein the training images of different modalities are obtained based on different imaging principles.
[0082] In the embodiment of the present application, the training images can be obtained by using an image acquisition device to acquire images of relevant areas. The image acquisition device may include, but is not limited to, one or more of a gun camera, a ball camera, a panoramic camera, etc. Any one or more of the above devices are used to monitor the crowd area, and a certain number of training images are acquired for the above crowd area.
[0083] As an embodiment, the user may also define a regular area, that is, the user may, but is not limited to, determine a part of a scene / area that the user is interested in in a certain scene as the above-mentioned relevant area, and then collect images for the relevant area.
[0084] Specifically, the camera can be installed top-mounted or diagonally, and the staff can use different installation methods according to the actual scene. For example, for installation heights greater than 2.5 meters, it is recommended to install the camera diagonally, and for installation heights less than 2.5 meters, it is recommended to install the camera top-mounted. The standard for the camera is to require the camera to capture the entire body of the person as much as possible, so that the posture can be detected later, and the surveillance video can be analyzed. The staff can also select a regular area in the surveillance video screen to indicate the area to be detected by the terminal device and the tracking system. The regular area allows polygons, and the specific shape can be manually entered by the staff or the default regular area can be set.
[0085] The camera in the embodiment of the present application can be a collection of multiple image acquisition devices, and each image acquisition device can collect training images of one modality. For example, most of the current visual systems are based on visible light, so common cameras can be used to collect training images of visible light modality. However, the above-mentioned visual system has great disadvantages when detecting at night. Infrared light can work normally at night, has the ability to penetrate fog, haze, and smoke, and is not affected by flash and strong light. Therefore, in the embodiment of the present application, images of the two modalities can be collected by combining visible light and infrared light, and then the impact of the above problems can be reduced based on the images of the two modalities; that is, in the embodiment of the present application, an infrared camera can be used to collect training images of infrared light modality. In addition, low-light technology has better adaptability in various environments, and corresponding training images of low-light modality can be collected; depth images provide three-dimensional information of the scene, and corresponding training images of depth modality can be collected; SAR (Synthetic Aperture Radar) images have unique phase information that cannot be obtained by other sensors; image information obtained by different sensors contains rich information, and corresponding training images of different sensor modalities can be collected.
[0086] Among them, the above imaging principles include but are not limited to: infrared imaging, X-ray imaging, visible light imaging, low-light imaging, depth imaging, etc.
[0087] Therefore, in order to further broaden the application scenarios and improve the detection quality, the embodiments of the present application may use at least two or more training images of different modalities as training data.
[0088] In the embodiment of the present application, the terminal device uses the training images of multiple modalities collected in step S11 to implement Figure 2 For the multimodal distillation training in , please refer to step S12 to step S15 for the specific training process.
[0089] Step S12: input training images of different modalities into different target detection models to obtain output results of each target detection model, wherein the number of target detection models is consistent with the number of modalities.
[0090] In an embodiment of the present application, it is assumed that the training set includes training images of K modalities, and the image sequences or video frame sequences correspond one to one. For this, the terminal device constructs K target detection models respectively. The terminal device inputs the training images of K modalities into the corresponding K target detection models respectively to obtain the output results of each target detection model, including detection frames and / or detection classifications generated based on the features in the training images. Among them, the modality type and the target detection model correspond one to one.
[0091] In an embodiment of the present application, for the input of multimodal training images, the terminal device can adopt a combination of modality mutual distillation and self-distillation for training. This method can not only learn the complementary knowledge between different modalities, but also enable the model parameters of the target detection model to be optimally and stably updated during training.
[0092] Among them, self-distillation means that the teacher model and the student model are one model, that is, one model guides itself to learn and complete knowledge distillation. Mutual distillation means that the teacher model and the student model are two different models, that is, using other models to guide itself to learn and complete knowledge distillation.
[0093] Specifically, each target detection model needs to be trained on multiple frames of training images of the same modality. At the current moment, the terminal device obtains the current output result of the target detection model for one of the training images based on the current model parameters, that is, the target detection model uses the current model parameters to extract features from one of the training images to obtain the current feature map, and the terminal device can also obtain the historical output results of the target detection model for the same training image based on the historical model parameters, that is, the target detection model uses the historical model parameters to extract features from the same training image to obtain the historical feature map. The terminal device can generate the self-distillation loss of the target detection model at the current moment based on the gap between the current output result and the historical output result, that is, the difference between the current feature map and the historical feature map.
[0094] The historical model parameters may be the model parameters of the target detection model at the previous moment before the current moment, or may be the average value of the model parameters of the target detection model in the previous period before the current moment. The model parameters of the target detection model in the previous period include multiple ones, and in this case, the average value of the multiple model parameters may be calculated as the historical model parameters.
[0095] Step S13: Based on the output result of one target detection model and the output results of other target detection models, obtain the mutual distillation loss of the target detection model.
[0096] In an embodiment of the present application, when a terminal device trains a target detection model, it needs to obtain the output result of the target detection model on the one hand, and the output result of other target detection models other than the target detection model on the other hand. The terminal device can generate the mutual distillation loss of the target detection model based on the gap between the output result of the target detection model and the output result of other target detection models.
[0097] Step S14: Perform model training according to the mutual distillation loss of the target detection model.
[0098] In an embodiment of the present application, the terminal device can perform model training according to the mutual distillation loss of the target detection model.
[0099] Furthermore, the terminal device can also construct a loss function of the target detection model according to the self-distillation loss and mutual distillation loss of the target detection model, thereby performing model training on the target detection model according to the constructed loss function.
[0100] Step S15: Fusing the trained target detection models to obtain a final target detection model.
[0101] In the embodiment of the present application, according to Figure 2 In the decision fusion step, the terminal device fuses the multiple target detection models based on different modalities trained in step S14 to construct a final target detection model.
[0102] Among them, each target detection model includes a feature extractor and a target detector. The feature extractor is mainly used to extract image features of training images, and the target detector is mainly used to detect targets in training images based on image features to generate corresponding detection frames.
[0103] After the terminal device of the embodiment of the present application is trained to obtain multiple target detection models, a simple decision-level fusion can be performed. Specifically, the terminal device can filter and fuse the outputs of all target detectors by means of WBF (weighted boxes fusion). The above method is simple and convenient, and the fusion difficulty is relatively low. In addition, in order to improve the fusion effect, the terminal device can also adopt a feature-level fusion method. For example, the terminal device can stack the outputs of multiple feature extractors according to channels, and then send them to a unified target detector for retraining to obtain the final target detection model. During the feature-level fusion training process, the terminal device needs to fix the model parameters of the feature extractor part, and only needs to fine-tune the model parameters of the unified target detector part. In addition, the loss function can use FocalLoss or other known mature loss functions.
[0104] Furthermore, although the final target detection model obtained by the above training can result in information of multiple modalities, there is still a certain risk of false detection, such as fire hydrants on the road, chairs in indoor scenes, etc. The false detection confidence of targets in such scenes is relatively high, which cannot be solved by simply relying on network optimization and increasing the confidence threshold.
[0105] Therefore, the embodiment of the present application adopts maintaining a positive sample library, and by comparing with the positive samples in the positive sample library, the weight of the confidence of target detection is weakened, which can reduce the occurrence of false detection to a certain extent.
[0106] Specifically, the establishment of the positive sample library may also depend on the detection results of the target detection model. For example, if the sample confidence output by the target detection model is greater than a first preset threshold, such as 0.95, then the sample is considered to be a positive sample and can be included in the positive sample library, thereby establishing and continuously updating the positive sample library.
[0107] In the specific prediction process, assuming that the detection confidence of a picture to be tested is 0.7, the terminal device measures the correlation between the picture to be tested and the positive sample pictures in the positive sample library. The correlation measurement can be calculated using KL divergence, L2 distance, correlation filtering, etc., and the value with the highest correlation is taken as the output. If the output of the correlation is 0.4, the final confidence of the picture to be tested is 1.1. Although the detection confidence of the picture to be tested is higher than the first confidence threshold, such as 0.6, it can be attributed to the positive sample, but the final confidence of the picture to be tested is lower than the second confidence threshold, such as 1.5. At this time, the terminal device can believe that the picture to be tested has a false detection.
[0108] Based on the fact that the confidence of the positive detection and the matching degree of the positive sample library are both very high, the terminal device can further detect whether there is a false detection in the detection result by setting a second confidence threshold, thereby achieving the purpose of reducing false detections.
[0109] It should be noted that the above processes are all based on the premise of the final target detection model, that is, one target detection model is used to detect multiple modal images to be detected. In other embodiments, according to the detection results, sample data can be obtained in different modalities, and the library construction and calculation of the matching degree of the sample library can also be performed on the basis of multiple modalities. This process will not be repeated here.
[0110] In order to further reduce false detection, the terminal device can also perform a series of combined filtering on the detection results. Because the tracking system in this application uses multimodal input, there are many ways to reduce false detection. For example, for different scenes, the filtering conditions in the RGB image include but are not limited to the target size, target color, target motion state, positional relationship between targets, etc. In addition, it also includes the target temperature in the infrared image, the target distance in the depth image, etc.
[0111] The embodiment of the present application introduces multimodal data in multi-target tracking, combines the mutual distillation between multimodal data and the self-distillation within the modality, and proposes a two-stage training method, which can effectively improve the feature quality, optimize the training process, and improve the robustness of the model; in addition, in response to the false detection problem in pedestrian tracking, a post-processing method of positive sample library comparison and multimodal data combination filtering is proposed, which effectively reduces the number of false detections.
[0112] In an embodiment of the present application, a terminal device obtains training images of several modalities; inputs training images of different modalities into different target detection models to obtain the output results of each target detection model, wherein the number of target detection models is consistent with the number of modalities; based on the output results of a target detection model, obtains the self-distillation loss of the target detection model; based on the output results of a target detection model and the output results of other target detection models, obtains the mutual distillation loss of the target detection model; performs model training according to the self-distillation loss and mutual distillation loss of the target detection model; fuses the trained target detection models to obtain the final target detection model. The target detection method of the present application introduces multimodal data and combines mutual distillation and self-distillation between modalities, which can effectively improve feature quality, optimize the training process, and improve model robustness.
[0113] Since the object detection model can be divided into feature extractors and object detectors according to their functions, the following further introduces the specific training process of self-distillation and mutual distillation from the perspective of feature extractors and object detectors. Figure 3 and Figure 4 , Figure 3 is a flow chart of another embodiment of the model training method provided by the present application, Figure 4 It is a schematic diagram of the framework of the target detection model provided in this application.
[0114] like Figure 4 As shown, the target detection model of the embodiment of the present application includes a feature extractor and a target detector, wherein the structure of the feature extractor can adopt different structures, but the target detector must adopt the same structure, that is, Figure 4 The detection head in.
[0115] Specifically, if a feature extractor with a different structure is used, the terminal device can Figure 4 The feature extractor output in the output is mutual distillation; if the feature extractor with the same structure is used, the terminal device can Figure 4 Each feature extraction block of the feature extractor, that is, the output of each Block, is distilled, and finally the output distillation losses of all Blocks are added together to complete the mutual distillation.
[0116] like Figure 3 As shown, the specific steps of the model training method in the embodiment of the present application are as follows:
[0117] Step S21: input the training image of the first modality into the first target detection model, obtain the first feature map extracted by the first feature extractor from the training image of the first modality, and the first target detection result of the first target detector based on the first feature map.
[0118] In an embodiment of the present application, the terminal device inputs a training image of a first modality, i.e., modality A, into a first target detection model, and obtains a first feature map output by a first feature extractor at an output, and a first target detector, i.e., a first target detection result output by a detection head based on the first feature map.
[0119] Step S22: Obtain a self-distillation loss of a first feature extractor based on the first feature map, and obtain a detector loss of a first object detector based on the first object detection result.
[0120] In an embodiment of the present application, the terminal device obtains the self-distillation loss of the first feature extractor based on the first feature map. The specific formula of the self-distillation loss is as follows:
[0121]
[0122] Among them, Output A represents the output result of the first feature extractor based on the model parameters at the current moment, Represents the output result of the first feature extractor based on the average value of the model parameters at historical moments.
[0123] The self-distillation loss uses the average value of the model parameters before the current moment to constrain the model parameters at the current moment. The advantage of this method is that it enhances the stability of the training process and obtains a high mean accuracy and a small variance.
[0124] In the embodiment of the present application, the terminal device obtains the detector loss of the first target detector based on the first target detection result, that is, L det , represents the loss function generated by the detection head. The detector loss can be constructed by the difference between the output result of the detection head and the actual annotation result.
[0125] Step S23: input the training image of the second modality into the second target detection model, obtain the second feature map extracted by the second feature extractor from the training image of the second modality, and the second target detection result of the second target detector based on the second feature map.
[0126] Step S24: Obtain a self-distillation loss of a second feature extractor based on the second feature map, and obtain a detector loss of a second object detector based on the second object detection result.
[0127] In the embodiment of the present application, steps S23 to S24 are substantially the same as the above-mentioned steps S21 to S22 and are not described again herein.
[0128] Step S25: Determine the inter-distillation loss of the first feature extractor based on the first feature map and the other second feature maps.
[0129] In an embodiment of the present application, the terminal device determines the mutual distillation loss of the first feature extractor based on the first feature map and other second feature maps, where the other second feature maps refer to feature maps output by all feature extractors except the feature extractor of the first feature map.
[0130] Among them, the specific formula for the mutual distillation loss is as follows:
[0131]
[0132] Among them, Output A Represents the first feature map, Output m Represents the second feature map. The mutual distillation loss means using the feature maps of other modalities to constrain the feature extractor where the first feature map of modality A is located. The advantage of this method is that modality A can indirectly learn the knowledge of other modal data and extract more robust features.
[0133] Furthermore, when the feature extractors of the same structure are used in multiple target detection models, the terminal device can also obtain the feature subgraph of each feature extraction block in the feature extractor, thereby calculating the mutual distillation loss of each feature extraction block, and adding them to obtain the mutual distillation loss of the feature extractor. The specific process is as follows: obtain the first feature subgraph output by each feature extraction block in the first feature extractor; obtain the second feature subgraph output by each feature extraction block in the second feature extractor; determine the mutual distillation loss of the feature extraction block at the same position according to the first feature subgraph and the second feature subgraph output by the feature extraction block at the same position; determine the mutual distillation loss of the first feature extractor based on the mutual distillation loss of all feature extraction blocks in the first feature extractor.
[0134] Step S26: Determine the inter-distillation loss of the second feature extractor based on the second feature map and other first feature maps.
[0135] In the embodiment of the present application, step S26 is substantially the same as the above-mentioned step S25 and will not be described in detail here.
[0136] Step S27: Perform model training on the first target detection model according to the self-distillation loss, mutual distillation loss, and detector loss of the first target detection model.
[0137] In the embodiment of the present application, the terminal device obtains the self-distillation loss L of the first target detection model according to the above steps. A→A , inter-distillation loss L O→A and the detector loss L det Construct a loss function of the first target detection model to train the first target detection model. The loss function of the first target detection model is expressed as the following formula:
[0138] L A =L det +L O→A +L A→A
[0139] Step S28: Perform model training on the second target detection model according to the self-distillation loss, mutual distillation loss, and detector loss of the second target detection model.
[0140] In the embodiment of the present application, step S28 is substantially the same as the above-mentioned step S27 and will not be described in detail here.
[0141] Please continue reading Figure 5 , Figure 5 It is a flow chart of an embodiment of a target detection method provided by the present application.
[0142] like Figure 5 As shown, the specific steps of the target detection method in the embodiment of the present application are as follows:
[0143] Step S31: Acquire the image to be detected.
[0144] In the embodiment of the present application, the modality of the image to be detected is the modality trained in the training process of the above embodiment, that is, before the image to be detected is detected using the target detection model, the modality information of the image to be detected needs to be trained in advance.
[0145] Furthermore, the image to be detected may be one or more images, and the multiple images to be detected may also be images of different modalities, but the relevant modalities all need to be trained in the above-mentioned model training process.
[0146] Step S32: input the image to be detected into a pre-trained target detection model.
[0147] In this embodiment of the present application, the pre-trained target detection model can be Figures 1 to 4 The final target detection model is obtained by training with the model training method.
[0148] Step S33: Detect the target in the image to be detected based on the target frame output by the target detection model.
[0149] In an embodiment of the present application, the target detection module can output a target box that marks the target, or directly output information about the target, such as the size, position, height, etc. of the target.
[0150] Furthermore, after obtaining the detection results of each frame of the image to be detected through steps S31 to S33, the terminal device also needs to link the detection results between frames to achieve a complete pedestrian tracking effect. Figure 6 , Figure 6 yes Figure 5 A schematic diagram of a specific flow chart of step S33 in the target detection method is shown.
[0151] like Figure 6 As shown, the specific steps of step S33 in the embodiment of the present application are as follows:
[0152] Step S331: Identify the current target in the current image to be detected based on the target frame output by the target detection model.
[0153] Step S332: Identify other targets in the area where the current target is located in the previous frame of the image to be detected.
[0154] Step S333: Calculate the correlation between other targets and the current target.
[0155] Step S334: Determine whether the correlation is greater than or equal to a third preset threshold.
[0156] In an embodiment of the present application, since the detection time difference between frames is very small, the displacement of the same target in the adjacent frames of the image to be detected is generally also very small. Therefore, the terminal device can use the trained twin network to calculate the correlation, that is, calculate the correlation between the current target and the detection target in the area where the current target is located in the previous frame of the image to be detected. The terminal device obtains the maximum value of the correlation. If the maximum value exceeds the preset threshold, it is considered that the two targets are the same target in the adjacent frames of the image to be detected. The terminal device can assign the target ID of the target corresponding to the maximum correlation in the previous frame of the image to be detected to the current target, and enter step S335. If the maximum value is lower than the preset threshold, a new target ID is reallocated and the process proceeds to step S336.
[0157] Step S335: setting the target identifier of the current target according to the target identifiers of other targets.
[0158] Step S336: reallocate the target identifier.
[0159] The target detection method in the embodiment of the present application avoids the limitations of artificially formulated rules by constructing a twin network to calculate the correlation, and the measurement of correlation is more accurate, solving the ID allocation problem in multi-target tracking.
[0160] The above embodiment is only one common case of the present application and does not limit the technical scope of the present application. Therefore, any minor modifications, equivalent changes or modifications made to the above content based on the essence of the present application are still within the scope of the technical solution of the present application.
[0161] Please continue to see Figure 7 , Figure 7 Schematic diagram of the structure of an embodiment of a terminal device provided by the present application. Figure 7 As shown, the terminal device 40 includes an acquisition module 41, an input module 42, a distillation module 43 and a training module 44.
[0162] The acquisition module 41 is used to acquire training images of different modalities, wherein the training images of different modalities are obtained based on different imaging principles.
[0163] The input module 42 is used to input training images of different modalities into different target detection models to obtain output results of each target detection model, wherein the number of the target detection models is consistent with the number of modalities.
[0164] The distillation module 43 is used to obtain the mutual distillation loss of the target detection model based on the output result of one target detection model and the output results of other target detection models.
[0165] The training module 44 is used to perform model training according to the mutual distillation loss of the target detection model, and obtain a final target detection model based on multiple trained target detection models.
[0166] See also Figure 8 , Figure 8 FIG. 5 is a schematic diagram of the structure of another embodiment of the terminal device provided by the present application. The target detection device includes a memory 52 and a processor 51 which are connected to each other.
[0167] The memory 52 is used to store program instructions for implementing any one of the above-mentioned model training methods and / or target detection methods.
[0168] The processor 51 is used to execute program instructions stored in the memory 52 .
[0169] The processor 51 may also be referred to as a CPU (Central Processing Unit). The processor 51 may be an integrated circuit chip having the ability to process signals. The processor 51 may also be a general-purpose processor, a digital signaling processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0170] The memory 52 can be a memory stick, a TF card, etc., which can store all the information in the terminal device, including the input raw data, computer programs, intermediate operation results and final operation results are all stored in the memory. It stores and retrieves information according to the location specified by the controller. With the memory, the string matching prediction device has a memory function and can ensure normal operation. The memory of the string matching prediction device can be divided into main memory (internal memory) and auxiliary memory (external memory) according to the purpose. There is also a classification method of dividing it into external memory and internal memory. External memory is usually a magnetic medium or an optical disk, etc., which can store information for a long time. Memory refers to the storage component on the motherboard, which is used to store the data and programs currently being executed, but is only used to temporarily store programs and data. If the power is turned off or the power is cut off, the data will be lost.
[0171] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0172] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0173] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0174] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a system server, or a network device, etc.) or a processor to execute all or part of the steps of the various implementation methods of the present application.
[0175] See also Fig. 9 , which is a schematic diagram of the structure of the computer-readable storage medium of the present application. The storage medium of the present application stores a program file 61 that can implement all the above-mentioned model training methods and / or target detection methods, wherein the program file 61 can be stored in the above-mentioned storage medium in the form of a software product, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to execute all or part of the steps of each implementation method of the present application. The aforementioned storage device includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk, etc., which can store program codes, or a target detection device such as a computer, a server, a mobile phone, or a tablet.
[0176] The above are only implementation methods of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A model training method based on multimodal data, characterized in that: The model training method comprises: Acquire training images of different modalities, wherein the training images of different modalities are obtained based on different imaging principles; Inputting the training images of different modalities into different target detection models to obtain output results of each target detection model, wherein the number of the target detection models is consistent with the number of modalities; Based on an output result of an object detection model and output results of other object detection models, obtain the mutual distillation loss of the object detection model; Performing model training according to the mutual distillation loss of the target detection model; Based on multiple trained target detection models, the final target detection model is obtained.
2. The model training method according to claim 1, characterized in that: The model training method also includes: Based on the output result of an object detection model, obtain the self-distillation loss of the object detection model; The model training is performed according to the mutual distillation loss and self-distillation loss of the target detection model.
3. The model training method according to claim 2, characterized in that: The step of obtaining the self-distillation loss of a target detection model based on an output result of the target detection model includes: Obtaining a current output result of the target detection model for the training image based on current model parameters, and a historical output result of the target detection model for the same training image based on historical model parameters; Based on the current output result and the historical output results, a self-distillation loss of the target detection model is generated.
4. The model training method according to claim 3, characterized in that: The historical output result is the average value of the training output results of the target detection model before the current training time.
5. The model training method according to claim 2, It is characterized in that Wherein, the target detection model includes a feature extractor and a target detector; The model training method further includes: Inputting a training image of a first modality into a first object detection model, obtaining a first feature map extracted by a first feature extractor from the training image of the first modality, and a first object detection result of a first object detector based on the first feature map; Obtaining a self-distillation loss of the first feature extractor based on the first feature map, and obtaining a detector loss of the first object detector based on the first object detection result; Inputting the training image of the second modality into the second object detection model, obtaining a second feature map extracted by a second feature extractor from the training image of the second modality, and a second object detection result of a second object detector based on the second feature map; Obtaining a self-distillation loss of the second feature extractor based on the second feature map, and obtaining a detector loss of the second object detector based on the second object detection result; Determining a mutual distillation loss of the first feature extractor based on the first feature map and the other second feature map; determining a mutual distillation loss of the second feature extractor based on the second feature map and other first feature maps; Performing model training on the first target detection model according to the self-distillation loss, the mutual distillation loss, and the detector loss of the first target detection model; The second target detection model is trained according to the self-distillation loss, the mutual distillation loss, and the detector loss of the second target detection model.
6. The model training method according to claim 5, characterized in that: The first feature extractor and the second feature extractor have the same structure, and each includes a plurality of feature extraction blocks; The determining the mutual distillation loss of the first feature extractor based on the first feature map and the other second feature map includes: Obtaining a first feature subgraph output by each feature extraction block in the first feature extractor; Obtaining a second feature subgraph output by each feature extraction block in the second feature extractor; Determine the mutual distillation loss of the feature extraction block at the same position according to the first feature subgraph and the second feature subgraph output by the feature extraction block at the same position; Based on the mutual distillation losses of all feature extraction blocks in the first feature extractor, the mutual distillation loss of the first feature extractor is determined.
7. The model training method according to claim 1, It is characterized in that Wherein, the target detection model includes a feature extractor and a target detector; The trained target detection models are integrated to obtain the final target detection model, including: The outputs of feature extractors of several object detection models are stacked according to channels and input into a unified object detector to obtain the final object detection model. The training images of several modalities are input into a final object detection model to train the object detector in the final object detection model.
8. The model training method according to claim 1, characterized in that: After fusing the trained target detection models to obtain a final target detection model, the method further includes: Constructing a positive sample library, wherein the positive sample library includes a number of positive samples whose detection confidence is greater than a first preset threshold; Inputting the test image into the final object detection model to output a detection result of the test image, wherein the detection result includes a confidence level; Obtaining the maximum correlation between the test image and the positive samples in the positive sample library; Calculating a correlation confidence of the test image based on the confidence and the maximum correlation; Determining whether the relevant confidence of the test image is greater than or equal to a second preset threshold; If so, determining that the detection result of the test image is a positive detection; If not, it is determined that the detection result of the test image is a false detection.
9. A target detection method, characterized in that: The target detection method comprises: Acquire the image to be detected; Inputting the image to be detected into a pre-trained target detection model, wherein the target detection model is trained by the model training method according to any one of claims 1 to 7, and the modality of the image to be detected is any modality among the different modalities; The target in the image to be detected is detected based on the target frame output by the target detection model.
10. The target detection method according to claim 9, characterized in that: The detecting the target in the image to be detected based on the target frame output by the target detection model includes: Identify a current target in a current image to be detected based on a target frame output by the target detection model; Identify other targets in the area where the current target is located in the previous frame of the image to be detected; Calculating the correlation between the other targets and the current target; Determining whether the correlation is greater than or equal to a third preset threshold; If yes, setting the target identifier of the current target according to the target identifiers of the other targets; If not, reallocate the target ID.
11. A model training device, characterized in that: The model training device includes an acquisition module, an input module, a distillation module and a training module; wherein, The acquisition module is used to acquire training images of several modalities; The input module is used to input training images of different modalities into different target detection models to obtain output results of each target detection model, wherein the number of the target detection models is consistent with the number of modalities; The distillation module is used to obtain the mutual distillation loss of the target detection model based on the output result of one target detection model and the output results of other target detection models; The training module is used to perform model training according to the mutual distillation loss of the target detection model; and obtain a final target detection model based on multiple trained target detection models.
12. A terminal device, characterized in that: The terminal device includes a processor and a memory connected to the processor, wherein: The memory stores program instructions; The processor is used to execute the program instructions stored in the memory to implement the model training method described in any one of claims 1 to 8 and / or the target detection method described in any one of claims 9 to 10.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are executed, the model training method described in any one of claims 1 to 8 and / or the target detection method described in any one of claims 9 to 10 are implemented.
Citation Information
Patent Citations
Multi-modal visual target tracking method based on self-distillation symmetric adapter
CN117710414A
Noise Tolerant Ensemble RCNN for Semi-Supervised Object Detection
US20220172456A1
Training method and apparatus for image processing model, electronic device, computer program product, and computer storage medium
US20240412374A1