Target detection method, system and device, electronic equipment, medium and program product
By jointly processing data by multiple models and determining the annotation results with the processing results of multiple models, the problems of low efficiency and low accuracy of existing object detection algorithms are solved, and more efficient and accurate object detection algorithm development is achieved.
Patent Information
- Application Number
- CN202510214230.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
The development efficiency of existing object detection algorithms is low, with low accuracy, and rely on manual operations of professional developers.
The data is processed through multiple models, and the data annotation results are determined based on the processing results of multiple models, improving the efficiency and accuracy of data annotation, thereby improving the development efficiency and accuracy of object detection algorithms.
It improves the efficiency and accuracy of data annotation and improves the development efficiency and accuracy of object detection algorithms.
Smart Images

Figure CN120219802A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to an object detection method, system, device, electronic device, medium, and program product. Background Art
[0002] Object detection algorithms are generally used to locate and classify objects of interest in the collected images or video streams, and can be specifically applied to industrial defect detection, behavior detection, etc. In the entire application process of commonly used object detection algorithms, it relies on the manual operations of professional developers. For example, during the model training process of object detection algorithms, developers need to manually annotate the data for model training, etc. This process that relies on the manual operations of professional developers results in low development efficiency and low accuracy of object detection algorithms. Summary of the Invention
[0003] This application provides an object detection method, system, device, electronic device, medium, and program product. By processing data through multiple models, multiple processing results of the data are obtained, and the annotation result of the data is determined through the multiple processing results of the data, which improves the efficiency and accuracy of annotating the data, further improves the development efficiency of the object detection algorithm, and improves the accuracy of object detection.
[0004] This application provides an object detection method, and the method includes:
[0005] Obtain an image to be processed;
[0006] Process the image to be processed through a first model to obtain an object detection result of the image to be processed; the first model is trained based on a first training data set; the annotation result of the first data in the first training data set is obtained based on the processing results of the first data by multiple second models; the difference between the annotation result of the first data and the standard result corresponding to the first data is less than a first threshold, and the first threshold represents the difference between the processing result of the first data by a second model and the standard result corresponding to the first data; the standard result represents the annotation result determined in advance through manual annotation; the first data is any sample data in the first training data set.
[0007] Based on the above object detection method, before obtaining the image to be processed, the method further includes: obtaining a sample data set, processing the first data in the sample data set through a plurality of second models to obtain the processing results of each second model on the first data; the first data is any data in the sample data set; determining the first weight of each second model, and based on the first weight of each second model and the processing results of each second model on the first data, determining the annotation result of the first data; constructing a first training data set based on the first data and the annotation result of the first data.
[0008] It can be seen that determining the annotation result of the first data based on the weight of each second model and the processing result of each second model is beneficial to reducing the error of data annotation by a single second model, improving the accuracy of the annotation result, and providing high-quality training data for the first model.
[0009] Based on the above object detection method, determining the first weight of each second model includes: determining the confidence of each second model based on the processing results of each second model on a plurality of data; determining the first weight of each second model through the confidence of each second model.
[0010] It can be seen that determining the weight of the second model through the confidence of the second model can reflect the reliability of the processing result of each second model. Based on the weight of each second model, it is beneficial to obtain an accurate annotation result of the first data.
[0011] Based on the above object detection method, before obtaining the image to be processed, the method further includes: training a plurality of first models through the first training data set; the step of obtaining the object detection result of the image to be processed by the first model includes: processing the image to be processed through the plurality of first models to obtain the processing results of each first model; determining the second weight of each first model, and based on the second weight of each first model and the processing results of each first model, obtaining the object detection result of the image to be processed.
[0012] It can be seen that obtaining the object detection result of the image to be processed by combining a plurality of first models is beneficial to reducing the processing error of a single first model on the image to be processed and improving the accuracy of the object detection result.
[0013] Based on the above object detection method, determining the second weight of each first model includes: determining the confidence of each first model through the object detection results of each first model on a plurality of images; determining the second weight of each first model through the confidence of each first model.
[0014] It can be seen that determining the weights of the first model based on the confidence of the first model can reflect the reliability of each first model in obtaining the object detection result through image processing, which is beneficial to improving the accuracy of the object detection result.
[0015] Based on the above object detection method, before processing the image to be processed by the first model, the method further includes: in the first model, determining the third weight and activation value of any layer of the first model; obtaining a plurality of weight quantization factors and a plurality of activation value quantization factors; the plurality of weight quantization factors are a plurality of quantization factors corresponding to the preset third weights, and the plurality of activation value quantization factors are a plurality of quantization factors corresponding to the preset activation values; based on the plurality of weight quantization factors and the plurality of activation value quantization factors, determining the target weight quantization factor and the target activation value quantization factor of any layer; the target weight quantization factor and the target activation value quantization factor can make the dequantization processing result of the quantized first model have a similarity greater than the similarity threshold with the output result of the first model before quantization; based on the target weight quantization factor of each layer in the first model and the target activation value quantization factor of each layer, performing quantization processing on the first model to obtain the target quantized first model; the processing of the image to be processed by the first model includes: processing the image to be processed by the target quantized first model.
[0016] It can be seen that through the model quantization method provided in this application, it is beneficial to make the dequantization processing result of the quantized first model closer to the processing result of the first model before quantization, reduce the impact of model quantization on the model accuracy, and maximize the retention of the information and accuracy of the original model.
[0017] Based on the above object detection method, the determining the target weight quantization factor and the target activation value quantization factor of any layer based on the plurality of weight quantization factors and the plurality of activation value quantization factors includes: obtaining a plurality of quantized first models based on the plurality of weight quantization factors and the plurality of activation value quantization factors; processing the first data through each quantized first model to obtain a plurality of first processing results; performing dequantization processing on each first processing result to obtain a plurality of dequantization processing results; obtaining the second processing result of the first model before quantization on the first data; among the plurality of dequantization processing results, determining the weight quantization factor corresponding to the dequantization processing result most similar to the second processing result as the target weight quantization factor; determining the activation value quantization factor corresponding to the dequantization processing result most similar to the second processing result as the target activation value quantization factor.
[0018] It can be seen that by using the above method to determine the quantization factors of the first model, it is beneficial to obtain the optimal quantization factors, improve the processing accuracy of the first model after quantization, and achieve the optimal model quantization effect.
[0019] Based on the above object detection method, obtaining multiple quantized first models based on the multiple weight quantization factors and the multiple activation value quantization factors includes: determining a first weight quantization factor among the multiple weight quantization factors, and obtaining quantization factor pairs formed by the first weight quantization factor and each activation value quantization factor among the multiple activation value quantization factors; obtaining a first set of quantization factor pairs corresponding to the multiple weight quantization factors based on the quantization factor pairs corresponding to each first weight quantization factor; performing quantization processing on the first model through the first set of quantization factor pairs to obtain multiple quantized first models; the first weight quantization factor is any one of the multiple weight quantization factors; and / or, determining a first activation value quantization factor among the multiple activation value quantization factors, and obtaining quantization factor pairs formed by the first activation value quantization factor and each weight quantization factor among the multiple weight quantization factors; obtaining a second set of quantization factor pairs corresponding to the multiple activation value quantization factors based on the quantization factor pairs corresponding to each first activation value quantization factor; performing quantization processing on the first model through the second set of quantization factor pairs to obtain multiple quantized first models; the first activation value quantization factor is any one of the multiple activation value quantization factors.
[0020] It can be seen that by traversing the multiple weight quantization factors and the multiple activation value quantization factors to obtain multiple quantized first models, and by finely searching layer by layer for the weight quantization factors and activation value quantization factors of each layer in the first model, it is possible to effectively reduce the error accumulation introduced by rough quantization, enabling the first model to better retain the parameter information and accuracy of each layer of the first model before quantization after quantization, having stronger adaptability in terms of accuracy, and being applicable to application scenarios with high precision requirements.
[0021] Based on the above object detection method, the parameter order of magnitude of the second model is greater than that of the first model.
[0022] It can be seen that using a model with a large parameter order of magnitude for data annotation helps to improve the accuracy of data annotation, and training the first model with accurate data annotation is beneficial to improving the processing ability of the first model. Using the first model with a small parameter order of magnitude to process the image to be processed is beneficial to improving the model processing efficiency.
[0023] Based on the above object detection method, the first model is a student model and the second model is a teacher model.
[0024] It can be seen that since the teacher model is usually a model with a large number of parameters, generating the annotation results of the training data of the student model through the teacher model can improve the accuracy of data annotation and the training effect of the student model. Training the student model with annotation results of higher accuracy and guiding the training of the student model through the teacher model is beneficial to further improving the training effect of the student model and enhancing the processing ability of the student model.
[0025] This application also provides an object detection system, which includes:
[0026] An interaction module, configured to receive data to be processed, where the data to be processed includes one or more of a plurality of first data, an image to be processed, and hyperparameters of the first model; the plurality of first data are used to construct a first training dataset of the first model; the first model is used to process the image to be processed to obtain an object detection result of the image to be processed;
[0027] A data processing module, configured to process each of the plurality of first data based on a plurality of second models to obtain an annotation result of each first data; construct a first training dataset through the plurality of first data and the annotation result of each first data in the plurality of first data; wherein, the difference between the annotation result of the first data and the standard result corresponding to the first data is less than a first threshold, and the first threshold represents the difference between the processing result of a second model on the first data and the standard result corresponding to the first data; the standard result represents an annotation result determined in advance through manual annotation;
[0028] A model training module, configured to train the first model based on the first training dataset and hyperparameters.
[0029] This application also provides an object detection device, which includes:
[0030] An acquisition module, configured to acquire an image to be processed;
[0031] A processing module, configured to process the image to be processed through a first model to obtain an object detection result of the image to be processed; the first model is trained based on a first training data set; the annotation result of the first data in the first training data set is obtained based on the processing results of a plurality of second models on the first data; the difference between the annotation result of the first data and the standard result corresponding to the first data is less than a first threshold, and the first threshold represents the difference between the processing result of a second model on the first data and the standard result corresponding to the first data; the standard result represents an annotation result determined in advance through manual annotation; the first data is any sample data in the first training data set.
[0032] This application provides an electronic device, which includes a processor and a memory for storing a computer program that can run on the processor; wherein,
[0033] The processor is configured to run the computer program to execute any one of the above object detection methods.
[0034] This application provides a computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above object detection methods.
[0035] This application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements any one of the above object detection methods.
[0036] This application provides an object detection method, system, device, electronic device, medium and program product. Based on the object detection method provided by this application, by combining the processing results of a plurality of second models, the annotation result of the training data of the first model is obtained, which helps to improve the annotation efficiency and accuracy of the training data. Further, by training the first model with the annotation result jointly obtained by a plurality of second models, it helps to improve the training effect of the first model, improve the processing ability of the first model, and improve the development efficiency of the object detection algorithm. Description of the Drawings
[0037] Figure 1 It is a flowchart of an object detection method provided by this application;
[0038] Figure 2 It is a flowchart of a data annotation method provided by this application;
[0039] Figure 3 It is a schematic structural diagram of an object detection system provided by this application;
[0040] Figure 4 It is a schematic structural diagram of another object detection system provided by this application;
[0041] Figure 5 A flowchart for the development of an object detection system provided for this application;
[0042] Figure 6 A schematic structural diagram of an object detection device provided for this application;
[0043] Figure 7 A schematic diagram of the composition structure of an electronic device provided for this application. Detailed implementation manners
[0044] Object detection algorithms are generally used to locate and classify objects of interest in the collected images or video streams, and can be specifically applied to downstream application tasks of object detection in computer vision fields such as industrial defect detection and behavior detection. Traditional object detection algorithms mainly rely on manually designed features, such as Sobel operator edge detection features, Haar-like features, Histogram of Oriented Gradients (HOG), etc. These features have weak generalization ability and poor performance in complex scenarios. With the rapid development of deep learning technology, object detection algorithms based on deep learning use convolutional neural networks to learn features. By applying convolutional neural networks, input information can be converted into high-dimensional image features. These obtained high-dimensional image features have powerful feature expression ability and generalization performance, and perform better in complex scenarios, gradually replacing traditional machine vision detection algorithms. In recent years, traditional machine vision algorithms or deep learning algorithms have been gradually used in object detection tasks in the industrial field, such as defect detection of industrial products or components, personnel safety detection and recognition, etc.
[0045] Currently, the general process of applying object detection algorithms in the industrial field is as follows: collect data for a specific project, annotate and clean it, build a neural network framework, and design data preprocessing and algorithm postprocessing modules according to task requirements. Use the collected data for model training, fine-tuning, and verification of the effect, optimize the algorithm according to the verification results, adjust the parameters of the preprocessing and postprocessing modules, and finally package it as an image or encapsulate it as a Software Development Kit (SDK) according to the target environment. The above processing process has many deficiencies in industrial applications. For example, in the field of industrial defect detection, when detecting industrial products produced on a production line, the detection performance of the algorithm often fluctuates due to factors such as production processes, product raw materials, and product specifications. Therefore, in order to maintain the accuracy of the algorithm, the algorithm often needs to be optimized according to the changes of objective factors during the actual production process. However, manual optimization will consume a large amount of manpower and material resources, resulting in low production efficiency.
[0046] In some industrial behavior recognition and detection scenarios, faster processing efficiency is often required, while the accuracy requirement for detection results is relatively low. However, the commonly used dynamic quantization and static quantization methods often cannot bring sufficient precision, and the quantization-aware training method often requires training and optimization, which takes a long time. In some edge deployment scenarios, when the computational complexity of the object detection algorithm is high, it will lead to high computing power requirements and long model inference time. Therefore, model lightweighting is required. However, the current model lightweighting operations often rely on quantization, distillation, and pruning, resulting in a serious decline in model performance.
[0047] In summary, the entire process of current object detection methods often relies on manual operations by professional developers, such as model selection, hyperparameter adjustment, model deployment, etc. The implementation cycle is long, the development efficiency is low, it cannot be automatically configured, the development threshold is high, and the model update and iteration cycle is long.
[0048] To solve these problems, the embodiments of this application propose an object detection method that can improve the efficiency and accuracy of automatically annotating data through multiple models, greatly reducing the workload of manual annotation correction, improving the accuracy of annotation results, and further improving the accuracy of the object detection model.
[0049] The following further details the embodiments of this application in conjunction with the accompanying drawings and embodiments. It should be understood that the embodiments provided here are only used to explain the embodiments of this application and are not used to limit the embodiments of this application. Additionally, the embodiments provided below are partial embodiments for implementing this application, rather than all embodiments for implementing this application. Without conflict, the technical solutions described in the embodiments of this application can be implemented in any combination.
[0050] It should be noted that in the embodiments of this application, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a method or device including a series of elements not only includes the elements explicitly recited, but also includes other elements not explicitly listed, or further includes elements inherent to the implementation of the method or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional related elements in the method or device including such element (such as steps in a method or units in a device, for example, a unit in a device can be a part of a circuit, a part of a processor, a part of a program, or software, etc.).
[0051] The object detection method provided by the embodiments of this application includes a series of steps. However, the object detection method provided by the embodiments of this application is not limited to the recorded steps. Similarly, the object detection system and device provided by the embodiments of this application include a series of modules. However, the system and device provided by the embodiments of this application are not limited to including the explicitly recorded modules, and may also include modules required for obtaining relevant information or processing based on information.
[0052] The embodiments of this application provide an object detection method, as Figure 1 shown, Figure 1 which shows a flowchart of an object detection method. Figure 1 The object detection method shown includes:
[0053] Step 101: Obtain an image to be processed.
[0054] Step 102: Process the image to be processed through a first model to obtain an object detection result of the image to be processed.
[0055] The first model is trained based on a first training dataset; the annotation result of the first data in the first training dataset is obtained based on the processing results of multiple second models on the first data; the difference between the annotation result of the first data and the standard result corresponding to the first data is less than a first threshold, and the first threshold represents the difference between the processing result of a second model on the first data and the standard result corresponding to the first data; the standard result represents an annotation result determined in advance through manual annotation; the first data is any sample data in the first training dataset.
[0056] Among them, the difference between the annotation result of each data in the first training dataset and the standard result corresponding to the data is less than the first threshold corresponding to the data.
[0057] In this embodiment, the first model is used to perform object detection on the image to be processed, and accurately identify the position, size, category, etc. of the target object in the image to be processed. The first model may specifically be a model obtained based on mainstream object detection algorithms, and the object detection algorithms may be algorithms based on deep learning, such as algorithms in the Regions with Convolutional Neural Networks (RCNN) series, YOLO (You Only Look Once) series algorithms, Single Shot MultiBox Detector (SSD) algorithms, etc.
[0058] The first model is trained based on the sample data in the first training dataset and the annotation results corresponding to the sample data. Among them, the annotation result corresponding to any sample data in the first training dataset is obtained from the processing results of multiple second models on the sample data; the annotation results of all sample data in the first training dataset are based on the processing of multiple second models; alternatively, the annotation results of a part of the sample data in the first training dataset are obtained through manual annotation, and the annotation results of another part of the sample data are based on the processing of multiple second models.
[0059] For any sample data in the first training dataset, such as the first data, the second model is used to process the first data to obtain the processing result of the first data. Here, the processing result of the first data represents the annotation result of the first data by a second model. In the target detection process, the sample data in the first training dataset can be the training images for training. Correspondingly, the processing result obtained by the second model is the annotation result of the training image. The annotation result of the training image includes one or more of class annotation, location annotation, and attribute annotation.
[0060] To improve the annotation ability of the second model for different types of images, each second model among multiple second models can be trained based on the second training dataset. The sample data in the second training dataset can include a large amount of basic data, covering as many data types, appearance features, and image data under different environmental conditions as possible, such as illumination, angle, background complexity, etc.
[0061] In the specific implementation process, for example, when it is necessary to determine the types of photovoltaic module defects, a batch of photovoltaic module images with a large enough scale can be collected, so that the number of training images for each type of defect in the photovoltaic module is not less than 1000, and professional personnel are invited to perform data annotation on the training images to ensure the accuracy of the training image annotation. Through the collected photovoltaic module images and the annotation results of manual annotation, the second training dataset is constructed. The sample data and annotation results in the second training dataset can be randomly divided into a training set, a validation set, and a test set, and each second model is trained and validated so that the trained second model can perform target detection on the photovoltaic module images, obtain the processing results, and annotate the defect information in the photovoltaic module images.
[0062] The multiple second models in this embodiment include two or more second models, where each second model among the multiple second models is a different model. To ensure the accuracy of the data annotation results for each second model, each second model should select a model with a relatively large parameter scale or a model that can process input data of different scales. For example, based on the current advanced YOLOv8, Fast Regions with Convolutional Neural Networks (Fast-RCNN), Collaborative Detection Transformer (Co-DETR), etc., the network with the largest model parameters is selected from each series of algorithms for training to obtain the second models corresponding to different algorithms.
[0063] After obtaining different second models through different algorithms, the first data is processed based on the multiple second models to obtain the annotation result of the first data. Since the processing results of the same data by different algorithms are different, therefore, among the processing results of the first data obtained by each second model, a result with the smallest difference value from the standard result of the first data can be determined as the final annotation result of the first data, that is, as the annotation result in the first training dataset. Or, based on the processing capabilities of each second model for different types of data, the weights corresponding to different types of data for each second model are determined. When processing the first data, first determine the type corresponding to the first data, and process the first data through each second model to obtain multiple processing results of the first data. Determine the weights corresponding to each second model according to the type of the first data, and perform weighted summation through the weights corresponding to each second model and the processing results of each second model to obtain the annotation result of the first data. Here, the standard result corresponding to the first data can be the annotation result corresponding to the first data determined manually.
[0064] It can be seen that the annotation result of the first data obtained from the processing results of the first data by multiple second models is at least better than the processing results of one second model because the relatively better processing results are selected in this process. Through the method for determining the annotation result of the first data given in this embodiment, the processing error of a single second model for the first data can be effectively reduced, and the accuracy of automatically annotating the first data in the first training dataset can be improved.
[0065] The embodiments of the present application provide an object detection method. By combining the processing results of multiple second models, the annotation result of the sample data of the first model is obtained, which can improve the accuracy of automatically obtaining the annotation result of the sample data. Further, by training the first model with training data of relatively high accuracy, the object detection ability of the first model can be improved.
[0066] In practical applications, steps 101 to 102 can be implemented based on a processor, and the processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor.
[0067] In some embodiments, the parameter order of magnitude of the second model is greater than that of the first model.
[0068] In the embodiments of the present application, the second model is mainly used to process sample data to obtain a processing result, and an annotation result of the sample data is generated through the processing result. It can be seen that by using a model with a relatively large parameter order of magnitude as the second model, an accurate annotation result can be obtained through the second model, improving the data annotation accuracy. Training multiple second models with relatively large parameter orders of magnitude can effectively improve the generalization ability of the second model in practical applications. Without considering the model processing efficiency and only selecting the model with the best detection performance, such as the model with the highest mean average precision (MAP), as the second model for data processing can ensure the advantages of the second model in feature extraction and object detection accuracy.
[0069] Obtaining the annotation result of the sample data from the processing results of multiple second models can integrate the processing results of multiple detection algorithms, ensure a more accurate annotation result of the sample data, reduce the error of a single model on some difficult classification targets, improve the accuracy of the annotation result, and thus provide high-quality training data for the first model. Using a model with a relatively small number of parameters as the first model for image processing to obtain the object detection result can effectively improve the speed of object detection and is applicable to scenarios with real-time requirements. The object detection method provided in the embodiments of this application effectively balances the detection accuracy and processing efficiency and provides the best solution for the real-time requirements in end-side applications. Compared with the manual annotation method, the efficiency and quality of data annotation are improved.
[0070] In some embodiments, the first model is a student model, and the second model is a teacher model.
[0071] In deep learning and knowledge distillation techniques, a student model is a smaller and simpler model designed to learn and imitate the behavior or prediction results of another more complex and better-performing model (i.e., the teacher model). The teacher model is a well-trained and well-performing model, usually a larger and more complex model, which can achieve a high performance level on a specific task. In the embodiments of this application, in combination with the knowledge distillation technique, while the student model learns by imitating the output or feature representation of the teacher model and gradually optimizes its parameters and structure during training, the annotation result of the sample data for training the student model is automatically obtained through the teacher model, further improving the training efficiency of the student model.
[0072] At the same time, on the basis of the knowledge distillation technique, in this embodiment, by applying multiple teacher models to participate in the training of the student model, the accuracy of training the student model is further improved, which is beneficial to improving the accuracy of the student model in generating object detection results for the image to be processed.
[0073] Based on the data annotation method given in the above embodiments, Figure 2 A flowchart of a data annotation method is shown. The teacher model 203 is trained through the data in the basic data set 201 and the defect annotation 202 corresponding to the data in the basic data set 201. Here, the defect annotation 202 can be obtained by manually annotating the data in the basic data set 201. The new data set 204 is obtained, and the teacher model 203 processes the data in the new data set 204 to obtain the annotation result 205 of each data in the new data set 204. The student model 206 is trained through the new data set 204 and the annotation result 205 of each data in the new data set 204.
[0074] In some embodiments, before obtaining the image to be processed, the method further includes: obtaining a sample data set, processing the first data in the sample data set through a plurality of second models to obtain the processing results of each second model on the first data; the first data is any data in the sample data set; determining the first weight of each second model, and based on the first weight of each second model and the processing results of each second model on the first data, determining the annotation result of the first data; and constructing a first training data set based on the first data and the annotation result of the first data.
[0075] Corresponding to Figure 2 the data annotation method shown, the sample data set may be Figure 2 the newly added data set 204 shown. The first data in the sample data set is processed by different second models respectively to obtain the processing results of different second models on the first data. Here, the processing result is the annotation result of the first data obtained by each second model on the first data before finally determining the annotation result of the first data.
[0076] In this embodiment, after obtaining the processing results of each second model on the first data, based on the first weight of each second model, the processing results corresponding to each second model are weighted and summed to obtain the annotation result of the first data, and further, the annotation results of each data in the sample data set can be obtained. Based on each first data in the sample data set and the annotation result corresponding to each first data, a first training data set is constructed, and the first model is trained through the first training data set.
[0077] In this embodiment, the first weight of each second model can be determined based on the construction algorithm of each second model. For example, by analyzing the algorithms of each second model, the weaknesses of each algorithm in data processing are determined, that is, the data types that each algorithm is not good at processing are determined, and based on the data types that each algorithm is not good at processing, the first weight of the second model constructed by each algorithm is determined. Among them, the higher the first weight of the data type that the algorithm of the second model is better at processing. It can be seen that in this embodiment, the first weight corresponding to the second model can be determined through the type of the first data, and further, the annotation result of the first data can be determined.
[0078] For example, the R-CNN series of models may encounter difficulties in processing small or dense objects, and perform better in detecting objects with clear boundaries and significant features, such as faces, vehicles, etc.; the YOLO model performs well in real-time object detection tasks, such as in the fields of autonomous driving and video surveillance, but it is difficult to accurately capture the features of small objects; the SSD model has good performance in processing small objects. Therefore, when processing small object information, the processing results of the SSD model can be given a higher first weight, while the processing results of the YOLO model and the R-CNN series of models can be given a lower first weight; when processing face-related object recognition, a higher weight can be given to the processing results of the R-CNN series of models. It can be seen that when processing different types of data, different first weights can be set for each second model, so that the first weight of the second model changes dynamically according to the type of data that the second model is good at processing.
[0079] It can be seen that by combining the processing results of multiple second models and the first weights of each second model, it is beneficial to integrate the performance gains brought by the algorithms of multiple different models, reduce the processing errors of a single model for certain types of data, and improve the accuracy of the annotation results.
[0080] In some embodiments, determining the first weight of each second model includes: determining the confidence level of each second model based on the processing results of each second model for multiple data; determining the first weight of each second model through the confidence level of each second model.
[0081] This embodiment provides a method for determining the first weight of each second model. Based on the processing results of each second model for multiple data, the similarity of the processing results of every two second models for the same data can be determined through the processing results of multiple data obtained by each second model, and the confidence level of each second model can be determined according to the similarity. When the processing result of a second model for data is closer to the processing results of other multiple second models, it is considered that the confidence level of this second model is higher, and a higher weight can be set for this second model.
[0082] The data used to determine the confidence level of the second model in this embodiment can be any one data, or it can be the data in the second training dataset.
[0083] Since the processing results of different second models for different data are different, it is possible to process multiple types of data through each second model, determine the data types for which the processing results of each second model are more accurate, determine the confidence of each second model, and further determine the first weight of each second model. Among them, the confidence of each second model can be directly used as its corresponding first weight, or the first weight of each second model can be positively correlated with the corresponding confidence. Alternatively, the first weight of the second model can be comprehensively obtained based on the accuracy of the processing results of each second model for different types of data.
[0084] In practical applications, the confidence of each second model can also be determined based on the Intersection over Union (IoU). IoU is an index for measuring the overlapping degree between the detection box and the ground truth box. Here, the ground truth box refers to the box manually marked. The higher the value of IoU, the higher the overlapping degree between the detection box and the ground truth box, that is, the more accurate the position of the detection box. When any second model processes any data (image) and obtains the processing result, taking the image as the reference, the processing results of multiple second models corresponding to an image can be read, the processing results of multiple second models are stored in a json file, the detection box information in all processing results is stored in a python list L, and the detection boxes in the list L are sorted using the sorting function of Python (such as the sorted function).
[0085] The detection boxes of each image are clustered using a clustering method and stored in another python list F separately according to the approximate positions. Among them, a list F includes the detection box information of an image, and the data in the list F is a subset of the data in the list L. Iterate over all the processing results in the list L, that is, check each detection box in the list L one by one, match each detection box in the list L with the clustered detection boxes in the list F, evaluate the overlapping degree between them through IoU, and specifically, the IoU evaluation can be performed based on formula (1). The detection boxes with the intersection over union results obtained from formula (1) greater than the specified preset threshold (Threshold, THR) are merged.
[0086]
[0087] Among them, F i represents the i-th data in the list F, and L j represents the j-th data in the list L.
[0088] Calculate the average value of the confidence of all detection boxes in the list F through formula (2), and calculate the confidence of the merged detection boxes in the list F:
[0089]
[0090] Among them, C i represents the confidence of the i-th merged detection box, and the finally obtained merged confidence C in formula (2) can be used as the first weight of a second model.
[0091] Finally, according to the first weight, the coordinates of the new detection box are calculated by weighted summation as shown in formulas (3) and (4):
[0092]
[0093] Among them, "*" represents performing a convolution operation, X1, 2 represent the two abscissas of the detection box, and Y1, 2 represent the two ordinates of the detection box, represents the two abscissas of the detection box obtained by the n-th second model, represents the two ordinates of the detection box obtained by the n-th second model; C n represents the first weight of the n-th second model.
[0094] In the embodiment of the present application, the second model is mainly used to automatically annotate data to obtain the annotation result of the data, and it is necessary to ensure the accuracy of the annotation result. When the time duration of the data annotation cycle allows, the processing time of the data, that is, the time to obtain the annotation result of the data, does not need to be considered too much. After obtaining the annotation result of the data based on multiple second models, the annotation result automatically obtained by the second model can also be reviewed manually, and the mislabeled, missed-labeled, and mislabeled data can be corrected to further obtain the training data required for parameter updating of each second model subsequently.
[0095] In practical applications, based on the existing labeled data set (the second training data set), the data in each data set can be randomly divided into a training set, a validation set, and a test set, and the parameters of each second model can be updated. If there is a new second model, transfer learning can be selected to fine-tune on the basis of the existing second model to reduce the training time of the new model.
[0096] In this embodiment, by ensuring that the outputs of different second models have sufficient consistency in detection boxes, class labels, and confidence scores, the difference in the output results is avoided from being too large, which affects the quality of the final annotation result. By dynamically allocating the weights of the second models, it can effectively prevent the weight of a single model from being too high, resulting in the neglect of the processing results of other models.
[0097] In some embodiments, before obtaining the image to be processed, the method further includes: training a plurality of first models with a first training dataset; processing the image to be processed through the first models to obtain the target detection result of the image to be processed, including: processing the image to be processed through a plurality of first models to obtain the processing result of each first model; determining the second weight of each first model, and obtaining the target detection result of the image to be processed based on the second weight of each first model and the processing result of each first model.
[0098] After obtaining the first training dataset based on the processing of a plurality of second models, the plurality of first models can also be trained based on the first training dataset. Here, the plurality of first models are different models obtained based on different algorithms. For example, the first model can be a first model obtained based on algorithms such as PP-PicoDet, NanoDet, YOLOv3-Tiny, etc.
[0099] After processing the image to be processed through each first model to obtain the processing result of each first model, perform weighted summation on the processing result of each first model based on the second weight of each first model to obtain the target detection result.
[0100] In this embodiment, the second weight of each first model can be determined based on the construction algorithm of each first model. For example, analyze the algorithm of each first model to determine the weakness of each algorithm in data processing, that is, determine the type that each algorithm is not good at processing, and determine the second weight of the first model constructed by each algorithm based on the type that each algorithm is not good at processing. Among them, the higher the second weight of the data type that the algorithm corresponding to the first model is better at processing. It can be seen that in this embodiment, the second weight corresponding to the first model can be determined through the type of the image to be processed, and further the target detection result can be determined.
[0101] For example, the first model obtained based on the PP-PicoDet algorithm usually has a fast processing speed and is very suitable for application scenarios that require real-time detection. However, in the case of extremely small target sizes, the detection accuracy may decrease; the first model obtained based on the NanoDet algorithm is usually suitable for application scenarios that require high-speed detection. However, in tasks with extremely high requirements for detection accuracy, a more complex model may be required to achieve the required accuracy; the first model obtained based on the YOLOv3-Tiny algorithm performs well in various object detection tasks, including the detection of common objects such as pedestrians and vehicles. However, in complex backgrounds, the accurate recognition and positioning of targets by this algorithm may be interfered with, resulting in a decrease in detection accuracy. Therefore, different second weights can be set corresponding to the types of images that different first models are good at processing. It can be seen that when processing different types of images, different second weights can be set for each first model, so that the second weight of the first model changes dynamically according to the type of image that the first model is good at processing. Or, based on the accuracy of the processing results of each first model for different types of data, the second weight of the first model can be comprehensively obtained.
[0102] It can be seen that by combining the processing results of multiple first models and the second weights of each first model, it is beneficial to integrate the performance gains brought by algorithms with multiple different network structures, reduce the processing errors of a single model for certain types of data, and improve the accuracy of object detection results.
[0103] In some embodiments, determining the second weight of each first model includes: determining the confidence of each first model through the object detection results of each first model for multiple images; determining the second weight of each first model through the confidence of each first model.
[0104] This embodiment provides a method for determining the second weight of each first model. Based on the object detection results of each first model for multiple images, the similarity of the processing results of every two first models for the same image can be determined through the object detection results of multiple images obtained by each first model. The confidence of each first model can be determined according to the similarity. When the object detection result of a first model for an image is closer to the object detection results of other multiple first models, it is considered that the confidence of this first model is higher, and a higher second weight can be set for this first model. In this embodiment, the image used to determine the confidence of each first model can be any image, or it can be the data in the first training dataset.
[0105] Since the processing results of different first models for different data are different, based on the method given in the above embodiments and the calculation methods given in Formulas (1) to (4), the confidence of each first model can be determined by the IoU method based on the object detection results of each first model.
[0106] After obtaining the first training dataset, the first training dataset can be randomly divided into a training set, a validation set, and a test set. Select an already integrated object detection algorithm, configure hyperparameters for each first model, such as the learning rate, the number of layers, the regularization parameter, etc. If there is a suitable pre-trained model, transfer learning can also be selected to fine-tune on the basis of the existing model to reduce the training time of the first model. After the hyperparameter configuration is completed, the model can be trained using the training set obtained by dividing the first training dataset. The training of the first model can be implemented based on various training modes such as a single machine, multiple machines, multiple Graphics Processing Units (GPUs), Neural Processing Units (NPUs), etc., or the training mode of the model can be reasonably selected according to the currently idle resources or queued for resources. After starting the model training, training messages can be obtained and sent to the background through a specified port. After the background algorithm receives the messages, complete training statements are formed to start training the model. If during the training process of the first model, the performance of processing the data in the validation set no longer improves, the model training can be automatically terminated to avoid overfitting of the first model.
[0107] After the training of the first model is completed, each first model can also be evaluated, and the performance of each first model on the test set can be analyzed in detail through a series of evaluation metrics, such as accuracy, recall rate, F1 score, statistical metric (Area Under the Curve, AUC), etc. According to the evaluation results, further adjust the model structure of each first model, reconfigure the hyperparameters, or adjust the data in the first training dataset to further improve the performance of each first model.
[0108] This embodiment provides a method for target detection by jointly using multiple first models to determine the target detection result. It can integrate the performance gains brought by multiple algorithms with different network structures, reduce the error of a single model in target detection, and improve the accuracy of the target detection result. By ensuring sufficient consistency in the detection boxes, class labels, and confidence scores of the outputs of different first models, it is possible to avoid excessive differences in the output results and affect the quality of the final target detection result. By dynamically allocating the weights of the first models, it can effectively prevent a single model from having too high a weight and ignoring the results of other models. In the specific implementation process, the model inference can also be accelerated through the high-performance deep learning inference software development kit NVIDIA TensorRT, which supports NVIDIA GPUs and Jetson series embedded devices.
[0109] In some embodiments, before the above-mentioned first model processes the image to be processed, the method further includes: in the first model, determining the third weight and activation value of any layer of the first model; obtaining a plurality of weight quantization factors and a plurality of activation value quantization factors; the plurality of weight quantization factors are the quantization factors corresponding to the preset third weights, and the plurality of activation value quantization factors are the quantization factors corresponding to the preset activation values; based on the plurality of weight quantization factors and the plurality of activation value quantization factors, determining the target weight quantization factor and the target activation value quantization factor of any layer; the target weight quantization factor and the target activation value quantization factor can make the dequantization processing result of the quantized first model more similar to the output result of the first model before quantization than a similarity threshold; based on the target weight quantization factor of each layer in the first model and the target activation value quantization factor of each layer, performing quantization processing on the first model to obtain the target quantized first model; the above-mentioned processing of the image to be processed by the first model includes: processing the image to be processed by the target quantized first model.
[0110] This embodiment provides a model quantization method. The third weight in this embodiment represents the weight of any layer of the first model, and the third weight is used to reflect the mapping relationship between the input data and the output data between every two adjacent layers in the first model. In the current model training process, the weights are usually adjusted according to the loss function to make the output result closer to the actual result. The activation value of any layer of the first model represents the output value generated after any layer receives the input signal. The activation value is calculated through an activation function, and the activation function can be sigmoid, Rectified Linear Unit (ReLU), tanh, etc.
[0111] In the process of performing model quantization processing in this embodiment, the optimal quantized first model, that is, the target quantized first model, is obtained by adjusting the third weights and activation values of each layer of the first model. Taking the adjustment of the third weights and activation values of any layer as an example, first, a plurality of preset weight quantization factors and a plurality of activation value quantization factors are obtained. Here, the plurality of weight quantization factors and the plurality of activation value quantization factors can be set according to empirical values. By combining and traversing each weight quantization factor and each activation value quantization factor in each layer, the model performance of the quantized first model corresponding to each weight quantization factor and each activation value quantization factor is determined, and the weight quantization factor in the quantized first model with the best performance is found as the target weight quantization factor, and the activation value quantization factor in the quantized first model with the best performance is found as the target activation value quantization factor. Alternatively, according to the dequantization processing results of each quantized first model, the dequantization processing result most similar to the output result of the first model before quantization is determined to obtain the target weight quantization factor and the target activation value quantization factor.
[0112] In some embodiments, determining the target weight quantization factor and the target activation value quantization factor for any layer based on the plurality of weight quantization factors and the plurality of activation value quantization factors includes: obtaining a plurality of quantized first models based on the plurality of weight quantization factors and the plurality of activation value quantization factors; processing the first data through each quantized first model to obtain a plurality of first processing results; performing dequantization processing on each first processing result to obtain a plurality of dequantization processing results; obtaining the second processing result of the first model before quantization for the first data; among the plurality of dequantization processing results, determining the weight quantization factor corresponding to the dequantization processing result most similar to the second processing result as the target weight quantization factor; and determining the activation value quantization factor corresponding to the dequantization processing result most similar to the second processing result as the target activation value quantization factor.
[0113] This embodiment further provides a method for determining the target weight quantization factor and the target activation value quantization factor. After obtaining the preset plurality of weight quantization factors and the plurality of activation value quantization factors, each weight quantization factor and each activation value quantization factor are combined, and each layer of the first model is quantized to obtain a plurality of different quantized first models.
[0114] Perform object detection processing on the first data or the image to be detected through each quantized first model to obtain a first processing result, that is, obtain the object detection result corresponding to each quantized first model. Perform dequantization processing on each first processing result to obtain a dequantization processing result corresponding to each first processing result, and obtain multiple dequantization processing results. Perform object detection processing on the same first data or the same image to be detected through the first model before quantization to obtain the standard object detection result of the first model before quantization, that is, the second processing result. By comparing each dequantization processing result with the second processing result, determine the similarity between each dequantization processing result and the second processing result, and evaluate the processing ability of the quantized first model. The closer the dequantization processing result is to the second processing result, the better the model quantization effect of the quantized first model.
[0115] It can be determined that the quantized first model corresponding to the dequantization processing result most similar to the second processing result is the target quantized first model, and the weight quantization factor of each layer in the model is used as the target weight quantization factor corresponding to each layer, and the activation value quantization factor of each layer in the model is used as the target activation value quantization factor corresponding to each layer.
[0116] In some embodiments, obtaining multiple quantized first models based on multiple weight quantization factors and multiple activation value quantization factors includes: determining a first weight quantization factor among the multiple weight quantization factors, and through the first weight quantization factor and each activation value quantization factor among the multiple activation value quantization factors, obtaining quantization factor pairs formed by the first weight quantization factor and each activation value quantization factor; based on the quantization factor pairs corresponding to each first weight quantization factor, obtaining a first set of quantization factor pairs corresponding to the multiple weight quantization factors; performing quantization processing on the first model through the first set of quantization factor pairs to obtain multiple quantized first models; the first weight quantization factor is any one of the multiple weight quantization factors; and / or, determining a first activation value quantization factor among the multiple activation value quantization factors, and through the first activation value quantization factor and each weight quantization factor among the multiple weight quantization factors, obtaining quantization factor pairs formed by the first activation value quantization factor and each weight quantization factor; based on the quantization factor pairs corresponding to each first activation value quantization factor, obtaining a second set of quantization factor pairs corresponding to the multiple activation value quantization factors; performing quantization processing on the first model through the second set of quantization factor pairs to obtain multiple quantized first models; the first activation value quantization factor is any one of the multiple activation value quantization factors.
[0117] This embodiment further provides a method for traversing each quantization factor. According to the actual deployment requirements of the first model, each first model can be quantized and compressed to export a smaller and more efficient first model to adapt to a deployment environment with low computing resources, such as mobile devices, embedded systems, etc. Specifically, the quantization method in this embodiment can use the cosine similarity as the objective function, and alternately optimize the weight quantization factor and the activation value quantization factor to improve the model quantization accuracy, thereby achieving a faster processing speed.
[0118] For example, taking the quantization of an L-layer neural network as an example, where A l represents the activation value of the l-th layer, W l represents the weight of the l-th layer, S l represents the quantization factor of the l-th layer, and S l includes the weight quantization factor and the activation value quantization factor The result of the inverse quantization process of the l-th layer can be expressed as:
[0119]
[0120] where represents the quantization process of the activation value A based on the activation value quantization factor l , represents the quantization process of the weight W based on the weight quantization factor l ; "*" is the convolution operation. The output of each layer in the first model before quantization can be expressed as:
[0121] Q l =A l *W l (6)
[0122] To maximize the cosine similarity between the original floating-point activation value output of the l-th layer of the first model and the dequantized output obtained after quantization, the cosine similarity optimization objective function is obtained:
[0123]
[0124] In the specific implementation process, first fix the activation value quantization factor, and optimize the weight quantization factor of each layer within the search range by maximizing the objective function in formula (7). Here, the coefficients α and β can be set according to empirical values.
[0125] Then fix the weight quantization factor, search for the activation value quantization factor within the search range to maximize the objective function and optimize the activation value quantization factor.
[0126] Alternately optimize the activation value quantization factor and the weight quantization factor until the objective function in optimization formula (7) converges or exceeds the maximum number of optimization rounds.
[0127] The embodiment of the present application provides an object detection method. By integrating multiple advanced models with a large number of parameters in the academic community or the field for training, the best model is selected through comparison. At the same time, the effective information in the teacher model can be transmitted to a small model suitable for fast inference on the edge side, so as to significantly improve the detection accuracy of the small model while ensuring the inference efficiency. This processing method can effectively balance the detection accuracy and the processing efficiency, and fuse the results of multiple models through an adaptive weighting algorithm, integrate the performance gains brought by multiple different network structure algorithms, reduce the processing error of a single model on some difficult-to-classify data, and improve the accuracy of the annotation results. Different from the current manual annotation method or the method of annotating through a single pre-trained model, it greatly reduces the workload of manual annotation and has the characteristics of automatic annotation and accurate annotation.
[0128] The embodiment of the present application optimizes the model quantization effect by finely searching the quantization factor values of each layer of the neural network. By alternately searching the weight quantization factor and the activation value quantization factor of each layer of the model, the similarity of the parameter values before and after model quantization is maximized, and the optimal weight quantization factor and activation value quantization factor are found. By alternately searching the weight quantization factor and the activation value quantization factor of each layer, the error accumulation caused by extreme changes in a single quantization factor is prevented. The optimization direction of the weight quantization factor and the activation value quantization factor is controlled by dynamically adjusting the alternating step size. An error tolerance is set for the quantized model of each layer to control the error accumulation within a certain range, and protection is triggered when a large drop in accuracy occurs.
[0129] In the process of multi-model processing in the embodiment of the present application, it is ensured that the outputs of different models have sufficient consistency in the detection box, category label, and confidence score, and the output differences are avoided from being too large to affect the accuracy of the final object detection result. By dynamically adjusting the second weight of the first model and monitoring the second weights of each first model, it is prevented that the weight of a single model is too high, resulting in the neglect of the processing results of other models.
[0130] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0131] Based on the object detection method proposed in the foregoing embodiment, the embodiment of the present application also provides an object detection system, as Figure 3 shown, the object detection system includes:
[0132] An interaction module, configured to receive data to be processed, where the data to be processed includes one or more of a plurality of first data, an image to be processed, and hyperparameters of a first model; the plurality of first data is used to construct a first training dataset of the first model; the first model is used to process the image to be processed to obtain an object detection result of the image to be processed.
[0133] The object detection system provided by the embodiments of this application supports training multiple object detection algorithms. In the interaction module of the object detection system, a user can select or input data to be processed through a Web interface. The data to be processed includes algorithms for training a model, the scale of the model, the number of GPUs, the number of model training rounds, the batch size (BatchSize) of the basic data related to model training, etc. The user can write the parameter information of the above data to be processed into a json file and send a request to a specified port through the FastAPI framework. After receiving the request, the relevant processing module of the system forms a command for model training and calls the configured environment to start model training. The user can also operate, configure, and manage the entire object detection system through the Web interface. The model here can specifically be the first model.
[0134] A data processing module, configured to process each of the plurality of first data based on a plurality of second models to obtain an annotation result for each of the plurality of first data; construct a first training dataset through the plurality of first data and the annotation result for each of the plurality of first data; where the difference between the annotation result of the first data and the standard result corresponding to the first data is less than a first threshold, and the first threshold represents the difference between the processing result of a second model for the first data and the standard result corresponding to the first data; the standard result represents an annotation result determined in advance through manual annotation.
[0135] Here, the data processing module can manage and process the data in the second training dataset in its own database, and can also manage and process the plurality of first data in a specific scenario sent by the user. The data in the data processing module is stored and managed in a distributed storage form, responsible for storing and managing large-scale data, ensuring that the object detection system can process a large amount of data and supporting parallel access.
[0136] The automatic annotation module in the data processing module can process each first data through a plurality of second models, and automatically obtain the annotation result of the first data through the method provided by the above embodiments. At the same time, the manual intervention annotation module in the data processing module can also allow the data processing module to allow manual annotation intervention to ensure the quality and accuracy of data annotation. The annotation results of the data obtained by the data processing module can be stored in a json file in the yolo data format.
[0137] The automated annotation module in the data processing module also provides flexible task allocation and management tools, adding permission control to the operation of manually allocating tasks to ensure that only administrators or authorized users can perform task allocation operations. Task allocation is carried out through the Application Programming Interface (API). This interface receives the task Identification (ID) and the annotator ID, assigns the task to the specified annotator, and updates the database record.
[0138] The model training module is used to train the first model based on the first training dataset and hyperparameters.
[0139] The model training module allows users to manage and call different models, and configure the hyperparameters required for model training, supporting efficient storage and management of model training data in a distributed environment.
[0140] In the model training module, multiple advanced object detection algorithms in academia or industry are integrated. Each first model has a unique model ID, and the model is trained and called through the model ID. The model training module provides a unified algorithm call interface. The input and output of all first models are consistent. Model training, verification, evaluation, and inference are carried out through a unified API, simplifying the development process of the first model. The object detection system given in this embodiment supports the application of multiple algorithm libraries and frameworks, integrating common machine learning and deep learning frameworks such as TensorFlow, PyTorch, Scikit-learn, XGBoost, etc. Developers can pre-convert the format of the first model, and after verification processing, form a fixed model conversion script. Users can call the first model across frameworks without manually converting the data format or model format.
[0141] The model training module supports distributed training of large-scale datasets. Initialize the distributed environment on each node or GPU, specifying the globally unique identifier rank and the total number of processes world size. Then, load the first model and the first training dataset onto each GPU. Adopt the Data Parallel strategy. After calculating the gradients of each GPU node, aggregate the gradients of all nodes through the torch.distributed.all_reduce operation to complete the synchronous update of the first model parameters. The model training module also provides manual or automatic hyperparameter tuning tools such as grid search, random search, Bayesian optimization, etc. to improve the training efficiency of the first model.
[0142] The model training module also provides an automatic stopping mechanism during the training process of the first model. When the loss function of the first model converges after multiple consecutive rounds of training (the number of rounds can be specified by hyperparameters), the training can be terminated in advance. At the same time, the model training resume parameter can be used to resume training from the interrupted training progress, avoiding repeated work. After each round of model training is completed, multiple model evaluation metrics, such as accuracy, precision, recall, F1 value, AUC, etc., are automatically calculated to quickly evaluate the performance of the first model. After the model training is completed, the generalization ability of the first model is evaluated through the test set to ensure that the first model can achieve good performance in real data.
[0143] In the specific implementation process, the training of the second model can also be completed through the model training module. When training the second model through multiple algorithms, one or more algorithms with the highest detection accuracy (the highest mean average precision (MAP) value on the test set) can be selected as the second model for data annotation. In specific application scenarios, for example, in the defect recognition scenario in the industrial scenario, the physical characteristics of different defect types of different targets vary greatly, and the robustness of the same algorithm to different defects is poor. Therefore, the adaptive weighted summation method given in the above embodiments can be added during the data annotation process, and the annotation result of the data can be obtained through the processing results of multiple second models, improving the accuracy of automatic defect annotation.
[0144] Based on the object detection system given in this embodiment, taking the industrial scenario application as an example, developers can analyze the photovoltaic module data, design a post-processing algorithm according to the filtering criteria for different defects by users (such as brightness, position, area, etc.), filter the result (bounding box) obtained by model inference through the object box of the detection result and the post-processing algorithm by reading the unified filtering criteria interface file, and merge the results through the Non-Maximum Suppression (NMS) algorithm or the Weighted Boxes Fusion (WBF) algorithm. For example, for the virtual soldering defect, the bounding box coordinates obtained by model processing are represented as [x, y, w, h, score], where x, y, w, h respectively represent the horizontal and vertical coordinates of the center of the detection box, the width of the detection box, the height, and the result confidence. Reading the filtering criteria json file can determine whether the virtual soldering defect meets the conditions (height h > 10, width w > 20), and the determination result is failure or success.
[0145] In the object detection system provided by the embodiments of the present application, model migration can also be performed for a specific model of computing chip. Developers convert the first model, form a fixed conversion command after the test accuracy and speed of the model meet the requirements, and incorporate the function of the migrated first model into the overall code framework. For example, when the user selects to deploy the onnx model, the trained pytorch model is converted through a fixed command and called through a framework compatible with the onnx model. For the model acceleration function, a quantization script is written, and the deployed quantization tool is called to quantize the trained model. The function of the quantized model is incorporated into the overall code framework, and thus the complete algorithm process for photovoltaic module quality detection is obtained.
[0146] When the production line updates the production process, raw materials, or image acquisition equipment, the detection rate of a new batch of photovoltaic modules is likely to be unstable or decrease. The user can collect and upload the images that are missed or over-detected by the algorithm in the current batch of products to the object detection system. Select the second model given in the above embodiments to detect the collected data, and save the detection results in the yolo data format. The specific execution strategy is as follows: According to the training results of different algorithms, the second model with the best detection effect is selected for each type of defect and saved. During the annotation process, multiple second models are used to annotate the new data respectively to obtain annotation results, and the annotation results are fused by the method of adaptive weighted summation.
[0147] Based on the object detection system given in the above embodiments, Figure 4 The complete structural schematic diagram of the object detection system is further shown. It can be seen that in the object detection system given in the above embodiments, it also includes a processing and execution module, a deployment and monitoring module, a resource management and scheduling module, and a security and permission management module.
[0148] Among them, the online processing engine module in the processing and execution module can process the to-be-processed images input by the user in real time to obtain the object detection results, and can achieve a fast response to the application scenario. The model acceleration module and the heterogeneous computing power environment compatibility module in the processing and execution module support the compatibility of different computing resources, such as GPUs, NPUs, etc., improve the model processing efficiency through quantization operations, and at the same time are compatible with the computing capabilities of different hardware environments.
[0149] The model acceleration module in the processing and execution module can perform quantization processing on the model after the model training is completed. For example, first select 50 typical sample images in the first training dataset to calculate the objective function, and process them on the first model before quantization to obtain the second processing result. Quantize the first model parameters according to the quantization method given in the above embodiments. Assume that the search range of the weight quantization factor is set to The search range of the activation value quantization factor is The maximum number of iterations is 100 and 200 respectively. After the objective function given by formula (7) converges during quantization, or both the activation value quantization factor and the weight quantization factor reach the maximum number of iterations, the final target activation value quantization factor and the target weight quantization factor are obtained.
[0150] In specific applications, the pytroch model can also choose to perform model quantization through NVIDIA TensorRT, specifying the deployment device model as NVIDIA Jetson EVO Orin. First, convert the.pt model to an onnx file through the onnx method of pytorch, read the onnx model using the onnx parser of TensorRT, generate a TensorRT engine and save it to a.trt file, and directly load the TensorRT engine for processing during the inference process.
[0151] The deployment and monitoring module is used to manage the deployment of the model, monitor its performance in the production environment, and provide relevant logs and feedback mechanisms. Among them, the continuous integration (CI) / continuous deployment (CD) automation pipeline real-time monitoring and alerting module is used to provide CI / CD functions, monitor the running status of the system in real time during the automated model deployment process, and send alerts in a timely manner when abnormal situations occur.
[0152] The deployment and monitoring module can develop scenario-based image data preprocessing and postprocessing modules according to the needs of specific groups and users (such as the quality inspection department and the work safety department of some production enterprises). Users write filtering criteria for various detection targets (such as length, width, area, brightness, confidence, etc.) through the Web interface and write the postprocessing parameters into a json file in a fixed format. When the algorithm performs postprocessing, the target detection results of the model are filtered according to the parameters of the corresponding detection targets. After the model training is completed, the deployment and monitoring module supports exporting the trained model into various common formats (such as onnx, TensorFlow SavedModel, TorchScript, etc.) to facilitate model deployment and processing in different deployment environments.
[0153] For specific computing chips or other computing hardware, the system given in this embodiment supports customized development. Develop a fixed script to convert the neural network model trained based on the pytorch framework and call it in the customized algorithm package. The first model given in this embodiment also supports containerized deployment. Use a fixed script to package the first model into a container (such as a Docker image file) to facilitate deployment in the cloud environment or edge devices and ensure the consistency of the running environment of the first model.
[0154] Select a chip of the corresponding model according to the existing device hardware. For example, when deploying the model on the local edge device NVIDIA Jetson EVO Orin, directly save the first model as a.pt file. When deploying the model on the local edge device Atlas A500 A2 (310b chip), the target detection system converts the first model into an onnx model through a script, and then converts the onnx model into an.om model supported by the device, and imports it into the algorithm of the model that has been adapted. Users can select the parameters of preprocessing and postprocessing in the deployment and monitoring module, input the types of component defects and the corresponding filtering criteria into the platform, and form a json file recognizable by the python language. Deploy the model as an API service, use the FastAPI framework to build a RESTful API, and utilize its asynchronous characteristics to process concurrent requests.
[0155] Finally, export and save the complete trained first model on the cloud resource, mark and distinguish each first model by the version number, and replace the original first model in the original edge device after the user downloads it.
[0156] The first model retains the original version on the cloud resource, and users can download and replace it at any time.
[0157] In the resource management and scheduling module, the computing resource management module and the task scheduling module are responsible for managing and scheduling the underlying computing resources to ensure that tasks can be executed efficiently. The algorithm optimization strategy module provides optimization suggestions and automated optimization solutions for the algorithm to improve the overall performance of the system.
[0158] The security and permission management module is responsible for managing the security policies of different parts in the target detection system to ensure the overall security of the target detection system. The identity authentication and authorization module in the security and permission management module is used to ensure the security of the system, provide user authentication and permission management, and ensure the security of data and model access.
[0159] Based on the target detection method and target detection system given in the above embodiments, Figure 5 A flowchart for developing a target detection system is shown, including:
[0160] Step 501: Basic model development.
[0161] Before developing the user's personalized target detection algorithm, it is necessary to first develop the basic model. Here, the basic model can be the second model, that is, train multiple second models to enable the second model to have the ability to annotate sample data.
[0162] Step 502: Algorithm process development.
[0163] Developers analyze the data in specific application scenarios, determine the overall design logic for object detection, select appropriate algorithms for model training, and can further train the annotation capabilities of the base models.
[0164] Step 503: Data uploading and annotation.
[0165] Receive the data of specific application scenarios uploaded by users. In specific application scenarios, according to the annotation capabilities of different base models, calculate the second weight of each base model through the calculation methods given in formulas (1) to (4), and obtain the annotation result of the data based on the second weight of each base model and the processing result of each base model.
[0166] Step 504: Model training.
[0167] Construct the first training dataset based on the data annotation results obtained through the above steps, and perform model training on one or more first models based on the first training dataset, so that each first model can process the image to be processed and obtain the object detection result.
[0168] Step 505: Model deployment.
[0169] Deploy the one or more first models trained according to the hardware resources of the existing devices. Deploy the model as an API service, use the FastAPI framework to build a RESTful API, and utilize its asynchronous characteristics to process concurrent requests.
[0170] Step 506: Model acceleration.
[0171] Based on the method given in the above embodiments, perform quantization processing on one or more first models according to actual requirements.
[0172] Step 507: Model export.
[0173] Export and save one or more first models, or one or more quantized first models, on cloud resources, and mark and distinguish different first models through version numbers. After the user downloads, replace the first models in the original edge devices. Each first model retains its original version on the cloud resources, and the user can download and replace different versions of the first models at any time.
[0174] The embodiments of this application provide an object detection algorithm and an object detection system, which can integrate multiple advanced object detection models, and can be called in a standardized manner through the model input and output interfaces, allowing users to make independent selections according to their preferences. At the same time, developers can perform real-time update and maintenance on the model library containing multiple models, update the model version or supplement the latest models.
[0175] By building a target detection system, initial model development, algorithm process development, data uploading and annotation, model training, model deployment, model acceleration, final packaging and distribution of the trained model, rapid data annotation is carried out according to the customized target detection system to achieve tasks such as effective training, model deployment, and model acceleration of the model, and multiple models that can be directly deployed to the specified platform or hardware are output, achieving the effect of quickly replacing the old model.
[0176] The embodiments of the present application provide a target detection algorithm and a target detection system. By collecting a large-scale dataset with strong diversity and wide coverage and using a second model with excellent performance to generate data annotation results, more accurate training data can be obtained, thereby improving the generalization ability of the small model. Through knowledge distillation, the information of the high-precision teacher model is transmitted to the lightweight student model, enabling the student model to maintain a high detection accuracy in edge-side processing. After the student model receives the knowledge transfer from the teacher model, it can achieve fast processing in edge-side applications, and at the same time, the detection accuracy is close to that of the model with a large number of parameters, meeting the dual requirements of real-time performance and accuracy. The processing results generated by the teacher network not only contain category information but also reflect the confidence of each teacher model in processing data of different categories, enabling the student model to better understand the relationship between different data categories and improving the learning effect. According to the characteristics of each detection box, such as confidence, position, and area, the importance of the outputs of different second models is dynamically adjusted, reducing the workload of subsequent manual review and correction.
[0177] The embodiments of the present application improve the accuracy of automatic defect annotation in industrial scenarios by dynamically adjusting the weights of model outputs, enhancing the robustness of the target detection model; through an adaptive weighted summation method, such as data annotation by multiple second models, and processing the same input image by multiple trained first models to generate their respective target detection results, including detection boxes, category labels, and confidence scores; through dynamic weight assignment, according to the performance of each model on a specific task or dataset, weights are dynamically assigned to each model, and the weight assignment considers factors such as the historical performance and input features of the model, improving the accuracy of the results; through the weighted summation method, the outputs of each model are fused by weighted summation to calculate the weighted confidence score, and based on this, the final detection box and category label are selected, reducing the processing error of a single model and effectively combining the advantages of multiple models.
[0178] In the embodiments of the present application, by finely searching layer by layer for the weight quantization factor values and activation value quantization factors of each layer in the first model, and alternately searching for the weight quantization factor values and activation value quantization factors, the difference in parameter distributions before and after quantization of the first model can be effectively reduced, the model quantization effect can be improved, rather than simply taking the maximum value as the quantization factor, effectively reducing the error accumulation of each layer of the model introduced by rough quantization, enabling the model to better retain the information and accuracy of the model before quantization after quantization, having stronger adaptability in the accuracy of model quantization, and being particularly suitable for scenarios with high accuracy requirements.
[0179] Based on the object detection method proposed in the foregoing embodiments, the embodiments of the present application also provide an object detection device, as Figure 6 shown, the object detection device includes:
[0180] An acquisition module 601, configured to acquire an image to be processed.
[0181] A processing module 602, configured to process the image to be processed through a first model to obtain an object detection result of the image to be processed; the first model is trained based on a first training data set; the annotation result of the first data in the first training data set is obtained based on the processing results of the first data by multiple second models; the difference between the annotation result of the first data and the standard result corresponding to the first data is less than a first threshold, and the first threshold represents the difference between the processing result of the first data by a second model and the standard result corresponding to the first data; the standard result represents the annotation result determined in advance through manual annotation; the first data is any sample data in the first training data set.
[0182] Through the object detection device provided by the embodiments of the present application, the processing results of multiple second models can be fused, the performance gains brought by integrating multiple different model algorithms can be integrated, the calculation errors of a single model on some difficult-to-classify data can be reduced, and the accuracy of data annotation results can be improved.
[0183] In practical applications, the acquisition module 601 and the processing module 602 can be implemented based on a processor and a communication device.
[0184] In some embodiments, before acquiring the image to be processed, the acquisition module 601 is further configured to acquire a sample data set; the processing module 602 is further configured to process the first data in the sample data set through multiple second models to obtain the processing results of the first data by each second model; the first data is any data in the sample data set; determine the first weight of each second model, and based on the first weight of each second model and the processing result of the first data by each second model, determine the annotation result of the first data; based on the first data and the annotation result of the first data, construct a first training data set.
[0185] In some embodiments, the processing module 602 is specifically configured to determine the confidence level of each second model based on the processing results of multiple data by each second model; and determine the first weight of each second model based on the confidence level of each second model.
[0186] In some embodiments, before obtaining the image to be processed, the processing module 602 is further configured to train multiple first models with a first training dataset; process the image to be processed with the multiple first models to obtain the processing results of each first model; determine the second weight of each first model, and obtain the target detection result of the image to be processed based on the second weight of each first model and the processing result of each first model.
[0187] In some embodiments, the processing module 602 is specifically configured to determine the confidence level of each first model based on the target detection results of multiple images by each first model; and determine the second weight of each first model based on the confidence level of each first model.
[0188] In some embodiments, before processing the image to be processed with the first model, the processing module 602 is further configured to determine the third weight and the activation value of any layer of the first model in the first model; obtain multiple weight quantization factors and multiple activation value quantization factors; the multiple weight quantization factors are multiple quantization factors corresponding to the preset third weights, and the multiple activation value quantization factors are multiple quantization factors corresponding to the preset activation values; determine the target weight quantization factor and the target activation value quantization factor of any layer based on the multiple weight quantization factors and the multiple activation value quantization factors; the target weight quantization factor and the target activation value quantization factor can make the anti-quantization processing result of the quantized first model have a similarity greater than the similarity threshold with the output result of the first model before quantization; perform quantization processing on the first model based on the target weight quantization factor of each layer and the target activation value quantization factor of each layer in the first model to obtain the target quantized first model; process the image to be processed with the target quantized first model.
[0189] In some embodiments, the processing module 602 is specifically configured to obtain multiple quantized first models based on the multiple weight quantization factors and the multiple activation value quantization factors; process the first data with each quantized first model to obtain multiple first processing results; perform anti-quantization processing on each first processing result to obtain multiple anti-quantization processing results; obtain the second processing result of the first model before quantization for the first data; among the multiple anti-quantization processing results, determine the weight quantization factor corresponding to the anti-quantization processing result most similar to the second processing result as the target weight quantization factor; determine the activation value quantization factor corresponding to the anti-quantization processing result most similar to the second processing result as the target activation value quantization factor.
[0190] In some embodiments, the processing module 602 is specifically configured to determine a first weight quantization factor among a plurality of weight quantization factors, and obtain quantization factor pairs formed by the first weight quantization factor and each activation value quantization factor among the plurality of activation value quantization factors; based on the quantization factor pairs corresponding to each first weight quantization factor, obtain a first set of quantization factor pairs corresponding to the plurality of weight quantization factors; perform quantization processing on the first model through the first set of quantization factor pairs to obtain a plurality of quantized first models; the first weight quantization factor is any one of the plurality of weight quantization factors; and / or, determine a first activation value quantization factor among the plurality of activation value quantization factors, and obtain quantization factor pairs formed by the first activation value quantization factor and each weight quantization factor among the plurality of weight quantization factors; based on the quantization factor pairs corresponding to each first activation value quantization factor, obtain a second set of quantization factor pairs corresponding to the plurality of activation value quantization factors; perform quantization processing on the first model through the second set of quantization factor pairs to obtain a plurality of quantized first models; the first activation value quantization factor is any one of the plurality of activation value quantization factors.
[0191] It should be noted that the description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the same method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0192] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a terminal, a server, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0193] The embodiments of the present application further provide an electronic device. Figure 7 As shown in the schematic diagram of the composition structure of an electronic device provided in the embodiments of the present application, Figure 7 as shown, the electronic device 70 may include:
[0194] A memory 701 for storing executable instructions.
[0195] A processor 702, when executing the executable instructions stored in the memory 701, implements any of the above target detection methods.
[0196] The above processor 702 can be at least one of an ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.
[0197] The above computer-readable storage medium or memory 701 can be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; it can also be various terminals including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0198] The embodiments of the present application further provide a computer storage medium, on which computer-executable instructions are stored, and the computer-executable instructions are used to implement any of the target detection methods provided in the above embodiments.
[0199] Correspondingly, the embodiments of the present application further provide a computer program product, which includes computer-executable instructions, and the computer-executable instructions are used to implement any of the target detection methods provided in the above embodiments.
[0200] In some embodiments, the functions or modules included in the device provided by the embodiments of the present application can be used to execute the methods described in the method embodiments above, and its specific implementation can refer to the description of the method embodiments above. For the sake of brevity, it will not be repeated here.
[0201] The descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.
[0202] The methods disclosed in the method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0203] The features disclosed in the product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0204] The features disclosed in the method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0205] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0206] The embodiments of this application have been described above in conjunction with the accompanying drawings. However, this application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims. All of these are within the protection scope of this application.
Claims
1. A target detection method, characterized in that: The method comprises: Get the image to be processed; The image to be processed is processed by a first model to obtain a target detection result of the image to be processed; the first model is trained based on a first training data set; the labeling result of the first data in the first training data set is obtained based on the processing results of the first data by multiple second models; the difference between the labeling result of the first data and the standard result corresponding to the first data is less than a first threshold, and the first threshold represents the difference between the processing result of the first data by a second model and the standard result corresponding to the first data; the standard result represents the labeling result determined in advance by manual labeling; the first data is any sample data in the first training data set.
2. The method according to claim 1, characterized in that Before obtaining the image to be processed, the method further includes: Acquire a sample data set, and process first data in the sample data set through a plurality of second models to obtain a processing result of each second model on the first data; the first data is any data in the sample data set; Determine a first weight of each second model, and determine a labeling result of the first data based on the first weight of each second model and a processing result of each second model on the first data; A first training data set is constructed based on the first data and the annotation results of the first data.
3. The method according to claim 2, characterized in that The determining of the first weight of each second model comprises: Determine the confidence of each second model based on the processing result of each second model on the multiple data; The first weight of each second model is determined according to the confidence of each second model.
4. The method according to claim 1, characterized in that: Before obtaining the image to be processed, the method further includes: Training a plurality of first models using a first training data set; The step of processing the image to be processed by using the first model to obtain a target detection result of the image to be processed includes: Processing the image to be processed by using the multiple first models to obtain a processing result of each first model; The second weight of each first model is determined, and based on the second weight of each first model and the processing result of each first model, the target detection result of the image to be processed is obtained.
5. The method according to claim 4, characterized in that The determining of the second weight of each first model comprises: Determine the confidence of each first model based on the target detection results of each first model on the multiple images; The second weight of each first model is determined according to the confidence of each first model.
6. The method according to claim 1, characterized in that Before processing the image to be processed by the first model, the method further includes: In the first model, determining a third weight and an activation value of any layer of the first model; Acquire a plurality of weight quantization factors and a plurality of activation value quantization factors; the plurality of weight quantization factors are a plurality of quantization factors corresponding to the preset third weight, and the plurality of activation value quantization factors are a plurality of quantization factors corresponding to the preset activation value; Based on the multiple weight quantization factors and the multiple activation value quantization factors, determine a target weight quantization factor and a target activation value quantization factor of any layer; the target weight quantization factor and the target activation value quantization factor can make the similarity between the dequantization processing result of the first model after quantization and the output result of the first model before quantization greater than a similarity threshold; Based on the target weight quantization factor of each layer in the first model and the target activation value quantization factor of each layer, the first model is quantized to obtain a target quantized first model; The processing of the image to be processed by using the first model includes: processing the image to be processed by using the first model after the target is quantized.
7. The method according to claim 6, characterized in that The determining of a target weight quantization factor and a target activation value quantization factor of any layer based on the multiple weight quantization factors and the multiple activation value quantization factors includes: Based on the multiple weight quantization factors and the multiple activation value quantization factors, obtaining multiple quantized first models; Processing the first data by using each quantized first model to obtain a plurality of first processing results; performing inverse quantization processing on each first processing result to obtain a plurality of inverse quantization processing results; Obtain a second processing result of the first model on the first data before quantization; Among the multiple inverse quantization processing results, the weight quantization factor corresponding to the inverse quantization processing result most similar to the second processing result is determined as the target weight quantization factor; the activation value quantization factor corresponding to the inverse quantization processing result most similar to the second processing result is determined as the target activation value quantization factor.
8. The method according to claim 7, characterized in that The step of obtaining a plurality of quantized first models based on the plurality of weight quantization factors and the plurality of activation value quantization factors comprises: Among the multiple weight quantization factors, a first weight quantization factor is determined, and a quantization factor pair consisting of the first weight quantization factor and each activation value quantization factor is obtained through the first weight quantization factor and each activation value quantization factor in the multiple activation value quantization factors; based on the quantization factor pair corresponding to each first weight quantization factor, a first quantization factor pair set corresponding to multiple weight quantization factors is obtained; through the first quantization factor pair set, the first model is quantized to obtain multiple quantized first models; the first weight quantization factor is any quantization factor in the multiple weight quantization factors; And / or, among the multiple activation value quantization factors, determine a first activation value quantization factor, obtain the first activation value quantization factor through the first activation value quantization factor and each weight quantization factor among the multiple weight quantization factors, and form a quantization factor pair with each weight quantization factor respectively; based on the quantization factor pair corresponding to each first activation value quantization factor, obtain a second quantization factor pair set corresponding to multiple activation value quantization factors; through the second quantization factor pair set, quantize the first model to obtain multiple quantized first models; the first activation value quantization factor is any one of the multiple activation value quantization factors.
9. The method according to claim 1, characterized in that: The magnitude of parameters of the second model is greater than the magnitude of parameters of the first model.
10. The method according to claim 1, characterized in that The first model is a student model, and the second model is a teacher model.
11. A target detection system, characterized in that: The system comprises: An interactive module is used to receive data to be processed, wherein the data to be processed includes one or more of a plurality of first data, an image to be processed, and a hyperparameter of a first model; the plurality of first data are used to construct a first training data set for the first model; the first model is used to process the image to be processed to obtain a target detection result of the image to be processed; A data processing module, configured to process each of the plurality of first data based on a plurality of second models to obtain a labeling result of each of the first data; construct a first training data set through the plurality of first data and the labeling result of each of the plurality of first data; wherein the difference between the labeling result of the first data and the standard result corresponding to the first data is less than a first threshold, and the first threshold represents the difference between a processing result of the first data by a second model and the standard result corresponding to the first data; the standard result represents a labeling result determined in advance by manual labeling; A model training module is used to train the first model based on the first training data set and hyperparameters.
12. A target detection device, characterized in that: The device comprises: An acquisition module, used for acquiring an image to be processed; A processing module is used to process the image to be processed through a first model to obtain a target detection result of the image to be processed; the first model is obtained by training based on a first training data set; the labeling result of the first data in the first training data set is obtained based on the processing results of the first data by multiple second models; the difference between the labeling result of the first data and the standard result corresponding to the first data is less than a first threshold, and the first threshold represents the difference between the processing result of the first data by a second model and the standard result corresponding to the first data; the standard result represents the labeling result determined in advance by manual labeling; the first data is any sample data in the first training data set.
13. An electronic device, characterized in that: The electronic device comprises a processor and a memory for storing a computer program that can be run on the processor; wherein, The processor is configured to run the computer program to perform the method according to any one of claims 1 to 10.
14. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
15. A computer program product comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 10 when executed by a processor.