A training method and device of a detection model, a computing device, and a storage medium
By performing image enhancement and confidence correction on the target detection model, the problem of frequent adjustment of the confidence threshold during training is solved, the model training efficiency and accuracy are improved, and the training cycle is shortened.
Patent Information
- Application Number
- CN202310008084.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-04
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-01-04
AI Technical Summary
When training target detection models in image processing, existing technologies require frequent adjustment of confidence thresholds to screen predicted objects, resulting in high difficulty, heavy workload, and long training cycles.
By performing image enhancement processing on the sample image, the first and second reference images are obtained, the initial confidence of the candidate object is predicted using the first detection model, and correction is performed according to its current prediction accuracy. The target object with the corrected confidence greater than the threshold is selected, and the model parameters are adjusted in combination with the predicted object of the second detection model.
This improves the accuracy of the corrected confidence of candidate objects, reduces the need to adjust the confidence threshold at different training periods, reduces the workload and difficulty, and shortens the model training cycle.
Smart Images

Figure CN116977765B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, apparatus, computing device, and storage medium for training a detection model. Background Art
[0002] In the field of image processing technology, the learning paradigm of teacher and student models can be applied to object detection in images. The task of object detection is to find all targets (objects) of interest in an image. To train an object detection model, it is often necessary to annotate a large number of images to obtain the training data used to train the model. For example, if the target object is an elephant, all elephants in the image need to be annotated. Due to the uncontrollable nature of manual annotation, target objects are often omitted from the training data. Therefore, the teacher model predicts the training data to obtain the predicted objects (the omitted target objects) of the training data. The student model is then trained based on the training data containing the annotated target objects and the predicted objects. The teacher and student models are then optimized based on the student model's prediction results. In this process, the predicted objects output by the teacher model based on the training data are not completely correct. Therefore, it is necessary to screen out predicted objects with higher confidence based on an appropriate confidence threshold for training the student model. The setting of the confidence threshold will greatly affect the performance of the student model. For example, if the confidence threshold is too low, the training data of the student model will contain incorrect predicted objects, and the model prediction accuracy will be poor. If the confidence threshold is too high, some predicted objects will be missed during the student model prediction.
[0003] In related technologies, the prediction accuracy of the teacher model changes at different stages of training. For example, the prediction accuracy of the teacher model is very low during the initial training. As the number of training iterations increases, the prediction accuracy will also increase. Therefore, different confidence thresholds need to be set to screen prediction objects during different model training stages. However, since the model parameters need to be frequently adjusted to select the appropriate confidence threshold, the difficulty and workload of training the model increase, and the training cycle of the model is long.
[0004] Therefore, there is an urgent need for an image processing method, apparatus, computing device and storage medium that can reduce the difficulty and workload of training models and shorten the model training cycle. Summary of the Invention
[0005] The embodiments of the present application provide a method, apparatus, computing device, and storage medium for training a detection model, which can reduce the difficulty and workload of training the model and shorten the model training cycle.
[0006] In a first aspect, the embodiments of the present application provide a training method of a detection model, the method comprising:
[0007] iteratively training a first detection model and a second detection model based on a sample image set, each sample image containing a plurality of sample objects that have been labeled; wherein in one iteration process, the following operations are performed:
[0008] performing twice image enhancement processing on the selected sample image to obtain a corresponding first reference image and a second reference image, respectively;
[0009] obtaining, by the first detection model, at least one candidate object other than the sample objects in the first reference image and an initial confidence of each of the at least one candidate object;
[0010] correcting the obtained at least one initial confidence according to a current prediction accuracy of the first detection model to obtain a corresponding corrected confidence, and selecting at least one target object from the at least one candidate object, the corrected confidence of which is greater than a confidence threshold;
[0011] obtaining, by the second detection model, a plurality of predicted objects in the second reference image, and adjusting parameters of the first detection model and the second detection model based on a loss value obtained based on the plurality of predicted objects, the at least one target object, and the sample objects in the first reference image.
[0012] In a second aspect, the embodiments of the present application provide a training device of a detection model, the device comprising: an image enhancement unit, a first detection unit, a correction unit, a second detection unit, and an update unit, the training device being configured to iteratively train a first detection model and a second detection model based on a sample image set, each sample image containing a plurality of sample objects that have been labeled; wherein in one iteration process, the following operations are performed:
[0013] performing twice image enhancement processing on the selected sample image to obtain a corresponding first reference image and a second reference image, respectively, based on the image enhancement unit;
[0014] obtaining, by the first detection model, at least one candidate object other than the sample objects in the first reference image and an initial confidence of each of the at least one candidate object, based on the first detection unit;
[0015] correcting the obtained at least one initial confidence according to a current prediction accuracy of the first detection model to obtain a corresponding corrected confidence, and selecting at least one target object from the at least one candidate object, the corrected confidence of which is greater than a confidence threshold, based on the correction unit;
[0016] Based on the second detection unit, multiple predicted objects in the second reference image are obtained through the second detection model, and through the update unit, the loss values obtained based on the multiple predicted objects, the at least one target object and the sample objects in the first reference image are used to adjust the parameters of the first detection model and the second detection model respectively.
[0017] Optionally, the first detection unit is specifically configured to:
[0018] Predicting, by using the first detection model, a plurality of objects to be verified contained in the first reference image, wherein the plurality of objects to be verified include sample objects in the first reference image;
[0019] respectively obtaining a degree of overlap between the plurality of objects to be verified and at least one sample object in the first reference image;
[0020] From the multiple objects to be verified, objects to be verified whose overlap degree is less than a first set threshold are selected as candidate objects.
[0021] Optionally, the first detection unit is further configured to:
[0022] Selecting at least one discarded object whose overlap degree is not less than the first set threshold from the multiple objects to be verified, wherein the at least one discarded object includes a sample object in the first reference image;
[0023] Putting the at least one discarded object into a discarded object queue, and obtaining the sum of the number of the at least one discarded object and the number of historical discarded objects in the discarded object queue, where the historical discarded objects are discarded objects obtained in each round of training before the current round of training;
[0024] If the sum of the quantities is not less than the quantity threshold, the current prediction accuracy of the first detection model is determined based on the degree of overlap of the at least one discarded object and the degree of overlap of each historical discarded object, and the prediction accuracy represents: the probability that the at least one discarded object and each historical discarded object each belong to a sample object.
[0025] Optionally, the first detection unit is specifically configured to perform the following operations respectively for multiple preset value ranges:
[0026] Determining a first number of discarded objects having initial confidence levels belonging to a value range from the discarded object queue;
[0027] determining, from the discarded objects, a second number of discarded objects whose overlap exceeds a second set threshold, where the second set threshold is greater than the first set threshold;
[0028] Using the ratio of the second number to the first number as the value accuracy of the one value range;
[0029] The prediction accuracy is obtained based on the value accuracy rates corresponding to each of the multiple value ranges.
[0030] Optionally, the first detection unit is also used to, if the sum of the quantities is greater than the quantity threshold, delete the corresponding number of historical discarded objects that were placed earliest from the discarded object queue based on the sum of the quantities, so that the number of each discarded object in the discarded object queue is equal to the quantity threshold.
[0031] Optionally, the current prediction accuracy of the first detection model includes multiple value accuracy rates, each value accuracy rate corresponds to a preset value range;
[0032] The correction unit is specifically used to:
[0033] Determining the accuracy of the values corresponding to the initial confidence levels of the at least one candidate object according to the value ranges to which the initial confidence levels of the at least one candidate object belong respectively;
[0034] According to the value accuracy corresponding to each of the initial confidence levels, each of the initial confidence levels is corrected to obtain a corresponding corrected confidence level.
[0035] Optionally, the initial models of the first detection model and the second detection model are the same;
[0036] The updating unit is specifically configured to:
[0037] Obtaining a loss value based on differences between the multiple predicted objects and the at least one target object and a sample object in the first reference image;
[0038] adjusting a second model parameter of the second detection model according to the loss value;
[0039] The first model parameters of the first detection model are adjusted according to the adjusted second model parameters and the set coefficients.
[0040] In a third aspect, an embodiment of the present application provides a computer device comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes any one of the detection model training methods in the first aspect above.
[0041] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium including a computer program. When the computer program is run on a computer device, the computer program is configured to cause the computer device to perform the training method of any one of the detection model in the first aspect.
[0042] In a fifth aspect, an embodiment of the present application provides a computer program product including a computer program. The computer program is stored in a computer readable storage medium. When a processor of a computer device reads the computer program from the computer readable storage medium, the processor executes the computer program, so that the computer device performs the training method of any one of the detection model in the first aspect.
[0043] The present application has the following beneficial effects:
[0044] The training method of the detection model, the device, the computer device and the storage medium provided by the embodiments of the present application can perform image enhancement on the selected sample image, and obtain the corresponding first reference image and the second reference image in the iterative training process of the first detection model and the second detection model based on the sample image set. Further, the first detection model can be used to predict the first reference image, and at least one target object can be obtained. The at least one target object and the sample image in the first reference image are assumed to be the standard prediction result of the current training. According to the difference between the standard prediction result and the multiple prediction objects (including the sample image in the second reference image) predicted by the second detection model from the second reference image, a loss value is obtained. The first detection model and the second detection model are adjusted according to the loss value, and the current training is completed. In addition, in the process of predicting the first reference image by the first detection model to obtain at least one target object, the first detection model predicts the first reference image to obtain at least one candidate object that does not contain the sample image in the first reference image. Further, according to the prediction accuracy of the first detection model, the initial confidence of the at least one candidate object is corrected to obtain a corresponding corrected confidence. From the at least one candidate object, at least one target object with a corrected confidence greater than a confidence threshold is selected. Compared with the related art, different confidence thresholds are required in different periods to select more accurate target objects from candidate objects, which leads to complicated operation, large workload and long training period. In the present application, the initial confidence of the candidate object is corrected based on the prediction accuracy of the first detection model, so that the corrected confidence of the candidate object is more accurate, and the confidence threshold does not need to be changed in different periods, thereby solving the problems of complicated operation, large workload and long training period in the related art.
[0045] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0047] Figure 1 An optional schematic diagram of an application scenario provided in an embodiment of the present application;
[0048] Figure 2 A flowchart of a method for training a detection model provided in an embodiment of the present application;
[0049] Figure 3 A schematic diagram of a sample image provided in an embodiment of the present application;
[0050] Figure 4 A schematic diagram of a target object in a sample image provided in an embodiment of the present application;
[0051] Figure 5 A schematic diagram of at least one candidate object in a sample image provided in an embodiment of the present application;
[0052] Figure 6 A schematic diagram of at least one candidate object in a sample image provided in an embodiment of the present application;
[0053] Figure 7 A schematic diagram of an iterative process of a detection model provided in an embodiment of the present application;
[0054] Figure 8 A schematic diagram of an iterative process of a detection model provided in an embodiment of the present application;
[0055] Figure 9 A flowchart of a method for training a detection model provided in an embodiment of the present application;
[0056] Figure 10 A schematic diagram of multiple objects to be verified contained in a first reference image provided in an embodiment of the present application;
[0057] Figure 11A schematic diagram of the degree of overlap between an object to be verified and a corresponding sample object provided in an embodiment of the present application;
[0058] Figure 12a A schematic diagram of the degree of overlap between an object to be verified and a corresponding sample object provided in an embodiment of the present application;
[0059] Figure 12b A schematic diagram of the degree of overlap between an object to be verified and a corresponding sample object provided in an embodiment of the present application;
[0060] Figure 13a A schematic diagram of the degree of overlap between an object to be verified and a corresponding sample object provided in an embodiment of the present application;
[0061] Figure 13b A schematic diagram of the degree of overlap between an object to be verified and a corresponding sample object provided in an embodiment of the present application;
[0062] Figure 14a A schematic diagram of the degree of overlap between an object to be verified and a corresponding sample object provided in an embodiment of the present application;
[0063] Figure 14b A schematic diagram of the degree of overlap between an object to be verified and a corresponding sample object provided in an embodiment of the present application;
[0064] Figure 15 A flowchart of a method for training a detection model provided in an embodiment of the present application;
[0065] Figure 16 A flowchart of a method for training a detection model provided in an embodiment of the present application;
[0066] Figure 17 A flowchart of a method for training a detection model provided in an embodiment of the present application;
[0067] Figure 18 A schematic diagram of a discarded object queue provided in an embodiment of the present application;
[0068] Figure 19 A schematic diagram of a device for training a detection model provided in an embodiment of the present application;
[0069] Figure 20 A schematic diagram of the hardware structure of a computer device to which the embodiments of the present application are applied;
[0070] Figure 21 A schematic diagram of the hardware structure of another computer device to which an embodiment of the present application is applied. DETAILED DESCRIPTION
[0071] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application. The embodiments in the present application and the features in the embodiments can be combined with each other in a non-conflicting manner. Moreover, although a logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in an order different from that herein.
[0072] It can be understood that, in the following specific embodiments of the present application, data related to a sample image set is involved, and when the embodiments of the present application are applied to specific products or technologies, relevant permissions or consents need to be obtained, and the collection, use and processing of the relevant data need to comply with relevant laws, regulations and standards of countries and regions. For example, when relevant data needs to be obtained, relevant agreements of volunteer authorization data can be signed by recruiting relevant volunteers, and then the data of the volunteers can be used for implementation; or implementation is performed within the internal scope of an organization that has been authorized to allow, and the implementation of the following embodiments is performed by using the data of internal members to make relevant recommendations to the internal members; or the relevant data used in the specific implementation is all simulation data, for example, simulation data generated in a virtual scene.
[0073] In order to facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used by the embodiments of the present application are explained first:
[0074] Image enhancement technology is used to enhance useful information in an image. It can be a distortion process, and the purpose is to improve the visual effect of the image, aiming at the application occasion of a given image. The overall or local characteristics of the image are emphasized purposefully, the original unclear image becomes clear or some features of interest are emphasized, the differences between different object features in the image are enlarged, the features of no interest are suppressed, the image quality is improved, the information amount is enriched, the image interpretation and recognition effect is strengthened, and the needs of some special analysis are met. For example, randomly reducing or enlarging an image, randomly translating part of an image, randomly blurring part of an image, etc.
[0075] The intersection over union (loU) function calculates the ratio of the intersection and the union of two bounding boxes. The union of the two bounding boxes is the area containing the two bounding boxes, and the intersection is the area of the two bounding boxes that overlap. Then, the intersection over union is the ratio of the area of the intersection to the area of the union.
[0076] EMA (Exponential Moving Average) is an exponential moving average. Also known as the EXPMA indicator, it is also a trend indicator. The exponential moving average is an exponentially decreasing weighted moving average.
[0077] The embodiments of the present application relate to artificial intelligence and machine learning (ML) technology, and are mainly designed based on machine learning in artificial intelligence.
[0078] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0079] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0080] Machine learning is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0081] Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. Artificial neural networks (ANNs) abstract the neural networks in the human brain from an information processing perspective, building a simple model that forms different networks based on different connection structures. A neural network is a computational model composed of a large number of interconnected nodes (or neurons). Each node represents a specific output function, called an activation function. Each connection between two nodes represents a weighted value for the signal passing through that connection, called a weight. This serves as the memory of the artificial neural network. The network's output varies depending on the network's connection structure, weights, and activation function. The network itself is often an approximation of a natural algorithm or function, or it may express a logical strategy.
[0082] In the application of target detection on images, the embodiments of the present application do not limit the detection method used. For example, the first detection model and the second detection model can be applied convolutional neural networks, support vector machines (SVM), Bayesian classification, etc. and their combinations. The specific structures of the first detection model and the second detection model are not limited here. The training part of the first detection model and the second detection model in the embodiments of the present application relates to the technical field of machine learning. In the training part, the first detection model and the second detection model are trained by the machine learning technology, so that the first detection model and the second detection model are trained based on each sample image in the sample image set given in this embodiment, and the model parameters are continuously adjusted by the optimization algorithm until the model converges, so that the obtained first detection model and the second detection model can accurately predict the target object in the image. In the application part, the first detection model or the second detection model obtained in the training part can be used to predict the target object of the image, etc. In addition, it should be noted that the first detection model and the second detection model in the embodiments of the present application can be either online training or offline training, which is not specifically limited here. In this article, offline training is used as an example for illustration.
[0083] The following is a brief introduction to the design concept of the embodiment of this application:
[0084] In deep learning AI, knowledge distillation technology is commonly used as a learning paradigm to transfer knowledge from a large teacher model to a shallow student model, thereby improving the student's performance. The teacher model is often complex in structure and possesses good performance and generalization capabilities, while the student model is simple in structure and has limited expressive power. This allows the knowledge learned by the teacher model to guide the training of the student model, resulting in a student model with comparable performance but with significantly fewer parameters, thus compressing the student model while maintaining its processing efficiency.
[0085] In related technologies, knowledge distillation technology is applied to the field of image processing technology, and target detection is performed through the learning paradigm of the teacher model and the student model in the knowledge distillation technology. Since the accuracy of the confidence of the prediction results obtained by the teacher-student model in different training periods is different, in order to accurately screen out the missed target objects from the prediction results (including the labeled target objects) obtained by predicting the corresponding image from the teacher model, different confidence thresholds need to be set at different times to ensure the accuracy of the screened target objects. This method requires frequent adjustment of model parameters to select a suitable confidence threshold. In addition, under different loss functions, different teacher model structures, and different student model structures, it is necessary to perform the operation of selecting a suitable confidence threshold multiple times. As a result, not only is the training model difficult and labor-intensive, but the model training cycle is long, and when the loss function changes or the model structure changes, the confidence thresholds for different training periods need to be re-determined, further increasing the intensity and difficulty of the confidence threshold selection step.
[0086] In view of this, the embodiment of the present application provides a training method and device of a detection model, a computer device and a storage medium. In a model training process, at least one candidate object and its respective initial confidence in a first reference image except a sample object are obtained by a first detection model. According to the current prediction accuracy of the first detection model, the initial confidence of each of the obtained at least one candidate object is corrected to obtain a corresponding corrected confidence. At least one target object with a corrected confidence greater than a confidence threshold is selected from the at least one candidate object. Subsequently, based on the at least one target object, the sample object in the first reference image, and a plurality of prediction objects predicted by a second detection model for a second reference image, the first detection model and the second detection model are parameterized. Thus, the initial confidence of the candidate object is corrected in the model training, the accuracy of the corrected confidence is higher, and more missed labeled target objects can be obtained in the early training stage. In the later training stage, the false positives of the target objects can be suppressed. Therefore, the loss value accuracy is higher based on the plurality of prediction objects corresponding to the second detection model, the at least one target object corresponding to the first detection model, and the sample object in the first reference image. Thus, the parameterization of the first detection model and the second detection model is more accurate, the model training efficiency is improved, and different confidence thresholds do not need to be changed in different training periods of the model. The workload, work difficulty, and model training period are reduced. This adjustment method of the initial confidence can also be applied to different loss functions and different model structures, further reducing the workload, work difficulty, and period of model training of different types and structures.
[0087] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application, and the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0088] As shown in Figure 1 , it is an application scenario diagram of the embodiment of the present application. The application scenario diagram includes any one of a plurality of terminal devices 110 and any one of a plurality of servers 120. The training of the detection model in the embodiment of the present application can be applied to various scenarios, such as intelligent transportation (such as detection of roads and signboards in autonomous driving, target detection of vehicles involved in accidents, etc.), face recognition, industrial detection (such as chip defect detection under a focused ion beam and a scanning electron microscope), etc.
[0089] As shown in Figure 1 , it is an application scenario diagram of the embodiment of the present application. It is an application scenario diagram of the embodiment of the present application. The application scenario diagram includes two terminal devices 110 and one server 120.
[0090] The terminal device 110 and the server 120 can communicate through a communication network. In a face recognition scenario, a user can collect a face image through the terminal device 110. The terminal device 110 can be installed with an application related to face recognition, such as payment software, personal information management software, and the like. The application involved in the embodiments of the present application can be software, a webpage, a mini-program, or the like, and the background server is a background server corresponding to the software or the webpage, the mini-program, or the like, without limitation on the specific type of the client. In a smart traffic scenario, a user can collect a road driving video through the terminal device 110. The terminal device 110 can be installed with an application related to vehicle driving record management. The application involved in the embodiments of the present application can be software, a webpage, a mini-program, or the like, and the background server is a background server corresponding to the software or the webpage, the mini-program, or the like, without limitation on the specific type of the client. The specific application scenario of the detection model is not limited herein.
[0091] In an optional implementation, the communication network is a wired network or a wireless network. The terminal device 110 and the server 120 can be directly or indirectly connected through wired or wireless communication, without limitation in the present application.
[0092] In the embodiments of the present application, the terminal device 110 is an electronic device used by a user. The electronic device can be a personal computer, a mobile phone, a tablet computer, a notebook computer, an electronic book reader, a smart home, or the like, which has a certain computing capability and runs instant messaging software and websites or social software and websites. Each terminal device 110 is connected to the server 120 through a wireless network. The server 120 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs (Content Delivery Networks), and basic cloud computing services such as big data and artificial intelligence platforms.
[0093] The first detection model and the second detection model can be deployed on the server 120 for training. The server 120 can store a large number of sample images for training the first detection model and the second detection model. Optionally, after the first detection model and the second detection model are trained based on the training method in the embodiments of the present application, the trained first detection model or the second detection model can be directly deployed on the server 120 or the terminal device 110. Generally, the first detection model / second detection model is directly deployed on the server 120.
[0094] In one possible application scenario, the training samples in this application can be stored using cloud storage technology. Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (also known as storage nodes) in the network through application software or application interfaces to work together and provide external data storage and service access functions.
[0095] It should be noted that Figure 1 The examples shown are only for illustration. In fact, the number and communication mode of terminal devices and servers are not limited and are not specifically limited in the embodiments of this application.
[0096] The following describes the training method of the detection model provided by the exemplary embodiment of the present application in combination with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present application, and the implementation of the present application is not limited in this respect.
[0097] See also Figure 2 As shown, it is a flow chart of the training method of the detection model provided in the embodiment of the present application. Here, the server is used as an example for illustration. The specific implementation process of the method is as follows:
[0098] Iteratively train the first detection model and the second detection model based on a set of sample images, each sample image comprising: a plurality of labeled sample objects; wherein, during one iteration, the following operations are performed:
[0099] Step 201: performing image enhancement processing twice on the selected sample image to obtain a corresponding first reference image and a second reference image respectively;
[0100] In the embodiment of the present application, the sample image can be a face image, a person image, an animal image, an animation image, etc., and there is no specific limitation on the sample image. Accordingly, the sample object in the sample image can be a face, a person, an animal, an object in an animation, a person, etc. The sample image contains the sample objects that have been marked, and can also contain the target objects that have not been marked. For the sake of understanding, a sample image is given here, such as Figure 3 As shown in FIG, a sample image provided by an embodiment of the present application, the sample objects that have been labeled are cat a, cat b, cat c, and cat d; Figure 3 There is no annotation in the target object, that is, the target object that is missed is as follows Figure 4As shown, there are cats e, f, and g. It should be noted that the sample images and the target objects not labeled in the sample images are only examples and do not limit the specific sample images of this solution.
[0101] In an embodiment of the present application, the method of performing two image enhancement processes on the selected sample image may include: performing a random brightness enhancement on the sample image once to obtain a first reference image, and performing a random brightness reduction on the sample image once to obtain a second reference image; or, performing a random scaling on the sample image once to obtain a first reference image, and performing another random scaling on the sample image to obtain a second reference image; or, performing a random blurring on the sample image once to obtain a first reference image, and performing another random blurring on the sample image to obtain a second reference image, etc., wherein the obtained first reference image and second reference image may be exactly the same or different, and there is no specific limitation on the first reference image and the second reference image.
[0102] Step 202: Obtaining at least one candidate object other than the sample object in the first reference image and their respective initial confidence scores using the first detection model;
[0103] Based on the sample image in the above example, assuming that the sample image in the above example also contains the sun h, the embodiment of the present application provides an example, assuming that the sample image is scaled by a certain ratio to obtain a first reference image, then in one possibility, at least one candidate object other than the sample objects (cat a, cat b, cat c, cat d) in the first reference image can be obtained through the first detection model, such as Figure 5 As shown, at least one candidate object includes: cat e, cat f, cat g, and possibly sun h. The initial confidence level of cat e is 70%, the initial confidence level of cat f is 60%, the initial confidence level of cat g is 75%, and the initial confidence level of sun h is 30%.
[0104] Based on the sample image in the above example, the embodiment of the present application provides an example. Assuming that the sample image is randomly locally translated to obtain a first reference image, then in one possibility, at least one candidate object other than the sample objects (cat a, cat b, cat c, cat d) in the first reference image can be obtained through the first detection model, such as Figure 6 As shown, at least one candidate object includes cat e, cat f, cat g, and possibly sun h. The initial confidence level for cat e is 70%, for cat f is 61%, for cat g is 75%, and for sun h is 30%. It should be noted that the candidate objects in the above example are only used to clarify this solution and do not limit it.
[0105] Step 203: correcting the obtained at least one initial confidence according to the current prediction accuracy of the first detection model respectively, obtaining a corresponding corrected confidence, and selecting at least one target object with a corrected confidence greater than a confidence threshold from the at least one candidate object;
[0106] In the embodiment of the present application, the current prediction accuracy of the first detection model can be determined according to the prediction accuracy of the historical detection model. For example, the accuracy of the historical detection model in predicting the sample object and the target object in the sample image is different at different training periods during the training process. Therefore, the accuracy of the historical detection model at the corresponding training period can be used as the prediction accuracy of the first detection model at the corresponding training period.
[0107] In the embodiment of the present application, the current prediction accuracy of the first detection model can be determined according to the accuracy of the first detection model in predicting the sample object in the last iteration or the previous multiple iterations. For example, a method can be set in the first detection model to collect the prediction results of a set number of sample images, and the hit rate of the prediction results hitting the sample object in the sample image (for example, the sample image contains 4 sample objects, and the prediction result predicts 2 sample objects, so the hit rate is 50%). The average of the obtained set number of hit rates is obtained, and the average hit rate is used as the prediction accuracy of the next round of the first detection model.
[0108] In the embodiment of the present application, the current prediction accuracy of the first detection model is different at different training periods, which can be determined according to the test sample image (the test sample image does not have a missing labeled target object). For example, the iteration training of the first detection model can be divided into multiple training periods, and a set number of test sample images of the training period are set for each training period. Before the first detection model is trained in any training period of the multiple training periods, the set number of test sample images of the training period are predicted by the first detection model to obtain the predicted sample object of each test sample image. Further, the hit rate of the sample object prediction of each test sample image is obtained, the hit rate of the set number of test sample images is obtained, and the average hit rate is obtained by averaging the obtained set number of hit rates. The average hit rate is used as the prediction accuracy of the corresponding training period of the first detection model. It should be noted that the prediction accuracy can be set as needed, and the specific method of obtaining the prediction accuracy is not limited here.
[0109] In an embodiment of the present application, assuming that the current prediction accuracy of the first detection model is 70%, three candidate objects are predicted for the first reference image, the initial confidence of candidate object 1 is 80%, the initial confidence of candidate object 2 is 40%, and the initial confidence of candidate object 3 is 60%. Then the corrected confidence of candidate object 1 can be 80%*70%, the corrected confidence of candidate object 2 can be 40%*70%, and the corrected confidence of candidate object 3 can be 60%*70%. If the confidence threshold is 50%, the corrected confidence of candidate object 1 80%*70% is greater than the confidence threshold of 50%, and candidate object 1 is the target object.
[0110] In the embodiment of the present application, assuming that the current prediction accuracy of the first detection model is 70%, three candidate objects are predicted for the first reference image, the initial confidence of candidate object 1 is 80%, the initial confidence of candidate object 2 is 40%, and the initial confidence of candidate object 3 is 60%, then the corrected confidence of candidate object 1 can be = 80% * 70% * α + 0.1, the corrected confidence of candidate object 2 can be = 40% * 70% * α + 0.1, and the corrected confidence of candidate object 3 can be = 60% * 70% * α + 0.1. If the confidence threshold is 70% and α = 1.2, then the corrected confidence of candidate object 1 80% * 70% * α + 0.1 is greater than the confidence threshold of 70%, and candidate object 1 is the target object. It should be noted that the specific correction method for the initial confidence in this solution is not limited here and can be set as needed.
[0111] Step 204: Obtain multiple predicted objects in the second reference image through the second detection model, and use the loss values obtained based on the multiple predicted objects, the at least one target object and the sample objects in the first reference image to adjust the parameters of the first detection model and the second detection model respectively.
[0112] In the embodiment of the present application, based on the sample images in the above example (including sample objects: cat a, cat b, cat c, cat d), the first reference image and the second reference image, as shown in FIG. Figure 7 As shown, the sample image is randomly scaled twice by image enhancement technology to obtain a first reference image and a second reference image respectively. The first reference image is predicted by the first detection model to obtain at least one candidate object (such as Figure 7 As shown, at least one candidate object includes: cat e, cat f, cat g, sun h), after correcting the initial confidence of at least one candidate object according to the prediction accuracy, the corrected confidence is screened to obtain at least one target object (such as Figure 7As shown, at least one target object includes: cat e, cat f, cat g). In addition, the second reference image is predicted by the second detection model to obtain multiple predicted objects (multiple predicted objects include: cat a, cat b, cat c, cat d, cat e, cat f, cat g, sun h), and the sample objects marked in the sample image: cat a, cat b, cat c, cat d and at least one target object: cat e, cat f, cat g, and the loss values of the multiple predicted objects: cat a, cat b, cat c, cat d, cat e, cat f, cat g, sun h are obtained. The first detection model and the second detection model are adjusted according to the loss values. It should be noted here that the image enhancement method, sample image, at least one candidate object predicted, at least one target object, and multiple predicted objects involved in this embodiment are only examples provided by this application to clearly illustrate this solution, and do not limit the specific implementation of this solution. They can be specifically set according to the model structure, model parameters, model training period, etc. of the first detection model and the second detection model. For example, in Figure 8 In the method, the sample image is randomly scaled by image enhancement technology to obtain a first reference image, and the sample image is randomly locally translated to obtain a second reference image. After that, the first reference image is predicted by the first detection model, and the second reference image is predicted by the second detection model; wherein the first detection model predicts at least one candidate object (such as Figure 8 As shown, at least one candidate object includes: cat e, cat f, cat g, sun h), after correcting the initial confidence of at least one candidate object according to the prediction accuracy, the corrected confidence is screened to obtain at least one target object (such as Figure 8 As shown, at least one target object includes: cat e, cat f, cat g). The second detection model predicts multiple predicted objects (the multiple predicted objects include: cat a, cat b, cat c, cat d, cat g, and sun h). The sample objects labeled in the sample image: cat a, cat b, cat c, cat d and at least one target object: cat e, cat f, cat g, and sun h are obtained. The loss values of the multiple predicted objects: cat a, cat b, cat c, cat d, cat g, and sun h are obtained, and the parameters of the first detection model and the second detection model are adjusted according to the loss values.
[0113] In the above method, during the iterative training of the first detection model and the second detection model based on the sample image set, the selected sample image can be enhanced to obtain the corresponding first reference image and the second reference image. Furthermore, for the first reference image and the second reference image predicted from a sample image, the first detection model can be used to predict the first reference image to obtain at least one target object, and the at least one target object and the sample image in the first reference image are assumed to be the standard prediction results of this round of training (e.g., in the above example, at least one target object: cat e, cat f, cat g and sample objects: cat a, cat b, cat c, cat d). Based on the difference between this standard prediction result and the multiple predicted objects predicted by the second detection model for the second reference image (including the sample image in the second reference image, such as, in the above example, the multiple predicted objects: cat a, cat b, cat c, cat d, cat g, sun h), a loss value is obtained, and the parameters of the first detection model and the second detection model are adjusted according to the loss value to complete this round of training. In this process, the initial confidence of at least one candidate object is corrected based on the current prediction accuracy of the first detection model to obtain a corresponding corrected confidence, and at least one target object whose corrected confidence is greater than a confidence threshold is selected from the at least one candidate object. Compared with the related art, different confidence thresholds need to be adopted at different times to screen out more accurate target objects from the candidate objects, resulting in cumbersome and complicated operations, large workload, and long training cycles. In this application, the initial confidence of the candidate object can be corrected based on the prediction accuracy of the first detection model, so that the corrected confidence of the candidate object is more accurate, without the need to change the confidence threshold at different times, thereby solving the problems of cumbersome and complicated operations, large workload, and long training cycles in the related art.
[0114] Based on the above detection model training method, the embodiment of the present application provides a detection model training method, such as Figure 9 As shown, in step 202, obtaining at least one candidate object contained in the first reference image through the first detection model includes:
[0115] Step 901: Predicting multiple objects to be verified contained in the first reference image using the first detection model, where the multiple objects to be verified include sample objects in the first reference image;
[0116] In the embodiment of the present application, the multiple objects to be verified are the prediction results of the first detection model on the first reference image, and the multiple objects to be verified include sample objects in the first reference image, such as Figure 10As shown, the multiple objects to be verified may include: sample objects: (marked with dotted boxes to facilitate distinction from candidate objects) cat a, cat b, cat c, cat d, and at least one candidate object: cat e, cat f, cat g, sun h. It should be noted that the multiple objects to be verified here are only examples. The multiple objects to be verified may only include: sample objects: cat b, cat c, cat d, and at least one candidate object: cat e, cat f, cat g, sun h; or, the multiple objects to be verified may only include: sample objects: cat b, cat c, cat d, and at least one candidate object: cat g, sun h. There is no limitation on the specific sample objects and candidate objects included in the multiple objects to be verified.
[0117] Step 902: Obtain the degree of overlap between the multiple objects to be verified and at least one sample object in the first reference image respectively;
[0118] based on Figure 10 In the embodiment of the present application, taking cat a among multiple objects to be verified as an example, and taking the image area of cat a as a reference for the degree of overlap, the degree of overlap between cat a among multiple objects to be verified and at least one sample object in the first reference image is obtained: Figure 11 As shown, the degree of overlap between cat a in multiple objects to be verified and cat a in the sample object is completely overlapped (the overlapping part is the shaded part), as shown in Figure 12a As shown in , the degree of overlap between cat a in the multiple objects to be verified and cat b in the sample objects is: only one-sixth of the image area of cat a overlaps (the overlapping part is the shaded part), as shown in Figure 13a As shown in the figure, the degree of overlap between cat a in the multiple objects to be verified and cat c in the sample object is: only one-seventh of the image area of cat a overlaps (the overlapping part is the shaded part), as shown in the figure. Figure 14a As shown, the degree of overlap between cat a among the multiple objects to be verified and cat d among the sample objects is: only one-sixth of the image area of cat a overlaps (the overlapping part is the shaded part).
[0119] Alternatively, the overlap degree can also be the intersection-and-union ratio of the object to be verified and the corresponding sample object. For example, cat a among multiple objects to be verified is: Figure 11 As shown, the degree of overlap between cat a in the multiple objects to be verified and cat a in the sample object is completely overlapped (the overlapping part is the shaded part), that is, the intersection and union of cat a in the multiple objects to be verified and cat a in the sample object are equal, then the intersection and union ratio of cat a in the multiple objects to be verified and cat a in the sample object is 100%, as shown in Figure 12bAs shown, the degree of overlap between cat a in the multiple objects to be verified and cat b in the sample object is: the ratio of the intersection (shaded by the diagonal lines) and the union (shaded by the diagonal lines and the shaded by the dots) of cat a in the multiple objects to be verified and cat b in the sample object is 10%. Figure 13b As shown, the degree of overlap between cat a in the multiple objects to be verified and cat c in the sample object is: the ratio of the intersection (shaded by the oblique lines) and the union (shaded by the oblique lines and the shadows of the dots) of cat a in the multiple objects to be verified and cat c in the sample object is 8%, as shown in Figure 14b As shown, the degree of overlap between cat a among the multiple objects to be verified and cat d among the sample objects is: the ratio of the intersection (shaded by the diagonal lines) to the union (shaded by the diagonal lines and the dot) of cat a among the multiple objects to be verified and cat d among the sample objects is 10%. It should be noted that the degree of overlap can be set as needed, and the method for obtaining the degree of overlap here is only an example, and the specific method for obtaining the degree of overlap is not limited.
[0120] Step 903: Select, from the multiple objects to be verified, objects to be verified whose overlap degree is less than a first set threshold as candidate objects.
[0121] In the embodiment of the present application, assuming that the first set threshold is that the overlap degree is nine-tenths, then in the above Figure 10 In the sample image, it can be obtained that the objects to be verified whose overlap degree is less than the first set threshold include: cat e, cat f, cat g, and sun h, then cat e, cat f, cat g, and sun h are candidate objects. It should be noted that the first set threshold and candidate objects in the examples of this application are only used to clearly illustrate this scheme, and do not limit the selection of the first set threshold in this scheme, nor the number of candidate objects. For example, the first set threshold can be set according to the overlap between the sample object and the candidate object in the historical sample image, and the number of candidate objects can also depend on the model structure and model parameters of the first detection model and the relevant content in the first reference image. There is no specific restriction on the first set threshold and candidate objects here.
[0122] Based on the above Figure 9 The detection model training method in this application embodiment provides another detection model training method, such as Figure 15 As shown, the method further includes:
[0123] Step 1501: Select at least one discarded object whose overlap degree is not less than a first set threshold from the plurality of objects to be verified, wherein the at least one discarded object includes a sample object in the first reference image;
[0124] Based on the above example Figure 7 In the embodiment of the present application, Figure 16As shown, the first detection model detects multiple objects to be verified, including: cat a, cat b, cat c, cat d, and at least one candidate object: cat e, cat f, cat g, sun h, object i. Based on the degree of overlap, at least one discarded object with a degree of overlap not less than a first set threshold is selected from the multiple objects to be verified: cat a, cat b, cat c, cat d, object i, wherein at least one discarded object: cat a, cat b, cat c, cat d, object i includes sample objects: cat a, cat b, cat c, cat d. The examples here are only used to clearly illustrate this solution and do not limit the multiple objects to be verified, at least one discarded object, and the method of obtaining multiple objects to be verified and at least one discarded object in this solution.
[0125] Step 1502: Place the at least one discarded object into a discarded object queue, and obtain the sum of the number of the at least one discarded object and the number of historical discarded objects in the discarded object queue, where the historical discarded objects are discarded objects obtained in each training round before the current training round.
[0126] Based on the above example Figure 16 , put at least one discarded object: cat a, cat b, cat c, cat d, object i into the discarded object queue, and determine after putting at least one discarded object, the sum of the number of at least one discarded object in the discarded object queue and the number of each historical discarded object in the discarded object queue.
[0127] Step 1503: If the sum of the quantities is not less than the quantity threshold, the current prediction accuracy of the first detection model is determined based on the degree of overlap of the at least one discarded object and the degree of overlap of each historical discarded object. The prediction accuracy represents: the probability that the at least one discarded object and each historical discarded object each belong to a sample object.
[0128] In order to explain this plan more clearly, based on Figure 16 ,Here is a simplified flowchart of the detection model training process, such as Figure 17As shown, during the training of the detection model, during the Lth iteration, a sample image is selected from the sample image set for two image enhancements to obtain a first reference image and a second reference image. The first detection model predicts multiple objects to be verified for the first reference image, and the second detection model predicts multiple predicted objects for the second reference image. According to the degree of overlap between the multiple objects to be verified and the sample objects in the sample image, at least one discarded object with an overlap degree not less than a first set threshold and at least one candidate object with an overlap degree less than the first set threshold are screened out from the multiple objects to be verified, the at least one discarded object is put into a discarded object queue, and the number of at least one discarded object and the number of candidates for each historical image are determined. The sum of the number n of historical discarded objects (discarded objects obtained in the previous L-1 iterations), if the sum is not less than the quantity threshold, then based on the degree of overlap of at least one discarded object and the degree of overlap of each historical discarded object, the current prediction accuracy of the first detection model is determined, and according to the prediction accuracy, the initial confidence of at least one candidate object can be corrected, and the corrected confidence of at least one candidate object can be obtained. At least one target object whose corrected confidence is greater than the confidence threshold is determined, and the first detection model and the second detection model are adjusted according to the loss values of the at least one target object and the sample object in the first reference image and the multiple predicted objects predicted by the second detection model.
[0129] In an embodiment of the present application, a method for determining the current prediction accuracy of the first detection model can be: taking the degree of overlap of at least one discarded object and the average value of the degrees of overlap of each historical discarded object. For example, the degree of overlap of discarded object 1 included in at least one discarded object is complete overlap, and the degree of overlap of discarded object 2 is such that the overlapping area with the corresponding sample object accounts for eight-tenths of the image area of discarded object 2. The degree of overlap of historical discarded object 1 included in each historical discarded object is complete overlap, and the degree of overlap of historical discarded object 2 is such that the overlapping area with the corresponding sample object accounts for eight-tenths of the image area of historical discarded object 2. Then, the average value of the degree of overlap of at least one discarded object and the degree of overlap of each historical discarded object is nine-tenths, and the current prediction accuracy of the first detection model can be 90%.
[0130] In an embodiment of the present application, another method for determining prediction accuracy is provided. In step 1503, determining the current prediction accuracy of the first detection model based on the overlap degree of the at least one discarded object and the overlap degree of each historically discarded object includes:
[0131] For multiple preset value ranges, perform the following operations:
[0132] Determining a first number of discarded objects having initial confidence levels belonging to a value range from the discarded object queue;
[0133] determining, from the discarded objects, a second number of discarded objects whose overlap exceeds a second set threshold, where the second set threshold is greater than the first set threshold;
[0134] Using the ratio of the second number to the first number as the value accuracy of the one value range;
[0135] The prediction accuracy is obtained based on the value accuracy rates corresponding to each of the multiple value ranges.
[0136] In the embodiment of the present application, based on the above Figure 17 , the discarded object queue contains at least one discarded object: discarded object 1 (whose initial confidence is 90 and the overlap degree is 80%), discarded object 2 (whose initial confidence is 80 and the overlap degree is 79%), discarded object 3 (whose initial confidence is 20 and the overlap degree is 30%), as well as historical discarded object 1 (whose initial confidence is 90 and the overlap degree is 80%), historical discarded object 2 (whose initial confidence is 80 and the overlap degree is 78%), historical discarded object 3 (whose initial confidence is 90 and the overlap degree is 81%) and historical discarded object 4 (whose initial confidence is 4%). 0, the overlap degree is 35%), and the value range of the initial confidence can be divided into four categories: [0, 25), [25, 50), [50, 75), and [75, 100]. Then, from the discarded object queue, the first number of discarded objects belonging to a value range with the initial confidence is determined; that is, the number of discarded objects belonging to the value range [0, 25) is 1, the number of discarded objects belonging to the value range [25, 50) is 1, the number of discarded objects belonging to the value range [50, 75) is 0, and the number of discarded objects belonging to the value range [75, 100] is 5;
[0137] If the overlap threshold is 79%, the second number of discarded objects whose overlap exceeds the second set threshold is determined from each discarded object in each value range. That is, if the number of discarded objects whose overlap exceeds the second set threshold among the discarded objects in the value range [0, 25) is 0, then the value accuracy corresponding to the value range [0, 25) can be 0, or a default value greater than 0, or the value accuracy of the value range in the previous round of iterative training. If the number of discarded objects whose overlap exceeds the second set threshold among the discarded objects in the value range [25, 50) is 0, then the value accuracy corresponding to the value range [25, 50) can be 0, or a default value greater than 0, or the value accuracy of the value range in the previous round of iterative training. For each discarded object in the value range [50, 75), the number of discarded objects whose overlap exceeds the second set threshold is 0. The accuracy rate corresponding to the value range [50, 75] can be 0, or a default value greater than 0, or the accuracy rate of the value range in the previous round of iterative training. For each discarded object in the value range [75, 100], the number of discarded objects whose overlap exceeds the second set threshold is 3. The accuracy rate corresponding to the value range [50, 75] can be 60%.
[0138] Correspondingly, the prediction accuracy can be 0%, 0%, 0%, 60%, or the default value, default value, default value, 60%, or the accuracy of the value range in the previous round, the accuracy of the value range in the previous round, the accuracy of the value range in the previous round, and 60%. It should be noted that the initial confidence level, overlap level, value range, accuracy, and discarded object queue are only used to clarify this solution and do not limit its specific implementation.
[0139] Based on the above Figure 15 The training method process of the detection model in the embodiment of the present application also provides a training method for the detection model, and the method also includes: if the sum of the quantities is greater than the quantity threshold, then based on the sum of the quantities, deleting the corresponding number of historical discarded objects that were first put in from the discarded object queue, so that the number of each discarded object in the discarded object queue is equal to the quantity threshold.
[0140] In the embodiment of this application, Figure 18As shown, the discarded object queue only stores a threshold number m=n+1 discarded objects, and at least one discarded object includes discarded object 1 and discarded object 2. When at least one discarded object is placed in the discarded object queue, the historical discarded object n (the corresponding number of historical discarded objects placed earliest) is deleted from the discarded object queue, so that the discarded object queue always contains the most recently placed discarded object, and the number of discarded objects included is m.
[0141] Based on the training methods of the above-mentioned detection models, if the current prediction accuracy of the first detection model includes multiple value accuracy rates, each value accuracy rate corresponds to a preset value range;
[0142] Then in Figure 2 、 Figure 9 and Figure 15 The steps involved in the above step include correcting at least one initial confidence level obtained according to the current prediction accuracy of the first detection model to obtain a corresponding corrected confidence level, including:
[0143] Determining the accuracy of the values corresponding to the initial confidence levels of the at least one candidate object according to the value ranges to which the initial confidence levels of the at least one candidate object belong respectively;
[0144] According to the value accuracy corresponding to each of the initial confidence levels, each of the initial confidence levels is corrected to obtain a corresponding corrected confidence level.
[0145] In an embodiment of the present application, if the value accuracy corresponding to the value range [0,25) is 20%, the value accuracy corresponding to the value range [25,50) is 60%, the value accuracy corresponding to the value range [50,75) is 80%, and the value accuracy corresponding to the value range [75,100] is 50%, then if the initial confidence of candidate object 1 is 60, the corrected confidence of candidate object 1 is 60*80%, if the initial confidence of candidate object 2 is 80, the corrected confidence of candidate object 1 is 80*50%, if the initial confidence of candidate object 3 is 200, the corrected confidence of candidate object 1 is 20*20%, if the initial confidence of candidate object 4 is 40, the corrected confidence of candidate object 1 is 40*60%, if the initial confidence of candidate object 2. Alternatively, if the initial confidence of candidate 1 is 60, the corrected confidence of candidate 1 is 60*80%*0.7+0.1; if the initial confidence of candidate 2 is 80, the corrected confidence of candidate 1 is 80*50%*0.7+0.1; if the initial confidence of candidate 3 is 200, the corrected confidence of candidate 1 is 20*20%*0.7+0.1; if the initial confidence of candidate 4 is 40, the corrected confidence of candidate 1 is 40*60%*0.7+0.1; if the initial confidence of candidate 2 is 40. It should be noted that the specific method of correction is not limited here and can be set as needed.
[0146] Based on the training methods of the above detection models, the initial models of the first detection model and the second detection model are the same;
[0147] Then in Figure 2 、 Figure 9 and Figure 15 The method comprises: adjusting parameters of the first detection model and the second detection model based on the loss values obtained for the multiple predicted objects, the at least one target object, and the sample objects in the first reference image, respectively.
[0148] Obtaining a loss value based on differences between the multiple predicted objects and the at least one target object and a sample object in the first reference image;
[0149] adjusting a second model parameter of the second detection model according to the loss value;
[0150] The first model parameters of the first detection model are adjusted according to the adjusted second model parameters and the set coefficients.
[0151] In the embodiment of the present application, the model parameters of the first detection model can be updated according to the exponential moving average (EMA) method, and the second model parameters of the second detection model can be adjusted according to the loss value to obtain the adjusted second model parameters of this round (i-th round) of iterative training. If the first model parameter of the first detection model that has not been adjusted is (i-1th round of iterative training), the first model parameters of the adjusted first detection model are: Among them, α can be the coefficient of the exponential sliding average and can take any value between 0 and 1. In addition, it should be noted here that the first detection model and the second detection model obtained by the training method of the detection model here can both be used in production, and the second detection model obtained by the training method of any of the above detection models can both be used in production, and the first detection model or the second detection model used for production can both update the prediction accuracy of the detection model online, wherein the second detection model can also include a discarded object queue, based on which the prediction accuracy is obtained and applied in the same way as in the first detection model.
[0152] Based on the same concept, the present application embodiment provides a training device for a detection model, such as Figure 19 As shown, the apparatus includes: an image enhancement unit 1901, a first detection unit 1902, a correction unit 1903, a second detection unit 1904, and an updating unit 1905. The training apparatus is configured to iteratively train the first detection model and the second detection model based on a set of sample images, each sample image comprising: a plurality of labeled sample objects; wherein, during one iteration, the following operations are performed:
[0153] Based on the image enhancement unit 1901, the selected sample image is subjected to two image enhancement processes to obtain a corresponding first reference image and a second reference image respectively;
[0154] Based on the first detection unit 1902, using the first detection model, obtain at least one candidate object other than the sample object in the first reference image and their respective initial confidence scores;
[0155] Based on the correction unit 1903, according to the current prediction accuracy of the first detection model, the at least one obtained initial confidence is corrected to obtain a corresponding corrected confidence, and at least one target object whose corrected confidence is greater than a confidence threshold is selected from the at least one candidate object;
[0156] Based on the second detection unit 1904, multiple predicted objects in the second reference image are obtained through the second detection model, and through the update unit 1905, the loss values obtained based on the multiple predicted objects, the at least one target object and the sample objects in the first reference image are used to adjust the parameters of the first detection model and the second detection model respectively.
[0157] Optionally, the first detection unit 1902 is specifically configured to:
[0158] Predicting, by using the first detection model, a plurality of objects to be verified contained in the first reference image, wherein the plurality of objects to be verified include sample objects in the first reference image;
[0159] respectively obtaining a degree of overlap between the plurality of objects to be verified and at least one sample object in the first reference image;
[0160] From the multiple objects to be verified, objects to be verified whose overlap degree is less than a first set threshold are selected as candidate objects.
[0161] Optionally, the first detection unit 1902 is further configured to:
[0162] Selecting at least one discarded object whose overlap degree is not less than the first set threshold from the multiple objects to be verified, wherein the at least one discarded object includes a sample object in the first reference image;
[0163] Putting the at least one discarded object into a discarded object queue, and obtaining the sum of the number of the at least one discarded object and the number of historical discarded objects in the discarded object queue, where the historical discarded objects are discarded objects obtained in each round of training before the current round of training;
[0164] If the sum of the quantities is not less than the quantity threshold, the current prediction accuracy of the first detection model is determined based on the degree of overlap of the at least one discarded object and the degree of overlap of each historical discarded object, and the prediction accuracy represents: the probability that the at least one discarded object and each historical discarded object each belong to a sample object.
[0165] Optionally, the first detection unit 1902 is specifically configured to:
[0166] For multiple preset value ranges, perform the following operations:
[0167] Determining a first number of discarded objects having initial confidence levels belonging to a value range from the discarded object queue;
[0168] determining, from the discarded objects, a second number of discarded objects whose overlap exceeds a second set threshold, where the second set threshold is greater than the first set threshold;
[0169] Using the ratio of the second number to the first number as the value accuracy of the one value range;
[0170] The prediction accuracy is obtained based on the value accuracy rates corresponding to each of the multiple value ranges.
[0171] Optionally, the first detection unit 1902 is further configured to:
[0172] If the sum of the quantities is greater than the quantity threshold, based on the sum of the quantities, the earliest placed historical discarded objects of the corresponding quantity are deleted from the discarded object queue, so that the quantity of each discarded object in the discarded object queue is equal to the quantity threshold.
[0173] Optionally, the current prediction accuracy of the first detection model includes multiple value accuracy rates, each value accuracy rate corresponds to a preset value range;
[0174] The correction unit 1903 is specifically configured to:
[0175] Determining the accuracy of the values corresponding to the initial confidence levels of the at least one candidate object according to the value ranges to which the initial confidence levels of the at least one candidate object belong respectively;
[0176] According to the value accuracy corresponding to each of the initial confidence levels, each of the initial confidence levels is corrected to obtain a corresponding corrected confidence level.
[0177] Optionally, the initial models of the first detection model and the second detection model are the same;
[0178] The updating unit 1905 is specifically configured to obtain a loss value based on differences between the multiple predicted objects, the at least one target object, and the sample object in the first reference image;
[0179] adjusting a second model parameter of the second detection model according to the loss value;
[0180] The first model parameters of the first detection model are adjusted according to the adjusted second model parameters and the set coefficients.
[0181] Based on the same inventive concept as the above method embodiment, the present application embodiment also provides a computer device. In one embodiment, the computer device can be a server, such as Figure 1 In this embodiment, the structure of the computer device can be as follows: Figure 20As shown, it includes a memory 2001 , a communication module 2003 and one or more processors 2002 .
[0182] Memory 2001 is used to store computer programs executed by processor 2002. Memory 2001 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and programs required for running instant messaging functions, while the data storage area may store various instant messaging messages and operating instruction sets.
[0183] Memory 2001 may be a volatile memory, such as random-access memory (RAM); a non-volatile memory, such as read-only memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or any other medium capable of carrying or storing a desired computer program in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 2001 may be a combination of the aforementioned memories.
[0184] The processor 2002 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 2002 is configured to implement the above network acceleration method when calling the computer program stored in the memory 2001 .
[0185] The communication module 2003 is used to communicate with terminal devices and other servers.
[0186] The specific connection medium between the memory 2001, the communication module 2003 and the processor 2002 is not limited in the embodiment of the present application. Figure 20 In the embodiment, the memory 2001 and the processor 2002 are connected via a bus 2004. The bus 2004 is connected to the processor 2002 via a bus 2004. Figure 20 The connections between the other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The bus 2004 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 20 The diagram shows a single thick line, but this does not indicate that there is only one bus or one type of bus.
[0187] The computer storage medium stored in the memory 2001 stores computer executable instructions for implementing the multimedia information recommendation method of the embodiments of the present application. The processor 2002 is configured to execute the training method of the detection model as described above. Figure 2 or Figure 9 or Figure 15 as shown.
[0188] In another embodiment, the computer device can also be other computer devices, such as the terminal device 110 as shown. Figure 1 In this embodiment, the structure of the computer device can be as shown in Figure 21 , which includes a communication component 2110, a memory 2120, a display unit 2130, a camera 2140, a sensor 2150, an audio circuit 2160, a Bluetooth module 2170, a processor 2180, and the like.
[0189] The communication component 2110 is configured to communicate with the server. In some embodiments, a wireless fidelity (WiFi) module can be included, which belongs to a short-range wireless transmission technology. The computer device can help the user to send and receive information through the WiFi module.
[0190] The memory 2120 can be used to store software programs and data. The processor 2180 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 2120. The memory 2120 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. The memory 2120 stores an operating system that enables the terminal device 110 to run. In the present application, the memory 2120 can store an operating system and various application programs, and can also store a computer program for executing the training method of the detection model in the embodiments of the present application.
[0191] The display unit 2130 can also be used to display information input by the user or information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 2130 can include a display screen 2132 arranged on the front of the terminal device 110. The display screen 2132 can be configured in the form of a liquid crystal display, a light-emitting diode, and the like. The display unit 2130 can be used to display the user interface of the detection model detection process in the embodiments of the present application, and the like.
[0192] The display unit 2130 can also be used to receive inputted digital or character information, generate signal input related to user settings and function control of the terminal device 110, and specifically, the display unit 2130 can include a touch screen 2131 arranged on the front of the terminal device 110, which can collect touch operations of a user thereon or therearound, such as clicking a button, dragging a scroll box, and the like.
[0193] The touch screen 2131 can be overlaid on the display screen 2132, or the touch screen 2131 can be integrated with the display screen 2132 to realize the input and output functions of the terminal device 110, and after integration, the touch screen 2131 and the display screen 2132 can be referred to as a touch display screen. The display unit 2130 in the present application can display application programs and corresponding operation steps.
[0194] The camera 2140 can be used to capture still images, and a user can post comments on images captured by the camera 2140 through an application. The camera 2140 can be one or multiple. An object generates an optical image through a lens and projects the optical image onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 2180 to convert it into a digital image signal.
[0195] The terminal device can also include at least one sensor 2150, such as an acceleration sensor 2151, a distance sensor 2152, a fingerprint sensor 2153, and a temperature sensor 2154. The terminal device can also be configured with a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, a motion sensor, and other sensors.
[0196] The audio circuit 2160, the speaker 2161, and the microphone 2162 can provide an audio interface between a user and the terminal device 110. The audio circuit 2160 can convert received audio data into an electrical signal and transmit the electrical signal to the speaker 2161, which converts the electrical signal into a sound signal for output. The terminal device 110 can also be configured with a volume button for adjusting the volume of the sound signal. On the other hand, the microphone 2162 converts a sound signal collected into an electrical signal, which is received by the audio circuit 2160 and converted into audio data, which is then output to the communication component 2110 for transmission to another terminal device 110, for example, or to the memory 2120 for further processing.
[0197] The Bluetooth module 2170 is used to exchange information with other Bluetooth devices having a Bluetooth module through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable computer device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 2170 to exchange data.
[0198] The processor 2180 is the control center of the terminal device, which uses various interfaces and lines to connect various parts of the entire terminal. It performs various functions of the terminal device and processes data by running or executing software programs stored in the memory 2120 and calling data stored in the memory 2120. In some embodiments, the processor 2180 may include one or more processing units; the processor 2180 may also integrate an application processor and a baseband processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the baseband processor mainly processes wireless communications. It is understandable that the above-mentioned baseband processor may not be integrated into the processor 2180. In the present application, the processor 2180 can run the operating system, application programs, user interface display and touch response, as well as the training method of the detection model of the embodiment of the present application. In addition, the processor 2180 is coupled to the display unit 2130.
[0199] In some possible implementations, various aspects of the detection model training method provided in the present application may also be implemented in the form of a program product, which includes a computer program. When the program product is run on a computer device, the computer program is used to enable the computer device to execute the steps of the detection model training method according to various exemplary embodiments of the present application described above in this specification. For example, the computer device may execute the following steps: Figure 2 or Figure 3 or Figure 9 Follow the steps shown in .
[0200] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0201] The program product of the embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include a computer program, and can be run on a computer device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with a command execution system, apparatus, or device.
[0202] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with a command execution system, apparatus, or device.
[0203] The computer program embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0204] The computer program for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The computer program can be executed entirely on the user's computer device, partially on the user's computer device, as a stand-alone software package, partially on the user's computer device and partially on a remote computer device, or entirely on a remote computer device or server. In the case of a remote computer device, the remote computer device can be connected to the user's computer device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer device (e.g., through the Internet using an Internet service provider).
[0205] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.
[0206] In addition, although the operation of the present application method is described in a particular order in the accompanying drawings, this does not require or imply that these operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be merged into one step and executed, and / or a step can be decomposed into multiple steps and executed. Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0207] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0208] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0209] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0210] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A method for training a detection model, characterized in that: The method comprises: Iteratively train the first detection model and the second detection model based on a set of sample images, each sample image comprising: a plurality of labeled sample objects; wherein, during one iteration, the following operations are performed: Performing image enhancement processing twice on the selected sample image to obtain a corresponding first reference image and a second reference image respectively; Obtaining, by means of the first detection model, at least one candidate object other than the sample object in the first reference image and their respective initial confidence scores; Correcting the at least one initial confidence score obtained based on the current prediction accuracy of the first detection model to obtain a corresponding corrected confidence score, and selecting at least one target object having a corrected confidence score greater than a confidence threshold from the at least one candidate object; Obtaining, by means of the second detection model, a plurality of predicted objects in the second reference image, and adjusting parameters of the first detection model and the second detection model respectively using loss values obtained based on the plurality of predicted objects, the at least one target object, and sample objects in the first reference image; The current prediction accuracy of the first detection model includes multiple value accuracy rates, each value accuracy rate corresponds to a preset value range; Then, according to the current prediction accuracy of the first detection model, the at least one initial confidence degree obtained is corrected to obtain a corresponding corrected confidence degree, including: determining a value accuracy rate corresponding to each initial confidence degree according to the value range to which the initial confidence degree of the at least one candidate object belongs; each value accuracy rate corresponds to a preset value range; According to the value accuracy corresponding to each of the initial confidence levels, each of the initial confidence levels is corrected to obtain a corresponding corrected confidence level.
2. The method according to claim 1, wherein The obtaining, by using the first detection model, at least one candidate object other than the sample object in the first reference image includes: Predicting, by using the first detection model, a plurality of objects to be verified contained in the first reference image, wherein the plurality of objects to be verified include sample objects in the first reference image; respectively obtaining a degree of overlap between the plurality of objects to be verified and at least one sample object in the first reference image; From the multiple objects to be verified, objects to be verified whose overlap degree is less than a first set threshold are selected as candidate objects.
3. The method according to claim 2, wherein The method further comprises: Selecting at least one discarded object whose overlap degree is not less than the first set threshold from the multiple objects to be verified, wherein the at least one discarded object includes a sample object in the first reference image; Putting the at least one discarded object into a discarded object queue, and obtaining the sum of the number of the at least one discarded object and the number of historical discarded objects in the discarded object queue, where the historical discarded objects are discarded objects obtained in each round of training before the current round of training; If the sum of the quantities is not less than the quantity threshold, the current prediction accuracy of the first detection model is determined based on the degree of overlap of the at least one discarded object and the degree of overlap of each historical discarded object, and the prediction accuracy represents: the probability that the at least one discarded object and each historical discarded object each belong to a sample object.
4. The method according to claim 3, wherein The determining, based on the overlap degree of the at least one discarded object and the overlap degree of each historically discarded object, a current prediction accuracy of the first detection model includes: For multiple preset value ranges, perform the following operations: Determining a first number of discarded objects having initial confidence levels belonging to a value range from the discarded object queue; determining, from the discarded objects, a second number of discarded objects whose overlap exceeds a second set threshold, where the second set threshold is greater than the first set threshold; Using the ratio of the second number to the first number as the value accuracy of the one value range; The prediction accuracy is obtained based on the value accuracy rates corresponding to each of the multiple value ranges.
5. The method according to claim 3, wherein The method further comprises: If the sum of the quantities is greater than the quantity threshold, based on the sum of the quantities, the earliest placed historical discarded objects of the corresponding quantity are deleted from the discarded object queue, so that the quantity of each discarded object in the discarded object queue is equal to the quantity threshold.
6. The method according to any one of claims 1 to 5, wherein: The initial models of the first detection model and the second detection model are the same; Then, based on the loss values obtained for the multiple predicted objects, the at least one target object, and the sample objects in the first reference image, respectively adjusting the parameters of the first detection model and the second detection model, the method includes: Obtaining a loss value based on differences between the multiple predicted objects and the at least one target object and a sample object in the first reference image; adjusting a second model parameter of the second detection model according to the loss value; The first model parameters of the first detection model are adjusted according to the adjusted second model parameters and the set coefficients.
7. A training device for a detection model, characterized in that: The apparatus comprises: an image enhancement unit, a first detection unit, a correction unit, a second detection unit, and an update unit. The training apparatus is configured to iteratively train the first detection model and the second detection model based on a set of sample images, each sample image comprising: a plurality of labeled sample objects; wherein, during one iteration, the following operations are performed: Based on the image enhancement unit, performing image enhancement processing twice on the selected sample image to obtain a corresponding first reference image and a second reference image respectively; Obtaining, based on the first detection unit and using the first detection model, at least one candidate object other than the sample object in the first reference image and their respective initial confidence scores; Based on the correction unit, according to the current prediction accuracy of the first detection model, the at least one initial confidence obtained is corrected respectively to obtain a corresponding corrected confidence, and at least one target object whose corrected confidence is greater than the confidence threshold is selected from the at least one candidate object; wherein the current prediction accuracy of the first detection model includes multiple value accuracy rates, and each value accuracy rate corresponds to a preset value range; then, based on the correction unit, according to the current prediction accuracy of the first detection model, the at least one initial confidence obtained is corrected respectively to obtain a corresponding corrected confidence, including: determining the value accuracy rate corresponding to each initial confidence rate according to the value range to which the initial confidence rate of each candidate object belongs; each value accuracy rate corresponds to a preset value range; correcting each initial confidence rate according to the value accuracy rate corresponding to each initial confidence rate to obtain a corresponding corrected confidence; Based on the second detection unit, multiple predicted objects in the second reference image are obtained through the second detection model, and through the update unit, the loss values obtained based on the multiple predicted objects, the at least one target object and the sample objects in the first reference image are used to adjust the parameters of the first detection model and the second detection model respectively.
8. A computer-readable non-volatile storage medium, characterized in that: The computer-readable non-volatile storage medium stores a program, and when the program is executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 6.
9. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to call a computer program stored in the memory, and execute the method according to any one of claims 1 to 6 according to the obtained program.
10. A computer program product, characterized in that The method comprises a computer program stored in a computer-readable storage medium; when a processor of a computer device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the computer device performs the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Target detection model updating method and server
CN112766487A
Detection model training method, vehicle damage detection method and terminal equipment
CN114612744A