Target detection model training and target detection method and device, equipment and medium
By automatically adjusting network parameters based on performance consumption and detection accuracy during the target detection model training process, the problem of repeated trials in model training is solved, achieving a balance between speed and accuracy, and improving the efficiency and accuracy of model training.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2022-09-08
- Publication Date
- 2026-07-21
AI Technical Summary
Existing object detection models require repeated trials of different backbone networks and image types during training, resulting in wasted resources, high training costs, and poor model accuracy.
After training the target detection model in the current stage, the network parameters for the next stage are determined based on performance consumption and detection accuracy. An automatic search mechanism is used for training, and adjustments are made using a small amount of training data to obtain a converged model.
It shortens the time developers spend tuning parameters, achieves a balance between speed and accuracy, avoids wasting resources, and improves the efficiency and accuracy of model training.
Smart Images

Figure CN116152589B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of target detection technology, and in particular to a target detection model training and target detection method, apparatus, device, and medium. Background Technology
[0002] Multimedia content labeling platforms combine machine labeling and manual labeling. In the machine labeling process, target detection plays a crucial role in both automated labeling and assisting manual labeling.
[0003] In existing approaches, the backbone network of an object detection model and the types of images it is suitable for are typically selected. Then, the model is trained and updated using a training dataset. However, if this trained model proves unsuitable, the backbone network and the types of images it is suitable for are manually selected again, and the model is retrained using the training dataset. This undoubtedly leads to a waste of training resources, high training costs, and poor model accuracy. Summary of the Invention
[0004] This disclosure provides a target detection model training and target detection method, apparatus, device, and medium to solve the problem of repeated attempts required for model training, and enables rapid model training while balancing speed and accuracy.
[0005] In a first aspect, embodiments of this disclosure provide a target detection model training method, the training method comprising:
[0006] The first training task is performed on the object detection model at the current stage;
[0007] Based on the performance consumption and detection accuracy of the target detection model trained in the current stage, determine the network parameters to be used for the first training task of the target detection model in the next stage;
[0008] Based on the network parameters used in the first training task, a second training task is performed on the target detection model to obtain a converged target detection model;
[0009] The amount of training data for the first training task is less than the amount of training data for the second training task.
[0010] Secondly, this disclosure also provides a target detection method, the target detection method comprising:
[0011] Identify the multimedia content to be processed;
[0012] The multimedia content to be processed is input into the target detection model to obtain the target detection result of the multimedia content to be processed.
[0013] Thirdly, embodiments of this disclosure also provide a target detection model training device, the training device comprising:
[0014] The first training task execution module is used to execute the first training task on the object detection model at the current stage;
[0015] The network parameter determination module is used to determine the network parameters to be used for the first training task of the target detection model in the next stage, based on the performance consumption and detection accuracy of the target detection model trained in the current stage.
[0016] The target detection model determination module is used to perform a second training task on the target detection model based on the network parameters used in the first training task, so as to obtain a converged target detection model; wherein the amount of training data in the first training task is less than the amount of training data in the second training task.
[0017] Fourthly, embodiments of this disclosure also provide a target detection device, the target detection device comprising:
[0018] The multimedia content determination module is used to determine the multimedia content to be processed.
[0019] The target detection module is used to input the multimedia content to be processed into the target detection model to obtain the target detection result of the multimedia content to be processed.
[0020] Fifthly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0021] At least one processor; and
[0022] A memory communicatively connected to the at least one processor; wherein,
[0023] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the target detection model training method or target detection method as described in any of the above embodiments.
[0024] Sixthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a processor to execute and implement the target detection model training method or target detection method described in any of the above embodiments.
[0025] This disclosure provides a method for training an object detection model. The method involves performing a first training task on the object detection model in the current stage; determining the network parameters to be used for the next stage of training based on the performance consumption and detection accuracy of the model obtained in the current stage; and performing a second training task on the object detection model based on the network parameters used in the first training task to obtain a converged object detection model. This technical solution introduces an automatic search mechanism, solving the problem of repeated trials during model training and shortening the time required for parameter tuning by developers. By searching for network parameters that balance speed and accuracy, the method achieves faster speed while maintaining the accuracy of the object detection model, avoiding resource waste.
[0026] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0027] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0028] Figure 1 This is a schematic diagram of a target detection model training method provided in an embodiment of the present disclosure;
[0029] Figure 2 This is a schematic flowchart of an existing model training method provided in an embodiment of this disclosure;
[0030] Figure 3 This is a flowchart illustrating another object detection model training method provided in this embodiment of the disclosure;
[0031] Figure 4 This is a flowchart illustrating another object detection model training method provided in this embodiment of the present disclosure;
[0032] Figure 5 This is a schematic diagram of a target detection method provided in an embodiment of this disclosure;
[0033] Figure 6 This is a schematic diagram of the structure of a target detection model training device provided in an embodiment of the present disclosure;
[0034] Figure 7 This is a schematic diagram of the structure of a target detection device provided in an embodiment of the present disclosure;
[0035] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0036] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0037] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0038] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0039] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0040] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0041] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0042] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0043] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0044] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0045] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0046] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0047] Figure 1 This is a schematic flowchart of a target detection model training method provided in an embodiment of this disclosure. This embodiment is applicable to training a target detection model for processing multimedia content. The method can be executed by a target detection model training device, which can be implemented in software and / or hardware and is generally integrated into any electronic device with network communication capabilities. Such electronic devices include, but are not limited to, mobile terminals, PCs, or servers. Figure 1 As shown, the target detection model training method includes steps S110-S130:
[0048] S110. Perform the first training task on the object detection model at the current stage.
[0049] See Figure 2 Commonly used object detection models typically have manually defined network parameters. Due to limited computing resources in the deployment environment, algorithm developers often use a limited number of feature extraction network layers and parameters such as image input resolution. However, this parameter-defined training method often fails to achieve the optimal model accuracy under limited performance overhead.
[0050] See Figure 3The optimal network parameters can be obtained through the parameter search module, and then used to train the object detection model. For example, a training task can be performed on the object detection model at the current time to obtain the performance consumption and detection accuracy of the object detection model at the current time; the optimal network parameters can be obtained through the parameter search module, and the object detection model can be trained on the entire dataset to obtain the object detection model with the optimal detection accuracy and performance consumption.
[0051] The current stage can refer to the current moment when the object detection model is being trained. The current stage includes, but is not limited to, the stage of training the object detection model for the first time, and the stage of training the object detection model for subsequent times.
[0052] As an optional but non-limiting implementation, the first training task for the object detection model at this stage includes:
[0053] When the current stage is the first training stage, the network parameters are loaded and initialized in the object detection model, and the first training task is performed on the object detection model with the initialized network parameters loaded.
[0054] In the first training phase, it is necessary to load and initialize network parameters to initialize the object detection model. For example, load initial network parameters, such as the feature extraction network Darknet18 and a resolution of 640×640, into the object detection model, and then perform the first training task on the object detection model with the initialized network parameters loaded.
[0055] As an optional but non-limiting implementation, the first training task for the object detection model is performed at this stage, including steps A1-A2:
[0056] Step A1: If the current stage is not the first training stage, load the network parameters that were pre-searched in the previous stage and used to perform the first training task on the object detection model in the current stage into the object detection model.
[0057] Step A2: Perform the first training task on the object detection model loaded with the network parameters pre-searched in the previous stage.
[0058] If the current stage is not the first training stage, the network parameters required at the current moment can be determined based on the training results of the target detection model after performing the first training task in the previous stage. For example, if the training results of the network parameters used in the previous stage are poor, the network parameters need to be adjusted; then the network parameters pre-searched in the previous stage are loaded into the target detection model, and the first training task is performed on the target detection model loaded with the network parameters pre-searched in the previous stage.
[0059] As an optional but non-limiting implementation, the first training task for the object detection model is performed at this stage, including steps B1-B2:
[0060] Step B1: Collect data from the second training dataset used in the second training task to obtain the first training dataset; wherein the similarity between the label type distribution of the training samples in the first training dataset and the label type distribution of the training samples in the second training dataset is greater than a preset similarity threshold.
[0061] Step B2: Perform the first training task on the object detection model at the current stage using the first training dataset.
[0062] Here, "downsampling" can refer to randomly sampling the original full dataset to reduce its size. In this embodiment, the second training dataset used for the second training task is randomly sampled to obtain a smaller dataset. For example, 10 out of 100 images in the second training dataset are randomly sampled and selected as the first training dataset to improve data search speed.
[0063] To improve the learning efficiency of the object detection model, the similarity between the label type distribution of training samples in the first training dataset and the label type distribution of training samples in the second training dataset should be high. For example, if the label types of training samples in the first training dataset include 40% Class A labels, 30% Class B labels, and 30% Class C labels, then the label types of training samples in the second training dataset should also include Class A, Class B, and Class C labels, and the distribution of these label types should be highly similar to that of the first training dataset; for example, a preset label type distribution threshold should be met. For example, the label types of training samples in the second training dataset may include 38% Class A labels, 30% Class B labels, and 32% Class C labels; or the label type distribution of training samples in the second training dataset should be consistent with that of the first training dataset, for example, the label types of training samples in the second training dataset may also include 40% Class A labels, 30% Class B labels, and 30% Class C labels.
[0064] S120. Based on the performance consumption and detection accuracy of the target detection model trained in the current stage, determine the network parameters to be used for the first training task of the target detection model in the next stage.
[0065] Performance consumption refers to the energy consumed during model training, and this performance consumption is related to the training speed of the model in the object detection model. Performance consumption includes, but is not limited to, GFLOPs, which can refer to 1 billion floating-point operations per second and can be used to measure the complexity of the algorithm and / or model.
[0066] Detection accuracy can be a parameter that measures the precision of target detection. Higher detection accuracy generally indicates higher accuracy of the target detection model. However, excessively high detection accuracy can increase the training complexity of the target detection model, leading to a waste of model prediction resources.
[0067] Based on the performance consumption and detection accuracy of the target detection model at the current stage, it is determined whether the target detection model at the current stage is the optimal model, and the performance consumption and detection accuracy are appropriately adjusted to obtain the optimal target detection model that balances detection accuracy and detection speed. For example, if the performance consumption of the target detection model is high at the current moment, resulting in a decrease in the training speed of the target detection model, then network parameters that can improve the training speed of the target detection model can be selected in the next stage. If the detection accuracy of the target detection model is low at the current moment, resulting in low detection accuracy of the target detection model, then network parameters that can improve the detection accuracy of the target detection model can be selected in the next stage.
[0068] S130. Based on the network parameters used in the first training task, perform a second training task on the target detection model to obtain a converged target detection model.
[0069] Specifically, the network parameters used to determine the optimal object detection model are determined through a first training task, and these network parameters are then used to perform a second training task to obtain a converged object detection model. The amount of training data for the first training task is less than the amount of training data for the second training task.
[0070] As an optional but non-limiting implementation, based on the network parameters used in the first training task, a second training task is performed on the object detection model, including steps C1-C2:
[0071] Step C1: Select network parameters that meet the preset search conditions from the network parameters used to perform the first training task.
[0072] Step C2: Based on the network parameters that meet the preset search conditions, perform the second training task on the object detection model.
[0073] Among them, the network parameters with preset search conditions include the network parameters that minimize the computational complexity of the target detection model while keeping the performance consumption of the target detection model less than the preset performance consumption limit and the detection accuracy greater than the preset accuracy value.
[0074] The first training task determines the network parameters of the object detection model that satisfy preset search conditions and minimize computational complexity. These network parameters are then used to perform a second training task on the object detection model. The first training task yields network parameters that balance speed and accuracy, ensuring model accuracy while achieving faster speeds and avoiding resource waste.
[0075] As an optional but non-limiting implementation, a second training task is performed on the object detection model based on the network parameters that meet the preset search conditions, including steps D1-D2:
[0076] Step D1: Load network parameters that meet the preset search conditions into the object detection model.
[0077] Step D2: Using the second training dataset, perform the second training task on the object detection model loaded with network parameters that meet the preset search conditions.
[0078] The second training dataset used to perform the first training task is downsampled to obtain the first training dataset used to perform the first training task.
[0079] The network parameters that meet the preset search conditions determined by the first training task are loaded into the target detection model, and the target detection model is trained using the second training dataset to obtain a converged target detection model.
[0080] This disclosure provides a method for training an object detection model. The method involves performing a first training task on the object detection model in the current stage; determining the network parameters to be used for the next stage of training based on the performance consumption and detection accuracy of the model obtained in the current stage; and performing a second training task on the object detection model based on the network parameters used in the first training task to obtain a converged object detection model. Using the technical solution of this disclosure, in the first training stage, network parameters are loaded and initialized to initialize the object detection model; the first training task determines the network parameters corresponding to when the object detection model satisfies preset search conditions and has the minimum computational complexity; and the second training task is performed on the object detection model using these network parameters. This solves the problem of repeated trials required for model training and enables rapid model training while balancing speed and accuracy.
[0081] Figure 4 This is a flowchart illustrating another object detection model training method provided in this disclosure. The technical solution of this embodiment further optimizes the aforementioned embodiments based on the above embodiments. This embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 4As shown, the target detection model training method of this embodiment may include the following steps S410-S430:
[0082] S410. Perform the first training task on the object detection model at the current stage.
[0083] S420. If it is detected that the performance consumption of the target detection model trained in the current stage is greater than the preset performance consumption limit and / or the detection accuracy is greater than the preset accuracy value, then the network parameters used by the target detection model in the current stage are adjusted first to reduce the complexity of the feature extraction network in the target detection model.
[0084] If the performance consumption exceeds a preset performance consumption limit, exceeding a standard for the stress equipment, the equipment will experience performance consumption limitations. Exceeding these limits may cause equipment lag and prevent effective model training. Therefore, if the performance consumption exceeds the preset limit, the network parameters used in the current stage will be adjusted to improve the running speed.
[0085] If the detection accuracy is greater than the preset accuracy value, the detection accuracy can be appropriately reduced. Under the condition that the detection accuracy is not less than the preset accuracy value, the complexity of the feature extraction network in the target detection model can be reduced to make the running efficiency higher.
[0086] As an optional but non-limiting implementation, the network parameters used in the current object detection model are adjusted first, including:
[0087] The search direction is directed towards reducing the complexity of the feature extraction network in the object detection model. The nearest neighbor network parameters that are used in the current stage object detection model are searched from the preset network parameter library to complete the first adjustment of the network parameters used in the current stage object detection model.
[0088] As an optional but non-limiting implementation, the network parameters include image input resolution parameters and feature extraction network parameters that are applicable and supported in the object detection model.
[0089] As an optional but non-limiting implementation, the preset network parameter library includes feature extraction network parameters in at least two dimensions and applicable supported image input resolution parameters in at least two dimensions; the feature extraction network parameters in at least two dimensions are deployed in order of feature extraction network complexity; and the applicable supported image input resolution parameters in at least two dimensions are deployed in order of image resolution.
[0090] Table 1 Preset Network Parameter Library
[0091]
[0092]
[0093] Referring to Table 1, the feature extraction network parameters include, but are not limited to, darknet19, resent34, resen50, darknet53, and convnext-tiny; the image input resolution parameters include, but are not limited to, 384×384, 416×416, 512×512, 640×640, and 1280×1280. Here, P represents the detection accuracy, and F represents the performance cost. Higher feature extraction network complexity leads to higher performance costs, increasing resource consumption and increasing the risk of overfitting; conversely, insufficient feature extraction network complexity results in poor model accuracy. Higher image input resolution leads to higher detection accuracy, but exceeding a preset accuracy value reduces running speed, wasting resources; conversely, insufficient detection accuracy can lead to missed detections.
[0094] As an optional but non-limiting implementation, the nearest neighbor network parameters to the network parameters used by the current target detection model are searched from a pre-defined network parameter library, including:
[0095] Keeping the supported image input resolution parameters of the network parameters used by the target detection model in the current stage unchanged, the system searches for feature extraction network parameters that are close to the network parameters used by the target detection model in the current stage from the preset network parameter library.
[0096] In one optional embodiment of this disclosure, if the performance consumption of the target detection model trained in the current stage is detected to be greater than a preset performance consumption limit, then firstly, the image input resolution parameters supported by the network parameters used in the target detection model in the current stage are kept unchanged. Then, a feature extraction network parameter that is adjacent to the network parameters used in the target detection model in the current stage is searched from a preset network parameter library to reduce the complexity of the feature extraction network in the target detection model. For example, if the image input resolution parameter used in the current stage is 512×512 and the feature extraction network parameter is darknet53, then the image input resolution parameter in the current stage is kept unchanged, and the search direction is towards reducing the complexity of the feature extraction network in the target detection model. A network parameter adjacent to darknet53 is searched from the preset network parameter library, i.e., resen50 is used as the network parameter that can reduce the complexity of the feature extraction network in the target detection model, and a first adjustment is made.
[0097] In another optional embodiment of this disclosure, if the detection accuracy of the target detection model trained in the current stage is detected to be greater than a preset accuracy value, the applicable supported image input resolution parameters in the network parameters used by the target detection model in the current stage are kept unchanged, and a feature extraction network parameter that is close to the network parameter used by the target detection model in the current stage is searched from a preset network parameter library. For example, if the feature extraction network parameter in the current stage is convnext-tiny, a network parameter darknet53 that is close to convnext-tiny and can reduce the complexity of the feature extraction network in the target detection model is searched from the preset network parameter library and adjusted first.
[0098] As an optional but non-limiting implementation, searching for nearest-neighbor network parameters from a pre-defined network parameter library that match the network parameters used by the target detection model in the current stage also includes:
[0099] Keeping the feature extraction network parameters unchanged in the network parameters used by the target detection model at the current stage, the system searches from the preset network parameter library for suitable and supported image input resolution parameters that are close to the network parameters used by the target detection model at the current stage.
[0100] In one optional embodiment of this disclosure, if it is detected that the performance consumption of the target detection model trained in the current stage is greater than a preset performance consumption limit, and the detection accuracy is greater than a preset accuracy value, then first, the image input resolution parameter is kept unchanged, and the feature extraction network parameter is reduced; then, the feature extraction network parameter is kept unchanged, and the image input resolution parameter is reduced. For example, if the image input resolution parameter in the current stage is 640×640, and the feature extraction network parameter is darknet53, first, the image input resolution parameter is kept unchanged, and a network parameter that can reduce the complexity of the feature extraction network is searched; if it is determined that the finally searched feature extraction network parameter is resen34, then the feature extraction network parameter is kept unchanged, and the image input resolution parameter is reduced, so as to reduce the complexity of the feature extraction network in the target detection model and improve the computational efficiency.
[0101] S430. If it is detected that the performance consumption of the target detection model trained in the current stage is less than the preset performance consumption limit and / or the detection accuracy is less than the preset accuracy value, then the network parameters used by the target detection model in the current stage are adjusted for the second time to improve the complexity of the feature extraction network in the target detection model.
[0102] If the detection accuracy is less than a preset accuracy value, the detection accuracy can be improved to make target detection more accurate. If the performance consumption is less than a preset performance consumption limit, the performance consumption can be appropriately increased. Under the condition that the performance consumption is not greater than the preset performance consumption limit, the complexity of the feature extraction network in the target detection model can be increased to make the target detection accuracy higher.
[0103] As an optional but non-limiting implementation, a second adjustment is made to the network parameters used in the current object detection model, including:
[0104] The search direction is directed towards improving the complexity of the feature extraction network in the object detection model. The nearest neighbor network parameters that are used by the current object detection model are searched from the preset network parameter library to complete the second adjustment of the network parameters used by the current object detection model.
[0105] Referring to Table 1, network parameters with higher detection accuracy and higher performance consumption than those currently used can be searched from the preset network parameter library to improve the complexity of the feature extraction network in the target detection model.
[0106] In one optional embodiment of this disclosure, if it is detected that the performance consumption of the target detection model trained in the current stage is less than a preset performance consumption limit and the detection accuracy is less than a preset accuracy value, then a search is conducted in a direction that can improve the complexity of the feature extraction network in the target detection model. This involves searching a preset network parameter library for network parameters with higher detection accuracy and greater performance consumption than the network parameters used in the current stage. For example, if the current stage image input resolution parameter is 640×640 and the feature extraction network parameter is Resent34, first, the image input resolution parameter is kept constant while searching for network parameters that can improve the complexity of the feature extraction network. If the finally searched feature extraction network parameter is determined to be Darknet53, then the feature extraction network parameter is kept constant again while improving the image input resolution parameter. When the detection accuracy does not exceed the preset accuracy value, an image input resolution parameter with higher detection accuracy is selected to improve the complexity of the feature extraction network in the target detection model and increase the target detection accuracy.
[0107] In one optional embodiment of this disclosure, a search table combining image input resolution and feature extraction network parameters is introduced. Under performance constraints, the algorithm automatically determines the search direction to verify the effectiveness of the corresponding positional parameters in the table.
[0108] In another optional embodiment of this disclosure, GFLOPs are used as one of the search metrics. Alternatively, the running speed of the model under TRT can be used as a search metric to guide the search direction. This scheme uses a manually set sampling rate to sample the training set for the search, or it can combine Bayesian optimization to estimate a suitable sampling rate. In addition to searching for feature extraction network parameters and image input resolution parameters, the relationship between detection algorithms such as YOLO and resolution can also be searched; different FPNs can also be searched to achieve the optimal balance between accuracy and speed.
[0109] S440. The network parameters obtained through the first adjustment or the second adjustment are determined as the network parameters used in the next stage to perform the first training task on the object detection model.
[0110] This method involves adjusting the network parameters, either first or second, to obtain a network parameter balance between accuracy and speed. These network parameters are then used for the first training task in the next stage. This addresses the problem of relying too heavily on developers' experience in parameter tuning when manually defining network parameters. Experienced developers can quickly select suitable image resolutions and models with fewer trials, and train the model rapidly with different parameters, achieving good training results. However, for less experienced developers, it requires repeated iterations with different image input resolutions and feature extraction network parameters, and good convergence is not guaranteed. This solution searches for nearest-neighbor network parameters from a pre-defined network parameter library that are used by the current object detection model, and then performs first or second adjustments, resulting in a faster object detection model that balances speed and accuracy.
[0111] S450. Based on the network parameters used in the first training task, perform a second training task on the target detection model to obtain a converged target detection model.
[0112] This disclosure provides a method for training an object detection model. It introduces a preset network parameter library. While keeping the image input resolution parameter constant, it searches the preset network parameter library for feature extraction network parameters that are adjacent to the network parameters used by the object detection model in the current stage. The network parameters are then adjusted to obtain the optimal feature extraction network parameters under the given detection accuracy metric. Keeping the optimal feature extraction network parameters constant, the image input resolution parameter is adjusted to obtain an object detection model that balances accuracy and speed. This solves the problem of repeated trials during model training and enables rapid model training while maintaining a balance between speed and accuracy.
[0113] Figure 5 This is a schematic flowchart of an object detection method provided in an embodiment of this disclosure. This embodiment is applicable to situations where an object detection model is being trained for processing multimedia content. The method can be executed by an object detection device, which can be implemented in software and / or hardware, and is generally integrated into any electronic device with network communication capabilities. Such electronic devices include, but are not limited to, mobile terminals, PCs, or servers. Figure 5 As shown, the target detection method includes steps S510-S520:
[0114] S510. Determine the multimedia content to be processed.
[0115] S520. Input the multimedia content to be processed into the target detection model to obtain the target detection result of the multimedia content to be processed.
[0116] The target detection model is obtained using any of the target detection model training methods described in the above embodiments. After entering the model application stage, the multimedia content to be processed can be input into the trained target detection model, which then outputs whether the multimedia content contains the target content.
[0117] The technical solution of this disclosure embodiment selects the feature extraction network parameters of the current stage in the model training stage to adjust the feature extraction network parameters of the current stage, so as to obtain the optimal feature extraction network parameters under a specific image input resolution parameter; selects the image input resolution parameters of the current stage in the neighboring stage to adjust, so as to obtain the optimal image input resolution parameters under the optimal extraction network parameters, thereby further improving the model's detection accuracy and running speed of target content.
[0118] Figure 6 This is a schematic diagram of a target detection model training device provided in an embodiment of this disclosure. This embodiment is applicable to training a target detection model for multimedia content. This method can be executed by a target detection device, which can be implemented in software and / or hardware and is generally integrated into any electronic device with network communication capabilities. The electronic device includes, but is not limited to, mobile terminals, PCs, or servers. Figure 6 As shown, the target detection model training device includes: a first training task execution module 610, a network parameter determination module 620, and a target detection model determination module 630; wherein:
[0119] The first training task execution module 610 is used to execute the first training task on the object detection model at the current stage;
[0120] The network parameter determination module 620 is used to determine the network parameters to be used for the first training task of the target detection model in the next stage, based on the performance consumption and detection accuracy of the target detection model trained in the current stage.
[0121] The target detection model determination module 630 is used to perform a second training task on the target detection model based on the network parameters used in the first training task, so as to obtain a converged target detection model; wherein the amount of training data in the first training task is less than the amount of training data in the second training task.
[0122] In one optional embodiment of this disclosure, the first training task execution module may include:
[0123] The second training dataset used in the second training task is used to collect data to obtain the first training dataset; wherein the similarity between the label type distribution of the training samples in the first training dataset and the label type distribution of the training samples in the second training dataset is greater than a preset similarity threshold.
[0124] The first training task is performed on the object detection model at the current stage using the first training dataset.
[0125] In one optional embodiment of this disclosure, the first training task execution module may further include:
[0126] When the current stage is the first training stage, the network parameters are loaded and initialized in the object detection model, and the first training task is performed on the object detection model with the initialized network parameters loaded.
[0127] In one optional embodiment of this disclosure, the first training task execution module may further include:
[0128] When the current stage is not the first training stage, load the network parameters that were pre-searched in the previous stage and used to perform the first training task on the object detection model in the current stage into the object detection model;
[0129] Perform the first training task on the object detection model loaded with the network parameters pre-searched in the previous stage.
[0130] In one optional embodiment of this disclosure, the network parameter determination module may include:
[0131] If it is detected that the performance consumption of the target detection model trained in the current stage is greater than the preset performance consumption limit and / or the detection accuracy is greater than the preset accuracy value, then the network parameters used by the target detection model in the current stage will be adjusted first to reduce the complexity of the feature extraction network in the target detection model.
[0132] If it is detected that the performance consumption of the target detection model trained in the current stage is less than the preset performance consumption limit and / or the detection accuracy is less than the preset accuracy value, then the network parameters used by the target detection model in the current stage will be adjusted a second time to improve the complexity of the feature extraction network in the target detection model.
[0133] The network parameters obtained through the first or second adjustment are determined as the network parameters used in the next stage to perform the first training task on the object detection model.
[0134] In one optional embodiment of this disclosure, the network parameter determination module may further include:
[0135] The search direction is directed towards reducing the complexity of the feature extraction network in the object detection model. The nearest neighbor network parameters that are used in the current stage object detection model are searched from the preset network parameter library to complete the first adjustment of the network parameters used in the current stage object detection model.
[0136] In one optional embodiment of this disclosure, the network parameter determination module may further include:
[0137] The search direction is directed towards improving the complexity of the feature extraction network in the object detection model. The nearest neighbor network parameters that are used by the current object detection model are searched from the preset network parameter library to complete the second adjustment of the network parameters used by the current object detection model.
[0138] In one optional embodiment of this disclosure, the network parameter determination module may further include:
[0139] Keeping the supported image input resolution parameters of the network parameters used by the target detection model in the current stage unchanged, the system searches for feature extraction network parameters that are close to the network parameters used by the target detection model in the current stage from the preset network parameter library.
[0140] In one optional embodiment of this disclosure, the network parameter determination module may further include:
[0141] Keeping the feature extraction network parameters unchanged in the network parameters used by the target detection model at the current stage, the system searches from the preset network parameter library for suitable and supported image input resolution parameters that are close to the network parameters used by the target detection model at the current stage.
[0142] In one optional embodiment of this disclosure, the network parameter determination module may further include:
[0143] The network parameters include the image input resolution parameters and feature extraction network parameters that are applicable and supported in the object detection model.
[0144] In one optional embodiment of this disclosure, the network parameter determination module may further include:
[0145] The preset network parameter library includes feature extraction network parameters in at least two dimensions and applicable supported image input resolution parameters in at least two dimensions.
[0146] In one optional embodiment of this disclosure, the target detection model determination module may include:
[0147] Select network parameters that meet the preset search conditions from the network parameters used in the first training task;
[0148] Based on the network parameters that meet the preset search conditions, perform a second training task on the object detection model;
[0149] Among them, the network parameters with preset search conditions include the network parameters that minimize the complexity of the feature extraction network in the target detection model while keeping the performance consumption of the target detection model less than the preset performance consumption limit and the detection accuracy greater than the preset accuracy value.
[0150] In one optional embodiment of this disclosure, the target detection model determination module may further include:
[0151] Load network parameters that meet the preset search conditions into the object detection model;
[0152] Using the second training dataset, a second training task is performed on the object detection model loaded with network parameters that meet the preset search conditions;
[0153] The second training dataset used to perform the first training task is downsampled to obtain the first training dataset used to perform the first training task.
[0154] The target detection model training device provided in this disclosure can execute the target detection model training method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the target detection model training method.
[0155] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0156] Figure 7 This is a schematic diagram of a target detection device provided in an embodiment of this disclosure. This embodiment is applicable to training a target detection model for multimedia content. This method can be executed by a target detection device, which can be implemented in software and / or hardware and is generally integrated into any electronic device with network communication capabilities. The electronic device includes, but is not limited to, mobile terminals, PCs, or servers. Figure 7 As shown, the target detection device includes:
[0157] Multimedia content determination module 710 is used to determine the multimedia content to be processed;
[0158] The target detection module 720 is used to input the multimedia content to be processed into the target detection model to obtain the target detection result of the multimedia content to be processed.
[0159] The target detection model is obtained using any of the target detection model training methods described in the above embodiments. After entering the model application stage, the multimedia content to be processed can be input into the trained target detection model, which then outputs whether the multimedia content contains the target content.
[0160] The target detection device provided in this embodiment can execute the target detection method provided in any of the above embodiments of this disclosure, and has the corresponding functions and beneficial effects of executing the target detection method. For details, please refer to the relevant operations of the target detection method in the foregoing embodiments.
[0161] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 8 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 8 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0162] like Figure 8 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0163] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0164] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0165] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0166] The electronic device provided in this embodiment belongs to the same inventive concept as the target detection model training method or target detection method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0167] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the target detection model training method or target detection method provided in the above embodiments.
[0168] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0169] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0170] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0171] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: perform a first training task on the target detection model in the current stage; determine the network parameters to be used for performing the first training task on the target detection model in the next stage based on the performance consumption and detection accuracy of the target detection model trained in the current stage; and perform a second training task on the target detection model based on the network parameters used in the first training task to obtain a converged target detection model; wherein the amount of training data for the first training task is less than the amount of training data for the second training task.
[0172] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: determine the multimedia content to be processed; input the multimedia content to be processed into the target detection model; and obtain the target detection result of the multimedia content to be processed.
[0173] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0174] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0175] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0176] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0177] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0178] According to one or more embodiments of this disclosure, Example 1 provides a method for training an object detection model, the method comprising:
[0179] The first training task is performed on the object detection model at the current stage;
[0180] Based on the performance consumption and detection accuracy of the target detection model trained in the current stage, determine the network parameters to be used for the first training task of the target detection model in the next stage;
[0181] Based on the network parameters used in the first training task, a second training task is performed on the target detection model to obtain a converged target detection model;
[0182] The amount of training data for the first training task is less than the amount of training data for the second training task.
[0183] Example 2, based on the method described in Example 1, performs a first training task on the object detection model at the current stage, including:
[0184] The second training dataset used in the second training task is used to collect data to obtain the first training dataset; wherein the similarity between the label type distribution of the training samples in the first training dataset and the label type distribution of the training samples in the second training dataset is greater than a preset similarity threshold.
[0185] The first training task is performed on the object detection model at the current stage using the first training dataset.
[0186] Example 3, according to the method described in Example 1 or 2, is characterized in that, in the current stage, a first training task is performed on the object detection model, including:
[0187] When the current stage is the first training stage, the network parameters are loaded and initialized in the object detection model, and the first training task is performed on the object detection model with the initialized network parameters loaded.
[0188] Example 4, according to the method described in Example 1 or 2, is characterized in that, in the current stage, a first training task is performed on the object detection model, including:
[0189] When the current stage is not the first training stage, load the network parameters that were pre-searched in the previous stage and used to perform the first training task on the object detection model in the current stage into the object detection model;
[0190] Perform the first training task on the object detection model loaded with the network parameters pre-searched in the previous stage.
[0191] Example 5, according to the method described in Example 1, is characterized in that, based on the performance consumption and detection accuracy of the target detection model trained in the current stage, the network parameters used to perform the first training task on the target detection model in the next stage are determined, including:
[0192] If it is detected that the performance consumption of the target detection model trained in the current stage is greater than the preset performance consumption limit and / or the detection accuracy is greater than the preset accuracy value, then the network parameters used by the target detection model in the current stage will be adjusted first to reduce the complexity of the feature extraction network in the target detection model.
[0193] If it is detected that the performance consumption of the target detection model trained in the current stage is less than the preset performance consumption limit and / or the detection accuracy is less than the preset accuracy value, then the network parameters used by the target detection model in the current stage will be adjusted a second time to improve the complexity of the feature extraction network in the target detection model.
[0194] The network parameters obtained through the first or second adjustment are determined as the network parameters used in the next stage to perform the first training task on the object detection model.
[0195] Example 6, based on the method described in Example 5, is characterized by a first adjustment to the network parameters used in the target detection model at the current stage, including:
[0196] The search direction is directed towards reducing the complexity of the feature extraction network in the object detection model. The nearest neighbor network parameters that are used in the current stage object detection model are searched from the preset network parameter library to complete the first adjustment of the network parameters used in the current stage object detection model.
[0197] Example 7, based on the method described in Example 5, is characterized by a second adjustment to the network parameters used in the target detection model at the current stage, including:
[0198] The search direction is directed towards improving the complexity of the feature extraction network in the object detection model. The nearest neighbor network parameters that are used by the current object detection model are searched from the preset network parameter library to complete the second adjustment of the network parameters used by the current object detection model.
[0199] Example 8, according to the method described in Example 6 or 7, is characterized in that searching from a preset network parameter library for nearest neighbor network parameters that are used by the network parameters of the target detection model in the current stage includes:
[0200] Keeping the supported image input resolution parameters of the network parameters used by the target detection model in the current stage unchanged, the system searches for feature extraction network parameters that are close to the network parameters used by the target detection model in the current stage from the preset network parameter library.
[0201] Example 9, according to the method described in Example 6 or 7, is characterized in that, searching from a preset network parameter library for nearest neighbor network parameters that are the network parameters used by the target detection model in the current stage, further includes:
[0202] Keeping the feature extraction network parameters unchanged in the network parameters used by the target detection model at the current stage, the system searches from the preset network parameter library for suitable and supported image input resolution parameters that are close to the network parameters used by the target detection model at the current stage.
[0203] Example 10 describes the method according to Example 6 or 7, wherein the network parameters include applicable and supported image input resolution parameters and feature extraction network parameters in the object detection model.
[0204] Example 11, according to the method described in Example 6 or 7, is characterized in that the preset network parameter library includes feature extraction network parameters of at least two dimensions and applicable supported image input resolution parameters of at least two dimensions; the feature extraction network parameters of at least two dimensions are deployed in order of feature extraction network complexity; and the applicable supported image input resolution parameters of at least two dimensions are deployed in order of image resolution.
[0205] Example 12: According to the method described in Example 1, the second training task is performed on the target detection model based on the network parameters used in the first training task, including:
[0206] Select network parameters that meet the preset search conditions from the network parameters used in the first training task;
[0207] Based on the network parameters that meet the preset search conditions, perform a second training task on the object detection model;
[0208] Among them, the network parameters with preset search conditions include the network parameters that minimize the complexity of the feature extraction network in the target detection model while keeping the performance consumption of the target detection model less than the preset performance consumption limit and the detection accuracy greater than the preset accuracy value.
[0209] Example 13, according to the method described in Example 12, is characterized in that, based on network parameters that satisfy preset search conditions, a second training task is performed on the target detection model, including:
[0210] Load network parameters that meet the preset search conditions into the object detection model;
[0211] Using the second training dataset, a second training task is performed on the object detection model loaded with network parameters that meet the preset search conditions;
[0212] The second training dataset used to perform the first training task is downsampled to obtain the first training dataset used to perform the first training task.
[0213] According to one or more embodiments of this disclosure, Example 14 provides an object detection method, which uses the training method of any of the object detection models described in Examples 1-13 to obtain an object detection model, the detection method comprising:
[0214] Identify the multimedia content to be processed;
[0215] The multimedia content to be processed is input into the target detection model to obtain the target detection result of the multimedia content to be processed.
[0216] According to one or more embodiments of this disclosure, Example 15 also provides an object detection model training apparatus, the training apparatus comprising:
[0217] The first training task execution module is used to execute the first training task on the object detection model at the current stage;
[0218] The network parameter determination module is used to determine the network parameters to be used for the first training task of the target detection model in the next stage, based on the performance consumption and detection accuracy of the target detection model trained in the current stage.
[0219] The target detection model determination module is used to perform a second training task on the target detection model based on the network parameters used in the first training task, so as to obtain a converged target detection model; wherein the amount of training data in the first training task is less than the amount of training data in the second training task.
[0220] According to one or more embodiments of this disclosure, Example 16 provides an object detection device, which uses an object detection model obtained by training any of the object detection models described in Examples 1-13, the object detection device comprising:
[0221] The multimedia content determination module is used to determine the multimedia content to be processed.
[0222] The target detection module is used to input the multimedia content to be processed into the target detection model to obtain the target detection result of the multimedia content to be processed.
[0223] According to one or more embodiments of this disclosure, Example 17 also provides an electronic device, the electronic device comprising:
[0224] One or more processors;
[0225] Storage device for storing one or more programs.
[0226] When the one or more programs are executed by the one or more processors, the one or more processors implement the object detection model training method as described in any one of claims 1-13 or the object detection method as described in claim 14.
[0227] According to one or more embodiments of this disclosure, Example 18 also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the object detection model training method as described in any one of claims 1-13 or the object detection method as described in claim 14.
[0228] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0229] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0230] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for training an object detection model, characterized in that, The training method includes: The first training task is performed on the object detection model at the current stage; the object detection model is used to detect objects in the multimedia content to be processed, so as to label the multimedia content to be processed. Based on the performance consumption and detection accuracy of the target detection model trained in the current stage, determine the network parameters to be used for the first training task of the target detection model in the next stage; Based on the network parameters used in the first training task, a second training task is performed on the target detection model to obtain a converged target detection model; The amount of training data for the first training task is less than the amount of training data for the second training task. The step of determining the network parameters to be used for the first training task of the target detection model in the next stage, based on the performance consumption and detection accuracy of the target detection model trained in the current stage, includes: If the performance consumption of the target detection model trained in the current stage is found to be greater than the preset performance consumption limit and / or the detection accuracy is greater than the preset accuracy value, the search direction is shifted towards reducing the complexity of the feature extraction network in the target detection model. The applicable image input resolution parameters in the network parameters used by the target detection model in the current stage are kept unchanged. Feature extraction network parameters that are close to the network parameters used by the target detection model in the current stage are searched from the preset network parameter library to complete the first adjustment of the network parameters used by the target detection model in the current stage, so as to reduce the complexity of the feature extraction network in the target detection model. The network parameters obtained after the first adjustment are determined as the network parameters used for the first training task of the target detection model in the next stage. If the performance consumption of the target detection model trained in the current stage is less than the preset performance consumption limit and / or the detection accuracy is less than the preset accuracy value, then the network parameters used by the target detection model in the current stage are adjusted a second time to improve the complexity of the feature extraction network in the target detection model. Following a search direction that improves the complexity of the feature extraction network in the target detection model, while keeping the feature extraction network parameters unchanged, the model searches from the preset network parameter library for suitable and supported image input resolution parameters that are adjacent to the network parameters used by the target detection model in the current stage, thus completing the second adjustment of the network parameters used by the target detection model in the current stage. The network parameters obtained after the second adjustment are determined as the network parameters used in the next stage to perform the first training task on the target detection model. The network parameters include applicable and supported image input resolution parameters and feature extraction network parameters in the object detection model; the preset network parameter library includes feature extraction network parameters in at least two dimensions and applicable and supported image input resolution parameters in at least two dimensions; the feature extraction network parameters in at least two dimensions are deployed in order of feature extraction network complexity; the applicable and supported image input resolution parameters in at least two dimensions are deployed in order of image resolution.
2. The method according to claim 1, characterized in that, The first training task for the object detection model at this stage includes: The second training dataset used in the second training task is used to collect data to obtain the first training dataset; wherein the similarity between the label type distribution of the training samples in the first training dataset and the label type distribution of the training samples in the second training dataset is greater than a preset similarity threshold. The first training task is performed on the object detection model at the current stage using the first training dataset.
3. The method according to claim 1 or 2, characterized in that, The first training task for the object detection model at this stage includes: When the current stage is the first training stage, the network parameters are loaded and initialized in the object detection model, and the first training task is performed on the object detection model with the initialized network parameters loaded.
4. The method according to claim 1 or 2, characterized in that, The first training task for the object detection model at this stage includes: When the current stage is not the first training stage, load the network parameters that were pre-searched in the previous stage and used to perform the first training task on the object detection model in the current stage into the object detection model; Perform the first training task on the object detection model loaded with the network parameters pre-searched in the previous stage.
5. The method according to claim 1, characterized in that, Based on the network parameters used in the first training task, a second training task is performed on the object detection model, including: Select network parameters that meet the preset search conditions from the network parameters used in the first training task; Based on the network parameters that meet the preset search conditions, perform a second training task on the object detection model; Among them, the network parameters with preset search conditions include the network parameters that minimize the computational complexity of the target detection model while keeping the performance consumption of the target detection model less than the preset performance consumption limit and the detection accuracy greater than the preset accuracy value.
6. The method according to claim 5, characterized in that, Based on the network parameters that meet the preset search conditions, a second training task is performed on the object detection model, including: Load network parameters that meet the preset search conditions into the object detection model; Using the second training dataset, a second training task is performed on the object detection model loaded with network parameters that meet the preset search conditions; The second training dataset used to perform the first training task is downsampled to obtain the first training dataset used to perform the first training task.
7. A target detection method, characterized in that, An object detection model obtained by training the object detection model according to any one of claims 1-6, wherein the object detection method includes: Identify the multimedia content to be processed; The multimedia content to be processed is input into the target detection model to obtain the target detection result of the multimedia content to be processed.
8. A target detection model training device, characterized in that, The training device includes: The first training task execution module is used to execute the first training task on the object detection model at the current stage; the object detection model is used to perform object detection on the multimedia content to be processed, so as to label the multimedia content to be processed. The network parameter determination module is used to determine the network parameters to be used for the first training task of the target detection model in the next stage, based on the performance consumption and detection accuracy of the target detection model trained in the current stage. The target detection model determination module is used to perform a second training task on the target detection model based on the network parameters used in the first training task, so as to obtain a converged target detection model; wherein the amount of training data in the first training task is less than the amount of training data in the second training task. The network parameter determination module is specifically used for: If the performance consumption of the target detection model trained in the current stage is found to be greater than the preset performance consumption limit and / or the detection accuracy is greater than the preset accuracy value, the search direction is shifted towards reducing the complexity of the feature extraction network in the target detection model. The applicable image input resolution parameters in the network parameters used by the target detection model in the current stage are kept unchanged. Feature extraction network parameters that are close to the network parameters used by the target detection model in the current stage are searched from the preset network parameter library to complete the first adjustment of the network parameters used by the target detection model in the current stage, so as to reduce the complexity of the feature extraction network in the target detection model. The network parameters obtained after the first adjustment are determined as the network parameters used for the first training task of the target detection model in the next stage. If the performance consumption of the target detection model trained in the current stage is less than the preset performance consumption limit and / or the detection accuracy is less than the preset accuracy value, then the network parameters used by the target detection model in the current stage are adjusted a second time to improve the complexity of the feature extraction network in the target detection model. Following a search direction that improves the complexity of the feature extraction network in the target detection model, while keeping the feature extraction network parameters unchanged, the model searches from the preset network parameter library for suitable and supported image input resolution parameters that are adjacent to the network parameters used by the target detection model in the current stage, thus completing the second adjustment of the network parameters used by the target detection model in the current stage. The network parameters obtained after the second adjustment are determined as the network parameters used in the next stage to perform the first training task on the target detection model. The network parameters include applicable and supported image input resolution parameters and feature extraction network parameters in the object detection model; the preset network parameter library includes feature extraction network parameters in at least two dimensions and applicable and supported image input resolution parameters in at least two dimensions; the feature extraction network parameters in at least two dimensions are deployed in order of feature extraction network complexity; the applicable and supported image input resolution parameters in at least two dimensions are deployed in order of image resolution.
9. A target detection device, characterized in that, The target detection device includes: The multimedia content determination module is used to determine the multimedia content to be processed. The target detection module is used to input the multimedia content to be processed into the target detection model to obtain the target detection result of the multimedia content to be processed.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the object detection model training method as described in any one of claims 1-6 or the object detection method as described in claim 7.
11. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the object detection model training method as described in any one of claims 1-6 or the object detection method as described in claim 7.