Multi-target detection model training method, multi-target detection method, equipment and medium

By generating anchor boxes of target objects and training the model with image datasets, and using the attention boundary loss function to optimize Yolov5, the detection errors of Yolov5 in complex scenes and small target detection are solved, achieving higher detection accuracy and reliability.

CN120707988APending Publication Date: 2025-09-26艾索信息股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510827217.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing real-time target detection algorithm Yolov5 suffers from missed detection or false detection in complex scenes and small target detection, making it difficult to meet high-precision requirements, especially in situations with complex lighting and dense targets.

Method used

By acquiring an image dataset, generating anchor frames for at least one type of target object, and using the anchor frames and image datasets for model training, the loss function value is calculated using a preset attention boundary loss function, and the parameters are adjusted to obtain a multi-target detection model, including feature extraction, fusion and output modules, and image preprocessing is performed to enhance model adaptability.

Benefits of technology

It improves the accuracy and reliability of target detection, solves model detection errors, and enhances the detection capabilities of complex scenes and small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707988A_ABST
    Figure CN120707988A_ABST
Patent Text Reader

Abstract

The invention provides a multi-target detection model training method, a multi-target detection method, equipment and a medium, and relates to the technical field of target detection. The multi-target detection model training method comprises the following steps: acquiring an image data set, wherein each sample image in the image data set has category label data of at least one target object; then according to category label data of at least one target object in multiple sample images in the image data set, generating an anchor box of at least one category of target objects, and finally according to the anchor box of at least one category of target objects and the image data set, performing model training to obtain a multi-target detection model. The multi-target detection model is used for performing multi-target detection on the to-be-detected image of the preset scene, and according to the method provided by the invention, model training is performed through the anchor frame of the at least one type of target object and the image data set, so that the accuracy and reliability of target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and in particular to a multi-target detection model training method, a multi-target detection method, a device and a medium. Background Art

[0002] Object detection is a critical task in computer vision, widely used in numerous fields such as security surveillance, autonomous driving, and industrial inspection. The real-time object detection algorithm (You Only Look Once version 5, Yolov5) offers advantages such as fast detection speed and ease of edge deployment. However, it still has limitations when it comes to complex scenes, small object detection, and scenarios requiring high accuracy. For example, Yolov5 may miss or misdetect objects in complex lighting conditions and densely packed objects. Its accuracy for detecting tiny parts in industrial inspections still falls short of meeting high-precision requirements. For detecting targets such as port management vessels, the accuracy and recall of detecting small boats at long distances in foggy conditions also need improvement. Therefore, it is necessary to improve Yolov5's performance to better suit various practical application scenarios. Summary of the Invention

[0003] The purpose of the present invention is to address the deficiencies in the above-mentioned prior art and provide a multi-target detection model training method, multi-target detection method, device and medium, so as to perform model training through anchor frames and image data sets of at least one type of target object, thereby improving the accuracy and reliability of target detection.

[0004] To achieve the above objectives, the technical solutions adopted in the embodiments of the present application are as follows: In a first aspect, an embodiment of the present application provides a multi-target detection model training method, the method comprising: Acquire an image dataset, wherein each sample image in the image dataset has category label data of at least one target object; generating an anchor frame of at least one type of target object according to category label data of at least one target object in a plurality of sample images in the image dataset; Model training is performed based on the anchor frames of the at least one type of target object and the image dataset to obtain a multi-target detection model, which is used to perform multi-target detection based on the image to be detected in a preset scene.

[0005] In an optional embodiment, acquiring the image dataset includes: Acquire a plurality of images having the at least one target object from a preset image dataset; Annotating the at least one target object in each image to obtain a sample image having category label data of the at least one target object; The image dataset is constructed based on the multiple sample images.

[0006] In an optional embodiment, generating an anchor frame of at least one type of target object based on the category label data of the at least one target object in the plurality of sample images in the image dataset includes: generating a label file for the at least one target object according to the category label data of the at least one target object in the plurality of sample images; According to the label file of the at least one target object, a preset anchor frame clustering function is used to generate anchor frames of the at least one type of target object.

[0007] In an optional embodiment, the performing model training based on the anchor boxes of the at least one type of target object and the image dataset to obtain a multi-target detection model includes: Filtering positive sample images of the at least one target object from the image dataset according to the anchor frames of the at least one type of target object; Using a preset target detection model, perform target detection on the positive sample image to obtain a predicted position of the positive sample image for the at least one target object; Calculating a loss function value of the at least one target object using a preset attention boundary loss function according to the category label data of the at least one target object in the positive sample image and the predicted position of the at least one target object; Calculating a target loss function value based on the loss function value of the at least one target object; According to the target loss function value, the preset target detection model is adjusted to obtain the multi-target detection model.

[0008] In an optional embodiment, screening positive sample images of the at least one target object from the image dataset based on the anchor frames of the at least one type of target object includes: performing shape matching on the at least one target object in each sample image according to the anchor frame of the at least one type of target object, to obtain a shape matching degree of the at least one target object in each sample image; If the shape matching degree of the at least one target object is less than a preset threshold, each of the sample images is determined to be a positive sample image of the at least one target object.

[0009] In an optional embodiment, the target detection model includes: a feature extraction module, a feature fusion module and an output module, wherein the feature extraction module is used to extract features from an input sample image, the feature fusion module is used to fuse the extracted image features, and the output module is used to output target detection results based on the fused features; The feature extraction module includes: a two-dimensional convolution unit, a two-dimensional normalization unit and a linear activation function unit.

[0010] In an optional embodiment, before performing model training based on the anchor boxes of the at least one type of target object and the image dataset to obtain a multi-target detection model, the method further includes at least one of the following image preprocessing operations: Each sample image in the image data set is subjected to image color channel adjustment processing, image enhancement processing, noise processing, and image space transformation processing; wherein the image enhancement processing includes: image blur processing, limited contrast adjustment processing, image brightness contrast adjustment processing, image gamma correction processing, image fogging processing, and image color saturation processing; the noise processing includes: noise addition processing and noise filtering processing; and the image space transformation processing includes: image horizontal flip processing, cropping processing, and scaling processing.

[0011] In a second aspect, an embodiment of the present application further provides a multi-target detection method, the method comprising: Obtain the image to be detected of the preset scene; According to the image to be detected in the preset scene, a multi-target detection model is used to perform detection to obtain a target detection result of the image to be detected, wherein the target detection result includes: detection information of the image to be detected for at least one target object, and the multi-target detection model is any multi-target detection model described in the first aspect above.

[0012] In a third aspect, an embodiment of the present application further provides a computer device comprising: a processor, a storage medium, and a bus, wherein the storage medium stores program instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium through the bus, and the processor executes the program instructions to perform the steps of the multi-target detection model training method as described in any one of the first aspects, or to perform the steps of the multi-target detection method as described in any one of the second aspects.

[0013] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the computer program executes the steps of the multi-target detection model training method as described in any one of the first aspects, or executes the steps of the multi-target detection method as described in any one of the second aspects.

[0014] The beneficial effects of this application are: An embodiment of the present application provides a multi-target detection model training method, a multi-target detection method, an apparatus and a medium. The multi-target detection model training method includes: obtaining an image data set, wherein each sample image in the image data set has category label data of at least one target object; wherein the at least one target object is an object in the same scene whose size and shape meet preset difference conditions, and then generating anchor frames of at least one type of target object based on the category label data of at least one target object in multiple sample images in the image data set, and finally performing model training based on the anchor frames of at least one type of target object and the image data set to obtain a multi-target detection model. The multi-target detection model is used to perform multi-target detection on images to be detected in a preset scene. The method of the present application improves the accuracy and reliability of target detection by performing model training through anchor frames and image data sets of at least one type of target object, and solves the problem of detection errors in model detection caused by using the same set of anchor frames and image data sets for model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 One of the flow charts of a multi-target detection model training method provided in an embodiment of the present application; Figure 2 The second flowchart of a multi-target detection model training method provided in an embodiment of the present application; Figure 3 The third flowchart of a multi-target detection model training method provided in an embodiment of the present application; Figure 4 A fourth flowchart of a multi-target detection model training method provided in an embodiment of the present application; Figure 5 A flowchart of a multi-target detection method provided in an embodiment of the present application; Figure 6 A schematic diagram of the functional modules of a multi-target detection model training device provided in an embodiment of the present application; Figure 7 A schematic diagram of the functional modules of a multi-target detection device provided in an embodiment of the present application; Figure 8 A schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0018] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.

[0019] In the description of this application, it should be noted that if the terms "upper", "lower", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the accompanying drawings, or is the orientation or position relationship in which the product of the application is usually placed when in use. It is only for the convenience of describing this application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on this application.

[0020] In addition, the terms "first," "second," and the like in the description and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0021] It should be noted that, in the absence of conflict, the features in the embodiments of this application can be combined with each other.

[0022] In order to achieve accurate detection of multiple target objects in an image to be detected, an embodiment of the present application provides a multi-target detection model training method, which first obtains an image dataset, wherein each sample image in the image dataset has category label data of at least one target object; wherein, at least one target object is an object in the same scene whose size and shape meet preset difference conditions, and then generates an anchor frame of at least one type of target object based on the category label data of at least one target object in multiple sample images in the image dataset, and finally performs model training based on the anchor frame of at least one type of target object and the image dataset to obtain a multi-target detection model, which is used to perform multi-target detection on the image to be detected in a preset scene, so that target objects with size differences in the image to be detected can be accurately detected, thereby improving the accuracy and reliability of target detection.

[0023] The following describes a detailed explanation of the multi-target detection model training method provided in the embodiments of the present application using specific examples and accompanying drawings. The multi-target detection model training method provided in the embodiments of the present application can be implemented by running a computer device pre-installed with a preset multi-target detection model training algorithm or detection software. The computer device can be, for example, a server or terminal, where the terminal can be a user's computer. Figure 1 This is one of the flow charts of a multi-target detection model training method provided in an embodiment of the present application; Figure 1 As shown, the method includes: S101: Obtain an image dataset.

[0024] In this embodiment, each sample image in the image dataset has category label data of at least one target object.

[0025] Specifically, in the same scene, there are target objects with significantly different sizes and shapes that all meet the preset difference conditions. For example, in a seaport scene, each sample image includes at least one target object, which can be a ship and / or a buoy. The sizes and shapes of the ships and buoys differ significantly. If a sample image includes one ship and two buoys, the sample image will have class label data for the ship and class label data for the two buoys.

[0026] S102 : Generate an anchor frame of at least one type of target object according to category label data of at least one target object in a plurality of sample images in the image dataset.

[0027] S103 . Perform model training based on anchor frames and image datasets of at least one type of target object to obtain a multi-target detection model. The multi-target detection model is used to perform multi-target detection based on images to be detected in a preset scene.

[0028] Specifically, the category label data of each target object is screened out from multiple sample images in the image dataset, and then the anchor box of each type of target object is generated according to the category label data of each target object.

[0029] By training a preset target detection model based on anchor frames and image datasets of at least one type of target object, a multi-target detection model is finally trained. The multi-target detection model is used to accurately detect each target object in the image to be detected.

[0030] In summary, an embodiment of the present application provides a multi-target detection model training method, the method comprising: obtaining an image dataset, wherein each sample image in the image dataset has category label data of at least one target object; wherein the at least one target object is an object in the same scene whose size and shape meet preset difference conditions, and then generating anchor frames of at least one type of target object based on the category label data of at least one target object in multiple sample images in the image dataset, and finally performing model training based on the anchor frames of at least one type of target object and the image dataset to obtain a multi-target detection model, which is used to perform multi-target detection on images to be detected in a preset scene. The method of the present application improves the accuracy and reliability of target detection by performing model training through anchor frames and image datasets of at least one type of target object, and solves the problem of detection errors in model detection caused by using the same set of anchor frames and image datasets for model training.

[0031] The present application also provides another possible implementation of the multi-target detection model training method. Figure 2 The second flow chart of a multi-target detection model training method provided in the embodiment of the present application is as follows: Figure 2 As shown, obtain the image dataset, including: S201: Acquire multiple images having at least one target object from a preset image dataset.

[0032] S202 : Label at least one target object in each image to obtain a sample image having category label data of the at least one target object.

[0033] S203: Construct an image dataset based on multiple sample images.

[0034] In this embodiment, the preset image dataset includes a public dataset and / or a collected image dataset consisting of images pre-collected in the same scene, wherein the public dataset can be: an image recognition dataset (Common Objects in Context, COCO2017), a visual object category dataset (Visual Object Classes, VOC2017), a maritime vessel recognition dataset SeaShip (7000), etc.

[0035] If the preset image dataset includes a public dataset, and at least one target object includes a ship object and a floating object, then based on the category label data of the ship object, multiple images including the ship object are obtained from the public dataset, and an initial image dataset is constructed based on the multiple images including the ship object.

[0036] Then, floating object detection is performed on the images in the initial image dataset. This detection method can be performed using a single target detection model or manually by a human, without limitation. If an image is detected to contain a floating object, the floating object is annotated to obtain category label data for the floating object. The sample image corresponding to the image includes category label data for both the ship object and the floating object.

[0037] Specifically, the label format of the category label data is: [label type, normalized center point x-coordinate, normalized center point y-coordinate, normalized target box width, normalized target box height]. The label type of a ship object can be set to 0, and the label type of a float object can be set to 1.

[0038] The normalized coordinate calculation formula is as follows: norm_width = (xmax - xmin) / img_w norm_hight = (ymax - ymin) / img_h norm_centx = (xmax + xmin)*0.5 / img_w norm_centy = (ymax + ymin) *0.5 / img_h Among them, (xmin, ymin) and (xmax, ymax) represent the coordinates of the upper left corner and lower right corner of the target rectangular box corresponding to the target object, respectively; img_w represents the image width, img_h represents the image height; norm_centx represents the normalized center point x coordinate corresponding to the target object, norm_centy represents the normalized center point y coordinate corresponding to the target object, norm_width represents the normalized target box width corresponding to the target object, and norm_hight represents the normalized target box height corresponding to the target object.

[0039] If the preset image dataset also includes a collected image dataset composed of images pre-collected in the same scene, then each image in the collected image dataset is detected as a ship object and a float object, and the ship object and the float object are labeled respectively to obtain the category label data of the ship object and the float object, and obtain sample images. Finally, an image dataset is constructed based on all the sample images.

[0040] In the method provided in an embodiment of the present application, multiple images with at least one target object are obtained from a preset image dataset; at least one target object in each image is separately labeled to obtain a sample image with category label data of at least one target object; based on the multiple sample images, an image dataset is constructed, and by separately labeling at least one target object, each sample image in the image dataset has category label data of at least one target object.

[0041] The present application also provides another possible implementation of the multi-target detection model training method. Figure 3 The third flow chart of a multi-target detection model training method provided in the embodiment of the present application is as follows: Figure 3 As shown, generating an anchor frame of at least one type of target object based on the category label data of at least one target object in multiple sample images in an image dataset includes: S301 : Generate a label file for at least one target object according to category label data of at least one target object in a plurality of sample images.

[0042] S302: Generate anchor frames for at least one type of target object using a preset anchor frame clustering function according to a label file of at least one target object.

[0043] In this embodiment, the category label data of each target object is screened out from a plurality of sample images, and a label file of the target object is generated based on all the category label data of the same target object.

[0044] For example, all category label data of ship objects are filtered out from multiple sample images to generate a label file for ship objects, and all category label data of float objects are filtered out to generate a label file for float objects.

[0045] Based on the label file for the ship object, a preset anchor frame clustering function is used to generate anchor frames for the ship object. Based on the label file for the float object, a preset anchor frame clustering function is also used to generate anchor frames for the float object. This generates anchor frames for two types of target objects: anchor frames for the ship object and anchor frames for the float object. The preset anchor frame clustering function is a K-means clustering algorithm function. The anchor frames for the ship object are composed of a set of anchor frames, each of which has three different width and height dimensions. Similarly, the anchor frames for the float object are also composed of a set of anchor frames, each of which has three different width and height dimensions.

[0046] In the method provided in an embodiment of the present application, a label file of at least one target object is generated based on the category label data of at least one target object in multiple sample images. Based on the label file of at least one target object, a preset anchor frame clustering function is used to generate anchor frames of at least one category of target objects. Since there are differences in the shape and size of at least one target object, anchor frames of at least one category of target objects are generated based on the category label data of at least one target object for subsequent model training.

[0047] The present application also provides another possible implementation of the multi-target detection model training method. Figure 4 The fourth flow chart of a multi-target detection model training method provided in the embodiment of the present application is as follows: Figure 4 As shown, the model is trained based on the anchor boxes and image datasets of at least one type of target object to obtain a multi-target detection model, including: S401: Filter at least one positive sample image of a target object from an image dataset according to an anchor frame of at least one type of target object.

[0048] S402: Using a preset target detection model, perform target detection on the positive sample image to obtain a predicted position of at least one target object in the positive sample image.

[0049] S403: Calculate the loss function value of the at least one target object using a preset attention boundary loss function according to the category label data of the at least one target object in the positive sample image and the predicted position of the at least one target object.

[0050] S404. Calculate a target loss function value based on the loss function value of at least one target object.

[0051] S605. Adjust the parameters of the preset target detection model according to the target loss function value to obtain a multi-target detection model.

[0052] In this embodiment, before training the model, the image dataset is randomly divided into a training dataset, a validation dataset, and a test dataset according to a preset ratio, wherein the preset ratio may be 7:2:1.

[0053] It should be noted that in the loss function calculation part of model training, the existing target detection model only uses the same set of anchor frames to screen the real target, i.e., the target object, determine the positive sample image, and calculate the loss function value based on the screened positive sample image and the predicted position of the target object predicted by the preset target detection model.

[0054] However, when there is at least one target object in the training data set of the image data set, and the target size and shape of at least one target object are quite different, some target objects are larger in size, some target objects are smaller in size, some target objects are approximately rectangular in shape, and some target objects are approximately square in shape, the same set of anchor frames is used to screen the real targets, and the positive sample images obtained are not accurate enough, which affects the loss value and thus affects the detection accuracy of the model training.

[0055] Therefore, according to the anchor frame of at least one type of target object, the positive sample image of at least one target object is screened from the image dataset, and then the preset target detection model is trained to obtain the predicted position of the positive sample image for at least one target object. According to the category label data of at least one target object in the positive sample image and the predicted position of at least one target object, the preset attention boundary loss function (Weighted IOU, WIOU) is used to calculate the loss function value of at least one target object.

[0056] Optionally, shape matching is performed on at least one target object in each sample image according to the anchor frames of at least one type of target object to obtain a shape matching degree of at least one target object in each sample image.

[0057] If the shape matching degree of at least one target object is less than a preset threshold, each sample image is determined to be a positive sample image of at least one target object.

[0058] Specifically, the anchor frame of the target object is determined based on the target object's label type. For example, if the target object is a ship, shape matching is performed on the ship object in each sample image based on the ship's anchor frame. This involves calculating the width ratio of the ship's target frame to the width of the anchor frame, and the height ratio of the ship's target frame to the height of the anchor frame. The shape matching degree includes the maximum value of the width ratio or the height ratio. If the shape matching degree of the ship object in a sample image is less than a preset threshold, the sample image is determined to be a positive sample image of the ship object.

[0059] Similarly, if the target object is a floating object, shape matching is performed on the floating object in each sample image based on the floating object's anchor frame. This involves calculating the width ratio of the floating object's target frame to the width of the anchor frame, and the height ratio of the floating object's target frame to the height of the anchor frame. The shape matching degree includes the maximum value of the width ratio or the height ratio. If the shape matching degree of the floating object in a sample image is less than a preset threshold, the sample image is determined to be a positive sample image of the floating object.

[0060] According to the weight parameters corresponding to at least one target object and the loss function value of at least one target object, the target loss function value is calculated, and the preset target detection model is adjusted according to the target loss function value. The preset target detection model after parameter adjustment is continued to be trained until the preset number of training times is reached, and the training is stopped to obtain a multi-target detection model.

[0061] In the method provided in the embodiment of the present application, based on the anchor frame of at least one type of target object, a positive sample image of at least one target object is screened from an image dataset; a preset target detection model is used to perform target detection on the positive sample image to obtain the predicted position of the positive sample image for at least one target object; based on the category label data of at least one target object in the positive sample image and the predicted position of at least one target object, a preset attention boundary loss function is used to calculate the loss function value of at least one target object respectively; based on the loss function value of at least one target object, a target loss function value is calculated; based on the target loss function value, the preset target detection model is parameterized to obtain a multi-target detection model. By using the anchor frame of at least one type of target object to screen the positive sample image of at least one target object from an image dataset, the accuracy of determining the positive sample image can be improved, thereby improving the convergence speed of model training and the overall detection accuracy.

[0062] An embodiment of the present application also provides another possible implementation method of a multi-target detection model training method, wherein the target detection model includes: a feature extraction module, a feature fusion module and an output module, wherein the feature extraction module is used to extract features from the input sample image, the feature fusion module is used to fuse the extracted image features, and the output module is used to output the target detection results based on the fused features.

[0063] The feature extraction module includes: a two-dimensional convolution unit, a two-dimensional normalization unit, and a linear activation function unit.

[0064] In this embodiment, the target detection model is a model with an anchor box detection function. For example, the target detection model can be a target detection algorithm YOLOv5 based on a lightweight convolutional neural network (CNN). The YOLOv5 model includes: a feature extraction module Backbone, a feature fusion module Neck, and an output module Head. The main function of the Backbone module is to extract the features of the input image. The Neck module is located between the Backbone module and the Head module. Its main function is to integrate feature maps of different levels to improve detection performance. The Head module of the YOLOv5 model consists of three different output layers, which are responsible for detecting large, medium, and small-scale targets respectively. These output layers perform confidence calculation and bounding box regression on each pixel in the feature map through a preset prior box, and finally output a multidimensional array including object category, category confidence, box coordinates, width, and height information.

[0065] Among them, the Backbone module includes convolution layers and pooling layers. The convolution layer Conv includes a two-dimensional convolution unit Conv2d, a two-dimensional normalization unit (Batch Normalization, BatchNorm2d) and a linear activation function unit (RectifiedLinear Unit, ReLU).

[0066] In the method provided in the embodiment of the present application, an activation function (Sigmoid-Weighted Linear Unit, SiLU) is used in the convolutional layer of the existing YOLOv5 model. After replacing the activation function SiLU with the activation function ReLU, the inference speed is greatly improved while ensuring that the detection accuracy and recall rate are consistent with those before the replacement.

[0067] The present application also provides another possible implementation of a multi-target detection model training method. Model training is performed based on anchor frames and image datasets of at least one type of target object. Before obtaining the multi-target detection model, the method further includes at least one of the following image preprocessing operations: Each sample image in the image dataset is processed with image color channel adjustment, image enhancement, noise processing and image space transformation.

[0068] Among them, image enhancement processing includes: image blur processing, limited brightness and contrast adjustment processing, image brightness and contrast adjustment processing, image gamma correction processing, image fogging processing, image color saturation processing, noise processing includes: noise addition processing and noise filtering processing, image space transformation processing includes: image horizontal flip processing, cropping processing, and scaling processing.

[0069] In this embodiment, the image color channel adjustment processing is to convert the RGB image into a BGR image, convert the RGB image into a grayscale image, and randomly adjust the overall brightness values ​​of the three RGB channel images. The image color channel adjustment parameters include: the R channel change parameter is 20, the G channel change parameter is 20, the B channel change parameter is 20, and the probability value is 0.01.

[0070] Image blurring is to perform mosaic blurring on the sample image through image mosaic according to a certain probability, and enhance the model's detection of the mosaic blurred image, where the probability value is 0.01.

[0071] The restricted contrast adjustment process is to restrict the contrast of the adaptive histogram equalization enhancement model to target detection in the equalized image according to a certain probability, wherein the probability value is 0.01.

[0072] The image brightness and contrast adjustment process is to enhance the image contrast by adjusting the corresponding parameters according to a certain probability. The adjustment parameters include: a brightness (brightness_limit) parameter value of 0.2, a contrast (contrast_limit) parameter value of 0.3, and a probability of 0.01.

[0073] Image gamma correction processing is to enhance the image using a gamma correction algorithm according to a certain probability, and the probability can be set to 0.01.

[0074] Image fogging is to add random intensity fog within a specified range to the image with a certain probability, so as to improve the target detection effect in foggy scenes. The fogging parameters include: the fog coefficient low threshold (fog_coef_lower) is 0.1, the fog coefficient high threshold (fog_coef_upper) is 0.5, the fog circle coefficient (alpha_coef) is 0.08, and the probability is 0.01.

[0075] Image color saturation processing is to enhance the image according to the hue saturation algorithm. The color saturation parameters include: hue parameter (hue_shift_limit) value is 20, saturation (sat_shift_limit) parameter value is 30, brightness (val_shift_limit) parameter value is 20, and probability is 0.01.

[0076] Noise processing includes adding noise and filtering out noise, both of which are used to expand the training data set.

[0077] Noise processing mainly includes the addition of ISO noise and Gaussian noise. After training with images with added noise, the model's detection performance for images containing ISO noise and Gaussian noise is greatly improved; the Gaussian noise probability is 0.01.

[0078] The noise filtering process mainly includes two strategies: median filtering and Gaussian filtering to filter noise. The probability value of the median filtering algorithm is 0.01.

[0079] It should be noted that the above parameter values ​​and probability values ​​can be adjusted according to preprocessing requirements and are not limited here.

[0080] Image space transformation processing mainly includes the following aspects: image rotation simulates the state of the target in the image when the camera rotates, so as to improve the target detection performance when the camera rotates; image horizontal flipping, cropping, and scaling are all expansions of the image dataset; image copy and paste is to select small target areas from the original image, randomly copy and paste them to other appropriate locations, and increase the number and distribution diversity of small targets in the dataset.

[0081] The present application also provides a possible implementation of a multi-target detection method. Figure 5 A flowchart of a multi-target detection method provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the method includes: S501: Acquire an image to be detected of a preset scene.

[0082] S502: Based on the image to be detected in the preset scene, a multi-target detection model is used to perform detection to obtain a target detection result of the image to be detected.

[0083] The target detection result includes: detection information of at least one target object in the image to be detected.

[0084] In this embodiment, an image to be detected of a preset scene is obtained from an acquisition device interface, and the image to be detected is input into a pre-trained multi-target detection model, so that the multi-target detection model detects the image to be detected and obtains a target detection result of the image to be detected. If the image to be detected includes at least one target object, the target detection result includes detection information of at least one target object.

[0085] In summary, an embodiment of the present application provides a multi-target detection method, which includes: obtaining an image to be detected of a preset scene; using a multi-target detection model to perform detection based on the image to be detected of the preset scene, and obtaining a target detection result of the image to be detected, the target detection result including: detection information of the image to be detected for at least one target object. Since the multi-target detection model is a model obtained by model training based on the anchor frame and image data set of at least one type of target object, the multi-target detection model can accurately predict at least one target object in the image to be detected.

[0086] The following continues to explain the multi-target detection model training device, multi-target detection device, and computer equipment provided by any of the above embodiments of the present application. The specific implementation process and the technical effects produced are the same as those of the corresponding method embodiments mentioned above. For the sake of brief description, for the parts not mentioned in this embodiment, please refer to the corresponding content in the method embodiment.

[0087] Figure 6 This is a functional module diagram of a multi-target detection model training device provided in an embodiment of the present application. Figure 6 As shown, the multi-target detection model training device 100 includes: A first acquisition module 110 is configured to acquire an image dataset, wherein each sample image in the image dataset has category label data of at least one target object; A generating module 120, configured to generate an anchor frame of at least one type of target object based on category label data of at least one target object in a plurality of sample images in an image dataset; The training module 130 is used to perform model training based on the anchor frames and image datasets of at least one type of target object to obtain a multi-target detection model. The multi-target detection model is used to perform multi-target detection based on the images to be detected in a preset scene.

[0088] Optionally, the first acquisition module 110 is also used to obtain multiple images with at least one target object from a preset image dataset; mark at least one target object in each image separately to obtain a sample image with category label data of at least one target object; and construct an image dataset based on multiple sample images.

[0089] Optionally, the generation module 120 is further used to generate a label file of at least one target object based on the category label data of at least one target object in multiple sample images; and generate an anchor frame of at least one category of target objects using a preset anchor frame clustering function based on the label file of at least one target object.

[0090] Optionally, the training module 130 is further used to screen positive sample images of at least one target object from an image dataset based on an anchor frame of at least one type of target object; use a preset target detection model to perform target detection on the positive sample image to obtain a predicted position of the positive sample image for at least one target object; use a preset attention boundary loss function to calculate the loss function value of at least one target object based on the category label data of at least one target object in the positive sample image and the predicted position of at least one target object; calculate the target loss function value based on the loss function value of at least one target object; and adjust the parameters of the preset target detection model based on the target loss function value to obtain a multi-target detection model.

[0091] Optionally, the training module 130 is further used to perform shape matching on at least one target object in each sample image based on the anchor frame of at least one type of target object, and obtain the shape matching degree of at least one target object in each sample image; if the shape matching degree of at least one target object is less than a preset threshold, each sample image is determined to be a positive sample image of at least one target object.

[0092] Optionally, the target detection model includes: a feature extraction module, a feature fusion module and an output module, wherein the feature extraction module is used to extract features from the input sample image, the feature fusion module is used to fuse the extracted image features, and the output module is used to output the target detection results based on the fused features; the feature extraction module includes: a two-dimensional convolution unit, a two-dimensional normalization unit and a linear activation function unit.

[0093] Optionally, the apparatus further includes at least one of the following image preprocessing operations: Each sample image in the image data set is subjected to image color channel adjustment processing, image enhancement processing, noise processing and image space transformation processing; among which, image enhancement processing includes: image blur processing, limited contrast adjustment processing, image brightness contrast adjustment processing, image gamma correction processing, image fogging processing, image color saturation processing, noise processing includes: noise addition processing and noise filtering processing, image space transformation processing includes: image horizontal flip processing, cropping processing, and scaling processing.

[0094] Figure 7 This is a functional module diagram of a multi-target detection device provided in an embodiment of the present application. Figure 7 As shown, the multi-target detection device 200 includes: The second acquisition module 210 is used to acquire an image to be detected of a preset scene; The detection module 220 is used to perform detection using a multi-target detection model based on the image to be detected in a preset scene to obtain a target detection result of the image to be detected. The target detection result includes: detection information of the image to be detected for at least one target object.

[0095] The above-mentioned device is used to execute the method provided in the above-mentioned embodiment. Its implementation principle and technical effect are similar and will not be repeated here.

[0096] The above modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more microprocessors, or one or more field programmable gate arrays (FPGAs). For example, when a module is implemented by scheduling program code through a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0097] Figure 8 This is a schematic diagram of a computer device provided in an embodiment of the present application, which can be used for multi-target detection model training or multi-target detection. Figure 8 As shown, the computer device includes: a processor 310 , a storage medium 320 , and a bus 330 .

[0098] Storage medium 320 stores machine-readable instructions executable by processor 310. When the computer device is running, processor 310 communicates with storage medium 320 via bus 330, and processor 310 executes the machine-readable instructions to perform the steps of the above method embodiment. The specific implementation methods and technical effects are similar and will not be repeated here.

[0099] Optionally, the present application further provides a storage medium 320 on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method embodiment are executed. The specific implementation and technical effects are similar and will not be repeated here.

[0100] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0101] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0102] In addition, the functional units in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional units.

[0103] The aforementioned integrated unit implemented as a software functional unit can be stored in a computer-readable storage medium. The software functional unit, stored in a storage medium, includes instructions for causing a computer device (which may be a personal computer, server, or network device, etc.) or a processor to execute portions of the method steps described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a removable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0104] The above are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A multi-target detection model training method, characterized in that: The method comprises: Acquire an image dataset, wherein each sample image in the image dataset has category label data of at least one target object; generating an anchor frame of at least one type of target object according to category label data of at least one target object in a plurality of sample images in the image dataset; Model training is performed based on the anchor frames of the at least one type of target object and the image dataset to obtain a multi-target detection model, which is used to perform multi-target detection based on the image to be detected in a preset scene.

2. The method according to claim 1, wherein The acquiring of the image data set comprises: Acquire a plurality of images having the at least one target object from a preset image dataset; Annotating the at least one target object in each image to obtain a sample image having category label data of the at least one target object; The image dataset is constructed based on the multiple sample images.

3. The method according to claim 1, wherein The generating, based on the category label data of the at least one target object in the plurality of sample images in the image dataset, an anchor frame of at least one category of target objects comprises: generating a label file for the at least one target object according to the category label data of the at least one target object in the plurality of sample images; According to the label file of the at least one target object, a preset anchor frame clustering function is used to generate anchor frames of the at least one type of target object.

4. The method according to claim 1, wherein The performing model training based on the anchor frames of the at least one type of target object and the image dataset to obtain a multi-target detection model includes: Filtering positive sample images of the at least one target object from the image dataset according to the anchor frames of the at least one type of target object; Using a preset target detection model, perform target detection on the positive sample image to obtain a predicted position of the positive sample image for the at least one target object; Calculating a loss function value of the at least one target object using a preset attention boundary loss function according to the category label data of the at least one target object in the positive sample image and the predicted position of the at least one target object; Calculating a target loss function value based on the loss function value of the at least one target object; According to the target loss function value, the preset target detection model is adjusted to obtain the multi-target detection model.

5. The method according to claim 4, wherein The step of screening the positive sample images of the at least one target object from the image dataset according to the anchor frames of the at least one type of target object comprises: performing shape matching on the at least one target object in each sample image according to the anchor frame of the at least one type of target object, to obtain a shape matching degree of the at least one target object in each sample image; If the shape matching degree of the at least one target object is less than a preset threshold, each of the sample images is determined to be a positive sample image of the at least one target object.

6. The method according to claim 4, wherein The target detection model includes: a feature extraction module, a feature fusion module and an output module, wherein the feature extraction module is used to extract features from the input sample image, the feature fusion module is used to fuse the extracted image features, and the output module is used to output the target detection result based on the fused features; The feature extraction module includes: a two-dimensional convolution unit, a two-dimensional normalization unit and a linear activation function unit.

7. The method according to claim 1, wherein Before performing model training based on the anchor boxes of the at least one type of target object and the image dataset to obtain a multi-target detection model, the method further includes at least one of the following image preprocessing operations: Each sample image in the image data set is subjected to image color channel adjustment processing, image enhancement processing, noise processing, and image space transformation processing; wherein the image enhancement processing includes: image blur processing, limited contrast adjustment processing, image brightness contrast adjustment processing, image gamma correction processing, image fogging processing, and image color saturation processing; the noise processing includes: noise addition processing and noise filtering processing; and the image space transformation processing includes: image horizontal flip processing, cropping processing, and scaling processing.

8. A multi-target detection method, characterized in that: The method comprises: Obtain the image to be detected of the preset scene; According to the image to be detected in the preset scene, a multi-target detection model is used to perform detection to obtain a target detection result of the image to be detected, wherein the target detection result includes: detection information of the image to be detected for at least one target object, and the multi-target detection model is the multi-target detection model described in any one of claims 1 to 7 above.

9. A computer device, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores program instructions executable by the processor. When the electronic device is running, the processor and the storage medium communicate through the bus, and the processor executes the program instructions to perform the steps of the multi-target detection model training method described in any one of claims 1 to 7, or perform the steps of the multi-target detection method according to claim 8.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed by a processor, executes the steps of the multi-target detection model training method according to any one of claims 1 to 7, or executes the steps of the multi-target detection method according to claim 8.