Target detection model training method and device, target detection method and device and electronic equipment
By preprocessing the video dataset and training with the YOLOv5 model with improved loss function, the shortcomings of YOLOv5 in low resolution and small object detection are solved, and the model training efficiency and recognition accuracy are improved.
Patent Information
- Application Number
- CN202510098188.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
YOLOv5 has insufficient ability to handle low resolution and small object detection, resulting in low recognition accuracy and excessive waste of resources when data quality is low, affecting model training and recognition performance.
The image dataset is generated by preprocessing the video dataset and inputting it into a preset YOLOv5 model with improved loss function for training. The improved loss function includes IOU loss, center point loss, and length and width loss.
The efficiency and recognition accuracy of object detection model training are improved, ensuring that even if the predicted target box does not overlap with the anchor box, the loss function can provide effective gradients, promote stable training of the model and better positioning accuracy.
Smart Images

Figure CN120032332A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target recognition desensitization technology, and in particular to a target detection model training method, a target detection method, a device and an electronic device. Background Art
[0002] With the development of autonomous driving technology, vehicle systems increasingly rely on various sensors to collect data about the surrounding environment. These data contain a lot of personal privacy information, such as faces and license plates, which urgently need to be effectively protected through target recognition and data desensitization technology. In the field of data desensitization, efficient target detection models such as YOLOv5 (You Only Look Once version 5) usually use an anchor box mechanism to predict the bounding box of the target, and can perform target detection and classification simultaneously in a single forward propagation. As an advanced framework for deep learning target detection, it has significant advantages in the field of target detection compared to other target detection algorithms, which makes it possible to achieve real-time desensitization of vehicle data.
[0003] When training the model, YOLOv5 is slightly lacking in the ability to handle low resolution and small target detection, which will lead to low algorithm recognition accuracy. When the data quality is poor, limited resources are used to process large amounts of dirty data, which will affect model training, have little effect on improving model recognition performance, and result in excessive waste of resources. Summary of the invention
[0004] The problem solved by the present invention is how to improve the training efficiency and recognition accuracy of the target detection model.
[0005] To solve the above problems, the present invention provides a target detection model training method, a target detection method, a device and an electronic device.
[0006] In a first aspect, the present invention provides a method for training a target detection model, comprising:
[0007] Preprocess the video dataset to generate an image dataset;
[0008] The image data set is input into a preset YOLOv5 model for training to establish a target detection model, wherein the preset YOLOv5 model adopts an improved loss function, and the improved loss function includes IOU loss, center point loss and length and width loss.
[0009] Optionally, the preprocessing of the video data set includes:
[0010] Segmenting and cutting the video data set into frames to generate an initial image set;
[0011] A confidence calculation on data quality is performed on the initial image set, and images with data quality lower than a preset quality standard are filtered out from the initial image set according to the confidence calculation result to generate the image data set.
[0012] Optionally, after preprocessing the video data set, the method further includes:
[0013] Dividing the image dataset into a training set, a validation set, and a test set;
[0014] The Mosaic method is used to perform data enhancement on the image data in the training set.
[0015] Optionally, inputting the image data set into a preset YOLOv5 model for training includes:
[0016] Setting training parameters, and performing model training on a preset YOLOv5 model using the training set and the validation set based on the training parameters;
[0017] After the training is completed, the test set is used to perform a model test on the trained preset YOLOv5 model to establish the target detection model.
[0018] Optionally, the process of constructing the improved loss function includes:
[0019] Construct IOU loss function;
[0020] Determine the center point distance according to the distance between the predicted box and the real box of the preset YOLOv5 model, and determine the center point loss function according to the center point distance and the first hyperparameter;
[0021] Determine a length-width loss function according to a width difference, a height difference, and a second hyperparameter, wherein the width difference represents a relative difference in width between the predicted box and the true box, and the height difference represents a relative difference in height between the predicted box and the true box;
[0022] The improved loss function is constructed according to the IOU loss function, the center point loss function and the length-width loss function.
[0023] In a second aspect, the present invention provides a target detection method, comprising:
[0024] The target detection model established by the target detection model training method is used to perform target detection on the collected video frame by frame to determine the sensitive data information in each frame image of the video;
[0025] After erasing the sensitive data information in the image of each frame, the image of each frame is spliced to generate desensitized video information, and the desensitized video information is output.
[0026] Optionally, the output spliced desensitized video information includes:
[0027] KAFKA is used as the message transmission middleware to distribute and upload the desensitized video information to the cloud.
[0028] In a third aspect, the present invention provides a target detection model training device, comprising:
[0029] The first module is used to preprocess the video data set to generate an image data set;
[0030] The second module is used to input the image data set into a preset YOLOv5 model for training to establish a target detection model, wherein the preset YOLOv5 model adopts an improved loss function, and the improved loss function includes IOU loss, center point loss and length and width loss.
[0031] In a fourth aspect, the present invention provides a target detection device, comprising:
[0032] The third module is used to use the target detection model established by the target detection model training method to perform target detection on the collected video frame by frame to determine the sensitive data information in each frame image of the video;
[0033] The fourth module is used to erase the sensitive data information in the image of each frame, splice the image of each frame to generate desensitized video information, and output the desensitized video information.
[0034] In a fifth aspect, the present invention provides an electronic device, comprising a memory and a processor;
[0035] The memory is used to store computer programs;
[0036] The processor is used to implement the target detection model training method as described in the first aspect or the target detection method as described in the second aspect when executing the computer program.
[0037] In a sixth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the target detection model training method as described in the first aspect or the target detection method as described in the second aspect is implemented.
[0038] The object detection model training method of the present invention has the beneficial effects of: pre-processing a video data set to generate an image data set, which is used as a data set for training a YOLOv5 model; by inputting the image data set into a preset YOLOv5 model using an improved loss function for training, the improved loss function used by the preset YOLOv5 model introduces a center point loss and a length and width loss, so that even if the target box predicted by the YOLOv5 model and the two box areas of the anchor box in the image data set do not overlap, the loss function can also provide an effective gradient compared to the traditional YOLOv5 model, for example, the length and width loss can minimize the difference between the width and height of the target box predicted by the preset YOLOv5 model and the anchor box in the image data set, thereby making the convergence speed of the model training faster. In addition, the introduction of the center point loss and the length and width loss comprehensively considers the matching of the position and size, and has a better positioning effect for the target. Therefore, the present invention has a more stable training process and better positioning accuracy, and can improve the training efficiency and recognition accuracy of the object detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic diagram of a process of training a target detection model according to an embodiment of the present invention;
[0040] Figure 2 A schematic diagram of a process of preprocessing a video data set according to an embodiment of the present invention;
[0041] Figure 3 A schematic diagram of a process of performing data enhancement on image data according to an embodiment of the present invention;
[0042] Figure 4 A schematic diagram of the model training and testing process of an embodiment of the present invention;
[0043] Figure 5 A schematic diagram of a process for constructing an improved loss function according to an embodiment of the present invention;
[0044] Figure 6 A system architecture diagram of a target detection model training device according to an embodiment of the present invention;
[0045] Figure 7 is a system architecture diagram of a target detection device according to an embodiment of the present invention;
[0046] Figure 8 4 is a system architecture diagram of an electronic device according to an embodiment of the present invention.
[0047] Description of reference numerals:
[0048] 600 - target detection model training device, 610 - first module, 620 - second module, 700 - target detection device, 730 - third module, 740 - fourth module, 800 - electronic device, 810 - processor, 820 - memory. DETAILED DESCRIPTION
[0049] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.
[0050] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0051] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to"; the term "based on" means "based at least in part on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0052] It should be noted that the modifications of "one" and "plurality" mentioned in the present invention are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0053] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes, and are not used to limit the scope of these messages or information.
[0054] like Figure 1 As shown, an object detection model training method provided by an embodiment of the present invention includes:
[0055] S100: Preprocess the video dataset to generate an image dataset.
[0056] Specifically, different lighting, weather conditions and dynamic environments will affect the quality of data and the accuracy of the desensitization algorithm, and in order to ensure the smooth training of the vehicle target detection model, it is often necessary to collect a large number of video data sets, and then cut frames based on the large number of video data sets, and finally provide them to the target detection model for algorithm training. In the process, a large amount of dirty data and redundant data will be generated, which will take up too much disk space. Therefore, the video data set can be preprocessed, such as using distributed technology to batch intercept the collected videos and perform concurrent frame cutting to generate an image data set. Preprocessing can reduce the resource pressure in the data collection stage, reduce the computing burden of the server, and improve data transmission efficiency.
[0057] S200: Inputting the image data set into a preset YOLOv5 model for training to establish a target detection model, wherein the preset YOLOv5 model adopts an improved loss function, and the improved loss function includes IOU loss, center point loss and length-width loss.
[0058] Specifically, since the customized training of the traditional YOLOv5 model for specific scenarios and sensitive information still requires a large amount of labeled data and adjustment and optimization, it may be time-consuming and labor-intensive in actual operation. Therefore, the preset YOLOv5 model of this embodiment adopts an improved loss function, that is, the loss function used by the YOLOv5 target detection algorithm by default is improved. The default loss function based on IOU (Intersection over Union) such as CIOU (Complete IOU) and GIOU (Generalized IOU) is slightly insufficient in measuring the gap between the target box and the anchor point, resulting in a slow convergence speed of the BBR (Bounding Box Regression) model optimization during training, thereby affecting the positioning accuracy; therefore, this embodiment adopts the EIOU (Efficient-IOU) loss function, which can independently calculate the length and width of the predicted box and the true box based on the penalty term of CIOU, that is, separate the influencing factors of the aspect ratio of the predicted box and the true box, and add Focal Loss to solve the sample imbalance problem in BBox (Bounding Box) regression. The exemplary EIOU loss function is expressed as Formula 1:
[0059]
[0060] Among them, L EIOU is the EIOU loss function, L IOU is the IOU loss function, L dis is the center point loss, L asp is the length and width loss, ρ 2 (b,b gt) represents the Euclidean distance between the center point of the predicted box and the true box, ρ 2 (w,w gt ),ρ 2 (h,h gt ) represent the Euclidean distance of the width and height of the predicted box and the real box respectively (representing the difference between the predicted box and the real box in the width dimension and the height dimension respectively), c represents the diagonal distance of the minimum closed area containing both the predicted box and the real box, w c Indicates the width of the minimum enclosed area that contains both the predicted box and the true box, w h Indicates the height of the minimum enclosed area that contains both the predicted box and the true box.
[0061] Among them, IOU (Intersection over Union) is often used to evaluate the performance of target detection models, such as calculating the intersection size between the target box and the anchor box (between A and B). However, the traditional IOU has some limitations, especially when the target box and the anchor box have no intersection, the IOU value is 0, which makes it impossible to update the calculation model gradient and make it difficult to further learn and optimize the model. The loss function for IOU is expressed as follows:
[0062]
[0063] Among them, the center point distance is the Euclidean distance between the center points of the predicted box and the real box; let the center of the target box be (x p ,y p ), the center of the anchor box is (x g ,y g ), the center distance is calculated as:
[0064]
[0065] Among them, the width difference w diff and height difference h diff They are the relative differences in width and height between the predicted box and the real box respectively; the calculation method can be a simple difference or proportional difference.
[0066] EIOU integrates the above metrics into a loss function. Compared with the original IOU, it not only considers the overlapping area between the target box and the anchor box, but also introduces other metrics, such as the distance between the center points of the two boxes and the relative difference in width and height. This design allows the loss function to provide an effective gradient even if the two box areas do not overlap, thereby promoting better training and convergence of the model. The usual form is formula 2:
[0067] L EIOU =1-IOU+λ 1 D c +λ 2 (wdiff +h diff );
[0068] Among them, formula 2 is another form of formula 1, λ 1 and λ 2 They represent the hyperparameters for adjusting the center distance and aspect ratio difference between the predicted box and the true box. The value range of the hyperparameters is (0,1).
[0069] Since the EIOU loss function retains the advantages of CIOU loss and directly minimizes the difference between the width and height of the target box and the anchor box, it converges faster and has better positioning effect. Compared with other IOU loss functions, EIOU also has a more stable training process. By introducing penalty terms for center point distance and aspect ratio, it can ensure that effective gradient information is provided even when IOU is 0, ensuring that the model can continue to learn. It also has certain advantages in positioning accuracy. By comprehensively considering the matching of position and size, EIOU can significantly improve the positioning accuracy of the target detection model, especially when the size and shape of the target object vary greatly.
[0070] In this embodiment, the video data set is preprocessed to generate an image data set, which is used as a data set for training the YOLOv5 model; the image data set is input into a preset YOLOv5 model using an improved loss function for training, and the improved loss function used by the preset YOLOv5 model introduces a center point loss and a length and width loss. Compared with the traditional YOLOv5 model, even if the target box predicted by the YOLOv5 model and the two box areas of the anchor box in the image data set do not overlap, the loss function can also provide an effective gradient. For example, the length and width loss can minimize the difference between the width and height of the target box predicted by the preset YOLOv5 model and the anchor box in the image data set, thereby making the convergence speed of the model training faster. In addition, the introduction of the center point loss and the length and width loss comprehensively considers the matching of the position and size, and has a better positioning effect for the target. Therefore, the present invention has a more stable training process and better positioning accuracy, and can improve the training efficiency and recognition accuracy of the target detection model.
[0071] Optionally, the preprocessing of the video data set includes:
[0072] S110: segment and cut the video data set into frames to generate an initial image set.
[0073] Specifically, combined Figure 2 As shown, for the original video collected in the data collection stage, a pre-written video frame cutting program can be used to segment and cut the collected video, for example, cutting it into images at a frame rate of 3 seconds.
[0074] S120: performing confidence calculation on data quality of the initial image set, and filtering images with data quality lower than a preset quality standard from the initial image set according to the confidence calculation result, so as to generate the image data set.
[0075] Specifically, combined Figure 2 As shown in the figure, the data is annotated while cutting the frame, and the corresponding annotation result file is output. The confidence of the data quality is calculated. If the specified confidence is reached, it is included in the data set and transmitted downward. If it is lower than the confidence, it is not transmitted, so that images with poor data quality can be filtered out. In the data desensitization process, the data set required for the early stage of model training contains a large amount of redundant data and dirty data, which leads to resource problems such as low data processing capacity and excessive disk space occupation. When encountering bad weather or network fluctuations, the quality of data obtained by the vehicle is too low. If the data is sent to the cloud as a training data set according to the data volume at this time, it will lead to excessive waste of training resources. Through the above preprocessing process, the resource pressure in the data collection stage can be reduced, the computing burden of the server can be reduced, and the data transmission efficiency can be improved.
[0076] The filtering standard determines the accuracy of the model's target recognition. If the filtering standard is too high, the model may oscillate around the optimal solution; if the learning rate (which determines the step size of the parameter update in each iteration) is too small, the model may converge slowly, so the standard setting needs to be adjusted according to the specific situation.
[0077] In this optional embodiment, the video data set is segmented and frame-cut to generate an initial image set, and then images with poor data quality are filtered out from the initial image set, thereby effectively improving the accuracy of model recognition and reducing the computing burden on the server.
[0078] Optionally, after preprocessing the video data set, the method further includes:
[0079] S130: Divide the image dataset into a training set, a validation set, and a test set.
[0080] Specifically, combined Figure 3 As shown, the images preprocessed in the acquisition phase and the corresponding annotation result data can be randomly divided into a training set, a validation set, and a test set in a ratio of 6:2:2.
[0081] S140: Using the Mosaic method to perform data enhancement on the image data in the training set.
[0082] Specifically, combined Figure 3As shown in the figure, by using the Mosaic method for data enhancement, four pictures (or more pictures, depending on the required enhancement effect) can be randomly selected from the training set. These pictures can be from the same category or from different categories, and their sizes can be different. However, in order to ensure image alignment during stitching, the images can be cropped or scaled. For example, each selected picture can be scaled to make its size suitable for the resolution of the final stitched image. In order to ensure the uniform size of the stitched image, the four pictures can be scaled to the same size, or a fixed-size window can be pre-set to crop an area from each picture. Since the position and size of the target in the stitched image will change, the target annotation box needs to be adjusted accordingly. The Mosaic method can generate more diverse training samples, compress the data volume, and enhance the data quality strength, so that the model can extract more scene features and patterns during the training process, and enhance the generalization ability and robustness of the target detection model. The model improved by the Mosaic data enhancement technology has a significant effect on improving the ability to recognize complex scenes, and the recognition accuracy and target detection performance have been greatly improved.
[0083] In this optional embodiment, the Mosaic method is used to perform data enhancement on the image data in the training set, so that the model can extract features and patterns of more scenes during the training process, thereby enhancing the generalization ability and robustness of the target detection model.
[0084] Optionally, inputting the image data set into a preset YOLOv5 model for training includes:
[0085] S210: Setting training parameters, and performing model training on a preset YOLOv5 model using the training set and the validation set based on the training parameters.
[0086] Specifically, combined Figure 4 As shown in the figure, during model training, the input image size can be set to 640*640, the batch size can be set to 16, the number of training iterations can be set to 500, the initial learning rate can be set to 0.01, the learning rate momentum can be set to 0.937, and the weight decay coefficient can be set to 0.0005. After the training, the model can be verified on the verification set.
[0087] S220: After the training is completed, the trained preset YOLOv5 model is tested using the test set to establish the target detection model.
[0088] Specifically, combined Figure 4 As shown in Figure 1, after the training is completed, the test set is used to test the model, and finally a further optimized improved YOLOv5 model, namely the target detection model, is obtained.
[0089] In this optional embodiment, the model is trained using a training set and a validation set, and then the model is tested using a test set after the training is completed, which effectively improves the accuracy of the target detection model.
[0090] Optionally, the process of constructing the improved loss function includes:
[0091] S201: Construct an IOU loss function.
[0092] Specifically, combined Figure 5 As shown above, the loss function of IOU can be expressed as follows:
[0093]
[0094] S202: Determine a center point distance according to the distance between the predicted box and the real box of the preset YOLOv5 model, and determine a center point loss function according to the center point distance and a first hyperparameter.
[0095] Specifically, as mentioned above, the center point distance can be expressed as:
[0096]
[0097] The center point loss function can be expressed as λ 1 D c .
[0098] S203: Determine a length-width loss function according to a width difference, a height difference, and a second hyperparameter, wherein the width difference represents a relative difference in width between the predicted box and the true box, and the height difference represents a relative difference in height between the predicted box and the true box.
[0099] Specifically, as mentioned above, the length-width loss function can be expressed as λ 2 (w diff +h diff ).
[0100] S204: Constructing the improved loss function according to the IOU loss function, the center point loss function and the length-width loss function.
[0101] Specifically, as mentioned above, the improved loss function can be expressed as:
[0102] L EIOU =1-L IOU +λ 1 D c +λ 2 (w diff +h diff );
[0103] In this optional embodiment, an improved loss function is constructed based on the IOU loss function, the center point loss function and the length and width loss function, which can improve the positioning accuracy of the target detection model.
[0104] An embodiment of the present invention provides a target detection method, comprising:
[0105] The target detection model established by the target detection model training method is used to perform target detection on the collected video frame by frame to determine the sensitive data information in each frame image of the video;
[0106] After erasing the sensitive data information in the image of each frame, the image of each frame is spliced to generate desensitized video information, and the desensitized video information is output.
[0107] Specifically, the improved YOLOv5 model is added to the data desensitization project as a detection model to perform target recognition on the collected video frame by frame. After the algorithm detects sensitive data in the current image, it will mark the result area. Then, the marked layer can use the compiled grayscale processing function to erase the corresponding sensitive data information in the layer, and output the current frame data. After the detection is completed, all data are spliced into video information and output to the configured desensitization result directory; the comparative experimental results of the improved YOLOv5 model and the traditional YOLOv5 are shown in Table 1 below:
[0108] Table 1 Comparative experimental results
[0109] method Parameter quantity Calculation Amount Accuracy Traditional YOLOv5 123.35M 92.6G 89.91% Improving YOLOv5 125.88M 93.7G 95.42%
[0110] In this embodiment, as can be seen from Table 1, the target detection model is used to perform target detection on the collected video, which effectively improves the accuracy of target recognition, and can erase sensitive data in the image to achieve data desensitization.
[0111] Optionally, the outputting the desensitized video information includes:
[0112] KAFKA is used as the message transmission middleware to distribute and upload the desensitized video information to the cloud.
[0113] Specifically, the data desensitization result video output in the previous stage is uploaded to the data cloud. In order to prevent data volume from being too large, insufficient upload bandwidth resources, data transmission failure, data loss, data delay disorder, etc. due to a series of unstable factors such as network fluctuations and insufficient server memory resources during data upload, data distributed upload technology can be used to transmit data, and KAFKA message queue technology is used as message transmission middleware. KAFKA can handle large amounts of data and support high-throughput message delivery, which is suitable for scenarios that require fast processing and distribution of messages. KAFKA's message queues are distributed on multiple servers, which makes it highly scalable and fault-tolerant. KAFKA stores messages on disk instead of in memory, which allows it to process large amounts of data without losing information. The existence of middleware can reduce the coupling between the sender and the receiver, ensuring that both parties can transmit and receive data at their own rates, improving data transmission efficiency and system stability.
[0114] In this optional embodiment, distributed technology is used to upload the desensitized video information to the cloud, which effectively improves data transmission efficiency and system stability.
[0115] like Figure 6 As shown, an object detection model training device 600 provided by an embodiment of the present invention includes:
[0116] The first module 610 is used to preprocess the video data set to generate an image data set;
[0117] The second module 620 is used to input the image data set into a preset YOLOv5 model for training to establish a target detection model, wherein the preset YOLOv5 model adopts an improved loss function, and the improved loss function includes IOU loss, center point loss and length and width loss.
[0118] like Figure 7 As shown, an object detection device 700 provided by an embodiment of the present invention includes:
[0119] The third module 730 is used to perform target detection on the collected video frame by frame using the target detection model established by the target detection model training method to determine sensitive data information in each frame image of the video;
[0120] The fourth module 740 is used to erase the sensitive data information in the image of each frame, splice the image of each frame to generate desensitized video information, and output the desensitized video information.
[0121] like Figure 8As shown, an electronic device 800 provided by an embodiment of the present invention includes a memory 820 and a processor 810; the memory 820 is used to store a computer program; the processor 810 is used to implement the target detection model training method or target detection method as described above when executing the computer program.
[0122] In other words, an electronic device 800 includes a memory 820 and a processor 810 coupled to the memory 820; the memory 820 is configured to store a computer program; and the processor 810 is configured to perform the following operations when executing the computer program:
[0123] Preprocess the video dataset to generate an image dataset;
[0124] The image data set is input into a preset YOLOv5 model for training to establish a target detection model, wherein the preset YOLOv5 model adopts an improved loss function, and the improved loss function includes IOU loss, center point loss and length and width loss.
[0125] Alternatively, the target detection model established by the target detection model training method is used to perform target detection on the collected video frame by frame to determine the sensitive data information in each frame image of the video;
[0126] After erasing the sensitive data information in the image of each frame, the image of each frame is spliced to generate desensitized video information, and the desensitized video information is output.
[0127] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the target detection model training method or the target detection method as described above is implemented.
[0128] In other words, a non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor performs the following operations:
[0129] Preprocess the video dataset to generate an image dataset;
[0130] The image data set is input into a preset YOLOv5 model for training to establish a target detection model, wherein the preset YOLOv5 model adopts an improved loss function, and the improved loss function includes IOU loss, center point loss and length and width loss.
[0131] Alternatively, the target detection model established by the target detection model training method is used to perform target detection on the collected video frame by frame to determine the sensitive data information in each frame image of the video;
[0132] After erasing the sensitive data information in the image of each frame, the image of each frame is spliced to generate desensitized video information, and the desensitized video information is output.
[0133] An electronic device 800 that can be used as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device 800 is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 800 can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the present invention described herein and / or required.
[0134] The electronic device 800 includes a computing unit, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The computing unit, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0135] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc. In the present application, the unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present invention. In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0136] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will fall within the protection scope of the present invention.
Claims
1. A target detection model training method, characterized in that: include: Preprocess the video dataset to generate an image dataset; The image data set is input into a preset YOLOv5 model for training to establish a target detection model, wherein the preset YOLOv5 model adopts an improved loss function, and the improved loss function includes IOU loss, center point loss and length and width loss.
2. The target detection model training method according to claim 1, characterized in that: The preprocessing of the video data set includes: Segmenting and cutting the video data set into frames to generate an initial image set; A confidence calculation on data quality is performed on the initial image set, and images with data quality lower than a preset quality standard are filtered out from the initial image set according to the confidence calculation result to generate the image data set.
3. The target detection model training method according to claim 2, characterized in that: After the video data set is preprocessed, the method further includes: Dividing the image dataset into a training set, a validation set, and a test set; The Mosaic method is used to perform data enhancement on the image data in the training set.
4. The target detection model training method according to claim 3, characterized in that: The inputting the image data set into a preset YOLOv5 model for training comprises: Setting training parameters, and performing model training on a preset YOLOv5 model using the training set and the validation set based on the training parameters; After the training is completed, the test set is used to perform a model test on the trained preset YOLOv5 model to establish the target detection model.
5. The target detection model training method according to claim 1, characterized in that: The construction process of the improved loss function includes: Construct IOU loss function; Determine the center point distance according to the distance between the predicted box and the real box of the preset YOLOv5 model, and determine the center point loss function according to the center point distance and the first hyperparameter; Determine a length-width loss function according to a width difference, a height difference, and a second hyperparameter, wherein the width difference represents a relative difference in width between the predicted box and the true box, and the height difference represents a relative difference in height between the predicted box and the true box; The improved loss function is constructed according to the IOU loss function, the center point loss function and the length-width loss function.
6. A target detection method, characterized in that: include: The target detection model established by the target detection model training method according to any one of claims 1 to 5 is used to perform target detection on the collected video frame by frame to determine the sensitive data information in each frame image of the video; After erasing the sensitive data information in the image of each frame, the image of each frame is spliced to generate desensitized video information, and the desensitized video information is output.
7. The target detection method according to claim 6, characterized in that: The outputting the desensitized video information comprises: KAFKA is used as the message transmission middleware to distribute and upload the desensitized video information to the cloud.
8. A target detection model training device, characterized in that: include: The first module is used to preprocess the video data set to generate an image data set; The second module is used to input the image data set into a preset YOLOv5 model for training to establish a target detection model, wherein the preset YOLOv5 model adopts an improved loss function, and the improved loss function includes IOU loss, center point loss and length and width loss.
9. A target detection device, characterized in that: include: A third module is used to perform target detection on the collected video frame by frame using the target detection model established by the target detection model training method according to any one of claims 1 to 5, so as to determine the sensitive data information in each frame image of the video; The fourth module is used to erase the sensitive data information in the image of each frame, splice the image of each frame to generate desensitized video information, and output the desensitized video information.
10. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the target detection model training method as described in any one of claims 1 to 5 or the target detection method as described in claim 6 or 7 when executing the computer program.
11. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, it implements the target detection model training method as described in any one of claims 1 to 5 or the target detection method as described in claim 6 or 7.
Citation Information
Cited By
Detection method, device and equipment of slagging-off plate and medium
CN121280338A