Target detection model determination method, application method and related device
By introducing differential convolution module and improved noise candidate box generation strategy into the object detection model, the problem of low object detection accuracy in low light environments is solved, and higher robustness and accuracy are achieved, providing reliable technical support for object detection in complex environments.
Patent Information
- Application Number
- CN202510356295.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-05-13
AI Technical Summary
In low-light or night environments, traditional object detection methods face the problem of insufficient image details and edge information, resulting in a significant decline in detection performance.
By introducing a differential convolution module and an improved noise candidate box generation strategy, image edge features are enhanced, and the noise candidate box is optimized through multi-scale feature extraction and non-maximum suppression, the robustness and accuracy of the object detection model are improved.
It significantly improves the target detection accuracy in night scenes, reduces the misidentification problems caused by insufficient light, and provides reliable technical support for applications such as autonomous driving and security monitoring.
Smart Images

Figure CN119992071A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of target detection, and in particular to a method for determining a target detection model, an application method and related devices. Background Art
[0002] Object detection technology plays a key role in many computer vision applications such as autonomous driving, security monitoring, and industrial inspection. However, in low-light or nighttime environments, the lack of image details and edge information poses serious challenges to traditional object detection methods. Although modern deep learning methods have made significant progress in the use of global features and edge information, this information is still insufficient to meet the needs of accurate detection under low-light conditions, resulting in a significant decrease in detection performance. To address this problem, object detection methods based on generative models, such as DiffusionDet, have gradually attracted attention. Although DiffusionDet can effectively generate potential target areas, the generated noisy candidate boxes are often not accurate enough and have a high degree of overlap in nighttime scenes, which affects the final detection accuracy. Therefore, how to enhance edge features and reduce redundant candidate boxes to improve the robustness and accuracy of detectors in nighttime environments has become a technical problem that needs to be solved urgently. Summary of the invention
[0003] The purpose of this application is to provide a method for determining a target detection model, an application method and related devices, which can improve the target recognition accuracy in night scenes and reduce the problem of misrecognition caused by insufficient light.
[0004] To achieve the above objectives, this application provides the following solutions:
[0005] In a first aspect, the present application provides a method for determining a target detection model, the method for determining a target detection model comprising:
[0006] A data set is obtained; the data set includes the positions and categories of the bounding boxes of target objects in a plurality of RGB images.
[0007] The data set is preprocessed to obtain a preprocessed data set.
[0008] Multi-scale feature extraction is performed on the preprocessed data set to obtain a plurality of multi-scale feature maps.
[0009] The edge features of the last layer of features of the plurality of multi-scale feature maps are enhanced to obtain a plurality of enhanced multi-scale feature maps.
[0010] Based on the data set, a Gaussian noise candidate box is randomly generated, and the Gaussian noise candidate box is diffused to obtain a diffused Gaussian noise candidate box.
[0011] Non-maximum suppression (NMS) is performed on the diffused Gaussian noise candidate frame to obtain an optimized Gaussian noise candidate frame.
[0012] The enhanced multi-scale feature map and the optimized Gaussian noise candidate box are used as input, and the position and category of the bounding box of the corresponding target object are used as labels to train the constructed model to obtain a target detection model; the constructed model includes: a feature extractor, a differential convolution module, a noise candidate box generator, a candidate box optimization layer and a feature decoder.
[0013] In a second aspect, the present application provides an application method of a target detection model, and the application method of the target detection model includes:
[0014] Get the image to be detected.
[0015] The image to be detected is preprocessed to obtain a preprocessed image.
[0016] Multi-scale feature extraction is performed on the preprocessed image to obtain a multi-scale feature map.
[0017] The edge features of the last layer of features of the multi-scale feature map are enhanced to obtain an enhanced multi-scale feature map.
[0018] Based on the image to be detected, a Gaussian noise candidate frame is randomly generated, and the Gaussian noise candidate frame is diffused to obtain a diffused Gaussian noise candidate frame.
[0019] Non-maximum suppression is performed on the diffused Gaussian noise candidate frame to obtain an optimized Gaussian noise candidate frame.
[0020] The enhanced multi-scale feature map and the optimized Gaussian noise candidate box are input into the feature decoder of the target detection model to obtain the position and category of the bounding box of the target object; the target detection model is a model trained based on the determination method of the target detection model described above.
[0021] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for determining a target detection model or the method for applying a target detection model.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for determining a target detection model or the method for applying a target detection model.
[0023] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for determining a target detection model or the method for applying a target detection model.
[0024] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0025] The present application provides a method for determining a target detection model, an application method and a related device, the determination method comprising: obtaining a data set; the data set comprising the position and category of the bounding box of a target object in a plurality of RGB images; preprocessing the data set to obtain a preprocessed data set; performing multi-scale feature extraction on the preprocessed data set to obtain a plurality of multi-scale feature maps; performing edge feature enhancement on the last layer features of a plurality of the multi-scale feature maps to obtain a plurality of enhanced multi-scale feature maps; based on the data set, randomly generating a Gaussian noise candidate box, and diffusing the Gaussian noise candidate box to obtain a diffused Gaussian noise candidate box; performing non-maximum suppression on the diffused Gaussian noise candidate box to obtain an optimized Gaussian noise candidate box; taking the enhanced multi-scale feature map and the optimized Gaussian noise candidate box as input, taking the position and category of the bounding box of the corresponding target object as a label, training a constructed model to obtain a target detection model; the constructed model comprises: a feature extractor, a differential convolution module, a noise candidate box generator, a candidate box optimization layer and a feature decoder. By enhancing the edge features of the image, the accuracy of target detection in low-light conditions can be improved; by optimizing the generation strategy of the noise candidate frame, redundant information is reduced, and the robustness and accuracy of the feature decoder are improved. This application is suitable for target detection tasks in multiple low-light environments such as autonomous driving and security monitoring, and provides reliable technical support for the application of visual perception systems in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0027] Figure 1 A flowchart of a method for determining a target detection model provided in one embodiment of the present application.
[0028] Figure 2 A structural block diagram of a method for determining a target detection model provided in one embodiment of the present application.
[0029] Figure 3 A flowchart of an application method of a target detection model provided in one embodiment of the present application.
[0030] Figure 4 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0032] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0033] This application provides a target detection method based on the DiffusionDet model, which enhances the accuracy of target detection in night scenes by introducing a differential convolution module and an improved noise candidate box generation strategy.
[0034] In an exemplary embodiment, Figure 1 and Figure 2 As shown, a method for determining a target detection model is provided, and the method for determining a target detection model includes:
[0035] A1: Obtain a data set; the data set includes the positions and categories of the bounding boxes of target objects in a number of RGB images.
[0036] A2: Preprocess the data set to obtain a preprocessed data set.
[0037] A3: Perform multi-scale feature extraction on the preprocessed data set to obtain a plurality of multi-scale feature maps.
[0038] A4: Perform edge feature enhancement on the last layer features of several multi-scale feature maps to obtain several enhanced multi-scale feature maps.
[0039] A5: Based on the data set, randomly generate a Gaussian noise candidate box, and diffuse the Gaussian noise candidate box to obtain a diffused Gaussian noise candidate box.
[0040] A6: Perform non-maximum suppression on the diffused Gaussian noise candidate frame to obtain an optimized Gaussian noise candidate frame.
[0041] A7: Taking the enhanced multi-scale feature map and the optimized Gaussian noise candidate box as input, and taking the position and category of the corresponding target object's bounding box as labels, the constructed model is trained to obtain a target detection model; the constructed model includes: a feature extractor, a differential convolution module, a noise candidate box generator, a candidate box optimization layer and a feature decoder.
[0042] The implementation of steps A1 to A7 above effectively combines deep learning technology with complex computer vision strategies, significantly improves target recognition accuracy in night scenes, reduces misidentification problems caused by insufficient light, and provides reliable technical support for the safety of autonomous vehicles in visually restricted environments.
[0043] As an optional implementation, in step A1, RGB images are collected by a vehicle-mounted camera, security monitoring equipment or low-light imaging equipment. The images cover target categories such as pedestrians, vehicles, bicycles, etc. on the road, and the bounding boxes and category labels of the target objects are annotated in the data set.
[0044] As an optional implementation, in step A2, specifically including:
[0045] A21: Standardize the data set to obtain a standardized data set.
[0046] A22: Perform dynamic size alignment on the standardized data set to obtain a preprocessed data set.
[0047] Specifically, 1. Standardized processing.
[0048] The predefined parameters are used: the mean of the three RGB channels μ = [123.675, 116.280, 103.530], and the standard deviation σ = [58.395, 57.120, 57.375].
[0049] The expression of the standardization process is:
[0050]
[0051] Among them, Inorm is the standardized image; Iraw is the original image; μ is the mean of the three RGB channels; σ is the standard deviation; This is channel-by-channel division.
[0052] 2. Dynamic size alignment.
[0053] The width and height of the image are expanded to an integer multiple of 32 by padding with zeros.
[0054] As an optional implementation, in step A3, the feature extraction network can use ResNet-50, ResNet101 or VGG-16 network, and the preprocessed image is input into the feature extraction network to obtain a multi-scale feature map {P2, P3, P4, P5}. At this time, the multi-scale feature map will retain the key feature information in the image, including the shape, texture and other features of the target object.
[0055] Since the details and edge information of images in night scenes are limited, this application uses a differential convolution module to enhance the edge features in the image by performing a convolution operation on the differences between adjacent pixels. The purpose of the differential convolution module is to improve the recognition of edges so that the target detection algorithm can still accurately identify targets in low-light environments.
[0056] As an optional implementation, in step A4, specifically including:
[0057] A41: Extract the central features and global features of each of the multi-scale feature maps respectively.
[0058] A42: Subtract the central feature from the global feature corresponding to each of the multi-scale feature maps to obtain the corresponding edge feature.
[0059] A43: Add the last layer features of each of the multi-scale feature maps to the corresponding edge features to obtain a plurality of enhanced multi-scale feature maps.
[0060] Specifically, the last layer P5 of the feature extraction network output is used as input, 1×1 convolution is applied to extract central features, 3×3 convolution is applied to extract global features, the central features are subtracted from the global features to obtain edge features, and the edge features are added to P5 to achieve edge feature enhancement.
[0061] In night scenes, there are more noise and less details in the image. The traditional candidate frame generation method is prone to generate redundant and highly overlapping noise candidate frames, which is not conducive to the subsequent decoder to generate target frames. Therefore, the number of noise candidate frames is first expanded to ensure that the target to be detected is covered, and then NMS is used to process the expanded noise candidate frames to improve the subsequent decoder's ability to recognize the target.
[0062] Specifically, in step A5, during the diffusion process, Gaussian noise candidate boxes are randomly generated, and the size of the noise boxes follows a distribution with a mean of 32×32 and a variance of 0.2, increasing the number of noise candidate boxes from 500 to 1000.
[0063] In step A6, non-maximum suppression is performed on the 1000 noise candidate boxes obtained in the previous step, and the overlap threshold is designed to be 0.9 to enhance the feature decoder's ability to recognize the target box.
[0064] As an optional implementation, in step A7, specifically including:
[0065] A71: Input the enhanced multi-scale feature map into a feature decoder to obtain an output of the feature decoder; the output of the feature decoder is the predicted position and category of the target object.
[0066] A72: Determine a loss value based on the position and category of the bounding box of the real target object, the output of the feature decoder and the determined loss function. The expression of the loss function is:
[0067]
[0068] in, is the loss function; cls is the first weight parameter; is the classification loss based on Focal Loss, which is used to solve the problem of category imbalance; L1 is the second weight parameter; is the coordinate regression loss, which is used to directly optimize the box position parameters; giou is the third weight parameter; It is the GIoU loss, which increases the sensitivity to box overlap.
[0069] A73: According to the loss value, the network parameters of the feature extractor, the differential convolution module, the noise candidate box generator, the candidate box optimization layer and the feature decoder are optimized to obtain a target detection model.
[0070] In summary, this application provides a method for determining a target detection model, which specifically includes the following steps: first, using a feature extraction network to convert night images into multi-scale feature maps, and in the process, enhancing edge detail features through edge convolution; second, generating Gaussian noise and converting it into random Gaussian noise candidate frames, increasing the number of noise candidate frames, and then applying non-maximum suppression to these noise frames to remove repeated noise frames that are difficult for the decoder to identify; finally, the optimized noise frames are processed by the feature decoder to obtain the final target detection result. This application effectively combines deep learning technology with complex strategies of computer vision, significantly improves the accuracy of target recognition in night scenes, reduces the problem of misrecognition caused by insufficient light, and provides reliable technical support for the safety of autonomous driving vehicles in visually restricted environments.
[0071] like Figure 3 As shown, a method for applying a target detection model is provided, that is, target detection is performed based on the trained model. The specific steps are as follows:
[0072] B1: Get the image to be detected.
[0073] B2: Preprocessing the image to be detected to obtain a preprocessed image.
[0074] B3: Perform multi-scale feature extraction on the preprocessed image to obtain a multi-scale feature map.
[0075] B4: Perform edge feature enhancement on the last layer of features of the multi-scale feature map to obtain an enhanced multi-scale feature map.
[0076] B5: Based on the image to be detected, randomly generate a Gaussian noise candidate frame, and diffuse the Gaussian noise candidate frame to obtain a diffused Gaussian noise candidate frame.
[0077] B6: Perform non-maximum suppression on the diffused Gaussian noise candidate frame to obtain an optimized Gaussian noise candidate frame.
[0078] B7: Input the enhanced multi-scale feature map and the optimized Gaussian noise candidate box into the feature decoder of the target detection model to obtain the position and category of the bounding box of the target object; the target detection model is a model trained based on the determination method of the target detection model described above.
[0079] As an optional implementation, in step B1, the image to be detected is a night picture.
[0080] As an optional implementation, in step B2, specifically including:
[0081] B21: performing standardization processing on the image to be detected to obtain a standardized image.
[0082] B22: Performing dynamic size alignment on the standardized image to obtain a preprocessed image.
[0083] As an optional implementation, in step B3, the preprocessed image is input into a feature extraction network to obtain a multi-scale feature map {P2, P3, P4, P5}.
[0084] In night scenes, due to insufficient lighting, image details and edge information are easily lost. In order to make up for this deficiency, the present application introduces a differential convolution module in the feature extraction network. This module can effectively enhance the edge information in the image by performing a convolution operation on the differences between adjacent pixels, thereby improving the accuracy of target detection under low light conditions. That is, in step B4, the differential convolution module is used to enhance the edge features of P5 to obtain an enhanced multi-scale feature map.
[0085] In night scenes, due to the increase in image noise and the loss of details, traditional candidate box generation methods often lead to redundant and highly overlapping candidate boxes. To solve this problem, the present application increases the number of generated noise candidate boxes to enhance the ability of the feature decoder to identify targets from more candidate areas. Subsequently, the non-maximum suppression strategy is applied to effectively remove candidate boxes with high overlap, improve the detection capability of the feature decoder, and further improve the detection accuracy. That is, in steps B5 and B6, noise is generated by the diffusion noise candidate box generator and optimized using NMS, and the noise box is diffused and denoised by iterating T = 100 steps of the diffusion model to gradually approach the true target distribution.
[0086] As an optional implementation, in step B7, the category and refined coordinate values of each frame are generated for the denoised noise frame through a feature decoder.
[0087] Specifically, after generating the optimized noise candidate boxes, the feature decoder is used in combination with the feature map extracted above to optimize and determine the specific location and category of the target in each candidate box. The feature decoder further accurately locates the target based on the feature map of the candidate box and the optimized edge features and outputs the final detection result.
[0088] In addition, it should be noted that in actual applications, after obtaining the category and refined coordinate values of each box, post-processing will be performed, that is, the target box obtained by the feature decoder is filtered using NMS, and the overlap threshold is set to 0.5 to obtain the final detection target.
[0089] The trained and optimized object detection model can be directly deployed in practical applications such as self-driving cars and security monitoring. During the object detection process, after inputting night images, the model automatically outputs the location and category of the target and performs real-time analysis to support the decision-making system.
[0090] The technical solution of this application can significantly improve the accuracy of target detection in night scenes. By introducing the differential convolution module, the edge features in the image are effectively enhanced; and by optimizing the generation strategy of the noise candidate frame, the redundant information is reduced, and the robustness and accuracy of the feature decoder are improved. This technical solution is suitable for target detection tasks in multiple low-light environments such as autonomous driving and security monitoring, and provides reliable technical support for the application of visual perception systems in complex environments.
[0091] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 4As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data sets. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for determining a target detection model or an application method of a target detection model is implemented.
[0092] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0093] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the above-mentioned method embodiments are implemented when the processor executes the computer program.
[0094] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the above-mentioned method embodiments are implemented.
[0095] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the above-mentioned method embodiments are implemented.
[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0097] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0098] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0099] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0100] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for determining a target detection model, characterized in that: The method for determining the target detection model includes: Acquire a data set; the data set includes the location and category of the bounding box of the target object in a plurality of RGB images; Preprocessing the data set to obtain a preprocessed data set; Performing multi-scale feature extraction on the preprocessed data set to obtain a plurality of multi-scale feature maps; Performing edge feature enhancement on the last layer features of several multi-scale feature maps to obtain several enhanced multi-scale feature maps; Based on the data set, randomly generate a Gaussian noise candidate box, and diffuse the Gaussian noise candidate box to obtain a diffused Gaussian noise candidate box; Performing non-maximum suppression on the diffused Gaussian noise candidate frame to obtain an optimized Gaussian noise candidate frame; The enhanced multi-scale feature map and the optimized Gaussian noise candidate box are used as input, and the position and category of the bounding box of the corresponding target object are used as labels to train the constructed model to obtain a target detection model; the constructed model includes: a feature extractor, a differential convolution module, a noise candidate box generator, a candidate box optimization layer and a feature decoder.
2. The method for determining a target detection model according to claim 1, characterized in that: Preprocessing the data set to obtain a preprocessed data set specifically includes: Performing standardization on the data set to obtain a standardized data set; Dynamic size alignment is performed on the standardized data set to obtain a preprocessed data set.
3. The method for determining a target detection model according to claim 2, characterized in that: The expression of the standardization process is: Among them, Inorm is the standardized image; Iraw is the original image; μ is the mean of the three RGB channels; σ is the standard deviation; This is channel-by-channel division.
4. The method for determining a target detection model according to claim 1, characterized in that: Performing edge feature enhancement on the last layer features of the plurality of multi-scale feature maps to obtain a plurality of enhanced multi-scale feature maps, specifically comprising: Respectively extracting central features and global features of each of the multi-scale feature maps; Subtract the central feature from the global feature corresponding to each of the multi-scale feature maps to obtain the corresponding edge feature; The last layer feature of each of the multi-scale feature maps is added to the corresponding edge feature to obtain a plurality of enhanced multi-scale feature maps.
5. The method for determining a target detection model according to claim 1, characterized in that: Taking the multi-scale feature map as input, taking the enhanced multi-scale feature map and the optimized Gaussian noise candidate box as input, taking the position and category of the bounding box of the corresponding target object as a label, training the constructed model to obtain a target detection model, specifically including: Inputting the enhanced multi-scale feature map into a feature decoder to obtain an output of the feature decoder; the output of the feature decoder is the predicted position and category of the target object; Determining a loss value based on the position and category of the bounding box of the real target object, the output of the feature decoder and the determined loss function; According to the loss value, the network parameters of the feature extractor, the differential convolution module, the noise candidate box generator, the candidate box optimization layer and the feature decoder are optimized to obtain a target detection model.
6. The method for determining a target detection model according to claim 5, characterized in that: The expression of the loss function is: in, is the loss function; cls is the first weight parameter; is the classification loss based on FocalLoss, which is used to solve the problem of category imbalance; L1 is the second weight parameter; is the coordinate regression loss, which is used to directly optimize the box position parameters; giou is the third weight parameter; It is the GIoU loss, which increases the sensitivity to box overlap.
7. An application method of a target detection model, characterized in that: The application method of the target detection model includes: Acquire the image to be detected; Preprocessing the image to be detected to obtain a preprocessed image; Performing multi-scale feature extraction on the preprocessed image to obtain a multi-scale feature map; Performing edge feature enhancement on the last layer of features of the multi-scale feature map to obtain an enhanced multi-scale feature map; Based on the image to be detected, randomly generate a Gaussian noise candidate frame, and diffuse the Gaussian noise candidate frame to obtain a diffused Gaussian noise candidate frame; Performing non-maximum suppression on the diffused Gaussian noise candidate frame to obtain an optimized Gaussian noise candidate frame; The enhanced multi-scale feature map and the optimized Gaussian noise candidate box are input into the feature decoder of the target detection model to obtain the position and category of the bounding box of the target object; the target detection model is a model trained based on the target detection model determination method described in any one of claims 1-6.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for determining the target detection model described in any one of claims 1 to 6 or the method for applying the target detection model described in claim 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method for determining the target detection model described in any one of claims 1 to 6 or the method for applying the target detection model described in claim 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the method for determining the target detection model described in any one of claims 1 to 6 or the method for applying the target detection model described in claim 7.