Cross-domain adaptive target detection method for dynamic water area scene
By improving the image enhancement and structural optimization of the YOLOv11 model, combined with the adaptive convolutional fusion module and the improved coordinate loss function, the problem of low accuracy of traditional object detection methods in complex water environments is solved, and efficient and accurate water object detection is achieved, reducing costs.
Patent Information
- Application Number
- CN202510451212.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional object detection methods face interference such as light changes, large, water surface reflections and ripple in complex water environments, resulting in low detection accuracy. The application of deep learning neural network models in this environment requires a large amount of data and high-performance hardware, which is costly.
Improve the YOLOv11 model, and use the multi-overlapping window bicubital interpolation synthesis method to enhance image, maintain image edges and details, and reduce artificial boundaries. Multi-scale and adaptive convolution fusion modules are designed to adjust the target shape through adaptive convolution units, and combined with improved coordinate loss function, enhance the adaptability and robustness of the model.
Achieve efficient and accurate object detection and identification in complex water environments, improve the robustness and real-timeness of the model, reduce dependence on high-performance hardware, and reduce costs.
Smart Images

Figure CN120014248A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to image processing and machine learning technology, and in particular to a cross-domain adaptive target detection method for dynamic water scenes. Background Art
[0002] With the popularity of water activities and the growth of related economic activities, the demand for efficient and accurate water monitoring methods is increasing. Traditional video-based target detection methods often face multiple challenges in complex water environments, such as large changes in lighting, interference from water surface reflections and ripples, and these environmental factors seriously affect the detection accuracy. In addition, different water quality conditions and weather changes also bring additional difficulties to target detection. For example, in strong or low light conditions, the performance of traditional cameras will be limited, water surface reflections will cause overexposure or loss of details, and ripples and floating objects may cause incorrect target recognition, which makes it impossible to effectively adapt to the light reflections and visual disturbances on the water surface, resulting in false detections or missed detections, reducing practicality and reliability.
[0003] Existing deep learning neural network model methods, such as the deep learning-based detection method YOLOv11 model, perform well in general environments, but still show limitations when applied to dynamic and complex light environments on the water. They require a large amount of water data sets for training and high-performance hardware equipment for calculations, consuming a lot of manpower and material resources to cope with complex situations on the water.
[0004] Therefore, how to improve the structure of the neural network model and obtain better recognition accuracy and efficiency at a low cost has become a key research direction. Summary of the invention
[0005] In order to solve the problems existing in the prior art and improve the efficiency and accuracy of real-time target detection and recognition in aquatic environments, this paper improves the YOLOv11 model and proposes a cross-domain adaptive target detection method for dynamic water scenes. The method is improved mainly in two aspects:
[0006] 1. In the process of data enhancement, in order to avoid the loss of image continuity and naturalness caused by window division in the enhancement process in the traditional histogram equalization processing, a multi-overlapping window bicubic interpolation synthesis method is adopted to maintain the image edges and details, ensure the continuity and naturalness of image enhancement, reduce the artificial boundaries caused by window division, and enhance the adaptability of the model in various environments.
[0007] 2. Model structure adjustment and optimization: A new multi-scale and adaptive convolution fusion module (MSAK Conv) is designed to replace the shallow convolution module of the traditional YOLOv11 model. In particular, the newly designed adaptive convolution unit (AdpConv) is used to further adjust and adapt to the shape of the target, which helps to capture more complex and irregular target shapes and can simultaneously adapt to the detection of surface targets of different sizes and shapes. In addition, an improved coordinate loss function is used during model training to comprehensively consider the size and size of the bounding box itself, enhance the bounding box's ability to locate the real target, and reduce the impact of background noise on detection.
[0008] The technical solution of the present invention is:
[0009] A cross-domain adaptive target detection method for dynamic water scenes includes the following steps:
[0010] Step 1: Collect and annotate images of water surface objects to build a water target detection dataset;
[0011] Step 2: Perform image enhancement on the water target detection dataset constructed in step 1, and further process the image using the improved adaptive histogram equalization method:
[0012] Step 2.1: Set a sliding window on the image, and adjacent sliding windows overlap, and the overlapping area is greater than half of the window area;
[0013] Step 2.2: Use adaptive histogram equalization method to adjust the image contrast for each sliding window;
[0014] Step 2.3: For each pixel in the image, a bicubic interpolation method is used to synthesize the grayscale value of the pixel in each sliding window adjusted in step 2.2 to obtain the final grayscale value of the pixel;
[0015] Step 2.4: Reassemble the processed pixels into a complete image;
[0016] Step 3: Establish a detection model; the detection model is obtained by replacing the shallow convolution module of the YOLOv11 model with a multi-scale and adaptive convolution fusion module;
[0017] In the multi-scale and adaptive convolution fusion module, the input image is processed in parallel by three convolution kernels of different sizes, and the obtained feature maps are spliced in the channel dimension to complete feature stacking; the stacked features are input into the multi-level channel attention unit, and the feature maps processed by the multi-level channel attention unit are finally input into the adaptive convolution unit; the calculation formula of the adaptive convolution unit is
[0018]
[0019] Represents the weight of the corresponding area of the feature map calculated by a convolution kernel in the adaptive convolution unit, Represents the feature map of the input adaptive convolution unit, and They are the height and width of the feature map of the input adaptive convolution unit, respectively. Represents a convolution kernel in the adaptive convolution unit, with a size of , is a mask obtained by learning, wherein the mask is obtained by offset learning of a 3×3 convolution kernel;
[0020] After multiplying the weights of each region of the feature map calculated by each convolution kernel in the adaptive convolution unit with the corresponding region of the feature map, the feature map output by the adaptive convolution unit is obtained by combining them;
[0021] Step 4: Divide the water target detection dataset processed in step 2 into a training set and a validation set. Use the training set to train the detection model established in step 3, and evaluate the performance of the model on the validation set. Adjust the model parameters and structure according to the validation results to obtain the final trained detection model. The loss function in the training process is:
[0022]
[0023] in is the total loss, To control the weight of coordinate loss, is the weight of controlling the object confidence loss, To control the weight of category loss; is the object confidence loss, is the category loss, is the coordinate loss;
[0024]
[0025] in is the scaling factor, is the bounding box size loss function, is the bounding box shape loss function:
[0026]
[0027]
[0028] in For the The actual center coordinates of the bounding box, For the The predicted center coordinates of the bounding boxes, and They are respectively The actual width and actual height of the bounding box, and They are respectively The predicted width and height of the bounding box; N is the number of bounding boxes, is an indicator function, which takes the value 1 if the object exists, otherwise 0; and They represent the weight coefficients in the horizontal direction and the vertical direction respectively. is the set scale constant;
[0029] Step 5: Input the new image into the detection model trained in step 4, detect and identify the objects in it, and output the detection results including the bounding box and category confidence.
[0030] Furthermore, in step 2, the image enhancement operations include rotation, flipping, cropping, scaling, and lighting adjustment.
[0031] Furthermore, in step 2.1, the size of the sliding window is dynamically adjusted according to the local characteristics of the image content during the image processing process, using a smaller window in areas with rich details and a larger window in uniform areas.
[0032] Furthermore, in step 2.2, the formula for adjusting the image contrast using the adaptive histogram equalization method is:
[0033]
[0034]
[0035] in is the transformation function, is the gray value in the original image, is the new grayscale value after adjustment, Represents grayscale value Probability of occurrence.
[0036] Furthermore, in the multi-scale and adaptive convolution fusion module, the three convolution kernels with sizes of 3x3, 5x5 and 7x7 that process the input image in parallel capture details of different scales in the image through convolution kernels of different sizes.
[0037] Furthermore, in step 4, the object confidence loss
[0038]
[0039] in For the The confidence that an object actually exists in the bounding box, For the The confidence level of the predicted existence of an object in the bounding box, It is an indicator function when the object does not exist. If the object does not exist, the value is 1, otherwise it is 0. is the weight of the non-existent object.
[0040] Furthermore, in step 4, the category loss
[0041]
[0042] in is the true category label, The target predicted by the model belongs to the category The probability of A collection of categories.
[0043] In addition, the present invention also provides an electronic device and a readable storage medium:
[0044] An electronic device comprises a processor and a memory, wherein the memory is used to store one or more programs;
[0045] When the one or more programs are executed by the processor, the above method is implemented.
[0046] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0047] Beneficial effects:
[0048] The method proposed in the present invention ensures efficient and accurate detection and identification of people and obstacles in complex water environments. The present invention not only optimizes the data processing and model training process, but also improves the robustness and real-time performance of the model through structural adjustment and preprocessing enhancement, providing a solid technical foundation for practical applications.
[0049] Compared with the prior art, the present invention adopts a multi-overlapping window bicubic interpolation synthesis method in the data enhancement process to maintain image edges and details, ensure the continuity and naturalness of image enhancement, reduce artificial boundaries caused by window division, and enhance the adaptability of the model in various environments; in the established detection model, a new multi-scale and adaptive convolution fusion module is designed to replace the shallow convolution module of the traditional YOLOv11 model, and the shape of the target is further adjusted and adapted through the adaptive convolution unit therein, so as to help capture more complex and irregular target shapes, and can simultaneously adapt to the detection of water surface targets of different sizes and shapes; and in the model training process, an improved coordinate loss function is adopted, the size and size of the bounding box itself are comprehensively considered, the positioning ability of the bounding box for the real target is enhanced, and the influence of background noise on detection is reduced, so that it can effectively adapt to the complex and changeable water environment, and significantly improve the accuracy of water target detection under the conditions of unstable lighting and water surface ripple interference.
[0050] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0052] Figure 1 A schematic diagram of the process of the cross-domain adaptive target detection method for dynamic water scenes proposed by the present invention;
[0053] Figure 2 This is a schematic diagram of the improved YOLOv11 network;
[0054] Figure 3 This is the schematic diagram of the MSAK module;
[0055] Figure 4 Schematic diagram of the AdpConv unit architecture. DETAILED DESCRIPTION
[0056] Embodiments of the present invention are described in detail below. The embodiments are exemplary and intended to be used to explain the present invention, but should not be construed as limiting the present invention.
[0057] like Figure 1 As shown, the cross-domain adaptive target detection method for dynamic water scenes in this embodiment includes the following steps:
[0058] Step 1: Build a water target detection dataset.
[0059] Frame extraction is performed on the video of water surface objects captured by a high-definition camera to extract static images. Here, representative and diverse image frames can be selected through timed frame extraction or motion detection algorithms.
[0060] The image obtained after the frame extraction is manually labeled to label the objects in the image. In this embodiment, the labeling tool LabelImg is used for labeling, and each target object is labeled by a bounding box and a corresponding label is assigned.
[0061] The intersection-over-union ratio is used to evaluate the accuracy of the annotation and the degree of overlap between the bounding box predicted by the detection model and the true bounding box. The calculation formula is as follows:
[0062]
[0063] is the area of overlap between the predicted bounding box and the true bounding box. It is the total area covered by the predicted bounding box and the true bounding box, that is, the area of the predicted bounding box and the true bounding box minus the area of the overlap.
[0064] Step 2: Perform image enhancement operations on the water target detection dataset constructed in step 1, including rotation, flipping, cropping, scaling, and lighting adjustment.
[0065] By randomly rotating and flipping the image, the trained model is not dependent on the specific orientation of the target in the image, which enhances the model's adaptability to changes in the object's orientation. By randomly cropping and scaling different areas of the image, the model can learn to recognize objects from partial views, while also simulating different object sizes and perspectives.
[0066] Lighting adjustment, i.e. adjusting the brightness, contrast and hue of an image, allows you to simulate different lighting conditions, which is especially important for dealing with lighting changes in natural environments.
[0067] In order to better handle the problem of illumination change and water surface reflection, simulate data under different water quality (clear, turbid) and weather (sunny, cloudy, rainy) conditions, and enhance the adaptability of the model in various environments, this embodiment also uses an improved adaptive histogram equalization method to process images. The specific process is as follows:
[0068] Step 2.1: Set a sliding window on the image, and the size of the sliding window is dynamically adjusted according to the local characteristics of the image content during image processing. Use a smaller window in areas with rich details and a larger window in relatively uniform areas. Adjacent sliding windows overlap, and the overlapping area is greater than half of the window area to avoid visible edge effects when processing boundaries and ensure the reliability of statistical data.
[0069] Step 2.2: For each window in the image, calculate the grayscale histogram and adjust the contrast of the image by adaptive histogram equalization method.
[0070] Since the histogram of each window is calculated independently, that is, the grayscale histogram of each window is calculated, and the frequency of each grayscale value appearing in the window is counted. Then the contrast of the image is improved by the adaptive histogram equalization method (AHE), and its expression is:
[0071]
[0072] in is the transformation function, is the gray value in the original image, is the new grayscale value after adjustment.
[0073] Applying equalization to the histogram of each window involves redistributing the grayscale values to make the histogram distribution more uniform, thereby improving the local contrast of the window. The specific approach is to calculate the local adaptive cumulative distribution function (CDF) for transformation, that is,
[0074]
[0075] This function is a mathematical relationship of the cumulative frequency of each gray value, which is used to map the old gray value to the new gray value; in the function, Represents grayscale value The probability of occurrence, the integral is calculated from the gray value 0 to The cumulative probability, that is, the gray value The sum of the probabilities of occurrence of all gray values below it.
[0076] Step 2.3: For each pixel, it may belong to multiple overlapping windows, so each window will give an adjusted grayscale value. The transformed grayscale values of each window at the same pixel are synthesized by the bicubic interpolation method to determine the final grayscale value for each pixel.
[0077] This method ensures the continuity and naturalness of the enhancement by preserving edges and details through bicubic interpolation and reducing the artificial boundaries caused by window division.
[0078] Step 2.4: Reassemble the processed pixels into a complete image. The effect of image enhancement was evaluated. The results showed that after image enhancement, data under different water quality (clear, turbid) and weather conditions (sunny, cloudy, rainy) can be simulated, enhancing the adaptability of the model in a variety of environments.
[0079] Step 3: Establish a detection model. The detection model is obtained by optimizing the traditional YOLOv11 model. The main purpose is to design a new multi-scale and adaptive convolution fusion module to replace the shallow convolution module of the traditional YOLOv11 model.
[0080] Step 3.1: Based on the traditional YOLOv11 model, initialize the network structure of the detection model, including the input layer, feature extraction layer, detection head and other modules. According to the specific application scenario and dataset characteristics, configure model parameters such as input size, anchor boxes, number of categories, etc.
[0081] Step 3.2: Design a multi-scale and adaptive convolution fusion module (MSAK Conv) to replace the shallow convolution module of the traditional YOLOv11 model, such as Figure 4 As shown, it can simultaneously adapt to the detection of surface targets of different sizes and shapes. Specifically:
[0082] like Figure 3 As shown in the figure, in the multi-scale and adaptive convolution fusion module, the input image is processed in parallel by three convolution kernels of different sizes (3x3, 5x5, and 7x7) to extract features of different scales. Specifically, using convolution kernels of different sizes can capture details of different scales in the image. For example, small convolution kernels (3x3) pay more attention to details and local information, medium convolution kernels (5x5) process broader contextual information, and large convolution kernels (7x7) help capture a wider range of global information. These feature maps processed by different convolution kernels are spliced (Channel Concat) in the channel dimension to complete feature stacking, so that the model can comprehensively utilize this information in subsequent processing, thereby enhancing its understanding and analysis capabilities of complex scenes.
[0083] The stacked features are input into the Multi-Level Channel Attention (MLCA) unit. By calculating the attention weights between each channel, the model adaptively adjusts the weights of features of different scales, thereby fusing these feature maps processed by different convolution kernels across channels, so that the model can better understand and fuse information at each scale and integrate features of multiple scales.
[0084] The feature map after MLCA is finally input into the adaptive convolution unit (AdpConv) to further adjust and adapt to the shape of the target, which helps to capture more complex and irregular target shapes, especially when the target size and shape vary greatly.
[0085] like Figure 4As shown in the figure, in the adaptive convolution unit, based on a 3×3 rectangular convolution kernel, 9 trainable parameters are introduced to mark the displacement direction of each convolution position, so that the 3×3 convolution kernel can create a mask (MASK) through offset learning. Through this mask, a new convolution kernel can be constructed, and its size can be expanded to 4×4 at most. This 4×4 convolution kernel can adjust its shape according to different masks to better match the features in the data. As the gradient descent process proceeds, these shapes will continue to be optimized to better adapt to the data, and ultimately achieve the best convolution performance. The calculation formula of the adaptive convolution unit is
[0086]
[0087] Represents the weight of the corresponding area of the feature map calculated by a convolution kernel, Represents the feature map of the input adaptive convolution unit, with a size of , Represents a convolution kernel with a size of , is the mask obtained through learning. The weights of each region of the feature map calculated by each convolution kernel in the adaptive convolution unit are multiplied by the corresponding region of the feature map, and then the feature map output by the adaptive convolution unit is obtained.
[0088] Step 4: Divide the water target detection dataset processed in step 2 into a training set and a validation set. Use the training set to train the detection model established in step 3, optimize the model parameters through the back propagation algorithm, adjust the learning rate, batch size and other hyperparameters during the training process, and use early stopping and learning rate scheduling to improve the training effect. Evaluate the performance of the model on the validation set, adjust the model parameters and structure according to the validation results, and obtain the final trained detection model.
[0089] The overall loss function during traditional YOLOv11 model training is a weighted sum of several main loss functions, including coordinate loss, object confidence loss, and category loss:
[0090]
[0091] in is the total loss, which is the weighted sum of each sub-loss and is used for optimization during training. To control the weight of coordinate loss, is the weight of controlling the object confidence loss, is the weight for controlling the class loss. The object confidence loss measures the model's prediction accuracy of whether an object exists and how well the predicted confidence matches the actual situation. is the category loss, which is used to evaluate the accuracy of the model’s prediction of the target category. is the coordinate loss, which is responsible for reducing the difference between the predicted bounding box and the true bounding box.
[0092] In this invention, an improved coordinate loss function is proposed to solve the problem of false detection or missed detection caused by the influence of water surface background noise (such as water waves, light, etc.) on the detection accuracy:
[0093]
[0094] in is the scaling factor, which is set according to the target size; is the bounding box size loss function, The improved coordinate loss function is the bounding box shape loss function, which calculates the loss by focusing on the shape and size of the bounding box itself, thereby making the bounding box regression more accurate.
[0095] Specifically, the bounding box size loss function is:
[0096]
[0097] in For the The actual center coordinates of the bounding box, For the The predicted center coordinates of the bounding boxes, and They are respectively The actual width and actual height of the bounding box, and They are respectively The predicted width and predicted height of the bounding boxes; N is the number of bounding boxes, and is an indicator function that equals 1 if the object exists and 0 otherwise.
[0098] This part of the loss calculates the square of the Euclidean distance between the predicted center coordinates and the actual center coordinates, with the goal of minimizing the position error of the predicted center.
[0099] The square root of this loss is used to reduce the impact of the size error of the large-sized bounding box on the loss, so that the model has a more balanced sensitivity to the size error of bounding boxes of different sizes compared to the small-sized bounding box.
[0100] The bounding box shape loss function is:
[0101]
[0102] in and They represent the weight coefficients in the horizontal direction and the vertical direction respectively. is the set scale constant.
[0103] The improved coordinate loss function comprehensively considers the size and shape of the target, and can solve the problem of false detection or missed detection caused by the influence of water surface background noise (such as water waves, light, etc.) on the detection accuracy.
[0104] Object Confidence Loss Responsible for optimizing the confidence of the model in predicting whether an object exists within the bounding box:
[0105]
[0106] For the The confidence that an object actually exists in the bounding box, For the The confidence level of the predicted object in the bounding box, It is an indicator function when the object does not exist, which means that if the object does not exist, it is 1, otherwise it is 0. is the weight of the non-existent object, which is used to balance the imbalance of positive and negative samples.
[0107] Class loss Responsible for the performance of the model in multi-category recognition tasks, usually using cross entropy loss:
[0108]
[0109] here is the true category label, The target predicted by the model belongs to the category The probability of A collection of categories.
[0110] Step 5: Detection and recognition: Input the new image into the detection model trained in step 4. The model processes the input data, detects and recognizes the objects in it, and outputs the detection results including bounding boxes and category confidence.
[0111] Comparison and verification:
[0112] The detection model proposed in the present invention is compared with the traditional YOLOv11 model on the publicly available dataset WSODD. The detection model proposed in the present invention achieves 80% of mAP@0.5 and 45.3% of mAP@0.5:0.95, which are 4.5% and 3.4% higher than the traditional YOLOv11 model, respectively. Compared with the mainstream algorithm, the detection model proposed in the present invention has superior performance.
[0113] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention without departing from the principles and intent of the present invention.
Claims
1. A cross-domain adaptive target detection method for dynamic water scenes, characterized by: The following steps are involved: Step 1: Collect and annotate images of water surface objects to build a water target detection dataset; Step 2: Perform image enhancement on the water target detection dataset constructed in step 1, and further process the image using the improved adaptive histogram equalization method: Step 2.1: Set a sliding window on the image, and adjacent sliding windows overlap, and the overlapping area is greater than half of the window area; Step 2.2: Use adaptive histogram equalization method to adjust the image contrast for each sliding window; Step 2.3: For each pixel in the image, a bicubic interpolation method is used to synthesize the grayscale value of the pixel in each sliding window adjusted in step 2.2 to obtain the final grayscale value of the pixel; Step 2.4: Reassemble the processed pixels into a complete image; Step 3: Establish a detection model; the detection model is obtained by replacing the shallow convolution module of the YOLOv11 model with a multi-scale and adaptive convolution fusion module; In the multi-scale and adaptive convolution fusion module, the input image is processed in parallel by three convolution kernels of different sizes, and the obtained feature maps are spliced in the channel dimension to complete feature stacking; The stacked features are input into the multi-level channel attention unit, and the feature map processed by the multi-level channel attention unit is finally input into the adaptive convolution unit; the calculation formula of the adaptive convolution unit is: Represents the weight of the corresponding area of the feature map calculated by a convolution kernel in the adaptive convolution unit, Represents the feature map of the input adaptive convolution unit, and They are the height and width of the feature map of the input adaptive convolution unit, respectively. Represents a convolution kernel in the adaptive convolution unit, with a size of , is a mask obtained by learning, wherein the mask is obtained by offset learning of a 3×3 convolution kernel; After multiplying the weights of each region of the feature map calculated by each convolution kernel in the adaptive convolution unit with the corresponding region of the feature map, the feature map output by the adaptive convolution unit is obtained by combining them; Step 4: Divide the water target detection dataset processed in step 2 into a training set and a validation set. Use the training set to train the detection model established in step 3, and evaluate the performance of the model on the validation set. Adjust the model parameters and structure according to the validation results to obtain the final trained detection model. The loss function in the training process is: in is the total loss, To control the weight of coordinate loss, is the weight of controlling the object confidence loss, To control the weight of category loss; is the object confidence loss, is the category loss, is the coordinate loss; in is the scaling factor, is the bounding box size loss function, is the bounding box shape loss function: in For the The actual center coordinates of the bounding box, For the The predicted center coordinates of the bounding boxes, and They are respectively The actual width and actual height of the bounding box, and They are respectively The predicted width and predicted height of the bounding box; N is the number of bounding boxes, is an indicator function, which takes the value 1 if the object exists, otherwise 0; and They represent the weight coefficients in the horizontal direction and the vertical direction respectively. is the set scale constant; Step 5: Input the new image into the detection model trained in step 4, detect and identify the objects in it, and output the detection results including the bounding box and category confidence.
2. According to claim 1, a cross-domain adaptive target detection method for dynamic water scenes is characterized by: In step 2, image enhancement operations include rotation, flipping, cropping, scaling, and lighting adjustment.
3. According to claim 1, a cross-domain adaptive target detection method for dynamic water scenes is characterized by: In step 2.1, the size of the sliding window is dynamically adjusted according to the local characteristics of the image content during the image processing process, using a smaller window in areas with rich details and a larger window in uniform areas.
4. According to claim 1, a cross-domain adaptive target detection method for dynamic water scenes is characterized by: In step 2.2, the formula for adjusting the image contrast using the adaptive histogram equalization method is: in is the transformation function, is the gray value in the original image, is the new grayscale value after adjustment, Represents grayscale value Probability of occurrence.
5. According to claim 1, a cross-domain adaptive target detection method for dynamic water scenes is characterized by: In the multi-scale and adaptive convolution fusion module, the three convolution kernels with sizes of 3x3, 5x5, and 7x7 that process the input image in parallel capture details of different scales in the image through convolution kernels of different sizes.
6. According to claim 1, a cross-domain adaptive target detection method for dynamic water scenes is characterized by: In step 4, object confidence loss in For the The confidence that an object actually exists in the bounding box, For the The confidence level of the predicted object in the bounding box, It is an indicator function when the object does not exist. If the object does not exist, the value is 1, otherwise it is 0. is the weight of the non-existent object.
7. The cross-domain adaptive target detection method for dynamic water scenes according to claim 1 is characterized by: In step 4, the category loss in is the true category label, The target predicted by the model belongs to the category The probability of A collection of categories.
8. An electronic device, comprising a processor and a memory, wherein the memory is used to store one or more programs; characterized in that: When the one or more programs are executed by the processor, the method described in any one of claims 1 to 7 is implemented.
9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Water area target detection method based on multi-modal feature fusion and dynamic candidate box optimization
CN120580554A
Farmland obstacle detection method and device, storage medium and program product
CN120726599A