Lightweight low-light target detection method and system
By building a lightweight target detection model based on YOLOv8s, the problem of poor detection in low-light environments is solved, and high-precision and fast target detection is achieved, which is suitable for real-time applications in complex environments.
Patent Information
- Application Number
- CN202510666717.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional target detection models have poor detection effects in low-light environments, are difficult to meet real-time requirements, and are not robust enough to image interference such as occlusion, blur, and noise.
A lightweight target detection model is constructed using the YOLOv8s framework, including a multi-scale phantom convolution module, a backbone network module, a neck network module, and a detection head module. The model is trained on an annotated low-light image dataset to improve its detection accuracy and robustness in low-light conditions.
It improves the accuracy and speed of target detection in low-light conditions, enhances the robustness of the model, meets real-time requirements, and is suitable for scenarios such as night monitoring, unmanned driving, and security systems.
Smart Images

Figure CN120655889A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a lightweight low-light target detection method and system. Background Art
[0002] Lightweight refers to methods or systems optimized for low computing resource consumption, small model size, and fast execution speed. Low-light refers to conditions where the image or scene is dim or lacks brightness, such as nighttime surveillance or underground garages. Object detection, a key task in computer vision, involves identifying the location and category of objects of interest (such as people, vehicles, and animals) in an image. Lightweight low-light object detection refers to a technical approach that effectively detects objects in low-light environments.
[0003] In reality, many image acquisition scenarios are under low-light conditions, such as streets at night, underground garages, mines, etc. Therefore, a lightweight target detection method suitable for low-light environments is proposed, which can be more widely used in edge devices, real-time monitoring, security protection and other scenarios.
[0004] However, traditional object detection models are prone to failure in low-light environments. While large models offer strong performance, they are difficult to deploy. Furthermore, traditional methods lack multi-scale feature extraction, context understanding, and adaptability to complex scenes, limiting their widespread adoption in practical applications. This leads to poor detection results and inadequate robustness to image interference such as occlusion, blur, and noise, reducing detection accuracy and making it difficult to meet the demands of real-time scenarios. Summary of the Invention
[0005] In order to address the technical problems that traditional methods have shortcomings in multi-scale feature extraction, context understanding and adaptability to complex scenes, which limit their widespread promotion in practical applications, resulting in poor detection effect, insufficient robustness to image interference such as occlusion, blur, and noise, reduced detection accuracy, and difficulty in meeting scenes with high real-time requirements, the present invention provides a lightweight low-light target detection method and system.
[0006] The technical solutions provided by the embodiments of the present invention are as follows:
[0007] First aspect:
[0008] An embodiment of the present invention provides a lightweight low-light target detection method, including:
[0009] S1: Acquire low-light image dataset;
[0010] S2: Annotate low-light image dataset;
[0011] S3: Using YOLOv8s as the framework, a lightweight target detection model is constructed. The lightweight target detection model includes a multi-scale phantom convolution module, a backbone network module, a neck network module, and a detection head module.
[0012] S4: Input the labeled low-light image dataset into the lightweight object detection model for training;
[0013] S5: Acquire the low-light image to be detected;
[0014] S6: Input the low-light image to be detected into the trained lightweight target detection model, and output the target detection result of the low-light image to be detected.
[0015] Second aspect:
[0016] An embodiment of the present invention provides a lightweight low-light target detection system, comprising:
[0017] processor;
[0018] A memory stores computer-readable instructions, which, when executed by a processor, implement the lightweight low-light target detection method of the first aspect.
[0019] The third aspect:
[0020] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the lightweight low-light target detection method according to the first aspect is implemented.
[0021] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0022] In an embodiment of the present invention, a low-light image dataset is obtained and then annotated to ensure that the target information in the image can be learned. A lightweight target detection model is constructed using YOLOv8s as a framework, and the annotated low-light image dataset is input into the lightweight target detection model for training. Finally, in the actual application stage, the low-light image to be detected is obtained and input into the trained lightweight target detection model, which outputs the target detection result of the low-light image to be detected. This improves the detection accuracy and speed, enhances the robustness of the model, meets the needs of scenarios with high real-time requirements, and improves the recognition ability of the model in complex environments. It is of great significance for various scenarios such as night monitoring, unmanned driving, and security systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 A schematic diagram of a flow chart of a lightweight low-light target detection method provided by an embodiment of the present invention;
[0025] Figure 2 A schematic structural diagram of a lightweight low-light target detection system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0027] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.
[0028] In the embodiments of the present invention, the terms "image" and "picture" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same. The terms "of," "corresponding," and "corresponding" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same.
[0029] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0030] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0031] Reference Manual Figure 1 , shows a flow chart of a lightweight low-light target detection method provided by an embodiment of the present invention.
[0032] An embodiment of the present invention provides a lightweight low-light target detection method. This method can be implemented by a lightweight low-light target detection device, which can be a terminal or a server. The processing flow of the lightweight low-light target detection method may include the following steps:
[0033] S1: Obtain a low-light image dataset.
[0034] Among them, the low-light image dataset refers to a collection of images with poor lighting conditions and insufficient brightness, which is used for training and testing artificial intelligence models.
[0035] It should be noted that by collecting diverse low-light images, the generalization and robustness of the model under complex lighting conditions can be significantly enhanced, thereby improving its practical value and reliability in actual night monitoring, security and other application scenarios.
[0036] S2: Annotate low-light image dataset.
[0037] It should be noted that labeling the low-light image dataset is a key preparatory step before model training, which directly determines whether the model can correctly learn the visual features of the target under low-light conditions.
[0038] In a possible implementation, S2 specifically includes: labeling the low-light image dataset into a Yolo format.
[0039] The yolo format specifically includes: category number, target center position, target width and height information.
[0040] Among them, the YOLO format is a lightweight data annotation format commonly used in target detection. Each image corresponds to a .txt file, and each line represents a target.
[0041] The category number refers to the identifier of the category to which the target belongs, usually an integer.
[0042] The target center position refers to the coordinates of the center point of the target, which is usually expressed as a normalized ratio of the width and height of the image.
[0043] The target width refers to the width of the target frame, that is, the horizontal size of the target in the image. It is also a normalized value, which represents the ratio of the target width to the width of the entire image.
[0044] The height information refers to the height of the target frame, that is, the vertical size of the target in the image, which is usually expressed as a normalized value, that is, the ratio of the target height to the height of the entire image.
[0045] It should be noted that the YOLO format annotation is not only standardized and efficient to read, but also highly compatible, making it easy to directly apply to the training process of the YOLO series of models. This structured annotation method improves training efficiency and reduces the model's false detection and missed detection rate during actual detection, thus laying a solid data foundation for subsequent model construction and training.
[0046] S3: Using YOLOv8s as the framework, a lightweight target detection model is constructed. The lightweight target detection model includes: a multi-scale phantom convolution module, a backbone network module, a neck network module, and a detection head module.
[0047] YOLO (You Only Look Once) is a widely used target detection algorithm. YOLOv8s is a version of the YOLO series. v8 indicates the eighth generation, and s usually indicates a small variant of the version, which is suitable for low-resource devices or application scenarios that require fast processing.
[0048] Among them, the lightweight target detection model refers to a machine learning model with small computing requirements and low memory usage, which is mainly used to identify target objects in images and output their categories and locations.
[0049] The Multi-Scale Ghost Convolution Module is a custom convolution module that processes images at different scales to improve the accuracy and robustness of object detection, especially in low light or complex backgrounds. It extracts image features through convolution at different scales.
[0050] The backbone network module is responsible for feature extraction in the object detection model and typically includes classic convolutional neural network structures (such as ResNet and VGG). The neck network module connects the backbone network and the detection head in the detection model. It is typically used for feature fusion and optimization, enabling the model to perform better on multi-level features. The detection head module is the final part of the model, responsible for final object classification and localization based on the features extracted by the backbone network.
[0051] It should be noted that by building a lightweight target detection model based on the YOLOv8s framework, it is possible to significantly reduce the consumption of computing resources while maintaining high accuracy, making it suitable for deployment on edge devices or in low-power environments.
[0052] In a possible implementation, the backbone network module includes: an MA-SPPF unit and a C2fMSG unit.
[0053] The neck network module includes: RepNCSPELAN4 unit.
[0054] Among them, the RepNCSPELAN4 unit combines the characteristics of the CSP module and the ELAN module, and uses the RepConv module as a calculation module.
[0055] The detection head module is specifically a PDetect detection head.
[0056] The MA-SPPF unit is a module that combines maximum pooling and average pooling to extract spatial features from different scales. The C2fMSG unit is a feature fusion structure based on the C2f module (Improved Cross-Stage Partial) in YOLOv8.
[0057] The RepNCSPELAN4 unit is a composite structure that combines the advantages of multiple modules. The CSP module (CrossStage Partial) is used to split feature processing into two paths, improving gradient flow efficiency and reducing redundant computation. The ELAN module (Efficient Layer Aggregation Network) enhances feature reuse through efficient inter-layer connections. The RepConv module (Re-parameterized Convolution) is used to merge multiple convolutional structures into a single structure during inference, accelerating inference.
[0058] Among them, the RepNCSPELAN4 unit combines the characteristics of the segmentation input of the CSP module and the splicing of the branch results of the ELAN module.
[0059] Among them, the PDetect detection head (PartialDetection Head) is a simplified and optimized detection head structure.
[0060] In one possible implementation, the multi-scale phantom convolution module processes the input feature map obtained after processing the low-light image, specifically including:
[0061] Through the convolution layer, the number of channels of the input feature map is compressed to generate the intrinsic feature map.
[0062] The convolutional layer is the basic computational unit in a neural network, extracting features from an image or feature map by sliding a convolution kernel. The intrinsic feature map is a compressed, preliminary feature map that retains the core information of the input image and serves as the basis for subsequent operations.
[0063] The intrinsic feature map is split into a preset number of feature maps with the same number of channels.
[0064] In the present invention, the preset number of groups is specifically 2.
[0065] Perform convolution on the feature map to generate an output feature map.
[0066] Among them, the output feature map refers to the result after the group convolution operation, which is used to fuse with the intrinsic feature map.
[0067] The intrinsic feature map and the output feature map are concatenated in the channel dimension.
[0068] Through convolution operation, multi-dimensional features are extracted from the spliced feature maps.
[0069] Among them, multi-dimensional feature extraction is to further convolve the spliced feature map to extract information in multiple dimensions such as width, height, and depth, so that the model can understand the image content more comprehensively.
[0070] The extracted multi-dimensional features are fused to generate the final output feature map:
[0071]
[0072] Among them, output 多尺度幻影卷积 Represents the output of the multi-scale phantom convolution module, f1 and f2 both represent ordinary convolution, and input1 represents the input received by the multi-scale phantom convolution. represents a 3×3 group convolution, represents a 5×5 group convolution, Represents concatenation in the channel dimension.
[0073] Among them, the final output feature map is the final result of the multi-scale phantom convolution module, and the number of channels of the final output feature map is C2.
[0074] It should be noted that, first, by compressing the channels of the input image, the computational complexity is effectively reduced while retaining key intrinsic features. Secondly, grouping the feature maps and performing independent convolution processing can simulate multi-scale perception and enhance the model's ability to perceive objects of different sizes. After splicing the intrinsic features with the newly extracted features, convolution extraction and fusion are performed again, enabling the model to extract richer and more stable multi-dimensional features under low-light conditions. The overall design takes into account both computational efficiency and feature expression capabilities, improving the detection accuracy and robustness of the model in low-light scenes.
[0075] In a possible implementation, the MA-SPPF unit includes three branches.
[0076] The first branch is used for global maximum pooling processing:
[0077] y1=Maxpool(input2)
[0078] Among them, y1 represents the output feature of the first branch, Maxpool represents maximum pooling, and input2 represents the input feature received by the MA-SPPF unit.
[0079] It should be noted that global maximum pooling is a downsampling method that extracts the maximum value of each channel from the entire feature map and emphasizes the most significant local features.
[0080] The second branch is used for global average pooling:
[0081] y2=Avgpool(input2)
[0082] Where y2 represents the output feature of the second branch, and Avgpool represents average pooling. It should be noted that global average pooling is used to extract the average value of each channel from the entire feature map, focusing more on the overall information.
[0083] The third branch is the SPPF module branch:
[0084] y3=f(input2)
[0085] Among them, y3 represents the output feature of the third branch, and f() represents the branch processing of the SPPF module.
[0086] The feature map processed by the global maximum pooling, the feature map processed by the global average pooling, and the feature map processed by the third branch are spliced in the depth direction:
[0087]
[0088] in, Indicates splicing in the channel dimension, output MA-SPPF Represents the output of the MA-SPPF unit.
[0089] Among them, the third branch refers to an additional path other than maximum pooling and average pooling, which usually directly performs serial pooling on the input to maintain the original information flow.
[0090] It should be noted that the input feature map is subjected to global maximum pooling and global average pooling respectively to extract the significant features and overall statistical features in the image. The two complement each other and help enhance the model's ability to perceive the significant areas and background information in the image. At the same time, the features processed by the remaining branches are combined with the Figure 1 The combined features are spliced together in the channel dimension to form a richer fused feature map, enabling the model to better understand spatial contextual relationships. This overall design not only improves the model's ability to express different features, but also maintains a low computational cost, helping to improve detection stability and accuracy in complex environments such as low light.
[0091] In one possible implementation, the processing of the RepNCSPELAN4 unit specifically includes:
[0092] Perform convolution on the input feature map:
[0093] F1=conv 1×1 (input3)
[0094] Among them, F1 represents the feature after 1×1 convolution processing, conv 1×1 Represents a 1×1 convolution kernel, and input3 represents the input features received by the RepNCSPELAN4 unit.
[0095] The input feature map after convolution is divided into the first part and the second part.
[0096] Combine the RepNCSP module and the convolutional layer to process the second part:
[0097] F2=conv 3×3 (f(F 1b ))
[0098] Among them, F2 means F 1b The features obtained after RepNCSP module and 3×3 convolution processing, conv 3×3 represents the convolution kernel of size 3×3, F 1b Indicates the second part.
[0099] According to the processed second part, feature fusion is performed by combining the RepNCSP module and the convolution layer to generate a fused feature map:
[0100] F3=conv 3×3 (f(F2))
[0101] Among them, F3 represents the feature obtained by processing F2 through the RepNCSP module and 3×3 convolution. 3×3 Represents a convolution kernel of size 3×3.
[0102] Concatenate the fused feature maps:
[0103]
[0104] Among them, output RepNCSPELAN4 Represents the final output of the RepNCSPELAN4 unit, Indicates splicing in the channel dimension, F 1a Indicates the first part.
[0105] It should be noted that after the basic information in the input feature map is extracted through the initial convolution, it is divided into two parts for processing: one part maintains the original feature stream, and the other part further mines deep features by introducing the RepNCSP module and the convolution operation. This strategy effectively utilizes the cross-stage branching advantages of CSP and the lightweight feature expression capabilities of the Rep module, improving the model's ability to capture feature details. Finally, by splicing and fusing the feature maps of multiple sub-branches, a fused feature map containing richer semantic information is constructed. This modular, parallel processing structure not only enhances the expressive power of the model, but also ensures a low number of parameters and computational load. It is very suitable for lightweight and efficient target detection tasks, especially in complex or low-light environments. It performs more stably.
[0106] In one possible implementation, the RepNCSP module is used to split the input into two branches.
[0107] The normal branch is used to perform normal convolution on the input.
[0108] The complex branch processes the input by combining convolution processing and RepNBottleneck modules.
[0109] Concatenate the results of ordinary branch processing and complex branch processing.
[0110] It should be noted that the RepNCSP module uses a dual-branch structure to process input feature maps, emphasizing its advantages in balancing feature expressiveness and computational efficiency. The standard branch rapidly extracts basic features through standard convolution, ensuring the speed and stability of information transfer. The complex branch combines convolution operations with the RepNBottleneck module to extract richer, deeper semantic features. The two paths complement each other in terms of expressiveness and efficiency, capturing both local details and high-level features. Ultimately, the results of the two branches are concatenated in the channel dimension, enhancing the feature map's information diversity and spatial understanding capabilities. This design significantly improves the model's ability to recognize targets in complex scenes and low-light environments while maintaining a lightweight structure, offering both speed and accuracy.
[0111] In one possible implementation, the RepNBottleneck module processes the input feature map through the RepConvN module and the ordinary convolution module.
[0112] Based on the processed input feature map, the output feature map is generated through residual connection:
[0113] output RepNBottleneck =input4+conv(RepConvN(input4))
[0114] Among them, conv means ordinary convolution processing, RepConvN means processing by RepConvN module, input4 means the input features received by RepNBottleneck module, output RepNBottleneck Represents the features output by the RepNBottleneck module.
[0115] It should be noted that the RepConvN module is divided into three branches, two branches perform ordinary convolution operations, and one branch performs normalization operations.
[0116] The output feature maps after convolution and normalization are added together and nonlinear activation is performed:
[0117] output RepConvN =SiLU(conv 1×1 (input5)+conv 3×3 (input5)+bn(input5))
[0118] Among them, SiLU represents the SiLU activation function, bn represents the bn normalization process, input5 represents the input received by the RepConvN module, and output RepConvN Represents the output features of the RepConvN module.
[0119] It should be noted that the RepNBottleneck module combines RepConvN and ordinary convolution, effectively reducing computational overhead while maintaining feature extraction capabilities. Through the design of residual connections, the model can retain more original information and avoid the gradient vanishing problem, improving training stability and efficiency. After the output feature map is convolutional and normalized, the accuracy and consistency of the features are further guaranteed. Finally, by using the SiLU activation function, the nonlinear mapping capability is enhanced, giving the model stronger fitting capabilities in complex environments. Overall, this structural optimization enables the model to find a good balance between computational efficiency, training stability, and performance improvement, and is particularly suitable for tasks that require efficient processing and deep feature learning.
[0120] In one possible implementation, the PDetect detection head performs filtering processing on key channels through partial convolution.
[0121] It should be noted that the PDetect detection head uses partial convolution to filter the key channels of the input features. This structural design has significant advantages. First, by focusing on key channels rather than all channels, redundant calculations are reduced, the operating efficiency of the model is improved, and it is suitable for lightweight deployment needs. Secondly, partial convolution helps to highlight the feature information that is most sensitive to target detection, increase the model's attention to the core target area, and reduce invalid background interference. This strategy not only speeds up the detection speed, but also improves the detection accuracy and robustness of the model in low-light or complex environments to a certain extent. Overall, the PDetect detection head effectively reduces the computational burden while maintaining detection performance, providing good support for small devices or real-time application scenarios.
[0122] S4: Input the labeled low-light image dataset into the lightweight object detection model for training.
[0123] It's important to note that by training a lightweight object detection model on a dataset of annotated low-light images, the model learns target features while maintaining a compact structure and high computational efficiency. The weight control mechanism during training not only ensures the model's recognition accuracy but also helps limit its complexity, making it more suitable for deployment on resource-constrained devices. By continuously optimizing model parameters until the weights meet preset standards, a good balance between performance and efficiency can be achieved.
[0124] S5: Acquire a low-light image to be detected.
[0125] Low-light images to be detected refer to images captured in poor lighting conditions, such as nighttime streets, basements, and tunnels, that have not yet undergone target recognition processing. These images typically have low contrast, high noise, and a lack of detail, placing higher demands on target detection.
[0126] It's important to note that by acquiring real-world low-light images as model input, we can verify the trained lightweight model's detection performance and adaptability in real-world scenarios. This step ensures the system's real-time performance and practicality, enabling its widespread application in tasks such as nighttime surveillance, autonomous driving, and security inspections.
[0127] S6: Input the low-light image to be detected into the trained lightweight target detection model, and output the target detection result of the low-light image to be detected.
[0128] The trained lightweight object detection model is trained on labeled data and features a small number of parameters, low computational cost, and the ability to run on resource-constrained devices. The object detection result is the final output of locating and classifying various objects appearing in the input image, indicating which objects the model has identified, their locations, and their categories.
[0129] It's important to note that by feeding the image to be detected into a trained, lightweight model, the system efficiently outputs accurate object detection results, completing a closed-loop process from image acquisition to intelligent analysis. This process is not only responsive and suitable for real-time processing, but also boasts excellent deployment adaptability due to the lightweight model, running smoothly on mobile devices, monitoring terminals, and edge computing platforms.
[0130] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0131] In an embodiment of the present invention, a low-light image dataset is obtained and then annotated to ensure that the target information in the image can be learned. A lightweight target detection model is constructed using YOLOv8s as a framework, and the annotated low-light image dataset is input into the lightweight target detection model for training. Finally, in the actual application stage, the low-light image to be detected is obtained and input into the trained lightweight target detection model, which outputs the target detection result of the low-light image to be detected. This improves the detection accuracy and speed, enhances the robustness of the model, meets the needs of scenarios with high real-time requirements, and improves the recognition ability of the model in complex environments. It is of great significance for various scenarios such as night monitoring, unmanned driving, and security systems.
[0132] Reference Manual Figure 2 , shows a structural schematic diagram of a lightweight low-light target detection system provided by the present invention.
[0133] The present invention further provides a lightweight low-light target detection system 20, which is applied to the above-mentioned lightweight low-light target detection method, comprising:
[0134] Processor 201.
[0135] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201 , the lightweight low-light target detection method of the method embodiment is implemented.
[0136] The lightweight low-light target detection system 20 provided by the present invention can execute the above-mentioned lightweight low-light target detection method and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate on them.
[0137] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0138] In an embodiment of the present invention, a low-light image dataset is obtained and then annotated to ensure that the target information in the image can be learned. A lightweight target detection model is constructed using YOLOv8s as a framework, and the annotated low-light image dataset is input into the lightweight target detection model for training. Finally, in the actual application stage, the low-light image to be detected is obtained and input into the trained lightweight target detection model, which outputs the target detection result of the low-light image to be detected. This improves the detection accuracy and speed, enhances the robustness of the model, meets the needs of scenarios with high real-time requirements, and improves the recognition ability of the model in complex environments. It is of great significance for various scenarios such as night monitoring, unmanned driving, and security systems.
[0139] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), but may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0140] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0141] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function according to the embodiments of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available media can be magnetic media (such as floppy disks, hard disks, tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0142] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0143] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0144] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0145] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0146] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0147] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical, or other forms.
[0148] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0149] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0150] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program codes.
[0151] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the lightweight low-light target detection method according to the method embodiment is implemented.
[0152] The computer-readable storage medium provided by the present invention can implement the steps and effects of the lightweight low-light target detection method of the above method embodiment. To avoid repetition, the present invention will not elaborate on them.
[0153] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0154] In an embodiment of the present invention, a low-light image dataset is obtained and then annotated to ensure that the target information in the image can be learned. A lightweight target detection model is constructed using YOLOv8s as a framework, and the annotated low-light image dataset is input into the lightweight target detection model for training. Finally, in the actual application stage, the low-light image to be detected is obtained and input into the trained lightweight target detection model, which outputs the target detection result of the low-light image to be detected. This improves the detection accuracy and speed, enhances the robustness of the model, meets the needs of scenarios with high real-time requirements, and improves the recognition ability of the model in complex environments. It is of great significance for various scenarios such as night monitoring, unmanned driving, and security systems.
[0155] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
[0156] There are a few points to note:
[0157] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention. Other structures may refer to conventional designs.
[0158] (2) For the sake of clarity, the thickness of layers or regions in the drawings used to describe the embodiments of the present invention are exaggerated or reduced, that is, these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or intervening elements may be present.
[0159] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to form new embodiments.
[0160] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A lightweight low-light target detection method, characterized in that: include: S1: Acquire low-light image dataset; S2: Annotate the low-light image dataset; S3: Using YOLOv8s as a framework, a lightweight target detection model is constructed, wherein the lightweight target detection model includes: a multi-scale phantom convolution module, a backbone network module, a neck network module, and a detection head module; S4: Inputting the labeled low-light image dataset into the lightweight object detection model for training; S5: Acquire the low-light image to be detected; S6: Input the low-light image to be detected into the trained lightweight target detection model, and output the target detection result of the low-light image to be detected.
2. The lightweight low-light target detection method according to claim 1, characterized in that: The S2 specifically includes: marking the low-light image dataset into a Yolo format; The YOLO format specifically includes: category number, target center position, target width and height information.
3. The lightweight low-light target detection method according to claim 1, characterized in that: The backbone network module includes: a MA-SPPF unit and a C2fMSG unit; The neck network module includes: RepNCSPELAN4 unit; The RepNCSPELAN4 unit combines the features of the CSP module and the ELAN module, and uses the RepConv module as a calculation module; The detection head module is specifically a PDetect detection head.
4. The lightweight low-light target detection method according to claim 1, characterized in that: The multi-scale phantom convolution module processes the input feature map obtained after processing the low-light image, specifically including: Through the convolution layer, the number of channels of the input feature map is compressed to generate an intrinsic feature map; Splitting the intrinsic feature map into a preset number of feature maps with the same number of channels; Performing convolution processing on the feature map to generate an output feature map; Splicing the intrinsic feature map and the output feature map in the channel dimension; Through convolution operation, multi-dimensional feature extraction is performed on the spliced feature map; The extracted multi-dimensional features are fused to generate the final output feature map, where the number of channels of the final output feature map is C2: Among them, output 多尺度幻影卷积 Represents the output of the multi-scale phantom convolution module, f1 and f2 both represent ordinary convolution, and input1 represents the input received by the multi-scale phantom convolution. represents a 3×3 group convolution, represents a 5×5 group convolution, Represents concatenation in the channel dimension.
5. The lightweight low-light target detection method according to claim 3, characterized in that: The MA-SPPF unit includes three branches; The first branch is used for global maximum pooling processing: y1 = Maxpool(input2); Among them, y1 represents the output feature of the first branch, Maxpool represents maximum pooling, and input2 represents the input feature received by the MA-SPPF unit; The second branch is used for global average pooling: y2 = Avgpool(input2); Among them, y2 represents the output feature of the second branch, and Avgpool represents average pooling; The third branch is the SPPF module branch: y3 = f(input2); Among them, y3 represents the output feature of the third branch, and f() represents the branch processing of the SPPF module; The feature map processed by the global maximum pooling, the feature map processed by the global average pooling, and the feature map processed by the third branch are spliced in the depth direction: in, Indicates splicing in the channel dimension, output MA-SPPF Represents the output of the MA-SPPF unit.
6. The lightweight low-light target detection method according to claim 3, characterized in that: The processing process of the RepNCSPELAN4 unit specifically includes: Perform convolution on the input feature map: F1=conv 1×1 (input3); Among them, F1 represents the feature after 1×1 convolution processing, conv 1×1 Represents a 1×1 convolution kernel, and input3 represents the input features received by the RepNCSPELAN4 unit; Split the input feature map after convolution into the first part and the second part; Combine the RepNCSP module and the convolutional layer to process the second part: F2=conv 3×3 (f(F 1b )); Among them, F2 means F 1b The features obtained after RepNCSP module and 3×3 convolution processing, conv 3×3 represents the convolution kernel of size 3×3, F 1b Indicates the second part; According to the processed second part, feature fusion is performed by combining the RepNCSP module and the convolution layer to generate a fused feature map: <h2 style=";text-align:left;direction:ltr">F3=conv<h2 style=";text-align:left;direction:ltr"> 3×3 <h2 style=";text-align:left;direction:ltr"> (f(F2)); Among them, F3 represents the feature obtained by processing F2 through the RepNCSP module and 3×3 convolution. 3×3 Represents a convolution kernel of size 3×3; The fused feature maps are spliced: Among them, output RepNCSPELAN4 Represents the final output of the RepNCSPELAN4 unit, Indicates splicing in the channel dimension, F 1a Indicates the first part.
7. The lightweight low-light target detection method according to claim 3, characterized in that: The RepNCSP module is used to split the input into two branches; The ordinary branch is used to perform ordinary convolution on the input; The complex branch processes the input by combining convolution processing and RepNBottleneck modules; Concatenate the results of ordinary branch processing and complex branch processing.
8. The lightweight low-light target detection method according to claim 7, characterized in that: The RepNBottleneck module processes the input feature map through the RepConvN module and the ordinary convolution module; Based on the processed input feature map, the output feature map is generated through residual connection: output RepNBottleneck =input4+conv(RepConvN(input4)) Among them, conv means ordinary convolution processing, RepConvN means processing by RepConvN module, input4 means the input features received by RepNBottleneck module, output RepNBottleneck Represents the features output by the RepNBottleneck module; Performing convolution and normalization operations on the output feature maps respectively; The output feature maps after convolution and normalization are added together and nonlinear activation is performed: output R epConv N =SiLU(conv 1×1 (input5)+conv 3×3 (input5)+bn(input5)) Where SiLU represents the SiLU activation function, bn represents the bn normalization process, input5 represents the input received by the RepConvN module, and output R epConv N Represents the output features of the RepConvN module.
9. The lightweight low-light target detection method according to claim 3, characterized in that: The PDetect detection head performs filtering processing on key channels through partial convolution.
10. A lightweight low-light target detection system, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the lightweight low-light target detection method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Graphite electrode defect detection method and system
CN121033006A
Lightweight target detection method and system for construction robot in tunnel dusk dust environment
CN121505243A
Tunnel dust environment construction robot lightweight target detection method and system
CN121505243B