Target detection method and device, equipment and medium

By combining YOLOv8 and the Spatial Depth Transform Convolutional Module (SPD-Conv), the problem of fine-grained information loss in the detection of tiny objects in the laboratory is solved, enabling accurate and real-time detection of tiny objects in the laboratory, and improving the accuracy of detection and management efficiency.

CN120876832APending Publication Date: 2025-10-31INSPUR (SHANDONG) COMPUTER TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510998106.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional manual monitoring in laboratories suffers from slow response and high false negative rates, making it difficult to effectively detect small items, especially in low-resolution images or small object detection, where fine-grained information is severely lost, affecting the efficiency and quality of laboratory safety management.

Method used

By employing the YOLOv8 architecture combined with the Spatial Depth Transformation Convolutional Module (SPD-Conv) and utilizing target data augmentation algorithms, a target detection model is constructed. By replacing traditional stride and pooling operations, data diversity is enhanced, enabling accurate and real-time detection of tiny laboratory objects.

Benefits of technology

It enables precise, real-time detection of tiny laboratory items, improves the robustness of the model, reduces the probability of accidental damage, and enhances management standardization and root cause tracing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876832A_ABST
    Figure CN120876832A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method and device, equipment and a medium, and relates to the technical field of image processing. The method comprises the following steps: acquiring real-time image data of a to-be-detected area; inputting the real-time image data into a pre-trained target detection model to obtain a target detection result; the target detection model is a model constructed based on a YOLOv8 architecture and combined with a space depth conversion convolution module, and the target detection model performs data enhancement by adopting a target data enhancement algorithm; and performing multi-path display on the target detection result on a webpage management interface, and judging whether to give an alarm for the to-be-detected area or not according to the target detection result. According to the technical scheme of the invention, small objects in a laboratory scene can be detected, and whether potential safety hazards exist or not can be judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a target detection method, apparatus, device, and medium. Background Technology

[0002] In laboratory safety management, traditional manual monitoring suffers from drawbacks such as slow response and high false negative rates, increasing accident risks and impacting work quality and efficiency. While traditional CNNs (Convolutional Neural Networks) (including some algorithms based on the YOLO series) perform well in behavior detection with the development of deep learning technology, their performance degrades in low-resolution images or small object detection due to the loss of fine-grained information caused by stride convolution and pooling layer downsampling operations, making them unsuitable for detecting tiny items such as screws and optical modules in laboratories. Although techniques such as MultiSEAM (Multi-Scale Spatially Enhanced Attention Module) improve the resolution of shallow feature maps through multi-level feature pyramids, their specificity and effectiveness in detecting tiny items in laboratory settings remain insufficient. Therefore, solving the problem of fine-grained information loss and achieving accurate, real-time detection of tiny laboratory items to compensate for the shortcomings of manual monitoring is a problem that needs to be addressed by those skilled in the art. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a target detection method, apparatus, device, and medium capable of achieving accurate and real-time detection of tiny laboratory items. The specific solution is as follows:

[0004] Firstly, this application discloses a target detection method, comprising:

[0005] Acquire real-time image data of the area to be detected;

[0006] Real-time image data is input into a pre-trained object detection model to obtain object detection results. The object detection model is a model built based on the YOLOv8 architecture, combined with a spatial depth transformation convolution module, and the object detection model uses object data augmentation algorithms for data augmentation.

[0007] The target detection results are displayed in multiple channels on the web management interface, and an alarm is triggered for the area to be detected based on the target detection results.

[0008] Optionally, before acquiring real-time image data of the region to be detected, the following steps are also included:

[0009] Historical image data of the area to be detected is acquired, and the historical image data is frame-by-frame acquired to obtain different frame images;

[0010] The different states in which the target object appears in the frame image are labeled, and the labeled frame images are used to construct a training dataset;

[0011] The training dataset is augmented using targeted data augmentation algorithms to generate an expanded dataset;

[0012] Obtain the source code files of the YOLOv8 model and modify them to create an environment for using the spatial depth transformation convolution module;

[0013] After modifying the source code file, obtain the initial configuration file of the YOLOv8 model and add the target structure to the initial configuration file to obtain the modified configuration file; where the target structure is the operation structure in the spatial depth transformation convolution module used to transform the image size spatial dimension to the depth dimension;

[0014] The model to be trained is built based on the modified configuration file, and the code used to train the model is run. The model is trained using the expanded dataset to obtain the object detection model.

[0015] Optionally, after acquiring different frame images from historical image data, the process may also include:

[0016] Select a frame image from the frame image that includes multiple labelable objects and each object has distinguishable feature information as the target frame image, and standardize the format, pixels and aspect ratio of the target frame image.

[0017] Accordingly, the different states in which the target object appears in the frame images are labeled, and the labeled frame images are used to construct a training dataset, including:

[0018] The different states of the target object in the target frame image after uniform normalization are labeled, and the labeled target frame image is used to construct a training dataset.

[0019] Optionally, the training dataset can be augmented using a targeted data augmentation algorithm to generate an expanded dataset, including:

[0020] Based on the Mosaic algorithm, and by adjusting the parameters in the Mosaic algorithm, simulated environmental features under the region to be detected are generated using the training dataset.

[0021] Image data corresponding to the simulated environmental features are added to the training dataset to generate an expanded dataset.

[0022] Optional object detection methods also include:

[0023] Acquire a preset number of new training data points in the region to be detected based on a preset time period, and use the new training data to incrementally train the already trained target detection model.

[0024] Optionally, the target detection results can be displayed in multiple ways on the web management interface, including:

[0025] A web-based management interface is built using a graphical user interface framework, and the target detection results are displayed in multiple ways on the web-based management interface; the graphical user interface framework is used to provide visual controls and event response mechanisms.

[0026] Optionally, based on the target detection results, determine whether to issue an alarm for the area to be detected, including:

[0027] When it is determined that an alarm should be triggered for the area to be detected based on the target detection results, the dangerous target in the alarm event should be identified.

[0028] Hazardous targets are classified and screened based on an event response mechanism to generate corresponding alarm logs, which are then displayed using visual controls.

[0029] Secondly, this application discloses a target detection device, comprising:

[0030] The image acquisition module is used to acquire real-time image data of the area to be detected;

[0031] The detection processing module is used to input real-time image data into a pre-trained target detection model to obtain target detection results. The target detection model is a model built based on the YOLOv8 architecture, combined with a spatial depth transformation convolution module, and the target detection model uses a target data augmentation algorithm for data augmentation.

[0032] The display output module is used to display the target detection results in multiple channels on the web management interface, and to determine whether to issue an alarm for the area to be detected based on the target detection results.

[0033] Thirdly, this application discloses an electronic device, which includes a processor and a memory; wherein the memory is used to store a computer program, which is loaded and executed by the processor to implement the aforementioned target detection method.

[0034] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the aforementioned target detection method.

[0035] This application provides a target detection method, comprising: acquiring real-time image data of the region to be detected; inputting the real-time image data into a pre-trained target detection model to obtain target detection results; the target detection model is a model constructed based on the YOLOv8 architecture and combined with a spatial depth transformation convolution module, and the target detection model uses a target data augmentation algorithm for data augmentation; displaying the target detection results in multiple channels on a web page management interface, and determining whether to issue an alarm for the region to be detected based on the target detection results.

[0036] Beneficial Effects: By combining YOLOv8 with the Spatial Depth Transformation Convolutional Module (SPD-Conv), the spatial-to-depth transformation and non-stretch convolution of SPD-Conv can replace the traditional stretch and pooling operations of YOLOv8, solving the problem of fine-grained information loss and achieving accurate, real-time detection of small objects in the laboratory. Simultaneously, target data augmentation algorithms increase data diversity and improve model robustness, particularly significantly enhancing the detection performance of small targets. This allows for the timely detection and handling of accidentally scattered or improperly stored small objects, effectively preventing laboratory accidents and reducing the probability of accidental damage to components. Furthermore, a web-based management interface was developed to display the target detection results in multiple channels, retaining corresponding image data to form logs for further alarm functions, playing a significant role in management standardization and root cause tracing.

[0037] In addition, the target detection device, equipment and storage medium provided in this application correspond to the above-mentioned target detection method and have the same effect. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0039] Figure 1 This is a flowchart of a target detection method disclosed in this application;

[0040] Figure 2 This is an example diagram of an interface for displaying test results disclosed in this application;

[0041] Figure 3 This application discloses a flowchart for detecting tiny objects in a laboratory setting.

[0042] Figure 4 This is a schematic diagram of the structure of a target detection device disclosed in this application;

[0043] Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] The demand for intelligent transformation in laboratory safety management has spurred the application of deep learning in areas such as object recognition. Compared to deep learning methods, traditional manual monitoring suffers from drawbacks such as slow response times (average alarm delay of 3-5 seconds) and high false negative rates. These shortcomings significantly increase the probability of laboratory accidents, reduce the quality of laboratory outputs, and affect the work efficiency of laboratory staff. Deep learning, through its visual data analysis capabilities, can capture the characteristics of violations in real time, greatly improving staff behavior in the laboratory and compensating for the shortcomings of manual monitoring.

[0046] With the development of deep learning technology, especially the widespread application of convolutional neural networks in the field of behavior detection, the accuracy and efficiency of behavior detection have been significantly improved. The YOLO series of algorithms, as one of the pioneers in behavior detection, has become an ideal choice for behavior detection due to its high recognition speed and good accuracy.

[0047] Traditional CNNs suffer from performance degradation in low-resolution image or small object detection tasks, primarily because the downsampling operations of stride convolutions and pooling layers lead to the loss of fine-grained information. For example, convolutions with a stride greater than 1 skip some pixels, while pooling layers compress the spatial dimension of feature maps. These operations can easily remove key features in small object or low-resolution scenes. SPD-Conv (Spatial Depth Transform Convolution) addresses the bottleneck problem of traditional CNNs in low-resolution and small object detection by replacing the stride and pooling operations in the original CNN algorithm with spatial-to-depth transformation and non-stride convolutions. It combines high efficiency and ease of use and has become an important improvement module in the field of computer vision.

[0048] The MultiSEAM algorithm module improves the resolution of shallow feature maps by constructing multi-level feature pyramids, thereby preserving pixel-level details of small objects. Compared to traditional pooling or downsampling methods, MultiSEAM enhances feature representation through channel recombination and residual connections. Current applications include: 1) Abnormal behavior detection: YOLOv8 can identify abnormal behaviors such as fighting and falling by analyzing real-time video streams; 2) Defect detection: In 3C electronic product production lines, YOLOv8 achieves micron-level scratch detection. Through multi-scale feature fusion technology, it classifies and detects defects such as metal surface oxidation and coating peeling. 3) Medical image analysis, tumor identification in CT / MRI images, supporting three-dimensional lesion volume calculation; microscopic cell classification and counting; 4) Ecological protection: Identification of nighttime poaching activities is achieved through thermal imaging and YOLOv8 fusion detection. 5) Smart home, bathroom anti-fog coating detection system is deployed through edge computing to achieve localized real-time monitoring; 6) Environmental perception, real-time detection of the number of pedestrians, vehicles, etc.; traffic sign recognition.

[0049] Currently, research on behavior detection is deepening. On the one hand, researchers are dedicated to optimizing deep learning models to improve the accuracy and real-time performance of behavior detection. For example, some studies utilize the network structure of the YOLO model, or combine it with other deep learning techniques, such as recurrent neural networks (RNNs), to enhance the model's ability to recognize behavioral features. On the other hand, the quality and diversity of datasets are also key factors affecting detection performance. In recent years, with the increase in publicly available video surveillance datasets, researchers have had the opportunity to train and test more accurate behavior detection models. Furthermore, applying object detection technology to mobile devices and edge computing has become a new research direction, which helps to achieve more flexible and widespread behavior detection applications.

[0050] It is evident that YOLO and other related deep learning technologies cover various scenarios such as abnormal behavior detection, defect detection, and medical image analysis, demonstrating strong practical value and feasibility in multiple fields.

[0051] To address this, this application provides a target detection scheme that combines YOLOv8 with SPD-Conv, utilizing SPD-Conv's space-to-depth transformation and non-stretch convolution to replace traditional stretch and pooling operations. This solves the problem of fine-grained information loss, enabling accurate and real-time detection of tiny laboratory items and compensating for the shortcomings of manual monitoring.

[0052] This invention discloses a target detection method, see [link to relevant documentation]. Figure 1 As shown, the method includes:

[0053] Step S11: Obtain real-time image data of the area to be detected.

[0054] In this embodiment, a scenario involving the detection of small objects in a laboratory is used as an example. Therefore, the real-time image data of the area to be detected can be video stream data of the laboratory monitored in real time by a camera. For example, a high-resolution camera, such as a camera with approximately 5 megapixels, can be used to monitor six tables arranged in two rows in the laboratory scene (each table can accommodate two servers horizontally) and the corresponding floor.

[0055] It is understandable that small items in a laboratory, such as screws, baffles, screwdriver bits, and optical modules located on lab tables, floors, and inside computer cases, could potentially damage circuit boards, cause workers to slip and fall, or dislodge resistors, capacitors, and other components if not promptly stored. Therefore, inspecting small items in a laboratory setting and determining whether they pose a safety hazard can effectively prevent laboratory accidents and reduce the probability of accidental component damage.

[0056] Step S12: Input the real-time image data into the pre-trained object detection model to obtain the object detection result; the object detection model is a model built based on the YOLOv8 architecture and combined with the spatial depth transformation convolution module, and the object detection model uses the object data augmentation algorithm for data augmentation.

[0057] As the foregoing is clear, YOLOv8 and related deep learning technologies possess strong capabilities across multiple domains. In this embodiment, a YOLOv8-SPD-based object detection model is pre-trained. By combining YOLOv8 with the Spatial Depth Transformation Convolutional Module (SPD-Conv), and utilizing SPD-Conv's spatial-to-depth transformation and non-stretch convolution to replace traditional stretch and pooling operations, the problem of information loss in low-resolution image and small object detection is addressed.

[0058] It is understandable that the detection accuracy of an object detection model directly affects its effectiveness; therefore, the quantity and quality of the dataset play a crucial role in model training, and the dataset should be as large as possible. In this embodiment, during model training, an object data augmentation algorithm is used to optimize the model. In one feasible implementation, the object data augmentation algorithm is the Mosaic data augmentation algorithm. Mosaic data augmentation generates new training samples by scaling, cropping, and stitching together four images. This method can increase data diversity and improve the robustness of the model, especially significantly improving the detection performance of small objects.

[0059] Updating and optimizing the model allows it to handle long-term detection tasks. The optimization algorithm for the model is constantly evolving. As people develop in the field of behavior detection, new algorithms can be added to the original algorithm to improve accuracy or speed up recognition.

[0060] Step S13: Display the target detection results in multiple channels on the web management interface, and determine whether to issue an alarm for the area to be detected based on the target detection results.

[0061] In this embodiment, after the target detection model detects small objects, the results are output to the terminal's web management interface for display. Therefore, the web management interface is built based on a graphical user interface framework, which provides visual controls and event response mechanisms.

[0062] In one feasible implementation, a web management interface is developed based on PyQt5. PyQt5 is a Python library for creating graphical user interfaces (GUIs). It is based on the Qt library, a C++ library for creating cross-platform applications. PyQt5 allows developers to create powerful applications using the Python language. Figure 2 The example shown is a web management interface developed using PyQt5 for recognizing smoking and drinking.

[0063] Furthermore, this graphical user interface framework provides visual controls and an event response mechanism. The visual controls can display multiple streams of data collected from simultaneous detection by multiple cameras; the event response mechanism can classify hazardous materials or behaviors, determine the type of behavior, and filter them. Past detection results and corresponding image data are retained to form a log, which, combined with the visual controls, generates an alarm log panel displaying the time of the event, improving management standardization and enabling root cause analysis. Specifically, when an alarm is triggered for a target area based on the detection results, the hazardous targets in the alarm event are identified; the hazardous targets are classified and filtered based on the event response mechanism to generate corresponding alarm logs, which are then displayed using visual controls.

[0064] It should be noted that since the detection scenario may change dynamically over time, the trained model should undergo incremental training with new data to adapt to the changed detection scenario. Specifically, a predetermined number of new training data points should be acquired within a preset time period for the region to be detected, and the trained target detection model should be incrementally trained using this new data. For example, approximately 200 new data points should be added each month for incremental training.

[0065] Beneficial Effects: By combining YOLOv8 with the Spatial Depth Transformation Convolutional Module (SPD-Conv), the spatial-to-depth transformation and non-stretch convolution of SPD-Conv can replace the traditional stretch and pooling operations of YOLOv8, solving the problem of fine-grained information loss and achieving accurate, real-time detection of small objects in the laboratory. Simultaneously, target data augmentation algorithms increase data diversity and improve model robustness, particularly significantly enhancing the detection performance of small targets. This allows for the timely detection and handling of accidentally scattered or improperly stored small objects, effectively preventing laboratory accidents and reducing the probability of accidental damage to components. Furthermore, a web-based management interface was developed to display the target detection results in multiple channels, retaining corresponding image data to form logs for further alarm functions, playing a significant role in management standardization and root cause tracing.

[0066] like Figure 3 The diagram illustrates the execution flow of a laboratory micro-object detection system based on the SPD-Conv algorithm-optimized YOLOv8 model. First, image data is collected by a monocular camera and input into a computer. Then, the YOLOv8-SPD model on the computer detects the objects. It determines whether the marked items in the image data are installed or placed correctly. If correctly placed, the detection box is green; if improperly placed, the detection box is red and an alarm is triggered. The final behavior type and classification results are output to the terminal's web management page for display. This allows for the timely detection and retrieval of accidentally scattered or improperly stored micro-objects.

[0067] Based on the above embodiments, this embodiment describes the construction and training process of the object detection model. Specifically, it includes the following steps: Step 1: Acquire historical image data of the area to be detected, and perform frame acquisition on the historical image data to obtain different frame images; Step 2: Label the different states in which the target object appears in the frame images, and use the labeled frame images to construct a training dataset.

[0068] In this embodiment, a dataset needs to be collected before model training. This can be done by communicating with the lab administrator and accessing historical video data from the lab, capturing frames to obtain different frame images as raw data. Further, the target objects in the frame images, i.e., small items such as screws, baffles, screwdriver bits, and optical modules, are labeled, including different states of the target objects. Taking a screw as an example, it can be labeled with three states: screw-on, screw-loss, and screw-box. Finally, the labeled frame images are used to construct the training dataset.

[0069] Understandably, the detection accuracy of a model directly affects its effectiveness; therefore, the quantity and quality of the dataset play a crucial role in model training. In one feasible implementation, to improve dataset quality, the acquired frame images are further screened. Each frame image should contain multiple annotable objects with distinct object features. Specifically, frame images containing multiple annotable objects, each with distinguishable feature information, are selected as target frame images. Then, the format, pixel count, and aspect ratio of the target frame images are standardized.

[0070] Step 3: Use target data augmentation algorithms to augment the training dataset to generate an expanded dataset.

[0071] In another feasible implementation, to increase the amount of data in the dataset, a target data augmentation algorithm is used to augment the training dataset to generate an expanded dataset. Specifically, based on the Mosaic algorithm, the parameters of the Mosaic algorithm are adjusted to generate simulated environmental features under the region to be detected using the training dataset; the image data corresponding to the simulated environmental features are then added to the training dataset to generate the expanded dataset.

[0072] In this embodiment, a Mosaic module can be added to the model configuration file to achieve specific functions by changing parameters. For example, it can simulate rack arrangement and cable obstruction. Exemplary parameter additions are as follows: transforms=[ Mosaic(prob=0.8, img_scale=(640,640), border=(-320,-320), # Increase the splicing area mixup_scale=(0.8,1.2)), # Simulate rack arrangement RandomPerspective(scale=(0.01,0.15), # Simulate cable obstruction degrees=0, translate=0.1).

[0073] The stitching probability was set to 0.8, the image scaling size was set to (640, 640), and the boundary padding value was set to (-320, -320) to increase the image stitching range and simulate the dense arrangement of laboratory cabinets. The mixed scale parameter was set to (0.8, 1.2) to introduce different scale transformations during image stitching, enhancing the model's adaptability to the diversity of cabinet arrangements. In conjunction with the RandomPerspective transformation, the scaling factor was set to (0.01, 0.15), the rotation angle to 0 degrees, and the translation factor to 0.1 to simulate the occlusion effect of cables on small objects. The scaling factor was used to control the intensity of the perspective transformation, enabling the model to learn the target features under different degrees of occlusion. The translation factor was used to introduce random displacement, enhancing the model's robustness to changes in target position.

[0074] Step 4: Obtain the source code file of the YOLOv8 model and modify the source code file to create the environment for using the spatial depth transformation convolution module;

[0075] Step 5: After modifying the source code files, obtain the initial configuration file of the YOLOv8 model and add the target structure to the initial configuration file to obtain the modified configuration file;

[0076] Step 6: Build the model to be trained based on the modified configuration file, and run the code to train the model. Use the expanded dataset to train the model to obtain the object detection model.

[0077] In this embodiment, the object detection model is the YOLOv8-SPD architecture. This architecture is based on the YOLOv8 architecture, with the original strided convolution replaced by the SPD-Conv module (space-to-depth transformation + non-strided convolution).

[0078] Specifically, SPD-Conv is added to the YOLOv8 model. First, the YOLOv8 source code files are modified to add support for SPD-Conv at the framework level, creating an environment for its use. This involves modifying the block.py, _init_.py, and tasks.py files in the source code.

[0079] After modification, open the initial YOLOv8 configuration file: yolov8.yaml, which defines the model's network architecture, including the type, parameters, and connection methods of each layer. Add the target structure under the network structure module in the yolov8.yaml file to obtain the modified configuration file. The network structure module defines the various components of the model and their connections. The target structure is the operation structure in the spatial depth transformation convolution module used to transform the image size spatial dimension to the depth dimension, named the "space_to_depth" structure. Specifically, it rearranges the spatial dimensions (height and width) of the input tensor and merges them into the depth dimension. This operation helps the model better capture local and global features, especially when processing low-resolution images and small objects, improving the detection performance for low-resolution images and small objects. This allows the model's SPD-Conv to be implemented.

[0080] Finally, load the configured model and run the code for training. Train the model using pre-labeled data, enabling it to detect tiny objects in real time using image data input from a monocular camera (approximately 5 megapixels).

[0081] In one feasible implementation, to optimize the model, the localization of specific parts can be added to the original model, such as the chassis, nearby tables, cabinets, etc., and additional detection can be performed on these specific parts. This optimizes the deployment of computing power and improves detection efficiency. Visual data and environmental perception data can also be integrated to increase the variety and quantity of data input. For example, adding infrared cameras and multiple cameras in the same detection scenario can transform the detection space into three dimensions, and the detection method can be changed from simple images to a combination of images and heat sources. Furthermore, a dynamic alarm mechanism with adaptive thresholds can be set based on the detection results of the target detection model. The alarm threshold is dynamically adjusted based on the type of laboratory equipment, operating time, and historical risk data. By constructing a risk level model, dynamic threshold coefficients are generated by combining equipment importance (e.g., precision instruments, ordinary operating tables), time period (peak experimental periods, unattended periods), and historical alarm accuracy. When a minor abnormality is detected, an alarm is triggered based on the current threshold coefficient (e.g., stricter thresholds for high-risk equipment, and appropriately relaxed thresholds for low-risk areas). The alarm processing results are recorded, and the threshold model is updated using reinforcement learning to continuously reduce the false alarm rate. This can help avoid excessive false alarms or omissions of key risks.

[0082] Accordingly, this application also discloses a target detection device, see [link to relevant documentation]. Figure 4 As shown, the device includes:

[0083] Image acquisition module 11 is used to acquire real-time image data of the area to be detected;

[0084] The detection processing module 12 is used to input real-time image data into a pre-trained target detection model to obtain target detection results. The target detection model is a model built based on the YOLOv8 architecture and combined with a spatial depth transformation convolution module. The target detection model also uses a target data augmentation algorithm for data augmentation.

[0085] The display output module 13 is used to display the target detection results in multiple channels on the web page management interface, and to determine whether to issue an alarm for the area to be detected based on the target detection results.

[0086] For more detailed information on the working process of each of the above modules, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0087] Therefore, the above-described solution in this embodiment, by combining YOLOv8 with the Spatial Depth Transformation Convolutional Module (SPD-Conv), utilizes SPD-Conv's spatial-to-depth transformation and non-strut convolution to replace YOLOv8's traditional strut and pooling operations, solving the problem of fine-grained information loss and achieving accurate, real-time detection of tiny objects in the laboratory. Simultaneously, target data augmentation algorithms increase data diversity and improve model robustness, particularly significantly enhancing the detection performance of small targets. This allows for the timely detection and retrieval of accidentally scattered or improperly stored tiny objects, effectively preventing laboratory accidents and reducing the probability of accidental damage to components. Furthermore, a web-based management interface is developed for multi-channel display of target detection results, which can retain corresponding image data to form logs for further alarm functions, playing a significant role in management standardization and root cause tracing.

[0088] In one specific embodiment, the target detection device further includes: a model building module, used for:

[0089] Historical image data of the area to be detected is acquired, and the historical image data is frame-by-frame acquired to obtain different frame images;

[0090] The different states in which the target object appears in the frame image are labeled, and the labeled frame images are used to construct a training dataset;

[0091] The training dataset is augmented using targeted data augmentation algorithms to generate an expanded dataset;

[0092] Obtain the source code files of the YOLOv8 model and modify them to create an environment for using the spatial depth transformation convolution module;

[0093] After modifying the source code file, obtain the initial configuration file of the YOLOv8 model and add the target structure to the initial configuration file to obtain the modified configuration file; where the target structure is the operation structure in the spatial depth transformation convolution module used to transform the image size spatial dimension to the depth dimension;

[0094] The model to be trained is built based on the modified configuration file, and the code used to train the model is run. The model is trained using the expanded dataset to obtain the object detection model.

[0095] In one specific implementation, the model building module includes: a preprocessing unit, used for:

[0096] Select a frame image from the frame image that includes multiple annotable objects and each object has distinguishable feature information as the target frame image, and standardize the format, pixels and aspect ratio of the target frame image.

[0097] In one specific implementation, the detection processing module includes: a data enhancement module, used for:

[0098] Based on the Mosaic algorithm, and by adjusting the parameters in the Mosaic algorithm, simulated environmental features under the region to be detected are generated using the training dataset.

[0099] Image data corresponding to the simulated environmental features are added to the training dataset to generate an expanded dataset.

[0100] In one specific embodiment, the target detection device further includes: an incremental training module, used for:

[0101] Acquire a preset number of new training data points in the region to be detected based on a preset time period, and use the new training data to incrementally train the already trained target detection model.

[0102] In one specific implementation, the display output module is specifically used for:

[0103] A web-based management interface is built using a graphical user interface framework, and the target detection results are displayed in multiple ways on the web-based management interface; the graphical user interface framework is used to provide visual controls and event response mechanisms.

[0104] When it is determined that an alarm should be triggered for the area to be detected based on the target detection results, the dangerous target in the alarm event should be identified.

[0105] Hazardous targets are classified and screened based on an event response mechanism to generate corresponding alarm logs, which are then displayed using visual controls.

[0106] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0107] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the target detection method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be a computer.

[0108] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0109] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored on it can include an operating system 221, computer programs 222, and data 223, etc. The data 223 can include various types of data. The storage method can be temporary storage or permanent storage.

[0110] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the target detection method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0111] Furthermore, this application also discloses a computer-readable storage medium, which includes random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, magnetic disks, optical disks, or any other form of storage medium known in the art. The computer program, when executed by a processor, implements the aforementioned target detection method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0112] Furthermore, embodiments of this application also provide a computer program product, including a computer program / instructions, which, when executed by a processor, implements any of the above-described methods for the target detection method.

[0113] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0114] The steps of the target detection method or algorithm described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0115] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0116] The above provides a detailed description of the target detection method, apparatus, device, and medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A target detection method, characterized in that, include: Acquire real-time image data of the area to be detected; The real-time image data is input into a pre-trained target detection model to obtain the target detection result; The target detection model is a model built on the YOLOv8 architecture, combined with a spatial depth transformation convolution module, and the target detection model uses a target data augmentation algorithm for data augmentation. The target detection results are displayed in multiple channels on the web page management interface, and an alarm is triggered for the area to be detected based on the target detection results.

2. The target detection method according to claim 1, characterized in that, Before acquiring the real-time image data of the region to be detected, the method further includes: Historical image data of the region to be detected is acquired, and the historical image data is frame-by-frame acquired to obtain different frame images; The different states in which the target object appears in the frame image are labeled, and a training dataset is constructed using the labeled frame images; The training dataset is augmented using a target data augmentation algorithm to generate an expanded dataset; Obtain the source code file of the YOLOv8 model and modify the source code file to create a usage environment for the spatial depth transformation convolution module; After the source code file is modified, the initial configuration file of the YOLOv8 model is obtained, and a target structure is added to the initial configuration file to obtain the modified configuration file; wherein, the target structure is the operation structure in the spatial depth transformation convolution module used to transform the image size spatial dimension to the depth dimension; The model to be trained is constructed based on the modified configuration file, and the code for training the model is run. The model is trained using the expanded dataset to obtain the object detection model.

3. The target detection method according to claim 2, characterized in that, After acquiring frames from the historical image data to obtain different frame images, the process further includes: Select a frame image from the frame image that includes multiple annotable objects and each object has distinguishable feature information as the target frame image, and standardize the format, pixels and aspect ratio of the target frame image. Accordingly, the step of labeling the different states in which the target object appears in the frame image, and constructing a training dataset using the labeled frame images, includes: The different states of the target object in the target frame image after uniform normalization are labeled, and the labeled target frame image is used to construct a training dataset.

4. The target detection method according to claim 2, characterized in that, The step of augmenting the training dataset using a target data augmentation algorithm to generate an expanded dataset includes: Based on the Mosaic algorithm, and by adjusting the parameters in the Mosaic algorithm, simulated environmental features under the region to be detected are generated using the training dataset. The image data corresponding to the simulated environment features are added to the training dataset to generate an expanded dataset.

5. The target detection method according to claim 1, characterized in that, Also includes: Based on a preset time period, a preset number of new training data are acquired in the region to be detected, and the new training data is used to incrementally train the target detection model that has already been trained.

6. The target detection method according to any one of claims 1 to 5, characterized in that, The step of displaying the target detection results in multiple channels on the web page management interface includes: A web page management interface is constructed based on a graphical user interface framework, and the target detection results are displayed in multiple ways on the web page management interface; the graphical user interface framework is used to provide visual controls and event response mechanisms.

7. The target detection method according to claim 6, characterized in that, Determining whether to issue an alarm for the area to be detected based on the target detection results includes: When it is determined, based on the target detection results, to trigger an alarm for the area to be detected, the dangerous target in the alarm event is identified. The dangerous targets are classified and screened based on the event response mechanism to generate corresponding alarm logs, and the alarm logs are displayed using the visualization control.

8. A target detection device, characterized in that, include: The image acquisition module is used to acquire real-time image data of the area to be detected; The detection processing module is used to input the real-time image data into a pre-trained target detection model to obtain target detection results; the target detection model is a model built based on the YOLOv8 architecture and combined with a spatial depth transformation convolution module, and the target detection model uses a target data augmentation algorithm for data augmentation; The display output module is used to display the target detection results in multiple channels on the web page management interface, and to determine whether to issue an alarm for the area to be detected based on the target detection results.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory; wherein the memory is used to store a computer program, which is loaded and executed by the processor to implement the target detection method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein the computer program, when executed by a processor, implements the target detection method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Failure identification method and device for lower lock pin component parts of railway wagon

    CN115527018A

  • Traffic sign board detection method based on improved YOLOv8s

    CN117392640A

  • Switch cabinet state identification method based on small target perception

    CN119810424A

  • X-ray security check image hazardous article segmentation and detection system based on improved YOLOv8

    CN119919664A