Target detection method and system based on background difference, terminal and storage medium

By combining background subtraction-based target detection methods with Gaussian mixture background subtraction and transfer learning, the target detection model is optimized, solving the problem of low accuracy in personnel detection in complex environments and achieving high-precision target localization and detection.

CN121095623APending Publication Date: 2025-12-09SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511051901.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing target detection technologies have low positioning accuracy for personnel detection in complex monitoring scenarios, especially in enclosed and dimly lit environments with variable lighting, uniform clothing, frequent occlusion, and dynamically changing backgrounds, making accurate identification difficult.

Method used

A background subtraction-based target detection method is adopted. The background is modeled by acquiring the environmental characteristic parameters of the scene image, the target image is preprocessed by the Gaussian mixture background subtraction algorithm, a class-balanced target image dataset is constructed, and the target detection model is optimized and trained by the transfer learning strategy. The background subtraction results are then fused to output the target detection localization.

Benefits of technology

It improves the accuracy of personnel detection in complex environments, solves the problems of missed detection and false detection, and provides an intelligent detection solution for industrial safety management applicable to a variety of complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095623A_ABST
    Figure CN121095623A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection processing, and discloses a target detection method and system based on background differencing, a terminal and a storage medium, and the method comprises the steps: obtaining a scene image, analyzing environment characteristic parameters in the scene image, and determining background modeling parameters according to the environment characteristic parameters; obtaining a target image, preprocessing the target image by using the background modeling parameters and a Gaussian mixture background difference algorithm, and constructing a category-balanced target image data set according to the preprocessed image data; performing optimization training on the target detection model by adopting a transfer learning strategy according to the target image data set to obtain a target detection model after transfer learning; detecting an input monitoring image based on the target detection model after transfer learning, fusing a detection result with a background difference extracted by the Gaussian mixture model, and outputting a positioning result of target detection; according to the method, the personnel detection precision in complex environments such as closed dark environments and changeable illumination environments is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of target detection processing, and in particular to a target detection method and system based on background difference, a terminal and a storage medium. BACKGROUND

[0002] Target detection technology is widely used in intelligent security, autonomous driving, industrial production and other fields. In personnel management related scenarios, accurate personnel detection and positioning is the key to efficient management. For example, in places with dense crowds such as large shopping malls, factories, and schools, accurate personnel location information can optimize personnel scheduling and ensure the safe and orderly operation of the place. In terms of personnel trajectory analysis, it helps to uncover personnel activity patterns and provide data support for business layout adjustment and security planning.

[0003] In complex monitoring scenarios such as underground parking lots, mine tunnels, and old factory sites, personnel detection faces many challenges. For example, in underground parking lots, the interior is dim and uneven, vehicles frequently enter and move, causing the background to constantly change, and light reflection and shadows to overlap, which easily interferes with the detection algorithm's judgment of personnel targets; in mine tunnels, dust is pervasive, personnel wear uniform clothing, making visual features similar, and the narrow underground environment causes frequent mutual occlusion between personnel, making it difficult for traditional detection methods to accurately identify individuals; in old factory sites, equipment is dense and layout is chaotic, and complex background textures and irregular shapes of equipment easily cause personnel target features to be submerged, and the start and stop of equipment also causes dynamic changes in the background. These scenarios commonly have problems such as closed and dim environment, variable lighting, single personnel clothing, frequent occlusion, and dynamic background changes, which seriously affect the accuracy and reliability of personnel detection.

[0004] Therefore, the prior art still needs to be improved. SUMMARY

[0005] The technical problem to be solved by the present application is that, in view of the defects of the prior art, the present application provides a target detection method and system based on background difference, a terminal and a storage medium, to solve the problem of low positioning accuracy of existing target detection technology in complex monitoring scenarios.

[0006] The technical solution adopted by the present application to solve the technical problem is as follows: In a first aspect, the present application provides a target detection method based on background difference, comprising: obtaining a scene image, analyzing environmental characteristic parameters in the scene image, and determining background modeling parameters according to the environmental characteristic parameters; obtaining a target image, preprocessing the target image using the background modeling parameters and a mixed Gaussian background difference algorithm, and constructing a class-balanced target image dataset according to the preprocessed image data; Based on the target image dataset, the target detection model is optimized and trained using a transfer learning strategy to obtain the target detection model after transfer learning. The target detection model based on the transfer learning is used to detect the input surveillance image, and the detection results are fused with the background difference extracted by the Gaussian mixture model to output the target detection localization result.

[0007] In one implementation, acquiring a scene image, analyzing environmental characteristic parameters in the scene image, and determining background modeling parameters based on the environmental characteristic parameters include: The scene image is acquired, and the background change frequency, illumination stability, and target operation mode of the scene image are analyzed to obtain the environmental characteristic parameters. Based on the environmental characteristic parameters, a dynamic parameter adjustment strategy is used to determine the initial number of modeling frames, the number of Gaussian distributions, the learning rate update strategy, and the foreground judgment threshold for Gaussian mixture background modeling, thereby obtaining the background modeling parameters.

[0008] In one implementation, determining the initial number of modeling frames, the number of Gaussian distributions, the learning rate update strategy, and the foreground judgment threshold for Gaussian mixture background modeling based on the environmental characteristic parameters using a dynamic parameter adjustment strategy includes: Based on the illumination stability, a preset number of frames is selected as the initial modeling data to obtain the initial modeling frame number; Based on the multimodal features of the scene image, determine the number of Gaussian distributions superimposed to describe the pixel value distribution, and obtain the number of Gaussian distributions; Based on the dwell time of the target task, a dynamic adjustment mechanism for the corresponding learning rate is determined, resulting in the learning rate update strategy; Based on the pixel difference features between the target and the background in the scene image, a threshold for distinguishing dynamic targets from static background regions is determined through statistical analysis, thus obtaining the foreground judgment threshold.

[0009] In one implementation, acquiring the target image, preprocessing the target image using the background modeling parameters and the Gaussian mixture background difference algorithm, and constructing a class-balanced target image dataset based on the preprocessed image data includes: Acquire the target image; wherein, the target image is target operation image data covering different operation areas, multiple angles, and multiple operation postures; Based on the Gaussian mixture background difference algorithm, the target image is subjected to color space conversion, and the Gaussian mixture model is initialized according to the background modeling parameters to obtain the background model; The matching degree between the pixels in the current image and the background model is calculated frame by frame. A binary mask image is generated by threshold discrimination, and the mask edge is optimized by morphological operations to obtain the preprocessed image data. Using an image annotation tool, rectangular boxes are used to annotate the moving target regions in the preprocessed image data to obtain the class-balanced target image dataset.

[0010] In one implementation, the step of optimizing and training the target detection model using a transfer learning strategy based on the target image dataset to obtain a transfer-learned target detection model includes: Select a pre-training dataset with similar features to the current scene, and pre-train the target detection model using the pre-training dataset; Based on the class-balanced target image dataset, the pre-trained target detection model is fine-tuned using the transfer learning strategy to obtain the target detection model after transfer learning.

[0011] In one implementation, the step of fine-tuning the pre-trained target detection model using the transfer learning strategy to obtain the transferred-learned target detection model includes: The detection box loss weight, category loss weight, and confidence loss weight are dynamically adjusted based on the output of the pre-trained target detection model, and the model parameters are updated using an incremental training strategy to obtain the target detection model after transfer learning.

[0012] In one implementation, the target detection model based on the transfer learning performs detection on the input surveillance image, and fuses the detection result with the background difference extracted by the Gaussian mixture model to output the target detection localization result, including: The monitoring image is preprocessed based on the Gaussian mixture background difference algorithm, and a binary mask image is generated by the background modeling method to obtain the background difference mask. The target detection model after transfer learning is used to perform target detection on the monitoring image, and the target detection result includes the detection box coordinates and confidence score. Based on the integrity of the target region in the background difference mask, the confidence level output by the target detection model after transfer learning is corrected, and the corrected detection results are fused using a non-maximum suppression algorithm to output the final target detection localization result.

[0013] Secondly, the present invention provides a target detection system based on background difference, comprising: The background modeling module is used to acquire scene images, analyze environmental characteristic parameters in the scene images, and determine background modeling parameters based on the environmental characteristic parameters. The preprocessing module is used to acquire the target image, preprocess the target image using the background modeling parameters and the Gaussian mixture background difference algorithm, and construct a class-balanced target image dataset based on the preprocessed image data. The transfer learning module is used to optimize and train the target detection model using a transfer learning strategy based on the target image dataset, so as to obtain the target detection model after transfer learning. The fusion detection module is used to detect the input monitoring image based on the target detection model after transfer learning, and to fuse the detection result with the background difference extracted by the Gaussian mixture model to output the target detection localization result.

[0014] Thirdly, the present invention provides a terminal, comprising: a processor and a memory, wherein the memory stores a background difference-based target detection program, and the background difference-based target detection program, when executed by the processor, is used to implement the background difference-based target detection method as described in the first aspect.

[0015] Fourthly, the present invention also provides a medium, which is a computer-readable storage medium, storing a background difference-based target detection program, which, when executed by a processor, is used to implement the background difference-based target detection method as described in the first aspect.

[0016] The present invention, by employing the above technical solution, has the following effects: This invention effectively solves the problems of missed detection and false detection in traditional detection methods in complex environments such as closed and dimly lit areas and variable lighting conditions by synergistically optimizing background subtraction and target detection models. It improves the detection accuracy of personnel in uniform and occluded scenarios, and provides an intelligent detection solution for industrial safety management applicable to a variety of complex scenarios. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the target detection method based on background difference in this invention.

[0019] Figure 2 This is a schematic diagram of the steps involved in the target detection model for detecting personnel in this invention.

[0020] Figure 3 This is a schematic diagram illustrating the transfer learning of the YOLOv8 model based on a self-built dataset in this invention.

[0021] Figure 4 This is a functional schematic diagram of the terminal in one implementation of the present invention.

[0022] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0024] Exemplary methods Target detection technology has wide applications in fields such as intelligent security, autonomous driving, and industrial production. In personnel management scenarios, accurate personnel detection and positioning are key to achieving efficient management. For example, in densely populated places such as large shopping malls, factories, and schools, accurately obtaining personnel location information can optimize personnel scheduling and ensure the safe and orderly operation of the venue. In terms of personnel trajectory analysis, it helps to uncover patterns in personnel activity, providing data support for business layout adjustments and security planning.

[0025] Personnel detection faces numerous challenges in complex monitoring scenarios such as underground parking lots, mine tunnels, and old factory areas. Taking underground parking lots as an example, the dim and uneven lighting, frequent vehicle entry and exit, and constantly changing backgrounds, along with interplay of light reflection and shadow, easily interfere with detection algorithms' ability to identify personnel. In mine tunnels, dust fills the air, and uniform work clothes make personnel appear similar. The confined underground environment and frequent occlusion between personnel further complicate the identification process, making it difficult for traditional detection methods to accurately identify individuals. In old factory areas, densely packed and cluttered equipment, along with complex background textures and irregular shapes, can easily obscure personnel features. Furthermore, the starting and stopping of equipment dynamically alters the background. These scenarios commonly present problems such as enclosed and dimly lit environments, variable lighting, uniform personnel clothing, frequent occlusion, and dynamically changing backgrounds, severely impacting the accuracy and reliability of personnel detection.

[0026] To address the above technical problems, this invention provides a target detection method based on background subtraction. The method includes: acquiring a scene image, analyzing environmental characteristic parameters in the scene image, and determining background modeling parameters based on the environmental characteristic parameters; acquiring a target image, preprocessing the target image using the background modeling parameters and a Gaussian mixture background subtraction algorithm, and constructing a class-balanced target image dataset based on the preprocessed image data; optimizing and training a target detection model using a transfer learning strategy based on the target image dataset to obtain a transfer-learned target detection model; detecting an input surveillance image based on the transfer-learned target detection model, and fusing the detection result with the background subtraction extracted by the Gaussian mixture model to output the target detection localization result. This invention improves the accuracy of personnel detection in complex environments such as enclosed, dimly lit, and variable lighting conditions.

[0027] like Figure 1 As shown, this embodiment of the invention provides a target detection method based on background subtraction, including the following steps: Step S100: Acquire a scene image, analyze the environmental characteristic parameters in the scene image, and determine the background modeling parameters based on the environmental characteristic parameters.

[0028] In this embodiment, the background subtraction-based target detection method can be used in complex environments such as closed and dimly lit environments, variable lighting, uniform clothing, frequent occlusion, and dynamic background changes. In such complex environments, firstly, the monitoring scene is processed by background subtraction using a Gaussian mixture background modeling method to extract the moving target region to eliminate static background interference. Then, a personnel dataset based on a specific scenario is constructed, and the YOLOv8 model (a target detection model) is optimized and trained using a transfer learning strategy. By fusing the background subtraction results with the model output results after transfer learning, the accuracy and robustness of personnel detection are improved.

[0029] In this embodiment, the method mainly includes the following four stages: 1) Scene detection characteristic analysis and background modeling parameter determination stage; 2) The personnel dataset construction stage based on background difference; 3) Optimization stage of target detection model based on transfer learning; 4) Fusion detection stage of the transfer model in background difference processing.

[0030] In the scene detection feature analysis and background modeling parameter determination stage, the main task is to obtain the environmental characteristics of the shooting scene, and then determine the key parameters of Gaussian mixture background modeling based on the environmental characteristics of the shooting scene, so as to build the background model based on these key parameters.

[0031] Specifically, in one implementation of this embodiment, step S100 includes the following steps: Step S101: Acquire the scene image, analyze the background change frequency, illumination stability and target operation mode of the scene image, and obtain the environmental characteristic parameters; Step S102: Based on the environmental characteristic parameters, a dynamic parameter adjustment strategy is adopted to determine the initial modeling frame number, Gaussian distribution number, learning rate update strategy, and foreground judgment threshold for Gaussian mixture background modeling, thereby obtaining the background modeling parameters.

[0032] In this embodiment, it is first necessary to acquire scene images of the shooting location, and then obtain environmental characteristic parameters of the shooting scene based on these scene images. These environmental characteristic parameters include: background change frequency, illumination stability, and personnel operation mode (i.e., target operation mode). Based on these environmental characteristic parameters, the key parameters for Gaussian mixture background modeling are determined by using the personnel operation mode, such as the length of time the personnel stay, the background change frequency, and the illumination stability, to adapt to the dynamic background changes of complex scenes (e.g., tunnel scenes). These key parameters include: initial modeling frame count, number of Gaussian distributions, learning rate update strategy, and foreground judgment threshold.

[0033] In one implementation of this embodiment, step S102 includes the following steps: Step S102a: Select a preset number of frames as the initial modeling data based on the illumination stability to obtain the initial modeling frame number; Step S102b: Based on the multimodal features of the scene image, determine the number of Gaussian distributions superimposed to describe the pixel value distribution, and obtain the number of Gaussian distributions; Step S102c: Based on the dwell time of the target task, determine the corresponding dynamic adjustment mechanism of the learning rate to obtain the learning rate update strategy; Step S102d: Based on the pixel difference features between the target and the background in the scene image, a threshold for distinguishing dynamic targets from static background areas is determined through statistical analysis, thereby obtaining the foreground judgment threshold.

[0034] Specifically, in this embodiment, a dynamic parameter adjustment strategy can be used to determine the key parameters for Gaussian mixture background modeling: Initial modeling frame count: Based on the stability of the shooting scene background, select a sufficient number of frames (e.g., hundreds of frames) as the initial modeling data to ensure that the constructed background model can cover the initial scene information; Number of Gaussian distributions: Considering the multimodal features of the tunnel background (e.g., fixed equipment, slowly moving objects), 3-5 Gaussian distributions are superimposed to describe the pixel value distribution, balancing model complexity and adaptability; Learning rate update strategy: Design a learning rate that is related to the time people spend on the job. The dynamic adjustment mechanism, in which the learning rate formula is: ,in The learning rate for Gaussian mixture modeling, The frame rate of the video recording.

[0035] The Gaussian mixture background model updates the background slowly with a low learning rate to avoid misclassifying people who linger for a long time as background. Foreground judgment threshold: The threshold is determined through statistical analysis by combining the pixel difference features between people and the background in the tunnel scene. Distinguish between dynamic personnel targets and static background areas.

[0036] As an example, in a practical application scenario, during the tunnel scene detection characteristic analysis and background modeling parameter determination stage, the background changes in the current shooting scene are analyzed to determine the number of Gaussian components (K) in the Gaussian background difference mixture and the number of frames for initial background modeling. Then, based on the movement and changes of the foreground target in the shooting scene, the dynamic update learning rate in the Gaussian background difference is determined. Furthermore, based on the complexity of the background texture in the shooting scene, the foreground judgment threshold T is determined. Based on the sensor noise characteristics in the shooting scene (the higher the camera quality, the less noise in the shooting scene image), the variance range coefficient in the Gaussian background modeling is determined.

[0037] The following section, using the above application scenarios as an example, explains the key parameters for Gaussian mixture background modeling through specific implementation methods: First, based on the characteristic of slow background changes in a closed tunnel environment, the number of Gaussian components in the mixed Gaussian background difference is determined. To adapt to the multimodal features of the background and balance computational efficiency.

[0038] Then, based on the stable background characteristics in the initial stage of tunnel monitoring, 300 frames of images were selected as the initial background modeling frame number to ensure coverage of subtle changes in illumination and to eliminate interference from dynamic targets.

[0039] Then, considering the characteristic of long working hours for tunnel construction workers, the formula is used: Calculate the dynamically updated learning rate This ensures a balance between static target recognition and gradual background updates.

[0040] Finally, given the simple background texture of the tunnel and the high contrast between the foreground and background, a foreground judgment threshold was determined. This is to reduce noise interference and avoid misjudging background pixels.

[0041] Based on the high image quality and low noise sensor characteristics of the Ezviz CS-CB2 camera used in the experiment, the variance range coefficient λ=3 was set to accommodate illumination fluctuations and eliminate random noise through the 3σ principle.

[0042] like Figure 1 As shown, this embodiment of the invention provides a target detection method based on background subtraction, including the following steps: Step S200: Obtain the target image, preprocess the target image using the background modeling parameters and Gaussian mixture background difference algorithm, and construct a class-balanced target image dataset based on the preprocessed image data.

[0043] In this embodiment, during the construction phase of the personnel dataset based on background subtraction, firstly, on-site images of personnel (i.e., target images) are acquired, and the images are preprocessed using the Gaussian mixture background subtraction algorithm; then, moving target regions are extracted and static background interference is removed before classification and labeling are performed to construct a scene personnel image dataset (i.e., target image dataset) with balanced categories.

[0044] Specifically, in one implementation of this embodiment, step S200 includes the following steps: Step S201: Obtain the target image; wherein, the target image is target operation image data covering different operation areas, multiple angles, and multiple operation postures; Step S202: Based on the Gaussian mixture background difference algorithm, the target image is converted to a different color space, and the Gaussian mixture model is initialized according to the background modeling parameters to obtain the background model; Step S203: Calculate the matching degree between the pixels in the current image and the background model frame by frame, generate a binary mask image by threshold discrimination, and optimize the mask edge by morphological operations to obtain the preprocessed image data; Step S204: Based on the image annotation tool, use rectangular boxes to annotate the moving target regions in the preprocessed image data to obtain the class-balanced target image dataset.

[0045] Specifically, in this embodiment, an image sequence containing personnel working scenes is first collected by on-site monitoring equipment to obtain image data covering different working areas, multiple angles (front / side / back) and typical working postures.

[0046] Secondly, a Gaussian mixture background subtraction algorithm is used for image preprocessing: the RGB image is converted to the HSV color space to reduce the impact of illumination changes. A Gaussian mixture model is initialized based on the characteristics of the tunnel scene (e.g., setting 3-5 Gaussian distributions). The matching degree between the current pixel and the background model is calculated frame by frame. A binary mask image is generated through thresholding, where white areas represent moving targets (personnel) and black areas represent static background. Morphological operations (such as dilation and erosion) are used to optimize the mask edges, eliminate noise holes, and ensure the integrity of the moving target's contour.

[0047] The preprocessed images are labeled: the LabelImg tool (i.e., the image labeling tool) is used to label the moving target regions in the mask image with rectangular boxes to form a basic dataset suitable for people detection, providing standardized input for subsequent target detection model training.

[0048] As an example, in a real-world application scenario, Gaussian mixture background difference preprocessing is used: , , The Gaussian mixture model is used to perform background subtraction on tunnel monitoring videos, removing static backgrounds and retaining moving personnel targets, thereby reducing background noise interference.

[0049] HSV color space conversion and contour detection: The differenced image is converted to the HSV space to enhance color stability. The contour detection algorithm is used to extract pixel information of the personnel area and accurately define the target boundary.

[0050] LabeleImg tool annotation and dataset construction: LabeleImg was used to annotate personnel targets according to the YOLO input format (e.g., {cls, xc, yc, w, h}). The 2600 instances were divided into a test set (260 cases) and a training set (2340 cases) in a 1:9 ratio to form a tunnel-specific detection dataset.

[0051] like Figure 1 As shown, this embodiment of the invention provides a target detection method based on background subtraction, including the following steps: Step S300: Based on the target image dataset, the target detection model is optimized and trained using a transfer learning strategy to obtain the target detection model after transfer learning.

[0052] In this embodiment, the optimization stage of the target detection model based on transfer learning mainly adopts the transfer learning strategy for optimization. Specifically, the target detection model is first pre-trained based on a similar scene dataset. Then, the pre-trained model is fine-tuned on the target image dataset constructed in step S200. By optimizing the detection box loss weight, category loss weight, and confidence loss weight, the model's ability to extract features of people and its robustness to small target detection are enhanced.

[0053] Specifically, in one implementation of this embodiment, step S300 includes the following steps: Step S301: Select a pre-training dataset with similar features to the current scene, and pre-train the target detection model using the pre-training dataset; Step S302: Based on the class-balanced target image dataset, the pre-trained target detection model is fine-tuned using the transfer learning strategy to obtain the target detection model after transfer learning.

[0054] In one implementation of this embodiment, step S302 includes the following steps: Step S302a: Dynamically adjust the detection box loss weight, category loss weight, and confidence loss weight according to the output of the pre-trained target detection model, and update the model parameters using an incremental training strategy to obtain the target detection model after transfer learning.

[0055] In this embodiment, a dataset with similar characteristics to the scene (i.e., a similar scene dataset, such as a construction worker safety management dataset in a construction scene) is first selected as the pre-training data source. This type of dataset contains labeled samples of uniformly dressed personnel, complex backgrounds, and partially occluded scenes, which can provide basic feature representations for personnel detection in the object detection model. The object detection model is pre-trained using this dataset, enabling the model to initially learn common features such as personnel contours, postures, and clothing, forming an initial parameter matrix with strong generalization ability.

[0056] Secondly, the pre-trained model is transferred to the personnel dataset (i.e., the target image dataset) constructed after step S200 for fine-tuning training. Specifically, the detection performance is optimized by dynamically adjusting the weights of the loss function: 1) Increase the weight of the detection box loss to enhance the target localization accuracy and ensure more accurate bounding box regression for uniformed personnel; 2) Reduce the weight of the category loss (ClsLoss) and adjust the classification threshold to avoid misclassification due to clothing similarity; 3) Optimize the weight of confidence loss (ObjLoss) to enhance the model's ability to judge the existence of occluded targets, thereby improving the model's detection accuracy of uniformly dressed people and recall rate of occluded targets in tunnel scenarios.

[0057] In the process of transfer learning, this embodiment adopts an incremental training strategy to cope with the dynamic changes in the tunnel construction scenario. When a new work area is added or the equipment layout is adjusted, the model parameters are updated through incremental training with small samples to maintain the model's adaptability to new occlusion patterns and environmental changes, and to ensure the continuous stability of detection performance.

[0058] As an example, in a practical application scenario, during the target detection model optimization stage based on transfer learning, the existing open-source construction worker target detection dataset, Safety Helmet Detection (which includes over 26,000 construction worker images), is selected to pre-train the target detection model, resulting in a pre-trained model. Further, during fine-tuning, the model's hyperparameters—the weights of the loss function, classification loss, and classification error—are dynamically adjusted. Background subtraction is used to reduce noise in the input data, improving the model's robustness against static backgrounds. This stage achieves a significant improvement in detection accuracy under small sample conditions through a hierarchical transfer learning strategy. The loss function weights are specifically adjusted to balance the requirements of detection accuracy and speed in tunnel scenarios.

[0059] The following section, using the above application scenarios as examples, explains the parameters for model optimization through specific implementation methods: like Figure 3 As shown, in the optimization model training phase: the open-source dataset SafetyHelmet Detection for construction worker detection is selected, and a pre-trained model based on YOLOv8 is trained to extract the common features of construction workers.

[0060] Further, the data from step S200 is used as the dataset (i.e., the self-built training set). The weight file of the pre-trained model is preloaded into the YOLOv8 model, and transfer training of the new object detection model begins. During training, the training hyperparameters are dynamically adjusted to suit the characteristics of the tunnel scene. The detection box loss weight (Box=0.7) is increased to enhance the accuracy of the bounding boxes and solve the problem of box overlap caused by dense crowds in the tunnel scene; the classification loss weight (Cls=0.2) is reduced because the scene only needs to identify the "Person" category (personal category), reducing interference from irrelevant features; the confidence loss weight (Obj=0.5) balances the judgment of object existence and detection efficiency. Learning rate and training epochs: A cosine annealing learning rate (initial lr=0.0001) is used for 200 training epochs to ensure that the model fully converges on the tunnel data, and finally, a new detection transfer optimization model is obtained.

[0061] like Figure 1 As shown, this embodiment of the invention provides a target detection method based on background subtraction, including the following steps: Step S400: Detect the input surveillance image based on the target detection model after transfer learning, and fuse the detection result with the background difference extracted by the Gaussian mixture model to output the target detection localization result.

[0062] In this embodiment, during the fusion detection stage of the transfer model in background subtraction processing, the moving target region extracted by Gaussian mixture background subtraction is fused with the detection results of the target detection model optimized by transfer learning. The detection box localization of the model is corrected by the background subtraction results. Furthermore, the transfer model is used to improve the target classification accuracy, realize collaborative detection of people, and improve the detection robustness in the shooting scene environment.

[0063] Specifically, in one implementation of this embodiment, step S300 includes the following steps: Step S401: The monitoring image is preprocessed based on the Gaussian mixture background difference algorithm, and a binary mask image is generated by the background modeling method to obtain the background difference mask; Step S402: Target detection is performed on the monitoring image using the target detection model after transfer learning, and the target detection result including the detection box coordinates and confidence score is output. Step S403: Based on the integrity of the target region in the background difference mask, the confidence level output by the target detection model after transfer learning is corrected, and the corrected detection results are fused through a non-maximum suppression algorithm to output the final target detection localization result.

[0064] In this embodiment, firstly, the real-time monitoring image input from the site is preprocessed using a Gaussian mixture background difference algorithm. A binary mask image is generated through background modeling to extract the moving target region and remove static background interference. Morphological operations are then used to optimize the mask contour, ensuring the integrity of the moving target and providing a high-recall foreground region benchmark for subsequent detection.

[0065] Secondly, the optimized object detection model, developed through transfer learning, is used to detect objects in the preprocessed image, outputting detection results including bounding box coordinates and confidence scores. This optimized object detection model enhances the ability to extract features of people in specific scenes through transfer learning strategies, improving classification accuracy and localization accuracy.

[0066] Finally, confidence level adjustment is performed: based on the integrity of the target region in the background difference mask, the confidence level of the model output is adaptively corrected. The corrected detection results are then fused using a non-maximum suppression algorithm to generate the final personnel detection output. This fusion strategy improves target recall through background difference and optimizes classification accuracy through transfer modeling, achieving a synergistic improvement in detection accuracy and recall in complex tunnel environments. This effectively solves the problems of missed detections and false positives in scenes with uniform clothing and occlusion.

[0067] As an example, in practical applications, the fusion detection stage of the transfer model for background subtraction mainly includes the following steps: input image to be detected, foreground target extraction, input after subtraction processing, and YOLOv8 detection (e.g., ...). Figure 2 As shown), specifically: Scene images (i.e., the images to be detected) are acquired through monitoring equipment. The image sequence is first input into a Gaussian mixture background model to detect foreground targets. The background target region image is processed frame by frame, and a binary mask is generated using dynamic thresholding. Then, morphological processing is performed: dilation and erosion operations are used to optimize the mask edges, eliminate noise holes, and remove background targets while ensuring the integrity of the target outline. Finally, the processed image (the image after differential processing) is input into the transfer detection model for detection, outputting detection boxes with coordinates and confidence scores, enhancing classification accuracy and localization accuracy. This stage uses a non-maximum suppression algorithm to fuse the corrected detection results, eliminate duplicate boxes, and generate the final detection output, achieving a synergistic improvement in accuracy and recall in complex tunnel environments.

[0068] The following section, based on the above application scenarios, explains the final detection process of tunnel monitoring image sequences through specific implementation methods: In the image processing and fusion detection stage of the Gaussian mixture background difference and migration detection model, the tunnel monitoring image sequence is input into the system for the following detection process: Foreground target region extraction: The static background is filtered by Gaussian background subtraction method to retain the moving person target, generate a binary background mask, and process it into an image that only retains the original pixel information of the foreground target; The processed image is input into the migration detection model for detection, and the output includes bounding boxes with coordinates and confidence scores, enhancing classification accuracy and localization accuracy. A non-maximum suppression algorithm is then used to fuse the corrected detection results, eliminating duplicate boxes and generating the final detection output. This embodiment achieves the following technical effects through the above technical solution: This embodiment effectively solves the problems of missed detection and false detection in traditional detection methods in complex environments such as closed and dimly lit areas and variable lighting by co-optimizing background subtraction and target detection models. It improves the detection accuracy of personnel in uniform and occluded scenes, and provides an intelligent detection solution for industrial safety management applicable to a variety of complex scenarios.

[0069] Exemplary device Based on the above embodiments, the present invention also provides a target detection system based on background difference, comprising: The background modeling module is used to acquire scene images, analyze environmental characteristic parameters in the scene images, and determine background modeling parameters based on the environmental characteristic parameters. The preprocessing module is used to acquire the target image, preprocess the target image using the background modeling parameters and the Gaussian mixture background difference algorithm, and construct a class-balanced target image dataset based on the preprocessed image data. The transfer learning module is used to optimize and train the target detection model using a transfer learning strategy based on the target image dataset, so as to obtain the target detection model after transfer learning. The fusion detection module is used to detect the input monitoring image based on the target detection model after transfer learning, and to fuse the detection result with the background difference extracted by the Gaussian mixture model to output the target detection localization result.

[0070] This embodiment achieves the following technical effects through the above technical solution: This embodiment effectively solves the problems of missed detection and false detection in traditional detection methods in complex environments such as closed and dimly lit areas and variable lighting by co-optimizing background subtraction and target detection models. It improves the detection accuracy of personnel in uniform and occluded scenes, and provides an intelligent detection solution for industrial safety management applicable to a variety of complex scenarios.

[0071] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 4 As shown.

[0072] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected via a system bus; wherein, the processor of the terminal provides computing and control capabilities; the memory of the terminal includes a storage medium and internal memory; the storage medium stores the operating system and computer programs; the internal memory provides an environment for the operation of the operating system and computer programs in the storage medium; the interface is used to connect to external devices; the display screen is used to display relevant information; and the communication module is used to communicate with a cloud server or other devices.

[0073] When executed by the processor, this computer program is used to implement the operation of a background difference-based target detection method.

[0074] It will be understood by those skilled in the art that Figure 4 The schematic diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0075] In one embodiment, a terminal is provided, comprising: a processor and a memory, the memory storing a background difference-based target detection program, which, when executed by the processor, is used to implement the operation of the background difference-based target detection method as described above.

[0076] In one embodiment, a storage medium is provided, wherein the storage medium stores a background difference-based target detection program, which, when executed by a processor, is used to implement the operation of the background difference-based target detection method described above.

[0077] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, database, or other media used in the embodiments provided by this invention can include both non-volatile and volatile memory.

[0078] In summary, this invention provides a target detection method, system, terminal, and storage medium based on background subtraction, comprising: acquiring a scene image, analyzing environmental characteristic parameters in the scene image, and determining background modeling parameters based on the environmental characteristic parameters; acquiring a target image, preprocessing the target image using the background modeling parameters and a Gaussian mixture background subtraction algorithm, and constructing a class-balanced target image dataset based on the preprocessed image data; optimizing and training a target detection model using a transfer learning strategy based on the target image dataset to obtain a target detection model after transfer learning; detecting an input surveillance image based on the target detection model after transfer learning, and fusing the detection result with the background subtraction extracted by the Gaussian mixture model to output the target detection localization result; this invention improves the accuracy of personnel detection in complex environments such as enclosed and dimly lit environments with variable lighting.

[0079] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A target detection method based on background subtraction, characterized in that, include: Acquire scene images, analyze environmental characteristic parameters in the scene images, and determine background modeling parameters based on the environmental characteristic parameters; The target image is acquired, and the target image is preprocessed using the background modeling parameters and the Gaussian mixture background difference algorithm. A class-balanced target image dataset is then constructed based on the preprocessed image data. Based on the target image dataset, the target detection model is optimized and trained using a transfer learning strategy to obtain the target detection model after transfer learning. The target detection model based on the transfer learning is used to detect the input surveillance image, and the detection results are fused with the background difference extracted by the Gaussian mixture model to output the target detection localization result.

2. The target detection method based on background difference according to claim 1, characterized in that, The process of acquiring a scene image, analyzing environmental characteristic parameters in the scene image, and determining background modeling parameters based on the environmental characteristic parameters includes: The scene image is acquired, and the background change frequency, illumination stability, and target operation mode of the scene image are analyzed to obtain the environmental characteristic parameters. Based on the environmental characteristic parameters, a dynamic parameter adjustment strategy is used to determine the initial number of modeling frames, the number of Gaussian distributions, the learning rate update strategy, and the foreground judgment threshold for Gaussian mixture background modeling, thereby obtaining the background modeling parameters.

3. The target detection method based on background difference according to claim 2, characterized in that, The step of determining the initial modeling frame number, Gaussian distribution number, learning rate update strategy, and foreground judgment threshold for Gaussian mixture background modeling based on the environmental characteristic parameters using a dynamic parameter adjustment strategy includes: Based on the illumination stability, a preset number of frames is selected as the initial modeling data to obtain the initial modeling frame number; Based on the multimodal features of the scene image, determine the number of Gaussian distributions superimposed to describe the pixel value distribution, and obtain the number of Gaussian distributions; Based on the dwell time of the target task, a dynamic adjustment mechanism for the corresponding learning rate is determined, resulting in the learning rate update strategy; Based on the pixel difference features between the target and the background in the scene image, a threshold for distinguishing dynamic targets from static background regions is determined through statistical analysis, thus obtaining the foreground judgment threshold.

4. The target detection method based on background difference according to claim 1, characterized in that, The process of acquiring the target image involves preprocessing the target image using the background modeling parameters and the Gaussian mixture background difference algorithm, and constructing a class-balanced target image dataset based on the preprocessed image data, including: Acquire the target image; wherein, the target image is target operation image data covering different operation areas, multiple angles, and multiple operation postures; Based on the Gaussian mixture background difference algorithm, the target image is subjected to color space conversion, and the Gaussian mixture model is initialized according to the background modeling parameters to obtain the background model; The matching degree between the pixels in the current image and the background model is calculated frame by frame. A binary mask image is generated by threshold discrimination, and the mask edge is optimized by morphological operations to obtain the preprocessed image data. Using an image annotation tool, rectangular boxes are used to annotate the moving target regions in the preprocessed image data to obtain the class-balanced target image dataset.

5. The target detection method based on background difference according to claim 1, characterized in that, The step of optimizing and training the target detection model using a transfer learning strategy based on the target image dataset to obtain a transfer-learned target detection model includes: Select a pre-training dataset with similar features to the current scene, and pre-train the target detection model using the pre-training dataset; Based on the class-balanced target image dataset, the pre-trained target detection model is fine-tuned using the transfer learning strategy to obtain the target detection model after transfer learning.

6. The target detection method based on background difference according to claim 5, characterized in that, The step of fine-tuning the pre-trained target detection model using the transfer learning strategy to obtain the transferred-learned target detection model includes: The detection box loss weight, category loss weight, and confidence loss weight are dynamically adjusted based on the output of the pre-trained target detection model, and the model parameters are updated using an incremental training strategy to obtain the target detection model after transfer learning.

7. The target detection method based on background difference according to claim 1, characterized in that, The target detection model based on the transfer learning performs detection on the input surveillance image, and fuses the detection results with the background difference extracted by the Gaussian mixture model to output the target detection localization result, including: The monitoring image is preprocessed based on the Gaussian mixture background difference algorithm, and a binary mask image is generated by the background modeling method to obtain the background difference mask. The target detection model after transfer learning is used to perform target detection on the monitoring image, and the target detection result includes the detection box coordinates and confidence score. Based on the integrity of the target region in the background difference mask, the confidence level output by the target detection model after transfer learning is corrected, and the corrected detection results are fused using a non-maximum suppression algorithm to output the final target detection localization result.

8. A target detection system based on background difference, characterized in that, include: The background modeling module is used to acquire scene images, analyze environmental characteristic parameters in the scene images, and determine background modeling parameters based on the environmental characteristic parameters. The preprocessing module is used to acquire the target image, preprocess the target image using the background modeling parameters and the Gaussian mixture background difference algorithm, and construct a class-balanced target image dataset based on the preprocessed image data. The transfer learning module is used to optimize and train the target detection model using a transfer learning strategy based on the target image dataset, so as to obtain the target detection model after transfer learning. The fusion detection module is used to detect the input monitoring image based on the target detection model after transfer learning, and to fuse the detection result with the background difference extracted by the Gaussian mixture model to output the target detection localization result.

9. A terminal, characterized in that, include: The processor and memory, the memory storing a background difference-based target detection program, which, when executed by the processor, is used to implement the background difference-based target detection method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a background difference-based target detection program, which, when executed by a processor, is used to implement the background difference-based target detection method as described in any one of claims 1-7.