Method for detecting moving target, electronic device and storage medium

CN122223047BActive Publication Date: 2026-09-08JIANGSU MANYUN LOGISTICS INFORMATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610668042.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-09-08
Estimated Expiration
2046-05-15

AI Technical Summary

Technical Problem

[0005]本申请提供一种运动目标的检测方法、电子设备及存储介质,以解决在图像的背景存在噪声时,如摇曳的树枝、波光粼粼的水面以及雨雪噪点等,PAWCS算法无法区分这些噪声与真实的运动目标,从而无法准确地进行前景分割,导致误检率较高的问题,实现了噪声的抑制,提高前景分割的准确性,降低误检率

Benefits of technology

[0019] In a sixth aspect, this application provides a computer program product comprising: execution instructions stored in a readable storage medium, at least one processor of an electronic device being able to read the execution instructions from the readable storage medium, and the at least one processor executing the execution instructions causing the electronic device to implement a moving target detection method as described in the first aspect and any possible design of the first aspect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223047B_ABST
    Figure CN122223047B_ABST
Patent Text Reader

Abstract

The application provides a moving target detection method, an electronic device and a storage medium, and relates to the technical field of image processing. The method comprises the following steps: determining a current key frame image from a video stream and obtaining a binary foreground mask of the current key frame image, the current key frame image being an image that may have background interference; inputting the current key frame image and the binary foreground mask of the current key frame image into a large visual language model to obtain a foreground analysis result; updating foreground segmentation parameters according to the foreground analysis result to obtain updated foreground segmentation parameters; and performing foreground segmentation on each non-key frame image located after the current key frame image and before a next key frame image in the video stream according to the updated foreground segmentation parameters to obtain a binary foreground mask corresponding to each non-key frame image respectively and used for indicating the position of a moving target. Thus, the noise is suppressed, the accuracy of foreground segmentation is improved, and the false detection rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method for detecting moving targets, an electronic device, and a storage medium. Background Technology

[0002] Moving target detection is a fundamental technology in fields such as video surveillance and autonomous driving. The purpose of moving target detection is to separate and identify targets (such as pedestrians and vehicles) that change position relative to the background from each frame of a video stream, providing a foundation for tasks such as target tracking and behavior analysis.

[0003] Currently, mainstream detection methods primarily employ background subtraction-based algorithms. One example is the pixel-based adaptive word consensus segmenter (PAWCS) algorithm. This algorithm maintains a bag-of-words (BOD) for each pixel, storing historical color samples. It determines whether a pixel belongs to the foreground or background based on the degree of matching between the current pixel's color and the historical color samples, thus achieving foreground segmentation and obtaining a binary foreground mask. This binary foreground mask reflects the position of moving targets in the image, enabling moving target detection.

[0004] However, when there is noise in the background of the image, such as swaying tree branches, shimmering water, and rain and snow noise, the PAWCS algorithm cannot distinguish these noises from real moving targets, thus failing to accurately segment the foreground and resulting in a high false detection rate. Summary of the Invention

[0005] This application provides a method, electronic device, and storage medium for detecting moving targets, to address the problem that when there is background noise in an image, such as swaying tree branches, shimmering water surfaces, and rain and snow noise, the PAWCS algorithm cannot distinguish between this noise and real moving targets, thus failing to accurately segment the foreground and resulting in a high false detection rate. The method achieves noise suppression, improves the accuracy of foreground segmentation, and reduces the false detection rate.

[0006] Firstly, this application provides a method for detecting a moving target, the method comprising: The current keyframe image is determined from the video stream, and a binarized foreground mask of the current keyframe image is obtained; the video stream includes multiple consecutive images; The current keyframe image and its binarized foreground mask are input into the large visual language model to obtain the foreground analysis results. Based on the foreground analysis results, the foreground segmentation parameters are updated to obtain the updated foreground segmentation parameters; Based on the updated foreground segmentation parameters, foreground segmentation is performed on each non-keyframe image in the video stream that is located after the current keyframe image and before the next keyframe image, to obtain a binarized foreground mask corresponding to each non-keyframe image; wherein, the binarized foreground mask of the non-keyframe image is used to indicate the position of the moving target in the corresponding non-keyframe image.

[0007] In one possible design, the method further includes: When the next keyframe image is determined from the video stream, the next keyframe image is used as the new current keyframe image, and the steps of obtaining the binarized foreground mask of the current keyframe image, inputting it into the large visual language model, updating the foreground segmentation parameters, and performing foreground segmentation are continued.

[0008] In one possible design, the foreground analysis result includes the foreground authenticity determination result of the current keyframe image. The foreground segmentation parameters include a consensus threshold, a background update rate, and a historical bag of words. The consensus threshold and the historical bag of words are used to determine whether a pixel is foreground during foreground segmentation. The background update rate is used to update the historical bag of words. The historical bag of words includes a color sample library corresponding to each pixel position in the video stream. Each color sample library records the color value of the pixel position in the historical image.

[0009] In one possible design, updating the foreground segmentation parameters based on the foreground analysis results to obtain updated foreground segmentation parameters includes: When the foreground authenticity determination result of the current keyframe image is dynamic background interference, the consensus threshold is updated according to Formula 1 to obtain the updated consensus threshold, and the background update rate is updated according to Formula 2 to obtain the updated background update rate. Formula 1 is as follows: ; The updated consensus threshold. To perform the minimum value operation, This is the consensus threshold. As the first coefficient, This is the first upper limit value; Formula 2 is as follows: ; The updated background update rate. To perform the maximum value operation, For background update rate, As the second coefficient, This is the second upper limit value; When the foreground authenticity determination result of the current keyframe image is a ghost region, the color sample library corresponding to each pixel position in the ghost region is reset to obtain an updated bag of historical words.

[0010] In one possible design, determining the current keyframe image from the video stream and obtaining the binarized foreground mask of the current keyframe image includes: Images are acquired from the video stream in real time; Upon acquiring each image, a binarized foreground mask and a confidence map for that image are calculated; the confidence map is used to indicate the confidence level at which each pixel in the image is identified as a foreground. When it is determined that the image meets the preset conditions based on the binarized foreground mask and confidence map of the image, the image is determined as the current keyframe image, and the binarized foreground mask of the image is determined as the binarized foreground mask of the current keyframe image.

[0011] In one possible design, the preset conditions include: The binarized foreground mask of the image undergoes a sudden change in area relative to the binarized foreground mask of the previous image; or, The average confidence score of the image and the average confidence scores of the N consecutive images preceding and following the image are both less than a preset confidence threshold; where N is a positive integer, and the average confidence score is determined based on the image confidence map; or, A new connected component exists in the binarized foreground mask of the image.

[0012] In one possible design, after obtaining the binarized foreground mask of the current keyframe image, the method further includes: The current keyframe image and its binarized foreground mask are input into the large visual language model to obtain semantic analysis results; the semantic analysis results are used to indicate whether the processing type of each connected component of the binarized foreground mask of the current keyframe image is filtering or retention. Based on the semantic analysis results, the binarized foreground mask of the current keyframe image is corrected to obtain the corrected binarized foreground mask of the current keyframe image.

[0013] In one possible design, the step of correcting the binarized foreground mask of the current keyframe image based on the semantic analysis result to obtain the corrected binarized foreground mask of the current keyframe image includes: For a connected component of the binarized foreground mask of the current keyframe image, when the processing type of the connected component is filtering, all pixels in the connected component are set as background; when the processing type of the connected component is retention, all pixels in the connected component are set as foreground. Based on the correction processing of all connected components of the binarized foreground mask of the current keyframe image, the corrected binarized foreground mask of the current keyframe image is obtained.

[0014] Using the method provided in the first aspect, the current keyframe image is determined from the video stream, and its binarized foreground mask is obtained. The current keyframe image is an image that may contain background interference, thus identifying potentially interference-prone keyframe images from the video stream. This current keyframe image is used as a reference for subsequent semantic analysis. The current keyframe image and its binarized foreground mask are input into a large-scale visual language model to obtain foreground analysis results. This allows the semantic understanding capabilities of the large-scale visual language model to accurately distinguish between different foreground states in the keyframe image, such as real moving targets, dynamic background interference, or ghosting regions. Based on the foreground analysis results, the foreground segmentation parameters are updated, resulting in updated foreground segmentation parameters. This transforms the semantic cognition of the large-scale visual language model into dynamic updates of the foreground segmentation parameters, achieving adaptive parameter optimization so that the updated parameters adapt to changes in the current scene. Based on the updated foreground segmentation parameters, foreground segmentation is performed on each non-keyframe image in the video stream that is located after the current keyframe image and before the next keyframe image, resulting in a binarized foreground mask corresponding to each non-keyframe image, thereby achieving accurate foreground segmentation and reducing the false detection rate.

[0015] Secondly, this application provides a moving target detection device, comprising: a module for performing the moving target detection method in the first aspect and any possible design of the first aspect.

[0016] Thirdly, this application provides an electronic device including a first processor, which, when executing a computer-executable program or instructions in a memory, implements a method for detecting moving targets as described in the first aspect and any possible design of the first aspect.

[0017] Fourthly, this application provides an electronic device including at least one memory and at least one second processor. The memory stores a computer-executable program or instructions, and the second processor executes the computer-executable program or instructions to implement a moving target detection method as described in the first aspect and any possible design of the first aspect.

[0018] Fifthly, this application provides a computer-readable storage medium storing a computer-executable program or instructions, which, when executed by a processor, implement a method for detecting moving targets as described in the first aspect and any possible design of the first aspect.

[0019] In a sixth aspect, this application provides a computer program product comprising: execution instructions stored in a readable storage medium, at least one processor of an electronic device being able to read the execution instructions from the readable storage medium, and the at least one processor executing the execution instructions causing the electronic device to implement a moving target detection method as described in the first aspect and any possible design of the first aspect.

[0020] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description

[0021] Figure 1 This is a flowchart of a method for detecting moving targets provided in an embodiment of this application.

[0022] Figure 2 This is a flowchart illustrating a method for determining a current keyframe image and obtaining a binarized foreground mask of the current keyframe image, as provided in an embodiment of this application.

[0023] Figure 3 This is a schematic diagram of the structure of a moving target detection device provided in an embodiment of this application.

[0024] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 1 .

[0025] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 2 . Detailed Implementation

[0026] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c alone can mean: a alone, b alone, c alone, a combination of a and b, a combination of a and c, a combination of b and c, or a, b, and c, where a, b, and c can be single or multiple. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0027] The terms “center,” “longitudinal,” “lateral,” “up,” “down,” “left,” “right,” “front,” and “rear,” etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0028] The terms "connected" and "connected" should be interpreted broadly. For example, in circuit structures, "connected" or "connected" can refer not only to physical connections but also to electrical or signal connections. This could be a direct connection (physical connection) or an indirect connection via at least one intermediate component, as long as the circuit is connected. It could also refer to the internal connection between two components. Similarly, a signal connection can refer to a connection via a circuit or a medium, such as radio waves. Those skilled in the art will understand the specific meaning of these terms in this application based on the specific circumstances.

[0029] For example, this application provides a method, electronic device, and storage medium for detecting moving targets. By identifying keyframe images in a video stream that may contain background interference, semantic analysis is performed on the keyframe images using a large visual language model to obtain foreground analysis results. This introduces semantic understanding capabilities through the large visual language model, accurately distinguishing between background interference and real moving targets in the keyframe images. The foreground analysis results are then used to update the parameters used for foreground segmentation in the PAWCS algorithm, adapting the updated parameters to changes in the current scene. The updated parameters are then used to perform foreground segmentation on subsequent non-keyframe images in the video stream, obtaining corresponding binarized foreground masks, achieving accurate foreground segmentation and reducing the false detection rate.

[0030] The moving target detection method provided in this application can be executed by an electronic device or by a moving target detection device (hereinafter referred to as the detection device) in an electronic device.

[0031] Among them, electronic devices can be servers, desktop computers, mobile phones, tablets, laptops, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, etc.

[0032] The detection device can be implemented through a combination of software and / or hardware. For example, the detection device can be an application (APP), a webpage, or a public account. For the sake of simplicity, this application embodiment uses the execution of a detection device as an example for explanation.

[0033] Below, in conjunction with Figures 1 to 2 The method for detecting moving targets provided in the embodiments of this application will be described.

[0034] Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for detecting a moving target according to an embodiment of this application. Figure 1 As shown, the method includes: S101, The detection device determines the current keyframe image from the video stream and obtains the binarized foreground mask of the current keyframe image.

[0035] The video stream consists of multiple consecutive images.

[0036] The current keyframe image is an image that may contain background interference.

[0037] Specifically, the detection device can acquire video streams in real time. When acquiring each image, it uses a background subtraction algorithm (such as the PAWCS algorithm) to extract the foreground and obtain a binary foreground mask for each image.

[0038] The binarized foreground mask includes pixels in the foreground and pixels in the background. Pixels in the foreground may correspond to moving targets, or they may be background noise that is misidentified as foreground pixels.

[0039] Thus, the detection device can identify the current keyframe image based on the binarized foreground mask of the image, which may contain background interference.

[0040] Based on this, the detection device can identify the current keyframe image that may have background interference from the video stream, and use the current keyframe image as a reference to facilitate subsequent semantic analysis based on the current keyframe image, thereby improving the foreground segmentation quality of subsequent images based on the results of semantic analysis.

[0041] S102. The detection device inputs the current keyframe image and the binarized foreground mask of the current keyframe image into the large visual language model to obtain the foreground analysis results.

[0042] Considering that while the detection device uses a background subtraction algorithm to obtain a binarized foreground mask when determining keyframe images, this algorithm may misidentify background as foreground. This would affect the accuracy of the binarized foreground mask in identifying images with potential background interference, and the detection device would be unable to determine whether background interference actually exists or what the specific cause of the interference is. Therefore, the detection device can input the current keyframe image and its binarized foreground mask into a large vision-language model (LVLM) for semantic analysis, thereby enabling foreground analysis.

[0043] Among them, the large visual language model is a multimodal large model that integrates visual encoders and language models, and has the ability to understand scenes, make logical inferences and perform semantic analysis.

[0044] In some examples, the foreground analysis results include the foreground authenticity determination results of the current keyframe image.

[0045] The foreground authenticity determination result is used to describe whether the foreground part in the current keyframe image is a real moving target, dynamic background interference, or ghosting region.

[0046] In addition, the foreground analysis results may also include a scene context description of the current keyframe image, an analysis of the causes of interference, and target category labels.

[0047] The scene context description is used to describe the scene corresponding to the current keyframe image, such as "rainy day" or "wind blowing leaves".

[0048] Among them, the interference cause analysis is used to describe the reasons why there is background in the current keyframe image that is misidentified as foreground. For example, "The interference is due to the high-frequency jitter of values ​​caused by wind, and the change in pixel value exceeds the background threshold."

[0049] The target category label is used to describe the type of real moving target, such as "person", "dog", "car", "bicycle", "tree" and "flag".

[0050] Based on this, the detection device can obtain richer evidence for parameter updates based on the foreground analysis results, thereby further contributing to more accurate parameter updates.

[0051] Specifically, the detection device can stitch together the current keyframe image and its binarized foreground mask to obtain the stitching result, and construct a multimodal prompt template. The stitching result and the multimodal prompt template are then input into the large visual language model.

[0052] The prompt template is as follows: "You are a professional video surveillance scene analysis expert. Please perform the following analysis on the input image:" Analysis dimensions: Scene context description: Describe the lighting conditions (sunlight / night / backlight / fog) and weather conditions (sunny / rainy / snowy / strong wind); Foreground authenticity determination results: Determine the type of each highlighted area (real moving target / dynamic background interference / ghosting area); Analysis of interference causes: If the background interference is dynamic, describe the specific physical phenomenon (swaying leaves / ripples on the water / changes in light and shadow, etc.). Target category label: If it is a real moving target, label the specific category (person / vehicle / animal / unknown object). Output requirements: Strictly output standard JSON format, containing only the two primary keys: Scene and Regions; The Scene field contains: • Light (value: Daylight / Night / Backlight / Foggy); • Weather (value: Sunny / Rainy / Snowy / Windy / Foggy); The Regions field is an array, and each element must contain: • ID: Unique identifier for the region (an integer starting from 1 and incrementing); • Type (value: Real_Object / Dynamic_Noise / Ghost); • Category (Real moving objects: Person, Vehicle, Animal, Unknown_Object; Dynamic background interference: Tree_Shadow, Water_Ripple, Light_Shadow, Sensor_Noise, Curtain_Shake; Ghosting region: Ghost). • Action (values: Keep / Filter). The large visual language model can perform foreground analysis on the current keyframe image and output the foreground analysis results, so that the detection device can obtain the foreground analysis results.

[0053] Based on the above prompt template, the foreground analysis results are as follows: Scene context description: In windy weather, the surveillance footage shows: Left side: Tree branches swaying violently (PAWCS misdetects it as a moving target) In the middle: A car suddenly stops (just coming to a standstill from a moving position). Right side: Pedestrians running across quickly { "Scene": { "Light": "Daylight", "Weather": "Windy" }, "Regions": [ { "ID": 1, "Type": "Dynamic_Noise", "Category": "Tree_Shadow", "State": "Oscillating", "Action": "Filter" }, { "ID": 2, "Type": "Real_Object", "Category": "Vehicle", "State": "Stopped", Action: "Keep" }, { "ID": 3, "Type": "Real_Object", "Category": "Person", "State": "Fast_Moving", Action: "Keep" } ] }” Based on this, the detection device obtains the foreground analysis results, so as to update the foreground segmentation parameters according to the foreground analysis results.

[0054] S103. The detection device updates the foreground segmentation parameters based on the foreground analysis results to obtain the updated foreground segmentation parameters.

[0055] The foreground segmentation parameters refer to a set of adjustable variables in the background subtraction algorithm (such as the PAWCS algorithm) used for foreground segmentation.

[0056] In some examples, foreground segmentation parameters include consensus threshold, background update rate, and historical bag of words.

[0057] The consensus threshold and historical bag-of-words are used to determine whether a pixel is foreground during foreground segmentation. The background update rate is used to update the historical bag-of-words.

[0058] The history bag includes a color sample library corresponding to each pixel position in the video stream, and each color sample library records the color value of the pixel position in the historical image.

[0059] The following section explains the consensus threshold, background update rate, and historical bag-of-words through the specific steps of foreground segmentation.

[0060] Specifically, during foreground segmentation, for a pixel in an image, the distance (e.g., Euclidean distance) between the pixel's color value and each color value in the color sample library corresponding to that pixel in the historical bag-of-words system is calculated. If this distance is less than a preset color threshold, and the absolute value of the difference between this distance and the preset color threshold is less than a consensus threshold, then the pixel is determined to be background; otherwise, the pixel is determined to be foreground. After processing an image, the color values ​​in the historical bag-of-words system are randomly updated according to the background update rate, which represents the probability of randomly updating the historical bag-of-words system.

[0061] The consensus threshold can affect the sensitivity of foreground segmentation. Increasing the consensus threshold can decrease the sensitivity of foreground segmentation, while decreasing the consensus threshold can increase the sensitivity of foreground segmentation.

[0062] The background update rate controls the update speed of the historical bag of words. Increasing the background update rate increases the update speed of the historical bag of words, while decreasing the background update rate decreases the update speed of the historical bag of words.

[0063] The bag-of-history is used to determine whether a pixel is part of the background. When a ghosting region appears in an image, the bag-of-history of the ghosting region can be checked to eliminate its influence.

[0064] Ghosting occurs when a real moving object is removed, but the area originally covered by the real moving object is still marked as foreground. This is because the bag-of-history stores the color values ​​of the real moving object when it was present, and the background color of the current image does not match these samples, causing some pixels to be judged as foreground and thus creating ghosting.

[0065] The detection device can update the foreground segmentation parameters based on the foreground analysis results, so as to obtain foreground segmentation parameters that are more adapted to the current scene.

[0066] For example, when foreground analysis indicates the presence of dynamic background interference in the image, the detection device can increase the consensus threshold, making it more difficult for pixels related to the dynamic background to meet the background determination criteria. This allows for the filtering out of dynamic background interference in subsequent foreground segmentation. The detection device can also reduce the background update rate to prevent the color values ​​of pixels related to the dynamic background from being added to the historical bag-of-words, thus avoiding contamination of the historical bag-of-words.

[0067] When the foreground analysis results indicate the presence of ghosting regions in the image, the detection device can clear the color sample library corresponding to the pixel positions of the ghosting regions, thereby eliminating the ghosting.

[0068] Based on this, the detection device updates the foreground segmentation parameters according to the foreground analysis results, resulting in updated foreground segmentation parameters that are more suitable for the scene indicated by the current keyframe image. This enables adaptive parameter adjustment based on the current scene, thereby helping to achieve accurate foreground segmentation.

[0069] S104. The detection device performs foreground segmentation on each non-keyframe image in the video stream that is located after the current keyframe image and before the next keyframe image, based on the updated foreground segmentation parameters, and obtains the binarized foreground mask corresponding to each non-keyframe image.

[0070] In some examples, the detection device can use the PAWCS algorithm to perform foreground segmentation on each non-keyframe image to obtain a binarized foreground mask corresponding to each non-keyframe image.

[0071] The binarized foreground mask of the non-keyframe image is used to indicate the position of the moving target in the corresponding non-keyframe image.

[0072] For example, the updated foreground segmentation parameters include the updated historical bag-of-words and the updated consensus threshold. For a pixel in a non-keyframe image, the detection device calculates the distance (e.g., Euclidean distance) between the color value of that pixel and each color value in the color sample library of the corresponding pixel position in the updated historical bag-of-words. If the distance is less than a preset color threshold, and the absolute value of the difference between the distance and the preset color threshold is less than the updated consensus threshold, then the pixel is determined to be background; otherwise, the pixel is determined to be foreground. After all pixels are determined, the binarized foreground mask corresponding to the non-keyframe image is obtained.

[0073] Based on this, the detection device can utilize the foreground segmentation parameters updated after semantic analysis through a large visual language model to achieve foreground segmentation of non-keyframe images, thereby achieving accurate foreground segmentation and reducing the false detection rate.

[0074] In this embodiment, the detection device determines the current keyframe image from the video stream and obtains its binarized foreground mask. The current keyframe image may contain background interference, allowing the detection device to identify potentially interfering keyframe images from the video stream. This keyframe image is used as a reference for subsequent semantic analysis. The detection device inputs the current keyframe image and its binarized foreground mask into a large visual language model to obtain foreground analysis results. This allows the large visual language model to leverage its semantic understanding capabilities to accurately distinguish between different foreground states, such as real moving targets, dynamic background interference, or ghosting regions. Based on the foreground analysis results, the detection device updates the foreground segmentation parameters, transforming the semantic understanding of the large visual language model into dynamic updates of the foreground segmentation parameters. This achieves adaptive parameter optimization, ensuring the updated parameters adapt to changes in the current scene. The detection device performs foreground segmentation on each non-keyframe image in the video stream that is located after the current keyframe image and before the next keyframe image, based on the updated foreground segmentation parameters. This results in a binarized foreground mask for each non-keyframe image, thereby achieving accurate foreground segmentation and reducing the false detection rate.

[0075] Based on the above exemplary description, when the detection device determines the next keyframe image from the video stream, it uses the next keyframe image as the new current keyframe image and continues to execute steps S101 to S104.

[0076] In this embodiment of the application, the video stream is a continuous sequence of images acquired in real time, such as those generated in real time by a surveillance camera at a preset frame rate and transmitted frame by frame to the detection device.

[0077] After receiving the video stream, the detection device processes it frame by frame, starting from the first image of the video stream. If the keyframe image is not determined, the detection device can use the default foreground segmentation parameters to perform foreground segmentation.

[0078] When the first keyframe image is determined, the detection device takes the first keyframe image as the current keyframe image and starts processing S101 to S103 to obtain the updated foreground segmentation parameters.

[0079] Before determining the next keyframe image, the detection device uses the updated foreground segmentation parameters to perform foreground segmentation on each non-keyframe image in the video stream that is located after the current keyframe image and before the next keyframe image. In other words, from the image after the first keyframe image until the detection device determines the keyframe image again, all images in between are non-keyframe images, and foreground segmentation is performed using the updated foreground segmentation parameters.

[0080] When the detection device determines the next keyframe image, it takes the next keyframe image as the new current keyframe image and repeats S101 to S104. By repeating this cycle, continuous adaptive updates based on changes in background interference in the video stream can be achieved, thereby further improving the quality of foreground segmentation and reducing the false detection rate.

[0081] Based on the above exemplary description, the detection device can update the foreground segmentation parameters in different ways based on different foreground authenticity discrimination results.

[0082] The foreground authenticity determination results include dynamic background interference, real moving targets, and ghosting regions.

[0083] Dynamic background interference refers to the misidentification of high-frequency pixel changes of non-real moving targets as foreground elements due to changes in the natural environment or lighting. Examples include leaves blowing in the wind, shimmering water, and falling rain or snow.

[0084] In this context, a real moving target refers to an object that actually undergoes displacement and needs to be detected and monitored. Examples include vehicles, pedestrians, or animals.

[0085] Ghosting regions refer to areas that were originally occupied by real moving objects, but after the moving objects move away, these areas are incorrectly continued to be marked as foreground because the color values ​​of the real moving objects are still retained in the bag-of-history. For example, after a car drives away, the area where the car was parked is still marked as foreground.

[0086] Therefore, the detection device can update the foreground segmentation parameters in different ways depending on the situation.

[0087] When the foreground authenticity determination result of the current keyframe image is dynamic background interference, the detection device can update the consensus threshold according to the following formula 1 to obtain the updated consensus threshold, and update the background update rate according to the following formula 2 to obtain the updated background update rate.

[0088] Formula 1; in, The updated consensus threshold. To perform the minimum value operation, This is the consensus threshold. As the first coefficient, This is the first upper limit value.

[0089] The first coefficient is used to increase the consensus threshold, and it is greater than 1. The first coefficient can be set according to the specific type of dynamic background interference. For example, if the dynamic background interference is wind blowing leaves, a smaller first coefficient, such as 1.5, can be set. If the dynamic background interference is rain or snow falling, the first coefficient can be appropriately increased, such as 2.0. This allows for a higher consensus threshold to be used for strong interference (such as rain or snow) to effectively filter out noise, and a moderate consensus threshold to be used for weak interference (such as leaves) to prevent real moving targets from being misjudged as background, thus achieving a dynamic balance.

[0090] Formula 2; in, The updated background update rate. To perform the maximum value operation, For background update rate, As the second coefficient, This is the second upper limit value.

[0091] The second coefficient is used to reduce the background update rate. The second coefficient is less than 1, for example, 0.5, to prevent noise from being updated into the background.

[0092] Based on this, when the foreground authenticity judgment result of the current keyframe image is dynamic background interference, the detection device can make high-frequency noise difficult to meet the background judgment conditions by adjusting the consensus threshold and background update rate, and prevent noise from polluting the historical bag of words, thereby reducing the false detection rate and improving the quality of foreground segmentation.

[0093] When the foreground authenticity determination result of the current keyframe image is a ghost region, the detection device can reset the color sample library corresponding to each pixel position in the ghost region to obtain an updated historical word bag.

[0094] Based on this, the detection device can eliminate ghosting when ghosting regions appear in the current keyframe image, thus avoiding affecting the foreground segmentation of subsequent non-keyframe images and improving the quality of foreground segmentation.

[0095] Based on the above exemplary description, the detection device can be as follows: Figure 2 The method shown determines the current keyframe image from the video stream and obtains the binarized foreground mask of the current keyframe image.

[0096] Please see Figure 2 , Figure 2 This is a flowchart illustrating a method for determining a current keyframe image and obtaining a binarized foreground mask of the current keyframe image, provided as an embodiment of this application. Figure 2 As shown, the method includes: S201, The detection device acquires images from the video stream in real time.

[0097] S202. When the detection device acquires an image, it calculates the binarized foreground mask and confidence map of the image.

[0098] The calculation method for the binarized foreground mask is similar to that in S104, and will not be repeated here.

[0099] The confidence map is a grayscale image of the same size as the image in the video stream. The confidence map is used to indicate the confidence level at which each pixel in the image is identified as a foreground.

[0100] The confidence map can be calculated by the detection device in the following way.

[0101] Specifically, for a single pixel in an image, the detection device calculates the distance (e.g., Euclidean distance) between the pixel's color value and each color value in the color sample library corresponding to that pixel in the bag-of-words system. It then determines the minimum distance and maps this minimum distance to a confidence score using linear normalization. For example, the ratio of the minimum distance to a preset maximum possible distance can be used to determine the confidence score for that pixel. Confidence scores are calculated for all pixels in the image, ultimately yielding a complete confidence score map.

[0102] S203. When the detection device determines that the image meets the preset conditions based on the image's binarized foreground mask and confidence map, it determines the image as the current keyframe image and the image's binarized foreground mask as the current keyframe image's binarized foreground mask.

[0103] In some examples, the preset conditions include: The binarized foreground mask of the image undergoes a sudden change in area relative to the binarized foreground mask of the previous image; or, The average confidence score of the image, as well as the average confidence scores of the N images preceding and consecutive to the image, are all less than a preset confidence threshold; where N is a positive integer, and the average confidence score is determined based on the image's confidence map; or, A new connected component exists in the binarized foreground mask of the image.

[0104] If the area of ​​the binarized foreground mask of an image changes abruptly compared to the binarized foreground mask of the previous image, it may correspond to a significant change in the scene, such as the sudden appearance / disappearance of a large number of objects, a drastic change in lighting, or camera obstruction. In such cases, a large visual language model is needed to intervene and adjust its parameters. Therefore, the detection device can determine the current keyframe image by judging whether an area change has occurred.

[0105] Specifically, the detection device can determine whether an area abrupt change has occurred by the rate of change of the number of foreground pixels in the binarized foreground mask of the image relative to the number of foreground pixels in the previous image. For example, if the rate of change is greater than a preset rate of change threshold, then an area abrupt change is determined to have occurred.

[0106] If the average confidence score of the image and the average confidence scores of the N preceding and consecutive images are all less than a preset confidence threshold, complex interference may have occurred, such as dense fog or heavy rain. In such cases, it becomes impossible to determine whether this interference is noise or a real moving target. Therefore, a large-scale visual language model is needed to intervene and adjust its parameters. Thus, the detection device can determine the current keyframe image by judging the average confidence score.

[0107] In some cases, if a new connected component exists in the binarized foreground mask of an image, it may correspond to a new object. However, it may be impossible to determine whether this object is interference or a real moving target. In such situations, a large visual language model is needed to intervene and adjust its parameters. Therefore, the detection device can determine the current keyframe image by identifying the new connected component.

[0108] Specifically, the detection device can perform connected component labeling when acquiring each image. It compares the center positions or bounding boxes of all connected components in the current image with all connected components that appeared in the past few frames (e.g., the most recent 10 frames) to determine if any new connected components have appeared. If a connected component is found whose location is more than a preset distance threshold from all historical connected components, then a new connected component is determined to have appeared.

[0109] When the detection device determines that an image does not meet the preset conditions based on the binarized foreground mask and confidence map of the image, it continues to acquire the next image and continues to calculate the binarized foreground mask and confidence map of the next image, and determines whether the next image meets the preset conditions based on the binarized foreground mask and confidence map of the next image.

[0110] Based on this, the detection device can accurately select the images that most need semantic intervention by setting preset conditions, ensuring that the parameters are updated at the most appropriate time, thereby reducing the calling frequency of the large visual language model while ensuring real-time performance and saving computing resources.

[0111] Based on the above exemplary description, after obtaining the binarized foreground mask of the current keyframe image, the detection device can also correct the binarized foreground mask of the current keyframe image, considering that the current keyframe image may be affected by background interference, so as to obtain a more accurate detection result.

[0112] Specifically, the detection device inputs the current keyframe image and its binarized foreground mask into the large visual language model to obtain semantic analysis results. Based on the semantic analysis results, the binarized foreground mask of the current keyframe image is corrected to obtain the corrected binarized foreground mask of the current keyframe image.

[0113] It should be noted that the detection device can obtain the semantic analysis results at the same time as the foreground analysis results in S102, or it can obtain the semantic analysis results using the large visual language model alone. This application does not impose any restrictions on this.

[0114] The semantic analysis results are used to indicate whether the processing type of each connected component of the binarized foreground mask of the current keyframe image is to filter or retain.

[0115] Connected components with the processing type "preserved" are those corresponding to the real moving target, while connected components with the processing type "filtered" are those corresponding to dynamic background interference or ghosting regions.

[0116] The following section uses specific examples to illustrate the process of obtaining semantic analysis results.

[0117] For example, the detection device takes the current keyframe image and its binarized foreground mask as multi-channel inputs and feeds them into the large visual language model. It then inputs the prompt, "Please analyze the image and its corresponding binarized foreground mask. For each connected component in the binarized foreground mask, determine whether it belongs to a real moving target (such as a person, vehicle, or animal) that needs attention." Finally, it provides a "keep" or "remove" instruction for each connected component.

[0118] The large visual language model receives multi-channel input and prompt words, performs semantic analysis on the binarized foreground mask, and outputs the semantic analysis results.

[0119] Semantic analysis results: {"Region_A": {"Label": "Tree_Shadow", "Action": "Filter"}, "Region_B": {"Label": "Person", "Action": "Keep"}} After the semantic analysis results are obtained, the detection device processes the corresponding connected components according to the processing type indicated in the semantic analysis results to obtain a corrected binarized foreground mask, thereby eliminating dynamic background interference and ghosting regions in the keyframe image and preserving the real moving target.

[0120] The following describes a specific method for correcting the binarized foreground mask of the current keyframe image.

[0121] The detection device sets all pixels in a connected component of a binarized foreground mask of the current keyframe image as background (e.g., setting the pixel value to 0) when the processing type of the connected component is screening, and sets all pixels in the connected component as foreground (e.g., setting the pixel value to 1) when the processing type of the connected component is retention.

[0122] The detection device performs correction processing on all connected components of the binarized foreground mask of the current keyframe image to obtain the corrected binarized foreground mask of the current keyframe image.

[0123] Based on this, for keyframe images, the detection device can eliminate dynamic background interference and ghosting areas in the keyframe images, retain the real moving targets, and thus obtain more accurate detection results.

[0124] Based on the above exemplary description, after acquiring the binarized foreground mask of each image, the detection device can also verify the consistency of motion vectors of pixels in the binarized foreground mask through sparse optical flow, thereby filtering out static noise. The images include keyframe images and non-keyframe images.

[0125] Optical flow refers to the instantaneous motion vector of pixels in an image over the time domain. Real moving targets typically possess spatially continuous optical flow fields with relatively consistent directions, while dynamic background interference exhibits chaotic and discontinuous motion patterns. Sparse optical flow refers to the motion vector of several feature points (such as corner points and edge points) in an image. By verifying the consistency of sparse optical flow vectors within the foreground region of the binarized foreground mask, we can distinguish between real moving targets and noise, thereby achieving noise filtering.

[0126] Specifically, for a binary foreground mask of an image, the detection device extracts sparse feature points within the pixel range of the foreground region using a Shi-Tomasi corner detector or random sampling. Using grayscale images from the previous and current images, it calculates the optical flow vector for each sparse feature point using, for example, the Lucas-Kanade optical flow method. For each connected component, it evaluates the consistency of the optical flow vectors within the component, including direction and magnitude consistency. The detection device processes the binary foreground mask based on the consistency judgment results. For example, the detection device can classify low-speed or stationary foreground pixels as static noise and remove them from the binary foreground mask, placing them as background. For connected components with chaotic motion directions, it classifies them as dynamic background interference and removes them from the binary foreground mask, placing them as background. For connected components with consistent motion directions, it retains them as real moving targets and does not process them.

[0127] Based on this, the detection device can improve the reliability and accuracy of the detection results by purifying the binary foreground mask of the image through sparse optical flow.

[0128] For example, this application also provides a device for detecting moving targets.

[0129] Figure 3 This is a schematic diagram of a moving target detection device provided in one embodiment of this application. Figure 3 As shown, the device includes: a determination module 101, an analysis module 102, an update module 103, and a foreground segmentation module 104.

[0130] The determination module 101 is used to determine the current keyframe image from the video stream and obtain the binarized foreground mask of the current keyframe image; the video stream includes multiple consecutive images; Analysis module 102 is used to input the current keyframe image and the binarized foreground mask of the current keyframe image into the large visual language model to obtain the foreground analysis results; The update module 103 is used to update the foreground segmentation parameters based on the foreground analysis results, so as to obtain the updated foreground segmentation parameters; The foreground segmentation module 104 is used to perform foreground segmentation on each non-keyframe image in the video stream that is located after the current keyframe image and before the next keyframe image, according to the updated foreground segmentation parameters, to obtain a binarized foreground mask corresponding to each non-keyframe image; wherein, the binarized foreground mask of the non-keyframe image is used to indicate the position of the moving target in the corresponding non-keyframe image.

[0131] It should be noted that the moving target detection device in this application embodiment can be used to execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.

[0132] In some examples, the device also includes a loop module; The loop module is used to take the next keyframe image as the new current keyframe image when the next keyframe image is determined from the video stream. The determination module 101, analysis module 102, update module 103 and foreground segmentation module 104 are also used to continue to execute the steps of obtaining the binarized foreground mask of the current keyframe image, inputting it into the large visual language model, updating the foreground segmentation parameters and performing foreground segmentation.

[0133] In some examples, the foreground analysis results include the foreground authenticity determination results of the current keyframe image. The foreground segmentation parameters include consensus threshold, background update rate, and historical bag of words. The consensus threshold and historical bag of words are used to determine whether a pixel is foreground during foreground segmentation. The background update rate is used to update the historical bag of words. The historical bag of words includes a color sample library corresponding to each pixel position in the video stream. Each color sample library records the color value of the pixel position in the historical image.

[0134] In some examples, the update module 103 is specifically used to update the consensus threshold according to Formula 1 when the foreground authenticity judgment result of the current keyframe image is dynamic background interference, and to update the background update rate according to Formula 2. Formula 1 is as follows: ; The updated consensus threshold. To perform the minimum value operation, This is the consensus threshold. As the first coefficient, This is the first upper limit value; Formula 2 is as follows: ; The updated background update rate. To perform the maximum value operation, For background update rate, As the second coefficient, This is the second upper limit value; When the foreground authenticity determination result of the current keyframe image is a ghost region, the color sample library corresponding to each pixel position in the ghost region is reset to obtain the updated historical bag of words.

[0135] In some examples, module 101 is specifically used to acquire images from the video stream in real time; For each image acquired, a binarized foreground mask and a confidence map are calculated; the confidence map is used to indicate the confidence level at which each pixel in the image is identified as the foreground. When an image meets the preset conditions based on its binarized foreground mask and confidence map, the image is determined as the current keyframe image, and its binarized foreground mask is determined as the binarized foreground mask of the current keyframe image.

[0136] In some examples, the preset conditions include: The binarized foreground mask of the image undergoes a sudden change in area relative to the binarized foreground mask of the previous image; or, The average confidence score of the image, as well as the average confidence scores of the N images preceding and consecutive to the image, are all less than a preset confidence threshold; where N is a positive integer, and the average confidence score is determined based on the image's confidence map; or, A new connected component exists in the binarized foreground mask of the image.

[0137] In some examples, the device also includes a correction module; The correction module is used to input the current keyframe image and its binarized foreground mask into the large visual language model after obtaining the binarized foreground mask of the current keyframe image, and obtain the semantic analysis results. The semantic analysis results are used to indicate whether the processing type of each connected component of the binarized foreground mask of the current keyframe image is to filter or retain. Based on the semantic analysis results, the binarized foreground mask of the current keyframe image is corrected to obtain the corrected binarized foreground mask of the current keyframe image.

[0138] In some examples, the correction module, specifically for a connected component of the binarized foreground mask of the current keyframe image, sets all pixels in the connected component as background when the processing type of the connected component is sieving, and sets all pixels in the connected component as foreground when the processing type of the connected component is retaining. Based on the correction processing of all connected components of the binarized foreground mask of the current keyframe image, the corrected binarized foreground mask of the current keyframe image is obtained.

[0139] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 1 .like Figure 4 As shown, the electronic device may include a first processor 201, which, when executing a computer-executable program or instruction stored in a memory, implements the embodiments of this application. Figures 1 to 2 The method for detecting moving targets is shown.

[0140] The electronic device can be used to perform the various steps and / or processes corresponding to the electronic devices in the above method embodiments.

[0141] Figure 5A schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 2 .like Figure 5 As shown, the electronic device may include a second processor 301 and a memory 302. The memory 302 stores a computer program. When the second processor 301 executes the computer program, it implements the embodiments of this application. Figures 1 to 2 The method for detecting moving targets is shown.

[0142] The electronic device can be used to perform the various steps and / or processes corresponding to the electronic devices in the above method embodiments.

[0143] The electronic device of this application can be used to execute the technical solutions of the method embodiments described above. Its implementation principle and technical effects are similar. The operations implemented by each module can be further referred to the relevant descriptions of the method embodiments, which will not be repeated here. The modules here can also be replaced by components or circuits.

[0144] This application can divide electronic devices into functional modules based on the above method examples. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in the embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0145] Another embodiment of this application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the embodiments of this application. Figures 1 to 2 The method for detecting moving targets is shown.

[0146] This application also provides a program product including executable instructions stored in a computer-readable storage medium. At least one processor of an electronic device can read the executable instructions from the computer-readable storage medium, and the at least one processor executes the executable instructions to cause the electronic device to implement embodiments of this application. Figures 1 to 2 The method for detecting moving targets is shown.

[0147] This application also provides a chip that is connected to a memory, or a chip that integrates a memory. When a software program stored in the memory is executed, it implements the embodiments of this application. Figures 1 to 2 The method for detecting moving targets is shown.

[0148] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0149] Those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the claims, any of the claimed embodiments can be used in any combination.

[0150] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for detecting a moving target, characterized in that, The method includes: The current keyframe image is determined from the video stream, and a binarized foreground mask of the current keyframe image is obtained; the video stream includes multiple consecutive images; The current keyframe image and its binarized foreground mask are input into the large visual language model to obtain the foreground analysis results. Based on the foreground analysis results, the foreground segmentation parameters are updated to obtain the updated foreground segmentation parameters. Based on the updated foreground segmentation parameters, foreground segmentation is performed on each non-keyframe image in the video stream that is located after the current keyframe image and before the next keyframe image, to obtain a binarized foreground mask corresponding to each non-keyframe image; wherein, the binarized foreground mask of the non-keyframe image is used to indicate the position of the moving target in the corresponding non-keyframe image; The foreground analysis result includes the foreground authenticity judgment result of the current keyframe image. The foreground segmentation parameters include consensus threshold, background update rate and historical bag of words. The consensus threshold and the historical bag of words are used to determine whether a pixel is foreground during foreground segmentation. The background update rate is used to update the historical bag of words. The historical bag of words includes a color sample library corresponding to each pixel position in the video stream. Each color sample library records the color value of the pixel position in the historical image. The step of updating the foreground segmentation parameters based on the foreground analysis results to obtain updated foreground segmentation parameters includes: When the foreground authenticity determination result of the current keyframe image is dynamic background interference, the consensus threshold is updated according to Formula 1 to obtain the updated consensus threshold, and the background update rate is updated according to Formula 2 to obtain the updated background update rate. Formula 1 is as follows: ; The updated consensus threshold. To perform the minimum value operation, This is the consensus threshold. As the first coefficient, The first upper limit value is set according to the specific type of the dynamic background interference. Formula 2 is as follows: ; The updated background update rate. To perform the maximum value operation, For background update rate, As the second coefficient, This is the second upper limit value; When the foreground authenticity determination result of the current keyframe image is a ghost region, the color sample library corresponding to each pixel position in the ghost region is reset to obtain an updated bag of historical words.

2. The method according to claim 1, characterized in that, The method further includes: When the next keyframe image is determined from the video stream, the next keyframe image is used as the new current keyframe image, and the steps of obtaining the binarized foreground mask of the current keyframe image, inputting it into the large visual language model, updating the foreground segmentation parameters, and performing foreground segmentation are continued.

3. The method according to claim 1 or 2, characterized in that, The step of determining the current keyframe image from the video stream and obtaining the binarized foreground mask of the current keyframe image includes: Images are acquired from the video stream in real time; Upon acquiring each image, a binarized foreground mask and a confidence map for that image are calculated; the confidence map is used to indicate the confidence level at which each pixel in the image is identified as a foreground. When it is determined that the image meets the preset conditions based on the binarized foreground mask and the confidence map of the image, the image is determined as the current keyframe image, and the binarized foreground mask of the image is determined as the binarized foreground mask of the current keyframe image.

4. The method according to claim 3, characterized in that, The preset conditions include: The binarized foreground mask of the image undergoes a sudden change in area relative to the binarized foreground mask of the previous image; or, The average confidence score of the image and the average confidence scores of the N consecutive images preceding and following the image are both less than a preset confidence threshold; where N is a positive integer, and the average confidence score is determined based on the image confidence map; or, A new connected component exists in the binarized foreground mask of the image.

5. The method according to claim 1 or 2, characterized in that, After obtaining the binarized foreground mask of the current keyframe image, the method further includes: The current keyframe image and its binarized foreground mask are input into the large visual language model to obtain semantic analysis results; the semantic analysis results are used to indicate whether the processing type of each connected component of the binarized foreground mask of the current keyframe image is filtering or retention. Based on the semantic analysis results, the binarized foreground mask of the current keyframe image is corrected to obtain the corrected binarized foreground mask of the current keyframe image.

6. The method according to claim 5, characterized in that, The step of correcting the binarized foreground mask of the current keyframe image based on the semantic analysis results to obtain the corrected binarized foreground mask of the current keyframe image includes: For a connected component of the binarized foreground mask of the current keyframe image, when the processing type of the connected component is filtering, all pixels in the connected component are set as background; when the processing type of the connected component is retention, all pixels in the connected component are set as foreground. Based on the correction processing of all connected components of the binarized foreground mask of the current keyframe image, the corrected binarized foreground mask of the current keyframe image is obtained.

7. An electronic device, characterized in that, include: First processor; The first processor is configured to execute a computer-executable program or instructions in the memory, causing the electronic device to perform the moving target detection method according to any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer-executable program or instructions, the computer-executable program or instructions being configured to perform the moving target detection method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Image processing method and device containing motion object, and electronic equipment

    CN107808388A

  • Foreground detection method based on adaptive background update and selective background update

    CN108010050A