Target detection method, apparatus, device, medium, and program product

By using a lightweight target detection model and multi-level screening technology on low-computing-power devices, the problem of insufficient target detection accuracy on low-computing-power devices is solved, and efficient and accurate target detection is achieved on devices such as smart cameras.

CN122200461APending Publication Date: 2026-06-12CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD
Filing Date
2026-01-16
Publication Date
2026-06-12

Smart Images

  • Figure CN122200461A_ABST
    Figure CN122200461A_ABST
Patent Text Reader

Abstract

The application provides a target detection method, device, equipment, medium and program product. The target detection method comprises: performing target detection on an original video to obtain a first image containing a target object and a plurality of image features of different scales of each first image; performing screening on the first image according to the plurality of image features of different scales to obtain a second image; and determining a detection result of the target object according to the second image. The application can first perform target detection on the original video through a lightweight target detection model, and then perform screening on the obtained first image according to a plurality of image features of different scales. Since the lightweight target detection model and the screening process do not require a large number of computing resources, and the screening process can significantly improve the accuracy of the detection result, the application can realize accurate detection of the target object on a low-power device, effectively solving the problems in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image detection technology, and in particular to a target detection method, apparatus, device, medium, and program product. Background Technology

[0002] The widespread adoption of smart cameras has brought great convenience to users' daily lives. For example, in smart home scenarios, users can use smart cameras to obtain and analyze information about their pets, including their location, emotions, and movements. However, in practical applications, some target objects do not cooperate with the camera's capture, resulting in inconsistent image quality from consecutive captures, making it difficult to accurately detect the target object.

[0003] Currently, many camera manufacturers use high-performance cameras, employing high-frame-rate CMOS sensors and high-performance chips in hardware, and complex neural network models with multiple layers in software, to achieve accurate object detection. However, this approach is limited by model complexity and cannot be used on low-performance chips. Therefore, how to achieve accurate object detection on low-performance devices has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides a target detection method, apparatus, device, medium, and program product to solve the technical problem that the prior art cannot achieve accurate detection of target objects on low computing power devices.

[0005] In a first aspect, embodiments of this application provide a target detection method, including: Target detection is performed on the original video to obtain a first image containing the target object and multiple image features of different scales for each of the first images; Based on the image features at multiple different scales, the first image is filtered to obtain the second image; The detection result of the target object is determined based on the second image.

[0006] Secondly, embodiments of this application provide a target detection device, comprising: The detection module is used to perform target detection on the original video to obtain a first image containing the target object and multiple image features of different scales for each of the first images; The filtering module is used to filter the first image based on the multiple image features at different scales to obtain the second image; The determination module is used to determine the detection result of the target object based on the second image.

[0007] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the target detection method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the target detection method described in the first aspect.

[0009] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the target detection method described in the first aspect.

[0010] Implementing the target detection method provided in this application allows for target detection on original videos using a lightweight target detection model, resulting in a first image containing the target object and multiple image features of different scales for each first image. Next, the first image is filtered based on these multiple image features to obtain a second image. Finally, the detection result of the target object is determined based on the second image. In this application, the lightweight target detection model does not require significant computing resources and can be directly deployed on low-computing-power devices (such as smart cameras). Furthermore, the filtering of the first image based on multiple image features of different scales requires minimal computing resources and effectively improves the quality of the second image (the second image contains a higher probability of containing a real target object), thereby enhancing the accuracy of the detection results. Therefore, the target detection method of this application enables accurate target object detection on low-computing-power devices, effectively solving the problems in related technologies. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of a target detection method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating the implementation process of a target detection method according to an embodiment of this application; Figure 3 This is a schematic diagram illustrating the multi-level screening process of the target object in an embodiment of this application; Figure 4 This is a structural block diagram of a target detection device shown in an embodiment of this application; Figure 5 This is a schematic diagram of the physical structure of an electronic device as shown in an embodiment of this application. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] The target detection method provided in this application is implemented by a target detection device or a target detection apparatus. The target detection device can be any type of electronic device, such as a smart camera, computer, or mobile phone, and this application does not impose any specific restrictions on it.

[0015] Figure 1 This is a flowchart of a target detection method provided in an embodiment of this application. (Refer to...) Figure 1 The target detection method of this application may include: Step 101: Perform target detection on the original video to obtain a first image containing the target object and multiple image features of different scales for each first image.

[0016] The original video is captured by a camera. The target object can be any living object, such as a person or pet, or any non-living object, such as a vehicle or drone. The type of target object can be set according to actual needs.

[0017] Step 101 involves first performing object detection on the original video to obtain initial detection results. These initial results include first images of the target object, confidence scores of the target object in each first image, and multiple image features at different scales for each first image. A first image refers to a local image containing the target object cropped from a video frame of the original video; each first image contains one target object. Next, the initial detection results are filtered based on the confidence scores. For example, first images and multiple image features at different scales related to target objects with confidence scores below a certain threshold are deleted, while the remaining first images and multiple image features related to the target objects are retained and input into subsequent steps 102-103.

[0018] Step 102: Based on image features at multiple different scales, filter the first image to obtain the second image.

[0019] Image features at various scales can be categorized into low-level, mid-level, and high-level features. Low-level features contain the most basic image characteristics, such as edges, corners, colors, and textures. Mid-level features contain complex local shapes, representing a transition from low-level details to high-level semantics. High-level features contain abstract semantic information, such as "this is a four-legged animal."

[0020] In this embodiment, the first image is filtered based on image features at multiple different scales to obtain the second image. Because the filtering process considers image features at multiple different scales of the first image, it can effectively eliminate first images containing false target objects, making the probability that the target object in the final second image is a real target object higher.

[0021] Step 103: Determine the detection result of the target object based on the second image.

[0022] In this embodiment, since the target object contained in the second image has a high probability of being a real target object, the detection result of the target object can be obtained directly from the second image. For example, the bounding box coordinates and confidence score of the target object in the second image can be used as the final detection result.

[0023] Implementing the target detection method provided in this application allows for target detection on original videos using a lightweight target detection model, resulting in a first image containing the target object and multiple image features of different scales for each first image. Next, the first image is filtered based on these multiple image features to obtain a second image. Finally, the detection result of the target object is determined based on the second image. In this application, the lightweight target detection model does not require significant computing resources and can be directly deployed on low-computing-power devices (such as smart cameras). Furthermore, the filtering of the first image based on multiple image features of different scales requires minimal computing resources and effectively improves the quality of the second image (the second image contains a higher probability of containing a real target object), thereby enhancing the accuracy of the detection results. Therefore, the target detection method of this application enables accurate target object detection on low-computing-power devices, effectively solving the problems in related technologies.

[0024] In conjunction with the above embodiments, in one implementation, step 102 may include: Step 1021: Obtain the pose features of the target object in the first image; Step 1022: Based on image features and pose features at multiple different scales, the first image is filtered to obtain the second image.

[0025] In this embodiment, the initial detection results output by the target detection model may also include the pose features of the target objects in each of the first images. Therefore, when performing step 1022, the first images can be filtered based on image features and pose features at multiple different scales to obtain the second images.

[0026] In this embodiment, since the pose characteristics of the target object are considered during the screening process, all first images containing mirrored objects (i.e., the target object in a mirror or its reflection) can be removed, so that the remaining first images are all images containing the real target object. The second image is this remaining portion of the first images containing the real target object.

[0027] For example, the target object is a cat. The second image obtained after steps 101-102 is an image containing a real cat, so the final detection result can be determined based on this part of the image containing a real cat.

[0028] In this embodiment, filtering the first image based on multiple image features and pose features of different scales can effectively improve the quality of the final second image containing the target object (the second image contains real target objects), thereby improving the accuracy of the detection results.

[0029] This application provides two object detection methods. The first method involves directly filtering a first image based on image features at multiple different scales to obtain a second image, and then determining the detection result of the target object based on the second image. The second method involves filtering a first image based on image features at multiple different scales and pose features to obtain a second image, and then determining the detection result of the target object based on the second image. Both methods can achieve accurate detection of target objects on low-computing-power devices. In practical implementation, either method can be used.

[0030] In this application, if the first object detection method is used, then filtering the first image based on image features at multiple different scales can include: Based on the image features of each first image at multiple different scales, the feature fingerprint of each first image is determined; Determine the deviation values ​​between each feature fingerprint and the standard feature fingerprint of the target object; The first image is filtered based on the deviation value.

[0031] In this embodiment, if a certain first image has Image features at each scale, then the feature fingerprint of the first image can be used [ [Represents] image features at a certain scale. ( It can be determined based on its corresponding multiple feature maps.

[0032] After obtaining the feature fingerprints of each first image, they are compared with the standard feature fingerprints of the target object to obtain the deviation values. Then, the relationship between each deviation value and a deviation threshold is determined: if the deviation value of a first image is less than the deviation threshold, it indicates a higher probability that the target object in that first image is the real target object; if the deviation value of a first image is not less than the deviation threshold, it indicates a higher probability that the target object in that first image is a false target object, and therefore that first image needs to be discarded. Finally, all first images with corresponding deviation values ​​not greater than the deviation threshold can be directly used as second images.

[0033] In this embodiment, the standard feature fingerprint is obtained by analyzing the feature fingerprints of a large number of target objects. The standard feature fingerprint is a universal feature fingerprint of the target objects, containing the common features of all target objects. For example, when the target object is a cat, the standard feature fingerprint can be obtained by analyzing the feature fingerprints of a large number of cats, which contains the common features of all cats.

[0034] In this embodiment, the standard feature fingerprint of the target object is obtained in advance, and then the feature fingerprint of each first image is compared with the standard feature fingerprint. This can eliminate first images that obviously contain false target objects, increase the probability that the target object contained in the second image is a real target object, and thus improve the accuracy of the detection results.

[0035] In conjunction with the above embodiments, in one implementation, determining the feature fingerprint of each first image based on multiple image features of different scales includes: For each first image, determine the average value of multiple feature maps included in the image features at each scale; The feature fingerprint of each first image is determined based on the average value of the image features at multiple different scales of each first image.

[0036] In this embodiment, an image feature at a certain scale is essentially a multi-dimensional digital matrix (with shape...). ),Include The image features at this scale are represented by the average of multiple feature maps. express.

[0037] In practical implementation, for any first image, the corresponding image features at a certain scale are determined. When this happens, the formula can be used: in, Indicates the first Image features at each scale include feature maps.

[0038] For example, if the first image X contains image features at three scales, the average values ​​of the corresponding feature maps are respectively as well as , then as well as By simply stitching the images together, we can obtain the feature fingerprint corresponding to the first image X.

[0039] In this embodiment, the feature fingerprint of each first image is determined by the average value of all feature maps corresponding to the image features at multiple different scales of each first image. This ensures that the feature fingerprint can accurately express the key information of the first image, thereby guaranteeing the accuracy of the final detection result.

[0040] In conjunction with the above embodiments, in one implementation, this application further extends the first target detection method to obtain a third target detection method. The difference between the third target detection method and the first target detection method is that, after filtering the first image based on the deviation value to obtain all corresponding first images with deviation values ​​not greater than a deviation threshold, these are no longer directly used as second images. Instead, they are further processed. Specifically, all corresponding first images with deviation values ​​not greater than the deviation threshold are determined as third images, and the third images are filtered based on the degree of feature correlation between them to obtain the second image.

[0041] In this embodiment, if the target object in a certain third image X is a real target object, then among all the third images, there must exist other third images Y that are similar to this third image, and the target objects in third images X and Y must be able to form a motion trajectory. Conversely, if the third image X is an isolated image, it can be determined that it must contain a false target object.

[0042] In view of the above considerations, this application can adopt a third target detection method, which eliminates isolated third images based on the degree of feature correlation (similarity) between each third image, and uses the remaining third images as second images. This can further increase the probability that the target objects contained in the second images are real target objects, and ensure the accuracy of the detection results.

[0043] Next, the second object detection method of this application will be described in detail. Based on the combination of the first and third object detection methods, the second object detection method further eliminates false object objects according to the pose characteristics of the target object, thereby ensuring that the target objects contained in the final second image are real object objects. Specifically, in step 1022, the first image is filtered according to multiple image features and pose features at different scales to obtain the second image, which may include: Step A1: Based on image features at multiple different scales, the first image is filtered to obtain the third image.

[0044] In this step, the principle of filtering the first image based on image features at multiple different scales is exactly the same as that in the first object detection method. The only difference is that the result obtained from the filtering is named the third image in this step.

[0045] For example, if the target object is a cat, and images 1-3 all contain cats, then analyzing the image features of image 1 at multiple scales can determine whether the cat in image 1 is a real cat. Similarly, analyzing the image features of image 2 at multiple scales can determine whether the cat in image 2 is a real cat, and analyzing the image features of image 3 at multiple scales can determine whether the cat in image 3 is a real cat. Assuming the cats in images 1-2 are real cats, then images 1-2 would be considered the third image.

[0046] Step A2: Based on the degree of feature correlation between each third image, filter the third images to obtain the fourth image.

[0047] In this step, the principle of filtering the third images based on the degree of feature correlation between each third image is exactly the same as that in the third target detection method. The only difference is that the result obtained by filtering is named the fourth image in this step.

[0048] Step A3: Based on the pose features, filter the fourth image to obtain the second image.

[0049] In this embodiment, if a portion of the first images contains a mirrored object, then this portion of the first images is likely to pass the filtering in steps A1 and A2 and ultimately become the fourth image. To remove mirrored objects, in step A3, the pose features of the target objects in each fourth image can be analyzed to determine which fourth images contain mirrored objects and delete them, leaving the remaining fourth images as the second images.

[0050] In this embodiment, the first image is first filtered based on image features at multiple different scales, removing images that clearly contain false target objects to obtain the third image. Then, based on the feature correlation between the third images, a second-level filtering is performed, removing isolated third images to obtain the fourth image. Finally, based on pose features, a third-level filtering is performed on the fourth image, removing images that contain mirrored objects to obtain the second image. This multi-level filtering method not only achieves accurate identification of real target objects but also eliminates the need for high-performance models and extensive computing resources, allowing the third target detection method to be deployed on any low-performance device, thus effectively solving the problems in related technologies.

[0051] In conjunction with the above embodiments, in one implementation, filtering the third images based on the degree of feature correlation between each third image to obtain a fourth image may include: Acquire at least one image group, wherein an image group comprises a first number of third images with consecutive timestamps; Obtain image chains in each image group. An image chain includes at least a second number of third images in the image group to which it belongs. The similarity of the feature fingerprints between adjacent third images in the image chain is greater than a first similarity threshold, and the second number is not greater than the first number. The third image included in each image chain is identified as the fourth image.

[0052] In this embodiment, after processing each third image in step A1, the third image is passed to step A2. Step A2 can set up a buffer pool to store the third images passed from step A1. Whenever a first number of third images (forming an image group) are stored, the similarity between each pair of these first number of third images is calculated. If the similarity between two third images is greater than a first similarity threshold, a connection is established between these two third images. Finally, based on the connection relationship, at least one image chain can be found, where the similarity of the feature fingerprints between adjacent third images in each image chain is greater than the first similarity threshold. Based on all the third images in an image chain, the motion trajectory of the target object can be roughly determined.

[0053] In this embodiment, if a third image is not connected to any other third image, that is, if there is an isolated third image, then the isolated third image is removed.

[0054] For example, if the first quantity is 6, during step A2, if 6 third images are detected in the buffer pool, in the order of their insertion (images 1-6), then connections are established among these 6 third images based on the pairwise similarity. Ultimately, if connections are established between image 1 and image 2, image 2 and image 4, image 4 and image 5, and image 5 and image 6, then image 1, image 2, and image 4-6 form an image chain. If the target object in this image chain is a cat, then the cat's movement trajectory can be determined based on this image chain. Since image 3 is not connected to any other image, image 3 needs to be removed.

[0055] In this embodiment, "timestamp continuity" means that the timestamps are arranged in chronological order. Step A1 processes each first image according to the chronological order of its timestamps, which ensures that the filtered third images are also stored in the buffer pool in chronological order. Therefore, when step A2 reads the first number of third images from the buffer pool, the first number of third images read are also arranged in chronological order.

[0056] In this embodiment, a sliding window is used to read each image group sequentially. After analyzing each image group, the third image with the earliest timestamp in that image group is deleted, which is also the first third image to enter the buffer pool. Then, the system waits for new third images to be added until the number of third images in the entire buffer pool reaches a certain number before proceeding to the next round of image group analysis.

[0057] In this embodiment, when determining the similarity between two third images, the L1 distance of the feature fingerprints between the two third images can be calculated first, and then the similarity can be determined based on the L1 distance. The smaller the L1 distance, the higher the similarity. If the similarity is greater than the first similarity threshold, it indicates that the two third images are extremely similar, and a connection can be established.

[0058] In this embodiment, by analyzing all the third images in the manner described above, at least one image group can be obtained, and then an image chain of all image groups can be obtained. Finally, the third image included in all image chains is the fourth image.

[0059] In this embodiment, the values ​​of the first quantity and the second quantity can be set according to actual needs, as long as the second quantity is not greater than the first quantity.

[0060] In this embodiment, by determining the image chain, all isolated images in the third image can be eliminated, thereby filtering out the third image containing false target objects. This makes the probability that the target object in the fourth image is a real target object higher, thus effectively improving the accuracy of the final detection result.

[0061] In one implementation, based on the above embodiments, the fourth image is filtered according to pose features to obtain the second image, including: Determine the average feature fingerprint of each image chain in the fourth image; Image chains are filtered based on average feature fingerprints and pose features; The fourth image, which is included in the remaining image chain, is identified as the second image.

[0062] In this embodiment, each image chain includes multiple fourth images. For a single image chain, the average value of the feature fingerprints of each of its included fourth images can be determined, and this average value is used as the average feature fingerprint of the image chain. Next, the image chains can be filtered based on the average feature fingerprint and pose features, eliminating image chains representing the motion trajectory of false target objects, ultimately obtaining image chains representing the motion trajectory of real target objects. The fourth images included in the image chains representing the motion trajectory of real target objects are the second images.

[0063] In this embodiment, the image chain is filtered based on the average feature fingerprint and pose features, so that all the second images obtained are images containing real target objects, which can effectively improve the accuracy of the detection results.

[0064] In one implementation, based on the above embodiments, filtering the image chain according to the average feature fingerprint and pose features includes: Obtain any two image chains whose average feature fingerprint similarity is greater than the second similarity threshold; Based on the pose features of each of the two image chains, determine the geometric relationship between the poses of the target objects contained in each of the two image chains; The image chain is filtered based on geometric relationships.

[0065] In this embodiment, for any two image chains whose average feature fingerprint similarity is greater than the second similarity threshold, the probability that one of them represents the motion trajectory of a false target object is extremely high. Therefore, for these two image chains, the following steps can be performed: based on the pose features of each of the two image chains, determine the geometric relationship between the poses of the target objects contained in each of the two image chains. The geometric relationship in this application is mainly a mirror relationship. If the determined geometric relationship is a mirror relationship, it means that one of the two image chains must represent the motion trajectory of a false target object (specifically, a mirror object), and therefore the image chain corresponding to the false target object can be further eliminated.

[0066] In this embodiment, when determining the similarity of the average feature fingerprints between two image chains, the L1 distance between the average feature fingerprints of the two image chains can be calculated first, and then the similarity can be determined based on the L1 distance.

[0067] In this embodiment, for any two image chains whose average feature fingerprint similarity is greater than the second similarity threshold, the geometric relationship of the poses between the target objects contained in them is determined, and the image chains are further filtered according to the geometric relationship. This ensures that all the second images obtained are second images containing real target objects, which can effectively improve the accuracy of the detection results.

[0068] In conjunction with the above embodiments, in one implementation, filtering the image chain based on geometric relationships includes: If the geometric relationship is a mirror relationship, obtain the confidence scores of the target objects contained in each of the two image chains; Remove image chains with low confidence scores.

[0069] In this embodiment, the confidence score has already been obtained in step 101. When two image chains are determined to be mirror images of each other, for each image chain, the average confidence score of all the second images contained within it can be used as the confidence score of the target object contained in that image chain. When the target object is a mirror image, its confidence score is lower than that of the real target object, therefore the image chain with the lower confidence score can be eliminated. The above steps are repeated until all image chains corresponding to mirror objects are eliminated, and the final image chain obtained is the image chain containing the real target object.

[0070] In this embodiment, after identifying two image chains that are mirror images of each other, the image chain corresponding to the mirror object is removed based on the confidence score, so that the remaining image chains are all image chains containing the real target object, which can effectively improve the accuracy of the detection results.

[0071] In conjunction with the above embodiments, in one implementation, filtering the image chain based on geometric relationships includes: If the geometric relationship is a mirror image, obtain the sharpness of the fourth image contained in each of the two image chains; Remove image chains containing low-resolution fourth images.

[0072] In this embodiment, when two image chains are determined to be mirror images of each other, the average sharpness of all the fourth images contained in each image chain can be used as the sharpness of the fourth images contained in that image chain. When the target object is a mirror image, its sharpness is lower than that of the real target object, so the image chain with low sharpness can be discarded. The above steps are repeated until the image chains corresponding to mirror objects in all image chains are discarded, and the final image chains are all image chains containing the real target object.

[0073] In this embodiment, after determining two image chains that are mirror images of each other, the image chain corresponding to the mirror object is removed based on the sharpness, so that the remaining image chains are all image chains containing the real target object, which can effectively improve the accuracy of the detection results.

[0074] In conjunction with the above embodiments, in one implementation, step 101 may include: The original video is subjected to object detection to obtain initial detection results. The initial detection results include the fifth image of the target object, multiple image features at different scales in each fifth image, and the confidence score of the target object in each fifth image. Based on the confidence score, the initial detection results are filtered to obtain a first image containing the target object and multiple image features of different scales for each first image. The first image is the fifth image containing the target object whose confidence score is greater than the score threshold.

[0075] In this embodiment, a lightweight object detection model is used to perform object detection on the original video, resulting in initial detection results including a fifth image containing the target object, multiple image features at different scales for each fifth image, and a confidence score for the target object in each fifth image. Next, based on the confidence score, fifth images containing target objects with confidence scores greater than a threshold are selected from the initial detection results as first images. Further selection yields the first images containing the target object and multiple image features at different scales for each first image.

[0076] In this embodiment, a lightweight object detection model is used to detect objects in the original video, and the initial detection results are filtered based on the confidence score to obtain a first image containing the target object and multiple image features of each first image at different scales. This enables subsequent filtering of the first image containing the real target object, ensuring the accuracy of the detection results.

[0077] In conjunction with the above embodiments, in one implementation method, target detection is performed on the original video, including: Retrieve video frames from the original video that contain the target object; Input the video frames into the target detection model to obtain the initial detection results; The object detection model is trained on the MobileNetV3 network in advance using sample images containing target objects and the labels of the target objects carried in the sample images as input.

[0078] In this embodiment, preliminary target detection can be performed on the original video (preliminary target detection can be performed by a smart camera) to obtain video frames containing the target object. Then, the video frames containing the target object are input into the target detection model, which further analyzes the video frames to obtain the initial detection results.

[0079] In this embodiment, MobileNetV3 is a lightweight, high-efficiency convolutional neural network architecture designed specifically for mobile and embedded devices (such as smartphones, smart cameras, IoT devices, etc.). It can achieve an optimal balance between speed (low latency), accuracy (high precision), and model size (small size) on devices with extremely limited computing resources, enabling complex object detection tasks to run in real time on terminal devices.

[0080] The object detection model trained on MobileNetV3 in this application includes a shared feature extraction backbone network and three parallel output branches, while also allowing feature extraction from the intermediate layers of the backbone network. Specifically, in one forward propagation, the object detection model can directly extract multiple image features of different scales (e.g., F1, F2, and F3 scales from nodes at different depths (e.g., layers 5, 10, and 13) of the feature extraction backbone network. These image features are mainly used for subsequent image filtering. Secondly, the high-level features extracted by the feature extraction backbone network are input to three independent head task branches: a classification branch, which outputs the confidence score of the target object through a fully connected layer and a Softmax activation function; a pose estimation branch, which outputs the pose features (e.g., pose angles) of the target object through a fully connected layer; and a localization and regression branch, which outputs the bounding box coordinates of the target object. This is used to accurately locate and crop the target object in the original video frame to obtain the first image.

[0081] In this embodiment, the object detection model trained using MobileNetV3 has advantages such as fewer parameters, lower computational complexity, smaller model size, and extremely low memory usage. It can not only improve running efficiency and achieve fast real-time inference on limited computing units, ensuring the smoothness of video stream processing, but also enable the entire object detection model to run directly on a low-cost, low-power embedded chip with a computing power of less than 1 TOPS. Therefore, it can effectively solve the problems in the prior art.

[0082] Figure 2 This is a schematic diagram illustrating the implementation process of a target detection method according to an embodiment of this application. The following will be combined with... Figure 2 This paper describes the target detection method of this application using a complete embodiment. In this embodiment, the executing entity is a target detection device, and the original video is captured by a smart camera. The target object is a dog, and the specific implementation process includes: Step S1 (Data Acquisition): Acquire the original video from the local machine or video storage server, perform video decoding to obtain a series of video frames, perform preliminary target detection on these video frames to obtain an image containing a dog.

[0083] In this step, the target detection device can pre-store the raw video captured by the smart camera locally, or directly request the raw video captured by the smart camera from the video storage server.

[0084] Among these methods, a smart camera can perform preliminary target detection on the original video to obtain an image containing a dog.

[0085] Step S2 (Feature Extraction): The video frames from Step S1 are input into the pre-trained target detection model to obtain initial detection results, specifically including the fifth image of the dog, multiple image features at different scales for each fifth image, the confidence score and pose features of the dog in each fifth image. The multiple image features at different scales include image features extracted from layer 5. Image features extracted from layer 10 and image features extracted from layer 13 . It includes lower-level features, such as edges and textures, which are related to the geometry of the image. Features containing intermediate levels, situated between low and high levels, help capture transitional information in an image. Including higher-level features, such as the dog's category and advanced semantic information, helps in understanding the overall content of the image.

[0086] Next, based on the confidence score, the fifth image containing a dog with a confidence score greater than the score threshold is selected from the initial detection results as the first image. Further filtering yields the first image containing a dog and multiple image features at different scales for each first image. , , (and the posture features of the dog in each of the first images)

[0087] Step S3 (Feature Comparison): Compare the features at three different scales from Step S2. , , The feature maps are extracted, and their shapes are as follows: , , Next, the average value of the feature at each scale is calculated using the following formula: in, Features The average value of the corresponding feature map, Features The average value of the corresponding feature map, Features The average value of the corresponding feature map.

[0088] For each first image, , , By simply piecing them together, we can obtain the characteristic fingerprint.

[0089] Next, the feature fingerprint of each first image is compared with the standard feature fingerprint of the dog. Specifically, the L1 distance between the feature fingerprint of each first image and the standard feature fingerprint of the dog is calculated as the deviation value between the two. Then, the first images are filtered according to the deviation value: if the deviation value of the first image is less than 0.1, it indicates that the deviation is small, and the first image is retained; if the deviation value of the first image is greater than or equal to 0.1 and less than 0.5, it indicates that the deviation is within an acceptable range, and the first image is retained; if the deviation value of the first image is greater than or equal to 0.5, it indicates that the deviation is large, and the first image is discarded. The deviation threshold is 0.5.

[0090] Among them, the standard feature fingerprint of a dog is derived by statistical analysis of a large number of known dog images and is used to represent the characteristics of a typical dog, such as the average facial features, body proportions and posture of a dog.

[0091] Finally, the first image obtained after step S3 is referred to as the third image. Step S3 is to remove target objects with high confidence but large feature bias, such as... Figure 3 As shown in 'a'. Figure 3 This is a schematic diagram illustrating the multi-level screening process of the target object in an embodiment of this application.

[0092] Step S4 (Similarity Evaluation): Read a first number of third images from the cache pool to form an image group. Calculate the pairwise similarity between each pair of the first number of third images in the image group. If the similarity between two third images is greater than a first similarity threshold, establish a connection between these two third images. Finally, based on the connection relationships, at least one image chain can be found, where the similarity of the feature fingerprints between adjacent third images in each image chain is greater than the first similarity threshold. If a third image does not have a connection with any other third image, it is removed. After analyzing each image group, delete the third image with the earliest timestamp in the image group, and then wait for new third images to be added until the number of third images in the entire cache pool reaches the first number before proceeding to the next round of image group analysis.

[0093] Step S4 is to remove target objects with high instantaneous confidence, such as... Figure 3 As shown in b in the figure.

[0094] Step S5 (Mirror Image Processing): First, determine the average feature fingerprint of each image chain. Any two image chains whose average feature fingerprint similarity is greater than a second similarity threshold are considered a group of suspected mirror image chains. Next, determine the geometric relationship between the poses of the target objects contained in each of these suspected mirror image chains. For example, if the dog's head angle is 30 degrees in one image chain and -30 degrees in another, these two image chains are symmetrical, meaning their geometric relationship is mirrored. Therefore, these two image chains can be identified as mirror image chains. Then, considering that the confidence score of a dog as a mirrored object is lower than that of a real dog, image chains with low confidence scores can be discarded. The final image chains obtained all contain real dogs.

[0095] Step S5 is to remove mirror targets with consistently high confidence levels, such as... Figure 3 As shown in c in the figure.

[0096] In this application, the object detection method can be executed by a smart camera, and the user can activate the object detection method on the smart camera via an app. Specifically, after binding the app to the smart camera, the user can set an on / off switch for the object detection method on the app side. If the user turns on this switch, after the smart camera captures an image containing the target object, it will process these images using the object detection method and implement multiple levels of filtering to make the final output detection result more accurate.

[0097] The target detection method of this application has at least the following technical effects: First, by adopting a multi-level filtering module in series, it can effectively filter high-confidence negative samples under various complex conditions, thereby improving the versatility of the false detection target filtering module.

[0098] Secondly, since only a lightweight feature extraction backbone network needs to be trained, it does not require excessive computing resources, making it easy to apply to consumer-grade or low-computing-power smart cameras. The object detection model in this application, after quantization to INT8, can achieve a running speed of nearly 20 FPS on low-computing-power chips (0.6 TOPS), effectively meeting the needs of low-computing-power devices. Furthermore, the method in this application can also be used normally at low frame rates.

[0099] Third, for target objects that are difficult to filter, such as mirror objects and transient objects, a clever judgment method has been designed based on their characteristics, which can effectively avoid false detection problems.

[0100] Figure 4This is a structural block diagram of a target detection device illustrated in an embodiment of this application. The target detection device provided in this application embodiment is described below, and the target detection device described below can be referred to in correspondence with the target detection method described above. For example... Figure 4 As shown, the target detection device 400 of this application may include: The detection module 401 is used to perform target detection on the original video to obtain a first image containing the target object and multiple image features of different scales for each of the first images; The filtering module 402 is used to filter the first image based on the multiple image features at different scales to obtain the second image; The determination module 403 is used to determine the detection result of the target object based on the second image.

[0101] According to the target detection device 400 provided in this application, the filtering module 402 is specifically used to: acquire the pose features of the target object in the first image; and filter the first image according to the multiple image features of different scales and the pose features to obtain a second image.

[0102] According to the target detection device 400 provided in this application, the filtering module 402 is specifically used to: filter the first image according to the multiple image features of different scales to obtain a third image; filter the third image according to the feature correlation degree between each of the third images to obtain a fourth image; and filter the fourth image according to the pose features to obtain a second image.

[0103] According to the target detection device 400 provided in this application, the screening module 402 is specifically used to: determine the feature fingerprint of each first image based on multiple image features of each first image at different scales; determine the deviation value between each feature fingerprint and the standard feature fingerprint of the target object; and screen the first images based on the deviation value.

[0104] According to the target detection device 400 provided in this application, the screening module 402 is specifically used for: determining the average value of multiple feature maps included in the image features at each scale for each first image; and determining the feature fingerprint of each first image based on the average value corresponding to the multiple image features at different scales of each first image.

[0105] According to the target detection device 400 provided in this application, the screening module 402 is specifically used for: acquiring at least one image group, wherein one image group includes a first number of third images with consecutive timestamps; acquiring image chains in each image group, wherein one image chain includes at least a second number of third images in its respective image group, wherein the similarity of the feature fingerprints between adjacent third images in the image chain is greater than a first similarity threshold, and the second number is not greater than the first number; and determining the third images included in each image chain as the fourth image.

[0106] According to the target detection device 400 provided in this application, the filtering module 402 is specifically used for: determining the average feature fingerprint of each image chain in the fourth image; filtering the image chain according to the average feature fingerprint and the pose feature; and determining the fourth image included in the remaining image chain as the second image.

[0107] According to the target detection device 400 provided in this application, the filtering module 402 is specifically used for: acquiring any two image chains whose average feature fingerprint similarity is greater than a second similarity threshold; determining the geometric relationship between the poses of the target objects contained in the two image chains according to their respective pose features; and filtering the image chains according to the geometric relationship.

[0108] According to the target detection device 400 provided in this application, the screening module 402 is specifically used for: if the geometric relationship is a mirror relationship, obtaining the confidence score of the target object contained in each of the two image chains; and removing the image chain with the low confidence score.

[0109] According to the target detection device 400 provided in this application, the screening module 402 is specifically used for: if the geometric relationship is a mirror relationship, obtaining the clarity of the fourth image contained in each of the two image chains; and removing the image chain containing the fourth image with low clarity.

[0110] According to the target detection device 400 provided in this application, the detection module 401 is specifically used for: performing target detection on the original video to obtain an initial detection result, the initial detection result including a fifth image containing a target object, multiple image features of different scales for each of the fifth images, and a confidence score of the target object in each of the fifth images; filtering the initial detection result according to the confidence score to obtain a first image containing a target object and multiple image features of different scales for each of the first images, wherein the first image is a fifth image containing a target object whose confidence score is greater than a score threshold.

[0111] According to the target detection device 400 provided in this application, the detection module 401 is specifically used for: acquiring video frames containing target objects in the original video; inputting the video frames into a target detection model to obtain the initial detection result; wherein, the target detection model is obtained by training the MobileNetV3 network in advance with sample images containing target objects and the labels of target objects carried by the sample images as input.

[0112] According to the target detection device 400 provided in this application, the determining module 403 is specifically used to: determine the detection result of the target object based on the bounding box coordinates and confidence score of the target object in the second image.

[0113] Figure 5 This is a schematic diagram of the physical structure of an electronic device according to an embodiment of this application. Figure 5 As shown, the electronic device may include a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call a computer program in the memory 530 to execute the steps of a target detection method, such as: performing target detection on the original video to obtain a first image containing the target object and multiple image features of different scales for each of the first images; filtering the first images according to the multiple image features of different scales to obtain a second image; and determining the detection result of the target object based on the second image.

[0114] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0115] On the other hand, embodiments of this application also provide a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of a target detection method provided in the above embodiments, such as: performing target detection on an original video to obtain a first image containing a target object and multiple image features of different scales for each of the first images; filtering the first images according to the multiple image features of different scales to obtain a second image; and determining the detection result of the target object according to the second image.

[0116] On the other hand, embodiments of this application also provide a processor-readable storage medium storing a computer program for causing a processor to execute the steps of a target detection method provided in the above embodiments, such as: performing target detection on an original video to obtain a first image containing a target object and multiple image features of different scales for each of the first images; filtering the first images according to the multiple image features of different scales to obtain a second image; and determining the detection result of the target object according to the second image.

[0117] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0118] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0119] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A target detection method, characterized in that, include: Target detection is performed on the original video to obtain a first image containing the target object and multiple image features of different scales for each of the first images; Based on the image features at multiple different scales, the first image is filtered to obtain the second image; The detection result of the target object is determined based on the second image.

2. The target detection method according to claim 1, characterized in that, The step of filtering the first image based on the multiple image features at different scales to obtain the second image includes: Obtain the pose features of the target object in the first image; Based on the image features at multiple different scales and the pose features, the first image is filtered to obtain the second image.

3. The target detection method according to claim 2, characterized in that, The step of filtering the first image based on the multiple image features at different scales and the pose features to obtain the second image includes: Based on the image features at multiple different scales, the first image is filtered to obtain the third image; Based on the degree of feature correlation between the various third images, the third images are filtered to obtain the fourth image; Based on the posture features, the fourth image is filtered to obtain the second image.

4. The target detection method according to any one of claims 1-3, characterized in that, The step of filtering the first image based on the multiple image features at different scales includes: Based on image features at multiple different scales of each first image, a feature fingerprint of each first image is determined; Determine the deviation value between each of the described feature fingerprints and the standard feature fingerprint of the target object; The first image is filtered based on the deviation value.

5. The target detection method according to claim 4, characterized in that, Determining the feature fingerprint of each first image based on multiple image features at different scales includes: For each of the first images, the average value of multiple feature maps included in the image features at each scale is determined; The feature fingerprint of each first image is determined based on the average value of image features at multiple different scales for each first image.

6. The target detection method according to claim 3, characterized in that, The step of filtering the third images based on the degree of feature correlation between each of the third images to obtain the fourth image includes: Acquire at least one image group, wherein one image group comprises a first number of third images with consecutive timestamps; Obtain image chains in each of the image groups, wherein an image chain includes at least a second number of third images in the image group to which it belongs, and the similarity of the feature fingerprints between adjacent third images in the image chain is greater than a first similarity threshold, and the second number is not greater than the first number; The third image included in each of the image chains is determined as the fourth image.

7. The target detection method according to claim 6, characterized in that, The step of filtering the fourth image based on the pose features to obtain the second image includes: Determine the average feature fingerprint of each image chain in the fourth image; The image chain is filtered based on the average feature fingerprint and the pose feature; The fourth image, which is included in the remaining image chain, is identified as the second image.

8. The target detection method according to claim 7, characterized in that, The step of filtering the image chain based on the average feature fingerprint and the pose feature includes: Obtain any two image chains whose average feature fingerprint similarity is greater than the second similarity threshold; Based on the pose features of the two image chains, determine the geometric relationship between the poses of the target objects contained in the two image chains. The image chain is filtered based on the geometric relationship.

9. The target detection method according to claim 8, characterized in that, The step of filtering the image chain based on the geometric relationship includes: If the geometric relationship is a mirror relationship, obtain the confidence score of the target object contained in each of the two image chains; Remove image chains with low confidence scores.

10. The target detection method according to claim 8, characterized in that, The step of filtering the image chain based on the geometric relationship includes: If the geometric relationship is a mirror image, obtain the sharpness of the fourth image contained in each of the two image chains; Remove image chains containing low-resolution fourth images.

11. The target detection method according to claim 1, characterized in that, The process of performing target detection on the original video to obtain a first image containing the target object and multiple image features of different scales for each of the first images includes: Target detection is performed on the original video to obtain initial detection results. The initial detection results include a fifth image containing the target object, multiple image features of different scales in each fifth image, and a confidence score of the target object in each fifth image. Based on the confidence score, the initial detection results are filtered to obtain the first image containing the target object and multiple image features of each first image at different scales, wherein the first image is the fifth image containing the target object whose confidence score is greater than a score threshold.

12. The target detection method according to claim 11, characterized in that, The step of performing target detection on the original video to obtain initial detection results includes: Obtain the video frames containing the target object from the original video; The video frame is input into the target detection model to obtain the initial detection result; The target detection model is obtained by training the MobileNetV3 network in advance with sample images containing target objects and the labels of the target objects carried in the sample images as input.

13. The target detection method according to claim 11, characterized in that, Determining the detection result of the target object based on the second image includes: The detection result of the target object is determined based on the bounding box coordinates and confidence score of the target object in the second image.

14. A target detection device, characterized in that, include: The detection module is used to perform target detection on the original video to obtain a first image containing the target object and multiple image features of different scales for each of the first images; The filtering module is used to filter the first image based on the multiple image features at different scales to obtain the second image; The determination module is used to determine the detection result of the target object based on the second image.

15. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the target detection method according to any one of claims 1 to 13.

16. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the target detection method as described in any one of claims 1 to 13.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the target detection method according to any one of claims 1 to 13.