Video detection method and device, storage medium and electronic equipment
Through the combined intercomparison comparison and center point matching method of anchorless object detection model and parallel backbone network, the problem of high missed detection rate of small and medium-sized targets in video detection is solved, and efficient and reliable video detection is achieved, suitable for complex scenarios and education fields.
Patent Information
- Application Number
- CN202510436124.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, manual labeling is low efficiency, high cost and high missed detection rate of small target prospect information during video detection, resulting in unreliable detection results.
An anchorless object detection model is used to extract features of the video image through a parallel backbone network, and match positive and negative samples with the interleaving ratio and center point matching method to directly predict the offset and width and height factors of the characteristics to be detected in the image relative to the preset key points to avoid missing detection of small targets.
It reduces the cost of video detection, improves the reliability and robustness of detection results, is suitable for complex scenarios, and promotes the deep integration of education and artificial intelligence.
Smart Images

Figure CN120472361A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a video detection method, device, storage medium, and electronic device. Background Art
[0002] The video detection process in related technologies involves manually labeling video data, then using a one-stage and two-stage object detection algorithm to detect logo information in the labeled video data, thereby achieving video detection. However, manual labeling is inefficient, costly, and has low accuracy. Furthermore, object detection methods have a high rate of underdetection of small foreground objects in videos, making video detection results unreliable. Summary of the Invention
[0003] The present invention aims to provide a video detection method, device, storage medium and electronic device, which can reduce the missed detection rate of small target foreground information in a video.
[0004] In order to achieve the above objectives, in a first aspect, the present disclosure provides a video detection method, comprising: Get the target video; Extract features from each frame of the target video to obtain a feature atlas corresponding to the target video; The feature atlas is input into the anchor-free target detection model to obtain the detection atlas output by the anchor-free target detection model. The anchor-free target detection model is used to predict the offset of the feature to be detected in the image relative to the preset key point and the width and height factor of the feature to be detected.
[0005] Optionally, extracting features from each frame of the target video to obtain a feature atlas corresponding to the target video includes: Perform feature extraction on each frame of the target video through a feature extraction network to obtain a feature atlas output by the feature extraction network; The feature extraction network model includes parallel backbone networks, and the sampling rate of each backbone network is different.
[0006] Optionally, performing feature extraction on each frame of the target video through a feature extraction network to obtain a feature atlas output by the feature extraction network includes: For each frame of the target video, extract features of the image through each backbone network to obtain multiple candidate feature images of different resolutions corresponding to the image, and fuse the multiple candidate feature images of different resolutions to obtain a feature image corresponding to the image; The feature images corresponding to each frame image in the target video are integrated into a feature atlas.
[0007] Optionally, the training process of the anchor-free object detection model includes: Acquire a training feature atlas, and perform positive and negative sample matching on each training feature image in the training feature atlas to balance the positive and negative samples in the training feature atlas; The target detection model is trained by training feature atlas after matching positive and negative samples to obtain an anchor-free target detection model.
[0008] Optionally, performing positive and negative sample matching on each training feature image in the training feature atlas includes: Determine a first sample set based on the training feature atlas by using an intersection-over-union matching method, and determine a second sample set based on the training feature atlas by using a center point matching method; An intersection of the first sample set and the second sample set is determined, and the training feature images in the intersection are marked as positive samples, and the other training feature images are marked as negative samples.
[0009] Optionally, determining the first sample set according to the training feature graph set by using an intersection-over-union matching method includes: For each training feature image in the training feature atlas, scaling the real frame of the training feature image according to the step size of the training feature image; Calculating the overlap rate between the scaled true frame and the predicted frame of the training feature image; When the overlap rate is less than an overlap threshold, the training feature image is determined to be a first sample.
[0010] Optionally, determining the second sample set according to the training feature atlas using a center point matching method includes: For each training feature image in the training feature atlas, scaling the real frame of the training feature image according to the step size of the training feature image; Calculating the distance between the center point of the scaled real frame and the center point of the feature to be detected in the training feature image; When the distance is less than a distance threshold, the training feature image is determined to be a second sample.
[0011] In a second aspect, the present disclosure provides a video detection device, comprising: Acquisition module, used to acquire target video; A feature extraction module is used to extract features from each frame of the target video to obtain a feature atlas corresponding to the target video; A detection module is used to input the feature atlas into the anchor-free target detection model to obtain the detection atlas output by the anchor-free target detection model. The anchor-free target detection model is used to predict the offset of the detection feature in the image relative to the preset key point and the width and height factors of the feature to be detected.
[0012] In a third aspect, the present disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned video detection method when executed by a processor.
[0013] In a fourth aspect, the present disclosure provides an electronic device, comprising: a memory having a computer program stored thereon; A processor is used to execute the computer program in the memory to implement the steps of the above-mentioned video detection method.
[0014] Through the above technical solution, a target video is acquired, a corresponding feature atlas is obtained, and the feature atlas is input into an anchor-free object detection model, resulting in a detection atlas output by the anchor-free object detection model. The anchor-free object detection model can directly predict the offset of the feature to be detected in the image relative to a preset key point, as well as the width and height factors of the feature to be detected, thus preventing small objects from being missed. The entire detection process eliminates the need for manual labeling, reducing video detection costs and ensuring the reliability of video detection results.
[0015] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following detailed description, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings: Figure 1 The figure is a flowchart of a video detection method according to an exemplary embodiment.
[0017] Figure 2 is a schematic diagram of a feature extraction network according to an exemplary embodiment.
[0018] Figure 3 The figure is a flowchart of positive sample matching according to an exemplary embodiment.
[0019] Figure 4 The figure is a flowchart of an intersection-over-union matching method according to an exemplary embodiment.
[0020] Figure 5 The figure is a flowchart of a center point matching method according to an exemplary embodiment.
[0021] Figure 6 is another flow chart of a video detection method according to an exemplary embodiment.
[0022] Figure 7 The figure is a block diagram of a video detection device according to an exemplary embodiment.
[0023] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0024] The following describes the specific embodiments of the present disclosure in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present disclosure and are not intended to limit the present disclosure.
[0025] It should be noted that all actions of obtaining video information in the present disclosure are performed in compliance with the corresponding data protection laws and policies of the country where the device is located and with the authorization given by the owner of the corresponding device.
[0026] As mentioned in the background art, in related technologies, video detection is achieved by first manually labeling video data, then using a first-stage and a second-stage object detection algorithm to detect identification information in the labeled video data. However, the inventors have discovered that the manual labeling stage suffers from low efficiency, high cost, and low accuracy. Furthermore, the object detection stage has a high rate of missed detection of small foreground objects in the video, resulting in unreliable video detection results.
[0027] In view of this, the present disclosure provides a video detection method, apparatus, storage medium and electronic device, which can reduce the missed detection rate of small target foreground information in a video, thereby improving the reliability of the video detection results.
[0028] Figure 1 is a flow chart showing a video detection method according to an exemplary embodiment. Figure 1 The video detection method can be used in various complex scenarios such as scanning and photographing. The video detection method may include the following steps: In step S11, a target video is acquired.
[0029] It is worth noting that the target video may be a lecture video, in which there are many small areas of foreground information, and the foreground information may include information such as the subject, type, and scene of the lecture video.
[0030] In step S12, feature extraction is performed on each frame image in the target video to obtain a feature atlas corresponding to the target video.
[0031] In step S13, the feature atlas is input into the anchor-free object detection (Anchor-Free Fully Convlutional One-Stage, Anchor-Free FCOS) model to obtain the detection atlas output by the anchor-free object detection model. The anchor-free object detection model is used to predict the offset of the feature to be detected in the image relative to the preset key point and the width and height factors of the feature to be detected.
[0032] It is worth noting that the preset key point can be the point at the upper left corner of the grid, and the grid is the real frame of the feature to be detected in the image during the image detection process of the anchor-free target detection model.
[0033] It is worth noting that the anchor-free target detection model will set several fixed-size anchors at the location of each feature to be detected on the output detection image, associate the predicted box with the real box, and there is only one predicted box at the location of each feature to be detected on the output detection image, thereby directly predicting the offset of the center point of the feature to be detected relative to the upper left corner of the grid and the width and height factors of the feature to be detected.
[0034] The present invention determines the feature atlas corresponding to the target video, inputs the feature atlas into the anchor-free target detection model, and obtains the detection atlas output by the anchor-free target detection model. The anchor-free target detection model can directly predict the offset of the feature to be detected in the image relative to the preset key point and the width and height factors of the feature to be detected, thereby avoiding the omission of small targets. No manual labeling is required during the entire detection process, which reduces the cost of video detection and ensures the reliability of the video detection results. It also has certain anti-watermark, anti-light and anti-noise capabilities, thereby ensuring the robustness of the anchor-free target detection model in complex scenes, making the video detection method highly applicable, and performing video detection on teaching videos can promote the deep integration of education and artificial intelligence.
[0035] In order to facilitate those skilled in the art to better understand the video detection method provided by the present disclosure, the relevant steps involved in the video detection method are described in detail below with examples.
[0036] In a possible embodiment, in step S12, feature extraction is performed on each frame of the target video to obtain a feature atlas corresponding to the target video, which may include: Perform feature extraction on each frame of the target video through the feature extraction network to obtain the feature atlas output by the feature extraction network; The feature extraction network model includes parallel backbone networks, and the sampling rate of each backbone network is different.
[0037] It's worth noting that small object foreground information carries very little RGB (red, green, and blue) information. If feature extraction is performed using a serial backbone network that includes multiple continuous convolutional pooling layers, this information will be severely lost during the convolutional pooling process, meaning that information will be lost as the resolution increases. Therefore, in the disclosed embodiments, a parallel backbone network is used for feature extraction. This parallel backbone network can include VGGNet (Visual Geometry Group Network, a deep convolutional neural network) or ResNet (Deep Residual Network, a deep residual neural network).
[0038] This disclosed embodiment uses a parallel backbone network with different sampling rates to extract features from each frame of the target video, preserving the features to be detected in the image to the greatest extent possible without loss, thereby reducing the rate of missed detection of small foreground objects in the video. Furthermore, both the parallel backbone network and the anchor-free object detection model support service deployment on high-performance GPUs (Graphics Processing Units). This deployment, based on a multi-process service framework, meets high concurrency requirements and ensures robust server calls.
[0039] In a possible embodiment, feature extraction is performed on each frame of the target video through a feature extraction network to obtain a feature atlas output by the feature extraction network, which may include: For each frame of the target video, each backbone network extracts features from the image to obtain multiple candidate feature images of different resolutions corresponding to the image, and then fuses the features of the multiple candidate feature images of different resolutions to obtain the feature image corresponding to the image; The feature images corresponding to each frame image in the target video are integrated into a feature atlas.
[0040] It is worth noting that by extracting features from the image through backbone networks with different sampling rates, multiple candidate feature images of different resolutions of the corresponding image are obtained, and feature fusion is performed on the candidate feature images of different resolutions, thereby improving the feature expression ability of the candidate feature images of different resolutions and realizing information interaction between backbone network branches with different resolutions, so that the features to be detected in the image can be retained losslessly to the greatest extent.
[0041] It's worth noting that traditional backbone networks lose information as they scale from high to low resolution during extraction, resulting in very low-resolution feature images and a loss of spatial structure. For example, the feature images obtained by VGGNet and ResNet are very low-resolution and lack spatial structure, requiring upsampling or deconvolution to restore them to high-resolution representations.
[0042] For example, see Figure 2The parallel backbone network includes three backbone networks with different sampling rates, each backbone network is 3×3 convolution. The image is upsampled and downsampled by the three backbone networks with different sampling rates to realize feature extraction, and three candidate feature images with different resolutions are obtained. The sizes of the three candidate feature images are 20×20, 40×40, and 80×80 respectively. Then, the three candidate feature images with different resolutions of 20×20, 40×40, and 80×80 are feature fused to obtain the feature image corresponding to the image.
[0043] The parallel backbone network in the embodiment of the present disclosure restores the candidate feature image to a high-resolution representation through upsampling or deconvolution, and runs multiple branches of different resolutions in parallel. Through information interaction between different branches, the features to be detected in the image can be retained losslessly to the greatest extent, reducing feature loss during the sampling process.
[0044] In one possible embodiment, the training process of the anchor-free object detection model includes: A training feature atlas is obtained, and positive and negative sample matching is performed on each training feature image in the training feature atlas, so that the positive and negative samples in the training feature atlas are balanced.
[0045] The target detection model is trained by training feature atlas after matching positive and negative samples to obtain an anchor-free target detection model.
[0046] For example, when the training feature images are sized 20×20, 40×40, and 80×80, the number of predicted boxes corresponding to the target video is 8,400 (20×20 + 40×40 + 80×80). Four target bounding box parameters (x, y, w, h) are directly predicted at each feature location to be detected. These four parameters correspond to the offset of the center point of the feature to be detected relative to the upper left corner of the grid cell, and the width and height factors of the feature to be detected. Because the target bounding box parameters are all relative to the scale of the predicted feature map, mapping back to the original image requires multiplying the stride of the feature map relative to the original image. A portion of positive sample boxes is then selected based on the true annotated boxes and the predicted boxes. By combining the intersection-over-union matching method with the center point matching method, predicted boxes that match the true boxes and can be used as positive samples are obtained.
[0047] It is worth noting that during the model training process, the removal of negative samples may have an impact on the performance of the model. Specifically, if there are a large number of negative sample images in the training set, and these images are not removed, the model may learn wrong features and fail to correctly identify the target object. Therefore, in the embodiment of the present disclosure, by matching positive and negative samples of the training feature atlas, the number and quality of positive samples in the training feature atlas are improved to avoid the imbalance of positive and negative samples in the training feature atlas, thereby improving the reliability of the detection results of the anchor-free target detection model. In addition, the anchor-free target detection model relies on massive and diversified data for training, which ensures the generalization ability of the anchor-free target detection model, and based on the network structure of the feature pyramid, it can effectively detect the area where the foreground information of small targets is located, and effectively solve the problem of large size span and large aspect ratio changes of the features to be detected, thereby ensuring the high accuracy and high recall rate of the anchor-free target model.
[0048] In one possible embodiment, see Figure 3 , performing positive and negative sample matching on each training feature image in the training feature atlas may include: In step S31, a first sample set is determined based on the training feature atlas by using an intersection-over-union matching method, and a second sample set is determined based on the training feature atlas by using a center point matching method.
[0049] In step S32, the intersection of the first sample set and the second sample set is determined, the training feature images in the intersection are marked as positive samples, and the other training feature images are marked as negative samples.
[0050] In the embodiment of the present disclosure, two sample sets are determined in two different ways, and the training feature images in the intersection of the two sample sets are marked as positive samples, which increases the number of high-quality positive samples and effectively improves the recall rate and accuracy of the features to be detected.
[0051] In one possible embodiment, see Figure 4 In step S31, determining the first sample set according to the training feature graph set by using the intersection-over-union matching method may include: In step S41 , for each training feature image in the training feature atlas, the ground truth frame of the training feature image is scaled according to the step size of the training feature image.
[0052] It is worth noting that in order to map the training feature image back to the original image, it is necessary to multiply the stride of the current training feature image relative to the original image. Therefore, before making a judgment, the real frame of the training feature image is scaled according to the stride of the training feature image.
[0053] In step S42, the overlap ratio between the scaled true frame and the predicted frame of the training feature image is calculated.
[0054] It is worth noting that the IOU (intersection-over-union) value of the scaled true box and the predicted box is calculated to determine the overlap rate between the true box and the predicted box.
[0055] In step S43 , when the overlap rate is less than the overlap threshold, the training feature image is determined to be the first sample.
[0056] It is worth noting that the overlap threshold can be preset according to the quality requirements of the positive samples, which is not limited in the embodiments of the present disclosure. The higher the quality requirements, the larger the overlap threshold; the lower the quality requirements, the smaller the overlap threshold.
[0057] In one possible embodiment, see Figure 5 In step S31, determining the second sample set based on the training feature graph set by using a center point matching method may include: In step S51 , for each training feature image in the training feature atlas, the ground truth frame of the training feature image is scaled according to the step size of the training feature image.
[0058] It is worth noting that in order to map the training feature image back to the original image, it is necessary to multiply the stride of the current training feature image relative to the original image. Therefore, before making a judgment, the real frame of the training feature image is scaled according to the stride of the training feature image.
[0059] In step S52 , the distance between the center point of the scaled real frame and the center point of the feature to be detected in the training feature image is calculated.
[0060] In step S53 , when the distance is less than the distance threshold, the training feature image is determined to be the second sample.
[0061] In the disclosed embodiment, two sample sets are determined by the intersection-over-union matching method and the center point matching method, respectively, and the training feature images in the intersection of the two sample sets are marked as positive samples, thereby increasing the number of high-quality positive samples and effectively improving the recall rate and accuracy of the features to be detected.
[0062] For example, see Figure 6 The video detection method provided by the present disclosure may further include the following steps: In step S61 , a target video is acquired.
[0063] In step S62, for each frame image in the target video, the image is feature extracted through a parallel backbone network to obtain multiple candidate feature images of different resolutions corresponding to the image, and the multiple candidate feature images of different resolutions are feature fused to obtain a feature image corresponding to the image.
[0064] In step S63, the feature images corresponding to each frame image in the target video are integrated into a feature atlas.
[0065] In step S64, the feature atlas is input into the anchor-free target detection model to obtain a detection atlas output by the anchor-free target detection model. The anchor-free target detection model is used to predict the offset of the feature to be detected in the image relative to the preset key point and the width and height factors of the feature to be detected.
[0066] The disclosed embodiment extracts features from each frame of the target video through a parallel backbone network to obtain multiple candidate feature images of different resolutions corresponding to the image, thereby avoiding the loss of foreground information of small targets in the video image. The multiple candidate feature images of different resolutions are feature-fused to obtain the feature image corresponding to the image, thereby improving the feature expression capability of the candidate images of different resolutions, so that the features to be detected in the image can be retained losslessly to the greatest extent. The feature atlas is input into the anchor-free target detection model to obtain a detection atlas, thereby avoiding the omission of foreground information of small targets, eliminating the need for manual labeling, and ensuring the accuracy and reliability of the video detection results at low cost.
[0067] Based on the same inventive concept, the present disclosure also provides a video detection device, see Figure 7 The video detection device includes an acquisition module 701, a feature extraction module 702 and a detection module 703.
[0068] The acquisition module 701 is used to acquire a target video.
[0069] The feature extraction module 702 is used to extract features from each frame of the target video to obtain a feature atlas corresponding to the target video.
[0070] The detection module 703 is used to input the feature atlas into the anchor-free target detection model to obtain the detection atlas output by the anchor-free target detection model. The anchor-free target detection model is used to predict the offset of the detection feature in the image relative to the preset key point and the width and height factors of the feature to be detected.
[0071] The present invention determines the feature atlas corresponding to the target video, inputs the feature atlas into the anchor-free target detection model, and obtains the detection atlas output by the anchor-free target detection model. The anchor-free target detection model can directly predict the offset of the feature to be detected in the image relative to the preset key point and the width and height factors of the feature to be detected, thereby avoiding the omission of small targets. No manual labeling is required during the entire detection process, which reduces the cost of video detection and ensures the reliability of the video detection results. It also has certain anti-watermark, anti-light and anti-noise capabilities, thereby ensuring the robustness of the anchor-free target detection model in complex scenes, making the video detection method highly applicable, and performing video detection on teaching videos can promote the deep integration of education and artificial intelligence.
[0072] In a possible embodiment, the feature extraction module 702 is configured to extract features from each frame of the target video through a feature extraction network to obtain a feature atlas output by the feature extraction network; The feature extraction network model includes parallel backbone networks, and the sampling rate of each backbone network is different.
[0073] In one possible embodiment, the feature extraction module 702 is configured to extract features of each frame of the target video through each backbone network to obtain a plurality of candidate feature images of different resolutions corresponding to the image, and perform feature fusion on the plurality of candidate feature images of different resolutions to obtain a feature image corresponding to the image; The feature images corresponding to each frame image in the target video are integrated into a feature atlas.
[0074] In a possible embodiment, the video detection device further includes a training module, the training module being configured to obtain a training feature atlas and perform positive and negative sample matching on each training feature image in the training feature atlas, so as to balance the positive and negative samples in the training feature atlas; The target detection model is trained by training feature atlas after matching positive and negative samples to obtain an anchor-free target detection model.
[0075] In a possible embodiment, the training module is configured to determine the first sample set based on the training feature atlas by using an intersection-over-union matching method, and to determine the second sample set based on the training feature atlas by using a center point matching method; An intersection of the first sample set and the second sample set is determined, and the training feature images in the intersection are marked as positive samples, and the other training feature images are marked as negative samples.
[0076] In a possible embodiment, the training module is configured to scale the ground truth frame of each training feature image in the training feature atlas according to the step size of the training feature image; Calculate the overlap rate between the scaled true frame and the predicted frame of the training feature image; When the overlap rate is less than the overlap threshold, the training feature image is determined to be the first sample.
[0077] In a possible embodiment, the training module is configured to scale the ground truth frame of each training feature image in the training feature atlas according to the step size of the training feature image; Calculate the distance between the center point of the scaled true frame and the center point of the feature to be detected in the training feature image; When the distance is less than the distance threshold, the training feature image is determined to be the second sample.
[0078] Regarding the video detection device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0079] Based on the same inventive concept, the present disclosure further provides an electronic device, including: a memory having a computer program stored thereon; The processor is used to execute the computer program in the memory to implement the steps of the above video detection method.
[0080] The present invention determines the feature atlas corresponding to the target video, inputs the feature atlas into the anchor-free target detection model, and obtains the detection atlas output by the anchor-free target detection model. The anchor-free target detection model can directly predict the offset of the feature to be detected in the image relative to the preset key point and the width and height factors of the feature to be detected, thereby avoiding the omission of small targets. No manual labeling is required during the entire detection process, which reduces the cost of video detection and ensures the reliability of the video detection results. It also has certain anti-watermark, anti-light and anti-noise capabilities, thereby ensuring the robustness of the anchor-free target detection model in complex scenes, making the video detection method highly applicable, and performing video detection on teaching videos can promote the deep integration of education and artificial intelligence.
[0081] Figure 8 FIG. 8 is a block diagram of an electronic device 800 according to an exemplary embodiment. Figure 8 As shown, the electronic device 800 may include: a processor 801 , a memory 802 , and may further include one or more of a multimedia component 803 , an input / output (I / O) interface 804 , and a communication component 805 .
[0082] The processor 801 is used to control the overall operation of the electronic device 800 to complete all or part of the steps in the above-mentioned video detection method. The memory 802 is used to store various types of data to support the operation of the electronic device 800. This data may include, for example, instructions for any application or method operating on the electronic device 800, as well as application-related data such as contact information, sent and received messages, images, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 804 provides an interface between the processor 801 and other interface modules, which may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 805 is used for wired or wireless communication between the electronic device 800 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more thereof, is not limited here. Therefore, the corresponding communication component 805 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.
[0083] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned video detection method.
[0084] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-described video detection method are implemented. For example, the computer-readable storage medium may be the aforementioned memory 802 including the program instructions. The program instructions may be executed by the processor 801 of the electronic device 800 to implement the above-described video detection method.
[0085] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device, and has a code portion for performing the above-mentioned video detection method when the computer program is executed by the programmable device.
[0086] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the scope of protection of the present disclosure.
[0087] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.
[0088] In addition, the various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.
Claims
1. A video detection method, characterized in that: include: Get the target video; Extract features from each frame of the target video to obtain a feature atlas corresponding to the target video; The feature atlas is input into the anchor-free target detection model to obtain the detection atlas output by the anchor-free target detection model. The anchor-free target detection model is used to predict the offset of the feature to be detected in the image relative to the preset key point and the width and height factor of the feature to be detected.
2. The video detection method according to claim 1, wherein: The step of extracting features from each frame of the target video to obtain a feature atlas corresponding to the target video includes: Perform feature extraction on each frame of the target video through a feature extraction network to obtain a feature atlas output by the feature extraction network; The feature extraction network model includes parallel backbone networks, and the sampling rate of each backbone network is different.
3. The video detection method according to claim 2, characterized in that: The feature extraction is performed on each frame of the target video through the feature extraction network to obtain a feature atlas output by the feature extraction network, including: For each frame of the target video, extract features of the image through each backbone network to obtain multiple candidate feature images of different resolutions corresponding to the image, and fuse the multiple candidate feature images of different resolutions to obtain a feature image corresponding to the image; The feature images corresponding to each frame image in the target video are integrated into a feature atlas.
4. The video detection method according to any one of claims 1 to 3, characterized in that: The training process of the anchor-free target detection model includes: Acquire a training feature atlas, and perform positive and negative sample matching on each training feature image in the training feature atlas to balance the positive and negative samples in the training feature atlas; The target detection model is trained by training feature atlas after matching positive and negative samples to obtain an anchor-free target detection model.
5. The video detection method according to claim 4, characterized in that: The performing positive and negative sample matching on each training feature image in the training feature atlas includes: Determine a first sample set based on the training feature atlas by using an intersection-over-union matching method, and determine a second sample set based on the training feature atlas by using a center point matching method; An intersection of the first sample set and the second sample set is determined, and training feature images in the intersection are marked as positive samples, and other training feature images are marked as negative samples.
6. The video detection method according to claim 5, characterized in that: The determining of the first sample set according to the training feature graph set by using the intersection-over-union matching method includes: For each training feature image in the training feature atlas, scaling the real frame of the training feature image according to the step size of the training feature image; Calculating the overlap rate between the scaled true frame and the predicted frame of the training feature image; When the overlap rate is less than an overlap threshold, the training feature image is determined to be a first sample.
7. The video detection method according to claim 5, characterized in that: The determining of the second sample set according to the training feature atlas by a center point matching method includes: For each training feature image in the training feature atlas, scaling the real frame of the training feature image according to the step size of the training feature image; Calculating the distance between the center point of the scaled real frame and the center point of the feature to be detected in the training feature image; When the distance is less than a distance threshold, the training feature image is determined to be a second sample.
8. A video detection device, characterized in that: include: Acquisition module, used to acquire target video; A feature extraction module is used to extract features from each frame of the target video to obtain a feature atlas corresponding to the target video; A detection module is used to input the feature atlas into the anchor-free target detection model to obtain the detection atlas output by the anchor-free target detection model. The anchor-free target detection model is used to predict the offset of the detection feature in the image relative to the preset key point and the width and height factors of the feature to be detected.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the video detection method according to any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the video detection method according to any one of claims 1 to 7.