Marking detection method, device, electronic device and storage medium

By identifying preset logos and candidate areas in video frames and combining neural networks with frame extraction detection, the problem of missed detection in existing logo detection technologies is solved, and efficient and accurate detection of complex and diverse logos is achieved.

CN114842379BActive Publication Date: 2025-09-16BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210454747.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-24
Publication Date
2025-09-16
Estimated Expiration
2042-04-24

AI Technical Summary

Technical Problem

Existing logo detection solutions have difficulty in fully identifying logos in videos and are prone to missed detections. In particular, they are unable to flexibly cope with complex and diverse logo changes and the high time cost of video frames.

Method used

By identifying the first area of ​​the preset logo and the first candidate area other than the preset logo in the video frame, a neural network is used to perform graphic and text detection, the second area is screened out based on the degree of overlap and distance, the logo area is determined, and frame extraction detection is used for time domain detection.

Benefits of technology

It achieves comprehensive detection of logos in the video, improves the comprehensiveness and accuracy of detection, avoids missed detection, and supports the detection of multiple types of logos while maintaining efficient detection, solving the problem of balancing detection speed and effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842379B_ABST
    Figure CN114842379B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device, electronic device, and storage medium for detecting a marker. The method comprises: obtaining a target video frame in a video to be detected; identifying a first region and a first candidate region in the target video frame, wherein the first region includes a preset marker and the first candidate region includes marker content other than the preset marker; determining a first candidate region in the first candidate region whose distance from the first region is less than a preset value as a second region; and determining the first region and the second region as marker regions in the target video frame. The method, device, electronic device, and storage medium for detecting a marker according to the present disclosure can solve the problem that existing detection schemes have difficulty in completely identifying markers and are prone to missed detection, and can achieve comprehensive detection of markers in videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a method and device for detecting an identifier, an electronic device, and a storage medium. Background Art

[0002] Videos circulating on the internet often carry various logos, such as icons and watermarks. For example, some video editing software sometimes automatically adds various tool icons and watermarks to user videos. During post-production video editing, it's common to want to avoid logos or use videos with logos, such as station logos.

[0003] Generally speaking, icons in videos have the following complex characteristics: there are many types and styles of logos; the logos are a mixture of graphics and text; the text near the logos can be changed; the logos are transparent or semi-transparent; the logo positions move in the video, etc.

[0004] However, due to the complexity of the logo, it is difficult for existing detection solutions to fully identify the logo, and it is easy to miss the detection. Summary of the Invention

[0005] The present disclosure provides a method, device, electronic device, and storage medium for detecting a marker, which at least solves the problem in related technologies that it is difficult to completely identify a marker and that misses a marker are likely to occur. The technical solution of the present disclosure is as follows:

[0006] According to a first aspect of an embodiment of the present disclosure, a method for detecting an identifier is provided, which includes: obtaining a target video frame in a video to be detected; identifying a first area and a first candidate area in the target video frame, wherein the first area includes a preset identifier and the first candidate area includes identifier content other than the preset identifier; determining a first candidate area in the first candidate area whose distance from the first area is less than a preset value as a second area; and determining the first area and the second area as identifier areas in the target video frame.

[0007] Optionally, the preset identifier includes at least one of a preset graphic identifier and a preset text identifier, wherein the step of determining the first area includes: performing graphic identifier detection and text identifier detection on the target video frame to obtain a second candidate area containing the preset graphic identifier and / or a second candidate area containing the preset text identifier; and fusing the obtained second candidate areas to determine the first area.

[0008] Optionally, the obtained second candidate areas are fused to determine the first area, including: merging second candidate areas whose overlap is greater than a first preset value to obtain a new second candidate area; when the overlap between any two second candidate areas is less than or equal to the first preset value, using the second candidate area as the first area.

[0009] Optionally, the obtained second candidate areas are fused to obtain the first area, including: determining the degree of overlap between each two second candidate areas, and when the degree of overlap is greater than a first preset value, merging the corresponding two second candidate areas and using the merged area as the first area; when the degree of overlap is less than or equal to the first preset value, using the corresponding two second candidate areas as the first areas respectively.

[0010] Optionally, the identification content other than the preset identification includes text, wherein the step of determining the first candidate area includes: performing text detection on the target video frame to obtain a first candidate area containing text.

[0011] Optionally, the step of determining a first candidate area in the first candidate area whose distance from the first area is less than a preset value as the second area includes: expanding the first area to obtain an expanded first area; determining the degree of overlap between the expanded first area and the first candidate area, and determining the first candidate area whose degree of overlap is greater than a second preset value as the second area.

[0012] Optionally, the step of determining a first candidate area in the first candidate area whose distance from the first area is less than a preset value as the second area includes: determining the first candidate area that meets a preset distance condition as the second area, wherein the preset distance condition includes: the distance from the first area in the height direction of the target video frame is less than a first preset distance; and / or the distance from the first area in the width direction of the target video frame is less than a second preset distance.

[0013] Optionally, the step of obtaining a target video frame in the video to be detected includes: extracting video frames from the video to be detected as the target video frames at preset time intervals, wherein the identification detection method also includes: determining, for the intermediate video frame between the current target video frame and the adjacent target video frames before and / or after the current target video frame, the regional similarity between the current target video frame and the intermediate video frame in the target area, wherein the target area is the first area and the second area corresponding to the current target video frame; when the intermediate video frame meets a preset condition, determining the target area as the identification area in the intermediate video frame, wherein the preset condition is: the proportion of the area of ​​the similar area in the intermediate video frame to the total area of ​​the target area is higher than a proportion threshold, wherein the similar area refers to the area where the regional similarity is higher than the similarity threshold.

[0014] According to a second aspect of an embodiment of the present disclosure, there is provided an identification detection device, comprising: an acquisition unit configured to acquire a target video frame in a video to be detected; a first determination unit configured to identify a first area and a first candidate area in the target video frame, wherein the first area contains a preset identification and the first candidate area contains identification content other than the preset identification; a second determination unit configured to determine a first candidate area in the first candidate area whose distance from the first area is less than a preset value as a second area; and a third determination unit configured to determine the first area and the second area as identification areas in the target video frame.

[0015] Optionally, the preset identifier includes at least one of a preset graphic identifier and a preset text identifier, wherein the first determination unit is further configured to: perform graphic identifier detection and text identifier detection on the target video frame to obtain a second candidate area containing the preset graphic identifier and / or a second candidate area containing the preset text identifier; and fuse the obtained second candidate areas to determine the first area.

[0016] Optionally, the first determination unit is further configured to: merge second candidate areas whose degree of overlap is greater than a first preset value to obtain a new second candidate area; when the degree of overlap between any two second candidate areas is less than or equal to the first preset value, use the second candidate area as the first area.

[0017] Optionally, the first determination unit is further configured to: determine the degree of overlap between each two second candidate areas, and when the degree of overlap is greater than a first preset value, merge the corresponding two second candidate areas and use the merged area as the first area; when the degree of overlap is less than or equal to the first preset value, use the corresponding two second candidate areas as the first areas respectively.

[0018] Optionally, the identification content other than the preset identification includes text, wherein the first determination unit is further configured to: perform text detection on the target video frame to obtain a first candidate region containing text.

[0019] Optionally, the second determination unit is further configured to: expand the first area to obtain an expanded first area; determine the degree of overlap between the expanded first area and the first candidate area, and determine the first candidate area whose overlap degree is greater than a second preset value as the second area.

[0020] Optionally, the second determination unit is further configured to: determine the first candidate area that meets the preset distance condition as the second area, wherein the preset distance condition includes: the distance between the first area and the target video frame in the height direction is less than a first preset distance; and / or the distance between the first area and the target video frame in the width direction is less than a second preset distance.

[0021] Optionally, the acquisition unit is further configured to: extract video frames from the video to be detected as the target video frames at preset time intervals, wherein the identification detection device also includes a fourth determination unit, and the fourth determination unit is configured to: determine the regional similarity between the current target video frame and the intermediate video frame in the target area for the intermediate video frame between the current target video frame and the adjacent target video frames before and / or after the current target video frame, wherein the target area is the first area and the second area corresponding to the current target video frame; when the intermediate video frame meets a preset condition, determine the target area as the identification area in the intermediate video frame, wherein the preset condition is: the proportion of the area of ​​the similar area in the intermediate video frame to the total area of ​​the target area is higher than a proportion threshold, wherein the similar area refers to the area where the regional similarity is higher than the similarity threshold.

[0022] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor, wherein the instructions executable by the processor, when executed by the processor, prompt the processor to execute the identification detection method according to the present disclosure.

[0023] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor, the processor is enabled to execute the identification detection method according to the present disclosure.

[0024] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes computer instructions, and when the computer instructions are executed by a processor, the identification detection method according to the present disclosure is implemented.

[0025] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:

[0026] According to the identification detection method of the embodiment of the present disclosure, the first area containing a preset identification and the first candidate area containing identification content other than the preset identification can be respectively determined, and the second area whose distance from the first area is less than a preset value can be screened out from the first candidate area, and the first area and the second area can be determined as identifications detected from the video frame. This solves the problem that existing detection schemes are difficult to fully identify identifications and are prone to missed detections, and can achieve comprehensive detection of identifications in videos.

[0027] In addition, according to the identification detection method of the embodiment of the present disclosure, by determining the first area and the second area as the identification detected from the video frame, the problem of being unable to detect icon types outside the range of a known identification data set is solved. Known identifications and other identifications outside the range of known identifications can be detected, thereby improving the comprehensiveness and accuracy of detection and avoiding missed detections.

[0028] In addition, according to the logo detection method of the embodiment of the present disclosure, by respectively determining the first area containing the preset logo and the first candidate area containing logo content other than the preset logo, it can support the detection of multiple types of logos and can process logos of icon type, text type, custom text style, and custom style.

[0029] In addition, the identification detection method according to the embodiment of the present disclosure can adopt a frame extraction detection method, which adopts a "time domain detection" scheme of first extracting the target video frame for detection and then detecting the intermediate video frame. It can obtain frame-by-frame detection results while maintaining efficient detection, solving the problem in related technologies that cannot simultaneously guarantee detection speed and detection effect, and achieving a balance between efficiency and effect.

[0030] In addition, the detection algorithm adopted by the identification detection method according to the embodiment of the present disclosure is relatively flexible, and each submodule in the detection process can be freely replaced with other alternatives.

[0031] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0033] Figure 1 is a schematic diagram illustrating a conventional icon detection method.

[0034] Figure 2A is a diagram illustrating a first example of identification in a video.

[0035] Figure 2B is a diagram illustrating a second example of identification in a video.

[0036] Figure 3 The figure is a flow chart showing a method for detecting a marker according to an exemplary embodiment.

[0037] Figure 4 The figure is a flowchart showing a step of determining a first area in a method for detecting a marker according to an exemplary embodiment.

[0038] Figure 5 It is a schematic diagram of a framework of an implementation example of a marker detection method according to an exemplary embodiment.

[0039] Figure 6 The figure is a block diagram showing an identification detection device according to an exemplary embodiment.

[0040] Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0041] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0042] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0043] It should be noted that the phrase "at least one of the items" in this disclosure includes three types of parallel situations: "any one of the items", "a combination of any multiple items of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. For another example, "performing at least one of step 1 and step 2" includes the following three parallel situations: (1) performing step 1; (2) performing step 2; and (3) performing steps 1 and 2.

[0044] It should be noted that although the application scenario of the station logo is explained below as an example, it should be understood that the application scenario of the identification detection method, device, electronic device and storage medium according to the present disclosure is not limited to this, and it can also be applied to any other video identification application scenario.

[0045] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0046] As mentioned above, existing detection schemes can usually only detect specific types of logos (such as icons and watermarks) in a single image. However, on the one hand, they cannot flexibly deal with logos outside the scope of the dataset; on the other hand, simply applying it to every frame of the video is too time-consuming.

[0047] In one existing logo detection solution, the logo detection of a single image can be achieved by training a convolutional neural network (CNN). Figure 1 Taking the logo detection method shown above as an example, a YOLO (You Only Look Once) network architecture can be used. Specifically, the network returns a series of rectangular coordinate boxes for the input image, indicating the location of each box and the corresponding logo category. When processing video, the video is typically sampled, for example, one frame every 10 seconds. Detection is then performed on each sampled video frame to achieve video detection.

[0048] Because this solution is not fast enough, and videos often contain a large number of video frames, applying this solution to all video frames can result in very slow detection speeds. Therefore, such solutions are usually only used in video review tasks. Specifically, in review tasks, the video may be sampled, and the logo is detected in the sampled video frames. The detection results of the sampled video frames are used as the detection results of the entire video for subsequent processing.

[0049] In addition, existing logo detection models and algorithms are usually only targeted at specific logo styles. Neural networks will record specific logo color, shape, and pattern information and detect it when it appears again. However, video logos such as network station logos are updated quickly and have many styles. Once a style update occurs, it is easy to miss the detection problem. In addition, in many online videos, there will be some user-related text near the video logo, such as Figure 2A and Figure 2B The text lines below the video logo shown may be, for example, lines corresponding to the user ID (e.g. Figure 2A as shown) or temporary activities (such as Figure 2B In addition, as mentioned above, the algorithm has a high time cost and low detection efficiency when detecting each video frame.

[0050] Another existing logo detection solution is a semi-manual approach, where reviewers and artists manually mark the logo area to determine the logo's location in each video frame. These manually marked logos are then partially blurred, for example, to render the watermark invisible. Currently, the vast majority of online logo restoration tools and video watermark removal services on most video websites utilize this approach.

[0051] However, this method has high labor costs and relies on manual frame-by-frame processing, which is prone to missed detections. In addition, directly adding local blurring or coding to manually labeled logos may affect the viewing experience of the video.

[0052] The following will describe an identification detection method, an identification detection device, an electronic device, a computer-readable storage medium, and a computer program product according to exemplary embodiments of the present disclosure with reference to the accompanying drawings to solve at least one of the above problems.

[0053] Figure 3It is a flowchart of a method for identification detection according to an exemplary embodiment. The execution subject of the identification detection method according to the exemplary embodiment of the present disclosure can execute the identification detection method for the acquired video to be detected to determine the identification area in the video frame of the video to be detected. Here, as an example, the execution subject of the identification detection method according to the exemplary embodiment of the present disclosure can be a personal computer, a tablet device, a personal digital assistant, a smart phone, a server or any combination thereof, but the execution subject of the identification detection method is not limited thereto, and it can also be any other type of electronic device capable of performing identification detection. The present disclosure does not impose any special restrictions on the execution subject of the identification detection method.

[0054] like Figure 3 As shown, the identification detection method according to an exemplary embodiment of the present disclosure may include the following steps:

[0055] In step S10, a target video frame in the video to be detected is obtained.

[0056] In this step, the video to be detected can refer to any type of dynamic image captured, recorded, processed, stored, transmitted and reproduced in the form of electrical signals. It can be a video transmitted through the Internet, such as a video on an Internet video platform; it can also be a video transmitted through other means such as radio, such as a video broadcast by a TV station.

[0057] The target video frame may be any one or more video frames that need to be detected in the video to be detected. For example, it may be all video frames in the video to be detected, or it may be a part of the video frames in the video to be detected.

[0058] As an example, video frames can be extracted from the video to be detected at preset time intervals as target video frames. This example can be applied to the video review application scenario described above, for example.

[0059] In step S20 , a first region and a first candidate region in the target video frame are identified.

[0060] Here, the first region may include a preset logo, and the first candidate region may include logo content other than the preset logo. In this context, a logo may refer to at least one of a graphic logo, a text logo, a watermark, a logo, or a trademark. It can be any type of graphic or text that appears in one or more frames of a video, such as a logo of a video playback platform, an identifier of a video playback user, or a video producer. Furthermore, the regions referred to herein may be represented, for example, by coordinate boxes.

[0061] In this step, the execution order of determining the first region and the first candidate region can be arbitrary. The first region can be determined first and then the first candidate region, or the first candidate region can be determined first and then the first region, or both can be performed simultaneously.

[0062] The preset identification may be an identification in a known identification database, for example, a known station logo, watermark, logo, trademark, etc.

[0063] For determining the first area, as an example, the preset identifier may include at least one of a preset graphic identifier and a preset text identifier. In this example, Figure 4 As shown, the step of determining the first area may include:

[0064] In step S21 , graphic mark detection and text mark detection are performed on the target video frame to obtain a second candidate region containing a preset graphic mark and / or a second candidate region containing a preset text mark.

[0065] In this step, as an example, a pre-trained graphic detection network can be used to perform graphic logo detection on the target video frame to obtain a second candidate region containing a preset graphic logo.

[0066] In this example, the pattern detection network can be used to detect a preset pattern mark in a video frame. The pattern detection network can be, for example, Figure 1 The YOLO structured graphic detection neural network shown is used to detect the category and location of station logo templates in the database.

[0067] As an example, a graphic detection network can be trained based on preset graphic identifiers in a known identifier database. For example, a training sample set can be established based on the known identifier database. The training sample set can include preset graphic identifier samples and labels indicating preset graphic identifier samples. In this way, the graphic detection network can be trained based on the training sample set to obtain a trained graphic detection network to detect whether a preset graphic identifier exists in a target video frame.

[0068] In this step, as an example, a pre-trained first text detection network can be used to perform text marker detection on the target video frame to obtain a second candidate region containing a preset text marker.

[0069] In this example, the first text detection network can be used to detect preset text identifiers in video frames. The first text detection network can be, for example, an optical character recognition (OCR) text recognition network, which is used to detect and recognize text in video images. It can detect certain text type identifiers, for example, and can be any application with video text recognition capabilities. As an example, the OCR text recognition network can use a recursive convolutional neural network algorithm for text classification.

[0070] As an example, the first text detection network can be trained based on preset text identifiers in a known identifier database. For example, a training sample set can be established based on the known identifier database. The training sample set can include preset text identifier samples and labels indicating the preset text identifier samples. In this way, the first text detection network can be trained based on the training sample set to obtain a trained first text detection network to detect whether there is a preset text identifier in the target video frame.

[0071] It should be noted that although examples of a graphic detection network, a first text detection network, and training methods thereof are given above, they are not limited thereto, and corresponding networks can also be obtained based on neural networks of other structures.

[0072] Furthermore, according to exemplary embodiments of the present disclosure, a neural network is used to identify preset graphic and text logos in a target video frame, and an artificial intelligence model can be applied to video logo recognition, which helps improve the accuracy and speed of video frame recognition. However, the method for detecting graphic and text logos in a target video frame is not limited to this method and can also be implemented using other algorithm modules with similar functions.

[0073] It should also be noted that the execution order of graphic logo detection and text logo detection can be arbitrary. Graphic logo detection can be performed first and then text logo detection, or text logo detection can be performed first and then graphic logo detection, or both can be performed at the same time.

[0074] In step S22, the obtained second candidate regions are fused to determine the first region.

[0075] In this step, the second candidate region including the preset graphic mark and the second candidate region including the preset text mark may be fused as a whole to obtain the first region.

[0076] As an example, the step of fusing the obtained second candidate areas and determining the first area may include: merging the second candidate areas whose overlap is greater than a first preset value to obtain a new second candidate area; when the overlap between any two second candidate areas is less than or equal to the first preset value, using the second candidate area as the first area.

[0077] For example, the degree of overlap between each pair of second candidate regions can be determined. When the degree of overlap exceeds a first preset value, the corresponding two second candidate regions are merged, and the merged region is used as the new second candidate region. In this way, the detection results can be rationally merged based on the degree of overlap between regions, avoiding the presence of duplicate regions in the detection results, which would lead to subsequent repeated processing and waste of computing resources, and ensuring faster detection speed.

[0078] Specifically, the second candidate area containing the preset graphic logo and the second candidate area containing the preset text logo both need to be detected as logos, for example, they may be locations that need to be eliminated from the video frame. Therefore, the coordinate frames of all the second candidate areas can be first merged into a set S, and then, for each two coordinate frames of the second candidate areas in the set S, the degree of overlap is calculated, for example, the intersection over union (IoU) index is calculated, that is, the intersection area of ​​the two coordinate frames divided by the union area.

[0079] When the IoU value is greater than the first preset value σ, the positions of the two second candidate regions are considered to overlap, and the range of the union of the two is taken to replace the coordinates of the original two second candidate regions as the determined first region. When the IoU value is less than or equal to the first preset value σ, the positions of the two second candidate regions are considered to be non-overlapping, and the coordinates of the original two second candidate regions remain unchanged and are respectively used as the determined first regions.

[0080] It should be noted that an example of expressing the degree of overlap based on the IoU indicator is given here. However, it is not limited to this and can also be replaced by other overlap calculation indicators.

[0081] It should also be noted that the above describes the fusion of the second candidate regions based on the degree of overlap between each two second candidate regions to obtain the first region. However, the exemplary embodiments of the present disclosure are not limited to this, and the first region can also be determined by fusion in other ways. For example, the overlapping area between each two second candidate regions can be removed, and the deduplicated second candidate region can be used as the first region.

[0082] In the above, combined with Figure 4An example is described of first determining a second candidate area and then fusing the second candidate area to determine the first area. In this way, the area containing the preset graphic logo and the area containing the preset text logo can be detected separately, and then the results of the separate detections can be fused. In this way, on the one hand, the differences between graphic detection and text detection can be taken into account, and the detection tasks can be performed separately to improve the accuracy of the detection results; on the other hand, the detection results of the separate detections can be reasonably fused to improve the accuracy of the detection results while avoiding the subsequent repeated processing caused by the presence of repeated areas in the detection results of the two aspects, wasting computing resources, ensuring a fast detection speed, and being able to complete the identification and positioning of multiple logos and their variants within an acceptable time, so as to ultimately achieve the elimination of the logos.

[0083] However, exemplary embodiments of the present disclosure are not limited to Figure 4 In the example shown, the first area containing the preset logo can also be determined by other methods. For example, the first area containing the preset logo can be directly detected without distinguishing between the preset graphic logo and the preset text logo. For example, a neural network can also be used to implement it.

[0084] In step S20, for determining the first candidate area, as an example, the identification content other than the preset identification may include text. In this example, the step of determining the first candidate area may include: performing text detection on the target video frame to obtain the first candidate area containing text.

[0085] As an example, a pre-trained second text detection network can be used to perform text detection on a target video frame to obtain a first candidate region containing text, wherein the second text detection network is used to detect text in the video frame.

[0086] In this example, the second text detection network can be used to detect text in the video frame. The second text detection network can be, for example, any text recognition neural network for detecting text in the video frame. Here, the second text detection network can detect text in the video frame that is not included in the known identifier database. In addition, as an example, the second text detection network can also detect text identifiers detected by the first text detection network.

[0087] The second text detection network can be, for example, a neural network using DBNet, and is used to detect text-like regions appearing in the video. For example, the second text detection network can be trained using any text training sample set, where the training sample set can include text samples and labels indicating the text samples. Thus, the second text detection network can be trained based on the training sample set to obtain a trained second text detection network for detecting whether text exists in the target video frame.

[0088] It should be noted that, although an example of the second text detection network and its training method is given above, it is not limited thereto, and a corresponding network can also be obtained based on neural networks of other structures.

[0089] Furthermore, according to exemplary embodiments of the present disclosure, text in a target video frame can be identified based on a neural network, and an artificial intelligence model can be applied to the recognition of video identifiers, which is beneficial for improving the accuracy and speed of video frame recognition. However, the method for identifying text in a target video frame is not limited to this, and other algorithm modules with similar functions can also be used to implement it.

[0090] Furthermore, while the above description describes an example in which the identifier content other than a preset identifier is text, the present disclosure is not limited thereto. The identifier content may also be any other type of identifier, for example, a graphic identifier. In this case, a pre-trained second graphic detection network may be used to detect the target video frame and obtain a first candidate region containing the graphic. Here, the second graphic detection network may detect graphics that are not included in the known identifier database. Furthermore, as an example, the second graphic detection network may also detect text identifiers detected by the aforementioned graphic detection network.

[0091] According to an exemplary embodiment of the present disclosure, in step S20, by determining the first candidate area including other than the preset mark, it is possible to determine, for example, Figure 2A and Figure 2B The text lines and temporary logos shown in the figure are used to avoid the situation where only referring to the known logo database for detection results in the inability to flexibly cope with the newly appeared, customized text or graphic logos.

[0092] It should also be noted that the execution order of determining the second candidate area containing the preset graphic logo and / or determining the second candidate area containing the preset text logo and determining the first candidate area can be arbitrary. The three can be executed successively in any order, and any two or three of the three can also be executed simultaneously.

[0093] In step S30 , a first candidate region whose distance from the first region is less than a preset value among the first candidate regions is determined as a second region.

[0094] Since the first candidate area contains identification content other than the preset identification, the first candidate area may contain both the identification and the video content. In step S30, the first candidate area containing the identification (i.e., the second area) can be screened out based on the distance between the first candidate area and the first area, while the first candidate area containing the video content is excluded. Here, the distance between the first candidate area and the first area can be determined in different ways, the field type of the preset value can be given according to the method for determining the distance between the two, and the size of the preset value can be arbitrarily set according to actual needs. An example of determining the second area will be given below.

[0095] In one example, the distance between the first candidate region and the first region can be determined by enlarging the first region and calculating the degree of overlap between the enlarged first region and the first candidate region. In this example, the preset value mentioned above can be a preset degree of overlap.

[0096] For example, the first region may be expanded to obtain an expanded first region; the degree of overlap between the expanded first region and the first candidate region is determined, and the first candidate region whose degree of overlap is greater than a second preset value is determined as the second region. Here, the direction of expansion of the first region may be arbitrarily specified. For example, the first region may be expanded in the height and width directions of the target video frame with the center point of the first region as a reference point.

[0097] Specifically, for the set S' obtained by optimizing the set S in step S23, the coordinate frame of each first region can be taken and compared with the coordinate frame of each first candidate region. If the two are in adjacent positions, the coordinate frame of the first candidate region is also retained as the second region. For example, the coordinate frame in the set S' (i.e., the first region) can be expanded by a preset multiple, and the IoU between the coordinate frame and each first candidate region can be calculated. If the IoU is greater than a second preset value, the corresponding first candidate region is determined as the second region.

[0098] In this example, the second area can be screened out from the first candidate area by expanding the first area, thereby excluding the first candidate area that does not belong to the video identification, improving the accuracy of the detection result, and improving the way of expanding the first area to determine the degree of overlap. The second area can be screened according to the size of the first area itself to avoid missing the second area, further improving the accuracy of the detection result.

[0099] In another example, a first candidate region that meets a preset distance condition can be determined as the second region, where the preset distance condition includes: a distance from the first region in the height direction of the target video frame being less than a first preset distance; and / or a distance from the first region in the width direction of the target video frame being less than a second preset distance. In this example, the preset value mentioned above can be the first preset distance and / or the second preset distance.

[0100] In this example, the second area can be screened out from the first candidate area based on the distance from the first area in two directions, thereby excluding the first candidate area that does not belong to the video identification, improving the accuracy of the detection result, and selecting two directions for calculation can improve the calculation speed while ensuring accuracy.

[0101] In the above example, the preset multiple, the first preset distance, and the second preset distance can be arbitrarily set according to actual needs. In addition, the method of determining the second area is not limited to the above example, and other methods can also be used to determine the second area that may be the logo in the video from the first candidate area.

[0102] In step S40 , the first region and the second region are determined as identification regions in the target video frame.

[0103] In this step, the first region and the second region can be used together as the identification region detected from the target video frame. In one example, the first region and the second region can be saved as the identification region detected from the target video frame. In another example, duplicate regions in the first region and the second region can be removed, and the deduplicated first region and the second region can be saved as the identification region in the target video frame.

[0104] In this way, according to the embodiment of the present disclosure, by determining the first area and the second area as the identifiers detected from the video frame, the problem of not being able to detect identifiers outside the range of a known identifier data set is solved. Known identifiers and other identifiers outside the range of known identifiers can be detected, thereby improving the comprehensiveness and accuracy of detection and avoiding missed detections.

[0105] In addition, according to an exemplary embodiment of the present disclosure, as in step S10, in one case, the steps described above can be applied to all video frames in the video to be detected to determine the identification area in each video frame, that is, the target video frame can be all video frames in the video to be detected; in another case, a part of the video frames in the video to be detected can be extracted to perform the above-mentioned detection. For this case, on the one hand, the detection results of the extracted video frames can be used for video review, and on the other hand, the identification areas of other video frames in the video can also be determined based on the detection results of these extracted video frames. This aspect will be described in detail below.

[0106] Specifically, when video frames are extracted from the video to be detected as target video frames at preset time intervals, the detection method according to an exemplary embodiment of the present disclosure may also include: determining the regional similarity between the current target video frame and the intermediate video frame in the target area for the intermediate video frame between the current target video frame and the adjacent target video frames before and / or after the current target video frame, wherein the target area is the first area and the second area corresponding to the current target video frame.

[0107] Here, the preset time interval may be, for example, 2 seconds, but is not limited thereto and may be adjusted according to actual needs. The preset condition may be: a ratio of the area of ​​the similar region in the middle video frame to the total area of ​​the target region is greater than a ratio threshold, wherein the similar region refers to a region whose region similarity is greater than the similarity threshold.

[0108] When the intermediate video frame meets the preset conditions, it can be considered that the intermediate video frame and the target video frame have the same identification area, and therefore, the target area can be determined as the identification area in the intermediate video frame; when the intermediate video frame does not meet the preset conditions, it can be considered that the intermediate video frame does not have the identification area of ​​the target video frame, and therefore, the above-mentioned detection steps S10 to S40 can be performed on the intermediate video frame.

[0109] For example, based on the first and second regions of each sampled target video frame, all intermediate video frames before and after the current target video frame can be traversed until the previous or next sampled target video frame is encountered. For all intermediate video frames within the traversed range, the start and end times of the appearance of the markers in the first and second regions are determined.

[0110] For example, assuming that the video to be detected has 100 video frames, the 10th, 20th, 30th, 40th, 50th, 60th, 70th, 80th, 90th, and 100th frames are selected from these 100 video frames as target video frames. Here, the 10th and 20th frames are adjacent target video frames, and the video frames between them are intermediate video frames, that is, the 11th to 19th frames are intermediate video frames. Similarly, the 20th and 30th frames are adjacent target video frames, that is, the 21st to 29th frames are intermediate video frames. And so on, other adjacent target video frames and the intermediate video frames between them can be determined.

[0111] Taking the processing of the 20th frame as an example, the first area and the second area in the 20th frame (which is the target video frame) may be determined first, and the determined first area and second area may be used as the target area.

[0112] Here, the adjacent target video frame before the 20th frame is the 10th frame, and the adjacent target video frame after the 20th frame is the 30th frame. For the intermediate video frames between the 20th frame and the 10th frame (i.e., the 11th to the 19th frames) and / or the intermediate video frames between the 20th frame and the 30th frame (i.e., the 21st to the 29th frames), the regional similarity between the 20th frame and each intermediate video frame in the target region can be determined.

[0113] When the intermediate video frame meets the preset conditions described above, it can be considered that the intermediate video frame has the same identification area as the 20th frame. For example, it can be determined that frames 11 to 19 and frames 21 to 25 all meet the above preset conditions. Therefore, it can be considered that frames 11 to 19 and frames 21 to 25 all have the same identification area as frame 20, while frames 26 to 29 do not meet the above preset conditions. It can be considered that frames 26 to 29 have different identification areas from frame 20. In this way, it can be determined that the starting frame of the identification area in frame 20 is frame 11 and the ending frame is frame 25, so that the start and end time of the identification can be determined.

[0114] Similar to the processing performed on the 20th frame above, the above process can be performed on all target video frames. It should be noted that, when it is determined that the above preset conditions are met between the 11th to 19th frames and the 20th frame, when processing the 10th frame (which is the target video frame), it may be determined that the above preset conditions are also met between the 11th to 19th frames and the 10th frame. In this case, it can be considered that the 11th to 19th frames and the 10th frame also have the same identification area, that is, the 11th to 19th frames as the intermediate video frames and the 10th and 20th frames as the target video frames all have the same identification area. Therefore, the identification area of ​​the 10th frame and the identification area of ​​the 20th frame can be determined as the identification area of ​​the 11th to 19th frames.

[0115] The regional similarity described above can be determined by calculating pixel similarity, and the similar region can be determined based on the pixel similarity in the region. Specifically, for each intermediate video frame within the traversal range, the portion within the coordinate frame of the target region of the intermediate video frame can be taken to calculate the pixel similarity with the sampled target video frame. The pixel similarity can be, for example, a mean squared error (MSE) indicator. When the ratio of the number of pixels with pixel similarity higher than the similarity threshold δ to the total number of pixels in the target region is higher than a ratio threshold γ, or when the ratio of the pixel area with pixel similarity higher than the similarity threshold δ to the total pixel area of ​​the target region is higher than a ratio threshold γ, then the intermediate video frame is considered to contain the same identifier as the target video frame. Here, the ratio threshold γ and the similarity threshold δ can be arbitrarily set according to actual needs, for example, they can be determined according to the accuracy requirements of the detection results of the actual detection. When the ratio threshold γ and the similarity threshold δ are higher, the detection result is more accurate.

[0116] In the above, the start time of the identifier's appearance corresponds to the video frame in which the identifier first appears in the time sequence, and the end time of the identifier's appearance corresponds to the video frame in which the identifier last appears in the time sequence. The video frame in which the identifier first appears can be the first video frame among the video frames containing the same identifier as the target video frame, and the video frame in which the identifier last appears can be the last video frame among the video frames containing the same identifier as the target video frame. Here, the first video frame can be an intermediate video frame before the target video frame, or the target video frame itself; the last video frame can be an intermediate video frame after the target video frame, or the target video frame itself.

[0117] Thus, according to the above method, the start and end times of the appearance of the markers in the first and second areas of the target video frame can be obtained, so as to facilitate subsequent batch processing of the video frames, such as batch erasing the markers in the video frames within the start and end time.

[0118] It should be noted that an example of expressing similarity based on the MSE indicator is given here, however, it is not limited to this and can also be replaced by other similarity calculation indicators.

[0119] As described above, according to the exemplary embodiments of the present disclosure, the "time domain detection" scheme of first extracting the target video frame for detection and then detecting the intermediate video frame is adopted. This can obtain frame-by-frame detection results while maintaining efficient detection, solving the problem in related technologies that cannot simultaneously guarantee detection speed and detection effect.

[0120] The above describes an identification detection method according to an exemplary embodiment of the present disclosure. The method involves technologies in the fields of computer vision, image and video processing, artificial intelligence, etc. The types of algorithms involved include image object detection (image detection), image segmentation (image segmentation), etc. Those skilled in the art can learn the specific calculation process of the relevant algorithms from the knowledge in the corresponding fields, which will not be described in detail here.

[0121] Refer to above Figure 3 and Figure 4 Described is a method of detecting an identification mark according to an exemplary embodiment of the present disclosure, Figure 5 A schematic diagram of a framework of an example of a method for detecting a marker according to an exemplary embodiment is shown below. Figure 5 An implementation example of the method is described.

[0122] like Figure 5 As shown, the target video frame can be input into the graphic detection network, the first text detection network and the second text detection network respectively. Among them, the graphic detection network is mainly for icon-type logo recognition, the first text detection network is mainly for text-type logo recognition, and the second text detection network is mainly used to detect all text-like areas in the entire image.

[0123] Specifically, on the one hand, a graphic detection network can be used to detect the graphic template category and its location in the database. On the other hand, a first text detection network can be used, for example, a recursive convolutional neural network algorithm for text classification can be used to perform OCR text recognition to detect the text in the video frame and perform text recognition on it. In addition, a second text detection network can be used, for example, a DBNet network can be used to detect the text area in the video frame to detect the text-like area appearing in the video frame. The above three network modules all return a series of coordinate boxes, marking the corresponding positions, that is, the second candidate area output by the graphic detection network and the first text detection network, and the first candidate area output by the second text detection network.

[0124] Specifically, the second candidate region may be fused to obtain the first region, and the first candidate region may be screened based on the first region to obtain the second region. In this way, the first region and the second region may be determined as the identification region in the target video frame.

[0125] When the target video frame is a sampled video frame, the time domain detection scheme described above can be adopted to determine the identification area in the intermediate video frame based on the detection result of the target video frame to complete the identification detection of the entire video to be detected, and the detected identification can be sent to the subsequent identification processing link.

[0126] In addition, according to the detection method of the exemplary embodiment of the present disclosure, the logo in the video can be automatically detected, and since the detection results of the target video frame and the intermediate video frame are automatically generated, the detection results can be sent to the subsequent video repair, logo erasure, etc., so as to be automatically processed, avoiding the problems of excessive repair, poor visual effects, etc. caused by the abrupt boundaries of the manually marked logo areas.

[0127] Figure 6 FIG. 1 is a block diagram of a data query device according to an exemplary embodiment. Figure 6 The device includes an acquiring unit 100, a first determining unit 200, a second determining unit 300 and a third determining unit 400.

[0128] The acquisition unit 100 is configured to acquire a target video frame in a video to be detected.

[0129] The first determining unit 200 is configured to identify a first region and a first candidate region in a target video frame, wherein the first region includes a preset identifier, and the first candidate region includes identifier content other than the preset identifier.

[0130] The second determining unit 300 is configured to determine, as the second region, a first candidate region whose distance from the first region is less than a preset value among the first candidate regions.

[0131] The third determining unit 400 is configured to determine the first region and the second region as identification regions in the target video frame.

[0132] As an example, the preset identifier includes at least one of a preset graphic identifier and a preset text identifier, wherein the first determination unit 200 is further configured to: perform graphic identifier detection and text identifier detection on the target video frame to obtain a second candidate area containing a preset graphic identifier and / or a second candidate area containing a preset text identifier; and fuse the obtained second candidate areas to determine the first area.

[0133] As an example, the first determination unit 200 is further configured to: merge second candidate areas whose overlap degree is greater than a first preset value to obtain a new second candidate area; when the overlap degree of any two second candidate areas is less than or equal to the first preset value, use the second candidate area as the first area.

[0134] As an example, the first determination unit 200 is also configured to: determine the degree of overlap between each two second candidate areas, and when the degree of overlap is greater than a first preset value, merge the corresponding two second candidate areas and use the merged area as the first area; when the degree of overlap is less than or equal to the first preset value, use the corresponding two second candidate areas as the first areas respectively.

[0135] As an example, the identification content other than the preset identification includes text, wherein the first determining unit 200 is further configured to: perform text detection on the target video frame to obtain a first candidate region containing text.

[0136] As an example, the second determination unit 300 is further configured to: expand the first area to obtain an expanded first area; determine the degree of overlap between the expanded first area and the first candidate area, and determine the first candidate area whose overlap degree is greater than a second preset value as the second area.

[0137] As an example, the second determination unit 300 is further configured to: determine the first candidate area that meets the preset distance condition as the second area, wherein the preset distance condition includes: the distance between the target video frame and the first area in the height direction is less than the first preset distance; and / or, the distance between the target video frame and the first area in the width direction is less than the second preset distance.

[0138] As an example, the acquisition unit 100 is also configured to: extract video frames from the video to be detected as target video frames at preset time intervals, wherein the identification detection device also includes a fourth determination unit, and the fourth determination unit is configured to: determine the regional similarity between the current target video frame and the intermediate video frame in the target area for the intermediate video frame between the current target video frame and the adjacent target video frames before and / or after the current target video frame, wherein the target area is the first area and the second area corresponding to the current target video frame; when the intermediate video frame meets the preset conditions, the target area is determined as the identification area in the intermediate video frame, wherein the preset conditions are: the proportion of the area of ​​the similar area in the intermediate video frame to the total area of ​​the target area is higher than the proportion threshold, wherein the similar area refers to the area whose regional similarity is higher than the similarity threshold.

[0139] Regarding the apparatus in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.

[0140] Figure 7 FIG. 1 is a block diagram of an electronic device according to an exemplary embodiment. Figure 7As shown, the electronic device 10 includes a processor 101 and a memory 102 for storing processor-executable instructions. Here, when the processor-executable instructions are executed by the processor, the processor is prompted to execute the identification detection method as described in the above exemplary embodiment.

[0141] As an example, the electronic device 10 does not necessarily need to be a single device, but may also be any collection of devices or circuits that can execute the above instructions (or instruction sets) individually or in combination. The electronic device 10 may also be part of an integrated control system or system manager, or may be configured as a server that is interconnected with a local or remote (e.g., via wireless transmission) interface.

[0142] In the electronic device 10, the processor 101 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 101 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0143] The processor 101 may execute instructions or codes stored in the memory 102, which may also store data. Instructions and data may also be sent and received over a network via a network interface device, which may employ any known transmission protocol.

[0144] Memory 102 may be integrated with processor 101, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. Furthermore, memory 102 may comprise a separate device, such as an external disk drive, a storage array, or any other storage device usable by a database system. Memory 102 and processor 101 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, or the like, such that processor 101 can access files stored in memory 102.

[0145] In addition, the electronic device 10 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 10 may be connected to each other via a bus and / or a network.

[0146] In an exemplary embodiment, a computer-readable storage medium may also be provided, and when the instructions in the computer-readable storage medium are executed by a processor, the processor is enabled to perform the identification detection method as described in the above exemplary embodiment. The computer-readable storage medium may be, for example, a memory including instructions. Alternatively, the computer-readable storage medium may be: read-only memory (ROM), random access memory (RAM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or ultra-fast digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device configured to store the computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.

[0147] In an exemplary embodiment, a computer program product may also be provided. The computer program product includes computer instructions. When the computer instructions are executed by a processor, the identification detection method as described in the exemplary embodiment above is implemented.

[0148] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0149] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for detecting a mark, characterized in that: The identification detection method comprises: Obtain the target video frame in the video to be detected; Identifying a first region and a first candidate region in the target video frame, wherein the first region includes a preset identifier, which is a known identifier, and the first candidate region includes identifier content other than the preset identifier; Determine a first candidate region in the first candidate regions whose distance from the first region is less than a preset value as a second region, wherein the second region includes an identifier other than the known identifier; determining the first area and the second area as identification areas in the target video frame, The step of determining a first candidate region in the first candidate region whose distance from the first region is less than a preset value as the second region includes: expanding the first region to obtain an expanded first region; determining the degree of overlap between the expanded first region and the first candidate region, and determining a first candidate region whose degree of overlap is greater than a second preset value as the second region; or Among them, the step of determining the first candidate area in the first candidate area whose distance from the first area is less than a preset value as the second area includes: determining the first candidate area that meets the preset distance condition as the second area, wherein the preset distance condition includes: the distance between the first area and the target video frame in the height direction is less than a first preset distance; and / or the distance between the first area and the target video frame in the width direction is less than a second preset distance.

2. The identification detection method according to claim 1, characterized in that: The preset identifier includes at least one of a preset graphic identifier and a preset text identifier, wherein the step of determining the first area includes: Performing graphic mark detection and text mark detection on the target video frame to obtain a second candidate region containing the preset graphic mark and / or a second candidate region containing the preset text mark; The obtained second candidate regions are fused to determine the first region.

3. The identification detection method according to claim 2, characterized in that: The fusing the obtained second candidate regions to determine the first region includes: Merge the second candidate regions whose overlap degree is greater than the first preset value to obtain a new second candidate region; When the degree of overlap between any two second candidate regions is less than or equal to the first preset value, the second candidate region is used as the first region.

4. The identification detection method according to claim 3, characterized in that: The step of merging the second candidate regions whose overlap degree is greater than a first preset value to obtain a new second candidate region includes: The degree of overlap between each two second candidate regions is determined, and when the degree of overlap is greater than a first preset value, the corresponding two second candidate regions are merged, and the merged region is used as a new second candidate region.

5. The identification detection method according to any one of claims 1 to 4, characterized in that: The identification content other than the preset identification includes text, wherein the step of determining the first candidate area includes: Perform text detection on the target video frame to obtain a first candidate region containing text.

6. The identification detection method according to claim 1, characterized in that: The step of obtaining a target video frame in the video to be detected includes: extracting a video frame from the video to be detected as the target video frame at a preset time interval, The identification detection method further includes: determining, for a current target video frame and an intermediate video frame between adjacent target video frames before and / or after the current target video frame, a regional similarity between the current target video frame and the intermediate video frame in a target region, wherein the target region is a first region and a second region corresponding to the current target video frame; When the intermediate video frame meets the preset condition, the target area is determined as the marked area in the intermediate video frame. The preset condition is that the ratio of the area of ​​the similar region in the middle video frame to the total area of ​​the target region is higher than a ratio threshold, wherein the similar region refers to a region whose region similarity is higher than a similarity threshold.

7. A marking detection device, characterized in that: The identification detection device comprises: An acquisition unit is configured to acquire a target video frame in a video to be detected; a first determining unit configured to identify a first region and a first candidate region in the target video frame, wherein the first region includes a preset identifier, which is a known identifier, and the first candidate region includes identifier content other than the preset identifier; a second determining unit configured to determine a first candidate region in the first candidate regions, the first region of which the distance therefrom is less than a preset value, as a second region, wherein the second region includes an identifier other than the known identifier; A third determining unit is configured to determine the first area and the second area as identification areas in the target video frame The second determining unit is further configured to: expand the first area to obtain an expanded first area; determine the degree of overlap between the expanded first area and the first candidate area, and determine the first candidate area whose degree of overlap is greater than a second preset value as the second area; or Wherein, the second determination unit is further configured to: determine the first candidate area that meets the preset distance condition as the second area, wherein the preset distance condition includes: the distance between the first area and the target video frame in the height direction is less than a first preset distance; and / or the distance between the first area and the target video frame in the width direction is less than a second preset distance.

8. The identification detection device according to claim 7, characterized in that: The preset identifier includes at least one of a preset graphic identifier and a preset text identifier, wherein the first determination unit is further configured to: perform graphic identifier detection and text identifier detection on the target video frame to obtain a second candidate area containing the preset graphic identifier and / or a second candidate area containing the preset text identifier; and fuse the obtained second candidate areas to determine the first area.

9. The identification detection device according to claim 8, characterized in that: The first determination unit is further configured to: merge second candidate areas whose overlap degree is greater than a first preset value to obtain a new second candidate area; and use the second candidate area as the first area when the overlap degree of any two second candidate areas is less than or equal to the first preset value.

10. The identification detection device according to claim 9, characterized in that: The first determining unit is further configured to: determine the degree of overlap between each two second candidate regions, and when the degree of overlap is greater than a first preset value, merge the corresponding two second candidate regions and use the merged region as the first region; When the degree of overlap is less than or equal to a first preset value, the corresponding two second candidate regions are respectively used as first regions.

11. The identification detection device according to any one of claims 7 to 10, characterized in that: The identification content other than the preset identification includes text, wherein the first determining unit is further configured to: perform text detection on the target video frame to obtain a first candidate region containing text.

12. The identification detection device according to claim 7, characterized in that: The acquisition unit is further configured to extract video frames from the video to be detected as the target video frames at preset time intervals. The identification detection device further includes a fourth determining unit, and the fourth determining unit is configured to: determining, for a current target video frame and an intermediate video frame between adjacent target video frames before and / or after the current target video frame, a regional similarity between the current target video frame and the intermediate video frame in a target region, wherein the target region is a first region and a second region corresponding to the current target video frame; When the intermediate video frame meets the preset condition, the target area is determined as the marked area in the intermediate video frame. The preset condition is that the ratio of the area of ​​the similar region in the middle video frame to the total area of ​​the target region is higher than a ratio threshold, wherein the similar region refers to a region whose region similarity is higher than a similarity threshold.

13. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor, Wherein, when the processor-executable instructions are executed by the processor, the processor is prompted to execute the identification detection method according to any one of claims 1 to 6.

14. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor, the processor is enabled to perform the identification detection method according to any one of claims 1 to 6.

15. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the identification detection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method and device for identifying target area, electronic equipment and roadside equipment

    CN111709354A

  • Fusion of bounding regions

    US10133951B1