Text security type detection method, device, equipment, medium, and product

By applying edge enhancement network layer and non-maximum suppression processing on advertising pictures, the problem of difficult to identify sensitive text in advertising pictures in the prior art is solved, and more efficient and accurate text security type detection is achieved.

CN114049319BActive Publication Date: 2025-06-13GUANGZHOU HUADUO NETWORK TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111299408.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-04
Publication Date
2025-06-13
Estimated Expiration
2041-11-04

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify target sensitive words, words, and information from advertising pictures containing a variety of text information, resulting in troubles in e-commerce platforms in text security type detection.

Method used

The edge enhancement network layer is used to perform edge enhancement processing in multi-image directions on advertising images to highlight the edge characteristics of text objects, and through text object detection and non-maximum suppression processing, the candidate boxes and corresponding security types of text objects are identified.

Benefits of technology

It improves the accuracy and efficiency of text security type detection, and can identify sensitive words, words, and information in advertising pictures with higher confidence, which is suitable for application scenarios such as e-commerce platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114049319B_ABST
    Figure CN114049319B_ABST
Patent Text Reader

Abstract

The present application discloses a text security type detection method, its device, equipment, medium, and product. The method includes: obtaining an advertisement picture to be detected; performing edge enhancement processing on the advertisement picture to be detected in multiple image directions by using an edge enhancement network layer to obtain a corresponding plurality of edge-enhanced pictures; performing text object detection on the advertisement picture and the edge-enhanced pictures to determine candidate boxes belonging to text objects therein and confidence information corresponding to each candidate box; performing non-maximum suppression processing on the candidate boxes to eliminate redundant candidate box information and improve the working efficiency of the text security detection system. Subsequently, text recognition is performed on the text objects corresponding to the remaining candidate boxes, and the corresponding security type is determined according to the recognized text. The present application can efficiently screen relevant pictures and text data constituting the above-mentioned model training set for training relevant models, making the model discrimination more accurate and having wide adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image recognition, and particularly to a method for detecting text security types, a corresponding device, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Using an artificial neural network model to detect text security types has become the mainstream technology. For application scenarios such as e-commerce platforms, a large number of advertising images are generated every day. Due to the restrictions of e-commerce platforms on sensitive words, phrases, information, etc. in advertising images, it is generally necessary to detect and identify the target text objects in the advertising images uploaded by users, and then make further processing according to different security types and requirements.

[0003] For application scenarios such as e-commerce platforms, if the advertising images uploaded by merchants for displaying advertising information contain target sensitive words, phrases, information, etc., due to uncertain factors such as different scales and densities of text objects, it is sometimes very difficult for existing text security type detection methods to accurately identify them from advertising images, which will cause great trouble for e-commerce platforms. Relying entirely on manual checking is even more unrealistic.

[0004] Therefore, how to accurately and efficiently identify target sensitive words, phrases, information, etc. from various advertising images to be detected containing multiple text information, and make the recognition results more accurate, has become a technical problem to be solved in this field. Summary of the Invention

[0005] The primary objective of the present application is to solve at least one of the above problems, and to provide a method for detecting text security types, a corresponding device, a computer device, a computer-readable storage medium, and a computer program product.

[0006] To meet the various objectives of the present application, the following technical solutions are adopted:

[0007] A method for detecting text security types includes the following steps:

[0008] Obtain an advertising image to be detected;

[0009] Perform edge enhancement processing on the advertising image in multiple image directions using an edge enhancement network layer to obtain a corresponding plurality of edge-enhanced images;

[0010] Perform text object detection according to the advertising image and the corresponding plurality of edge-enhanced images respectively, and determine the candidate boxes belonging to text objects and the confidence information corresponding to each candidate box;

[0011] Perform text recognition on the image corresponding to the remaining candidate boxes after non-maximum suppression, and determine the corresponding security type according to the recognized text.

[0012] In a further embodiment, obtaining an advertisement picture to be detected includes the following steps:

[0013] Respond to an advertisement release request triggered by a user, and obtain the corresponding submitted advertisement release information, where the advertisement release information includes an advertisement picture;

[0014] Obtain the advertisement picture therein from the advertisement release information.

[0015] In a further embodiment, performing edge enhancement processing on the advertisement picture in multiple image directions by using an edge enhancement network layer to obtain a corresponding plurality of edge-enhanced pictures includes the following steps:

[0016] Based on pixel shifting, process the advertisement picture to obtain a plurality of offset image data;

[0017] Perform difference on the advertisement picture and the offset image data one by one to obtain a plurality of difference image data;

[0018] Perform superposition combination processing on the advertisement picture and any of the difference image data multiple times to obtain a plurality of edge-enhanced pictures.

[0019] In a further embodiment, performing text object detection on the advertisement picture and the corresponding plurality of edge-enhanced pictures respectively to determine candidate boxes belonging to text objects and confidence information corresponding to each candidate box includes the following steps:

[0020] Perform standardized preprocessing on the advertisement picture and the edge-enhanced pictures;

[0021] Perform text object detection on the advertisement picture and the edge-enhanced pictures respectively to obtain candidate boxes belonging to text objects and confidence information corresponding to each candidate box.

[0022] In a further embodiment, performing text recognition on the image corresponding to the remaining candidate boxes after non-maximum suppression, and determining the corresponding security type according to the recognized text includes the following steps:

[0023] Perform non-maximum suppression processing on the candidate boxes and the confidence information corresponding to each candidate box, delete candidate boxes that point to the same text object and have a lower confidence, and obtain the remaining candidate boxes;

[0024] Perform text recognition on the text images corresponding to the remaining candidate boxes respectively to identify text objects;

[0025] Perform security type determination according to the text object and output a determination result.

[0026] In a specific embodiment, non-maximum suppression processing is performed on the candidate boxes and the confidence information corresponding to each candidate box, and the candidate boxes with lower confidence that point to the same text object are deleted, including the following steps:

[0027] The candidate boxes and their confidence information are sorted in reverse order as the initial candidate set, and at the same time, an empty target set is established;

[0028] Remove the candidate box with the highest confidence from the current candidate set and put it into the target set;

[0029] Based on pixels, calculate the intersection over union ratio value between the candidate box and all candidate boxes in the remaining candidate set;

[0030] Compare the intersection over union ratio value with a preset threshold. When the intersection over union ratio value is greater than the preset threshold, delete the corresponding candidate box in the candidate set;

[0031] Repeat the above operations until the candidate set is empty, and the candidate boxes in the target set are the output results.

[0032] A text security type detection device provided to meet one of the purposes of the present application includes an image acquisition module, an edge enhancement module, a text detection module, and a security discrimination module. Among them, the image acquisition module is used to acquire the advertisement picture therein from the advertisement release information; the edge enhancement module is used to perform edge enhancement processing on the advertisement picture in multiple image directions by using an edge enhancement network layer to obtain a corresponding plurality of edge-enhanced pictures; the text detection module is used to perform text object detection according to the advertisement picture and the corresponding plurality of edge-enhanced pictures respectively, and determine the candidate boxes belonging to the text object and the confidence information corresponding to each candidate box; the security discrimination module is used to perform text recognition on the image corresponding to the remaining candidate boxes after non-maximum suppression, and discriminate the corresponding security type according to the recognized text.

[0033] In a further embodiment, the image acquisition module includes: a trigger sub-module that responds to an advertisement release request triggered by a user and acquires the corresponding advertisement release information submitted thereby, and the advertisement release information includes an advertisement picture. An acquisition sub-module acquires the advertisement picture therein from the advertisement release information, which is a picture to be detected for identifying its text security type.

[0034] In a further example, the edge enhancement module includes: a bias sub-module configured to process the advertisement picture based on pixel shift to obtain a plurality of offset image data; a difference sub-module that performs difference on the advertisement picture and the offset image data one by one to obtain a plurality of difference image data; an enhancement sub-module that performs superposition combination processing on the advertisement picture and any of the difference image data multiple times to obtain a plurality of edge-enhanced pictures.

[0035] In the enhanced example, the text detection module includes: a preprocessing sub-module for performing standardized preprocessing on the advertisement picture and the edge-enhanced picture; a detection sub-module for respectively performing text object detection on the advertisement picture and the edge-enhanced picture to obtain candidate boxes belonging to text objects and confidence information corresponding to each candidate box.

[0036] In the enhanced example, the security discrimination module includes: a non-maximum suppression sub-module for performing non-maximum suppression processing on the candidate boxes and the confidence information corresponding to each candidate box, deleting candidate boxes that point to the same text object and have a lower confidence, and obtaining remaining candidate boxes; a text recognition sub-module for respectively performing text recognition on the text images corresponding to the remaining candidate boxes to recognize text objects; a determination sub-module for performing security type determination based on the text objects and outputting a determination result.

[0037] In the specific example, the non-maximum suppression sub-module includes: a set initialization unit for taking the candidate boxes and arranging them in reverse order according to their confidence information as an initial candidate set, and simultaneously establishing an empty target set; a preference unit for removing the candidate box with the highest confidence from the current candidate set and putting it into the target set; a calculation unit for calculating the intersection-over-union ratio value between the candidate box and all candidate boxes in the remaining candidate set based on pixels; a processing unit for comparing the intersection-over-union ratio value with a preset threshold, and when the intersection-over-union ratio value is greater than the preset threshold, deleting the corresponding candidate box in the candidate set; an iteration unit for repeating the above operations until the candidate set is empty, and the candidate boxes in the target set are the output results.

[0038] A computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory. The central processing unit is used to call and run a computer program stored in the memory to execute the steps of the text security type detection method described in the present application.

[0039] A computer-readable storage medium provided to meet another purpose of the present application stores a computer program implemented based on the text security type detection method in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the method.

[0040] A computer program product provided to meet another purpose of the present application includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps of the method described in any embodiment of the present application.

[0041] Compared with the prior art, the advantages of the present application are as follows:

[0042] This application uses an edge enhancement network layer to perform edge enhancement processing on the advertisement image to be detected in multiple image directions, highlighting the edge features of text objects, improving the detection effect, and obtaining corresponding multiple edge-enhanced images; performing text object detection on the advertisement image and the edge-enhanced images, determining the candidate boxes belonging to text objects and the confidence information corresponding to each candidate box; performing non-maximum suppression processing on the candidate boxes to eliminate redundant candidate box information and improve the working efficiency of the text security detection system, and then performing text recognition on the text objects corresponding to the remaining candidate boxes, and discriminating the corresponding security type according to the recognized text. If it is non-secure text, relevant users who publish the advertisement image can be alerted to carry out subsequent rectification work.

[0043] This application constructs an edge enhancement network layer. The pixel values in the edge-enhanced image can break through the conventional pixel value range, thereby improving the contrast of the pixels in the image, making the pixel edges more obvious, which is extremely beneficial to text detection work; at the same time, it provides richer edge features for subsequent text detection work, making its text detection ability stronger.

[0044] In addition, different from the traditional method of directly performing model training and inference on the original detection image, this application first proposes to perform model training and inference on the edge-enhanced image. The edge-enhanced image can enrich the edge features of text objects, which is beneficial to detecting them from the background of the original image to be detected, improving the detection accuracy and working efficiency of the text security type detection system.

[0045] In summary, the judgment result of whether the image to be detected contains sensitive words, phrases, information, etc. in this application has a higher confidence level and can be highly trusted, and is suitable for detecting target texts such as whether there are sensitive words, phrases, information, etc. in advertisement images in application scenarios such as e-commerce platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The above and / or additional aspects and advantages of this application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0047] Figure 1 is a schematic flowchart of a typical embodiment of the text security type detection method of this application;

[0048] Figure 2 is a schematic flowchart of constructing multiple edge-enhanced images in the embodiment of this application;

[0049] Figure 3 is a schematic flowchart of text object recognition and security type detection in the embodiment of this application;

[0050] Figure 4 is a schematic flowchart of non-maximum suppression processing in the specific embodiment of this application;

[0051] Figure 5 It is a principle block diagram of the text security type detection device of the present application;

[0052] Figure 6 It is a structural schematic diagram of a computer device adopted by the present application. Specific embodiments

[0053] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application.

[0054] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0055] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0056] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with wireless signal receivers that only have the ability to receive and no ability to transmit, and devices with both receiving and transmitting hardware that can conduct two-way communication on a two-way communication link. Such devices can include: cellular or other communication devices such as personal computers, tablet computers, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which can include a radio frequency receiver, a pager, Internet / intranet access, a web browser, a notepad, a calendar, and / or a GPS (Global Positioning System) receiver; conventional laptop and / or palm-top computers or other devices, which are conventional laptop and / or palm-top computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed manner at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, for example, it can be a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or it can also be a smart TV, a set-top box, etc.

[0057] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0058] It should be noted that the concept of "server" in this application can similarly be extended to the case applicable to a server cluster. According to the network deployment principle understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be invoked through an interface, or integrated into a physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by this in the implementation manner of the network deployment method of this application.

[0059] One or several technical features of this application, unless explicitly specified, can either be deployed on the server for implementation and accessed by the client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for implementation and access.

[0060] The neural network model cited or possibly cited in this application, unless explicitly specified, can either be deployed on a remote server and remotely invoked on the client, or deployed on a client capable of handling the device and directly invoked. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid over-occupying the client's hardware operating resources.

[0061] All kinds of data involved in this application, unless explicitly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being invoked by the technical solution of this application.

[0062] Those skilled in the art should be aware of this: Although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can all be executed independently. Similarly, for each embodiment disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.

[0063] For each embodiment to be disclosed in this application, unless explicitly pointed out that there is a mutually exclusive relationship between them, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0064] A text security type detection method of this application can be programmed into a computer program product and deployed to run on the client or the server. Thereby, by accessing the interface opened after the computer program product runs, human-computer interaction can be performed with the process of the computer program product through the graphical user interface to execute this method.

[0065] Please refer to Figure 1 , in the typical embodiment of the text security type detection method of the present application, it includes the following steps:

[0066] Step S1100: Obtain the advertisement picture to be detected;

[0067] The advertisement picture to be detected generally refers to an advertisement picture containing text objects. In an exemplary application scenario for assisting in the description of the present application, the advertisement picture to be detected can be an advertisement picture containing text objects on an e-commerce platform. The text objects shown in the picture are generally non-sensitive texts. The implementation of the present application is to identify non-security type texts, that is, pre-set text objects with sensitive information, from the advertisement picture. Therefore, it is necessary to identify the corresponding text objects in the advertisement picture and perform security type discrimination.

[0068] In the application scenario of the e-commerce platform, if it is necessary to obtain the picture to be detected, one implementation method is to receive the input of the e-commerce platform user, especially the input when the merchant instance user configures the advertisement information, and use the advertisement picture in the advertisement information as the picture to be detected; in another method, the server of the e-commerce platform can batch process the advertisement pictures in the e-commerce platform database in the background and use these advertisement pictures as the pictures to be detected for text security type detection.

[0069] Step S1200: Perform edge enhancement processing on the advertisement picture in multiple image directions by using an edge enhancement network layer to obtain corresponding multiple edge-enhanced pictures:

[0070] The edge enhancement network layer is divided into three layers. The first layer is the pixel shift processing layer. After the advertisement picture is processed by pixel eight-neighborhood shift, offset image data in multiple image directions is obtained, that is, the offset image data is the result of the overall movement of the advertisement picture; the second layer is the difference processing layer. The advertisement picture is subtracted from the offset image data in each image direction respectively, that is, difference image data in each image direction is obtained; the third layer is the enhancement processing layer. The advertisement picture is superimposed and combined with any N (N is not greater than 8) difference image data in the second layer to obtain the corresponding multiple edge-enhanced pictures. Wherein N can be set by relevant technical personnel according to experimental results comparison and prior knowledge application.

[0071] Step S1300: Perform text object detection according to the advertisement picture and the corresponding multiple edge-enhanced pictures respectively, and determine the candidate boxes belonging to text objects and the confidence information corresponding to each candidate box:

[0072] Perform text object detection on the advertisement picture using a pre-trained text detection model to obtain the candidate boxes of text objects in the advertisement picture and the confidence information corresponding to each candidate box; perform text object detection on the corresponding multiple edge-enhanced pictures, that is, the image data generated by the edge enhancement network layer, to obtain the candidate boxes of text objects in each edge-enhanced picture and the confidence information corresponding to each candidate box.

[0073] For the text detection model, various relatively excellent text detection models in existing technologies can be selected, including but not limited to: DBNet, PSENet, PANNet, EAST, CTPN, SegLink, PixelLink, TextBoxes, etc. They are all mature text detection models. As long as they are trained with sufficient corresponding training samples until convergence, they can be used as the text detection model of this application. In this application, it is recommended to use the DBNet model.

[0074] The candidate box refers to the position information of the text object in the picture to be detected, and the confidence information corresponding to each candidate box refers to the probability that the candidate box indicates correctly. A high confidence indicates a high probability that the candidate box indicates correctly, and vice versa.

[0075] Step S1400: Perform text recognition on the image corresponding to the remaining candidate boxes after non-maximum suppression, and discriminate the corresponding security type according to the recognized text:

[0076] During the process of text object detection in step S1300, a large number of candidate boxes will be generated at the position of the same target text object, and these candidate boxes overlap with each other; therefore, the candidate boxes need to be subjected to non-maximum suppression to eliminate redundant candidate boxes and obtain the remaining candidate boxes. Use a pre-trained text recognition model to perform text recognition on the image corresponding to the remaining candidate boxes; perform further security type discrimination according to the text recognition result. The discrimination method can be word library matching, text discrimination model, etc. In the embodiments of this application, a pre-trained text discrimination model can be used to perform security type discrimination on the recognized text to discriminate the corresponding security type.

[0077] For the text recognition model, various relatively excellent text recognition models in existing technologies can be selected, including but not limited to: CRNN, RARE, ASTER, DAN, SAR, etc. They are all mature text recognition models. As long as they are trained with sufficient corresponding training samples until convergence, they can be used as the text recognition model of this application. In this application, it is recommended to use the CRNN model.

[0078] For the text discrimination model, various excellent text recognition models in the prior art can be selected, including but not limited to: Fasttext, TextCNN, DPCNN, TextRCNN, BiLSTM+Attention, HAN, BERT, Capsule, TextGCN, etc. All of them can be used as the text recognition model. As long as they are trained with a sufficient amount of corresponding training samples until convergence, they can be used as the text recognition model of this application. In this application, the Fasttext model is recommended.

[0079] It can be seen from this embodiment that this application has multiple advantages, including but not limited to:

[0080] This application uses an edge enhancement network layer to perform edge enhancement processing on the to-be-detected advertisement image in multiple image directions, highlighting the edge features of the text object, improving the detection effect, and obtaining a corresponding number of edge-enhanced images; performing text object detection on the advertisement image and the edge-enhanced images, determining the candidate boxes belonging to the text object and the confidence information corresponding to each candidate box; performing non-maximum suppression processing on the candidate boxes to eliminate redundant candidate box information and improve the working efficiency of the text security detection system. Subsequently, text recognition is performed on the text objects corresponding to the remaining candidate boxes, and the corresponding security type is discriminated according to the recognized text. If it is non-secure text, relevant users who publish the advertisement image are alerted to carry out subsequent rectification work.

[0081] This application constructs an edge enhancement network layer. The pixel values in the edge-enhanced image break through the conventional range between 0 and 255, and the pixel value range is increased to between 0 and 2295, thereby improving the contrast of the pixels in the image, making the pixel edges more obvious, which is extremely beneficial to the text detection work; at the same time, it provides richer edge features for the subsequent text detection work, making its text detection ability stronger. In addition, different from the traditional method of directly performing model training and inference on the original image to be detected, this application first proposes to perform model training and inference on the edge-enhanced image. The edge-enhanced image can enrich the edge features of the text object, which is beneficial to detecting it from the background of the original image to be detected, and improving the detection accuracy and working efficiency of the text security type detection system.

[0082] In summary, the judgment result of this application on whether the to-be-detected image contains sensitive words, phrases, information, etc. has a higher confidence level, can be highly trusted, and is suitable for detecting target images such as advertisement images for the existence of sensitive words, phrases, and information in application scenarios such as e-commerce platforms.

[0083] Please refer to Figure 2, in the in - depth embodiment, in step S1200, an edge enhancement network layer is used to perform edge enhancement processing on the advertisement picture in multiple image directions to obtain a corresponding plurality of edge - enhanced pictures, including the following steps:

[0084] Step S1210: Based on pixel shifting, process the advertisement picture to obtain a plurality of offset image data;

[0085] In an image, each pixel has eight neighborhoods, which are, in clockwise order: upper - left, upper - center, upper - right, right - center, lower - right, lower - center, lower - left, and left - center. In the embodiment of the present application, the advertisement picture to be detected is scaled to a fixed size, defined as Img, and image shifting processing is performed in eight neighborhood directions: Each pixel point in the advertisement picture to be detected is moved one or several pixel positions in the upper - left direction to obtain an offset image ImgLeftTop; Each pixel point in the advertisement picture to be detected is moved one pixel position in the upper - center direction to obtain an offset image ImgTop; Each pixel point in the advertisement picture to be detected is moved one pixel position in the upper - right direction to obtain an offset image ImgRightTop; Each pixel point in the advertisement picture to be detected is moved one pixel position in the right - center direction to obtain an offset image ImgRight; Each pixel point in the advertisement picture to be detected is moved one pixel position in the lower - right direction to obtain an offset image ImgRightBottom; Each pixel point in the advertisement picture to be detected is moved one pixel position in the lower - center direction to obtain an offset image ImgBottom; Each pixel point in the advertisement picture to be detected is moved one pixel position in the lower - left direction to obtain an offset image ImgLeftBottom; Each pixel point in the advertisement picture to be detected is moved one pixel position in the left - center direction to obtain an offset image ImgLeft. In summary, after step S1210, a plurality of offset image data can be obtained for the original advertisement picture to be detected.

[0086] Step S1220: Perform difference operation on the advertisement picture and the offset image data one by one to obtain a plurality of difference image data;

[0087] There is only an offset of one pixel unit in different directions between the advertisement picture to be detected and the offset image data. Perform difference operation on the advertisement picture to be detected and the offset image data and take the absolute value. This value can reflect the pixel value difference between adjacent pixel positions in the image, that is, gradient or edge information. Perform difference processing on the advertisement picture to be detected and the offset image data respectively and take the absolute value, and then difference image data in different directions can be obtained as follows:

[0088] DfImgLeftTop = abs(Img - ImgLeftTop)

[0089] DfImgTop = abs(Img - ImgTop)

[0090] DfImgRightTop = abs(Img - ImgRightTop)

[0091] DfImgRight = abs(Img - ImgRight)

[0092] DfImgRightBottom = abs(Img - ImgRightBottom)

[0093] DfImgBottom = abs(Img - ImgBottom)

[0094] DfImgLeftBottom = abs(Img - ImgLeftBottom)

[0095] DfImgLeft = abs(Img - ImgLeft)

[0096] Step S1230: Multiple times, perform superposition and combination processing on the advertisement picture and the any number of differential image data to obtain multiple edge-enhanced pictures.

[0097] The advertisement picture is the original picture to be detected, and the differential image data extracts the edge information in different image directions of the advertisement picture. Superposing and combining the advertisement picture and the differential image data realizes edge enhancement of the foreground object in the advertisement picture. For text objects, this means can more prominently highlight the edge features of the text object, making the text object easier to detect and improving the detection effect.

[0098] Perform superposition and combination processing on the advertisement picture and the any N (maximum is 8) differential image data to obtain an image data with edge enhancement in the corresponding image direction. In the embodiments of the present application, the following six image data with edge enhancement can be constructed, and their superposition and combination are specifically as follows:

[0099] ImgBE1 = Img + DfImgLeft + DfImgTop + DfImgRight + DfImgBottom

[0100] ImgBE2 = Img + DfImgLeftTop + DfImgRightTop + DfImgLeftBottom + DfImgRightBottom

[0101] ImgBE3 = Img + DfImgLeft + DfImgTop + DfImgRight + DfImgBottom + DfImgLeftTop + DfImgRightTop + DfImgLeftBottom + DfImgRightBottom

[0102] ImgBE4 = DfImgLeft + DfImgTop + DfImgRight + DfImgBottom

[0103] ImgBE5 = DfImgLeftTop + DfImgRightTop + DfImgLeftBottom + DfImgRightBottom

[0104] ImgBE6 = DfImgLeft + DfImgTop + DfImgRight + DfImgBottom + DfImgLeftTop + DfImgRightTop + DfImgLeftBottom + DfImgRightBottom

[0105] In the embodiments of the present application, after the edge enhancement network layer performs edge enhancement on the to-be-detected picture, the pixel values of the edge-enhanced picture break through the conventional range between 0 and 255, and the pixel value range is increased to between 0 and 2295, improving the pixel contrast relative to the to-be-detected picture; the above steps also provide richer edge features in the picture. The above effects are beneficial to subsequent text detection work, making the text detection ability of the text detection model stronger.

[0106] In the in-depth example, the step S1300, performing text object detection on the advertisement picture and the corresponding multiple edge-enhanced pictures respectively, and determining candidate boxes belonging to text objects and confidence information corresponding to each candidate box, includes the following steps:

[0107] Step S1310, performing standardized preprocessing on the advertisement picture and the edge-enhanced picture;

[0108] The advertisement picture is the original picture to be detected, and the edge-enhanced picture is the image data obtained in step S1200, which enhances the edge information of its image relative to the original picture to be detected, making its text objects easier to detect.

[0109] Performing standardized preprocessing on the advertisement picture and the edge-enhanced picture realizes the processing of centralizing the image data by removing the mean value to make it conform to the data distribution law, and it is easier to obtain better generalization effects.

[0110] Step S1320: Perform text object detection on the advertisement picture and the edge-enhanced picture respectively to obtain candidate boxes belonging to text objects and confidence information corresponding to each candidate box.

[0111] Call a pre-trained text detection model to perform text object detection on the advertisement picture and the multiple edge-enhanced pictures respectively, and each picture obtains candidate boxes belonging to text objects and confidence information corresponding to each candidate box.

[0112] The candidate box refers to the position information of the text object in the picture to be detected, and the confidence information corresponding to each candidate box refers to the probability that the candidate box indicates correctly. A high confidence indicates a high probability that the candidate box refers to correctly, and vice versa.

[0113] For the text detection model, various excellent text detection models in existing technologies can be selected, including but not limited to: DBNet, PSENet, PANNet, EAST, CTPN, SegLink, PixelLink, TextBoxes, etc. They are all mature text detection models. As long as they are trained with sufficient corresponding training samples until convergence, they can be used as the text detection model of this application. In this application, the DBNet model is recommended.

[0114] The DBNet model is a segmentation-based method. First, it extracts progressive features of the advertisement picture to be detected through a progressive feature extraction backbone network. Secondly, it obtains feature maps of the same size through pyramid upsampling, fuses the feature maps in a feature concatenation manner, and predicts the fused feature maps to obtain a probability map and a threshold map at the same time. An approximate binary map is calculated through the probability map and the threshold map. During model training, it is supervised and trained through the probability map, the threshold map, and the binary map until the model converges. During model inference, only by predicting the probability map and the binary map through the model can the candidate boxes of text objects and their confidence information be obtained.

[0115] Please refer to Figure 3 , in a deepened example, Step S1400: Perform text recognition on the image corresponding to the remaining candidate boxes after non-maximum suppression, and distinguish the corresponding security types according to the recognized text, including the following steps:

[0116] Step S1410: Perform non-maximum suppression processing on the candidate boxes and the confidence information corresponding to each candidate box, delete the candidate boxes that point to the same text object and have a lower confidence, and obtain the remaining candidate boxes;

[0117] During the process of text object detection by the text detection model, a large number of candidate boxes will be generated at the position of the same target text object, and these candidate boxes all point to the same text object.

[0118] There may be no, one, or multiple text objects in the advertisement picture or the edge-enhanced picture. After each text object is detected by a text detection model, one or more candidate boxes may be output. Therefore, there is a situation where there is at least one candidate box in the candidate boxes that points to the same text object in the same picture. If subsequent processing is performed one by one on all candidate boxes that point to the same text object, there will be relatively serious redundant work, greatly reducing the working efficiency of text recognition. Assuming that the confidence corresponding to the candidate box can represent the matching probability of the true position information of the candidate box and its corresponding text object, the candidate box with the highest confidence among all candidate boxes that point to the same text object in the same picture can be selected as the only candidate box that points to the text object, and the remaining candidate boxes that point to the same text object in the same picture are deleted. It is worth mentioning that the candidate boxes that point to the same text object have overlapping regions with each other. The "duplicate removal" operation can be performed according to the confidence of the candidate boxes and the overlapping regions between them. Therefore, the non-maximum suppression method is used to process the candidate boxes, suppressing the candidate boxes with non-maximum confidence that point to the same text object, and leaving the candidate box with the highest confidence that points to the same text object. The non-maximum suppression processing can effectively eliminate redundant candidate box information, and the remaining candidate boxes with the highest confidence that point to the same text object make the subsequent text recognition work more efficient and accurate.

[0119] Please refer to Figure 4 , in a specific embodiment, the step S1410 includes the following steps:

[0120] Step S1411: Reverse the order of the candidate boxes according to their confidence information to form an initial candidate set, and at the same time establish an empty target set;

[0121] Reverse the order of all the candidate boxes according to their corresponding confidence levels, and set the result of the arrangement as the initial candidate set. At the same time, establish an empty target set. The initial candidate set contains candidate boxes with maximum confidence and candidate boxes with non-maximum confidence. The target set is used to store candidate boxes with maximum confidence.

[0122] Step S1412: Remove the candidate box with the highest confidence from the current candidate set and put it into the target set;

[0123] Select the candidate box with the highest confidence in the current candidate set, remove the record of the candidate box from the initial candidate set, and at the same time store the candidate box in the target set.

[0124] Step S1413: Calculate the intersection-over-union ratio value between the candidate box and all candidate boxes in the remaining candidate set based on pixels;

[0125] Based on the intersection over union (IOU) score between the candidate box with the highest confidence removed from the current candidate set as described in pixel calculation step S1412 and all candidate boxes remaining in the candidate set after the removal.

[0126] The IOU score reflects the likelihood that two candidate boxes refer to the same text object; however, in practical applications, only the binary results of "same" and "different" are needed. Therefore, a threshold needs to be set as the similarity boundary value. When the IOU value is greater than this value, it is determined that the two candidate boxes refer to the same text object, otherwise it is determined that they refer to two different text objects.

[0127] The setting of the threshold plays a crucial role and can indirectly affect the working efficiency of text recognition and text security type discrimination. If the value is set too high, text objects that are actually the same will be determined as "different", resulting in subsequent redundant work problems; if the value is set too low, text objects that are actually different will be determined as "same" and then subsequent deletion operations will be performed. If the removed text objects contain the only sensitive words, phrases, information, etc., it will lead to incorrect discrimination results. Therefore, the setting of this threshold needs to be reasonably set by relevant technical personnel through experimental observation, result analysis, and prior knowledge such as experience.

[0128] Step S1414: Compare the IOU score with a preset threshold. When the IOU score is greater than the preset threshold, delete the corresponding candidate box in the candidate set;

[0129] When the IOU score is greater than the preset threshold, it indicates that the two candidate boxes refer to the "same" text object. To avoid subsequent duplicate work, delete the corresponding candidate box information in the candidate set; when the value is not greater than the preset threshold, it indicates that the two candidate boxes point to two different text objects, and retain the corresponding candidate box information in the candidate set.

[0130] Step S1415: Repeat the above operations until the candidate set is empty, and the candidate boxes in the target set are the output results.

[0131] Repeat step S1412, step S1413, and step S1414 until the candidate set is empty. Then the target set contains all candidate boxes with extremely high confidence, and there is a one-to-one match between the candidate boxes and the text objects, which are used as the output results.

[0132] Step S1420: Perform text recognition on the text images corresponding to the remaining candidate boxes respectively to identify the text objects;

[0133] According to the remaining candidate boxes, extract the text images at the corresponding positions in the advertisement picture to be detected one by one. After image normalization processing, use a pre-trained text recognition model to perform text recognition to obtain the text content in the text images.

[0134] For the text recognition model, various excellent text recognition models in the prior art can be selected, including but not limited to: CRNN, RARE, ASTER, DAN, SAR, etc. These are all mature text recognition models. As long as they are trained with a sufficient amount of corresponding training samples until convergence, they can be used as the text recognition model of this application. In this application, the CRNN model is recommended.

[0135] The CRNN model includes three parts. The first is the convolutional layer, which uses a deep convolutional layer to extract deep features from the input image to obtain a feature sequence. The second is the recurrent layer, which uses a bidirectional RNN to learn and predict the feature sequence. The third is the transcription layer, which uses the CTC loss to convert the output result of the recurrent layer into a final label sequence.

[0136] Step S1430: Determine the security type according to the text object and output the determination result.

[0137] The text content recognized by the text recognition model may or may not contain the sensitive words, terms, information, etc. Therefore, further security type discrimination is performed according to the text recognition result to determine its corresponding security type. If the determined type is secure, advertising image promotion, publicity, etc. can be carried out according to the user's requirements. If the determined type is insecure, the user can be warned to make corresponding modifications to the text object in the advertising image according to the relevant rules of the relevant platform.

[0138] The discrimination method can be word library matching, text discrimination model, etc. In the embodiments of this application, a pre-trained text discrimination model can be used to perform security type discrimination on the recognized text.

[0139] For the text discrimination model, various excellent text recognition models in the prior art can be selected, including but not limited to: Fasttext, TextCNN, DPCNN, TextRCNN, BiLSTM+Attention, HAN, BERT, Capsule, TextGCN, etc. These can all be used as text recognition models. As long as they are trained with a sufficient amount of corresponding training samples until convergence, they can be used as the text recognition model of this application. In this application, the Fasttext model is recommended.

[0140] The Fasttext model encodes the text using a pre-trained word vector model and then uses a linear classifier to obtain its category.

[0141] In the embodiments of the present application, non-maximum suppression is used to perform redundant deletion of candidate boxes, leaving only the candidate box with the highest confidence level pointing to the same text object, which can effectively eliminate redundant candidate box information, reduce the repetitive work of subsequent text recognition and text discrimination, greatly improve the working efficiency of the text security type detection system, and make the text security type detection method more efficient and accurate.

[0142] A text security type detection device provided to meet one of the purposes of the present application includes an image acquisition module, an edge enhancement module, a text detection module, and a security discrimination module. Among them, the image acquisition module is used to acquire the advertisement picture therein from the advertisement release information; the edge enhancement module is used to perform edge enhancement processing on the advertisement picture in multiple image directions by using an edge enhancement network layer to obtain a corresponding plurality of edge-enhanced pictures; the text detection module is used to perform text object detection according to the advertisement picture and the corresponding plurality of edge-enhanced pictures respectively, and determine the candidate boxes belonging to the text object and the confidence information corresponding to each candidate box; the security discrimination module is used to perform text recognition on the image corresponding to the remaining candidate boxes after non-maximum suppression, and discriminate the corresponding security type according to the recognized text.

[0143] In a further embodiment, the image acquisition module includes: a trigger sub-module that responds to an advertisement release request triggered by a user and acquires the corresponding advertisement release information, and the advertisement release information includes an advertisement picture. An acquisition sub-module that acquires the advertisement picture therein from the advertisement release information, which is a picture to be detected for identifying its text security type.

[0144] In a further example, the edge enhancement module includes: a bias sub-module configured to obtain a plurality of bias image data by processing the advertisement picture based on pixel shift; a difference sub-module that performs difference on the advertisement picture and the bias image data one by one to obtain a plurality of difference image data; an enhancement sub-module that performs superposition combination processing on the advertisement picture and any of the difference image data multiple times to obtain a plurality of edge-enhanced pictures.

[0145] In a further example, the text detection module includes a preprocessing sub-module that performs standardized preprocessing on the advertisement picture and the edge-enhanced picture; a detection sub-module that performs text object detection on the advertisement picture and the edge-enhanced picture respectively to obtain candidate boxes belonging to the text object and the confidence information corresponding to each candidate box.

[0146] In a refined example, the security discrimination module includes a non-maximum suppression sub-module that performs non-maximum suppression processing on the candidate boxes and the confidence information corresponding to each candidate box, deletes the candidate boxes that point to the same text object and have a lower confidence, and obtains the remaining candidate boxes; a text recognition sub-module that respectively performs text recognition on the text images corresponding to the remaining candidate boxes to recognize the text objects; and a determination sub-module that determines the security type according to the text objects and outputs a determination result.

[0147] In a specific example, in the non-maximum suppression sub-module, there is an initialization unit that takes the candidate boxes and arranges them in reverse order according to their confidence information as the initial candidate set, and at the same time creates an empty target set; a selection unit that removes the candidate box with the highest confidence from the current candidate set and puts it into the target set; a calculation unit that calculates the intersection-over-union ratio value between the candidate box and all candidate boxes in the remaining candidate set based on pixels; a processing unit that compares the intersection-over-union ratio value with a preset threshold, and when the intersection-over-union ratio value is greater than the preset threshold, deletes the corresponding candidate box in the candidate set; and an iteration unit that repeats the above operations until the candidate set is empty, and the candidate boxes in the target set are the output results.

[0148] In a preferred embodiment, the text detection model is a DBNet model, the text recognition model is a CRNN model, the text discrimination model is a CNN model, the basic network architecture of the text detection model is a DBNet model, the basic network architecture of the text recognition model is a CRNN model, the basic network architecture of the text discrimination model is a CNN model, and the target text is sensitive words, phrases, information, etc.

[0149] To solve the above technical problems, an embodiment of the present application also provides a computer device. As Figure 6 shown, the internal structure schematic diagram of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a text security type detection method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the text security type detection method of the present application. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art can understand, Figure 6The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0150] In this embodiment, the processor is used to execute Figure 5 the specific functions of each module and its sub-modules. The memory stores the program codes and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program codes and data required to execute all modules / sub-modules in the text security type detection device of this application. The server can call the program codes and data of the server to execute the functions of all sub-modules.

[0151] This application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the text security type detection method according to any embodiment of this application.

[0152] This application also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by one or more processors, the steps of the method according to any embodiment of this application are implemented.

[0153] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of this application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a Read-Only Memory (ROM), or a Random Access Memory (RAM), etc.

[0154] In summary, this application uses an edge enhancement network layer to highlight the edge features of text objects, improving the detection effect; uses non-maximum suppression processing to eliminate redundant candidate box information, improving the working efficiency of the text security detection system; screens relevant pictures and text data for the text detection model, text recognition model, and text discrimination model for model training until convergence; enabling the text security type detection system to be efficient, accurate, and have wide adaptability.

[0155] Those skilled in the art can understand that the various operations, methods, steps, measures, and solutions in the processes discussed in this application can be alternated, changed, combined, or deleted. Further, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, those in the prior art having steps, measures, and solutions in the various operations, methods, and processes disclosed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted.

[0156] The above are only some embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for detecting text security types, characterized in that, it includes the following steps: Obtain an advertisement picture to be detected; Use an edge enhancement network layer to perform edge enhancement processing on the advertisement picture in multiple image directions to obtain corresponding multiple edge-enhanced pictures, including: obtaining multiple offset image data by performing pixel shift processing on the advertisement picture; performing difference on the advertisement picture and the offset image data one by one to obtain multiple difference image data; performing superposition and combination processing on the advertisement picture and any of the difference image data multiple times to obtain multiple edge-enhanced pictures; Perform text object detection based on the advertisement picture and the edge-enhanced pictures, and determine candidate boxes belonging to text objects and confidence information corresponding to each candidate box; Perform text recognition on the images corresponding to the remaining candidate boxes after non-maximum suppression, and determine their corresponding security types according to the recognized text.

2. The method for detecting text security types according to claim 1, characterized in that, obtaining an advertisement picture to be detected includes the following steps: Respond to an advertisement release request triggered by a user, and obtain the corresponding submitted advertisement release information, where the advertisement release information includes an advertisement picture; Obtain the advertisement picture therein from the advertisement release information.

3. The method for detecting text security types according to claim 1, characterized in that, performing text object detection based on the advertisement picture and the edge-enhanced pictures, and determining candidate boxes belonging to text objects and confidence information corresponding to each candidate box includes the following steps: Perform standardized preprocessing on the advertisement picture and the edge-enhanced pictures; Perform text object detection on the advertisement picture and the edge-enhanced pictures respectively to obtain candidate boxes belonging to text objects and confidence information corresponding to each candidate box.

4. The method for detecting text security types according to claim 1, characterized in that, performing text recognition on the images corresponding to the remaining candidate boxes after non-maximum suppression, and determining their corresponding security types according to the recognized text includes the following steps: Perform non-maximum suppression processing on the candidate boxes and the confidence information corresponding to each candidate box, delete candidate boxes pointing to the same text object and having a lower confidence, and obtain the remaining candidate boxes; Perform text recognition on the text images corresponding to the remaining candidate boxes respectively to recognize text objects; Determine the security type according to the text object and output the determination result.

5. The method for detecting text security types according to claim 4, characterized in that, performing non-maximum suppression processing on the candidate boxes and the confidence information corresponding to each candidate box, and deleting candidate boxes pointing to the same text object and having a lower confidence includes the following steps: Arrange the candidate boxes and their confidence information in reverse order as the initial candidate set, and at the same time establish an empty target set; Remove the candidate box with the highest confidence from the current candidate set and put it into the target set; Calculate the intersection-over-union ratio value between the candidate box and all candidate boxes in the remaining candidate set based on pixels; Compare the intersection-over-union ratio value with a preset threshold, and when the intersection-over-union ratio value is greater than the preset threshold, delete the corresponding candidate box in the candidate set; Repeat the above operations until the candidate set is empty, and the candidate bounding boxes in the target set are the output results.

6. A text security type detection device, characterized in that, it includes: An image acquisition module, configured to acquire an advertisement picture to be detected; An edge enhancement module, configured to perform edge enhancement processing on the advertisement picture in multiple image directions by using an edge enhancement network layer to obtain corresponding multiple edge-enhanced pictures, including: obtaining multiple offset image data by performing pixel shift processing on the advertisement picture; performing difference on the advertisement picture and the offset image data one by one to obtain multiple difference image data; performing superposition combination processing on the advertisement picture and any of the difference image data multiple times to obtain multiple edge-enhanced pictures; A text detection module, configured to perform text object detection according to the advertisement picture and the edge-enhanced pictures, and determine candidate bounding boxes belonging to text objects therein and confidence information corresponding to each candidate bounding box; A security discrimination module, configured to perform text recognition on the image corresponding to the remaining candidate bounding boxes after non-maximum suppression, and discriminate the corresponding security type according to the recognized text.

7. The text security type detection device according to claim 6, characterized in that, The image acquisition module further includes: A trigger sub-module, configured to respond to an advertisement release request triggered by a user, and acquire corresponding advertisement release information submitted thereby, where the advertisement release information includes an advertisement picture; An acquisition sub-module, configured to acquire the advertisement picture therein from the advertisement release information.

8. A computer device, including a central processing unit and a memory, characterized in that, The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that, It stores a computer program implemented according to the method according to any one of claims 1 to 5 in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.

10. A computer program product, including a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method and device for extracting text stroke images from image

    CN102810155A

  • Interaction platform based method for rapidly detecting text in complex background

    CN105404868A