Small target detection method and its device, equipment, medium and product
By using the target detection model on the e-commerce platform and the image recognition model that adds a classification head that can suppress weak neurons to identify advertising images, the false detection and missed detection of small target sensitive items in the e-commerce platform is solved, and higher recognition accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202111591509.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-23
AI Technical Summary
When e-commerce platforms identify small target sensitive items in advertising pictures, there is a problem of high false detection rate or high missed detection rate, resulting in the user's advertisements that meet the requirements being misjudged as containing sensitive items, or advertisements that do not meet the requirements being misjudged as normal, causing trouble.
The target detection model trained to convergence is used to detect the advertising image, obtain the target area image, and use the image recognition model with the classification head that can suppress weak neurons for image recognition to generate a small target recognition sequence, and finally output the small target recognition result according to the preset conditions of a specific instance scene.
It improves the accuracy of small-target items identification, reduces false detection and missed detection rates, helps e-commerce platforms to more effectively eliminate advertising images that violate relevant regulations, and significantly reduces time and labor costs.
Smart Images

Figure CN114332586B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image detection technology, and in particular to a small target detection method and its corresponding device, computer equipment, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of technology, the use of artificial neural network models for target detection has become the mainstream technology. For application scenarios such as e-commerce platforms, a large number of advertising images are generated every day. Due to the influence of risk control systems, advertising messages posted on e-commerce platforms must comply with relevant regulations. Due to the restrictions of e-commerce platforms on sensitive items in advertising images, such as cigarettes, it is generally necessary to detect and identify targets in advertising images uploaded by users, and then perform further processing.
[0003] For application scenarios such as e-commerce platforms, if the advertising images uploaded by merchants to display advertising information contain small sensitive items, the current main approach is to conduct manual screening, which is time-consuming and labor-intensive, and is prone to missed detections and false detections due to manual fatigue and small targets. Secondly, the application of existing target recognition models and target detection models in related application examples also has the problem of high false detection rates or high missed detection rates, especially in the detection and identification of small targets. For e-commerce platforms, this will cause advertisements posted by users that comply with relevant regulations to be mistakenly detected as containing sensitive items, or advertisements posted by users that do not comply with relevant regulations to be missed as normal, which will cause great trouble to e-commerce platforms.
[0004] In addition, it is unrealistic to completely rely on manual investigation, which is time-consuming. Nowadays, the products on e-commerce platforms are rapidly iterating, and so are the corresponding advertising information. Therefore, e-commerce users are required to quickly update advertising information in a short period of time in order to compete for sales share in the e-commerce market.
[0005] Therefore, how to accurately and efficiently identify small target objects from advertisement pictures to be detected that contain various sensitive objects so as to make the recognition results more accurate has become a technical problem that needs to be solved in this field. Summary of the invention
[0006] The primary purpose of the present application is to solve at least one of the above problems and provide a small target detection method and its corresponding device, computer equipment, computer-readable storage medium, and computer program product.
[0007] In order to meet the various objectives of this application, this application adopts the following technical solutions:
[0008] A small target detection method includes the following operations: obtaining an advertisement picture to be detected;
[0009] Using a target detection model that has been trained to convergence to perform target detection on the advertisement image to obtain a target area image;
[0010] Using an image recognition model that has been trained to convergence and to which a classification head capable of suppressing weak neurons is added, image recognition is performed on the advertisement image and the target area image to obtain a small target recognition sequence;
[0011] The small target recognition sequence is used to output the corresponding small target recognition result according to the preset conditions of a specific instance scene.
[0012] In a further embodiment, obtaining an advertisement image to be detected includes the following steps:
[0013] Responding to an advertisement publishing request triggered by a user, obtaining the corresponding advertisement publishing information submitted by the user, wherein the advertisement publishing information includes an advertisement picture;
[0014] The advertisement picture is obtained from the advertisement publishing information.
[0015] In a further embodiment, a target detection model that has been trained to convergence is used to perform target detection on the advertisement image to obtain a target area image, including the following steps:
[0016] A convolutional backbone network is used to extract features from the advertisement image to obtain multi-layer feature maps of different scales;
[0017] A region generation network is used to generate a plurality of candidate regions of interest for the multi-layer feature maps of different scales, and then an alignment operation of the regions of interest is performed;
[0018] The head network is used to perform three branch processing on the feature map after alignment of the region of interest, namely bounding box regression processing, recognition processing and mask map prediction, to obtain the detection result;
[0019] According to the detection result, the target area in the advertisement picture is captured to obtain a corresponding target area image.
[0020] In a further embodiment, an image recognition model that has been trained to convergence and to which a classification head capable of suppressing weak neurons is added is used to perform image recognition on the advertisement image and the target area image to obtain a small target recognition sequence, including the following steps:
[0021] Perform image block embedding on the input image to obtain multiple image block vectors, and add a classification vector to form multiple embedding vectors;
[0022] Adding a position encoding vector to the multiple embedding vectors to form an input vector, wherein the position encoding vector can maintain spatial position information between image blocks;
[0023] Stacking multiple encoding modules for the input vector to extract features, wherein the encoding modules include multi-head attention and multi-layer perceptron;
[0024] The common classification head is used to further transform the classification space of the final deep classification vector. At the same time, the new classification head is used to suppress weak neurons to transform the classification space of the deep classification vector, and two classification probabilities are obtained.
[0025] The above steps are performed respectively on the advertisement picture and the target area image to obtain a corresponding plurality of recognition results, forming a small target recognition sequence of the advertisement picture.
[0026] In a further embodiment, the small target recognition sequence is used to output a corresponding small target recognition result according to a preset condition of a specific instance scene, including the following steps:
[0027] Sorting the small target recognition sequence to obtain its maximum and minimum values;
[0028] Identify the actual needs of the application scenario of the example, use the minimum value as the probability value of whether the advertisement image contains the small target object in the high accuracy scenario, and use the maximum value as the probability value of whether the advertisement image contains the small target object in the high recall scenario;
[0029] The probability value is compared with a preset threshold value, and when the probability value is greater than the preset threshold value, it is determined that the advertisement image contains the small target object; otherwise, it is determined that the advertisement image does not contain the small target object.
[0030] In a preferred embodiment, the target detection model is a Mask-RCNN model, the image recognition model is a ViT model with a classification head that can suppress weak neurons added, the basic network architecture of the target detection model is a Mask-RCNN model, the basic network architecture of the image recognition model is a ViT model with a classification head that can suppress weak neurons added, and the target object is a cigarette.
[0031] A small target detection device provided to meet one of the purposes of the present application includes an image acquisition module, a target detection module, an image recognition module and a target discrimination module, wherein the image acquisition module is used to obtain an advertising image from the advertising information; the target detection module is configured to use a target detection model that has been trained to convergence to perform target detection on the advertising image to obtain a target area image; the image recognition module is configured to use an image recognition model that has been trained to convergence and to which a classification head that can inhibit weak neurons is added to perform image recognition on the advertising image and the target area image to obtain a small target recognition sequence; the target discrimination module is used to use the small target recognition sequence to output a corresponding small target recognition result according to preset conditions of a specific instance scene.
[0032] In a further embodiment, the image acquisition module includes: a response submodule, used to respond to an advertising publishing request triggered by a user, and obtain the corresponding advertising publishing information submitted, wherein the advertising publishing information includes advertising pictures; an acquisition submodule; used to obtain the advertising pictures from the advertising publishing information, and used to identify the pictures of small target objects to be detected.
[0033] In the in-depth example, the target detection module includes: a convolutional backbone submodule, which uses a convolutional backbone network to perform feature extraction on the advertising image to obtain multiple layers of feature maps of different scales; a region of interest submodule, which uses a region generation network to generate multiple candidate regions of interest for the multiple layers of feature maps of different scales, and then performs a region of interest alignment operation; a detection submodule, which uses a head network to perform three branch processing on the feature map after the region of interest alignment, namely, bounding box regression processing, recognition processing and mask map prediction, to obtain a detection result; a capture submodule, which is used to capture the target area in the advertising image according to the detection result to obtain the corresponding target area image.
[0034] In the further example, the image recognition module includes an embedding submodule, which is used to embed image blocks for the input image to obtain multiple image block vectors, and at the same time add a classification vector to form multiple embedding vectors; a position encoding submodule, which is used to add a position encoding vector to the multiple embedding vectors to form an input vector, and the position encoding vector can maintain the spatial position information between image blocks; a feature extraction submodule, which is used to stack multiple encoding modules for feature extraction on the above-mentioned input vector, and the encoding module includes multi-head attention and multi-layer perceptron; a classification submodule, which uses an ordinary classification head to perform further classification space transformation on the final deep classification vector, and at the same time uses a new classification head to suppress weak neurons to perform classification space transformation on the deep classification vector to obtain two classification probabilities; a sequence generation submodule, which performs the above steps on the advertising picture and the target area image respectively, obtains corresponding multiple recognition results, and constitutes a small target recognition sequence of the advertising picture.
[0035] In the further example, the target discrimination module includes: a sorting submodule, which is used to sort the small target recognition sequence to obtain its maximum and minimum values; a probability value calculation submodule, which is used to identify the actual needs of the example application scenario, and in a high-accuracy scenario, the minimum value is used as the probability value of whether the advertising image contains the small target item, and in a high-recall scenario, the maximum value is used as the probability value of whether the advertising image contains the small target item; a discrimination submodule, which is used to compare the probability value with a preset threshold value, and when the probability value is greater than the preset threshold value, it is determined that the advertising image contains the small target item, otherwise it is determined that the small target item is not contained.
[0036] In a preferred embodiment, the target detection model is a Mask-RCNN model, the image recognition model is a ViT model with a classification head that can suppress weak neurons added, the basic network architecture of the target detection model is a Mask-RCNN model, the basic network architecture of the image recognition model is a ViT model with a classification head that can suppress weak neurons added, and the target object is a cigarette.
[0037] A computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory, wherein the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the small target detection method described in the present application.
[0038] A computer-readable storage medium is provided to meet another purpose of the present application, which stores a computer program implemented according to the small target detection method in the form of computer-readable instructions. When the computer program is called and executed by a computer, the steps included in the method are executed.
[0039] A computer program product provided to meet another purpose of the present application includes a computer program / instruction, which, when executed by a processor, implements the steps of the method described in any embodiment of the present application.
[0040] Compared with the prior art, the advantages of this application are as follows:
[0041] The present application obtains an advertising image to be detected; uses a target detection model that has been trained to convergence to perform target detection on the advertising image, and captures a target area image in the advertising image based on the detection result; then uses an image recognition model that has been trained to convergence and has a classification head that can inhibit weak neurons to perform image recognition on the advertising image and the target area image respectively, and then obtains corresponding recognition results, and combines the recognition results to form a small target recognition sequence, which reflects multiple probability values of small target items contained in the advertising image; finally, the actual needs of the identification example application scenario are identified, and according to the actual needs, the probability value under preset conditions is obtained from the small target recognition sequence as the final probability value of whether the advertising image contains small target items, and a final judgment is made to output the result.
[0042] This application adopts enhanced detection of detection, recognition and scene conditions, and uses an image recognition model with a classification head that can suppress weak neurons to identify small target objects. On the one hand, the classification head added to the image recognition model uses the method of dynamically activating neurons to suppress the interference of weak neurons on the classification space transformation, realize the decoupling of features between different categories of small targets, and ultimately improve the accuracy of small target object recognition; on the other hand, enhanced detection can improve the accuracy of small target object recognition to a greater extent through hierarchical condition restrictions, and solve the business needs of high accuracy and high recall in the application scenarios of e-commerce platform examples. Ultimately, it helps e-commerce platforms to more effectively exclude advertising images that violate relevant regulations, while greatly reducing time and labor costs.
[0043] In summary, the judgment result made by this application on whether the advertising image contains small target objects has a higher confidence level and can be highly trusted. It is suitable for detecting whether there are sensitive target objects in advertising images in application scenarios such as e-commerce platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0045] Figure 1 A flowchart of a typical embodiment of the small target detection method of the present application is shown;
[0046] Figure 2This is a schematic diagram of the process of performing target detection on an advertisement image in an embodiment of the present application;
[0047] Figure 3 A schematic diagram of a process of performing image recognition on an input image in an embodiment of the present application;
[0048] Figure 4 This is a schematic diagram of the structure of a common classification head and a newly added classification head in an embodiment of the present application;
[0049] Figure 5 A schematic diagram of a process for outputting small target recognition results for an example scenario in an embodiment of the present application;
[0050] Figure 6 This is a principle block diagram of the small target detection device of the present application;
[0051] Figure 7 A schematic diagram of the structure of a computer device used in this application. DETAILED DESCRIPTION
[0052] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.
[0053] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0054] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0055] It will be understood by those skilled in the art that the "client", "terminal" and "terminal device" used herein include both devices with wireless signal receivers, which are devices with only wireless signal receivers without transmission capabilities, and devices with receiving and transmitting hardware, which are devices with receiving and transmitting hardware capable of two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers, tablet computers, which have single-line displays or multi-line displays or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service, personal communication system), which can combine voice, data processing, fax and / or data communication capabilities; PDA (Personal Digital Assistant, personal digital assistant), which may include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar and / or GPS (Global Positioning System, global positioning system) receiver; conventional laptop and / or palmtop computers or other devices, which have and / or include a conventional laptop and / or palmtop computer or other device with and / or including a radio frequency receiver. The "client", "terminal" and "terminal device" used herein may be portable, transportable, installed in a vehicle (air, sea and / or land), or suitable for and / or configured to run locally, and / or in a distributed form, at any other location on the earth and / or in space. The "client", "terminal" and "terminal device" used herein may also be a communication terminal, an Internet terminal, a music / video playing terminal, for example, a PDA, a MID (Mobile Internet Device) and / or a mobile phone with a music / video playing function, or a smart TV, a set-top box and other devices.
[0056] The hardware referred to by the names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann principle, such as a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit calls the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input and output devices to complete specific functions.
[0057] It should be pointed out that the concept of "server" referred to in this application can also be extended to the case of server clusters. According to the network deployment principle understood by those skilled in the art, the servers should be logically divided. In physical space, these servers can be independent of each other but can be called through interfaces, or integrated into a physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility, and should not use it to restrict the implementation of the network deployment method of this application.
[0058] Unless expressly specified, one or more technical features of the present application can be deployed on a server for implementation and accessed by a client through a remote call to obtain an online service interface provided by the server, or can be directly deployed and run on a client for access.
[0059] The neural network models referenced or may be referenced in this application, unless expressly specified, can be deployed on a remote server and remotely called on the client, or can be deployed and directly called on a client with sufficient device capabilities. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid excessive occupation of the client's hardware operating resources.
[0060] Unless explicitly specified, the various data involved in this application can be stored remotely on a server or on a local terminal device, as long as it is suitable for being called by the technical solution of this application.
[0061] Those skilled in the art should be aware that, although the various methods of the present application are described based on the same concept and thus present commonality with each other, unless otherwise specified, these methods can be independently executed. Similarly, for each embodiment disclosed in the present application, they are all proposed based on the same inventive concept, therefore, concepts with the same expression, and concepts that are appropriately changed for convenience despite different expressions, should be understood as equivalent.
[0062] Unless the mutually exclusive relationship between the embodiments to be disclosed in this application is explicitly stated, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct a new embodiment, as long as such combination does not deviate from the creative spirit of this application and can meet the needs of the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.
[0063] A small target detection method of the present application can be programmed as a computer program product and deployed in a client or server for execution. The method can be executed by accessing an interface opened after the computer program product is running and performing human-computer interaction with the process of the computer program product through a graphical user interface.
[0064] See also Figure 1 In a typical embodiment, the small target detection method of the present application includes the following steps:
[0065] Step S1100: obtaining an advertisement picture to be detected;
[0066] In an exemplary application scenario used to assist in explanation, the advertising image to be detected may be an advertising image containing items on an e-commerce platform. The items displayed in the image are generally non-sensitive items. The implementation of the present application is to identify sensitive items from the advertising images, i.e., pre-set small target items with sensitivity, such as cigarettes, etc.; therefore, it is necessary to identify whether small sensitive items are contained in the advertising images for further processing.
[0067] The small target object generally refers to the imaging size attribute of the target object in the picture, that is, the pixel area occupied by the target object in the picture is small. According to the definition of SPIE (International Society for Optical Engineering), an authoritative organization in the relevant international field, the small target object refers to an object whose area is less than 80 pixel values in a 256*256 image, that is, a target less than 0.12% of 256*256 is a small target. In the embodiment of the present application, the small target object is a smaller target object that appears in the advertising picture of the e-commerce platform, such as cigarettes, etc. The area ratio of the target object in the image is generally less than 0.12% of the image area. It is known by relevant technical personnel in this field that in different scene images, the image area ratio of the target object is different, such as close-up images, etc.; therefore, the image should be a scene picture under normal circumstances.
[0068] In the application scenario of the e-commerce platform, if it is necessary to obtain the pictures to be detected, one implementation method is to receive the input of the e-commerce platform users, especially the input when the merchant instance users configure the advertising information, and use the advertising pictures in the advertising information as the pictures to be detected; in another method, the server of the e-commerce platform can batch process the advertising pictures in the e-commerce platform database in the background, and use these advertising pictures as the pictures to be detected for target detection.
[0069] Step S1200: using a target detection model that has been trained to convergence to perform target detection on the advertisement image to obtain a target area image;
[0070] The target detection model is a model that takes an advertisement image as input and takes the target object label and its bounding box indicating the image location of the target area as output.
[0071] The target detection model is implemented as a preferred neural network model. For example, in the embodiment of the present application, the target detection model is a Mask-RCNN that has been trained to convergence. Alternatively, the neural network model can select a variety of relatively excellent target detection models in the prior art, including but not limited to: YOLO series models, other R-CNN series models, SSD models, DETR, etc., all of which are mature target detection models.
[0072] In this embodiment, the target detection model that has been trained to convergence includes three network components. The first network component is a convolutional backbone network, which is used to extract features from the advertising image to obtain multiple layers of feature maps of different scales; the second network component is a region generation network, which is used to generate multiple candidate regions of interest for the multiple layers of feature maps of different scales, and then perform region of interest alignment operations; the third network component is a head network, which is used to perform three branch processing on the feature map after the region of interest is aligned, namely, bounding box regression processing, recognition processing and mask map prediction, to obtain a detection result.
[0073] The detection result includes a bounding box indicating the area where the target item in the advertising image is located. The target area indicated by the bounding box in the advertising image is intercepted according to the bounding box to obtain a corresponding target area image. The number of bounding boxes depends on the detection results in the actual scene application. There may be no, one or more bounding boxes. Correspondingly, the number of target area images is consistent with the number of bounding boxes. The advertising image and the target area image are output to the next step for further processing.
[0074] Step S1300, using an image recognition model that has been trained to convergence and to which a classification head capable of suppressing weak neurons is added, to perform image recognition on the advertisement image and the target area image, and obtain a small target recognition sequence;
[0075] The image recognition model is a model that takes the advertisement picture or the target area image as input and outputs a probability value of containing a target object.
[0076] In an embodiment of the present application, the image recognition model is a ViT that has been trained to convergence and to which a classification head capable of suppressing weak neurons has been added. The image recognition model is implemented as a preferred neural network model. Alternatively, the neural network model can select a variety of relatively excellent image recognition models in the prior art, including but not limited to: VGG series models, Inception series models, ResNet series models, EfficientNet series models, HRNet, etc., and the like, which are all mature image recognition models. As long as the classification head capable of suppressing weak neurons is added and trained to convergence with sufficient corresponding training samples, they can be used as the image recognition model of the present application.
[0077] The image recognition model includes three components. The first component is an embedding component, which is used to embed image blocks for the input image, obtain multiple image block vectors, and add a classification vector to form multiple embedding vectors. Then, position encoding vectors are added to the multiple embedding vectors to form an input vector. The position encoding vector can maintain the spatial position information between image blocks; the second component is a feature extraction component, which is used to stack multiple encoding modules for feature extraction for the above input vector, and the encoding module includes multi-head attention and multi-layer perceptron; the third component is a classification component, which is configured to use an ordinary classification head to further perform classification space transformation on the final deep classification vector, and use a newly added classification head to suppress weak neurons to perform classification space transformation on the deep classification vector to obtain two classification probabilities;
[0078] The image recognition model is used to perform the above-mentioned image recognition operation on the advertising image and the target area image obtained in the previous step, respectively, to obtain corresponding multiple recognition results, that is, one recognition result is generated for each input image, that is, the probability value of whether the target object is contained, thereby constituting a target recognition sequence of the advertising image.
[0079] Step S1400: using the small target recognition sequence to output a corresponding small target recognition result according to preset conditions of a specific instance scene;
[0080] The target recognition sequence includes multiple probability values for whether the advertising image contains the target object. The small target recognition sequence is sorted to obtain its maximum and minimum values; then the actual needs of the instance application scenario are identified, and the minimum value is used as the probability value of whether the advertising image contains the small target object in a high-accuracy scenario, and the maximum value is used as the probability value of whether the advertising image contains the small target object in a high-recall scenario; finally, the probability value is compared with a preset threshold value. When the probability value is greater than the preset threshold value, it is determined that the advertising image contains the small target object, otherwise it is determined that the small target object is not contained.
[0081] In summary, this typical embodiment shows that the present application obtains an advertising picture to be detected; uses a target detection model that has been trained to convergence to perform target detection on the advertising picture, and captures the target area image in the advertising picture according to the detection result; then uses an image recognition model that has been trained to convergence and has a classification head that can inhibit weak neurons to perform image recognition on the advertising picture and the target area image respectively, and then obtains corresponding recognition results, and combines the recognition results to form a small target recognition sequence, which reflects multiple probability values of small target items contained in the advertising picture; finally, the actual needs of the identification example application scenario are identified, and according to the actual needs, the probability value under preset conditions is obtained from the small target recognition sequence as the final probability value of whether the advertising picture contains small target items, and a final judgment is made to output the result.
[0082] This application adopts enhanced detection of detection, recognition and scene conditions, and uses an image recognition model with a classification head that can suppress weak neurons to identify small target objects. On the one hand, the classification head added to the image recognition model uses the method of dynamically activating neurons to suppress the interference of weak neurons on the classification space transformation, realize the decoupling of features between different categories of small targets, and ultimately improve the accuracy of small target object recognition; on the other hand, enhanced detection can improve the accuracy of small target object recognition to a greater extent through hierarchical condition restrictions, and solve the business needs of high accuracy and high recall in the application scenarios of e-commerce platform examples. Ultimately, it helps e-commerce platforms to more effectively exclude advertising images that violate relevant regulations, while greatly reducing time and labor costs.
[0083] In summary, the judgment result made by this application on whether the advertising image contains small target objects has a higher confidence level and can be highly trusted. It is suitable for detecting whether there are sensitive small target objects in advertising images in application scenarios such as e-commerce platforms.
[0084] See also Figure 2 In a further embodiment, the step S1200, using a target detection model that has been trained to convergence to perform target detection on the advertisement image to obtain a target area image, includes the following steps:
[0085] Step S1210: extracting features from the advertisement image using a convolutional backbone network to obtain feature maps of multiple layers and different scales;
[0086] The convolutional backbone network adopts the structure of ResNet-FPN, which specifically includes two parts. The first part uses ResNet-101 as the skeleton network to extract features from bottom to top, and the second part uses the FPN structure, i.e., the feature pyramid network, to transmit strong semantic information from top to bottom. ResNet-FPN can fuse the features of each level so that it has both strong semantic information and strong spatial information, thereby enhancing the semantic expression and position expression on the feature maps of different scales.
[0087] The ResNet-101 skeleton network performs feature extraction on the advertisement image, which can be divided into five stages according to the size of the extracted feature map, wherein the feature layers output by stages 1, 2, 3, 4 and 5 are defined as C1, C2, C3, C4 and C5.
[0088] The FPN structure adopts a top-down structure with horizontal connections, which fuses the feature maps of each layer from the shallow layer to the deep layer, and makes full use of the semantic features and position features of each stage. The FPN structure obtains the feature layer P5 with a preset number of channels from the top layer C5 of the skeleton network through 1*1 convolution; then upsampling is performed on P5 to obtain feature layer one, and the resolution of feature layer one is consistent with that of feature layer C4; the next feature layer C4 is convolved through 1*1 convolution to obtain feature layer two with a preset number of channels; the feature layer one and the feature layer two are added and fused to obtain feature layer P4; and so on, feature layers P5, P4, P3, and P2 are obtained from feature layers C5, C4, C3, and C2. In addition, P5 is further feature extracted to obtain feature layer P6. P2-P6 are used for regional generation network components, and P2-P5 are used for head network components.
[0089] The larger feature map obtained by upsampling lacks edge detail information, and the feature map obtained by pooling operation will inevitably lose some edge features. The ResNet-FPN uses the high resolution of shallow feature maps and the high semantic information of deep feature maps at the same time. Fusion of these feature maps at different layers is equivalent to fusion of strong semantic information and strong edge information, thereby improving the effect of feature extraction.
[0090] Step S1220: using a region generation network to generate a plurality of candidate regions of interest for the multiple layers of feature maps of different scales, and then performing a region of interest alignment operation;
[0091] The region generation network uses a sliding window to slide the four feature maps P2-P6 one by one, and initializes the reference region at the point of each sliding window; that is, the specific coordinates of each corresponding basic anchor frame are calculated according to the coordinates of the sliding window; 1k basic anchor frames are generated on each feature layer. For each basic anchor frame, two confidences are generated, one for the foreground confidence and the other for the background confidence; and 4 coordinate deviation regression values are generated at the same time.
[0092] The region of interest alignment operation RolAlign is an improvement on the region of interest pooling RolPool; the RolPool uses two rounding operations when restoring the feature map to the original image size, which will lead to deviation results for multiple pixel points; while RolAlign directly cancels the rounding operation and obtains the pixel values of the fixed four point coordinates through bilinear interpolation, thereby reducing the restoration error.
[0093] The region generation network can generate multiple candidate regions of interest and is equipped with RolAlign to implement the alignment operation of the regions of interest.
[0094] Step S1230: Use the head network to perform three branch processing on the feature map after alignment of the region of interest, namely, bounding box regression processing, recognition processing and mask map prediction, to obtain a detection result.
[0095] The head network is used to perform the final detection on the feature map after alignment of the above-mentioned region of interest. The detection is divided into three branches. The bounding box regression processing branch and the recognition processing branch use a convolution layer with a convolution depth of 1024 instead of a fully connected layer for prediction, which can make fuller use of feature information; the mask map prediction branch is performed through a fully convolutional network FCN (Fully Convolution Network) to achieve semantic segmentation. Each ROI mask map has 80 categories, which can reduce the competition between categories and thus obtain better results.
[0096] Therefore, the detection results of the head network include the confidence of the target object, and the bounding box and mask map of the target area image indicated by it.
[0097] Step S1240: According to the detection result, intercept the target area in the advertisement image to obtain a corresponding target area image.
[0098] The detection result includes a bounding box indicating the target area image containing the target item in the advertisement image, a mask image and its confidence. In the embodiment of the present application, only the bounding box information in the detection result is called. According to the position of the target area image indicated by the bounding box information, it is cut out from the advertisement image, that is, the same number of target area images is obtained, and the number can be 0, 1, or more in the example application scenario, and the specific value is obtained from the example application.
[0099] The target area image and the advertisement picture are output together to step S1300 for further processing.
[0100] In summary, the target detection model adopts the Mask-RCNN model and uses the ResNet-FPN structure, while fusing the strong semantic information of deep features and the strong edge information of shallow features; the RolAlign operation is used to achieve pixel alignment in the generation of the region of interest to reduce the position deviation estimation of the bounding box; thereby effectively enhancing the detection effect of the target.
[0101] See also Figure 3 In the further example, the step S1300 uses the image recognition model that has been trained to convergence and to which a classification head capable of suppressing weak neurons is added to perform image recognition on the advertisement image and the target area image to obtain a small target recognition sequence, including the following steps:
[0102] Step S1310: embedding an image block for the input image to obtain multiple image block vectors, and adding a classification vector to form multiple embedding vectors;
[0103] The image recognition model is a ViT that has been trained to convergence and to which a classification head that can suppress weak neurons is added. The model is a model that is further improved based on the Transformer model applied to NLP problems and acts on the image field. The first step is to convert images in the image field into word structures in natural language processing, which is called image block embedding. Specifically, the image is standardized to obtain a picture of a standard size, which can be regarded as a complete sentence; then it is divided into small blocks of a fixed size, called Patches, and the pixel values of each small block are flattened to become a word in the sentence. Subsequently, each Patch is compressed into a vector of a certain dimension through a fully connected network to obtain multiple image block vectors. This process is image block embedding, called Patch Embedding. In the example application scenario of the present application, the standard size is 224*224, the fixed size of Patch is 16*16, and the dimension of the vector is 768.
[0104] After obtaining multiple image block vectors, a classification vector of the same dimension is concatenated. As the name implies, this vector is used for category information learning during the model training process. The classification vector is a learnable embedding vector.
[0105] Step S1320: adding a position encoding vector to the multiple embedding vectors to form an input vector, wherein the position encoding vector can maintain spatial position information between image blocks;
[0106] Multiple embedding vectors can be obtained by step S1310, but the embedding vectors lack position information, that is, except for the classification vector, other vectors have their corresponding position information in the picture. Therefore, in order to maintain the spatial position information between each patch in the input image, it is necessary to add a position encoding vector to these multiple embedding vectors. Specifically, a one-dimensional learnable position embedding vector is directly used, and the position embedding vector is directly added to the embedding vector to form an input vector.
[0107] Therefore, the input image has completed the vector embedding work here and can be put into Transformer for training and feature learning.
[0108] Step S1330: stack multiple encoding modules for the input vector to perform feature extraction, wherein the encoding modules include multi-head attention and multi-layer perceptron;
[0109] Multiple encoding modules are stacked for the above input vector, and the encoding module mainly includes multi-head attention and multi-layer perceptron. Specifically, the encoding module includes two parts, the first part is layer normalization (Layer Norm) -> Multi-Head Attention -> Dropout -> Short Cut. The second part is layer normalization (Layer Norm) -> Multi-Layer Perception -> Dropout -> Short Cut. The multi-layer perceptron includes full connection (Linear) -> activation function (GELU) -> Dropout -> full connection (Linear) -> Dropout.
[0110] The layer normalization is to perform normalization on a specified dimension of a single data.
[0111] The feature extraction enables the classification vector to fuse the semantic features of all image patches.
[0112] Step S1340: Use a common classification head to perform further classification space transformation on the finally obtained deep classification vector, and use a new classification head to suppress weak neurons to perform classification space transformation on the deep classification vector, to obtain two classification probabilities;
[0113] See also Figure 4 , using both common classification headers and new classification headers.
[0114] After multiple encoding modules in the above steps, the classification vector has extracted the feature information of each image block in the image. At this time, a common classification head is used to further transform the classification space for the final deep classification vector, that is, layer normalization (Layer Norm) -> fully connected (Linear) output classification probability; in addition, a new classification head is added at the same time, that is, a threshold activation function is connected after the original full connection (Linear), and then a full connection output classification probability is connected, that is: full connection (Linear) -> threshold activation function -> full connection (Linear) output classification probability. The activation function can dynamically activate strong neurons that can characterize the characteristics of the small target object, and on the contrary, it can inhibit weak neurons that interfere with the recognition of the small target; in the embodiment of the present application, the threshold for determining strong and weak neurons is artificially preset, and the preset threshold is set by relevant technical personnel in this field after experimental analysis and experience judgment in actual application. The activation state of the neuron less than the threshold is set to 0, that is, the neuron is inhibited; the activation state of the neuron not less than the threshold is set to 1, that is, the neuron is activated. Next, the activated neurons are connected to the subsequent fully connected layer for classification space transformation to obtain their probability values.
[0115] This step ultimately obtains two probability values, that is, two classification probabilities, both indicating whether the input image contains the target object.
[0116] Step S1350: perform the above steps on the advertisement image and the target area image respectively to obtain a corresponding plurality of recognition results, forming a small target recognition sequence of the advertisement image.
[0117] For the advertisement image and the target area image, the above steps are performed respectively, i.e., the processing of the image recognition model, to obtain multiple corresponding classification probabilities, wherein two classification probabilities can be obtained for each input image. The multiple classification probabilities are combined to form a sequence, called a small target recognition sequence, i.e., a small target recognition sequence of whether the advertisement image contains a small target object, for further processing in the next step.
[0118] In summary, the embodiment of the present application adopts ViT with a classification head that can suppress weak neurons as the network architecture of the image recognition model, divides the image into image blocks, completes image block embedding and position code embedding, and then uses the strong semantic feature extraction method of Transformer to extract the semantic features of the image, and finally realizes the final classification probability prediction by further classification space transformation of the two classification heads. The classification head that can suppress weak neurons can dynamically activate strong neurons that can characterize the deep features of small target items according to the recognition task, which is beneficial to enhancing the generalization ability of the image recognition model, making the image recognition effect significantly better than other image recognition models, thereby obtaining better detection results.
[0119] See also Figure 5 In the in-depth example, step S1400, using the small target recognition sequence to output the corresponding small target recognition result according to the preset conditions of the specific example scene, includes the following steps:
[0120] Step S1410, sorting the small target recognition sequence to obtain its maximum value and minimum value;
[0121] The small target recognition sequence includes multiple classification probabilities, and the value of each classification probability ranges from 0 to 1. The target recognition sequence is directly composed of multiple probability values obtained by image recognition of multiple input images, and its arrangement is disordered. Therefore, first sort the target recognition sequence so that the classification probabilities in the sequence are arranged from large to small. Then obtain its maximum and minimum values:
[0122] Prob_min=min(VitList)
[0123] Prob_max=max(VitList)
[0124] Among them, VitList represents the small target recognition sequence; Prob_min represents the minimum value in the small target recognition sequence; Prob_max represents the maximum value in the small target recognition sequence.
[0125] Step S1420, identifying the actual needs of the instance application scenario, using the minimum value as the probability value of whether the advertisement image contains the small target object in a high accuracy scenario, and using the maximum value as the probability value of whether the advertisement image contains the small target object in a high recall scenario;
[0126] The actual needs of the identification instance application scenario are different in different application scenarios. Specifically, in the embodiment of the present application, in the instance application scenario where high accuracy is required, the minimum value in the small target recognition sequence is used as the final probability value of whether the advertising image contains the target item; in the instance application scenario where high recall is required, the maximum value in the small target recognition sequence is used as the final probability value of whether the advertising image contains the target item.
[0127] Step S1430: compare the probability value with a preset threshold value, and when the probability value is greater than the preset threshold value, determine that the advertisement image contains the small target object; otherwise, determine that the advertisement image does not contain the small target object.
[0128] According to the previous step, the probability value in the example application scenario is obtained, and the probability value is compared with the preset threshold value. If the probability value is greater than the preset threshold value, it is determined that the target item is included in the advertisement image; if the probability value is not greater than the preset threshold value, it can be determined that the target person is not included in the advertisement image. In the example application scenario, further business processing can be performed on the advertisement image according to the determination result.
[0129] The preset threshold is a dividing value for determining whether the probability value indicates that the advertising image contains a small target item. The setting of this value directly affects the accuracy of the determination. If it is set too high, the advertising image that does not contain the small target item will be determined to contain it. If it is set too low, the advertising image that contains the small target item will be determined to not contain it. Therefore, the threshold needs to be set by relevant technical personnel based on the comparison of test results and the use of prior knowledge. In summary, the embodiment of the present application selects different probability values as the final probability estimate of whether the advertising image contains the target item and makes a final determination based on different example application scenarios for the obtained small target recognition sequence. The implementation of the steps can effectively improve the accuracy of target item recognition of the e-commerce platform under specific business requirements in specific application scenarios, and help the e-commerce platform to more effectively exclude advertising images that violate relevant regulations.
[0130] See also Figure 6 A small target detection device provided to meet one of the purposes of the present application includes an image acquisition module 1100, a target detection module 1200, an image recognition module 1300 and a target discrimination module 1400, wherein the image acquisition module 1100 is used to obtain an advertising image from the advertising information; the target detection module 1200 is configured to use a target detection model that has been trained to convergence to perform target detection on the advertising image to obtain a target area image; the image recognition module 1300 is configured to use an image recognition model that has been trained to convergence and to which a classification head that can inhibit weak neurons is added to perform image recognition on the advertising image and the target area image to obtain a small target recognition sequence; the target discrimination module 1400 is used to use the small target recognition sequence to output a corresponding small target recognition result according to the preset conditions of a specific instance scene.
[0131] In a further embodiment, the image acquisition module 1100 includes: a response submodule, used to respond to an advertising publishing request triggered by a user, and obtain the corresponding advertising publishing information submitted, wherein the advertising publishing information includes advertising pictures; an acquisition submodule; used to obtain the advertising pictures from the advertising publishing information, and used to identify the pictures of small target objects to be detected.
[0132] In the further example, the target detection module 1200 includes: a convolutional backbone submodule, which uses a convolutional backbone network to perform feature extraction on the advertising image to obtain multiple layers of feature maps of different scales; a region of interest submodule, which uses a region generation network to generate multiple candidate regions of interest for the multiple layers of feature maps of different scales, and then performs a region of interest alignment operation; a detection submodule, which uses a head network to perform three branch processing on the feature map after the region of interest alignment, namely, bounding box regression processing, recognition processing and mask map prediction, to obtain a detection result; a capture submodule, which is used to capture the target area in the advertising image according to the detection result to obtain a corresponding target area image.
[0133] In the further example, the image recognition module 1300 includes an embedding submodule, which is used to embed image blocks for the input image, obtain multiple image block vectors, and add a classification vector to form multiple embedding vectors; a position encoding submodule, which is used to add a position encoding vector to the multiple embedding vectors to form an input vector, and the position encoding vector can maintain the spatial position information between image blocks; a feature extraction submodule, which is used to stack multiple encoding modules for feature extraction on the above-mentioned input vector, and the encoding module includes multi-head attention and multi-layer perceptron; a classification submodule, which uses an ordinary classification head to perform further classification space transformation on the final deep classification vector, and uses a newly added classification head to suppress weak neurons to perform classification space transformation on the deep classification vector to obtain two classification probabilities; a sequence generation submodule, which performs the above steps on the advertising picture and the target area image respectively, obtains corresponding multiple recognition results, and constitutes a small target recognition sequence of the advertising picture.
[0134] In the further example, the target discrimination module 1400 includes: a sorting submodule, used to sort the small target recognition sequence to obtain its maximum and minimum values; a probability value calculation submodule, used to identify the actual needs of the example application scenario, and in a high-accuracy scenario, the minimum value is used as the probability value of whether the advertising image contains the small target item, and in a high-recall scenario, the maximum value is used as the probability value of whether the advertising image contains the small target item; a discrimination submodule, used to compare the probability value with a preset threshold value, and when the probability value is greater than the preset threshold value, it is determined that the advertising image contains the small target item, otherwise it is determined that the small target item is not contained.
[0135] In a preferred embodiment, the target detection model is a Mask-RCNN model, the image recognition model is a ViT model with a classification head that can suppress weak neurons added, the basic network architecture of the target detection model is a Mask-RCNN model, the basic network architecture of the image recognition model is a ViT model with a classification head that can suppress weak neurons added, and the target object is a cigarette.
[0136] In order to solve the above technical problems, the present application also provides a computer device. Figure 7 As shown, a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database may store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a small target detection method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the small target detection method of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0137] In this embodiment, the processor is used to execute Figure 6 The memory stores the program code and various data required to execute the above modules or submodules. The network interface is used to transmit data between user terminals or servers. The memory in this embodiment stores the program code and data required to execute all modules / submodules in the small target detection device of the present application, and the server can call the program code and data of the server to execute the functions of all submodules.
[0138] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the small target detection method of any embodiment of the present application.
[0139] The present application also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any embodiment of the present application when executed by one or more processors.
[0140] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments of the present application can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0141] In summary, this application uses a target detection model to perform target detection on advertising images, obtains target area images, and then performs image recognition on the advertising images and target area images respectively to obtain target recognition sequences, and then selects different probability values as the final probability estimate of whether the advertising image contains the target object according to different instance application scenarios. This greatly improves the accuracy of target object recognition, helps e-commerce platforms to more effectively exclude advertising images that violate relevant regulations, and greatly reduces time and labor costs.
[0142] This application uses enhanced detection of detection, recognition and scene conditions, and uses an image recognition model with a classification head that can suppress weak neurons to identify small target objects. It achieves the decoupling of features between different categories of small targets, and uses hierarchical condition restrictions to further improve the accuracy of small target object recognition, solving the business needs of high accuracy and high recall in the application scenarios of e-commerce platforms. Ultimately, it helps e-commerce platforms to more effectively exclude advertising images that violate relevant regulations, while greatly reducing time and labor costs.
[0143] Those skilled in the art will appreciate that the various operations, methods, steps, measures, and schemes in the processes discussed in this application may be alternated, altered, combined, or deleted. Further, other steps, measures, and schemes in the various operations, methods, and processes discussed in this application may also be alternated, altered, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and schemes in the prior art that are similar to those disclosed in this application may also be alternated, altered, rearranged, decomposed, combined, or deleted.
[0144] The above description is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A small target detection method, characterized in that: The steps include: Get the advertisement image to be detected; Using a target detection model that has been trained to convergence to perform target detection on the advertisement image to obtain a target area image; An image recognition model that has been trained to convergence and to which a classification head capable of suppressing weak neurons is added is used to perform image recognition on the advertising image and the target area image to obtain a small target recognition sequence, which includes: performing image block embedding on the input image to obtain multiple image block vectors, and adding a classification vector to form multiple embedding vectors; adding position encoding vectors to the multiple embedding vectors to form an input vector, and the position encoding vector can maintain the spatial position information between image blocks; stacking multiple encoding modules for feature extraction on the above input vector, and the encoding modules include multi-head attention and multi-layer perceptron; performing further classification space transformation on the finally obtained deep classification vector using an ordinary classification head, and performing classification space transformation on the deep classification vector using a newly added classification head to suppress weak neurons to obtain two classification probabilities; performing the above steps on the advertising image and the target area image respectively to obtain multiple corresponding recognition results to form a small target recognition sequence for the advertising image; The small target recognition sequence is used to output a corresponding small target recognition result according to the preset conditions of a specific instance scenario, and the specific instance scenario is a high-accuracy scenario or a high-recall scenario, which includes: sorting the small target recognition sequence to obtain its maximum and minimum values; identifying the actual needs of the instance application scenario, using the minimum value as the probability value of whether the advertising image contains a small target item in the high-accuracy scenario, and using the maximum value as the probability value of whether the advertising image contains a small target item in the high-recall scenario; comparing the probability value with a preset threshold value, and when the probability value is greater than the preset threshold value, determining that the advertising image contains a small target item, otherwise determining that it does not contain a small target item.
2. The small target detection method according to claim 1, characterized in that: Obtaining the advertisement image to be detected includes the following steps: Responding to an advertisement publishing request triggered by a user, obtaining the corresponding advertisement publishing information submitted by the user, wherein the advertisement publishing information includes an advertisement picture; The advertisement picture is obtained from the advertisement publishing information.
3. The small target detection method according to claim 1, characterized in that: Using a target detection model that has been trained to convergence to perform target detection on the advertisement image to obtain a target area image, the process includes the following steps: A convolutional backbone network is used to extract features from the advertisement image to obtain multi-layer feature maps of different scales; A region generation network is used to generate a plurality of candidate regions of interest for the multi-layer feature maps of different scales, and then an alignment operation of the regions of interest is performed; The head network is used to perform three branch processing on the feature map after alignment of the region of interest, namely bounding box regression processing, recognition processing and mask map prediction, to obtain the detection result; According to the detection result, the target area in the advertisement picture is captured to obtain a corresponding target area image.
4. The small target detection method according to any one of claims 1 to 3, characterized in that: The basic network architecture of the target detection model is the Mask-RCNN model, the basic network architecture of the image recognition model is the ViT model with a classification head added that can suppress weak neurons, and the small target object is a cigarette.
5. A small target detection device, characterized in that: include: An image acquisition module is used to acquire the advertisement image to be detected; A target detection module is configured to perform target detection on the advertisement image using a target detection model that has been trained to convergence, so as to obtain a target area image; The image recognition module is configured to use an image recognition model that has been trained to convergence and to which a classification head capable of suppressing weak neurons is added to perform image recognition on the advertising image and the target area image to obtain a small target recognition sequence, which includes: performing image block embedding on the input image to obtain multiple image block vectors, and adding a classification vector to form multiple embedding vectors; adding position encoding vectors to the multiple embedding vectors to form an input vector, and the position encoding vector can maintain the spatial position information between image blocks; stacking multiple encoding modules for feature extraction on the above input vector, and the encoding module includes multi-head attention and multi-layer perceptron; using an ordinary classification head to further perform classification space transformation on the final deep classification vector, and using a newly added classification head to suppress weak neurons to perform classification space transformation on the deep classification vector to obtain two classification probabilities; performing the above steps on the advertising image and the target area image respectively to obtain corresponding multiple recognition results to form a small target recognition sequence for the advertising image; The target discrimination module is used to use the small target recognition sequence to output the corresponding small target recognition result according to the preset conditions of a specific instance scenario, wherein the specific instance scenario includes a high-accuracy scenario and a high-recall scenario, and includes: sorting the small target recognition sequence to obtain its maximum and minimum values; identifying the actual needs of the instance application scenario, using the minimum value as the probability value of whether the advertising image contains a small target item in the high-accuracy scenario, and using the maximum value as the probability value of whether the advertising image contains a small target item in the high-recall scenario; comparing the probability value with a preset threshold value, and when the probability value is greater than the preset threshold value, determining that the advertising image contains a small target item, otherwise determining that it does not contain a small target item.
6. The small target detection device according to claim 5, characterized in that: The image acquisition module comprises: A response submodule is used to respond to the advertisement publishing request triggered by the user and obtain the corresponding advertisement publishing information submitted by the user, wherein the advertisement publishing information includes an advertisement picture; The acquisition submodule is used to acquire the advertisement pictures from the advertisement publishing information, and is used to identify the pictures of the small target objects to be detected.
7. The small target detection device according to claim 5, characterized in that: The target detection module comprises: The convolutional backbone submodule uses a convolutional backbone network to extract features from the advertisement image to obtain multi-layer feature maps of different scales; The region of interest submodule uses a region generation network to generate multiple candidate regions of interest for the multi-layer feature maps of different scales, and then performs a region of interest alignment operation; The detection submodule uses the head network to perform three branch processing on the feature map after alignment of the region of interest, namely bounding box regression processing, recognition processing and mask map prediction, to obtain the detection result; The capture submodule is used to capture the target area in the advertisement image according to the detection result to obtain a corresponding target area image.
8. A computer device comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 4.
9. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 4 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Weak texture object pose estimation method and system
CN113538569A
Image acquisition and processing methods for automatic vehicular exterior lighting control
US20040143380A1