An image recognition method and related device
By using a feature extraction model in the image recognition method for multiple downsampling and feature fusion, and dividing the downsampling layer with a size threshold, the problem of insufficient recognition accuracy for objects of different sizes in the prior art is solved, and efficient multi-size object detection is achieved.
Patent Information
- Application Number
- CN202110025048.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-01-08
AI Technical Summary
Existing network models are difficult to accurately identify specific objects of different sizes, especially in the fields of video and live broadcast, where the faces to be identified vary in size.
An image recognition method is adopted to extract N+M downsampled image feature through feature extraction model, fuse shallow and deep feature maps, and divide the downsampling layers with size thresholds to ensure the detection accuracy of objects of different sizes.
It realizes the balance of objects of different sizes, improves the detection accuracy of image recognition, and is suitable for a variety of application scenarios, including video and live broadcast.
Smart Images

Figure CN113536876B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to an image recognition method and related device. Background Art
[0002] Through a network model, specific objects such as human faces included in an image can be detected and recognized. This way of image object detection can provide reference data for many application scenarios.
[0003] However, in many scenarios, the images to be recognized have specific objects of various sizes. For example, in the fields of video and live broadcast, the human faces to be recognized are often of different sizes. Currently, most network models have a better recognition rate for specific objects of a certain size and are difficult to accurately recognize specific objects of different sizes. Summary of the Invention
[0004] To solve the above technical problems, this application provides an image recognition method and related device, which improves the recognition accuracy for objects of different sizes.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] On the one hand, the embodiments of this application provide an image recognition method, which includes:
[0007] Obtain a target image with an object to be recognized;
[0008] Perform image feature extraction on the target image according to a feature extraction model including N downsampling layers connected in sequence and M downsampling layers to obtain feature maps corresponding to the downsampling layers respectively. The size of the feature map output by the first downsampling layer in the M downsampling layers is smaller than a size threshold; the i-th downsampling layer and the (i + 1)-th downsampling layer in the N downsampling layers are adjacent downsampling layers, and the initial output feature of the i-th downsampling layer is fused with the initial output feature of the (i + 1)-th downsampling layer to obtain the feature map of the i-th downsampling layer;
[0009] Determine candidate detection frames corresponding to feature maps of different sizes according to the feature maps respectively determined by the downsampling layers. The candidate detection frames are used to identify the regions of the object to be recognized and the regions of non-objects to be recognized in the target image;
[0010] Determine the detection result of the object to be recognized for the target image according to the candidate detection frames.
[0011] On the other hand, the embodiments of this application provide an image recognition device, which includes an acquisition unit, an extraction unit, and a determination unit:
[0012] The obtaining unit is configured to obtain a target image having an object to be recognized;
[0013] The extraction unit is configured to perform image feature extraction on the target image according to a feature extraction model including N downsampling layers and M downsampling layers connected in sequence, to obtain feature maps corresponding to the downsampling layers respectively, and the size of the feature map output by the first downsampling layer in the M downsampling layers is smaller than a size threshold; the i-th downsampling layer and the (i + 1)-th downsampling layer in the N downsampling layers are adjacent downsampling layers, and the initial output feature of the i-th downsampling layer and the initial output feature of the (i + 1)-th downsampling layer are fused to obtain the feature map of the i-th downsampling layer;
[0014] The determination unit is configured to determine candidate detection frames corresponding to feature maps of different sizes according to the feature maps respectively determined by the downsampling layers, and the candidate detection frames are used to identify the region of the object to be recognized and the region of non-object to be recognized in the target image;
[0015] The determination unit is further configured to determine a detection result of the object to be recognized for the target image according to the candidate detection frames.
[0016] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory:
[0017] The memory is configured to store program code and transmit the program code to the processor;
[0018] The processor is configured to execute the method described in the above aspect according to the instructions in the program code.
[0019] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which is configured to store a computer program, and the computer program is used to execute the method described in the above aspect.
[0020] On the other hand, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described in the above aspect.
[0021] As can be seen from the above solution, for a target image with an object to be recognized, the feature extraction model is used to perform N+M times of downsampled image feature extraction on its sequence to obtain the feature maps corresponding to the downsampling layers respectively. Since the detection of small-sized objects requires the feature map to have a high resolution, and the feature maps obtained by the shallow downsampling layers of the first N layers can be used as the basis for recognizing small-sized objects, the initial output features of the i-th downsampling layer and the (i+1)-th downsampling layer among these N downsampling layers are fused. Based on the strong semantic information that can be obtained from the target image by the (i+1)-th downsampling layer, it is supplemented into the features extracted by the i-th downsampling layer to obtain the feature map of the i-th downsampling layer, enhancing the semantic information in the feature map of the i-th downsampling layer and ensuring the detection accuracy of small-sized objects based on this feature map. In addition, since the detection of large-sized objects requires the feature map to have strong semantic information and does not require a high resolution, therefore, based on a size threshold, the downsampling layers of the feature extraction model are divided, and the last M downsampling layers are determined therefrom. This size threshold ensures that the feature maps output by these M downsampling layers have strong semantic information, ensuring the detection accuracy of large-sized objects based on this feature map. In the process of determining the candidate detection boxes of the target image, the candidate detection boxes are respectively recognized based on the feature maps output by the foregoing respective downsampling layers. In this way, the consideration of objects to be recognized with different sizes is realized, playing a role in adapting to objects to be recognized with different sizes in the target image and ensuring the detection accuracy of objects with different sizes in the target image. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 FIG. is a schematic diagram of an application scenario of an image recognition method provided by an embodiment of the present application;
[0024] Figure 2 FIG. is a schematic flowchart of an image recognition method provided by an embodiment of the present application;
[0025] Figure 3 FIG. is a schematic diagram of another application scenario of an image recognition method provided by an embodiment of the present application;
[0026] Figure 4 FIG. is a schematic flowchart of another image recognition method provided by an embodiment of the present application;
[0027] Figure 5A schematic flowchart of a multi-classification recognition method using a maxout network with maximum output provided by an embodiment of the present application;
[0028] Figure 6 A schematic flowchart of a model training method provided by an embodiment of the present application;
[0029] Figure 7 A schematic flowchart of an image recognition device provided by an embodiment of the present application;
[0030] Figure 8 A schematic structural diagram of a server provided by an embodiment of the present application;
[0031] Figure 9 A schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0032] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0033] In the field of image recognition, the same deep model cannot balance the recognition performance for objects of different sizes in an image. Taking the Single Slot Scale-invariant Face Detector (S3FD) as an example, the S3FD model directly uses the features extracted from different layers for prediction, resulting in weak semantic information in the shallow features. Since small-sized objects are mainly detected in the shallow layers, the S3FD model has poor recognition effects on small-sized objects. For the Dual Shot Face Detector (DSFD) with high recognition accuracy for small-sized faces, due to the large number of model parameters of the DSFD model, the detection speed of DSFD is slow, which directly affects the user experience in real-time detection scenarios such as video and live broadcast face detection scenarios.
[0034] In view of this, the embodiments of the present application provide an image recognition method and related devices, which achieve the balance of objects to be recognized of different sizes, play the role of adapting to objects to be recognized of different sizes in the target image, and ensure the detection accuracy of objects of different sizes in the target image.
[0035] The image recognition method provided by the embodiments of this application is implemented based on artificial intelligence. Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0036] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0037] In the embodiments of this application, the artificial intelligence software technologies mainly involved include the above-mentioned computer vision technology, machine learning / deep learning, etc. For example, it can involve image processing, image semantic understanding (ISU), image recognition (IR), etc. in computer vision (Computer Vision). For example, it can involve deep learning in machine learning (Machine learning, ML), including various artificial neural networks (Artificial Neural Network, ANN).
[0038] The image recognition method provided by the embodiments of this application can be applied to image recognition devices with data processing capabilities, such as terminal devices or servers. This method can be independently executed by the terminal device, independently executed by the server, or applied to a network scenario where the terminal device and the server communicate, and is executed in cooperation with the terminal device and the server. Among them, the terminal device can be a mobile phone, a desktop computer, a portable computer, etc.; the server can be understood as an application server or a Web server. In actual deployment, the server can be an independent server or a cluster server. For the convenience of description, the following introduces the embodiments of this application with the server as the data processing device.
[0039] The image recognition device can be equipped with the ability to implement computer vision technology. Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. It also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0040] In the embodiments of the present application, the image recognition device can perform image processing, image recognition, etc. on the target image through computer vision technology.
[0041] The image recognition device can be equipped with machine learning capabilities. Machine learning is a multi-disciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks.
[0042] The image recognition method provided by the embodiments of the present application mainly involves the application of various artificial neural networks to identify the objects to be recognized in the target image.
[0043] For the sake of easy understanding, the following will introduce the image recognition method provided by the embodiments of the present application in combination with specific application scenarios.
[0044] See Figure 1 , Figure 1 which is a schematic diagram of the application scenario of an image recognition method provided by the embodiments of the present application. In the scenario shown in Figure 1 , it includes a server 101, on which a feature extraction model 102 is deployed, which is used to execute the image recognition method provided by the embodiments of the present application. Among them, the feature extraction model 102 includes N downsampling layers and M downsampling layers connected in sequence.
[0045] The server 101 obtains a target image with an object to be recognized. In Figure 1In the face recognition scenario shown, the object to be recognized included in the target image 103 is a human face.
[0046] During the image recognition process, the server 101 performs N+M times of image feature extraction on the target image by calling the trained feature extraction model, and obtains the feature maps corresponding to the downsampling layers respectively.
[0047] Since during the image recognition process, the detection of small-sized objects requires the feature map to have a high resolution, and the feature maps obtained by the first N shallow downsampling layers can be used as the basis for recognizing small-sized objects. Therefore, the initial output features of the i-th downsampling layer and the (i+1)-th downsampling layer among these N downsampling layers are fused to obtain the feature map of the i-th downsampling layer. Among them, the initial output feature of the N-th sampling layer is directly used as the feature map output.
[0048] Taking N = 4 as an example, as Figure 1 shown, for the first sampling layer among the first 4 downsampling layers, that is, taking i = 1, the initial output features of the first downsampling layer and the second downsampling layer are fused to obtain the feature map A of the first downsampling layer. After performing image feature extraction on the target image using the first 4 downsampling layers based on this process, the corresponding feature maps A, B, C, and D are respectively output.
[0049] The above process of fusing two downsampling layers is to add the strong semantic information that the (i+1)-th downsampling layer can obtain in the target image to the features extracted by the i-th downsampling layer, thereby enhancing the semantic information in the feature map of the i-th downsampling layer and ensuring the detection accuracy of small-sized objects based on this feature map.
[0050] Since the detection of large-sized objects requires the feature map to have strong semantic information and does not require a high resolution, the downsampling layers of the feature extraction model are divided based on a size threshold to determine M downsampling layers. This size threshold ensures that the feature maps output by these M downsampling layers have strong semantic information and ensures the detection accuracy of large-sized objects based on this feature map. As Figure 1 shown, if M = 2, then after continuing to perform image feature extraction using the last 2 downsampling layers, the corresponding feature maps E and F are respectively output.
[0051] Based on the feature maps respectively output by the above N+M downsampling layers, the corresponding candidate detection frames are determined, thereby determining the detection result of the object to be recognized in the target image. As Figure 1As shown, based on the feature maps A, B, C, D, E, and F output by each of the six downsampling layers included in the feature extraction model 102, the corresponding candidate detection frames a, b, c, d, e, and f are determined respectively, and then the face in the target image 103 is identified based on these candidate detection frames, and the position of the face in the target image 103 is determined.
[0052] The feature maps output by the N+M downsampling layers of the above feature extraction model have different sizes, such as Figure 1 The feature maps A, B, C, D, E and F shown are feature maps of different sizes, among which the feature maps output by the first N downsampling layers have strong resolution, and their semantic information is enhanced by feature fusion, ensuring the recognition of small-sized objects. The feature maps output by the last M downsampling layers have strong semantic information, and their high resolution is guaranteed by the size threshold, ensuring the recognition of large-sized objects. Therefore, candidate detection frames are identified based on these feature maps, respectively, and objects of different sizes to be identified are taken into account, which plays a role in adapting to objects of different sizes to be identified in the target image and ensures the detection accuracy of objects of different sizes in the target image.
[0053] Combine the following Figure 2 and Figure 3 The image recognition method provided in the embodiment of the present application is described in detail, wherein: Figure 2 A flowchart of an image recognition method provided in an embodiment of the present application is shown in FIG. Figure 3 FIG. 1 is a flowchart of an image recognition method in a face recognition scenario. Figure 2 As shown, the image recognition method includes the following steps:
[0054] S201: Acquire a target image having an object to be identified.
[0055] In the image recognition process, the server first obtains the target image to be recognized, which has the object to be recognized. The object to be recognized refers to the target object to be recognized, which can be objects of different categories, such as houses, cats, faces, etc., or different objects of the same category, such as male faces, female faces, etc. Figure 3 In the face scene shown, the object to be recognized is a face.
[0056] Before performing image recognition on the target image, the pixel values of the target image can be normalized, that is, the pixel values of the target image are transformed from the interval [0,255] to [-1,1] to improve the image recognition speed. Figure 4 The image data preprocessing 401 process is shown.
[0057] S202: Extract image features from the target image according to a feature extraction model including N downsampling layers and M downsampling layers connected in sequence, to obtain feature maps corresponding to the downsampling layers respectively.
[0058] This application performs image recognition of an object to be recognized on a target image based on a neural network model in artificial intelligence technology. The neural network model framework adopts a single-stage detection network combined with a backbone network. The single-stage detection network is used to generate and recognize candidate detection boxes, such as the Single Shot Multi-box Detector (SSD) model. The backbone network is used to extract image features of the object to be recognized from the target image, such as the MobileNetV2 model. Among them, the candidate detection box is a detection box generated according to the target image by a rectangular box for detecting the object to be recognized.
[0059] The aforementioned S3FD model uses the Visual Geometry Network-16 (VGG16) as the backbone network, while the DSFD model uses the Deep residual network-50 (ResNet50) as the backbone network and combines a feature enhancement module. Since the recognition speed of the model mainly depends on the computing speed of the backbone network, the floating-point operation amount of the VGG16 model is 16 GFLOPs, and that of ResNet50 is GFLOPs. In addition, the feature enhancement module adopted by the DSFD model doubles the number of parameters on the basis of the original model, which makes the detection speeds of the S3FD model and the DSFD slow and not applicable to real-time detection scenarios such as live broadcast and video. Among them, GFLOPs is Giga FloatingPoint of Operations, which is used to measure the model complexity.
[0060] Since the model complexities of VGG16 and ResNet50 are relatively high, that is, they are heavyweight models. To improve the image recognition speed, in the embodiments of this application, the backbone network adopts a lightweight feature extraction model. This feature extraction model includes a feature extraction model with N downsampling layers (basic convolutional layers) and M downsampling layers (extra convolutional layers) connected in sequence, and extracts features from the target image. Among them, the basic convolutional layer can be the MobileNetV2 model, and MobileNetV2 is only 0.57 GFLOPs. Thus, by adopting the lightweight feature extraction model, the model complexity is reduced and the image recognition rate is improved.
[0061] In practical applications, use the feature extraction model to extract image features from the target image to obtain feature maps corresponding to the downsampling layers respectively. Such as Figure 3As shown, the feature extraction model 300 includes a basic convolutional layer 301 and an additional convolutional layer 302.
[0062] It can be understood that for the detection of small-sized objects, a feature map with a higher resolution is required as a basis. The initial output features obtained by the basic convolutional layer including N shallow downsampling layers in the feature extraction model for feature extraction of the target image have a higher resolution. Therefore, it can be used to identify small-sized objects in the target image. Here, N is an integer greater than 1. In Figure 3 the scenario shown, the basic convolutional layer 301 includes downsampling layers C1, C2, C4, and C4, that is, N = 4.
[0063] Considering that the semantic information of the initial output features of the N shallow downsampling layers is weak, in the embodiments of the present application, feature enhancement is performed through feature fusion. That is, the neural network model for image recognition in the embodiments of the present application further includes a feature fusion layer for performing feature fusion on the initial output features of the basic convolutional layer.
[0064] Specifically, for the adjacent i-th downsampling layer and the (i + 1)-th downsampling layer among the N downsampling layers, the initial output features of the i-th downsampling layer and the initial output features of the (i + 1)-th downsampling layer are fused to obtain a fused feature as the feature map of the i-th downsampling layer, that is, the feature map obtained by fusing the feature map output by the i-th downsampling layer and the feature map output by the (i + 1)-th downsampling layer is used as the feature map of the i-th downsampling layer. For the N-th downsampling layer, the initial output features are directly output as the feature map of the N-th downsampling layer.
[0065] In practical applications, the feature map output by the (i + 1)-th downsampling layer can be upsampled first. After upsampling, the feature map output by the i-th downsampling layer has the same size as the feature map output by the (i + 1)-th downsampling layer, which is convenient for feature fusion. The upsampling operation can be a bilinear interpolation operation on the feature map output by the (i + 1)-th downsampling layer. Then, it is added to the feature map output by the i-th downsampling layer and summed, and then a non-linear process is performed using an activation layer, such as the Rectified Linear Unit (ReLU), to avoid gradient explosion during model training, improve the generalization ability of the model, and obtain the feature map of the i-th downsampling layer. It is expressed by the formula as follows:
[0066] Fconv i = ReLU(conv i + upsample(conv i+1 ))
[0067] where the value of i is 1, 2,..., N - 1, conv idenotes the initial output feature of the i-th downsampling layer, conv i+1 denotes the initial output feature of the (i + 1)-th downsampling layer, upsample denotes the bilinear interpolation operation, ReLU denotes the rectified linear unit function, Fconv i denotes the feature map corresponding to the i-th downsampling layer.
[0068] As Figure 3 shown, the target image 304 is input into the basic convolutional layer 301. After the basic convolutional layer 301 extracts shallow image features, feature fusion is performed through the feature fusion layer 304, and the feature maps V1, V2, and V3 corresponding to 4 downsampling layers are output. Among them, the 4th downsampling layer C4 in the basic convolutional layer 301 directly outputs the initial output feature as the feature map V4 corresponding to this layer.
[0069] The above feature extraction model is a lightweight backbone network. Compared with the heavyweight backbone network, it reduces the model complexity and improves the image recognition speed. At the same time, the initial output features of two adjacent downsampling layers in the first N downsampling layers are concatenated, increasing the receptive field of the basic extraction layer and enhancing the semantic information of the feature maps corresponding to the first N - 1 shallow downsampling layers, ensuring the detection accuracy for small-size objects in the target image.
[0070] In addition, the detection of large-size objects requires feature maps with strong semantic information as a basis. For this reason, the above feature extraction model adds M downsampling layers on the basis of N downsampling layers. That is, an additional convolutional layer is added after the basic convolutional layer. Since the feature maps output by the deep downsampling layers have strong semantic information and the feature maps with strong semantic information have a small size. Therefore, the last M downsampling layers are determined by the size threshold, that is, the size of the feature map output by the first downsampling layer in the M downsampling layers is less than the size threshold, ensuring that the M deep downsampling layers have strong semantic information, thus ensuring the detection accuracy for large-size objects in the target image.
[0071] In Figure 3 the shown scenario, the additional convolutional layer 302 includes downsampling layers C52 and C62, that is, M = 2. The additional convolutional layer 302 takes the initial output feature C51 of the 4th downsampling layer C5 of the basic convolutional layer 301 as the input, and uses the downsampling layers C52 and C62 to continue to extract image features from the target image 304, and outputs the feature maps V5 and V6 respectively. Among them, C61 is the initial output feature of the downsampling layer C52.
[0072] By ensuring the sizes of the output feature maps of the M deep downsampling layers through the size threshold, that is, ensuring that the output feature maps of the M deep downsampling layers have strong semantic information, the detection accuracy of large-size objects in the target image is guaranteed. Since identifying large-size objects in the target image does not require high-resolution feature maps, the additional convolutional layers in the feature extraction model can directly output feature maps without feature fusion. Thus, by combining the additional convolutional layers on the basis of the basic convolutional layer, the detection accuracy of both large-size and small-size objects is guaranteed, and the lightweight design of the feature extraction model is achieved, thereby improving the image recognition speed.
[0073] Based on the fact that the N+M feature maps output by the above feature extraction model have different sizes, they can be used as the basic data for subsequent image recognition. In a possible implementation, the ratio of the sizes of the feature maps output by adjacent downsampling layers in the feature extraction model is 1 / 2*1 / 2, that is, for adjacent i-th and (i+1)-th downsampling layers, the size of the feature map output by the (i+1)-th downsampling layer is 1 / 4 of the size of the feature map output by the i-th downsampling layer.
[0074] In Figure 3 In the shown scenario, the size of the feature map corresponding to each downsampling layer is 1 / 2*1 / 2 of the size of the feature map corresponding to the previous downsampling layer. For example, the size of feature map B is 1 / 2*1 / 2 of the size of feature map A, and the size of feature map F is 1 / 2*1 / 2 of the size of feature map E. Among them, the cumulative stride of the first layer of convolution is only (4*4), which guarantees the detection effect for multi-size objects.
[0075] The above feature extraction model uses a lightweight backbone network as the basic convolutional layer to extract image features from the target image, and enhances the semantic information of the feature maps corresponding to the shallow downsampling layers through feature fusion, ensuring the detection accuracy for small-size objects. In addition, by adding additional convolutional layers, deep feature extraction is performed on the target image to obtain feature maps with strong semantic information, ensuring the detection accuracy for large-size objects. Thus, the detection accuracy of the model for multi-size objects and the lightweight design of the model are achieved, thereby improving the image recognition rate.
[0076] S203: Determine the candidate detection frames corresponding to the feature maps of different sizes according to the feature maps respectively determined by the downsampling layer.
[0077] Based on the N+M feature maps determined by the N+M downsampling layers of the above feature extraction model, the candidate detection boxes corresponding to these N+M feature maps of different sizes are respectively determined. The candidate detection boxes are used to identify the regions of the object to be recognized and the regions other than the object to be recognized in the target image. Among them, the regions in the target image that only include objects other than the object to be recognized or do not include the object to be recognized in the target image are used as negative samples, and the regions in the target image that include the object to be recognized are used as positive samples. The size of the candidate detection box can be preset according to the image recognition scenario and is not limited here.
[0078] It should be noted that, in order to realize the detection of multi-size objects, in practical applications, rectangular boxes of multiple sizes can be designed, and candidate detection boxes of multiple sizes are respectively generated according to the above-obtained N+M feature maps of different sizes, that is, each feature map corresponds to candidate detection boxes of multiple sizes. Based on the candidate detection boxes, the object to be recognized in the target image is recognized, realizing an adaptive multi-size object candidate detection box generation mechanism, and ensuring the detection accuracy of multi-size objects.
[0079] The above neural network model for image recognition in the embodiments of the present application further includes a prediction layer, which is used to generate and recognize candidate detection boxes, and obtain the recognition results corresponding to the candidate detection boxes. The recognition results identify the category of the candidate detection box and the predicted position of the object to be recognized. That is Figure 4 The processing process of forward calculation 402 using the neural network model is shown.
[0080] In practical applications, the prediction layer takes the feature maps corresponding to the N+M downsampling layers of the above feature extraction model as inputs, performs convolution operations on these N+M feature maps using a fully convolutional network, generates candidate detection boxes for the target image, and then performs image recognition on the candidate detection boxes for the object to be recognized respectively, obtaining the recognition confidence corresponding to the candidate detection box for the object to be recognized and the predicted position of the object to be recognized. Among them, the convolution kernel used by the fully convolutional network can be 3*3.
[0081] It can be understood that among the candidate detection boxes generated by the same-size rectangular box according to feature maps of different sizes, the candidate detection boxes determined based on the feature maps with larger sizes include more negative samples. Since the imbalance in the number of positive and negative samples will affect the image recognition accuracy, in the embodiments of the present application, the maxout network is used to perform multi-classification recognition on the candidate detection boxes generated by the large-size feature maps, so as to reduce the impact of the imbalance in the number of positive and negative samples on the image recognition results.
[0082] Specifically, for the target downsampling layer, the feature map of the target downsampling layer is classified and recognized based on at least three classifications. According to the recognition result of this classification and recognition, the candidate detection box corresponding to the feature map of the target downsampling layer is determined. Among them, the target downsampling layer is any one of the first k downsampling layers among the N downsampling layers in the above feature extraction model, and k < N. The above at least three classifications include the category of the object to be recognized and at least two background categories, that is, non-object categories to be recognized. Therefore, the recognition result includes the recognition result of the object category to be recognized and the recognition results of at least two background categories.
[0083] In practical applications, candidate detection boxes can be generated for the target downsampling layer first, and then image recognition of at least three categories is performed on the candidate detection boxes to obtain the recognition confidence levels of the candidate detection boxes for at least three categories respectively. The maximum value of the recognition confidence levels among at least two background categories is selected as the recognition confidence level of the candidate detection box for the non-object region in the target image through the max function. Thus, the maximum value of this recognition confidence level and the recognition confidence level of the object category to be recognized are used to determine the detection result of the object to be recognized in the target image.
[0084] See Figure 5 , Figure 5 which is a schematic diagram of multi-class recognition using the maxout network provided by an embodiment of the present application. Figure 5 Take Figure 3 the face recognition scenario shown as an example. Take the target downsampling layer as the first downsampling layer among the N downsampling layers in the feature extraction model, that is, k = 1, and take the four-class classification for this target downsampling layer as an example to introduce the multi-class recognition process. As Figure 5 shown, the four-class classification includes the object to be recognized (i.e., face) and three background categories (i.e., background 1, background 2, and background 3).
[0085] In the image recognition process, four-class recognition is performed based on the feature map A output by the first downsampling layer, that is, classification and recognition of face, background 1, background 2, and background 3 are performed on the feature map A. Then, the background category with the highest recognition credibility is selected from background 1, background 2, and background 3 through the max function. Suppose it is background 1. Then, according to these two categories of background 1 and face, the candidate detection box corresponding to the feature map is determined. The candidate detection box identifies the recognition confidence level and predicted position of the face in the target image and the recognition confidence level and predicted position of background 1, so as to determine the detection result of the object to be recognized in the target image according to the candidate detection box.
[0086] In the binary classification scenario, an excessive imbalance in the number of positive and negative samples will affect the accuracy of image recognition. For the situation where there are too many negative samples in the candidate detection boxes generated based on the larger feature maps, by increasing the number of recognition categories, the sample ratio of the candidate detection boxes is effectively adjusted, and the ratio between different samples is adjusted, thus improving the recognition accuracy of the objects to be recognized in the target image.
[0087] In Figure 3 In the scenario shown, the prediction layer 305 performs convolution operations on the feature maps V1-V6 output by 6 downsampling layers using a fully convolutional network with a convolution kernel of 3*3, and outputs the corresponding candidate detection boxes v1-v6 for each downsampling layer. Since there are too many negative samples in the candidate detection box v1 determined by the first downsampling layer C1 of the basic convolutional layer 301, the maxout network is used to perform 4-class recognition on the feature map V1, and the maximum value of the recognition confidence levels of the 3 background classes is used as the recognition confidence level of the background class.
[0088] The above-mentioned generation of multi-size candidate detection boxes based on feature maps realizes an adaptive multi-size object candidate detection box generation mechanism, ensures that multi-size objects can all match the corresponding-size candidate detection boxes, and adopts a multi-class recognition method to reduce the impact of positive and negative sample imbalance on image recognition accuracy, ensuring the recognition accuracy for multi-size objects.
[0089] S204: Determine the detection result of the object to be recognized in the target image according to the candidate detection box.
[0090] Based on the multiple candidate detection boxes determined above and their respective recognition confidence levels and predicted positions for the object to be recognized, the detection result of the object to be recognized in the target image is finally determined. This recognition result indicates whether the target image includes the object to be recognized, and if the target image includes the object to be recognized, the position of the object to be recognized in the target image.
[0091] Since the number of candidate detection boxes determined above is very large, if each candidate detection box is processed, the computational complexity for recognizing the object to be recognized in the target image will be very large, resulting in a slow image recognition speed.
[0092] To improve the speed of image recognition based on candidate detection boxes, in the embodiments of the present application, the above neural network model for image recognition further includes an output layer. By setting a limited recognition threshold to remove candidate detection boxes with low recognition confidence levels, and using the non-maximum suppression method to remove redundant candidate detection boxes, the recognition speed of the model is thus improved. That is Figure 4 The limited recognition threshold in
[0093] Specifically, the candidate detection boxes can be removed by using a defined recognition threshold, that is, comparing the recognition confidence of the candidate detection boxes for the object to be recognized with the defined recognition threshold, and removing the candidate detection boxes with a recognition confidence less than the defined recognition threshold. Then, non-maximum suppression (NMS) is performed on the detection boxes after candidate box removal to obtain the detection result of the object to be recognized for the target image.
[0094] If the recognition confidence of the candidate detection box is less than the defined recognition threshold, it indicates that the candidate detection box is untrustworthy. By removing this candidate box, the recognition accuracy of the image is improved. If the recognition confidence of the candidate detection box is greater than the defined recognition threshold, it indicates that the candidate detection box is trustworthy and can be used as the basis for recognizing the object to be recognized in the target image. In practical applications, the defined recognition threshold can be preset. For example, the defined recognition threshold is set to 0.05, which is not limited here.
[0095] The non-maximum suppression operation process includes: sorting according to the recognition confidence of the candidate detection boxes, that is, the confidence score, then adding the candidate detection box with the highest recognition confidence to the output list and deleting it from the candidate detection box list. Calculate the overlap degree recognition parameter (such as IOU) between the candidate detection box with the highest recognition confidence and other candidate detection boxes, and delete the other candidate detection boxes with an overlap degree recognition parameter greater than the preset threshold. Repeat this process until the candidate detection box list is empty. Among them, the intersection over union (IOU) refers to the ratio of the overlapping area of two candidate detection boxes to the total area of these two candidate detection boxes.
[0096] The above method removes the candidate detection boxes with relatively low recognition confidence by setting the defined recognition threshold, avoiding the recognition of low-confidence candidate detection boxes and reducing the influence of low-confidence candidate detection boxes on the detection results of the object to be recognized. Since there are overlaps among multiple candidate detection boxes at the same position in the target image, non-maximum suppression is performed on the non-maximum values through non-maximum suppression, eliminating redundant candidate detection boxes and avoiding the processing process for redundant candidate detection boxes, thereby improving the image recognition speed.
[0097] In Figure 3 In the output layer 306 shown, the defined recognition threshold is set to 0.05, and the candidate detection boxes among the above candidate detection boxes v1 - v6 with a recognition confidence lower than this defined recognition threshold of 0.05 are removed, and non-maximum suppression is performed on the remaining candidate detection boxes to obtain and output the detection result of the target image 303 for the face. This detection result identifies that the target image 303 includes a face and the position of this face in the target image 303.
[0098] For the image recognition method provided in the foregoing embodiments, for a target image having an object to be recognized, the feature extraction model is used to sequentially perform N+M times of downsampled image feature extraction on it, and the feature maps corresponding to the downsampled layers are obtained. Since the detection of small-sized objects requires the feature map to have a high resolution, and the feature maps obtained by the first N shallow downsampled layers can be used as the basis for recognizing small-sized objects. For the initial output features of the i-th downsampled layer and the (i+1)-th downsampled layer among these N downsampled layers, they are fused. Based on the strong semantic information that the (i+1)-th downsampled layer can obtain in the target image, it is supplemented to the features extracted by the i-th downsampled layer, and the feature map of the i-th downsampled layer is obtained, enhancing the semantic information in the feature map of the i-th downsampled layer and ensuring the detection accuracy of small-sized objects based on this feature map. In addition, since the detection of large-sized objects requires the feature map to have strong semantic information and does not require a high resolution, therefore, based on a size threshold, the downsampled layers of the feature extraction model are divided, and the last M downsampled layers are determined therefrom. This size threshold ensures that the feature maps output by these M downsampled layers have strong semantic information, ensuring the detection accuracy of large-sized objects based on this feature map. When determining the candidate detection boxes of the target image, the candidate detection boxes are respectively recognized based on the feature maps output by the foregoing respective downsampled layers. In this way, the consideration of objects to be recognized of different sizes is realized, playing a role in adapting to objects to be recognized of different sizes in the target image and ensuring the detection accuracy of objects of different sizes in the target image.
[0099] For the feature extraction model used in the image recognition method provided in the foregoing embodiments, the embodiments of the present application provide a corresponding feature extraction model training method, and this training method includes the following steps:
[0100] S501: Determine the training samples for training the first initial model.
[0101] The first initial model is a pre-established neural network model. The first initial model includes an initial feature sub-model, and this initial feature sub-model includes N downsampled layers and M downsampled layers connected in sequence. The training samples include sample images having objects to be recognized and the true detection boxes of the objects to be recognized in the sample images marked.
[0102] In order to further improve the recognition performance of the model, in a possible implementation manner, the original samples are subjected to data preprocessing, and then the training samples for training the first initial model are determined according to the preprocessed original samples, that is, Figure 6 the process of image data augmentation preprocessing 601 included in the model process shown.
[0103] Among them, the original samples include the original images with the objects to be recognized and the true detection boxes of the objects to be recognized marked in the original images. The data preprocessing includes any one or a combination of multiple operations such as cropping, expanding, scaling, blurring, brightness modification, and contrast modification.
[0104] In the actual processing process, the original images in the original samples can be subjected to data preprocessing, and the true detection boxes corresponding to the processed original images can be adaptively modified, and these are used as training samples to train the first initial model, improving the generalization ability of the model for various environmental backgrounds and improving the recognition performance of the model for images.
[0105] In addition, the pixel values of the sample images in the training samples can be normalized, that is, the pixel values of the sample images are transformed from the interval [0, 255] to [-1, 1], accelerating the convergence speed during model training.
[0106] S502: Respectively determine the feature maps for the training samples through the downsampling layers of the initial feature sub-model.
[0107] Using the sample images in the training samples as inputs, the feature maps for the training samples are respectively determined through the N+M downsampling layers of the initial feature sub-model. This processing process is similar to S202 above and will not be elaborated here.
[0108] S503: According to the feature maps determined by the downsampling layers of the initial feature sub-model, determine the detection box recognition results of the objects to be recognized in the training samples.
[0109] Before model training, it is necessary to design candidate detection boxes for identifying the objects to be recognized. In the embodiments of the present application, the sizes of the candidate detection boxes adopted by the first initial model include N+M, which respectively correspond to the downsampling layers of the first initial model. For example, it is set that the sizes of the N+M candidate detection boxes are respectively proportional to the sizes of the feature maps corresponding to the N+M downsampling layers.
[0110] In the application process, according to the feature maps determined by the N+M downsampling layers of the initial feature sub-model, the designed candidate detection boxes are respectively used to generate candidate detection boxes for the sample images. These candidate detection boxes identify the regions of the objects to be recognized and the regions of the non-objects to be recognized in the sample images. Then, image recognition for the objects to be recognized is respectively performed on these candidate detection boxes to determine their respective corresponding recognition confidences and the predicted positions for the objects to be recognized, and these are used as the detection box recognition results of the objects to be recognized in the training samples.
[0111] S504: According to the degree of overlap in position between the detection box recognition results and the true detection boxes, train the first initial model, and use the trained initial feature sub-model as the feature extraction model.
[0112] In practical applications, the loss function can be used to calculate the difference between the detection box recognition result and the true detection box of the object to be recognized in the sample image, and the loss function value can be used to train the first initial model, and the trained initial feature sub-model is used as the feature extraction model. Among them, the loss function includes a classification loss function and a location loss function. The classification loss function is used to identify the category difference between the detection box recognition result and the object to be recognized in the sample image, and the location loss function is used to identify the difference between the predicted position of the object to be recognized indicated by the detection box recognition result and the true position of the object to be recognized in the sample image.
[0113] In the embodiment of the present application, the classification loss function is the cross-entropy function (CrossEntropy), and the location loss function is the absolute loss function (SmoothL1), which is expressed by the formula as follows:
[0114] loss crossentropy =-∑ i y i ·ln p i
[0115]
[0116] Where y i represents the true label of the i-th object to be recognized category, and p i represents the predicted result of the i-th object to be recognized category. x represents the difference between the true position and the predicted position of the object to be recognized in the sample image.
[0117] During the model training process, the model can be optimized based on the above-mentioned overlap degree recognition parameter IOU to match the candidate detection boxes. That is Figure 6 the process of matching the candidate detection box 602 shown.
[0118] Specifically, according to the overlap degree recognition parameter, determine the target detection box corresponding to the true detection box in the detection box recognition result. The overlap degree recognition parameter includes multiple gradient parameters that decrease in sequence. During the determination process, if the number of target detection boxes is less than the threshold, reduce the gradient parameter used for the overlap degree recognition parameter, and then, according to the overlap degree between the true detection box and the target detection box, perform model training on the first initial model. Among them, the gradient parameter can be set according to the actual image recognition scenario and is not limited here.
[0119] For Figure 3Taking the face recognition scenario as an example, since the aspect ratio of the face size is relatively fixed, the aspect ratio of the candidate detection frame is set to 1:1, and its size is set to [16, 32, 64, 128, 256, 512], which is proportional to the size of each layer of feature map. The gradient parameter of IOU is set to [0.5, 0.35, 0.1]. When matching the candidate detection frame, first match the target detection frame with IOU>0.5 from the detection frame recognition result. If the number of matched target detection frames is less than the threshold, the IOU is reduced to 0.35. If the number of matched target detection frames is still less than the threshold, the target detection frame is matched according to IOU>0.1, so that faces of different scales can match enough target detection frames, thereby realizing the detection of faces of different sizes. In model training, the difference between the target detection frame and the real detection frame is calculated as the optimization target of the detection position.
[0120] The first initial model is trained and optimized by matching target detection frames of different sizes, so that the trained first initial model can recognize objects of different sizes to be recognized, thereby realizing image recognition of multi-size objects by the model.
[0121] It can be understood that since the neural network model used for image recognition in the embodiment of the present application adopts a single-stage detection model combined with a lightweight backbone network (such as MobileNetV2), compared with the image recognition model that uses a heavyweight backbone network (such as ResNet50), it reduces the complexity of the model while also affecting the recognition accuracy of the model.
[0122] In order to improve the recognition accuracy of the neural network model provided in the embodiment of the present application, the neural network model can be trained by using the model distillation training method, that is, Figure 6 The model distillation training process 603 is shown. The model distillation training process includes the following steps:
[0123] S601: Determine a second initial model.
[0124] The second initial model includes an initial benchmark sub-model, and the initial benchmark sub-model has the same number of downsampling layers as the initial feature sub-model, that is, the initial benchmark sub-model includes N downsampling layers and M downsampling layers connected in sequence, and the model scale of the initial benchmark sub-model is larger than the initial feature sub-model. Among them, the model scale refers to the complexity of the model, which can be measured by model parameters, the number of model layers, etc. In an embodiment of the present application, the model scale can be measured by model parameters. Therefore, the initial benchmark sub-model and the initial feature sub-model have the same number of downsampling layers, but the number of parameters of the initial benchmark sub-model is larger than that of the initial feature sub-model.
[0125] S602: Train the second initial model according to the training samples, and obtain a benchmark feature extraction model based on the initial benchmark sub-model.
[0126] This process is similar to the above S502 - S504 process and will not be elaborated here.
[0127] S603: Obtain the feature maps respectively determined by the downsampling layers of the initial feature sub-model and the benchmark feature extraction model for the training samples.
[0128] S604: Adjust the parameters of the initial feature sub-model according to the differences between the obtained feature maps.
[0129] Based on the above S602, use the trained benchmark feature extraction model as the teacher model to train the initial feature sub-model. During the training process, for the training samples, obtain the feature maps corresponding to the N + M downsampling layers in the initial feature sub-model and the feature maps corresponding to the N + M downsampling layers in the benchmark feature extraction model respectively. Then, adjust the parameters of the initial feature sub-model according to the differences between the feature maps output by the two models corresponding to the same downsampling layer.
[0130] Taking MobileNetV2 as the initial feature sub-model in the above first initial model and ResNet50 as the initial benchmark sub-model of the second initial model as an example. Replace MobileNetV2 with the ResNet50 model. First, train the ResNet50 model well using the training samples as the teacher model. Then, during the training of MobileNetV2, by minimizing the differences between the feature maps determined by MobileNetV2 and the corresponding feature maps determined by the teacher model, the detection effect of the feature extraction model is improved.
[0131] The above method of obtaining the feature extraction model by using the model distillation training method ensures the lightweight of the neural network model for image recognition, and further improves the recognition accuracy of the neural network model for images, achieving a balance between model accuracy and efficiency.
[0132] The neural network model provided in the embodiments of this application uses SSD as a single-stage detection network combined with the lightweight MobileNetV2 as the backbone network. The average values of the mean of Average Precision (mAP) for face detection on the three validation sets of the publicly available face dataset WiderFace reach 0.932, 0.922, and 0.855 respectively. And the detection speed of a single image with a size of 640 * 640 is 200 milliseconds, with a fast detection speed, meeting the detection requirements in real-time scenarios such as videos and live broadcasts.
[0133] For the image recognition method provided in the above embodiments, an embodiment of the present application further provides an image recognition device.
[0134] See Figure 7 , Figure 7 which is a schematic structural diagram of an image recognition device provided in an embodiment of the present application. As Figure 7 shown, the image recognition device 700 includes an acquisition unit 701, an extraction unit 702, and a determination unit 703:
[0135] The acquisition unit 701 is configured to acquire a target image having an object to be recognized;
[0136] The extraction unit 702 is configured to perform image feature extraction on the target image according to a feature extraction model including N downsampling layers and M downsampling layers connected in sequence, to obtain feature maps corresponding to the downsampling layers respectively. The size of the feature map output by the first downsampling layer in the M downsampling layers is smaller than a size threshold; the i-th downsampling layer and the (i + 1)-th downsampling layer in the N downsampling layers are adjacent downsampling layers, and the initial output feature of the i-th downsampling layer is fused with the initial output feature of the (i + 1)-th downsampling layer to obtain the feature map of the i-th downsampling layer;
[0137] The determination unit 703 is configured to determine candidate detection frames corresponding to feature maps of different sizes according to the feature maps respectively determined by the downsampling layers. The candidate detection frames are used to identify the region of the object to be recognized and the region of non-object to be recognized in the target image;
[0138] The determination unit 703 is further configured to determine a detection result of the object to be recognized for the target image according to the candidate detection frames.
[0139] In a possible implementation manner, the target downsampling layer is one of the first k downsampling layers in the N downsampling layers, k < N. For the target downsampling layer, the determination unit 703 is configured to:
[0140] Perform classification recognition on the feature map of the target downsampling layer based on at least three classifications, where the at least three classifications include the category of the object to be recognized and at least two background categories;
[0141] Determine a candidate detection frame corresponding to the feature map of the target downsampling layer according to the recognition result of the classification recognition.
[0142] In a possible implementation manner, the determination unit 703 is further configured to:
[0143] Determine the training samples for training the first initial model, where the training samples with the object to be recognized are labeled with the true detection boxes of the object to be recognized; the initial feature sub-model in the first initial model includes N downsampling layers and M downsampling layers connected in sequence;
[0144] Determine the feature maps for the training samples through the downsampling layers of the initial feature sub-model respectively; among them, the sizes of the candidate detection boxes adopted by the first initial model include N + M, corresponding to the downsampling layers of the first initial model respectively;
[0145] Determine the detection box recognition results of the object to be recognized in the training samples according to the feature maps determined by the downsampling layers of the initial feature sub-model;
[0146] The device further includes a training unit:
[0147] The training unit is used to perform model training on the first initial model according to the overlapping degree between the detection box recognition result and the true detection box in position, and use the trained initial feature sub-model as the feature extraction model.
[0148] In a possible implementation manner, the training unit is used to:
[0149] According to the overlapping degree recognition parameter, determine the target detection box corresponding to the true detection box in the detection box recognition result. The overlapping degree recognition parameter includes multiple gradient parameters that decrease in sequence. During the determination process, if the number of target detection boxes is less than the threshold, reduce the gradient parameter used for the overlapping degree recognition parameter;
[0150] Perform model training on the first initial model according to the overlapping degree between the true detection box and the target detection box.
[0151] In a possible implementation manner, the determination unit 703 is further used to determine a second initial model, where the second initial model includes an initial reference sub-model, the initial reference sub-model has the same number of downsampling layers as the initial feature sub-model, and the model scale of the initial reference sub-model is larger than that of the initial feature sub-model;
[0152] The training unit is further used to train the second initial model according to the training samples, and obtain a reference feature extraction model based on the initial reference sub-model;
[0153] The acquisition unit 701 is further used to acquire the feature maps respectively determined by the downsampling layers of the initial feature sub-model and the reference feature extraction model for the training samples;
[0154] The device further includes an adjustment unit:
[0155] The adjustment unit is configured to adjust the parameters of the initial feature sub-model according to the differences between the acquired feature maps.
[0156] In a possible implementation, the determining unit 703 is configured to:
[0157] Perform data preprocessing on the original samples, where the data preprocessing includes any one or a combination of multiple operations such as cropping, padding, scaling, blurring, brightness modification, and contrast modification;
[0158] Determine the training samples according to the original samples after the data preprocessing.
[0159] In a possible implementation, the determining unit 703 is configured to:
[0160] Remove candidate boxes from the candidate detection boxes based on a defined recognition threshold;
[0161] Perform non-maximum suppression on the candidate detection boxes after the candidate box removal to obtain the detection result of the object to be recognized for the target image.
[0162] In a possible implementation, the ratio of the sizes of the feature maps output by adjacent downsampling layers in the feature extraction model is 1 / 2 * 1 / 2.
[0163] The image recognition device provided in the above embodiments performs N+M times of downsampled image feature extraction on a target image with an object to be recognized in sequence by using a feature extraction model, and obtains feature maps corresponding to the downsampling layers respectively. Since the detection of small-sized objects requires the feature maps to have a high resolution, and the feature maps obtained by the first N shallow downsampling layers can be used as the basis for recognizing small-sized objects, the initial output features of the i-th downsampling layer and the (i+1)-th downsampling layer among these N downsampling layers are fused. Based on the strong semantic information that the (i+1)-th downsampling layer can obtain in the target image, it is supplemented into the features extracted by the i-th downsampling layer, and the feature map of the i-th downsampling layer is obtained, enhancing the semantic information in the feature map of the i-th downsampling layer and ensuring the detection accuracy of small-sized objects based on this feature map. In addition, since the detection of large-sized objects requires the feature maps to have strong semantic information and does not require a high resolution, therefore, the downsampling layers of the feature extraction model are divided based on a size threshold, and the last M downsampling layers are determined therefrom. This size threshold ensures that the feature maps output by these M downsampling layers have strong semantic information and ensures the detection accuracy of large-sized objects based on this feature map. In the process of determining the candidate detection boxes of the target image, the candidate detection boxes are respectively recognized based on the feature maps output by the foregoing respective downsampling layers. In this way, the consideration of objects to be recognized with different sizes is realized, which plays a role in adapting to objects to be recognized with different sizes in the target image and ensures the detection accuracy of objects with different sizes in the target image.
[0164] An embodiment of the present application also provides a computer device. The computer device for image recognition provided in the embodiments of the present application will be introduced from the perspective of hardware implementation below.
[0165] See Figure 8 , Figure 8 FIG. is a schematic structural diagram of a server provided in an embodiment of the present application. The server 1400 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1422 (for example, one or more processors) and a memory 1432, and one or more storage media 1430 (for example, one or more mass storage devices) for storing application programs 1442 or data 1444. Among them, the memory 1432 and the storage media 1430 may be transient storage or persistent storage. The program stored in the storage media 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1422 may be configured to communicate with the storage media 1430 and execute a series of instruction operations in the storage media 1430 on the server 1400.
[0166] The server 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1458, and / or, one or more operating systems 1441, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0167] In the above embodiments, the steps executed by the server may be based on the Figure 8 server structure shown.
[0168] Among them, the CPU 1422 is used to execute the following steps:
[0169] Obtain a target image with an object to be recognized;
[0170] Perform image feature extraction on the target image according to a feature extraction model including N downsampling layers and M downsampling layers connected in sequence to obtain feature maps corresponding to the downsampling layers respectively. The size of the feature map output by the first downsampling layer in the M downsampling layers is smaller than a size threshold; the i-th downsampling layer and the (i + 1)-th downsampling layer in the N downsampling layers are adjacent downsampling layers, and the initial output feature of the i-th downsampling layer and the initial output feature of the (i + 1)-th downsampling layer are fused to obtain the feature map of the i-th downsampling layer;
[0171] Determine candidate detection frames corresponding to feature maps of different sizes according to the feature maps respectively determined by the downsampling layers. The candidate detection frames are used to identify the area of the object to be recognized and the area of non-object to be recognized in the target image;
[0172] Determine the detection result of the object to be recognized for the target image according to the candidate detection frames.
[0173] Optionally, the CPU 1422 may also execute the image recognition method provided in the above embodiments, which will not be elaborated here.
[0174] For the image recognition method described above, an embodiment of the present application also provides a terminal device for image recognition to implement and apply the above image recognition method in practice.
[0175] See Figure 9 , Figure 9A schematic structural diagram of a terminal device provided by an embodiment of the present application. For ease of description, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The terminal device may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), etc. Taking the terminal device as a mobile phone as an example:
[0176] Figure 9 The figure shows a block diagram of a part of the structure of a mobile phone related to the terminal device provided by an embodiment of the present application. Refer to Figure 9 , the mobile phone includes: a radio frequency (RF) circuit 1510, a memory 1520, an input unit 1530, a display unit 1540, a sensor 1550, an audio circuit 1560, a wireless fidelity (WiFi) module 1570, a processor 1580, and a power supply 1590 and other components. Those skilled in the art can understand that Figure 9 the structure of the mobile phone shown in
[0177] does not limit the mobile phone, and may include more or fewer components than shown in the figure, or combine some components, or arrange different components. Figure 9 The following specifically introduces each component of the mobile phone:
[0178] The RF circuit 1510 can be used for receiving and sending information or signals during communication. Specifically, after receiving the downlink information from the base station, it is processed by the processor 1580. Additionally, the uplink data is sent to the base station. Generally, the RF circuit 1510 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. Moreover, the RF circuit 1510 can also communicate with the network and other devices via wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0179] The memory 1520 can be used to store software programs and modules. The processor 1580 runs the software programs and modules stored in the memory 1520 to implement various functional applications and image recognition of the mobile phone. The memory 1520 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 1520 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0180] The input unit 1530 can be used to receive input numerical or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1530 can include a touch panel 1531 and other input devices 1532. The touch panel 1531, also known as a touch screen, can collect the touch operations of the user thereon or nearby (such as the operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel 1531), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 1531 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1580, and can receive and execute the commands sent by the processor 1580. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 1531. In addition to the touch panel 1531, the input unit 1530 can also include other input devices 1532. Specifically, the other input devices 1532 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0181] The display unit 1540 can be used to display the information input by the user or the information provided to the user and various menus of the mobile phone. The display unit 1540 can include a display panel 1541. Optionally, the display panel 1541 can be configured in the form of a liquid crystal display (LCD for short), an organic light-emitting diode (OLED for short), etc. Further, the touch panel 1531 can cover the display panel 1541. When the touch panel 1531 detects a touch operation thereon or nearby, it transmits it to the processor 1580 to determine the type of touch event. Subsequently, the processor 1580 provides a corresponding visual output on the display panel 1541 according to the type of touch event. Although in Figure 9 the touch panel 1531 and the display panel 1541 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 1531 and the display panel 1541 can be integrated to realize the input and output functions of the mobile phone.
[0182] The mobile phone may further include at least one sensor 1550, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1541 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1541 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.
[0183] The audio circuit 1560, the speaker 1561, and the microphone 1562 can provide an audio interface between the user and the mobile phone. The audio circuit 1560 can transmit the electrical signal converted from the received audio data to the speaker 1561, and the speaker 1561 converts it into a sound signal for output; on the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560 and then converted into audio data. After the audio data is output to the processor 1580 for processing, it is sent through the RF circuit 1510 to, for example, another mobile phone, or the audio data is output to the memory 1520 for further processing.
[0184] WiFi belongs to short - range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 1570, which provides users with wireless broadband Internet access. Although Figure 9 the WiFi module 1570 is shown, it can be understood that it does not belong to an essential component of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention according to needs.
[0185] The processor 1580 is the control center of the mobile phone, connecting various parts of the entire mobile phone using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1520, and by calling data stored in the memory 1520, it executes various functions of the mobile phone and processes data. Optionally, the processor 1580 may include one or more processing units; preferably, the processor 1580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 1580.
[0186] The mobile phone further includes a power source 1590 (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the processor 1580 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system.
[0187] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0188] In the embodiment of the present application, the memory 1520 included in the mobile phone can store program codes and transmit the program codes to the processor.
[0189] The processor 1580 included in the mobile phone can execute the image recognition method provided in the above embodiment according to the instructions in the program codes, which will not be elaborated here.
[0190] The embodiment of the present application further provides a computer-readable storage medium for storing a computer program, and the computer program is used to execute the image recognition method provided in the above embodiment.
[0191] The embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the image recognition method provided in various optional implementation manners of the above aspects.
[0192] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium can be at least one of the following media: read-only memory (English: read-only memory, abbreviation: ROM), RAM, magnetic disk, or optical disc, etc., which can store program codes.
[0193] It should be noted that the embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0194] As mentioned above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An image recognition method, characterized in that, the method includes: Obtaining a target image having an object to be recognized; Performing image feature extraction on the target image according to a feature extraction model including N downsampling layers and M downsampling layers connected in sequence, to obtain feature maps corresponding to the downsampling layers respectively. The size of the feature map output by the first downsampling layer in the M downsampling layers is smaller than a size threshold. The i-th downsampling layer and the (i + 1)-th downsampling layer in the N downsampling layers are adjacent downsampling layers. The initial output feature of the i-th downsampling layer and the initial output feature of the (i + 1)-th downsampling layer are fused to obtain the feature map of the i-th downsampling layer. For the N-th downsampling layer, the initial output feature is directly used as the feature map of the N-th downsampling layer for output; Determining candidate detection frames corresponding to feature maps of different sizes according to the feature maps respectively determined by the downsampling layers. The candidate detection frames are used to identify the area of the object to be recognized and the area of non-object to be recognized in the target image; Determining a detection result of the object to be recognized for the target image according to the candidate detection frames; Wherein, the target downsampling layer is one of the first k downsampling layers in the N downsampling layers, k < N. For the target downsampling layer, the step of determining candidate detection frames corresponding to feature maps of different sizes according to the feature maps respectively determined by the downsampling layers includes: Performing classification recognition based on at least three classifications on the feature map of the target downsampling layer. The at least three classifications include the category of the object to be recognized and at least two background categories; Determining candidate detection frames corresponding to the feature map of the target downsampling layer according to the recognition result of the classification recognition.
2. The method according to claim 1, characterized in that, the method further includes: Determining training samples for training a first initial model. The training samples having the object to be recognized are labeled with the true detection frames of the object to be recognized; the initial feature sub-model in the first initial model includes N downsampling layers and M downsampling layers connected in sequence; Determining feature maps for the training samples respectively through the downsampling layers of the initial feature sub-model; wherein, the sizes of the candidate detection frames adopted by the first initial model include N + M, corresponding to the downsampling layers of the first initial model respectively; Determining the detection frame recognition result of the object to be recognized in the training samples according to the feature maps determined by the downsampling layers of the initial feature sub-model; Performing model training on the first initial model according to the overlapping degree of the detection frame recognition result and the true detection frame in position, and taking the trained initial feature sub-model as the feature extraction model.
3. The method according to claim 2, characterized in that, the step of performing model training on the first initial model according to the overlapping degree of the detection frame recognition result and the true detection frame in position includes: Identify the target detection box corresponding to the true detection box in the detection box recognition result according to the overlap degree recognition parameter. The overlap degree recognition parameter includes multiple gradient parameters that decrease in sequence. During the determination process, if the number of target detection boxes is less than the threshold, decrease the gradient parameter used to decrease the overlap degree recognition parameter; Train the first initial model according to the overlap degree between the true detection box and the target detection box.
4. The method according to claim 2, wherein, the method further includes: Determine a second initial model, which includes an initial reference sub-model. The initial reference sub-model has the same number of downsampling layers as the initial feature sub-model, and the model scale of the initial reference sub-model is larger than that of the initial feature sub-model; Train the second initial model according to the training samples, and obtain a reference feature extraction model based on the initial reference sub-model; Obtain the feature maps respectively determined by the downsampling layers of the initial feature sub-model and the reference feature extraction model for the training samples; Adjust the parameters of the initial feature sub-model according to the differences between the obtained feature maps.
5. The method according to claim 2, wherein, The determination of the training samples for training the first initial model includes: Perform data preprocessing on the original samples, and the data preprocessing includes any one or a combination of clipping, expanding, scaling, blurring, brightness modification, and contrast modification; Determine the training samples according to the original samples after the data preprocessing.
6. The method according to any one of claims 1-5, wherein, The determination of the detection result of the object to be recognized for the target image according to the candidate detection boxes includes: Remove candidate boxes from the candidate detection boxes based on a defined recognition threshold; Perform a non-maximum suppression operation on the candidate detection boxes after the candidate box removal to obtain the detection result of the object to be recognized for the target image.
7. The method according to any one of claims 1-5, wherein, The ratio of the sizes of the feature maps output by adjacent downsampling layers in the feature extraction model is 1 / 2*1 / 2.
8. An image recognition device, wherein, the device includes an acquisition unit, an extraction unit, and a determination unit: The acquisition unit is used to acquire a target image with an object to be recognized; The extraction unit is used to perform image feature extraction on the target image according to a feature extraction model including N sequentially connected downsampling layers and M downsampling layers to obtain the feature maps corresponding to the downsampling layers respectively. The size of the feature map output by the first downsampling layer in the M downsampling layers is smaller than a size threshold; the i-th downsampling layer and the (i + 1)-th downsampling layer in the N downsampling layers are adjacent downsampling layers, and the initial output feature of the i-th downsampling layer and the initial output feature of the (i + 1)-th downsampling layer are fused to obtain the feature map of the i-th downsampling layer. For the N-th downsampling layer, directly output the initial output feature as the feature map of the N-th downsampling layer; The determining unit is configured to determine candidate detection frames corresponding to feature maps of different sizes according to the feature maps respectively determined by the downsampling layers, where the candidate detection frames are used to identify regions of the object to be recognized and regions other than the object to be recognized in the target image; The determining unit is further configured to determine a detection result of the object to be recognized for the target image according to the candidate detection frames; Wherein, the target downsampling layer is one of the first k downsampling layers among the N downsampling layers, k < N. For the target downsampling layer, the determining unit is configured to: Perform classification recognition on the feature map of the target downsampling layer based on at least three classifications, where the at least three classifications include the category of the object to be recognized and at least two background categories; Determine candidate detection frames corresponding to the feature map of the target downsampling layer according to the recognition result of the classification recognition.
9. The apparatus according to claim 8, wherein, the determining unit is further configured to: Determine training samples for training a first initial model, where the training samples with the object to be recognized are labeled with the true detection frames of the object to be recognized; the initial feature sub-model in the first initial model includes N downsampling layers and M downsampling layers connected in sequence; Determine feature maps for the training samples respectively through the downsampling layers of the initial feature sub-model; wherein, the sizes of the candidate detection frames adopted by the first initial model include N + M, and respectively correspond to the downsampling layers of the first initial model; Determine the detection frame recognition result of the object to be recognized in the training samples according to the feature maps determined by the downsampling layers of the initial feature sub-model; The apparatus further includes a training unit: The training unit is configured to perform model training on the first initial model according to the overlapping degree between the detection frame recognition result and the true detection frame in position, and use the trained initial feature sub-model as the feature extraction model.
10. The apparatus according to claim 9, wherein, the training unit is configured to: Determine a target detection frame corresponding to the true detection frame in the detection frame recognition result according to the overlapping degree recognition parameter, where the overlapping degree recognition parameter includes multiple gradient parameters that decrease in sequence. During the determination process, if the number of target detection frames is less than the threshold, reduce the gradient parameter used for the overlapping degree recognition parameter; Perform model training on the first initial model according to the overlapping degree between the true detection frame and the target detection frame.
11. The apparatus according to claim 9, wherein, the determining unit is further configured to determine a second initial model, where the second initial model includes an initial reference sub-model, the initial reference sub-model has the same number of downsampling layers as the initial feature sub-model, and the model scale of the initial reference sub-model is larger than that of the initial feature sub-model; The training unit is further configured to perform training on the second initial model according to the training samples, and obtain a reference feature extraction model based on the initial reference sub-model; The obtaining unit is further configured to obtain the feature maps respectively determined by the initial feature sub-model and the downsampling layer of the reference feature extraction model for the training samples; The apparatus further includes an adjustment unit: The adjustment unit is configured to adjust the parameters of the initial feature sub-model according to the differences between the obtained feature maps.
12. The apparatus according to claim 9, wherein, The determining unit is configured to: perform data preprocessing on the original sample, and the data preprocessing includes any one or a combination of multiple items of cropping, expanding, scaling, blurring, brightness modification, and contrast modification; determine the training sample according to the original sample after the data preprocessing.
13. The apparatus according to any one of claims 8-12, wherein, The determining unit is configured to: remove candidate boxes from the candidate detection boxes based on a defined recognition threshold; perform non-maximum suppression operation on the candidate detection boxes after the candidate boxes are removed to obtain the detection result of the object to be recognized for the target image.
14. The apparatus according to any one of claims 8-12, wherein, The ratio of the sizes of the feature maps output by adjacent downsampling layers in the feature extraction model is 1 / 2*1 / 2.
15. A computer device, wherein, The computer device includes a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor; The processor is configured to execute the method according to any one of claims 1-7 according to the instructions in the program codes.
16. A computer-readable storage medium, wherein, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1-7.
17. A computer program product, wherein, The computer program product includes instructions, and when the instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A small target detection method and a detection model based on a convolutional neural network
CN109886359A