Image scene recognition and model training methods, devices, and computer equipment
By extracting and fusing the foreground and background features of the image in image scene recognition, and using the self-attention mechanism to improve the recognition accuracy, the problem of low accuracy of image scene recognition in the prior art is solved.
Patent Information
- Application Number
- CN202110255557.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-03-09
AI Technical Summary
In the prior art, there is a problem of low accuracy in image scene recognition, especially when there is no detected object in the image, such as seasides and large forests.
By extracting the foreground area and background area in the image to be identified, the self-attention foreground weight and self-attention background weight are calculated based on the self-attention mechanism, the characteristics are adjusted and feature fusion is performed, and scene recognition is performed.
The accuracy of image scene recognition is improved, and sufficient scene recognition features are extracted through the joint action of the background area and the foreground area.
Smart Images

Figure CN113706550B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to an image scene recognition and model training method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of image processing technology, image recognition technology has emerged. Image recognition refers to the technology of using a computer to process, analyze, and understand images to identify various different patterns of targets and objects, and it is a practical application of applying deep learning algorithms. Currently, image recognition can recognize the scene in an image, that is, by detecting the objects in the image and then determining the scene where the image is located based on the objects in the image. However, the method of detecting the objects in the image and then performing scene recognition has a problem of low accuracy in image scene recognition because there may be no detectable objects in the image, such as the seaside, large forests, etc. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide an image scene recognition and model training method, apparatus, computer device, and storage medium that can improve the accuracy of image scene recognition.
[0004] An image scene recognition method, the method includes:
[0005] Obtain an image to be recognized;
[0006] Extract the foreground region and the background region from the image to be recognized;
[0007] Calculate the self-attention weight based on the foreground features corresponding to the foreground region to obtain the self-attention foreground weight, and adjust the foreground features through the self-attention foreground weight to obtain the self-attention foreground features;
[0008] Calculate the self-attention weight based on the background features corresponding to the background region to obtain the self-attention background weight, and adjust the background features through the self-attention background weight to obtain the self-attention background features;
[0009] Fuse the self-attention background features and the self-attention foreground features to obtain fused features, and perform scene recognition based on the fused features to obtain the image scene recognition result corresponding to the image to be recognized.
[0010] In one embodiment, partitioning the region based on the features of the image to be recognized to obtain the foreground region and the background region includes:
[0011] Calculate the mean value corresponding to the feature values in the features of the image to be recognized;
[0012] Perform binary partitioning on the features of the image to be recognized based on the mean value to obtain a foreground mask;
[0013] Calculate the product of the foreground mask and the pixel values of the image to be recognized to obtain the foreground region in the image to be recognized;
[0014] Invert the foreground mask to obtain a background mask, calculate the product of the background mask and the pixel values of the image to be recognized, and obtain the background region.
[0015] An image scene recognition device, the device includes:
[0016] An image acquisition module for acquiring an image to be recognized;
[0017] A region extraction module for extracting the foreground region and the background region in the image to be recognized;
[0018] A foreground feature extraction module for calculating self-attention weights based on the foreground features corresponding to the foreground region to obtain self-attention foreground weights, and adjusting the foreground features through the self-attention foreground weights to obtain self-attention foreground features;
[0019] A background feature extraction module for calculating self-attention weights based on the background features corresponding to the background region to obtain self-attention background weights, and adjusting the background features through the self-attention background weights to obtain self-attention background features;
[0020] A scene recognition module for fusing the self-attention background features and the self-attention foreground features to obtain fused features, and performing scene recognition based on the fused features to obtain an image scene recognition result corresponding to the image to be recognized.
[0021] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0022] Acquire an image to be recognized;
[0023] Extract the foreground region and the background region in the image to be recognized;
[0024] Calculate self-attention weights based on the foreground features corresponding to the foreground region to obtain self-attention foreground weights, and adjust the foreground features through the self-attention foreground weights to obtain self-attention foreground features;
[0025] Calculate self-attention weights based on the background features corresponding to the background region to obtain self-attention background weights, and adjust the background features through the self-attention background weights to obtain self-attention background features;
[0026] Fuse the self-attention background features and the self-attention foreground features to obtain fused features, and perform scene recognition based on the fused features to obtain an image scene recognition result corresponding to the image to be recognized.
[0027] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented:
[0028] Obtain an image to be recognized;
[0029] Extract the foreground region and the background region in the image to be recognized;
[0030] Calculate self-attention weights based on the foreground features corresponding to the foreground region to obtain self-attention foreground weights, and adjust the foreground features through the self-attention foreground weights to obtain self-attention foreground features;
[0031] Calculate self-attention weights based on the background features corresponding to the background region to obtain self-attention background weights, and adjust the background features through the self-attention background weights to obtain self-attention background features;
[0032] Fuse the self-attention background features and the self-attention foreground features to obtain fused features, and perform scene recognition based on the fused features to obtain an image scene recognition result corresponding to the image to be recognized.
[0033] The above image scene recognition method, device, computer device and storage medium extract the foreground region and the background region in the image to be recognized, then calculate self-attention weights based on the foreground features corresponding to the foreground region to obtain self-attention foreground weights, and adjust the foreground features through the self-attention foreground weights to obtain self-attention foreground features. Calculate self-attention weights based on the background features corresponding to the background region to obtain self-attention background weights, and adjust the background features through the self-attention background weights to obtain self-attention background features, and fuse the self-attention foreground features and the self-attention background features and then perform scene recognition, that is, jointly use the background region and the foreground region to recognize the image scene, so that sufficient scene recognition features can be extracted, thereby improving the accuracy of image scene recognition.
[0034] An image scene recognition model training method, the method includes:
[0035] Obtain training images and corresponding training scene labels, and input the training images into an initial image scene recognition model;
[0036] The initial image scene recognition model extracts the initial foreground training region and the initial background training region in the training image, inputs the initial foreground region into the initial foreground branch network, and inputs the initial background region into the initial background branch network;
[0037] The initial foreground branch network calculates self-attention weights based on the initial foreground features corresponding to the initial foreground training region to obtain the initial self-attention foreground weights, and adjusts the initial foreground features through the initial self-attention foreground weights to obtain the initial self-attention foreground features;
[0038] The initial background branch network calculates self-attention weights based on the initial background features corresponding to the initial background training region to obtain the initial self-attention background weights, and adjusts the initial background features through the initial self-attention background weights to obtain the initial self-attention background features;
[0039] The initial image scene recognition model fuses the initial self-attention background features and the initial self-attention foreground features to obtain the initial fused features, and performs scene recognition based on the initial fused features to obtain the initial image scene recognition result;
[0040] Calculate the loss information between the initial image scene recognition result and the training scene label, update the initial image scene recognition model based on the loss information, and return to iterate the step of inputting the training image into the initial image scene recognition model until the training completion condition is reached, and then obtain the trained image scene recognition model.
[0041] An image scene recognition model training device, the device includes:
[0042] A training data acquisition module, configured to acquire a training image and a corresponding training scene label, and input the training image into the initial image scene recognition model;
[0043] A model processing module, configured to extract the initial foreground training region and the initial background training region in the training image by the initial image scene recognition model, input the initial foreground region into the initial foreground branch network, and input the initial background region into the initial background branch network;
[0044] A foreground network processing module, configured to calculate self-attention weights by the initial foreground branch network based on the initial foreground features corresponding to the initial foreground training region to obtain the initial self-attention foreground weights, and adjust the initial foreground features through the initial self-attention foreground weights to obtain the initial self-attention foreground features;
[0045] A background network processing module, configured to calculate self-attention weights by the initial background branch network based on the initial background features corresponding to the initial background training region to obtain the initial self-attention background weights, and adjust the initial background features through the initial self-attention background weights to obtain the initial self-attention background features;
[0046] A model recognition module, which is used to perform feature fusion on the initial self-attention background feature and the initial self-attention foreground feature by an initial image scene recognition model to obtain an initial fusion feature, and perform scene recognition based on the initial fusion feature to obtain an initial image scene recognition result;
[0047] An iteration module, which is used to calculate the loss information between the initial image scene recognition result and the training scene label, update the initial image scene recognition model based on the loss information, and return to iteratively execute the step of inputting the training image into the initial image scene recognition model until the training completion condition is reached, and obtain a trained image scene recognition model.
[0048] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0049] Obtain a training image and a corresponding training scene label, and input the training image into an initial image scene recognition model;
[0050] The initial image scene recognition model extracts an initial foreground training area and an initial background training area in the training image, inputs the initial foreground area into an initial foreground branch network, and inputs the initial background area into an initial background branch network;
[0051] The initial foreground branch network calculates self-attention weights based on the initial foreground feature corresponding to the initial foreground training area to obtain initial self-attention foreground weights, and adjusts the initial foreground feature through the initial self-attention foreground weights to obtain an initial self-attention foreground feature;
[0052] The initial background branch network calculates self-attention weights based on the initial background feature corresponding to the initial background training area to obtain initial self-attention background weights, and adjusts the initial background feature through the initial self-attention background weights to obtain an initial self-attention background feature;
[0053] The initial image scene recognition model performs feature fusion on the initial self-attention background feature and the initial self-attention foreground feature to obtain an initial fusion feature, and performs scene recognition based on the initial fusion feature to obtain an initial image scene recognition result;
[0054] Calculate the loss information between the initial image scene recognition result and the training scene label, update the initial image scene recognition model based on the loss information, and return to iteratively execute the step of inputting the training image into the initial image scene recognition model until the training completion condition is reached, and obtain a trained image scene recognition model.
[0055] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0056] Obtain training images and corresponding training scene labels, and input the training images into the initial image scene recognition model;
[0057] The initial image scene recognition model extracts the initial foreground training region and the initial background training region from the training image, inputs the initial foreground region into the initial foreground branch network, and inputs the initial background region into the initial background branch network;
[0058] The initial foreground branch network calculates the self-attention weight based on the foreground features corresponding to the foreground region to obtain the self-attention foreground weight, and adjusts the foreground features through the self-attention foreground weight to obtain the self-attention foreground features;
[0059] The initial background branch network calculates the self-attention weight based on the background features corresponding to the background region to obtain the self-attention background weight, and adjusts the background features through the self-attention background weight to obtain the self-attention background features;
[0060] The initial image scene recognition model fuses the initial self-attention background features and the initial self-attention foreground features to obtain the initial fusion features, and performs scene recognition based on the initial fusion features to obtain the initial image scene recognition result;
[0061] Calculate the loss information between the initial image scene recognition result and the training scene label, update the initial image scene recognition model based on the loss information, and return to the step of inputting the training image into the initial image scene recognition model and iterate until the training completion condition is reached to obtain the trained image scene recognition model.
[0062] The above image scene recognition model training method, device, computer device and storage medium obtain training images and corresponding training scene labels, input the training images into the initial image scene recognition model. The initial image scene recognition model extracts the self-attention foreground features through the initial foreground branch network and extracts the self-attention background features through the initial background branch network. Then, the initial self-attention background features and the initial self-attention foreground features are fused to obtain the initial fusion features, and scene recognition is performed based on the initial fusion features to obtain the initial image scene recognition result. Calculate the loss information between the initial image scene recognition result and the training scene label, and update the initial image scene recognition model based on the loss information until the training completion condition is reached to obtain the trained image scene recognition model. Since the self-attention foreground features and the self-attention background features are respectively extracted through the foreground branch network and the background branch network, and then the initial self-attention background features and the initial self-attention foreground features are fused to identify the image scene recognition result, the trained image scene recognition model can improve the accuracy of image scene recognition. Brief Description of the Drawings
[0063] Figure 1 It is an application environment diagram of the image scene recognition method in an embodiment;
[0064] Figure 2 It is a schematic flowchart of the image scene recognition method in an embodiment;
[0065] Figure 3 It is a schematic flowchart of model recognition in an embodiment;
[0066] Figure 4 It is a schematic diagram of the module structure in a specific embodiment;
[0067] Figure 5 It is a schematic flowchart of obtaining the background region in an embodiment;
[0068] Figure 6 It is a schematic diagram of the obtained background mask in a specific embodiment;
[0069] Figure 7 It is a schematic flowchart of obtaining the self-attention foreground feature in an embodiment;
[0070] Figure 8 It is a schematic flowchart of obtaining the self-attention background feature in an embodiment;
[0071] Figure 9 It is a schematic flowchart of the image scene recognition model training method in an embodiment;
[0072] Figure 10 It is a schematic flowchart of obtaining the image scene recognition model in an embodiment;
[0073] Figure 11 It is a schematic flowchart of pre-training in an embodiment;
[0074] Figure 12 It is a schematic flowchart of image scene recognition in a specific embodiment;
[0075] Figure 13 It is a schematic diagram of the architecture of the image scene recognition model in a specific embodiment;
[0076] Figure 14 It is a schematic diagram of the application scenario of the image scene recognition method in a specific embodiment;
[0077] Figure 15 For Figure 14 It is a schematic diagram of the picture to be recognized in a specific embodiment;
[0078] Figure 16 It is a structural block diagram of the image scene recognition device in an embodiment;
[0079] Figure 17 It is a structural block diagram of an image scene recognition model training device in an embodiment;
[0080] Figure 18 It is an internal structure diagram of a computer device in an embodiment;
[0081] Figure 19 It is an internal structure diagram of a computer device in another embodiment. Detailed implementation manners
[0082] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0083] Computer Vision Technology (CV): Computer vision is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further performing graphics processing to make the computer process into images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0084] The solution provided in the embodiments of the present application relates to technologies such as image recognition in artificial intelligence, and is specifically described through the following embodiments:
[0085] The image scene recognition method provided by the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The server 104 obtains the image to be recognized uploaded by the terminal 102; extracts the foreground area and the background area in the image to be recognized; the server 104 calculates the self-attention weight based on the foreground features corresponding to the foreground area to obtain the self-attention foreground weight, and adjusts the foreground features through the self-attention foreground weight to obtain the self-attention foreground features; the server 104 calculates the self-attention weight based on the background features corresponding to the background area to obtain the self-attention background weight, and adjusts the background features through the self-attention background weight to obtain the self-attention background features; the server 104 fuses the self-attention background features and the self-attention foreground features to obtain the fused features, and performs scene recognition based on the fused features to obtain the image scene recognition result corresponding to the image to be recognized. The server 104 can return the image scene recognition result to the terminal 102 for display. Among them, the terminal 102 can be, but is not limited to, a laptop computer, a smart phone, a tablet computer, a desktop computer, a smart TV, and a portable wearable device. The server 104 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0086] In one embodiment, as Figure 2 shown, a method for image scene recognition is provided. Taking the server in Figure 1 as an example for illustration, it can be understood that this method can also be applied to the terminal, or jointly executed by the terminal and the server. In this embodiment, the method includes the following steps:
[0087] Step 202, obtain the image to be recognized.
[0088] Step 204, extract the foreground area and the background area in the image to be recognized.
[0089] Among them, the image to be recognized refers to the image that needs to be subjected to scene recognition. The foreground area refers to the partial area in the image to be recognized where the foreground is located. The foreground refers to the scenery or characters in the picture that are in front of the main body and even close to the camera lens, showing a certain spatial relationship or character relationship. The background area refers to the partial area in the image to be recognized where the background is located. The background is located behind the main body and is the scenery far from the camera, which is an important part of the environment.
[0090] Specifically, the server obtains the image to be recognized. The image to be recognized can be uploaded to the server by the terminal, obtained by the server from the database, or sent by the service server, which is used to process image-related services. The server segments the image to be recognized and extracts the foreground area and the background area in the image to be recognized. Among them, the server can segment the image to be recognized using a threshold-based segmentation algorithm to obtain the foreground area and the background area. The server can also use pixel clustering for segmentation. The server can also use the maximum entropy algorithm to segment the image to be recognized to obtain the foreground area and the background area. The server can also segment the image to be recognized based on a deep neural network algorithm to obtain the foreground area and the background area. In one embodiment, the server can extract the image features corresponding to the image to be recognized and then use the image features to segment the image to be recognized to obtain the foreground area and the background area. Among them, the deep neural network can be used to extract the image features corresponding to the image to be recognized. In a specific embodiment, the server can also obtain the image to be recognized through an application (APP, Application) or a client on the terminal.
[0091] Step 206: Calculate the self-attention foreground weight based on the foreground features corresponding to the foreground area, and adjust the foreground features through the self-attention foreground weight to obtain the self-attention foreground features.
[0092] Among them, the foreground features refer to the regional features corresponding to the foreground area, and the self-attention foreground weight refers to the weight corresponding to the foreground features calculated through the self-attention mechanism. Self-attention means finding the corresponding attention weight for each feature element. The self-attention foreground features refer to the features obtained by weighting the foreground features with the self-attention foreground weight.
[0093] Specifically, the server extracts the foreground features corresponding to the foreground area. The features of the foreground area can be extracted through a deep neural network to obtain the foreground features. The features can also be obtained through a feature extraction algorithm. For example, the features can be color features, texture features, shape features, spatial relationship features, etc. Then, the self-attention foreground weight is calculated using the foreground features. Among them, the self-attention weight can be calculated through a neural network established by the self-attention mechanism. The server uses the self-attention foreground weight to weight the foreground features to obtain the self-attention foreground features.
[0094] Step 208: Calculate the self-attention background weight based on the background features corresponding to the background area, and adjust the background features through the self-attention background weight to obtain the self-attention background features.
[0095] Among them, the background feature refers to the feature corresponding to the background area, and the self-attention background weight refers to the weight corresponding to the background feature calculated through the self-attention mechanism (self-attention mechanism). The self-attention background feature refers to the feature obtained by weighting the background feature with the self-attention foreground weight.
[0096] Specifically, the server extracts the background feature corresponding to the background area, and can extract the background feature corresponding to the background area through a deep neural network. Then, the self-attention weight is calculated using the background feature to obtain the self-attention background weight. The background feature is weighted with the self-attention background weight to obtain the self-attention background feature.
[0097] Step 210, fuse the self-attention background feature and the self-attention foreground feature to obtain a fused feature, and perform scene recognition based on the fused feature to obtain the image scene recognition result corresponding to the image to be recognized.
[0098] Among them, the fused feature refers to the feature obtained by fusing the self-attention background feature and the self-attention foreground feature. The image scene recognition result refers to the specific scene category corresponding to the image to be recognized. This scene category can be a scene name, a scene label, or a scene number, etc. For example, the image scene recognition result can be a city street, a highway, a park, a coffee shop, an office, a restaurant, etc.
[0099] Specifically, the server fuses the self-attention background feature and the self-attention foreground feature to obtain a fused feature. Among them, the fused feature can be a feature obtained by directly splicing the self-attention background feature and the self-attention foreground feature. The fused feature can also be a feature obtained by performing vector operations on the vectors corresponding to the features. For example, the sum, product, etc. of the vectors corresponding to the self-attention background feature and the self-attention foreground feature can be calculated to obtain the fused feature. Then, the fused feature is used for scene recognition to obtain the image scene recognition result corresponding to the image to be recognized. For example, a convolutional neural network can be used to perform scene recognition on the fused feature to obtain the image scene recognition result corresponding to the image to be recognized.
[0100] In the above image scene recognition method, the foreground area and the background area in the image to be recognized are extracted, and then the self-attention weight is calculated based on the foreground features corresponding to the foreground area to obtain the self-attention foreground weight, and the foreground features are adjusted by the self-attention foreground weight to obtain the self-attention foreground features. The self-attention weight is calculated based on the background features corresponding to the background area to obtain the self-attention background weight, and the background features are adjusted by the self-attention background weight to obtain the self-attention background features. The self-attention foreground features and the self-attention background features are fused and then scene recognition is performed, that is, the image scene is recognized by the combined action of the background area and the foreground area, so that sufficient scene recognition features can be extracted, thereby improving the accuracy of image scene recognition.
[0101] In one embodiment, as Figure 3 shown, the image scene recognition method includes:
[0102] Step 302, input the image to be recognized into the image scene recognition model.
[0103] Step 304, the image scene recognition model extracts the foreground area and the background area in the image to be recognized, inputs the foreground area into the foreground branch network, and inputs the background area into the background branch network.
[0104] Among them, the image scene recognition model refers to a model used for image scene recognition obtained by training through a neural network algorithm using training data. The image scene recognition model includes two branch networks, namely the foreground branch network and the background branch network. The foreground branch network is a network for extracting self-attention foreground features, and the background branch network is a network for extracting self-attention foreground features. Both of these two branch networks are networks with self-attention mechanisms. In one embodiment, the network structures of the foreground branch network and the background branch network are the same, but the network parameters are different. In one embodiment, the network structures and network parameters of the foreground branch network and the background branch network are both different. The network structure of the branch network can be set as needed, and the network parameters are obtained through training.
[0105] Specifically, the server pre-trains an image scene recognition model and deploys the image scene recognition model to the server for use. When the server obtains an image to be recognized, it can directly input the image to be recognized into the image scene recognition model. The image scene recognition model receives the input image to be recognized, extracts the image features corresponding to the image to be recognized, and performs region division based on the image features to obtain the foreground area and the background area. Then the foreground area is input into the foreground branch network of the image scene recognition model, and at the same time the background area is input into the background branch network of the image scene recognition model.
[0106] Step 306: The foreground branch network extracts the foreground features corresponding to the foreground region, calculates the self-attention weights using the foreground features to obtain the self-attention foreground weights, and weights the foreground features with the self-attention foreground weights to obtain the self-attention foreground features.
[0107] Specifically, the foreground branch network can extract the foreground features corresponding to the foreground region through a foreground feature extraction network. This foreground feature extraction network can be a convolutional neural network, a recurrent neural network, a long short-term memory network, a feedforward neural network, etc. Then, the self-attention weights are calculated for the foreground features to obtain the self-attention foreground weights. For example, the foreground features can be pooled and then compressed to extract the important information in the foreground features, and then the compressed features are weight-mapped to obtain the self-attention foreground weights. Then, the foreground features are weighted with the self-attention foreground weights to obtain the self-attention foreground features.
[0108] Step 308: The background branch network extracts the background features corresponding to the background region, calculates the self-attention weights using the background features to obtain the self-attention background weights, and weights the background features with the self-attention background weights to obtain the self-attention background features.
[0109] Specifically, the background branch network can also extract the background features corresponding to the background region through a background feature extraction network. This background feature extraction network can be a convolutional neural network, a recurrent neural network, a long short-term memory network, a feedforward neural network, etc. The network structure of this background feature extraction network is the same as that of the foreground feature extraction network, but the network parameters are different. In one embodiment, the network structure of the background feature extraction network and the foreground feature extraction network can also be different. When the background features are obtained, the background branch network calculates the self-attention weights based on the background features. For example, the background branch network can pool and then compress the background features to extract the important information in the foreground features, and then weight-map the compressed features to obtain the self-attention background weights. Then, the background features are weighted with the self-attention background weights to obtain the self-attention background features.
[0110] Step 310: The image scene recognition model fuses the self-attention background features and the self-attention foreground features to obtain the fused features, and performs scene recognition based on the fused features to obtain the image scene recognition result.
[0111] Specifically, the image scene recognition model can splice the self-attention background feature and the self-attention foreground feature, that is, connect the beginning and the end of the self-attention background feature and the self-attention foreground feature to obtain a fused feature, perform scene recognition on the fused feature to obtain an image scene recognition result, and then output the image scene recognition result. In one embodiment, the image scene recognition model can also perform vector operations on the self-attention background feature and the self-attention foreground feature. For example, perform a vector product operation, a vector sum operation, a dot product operation, etc. to obtain a fused feature.
[0112] In the above embodiment, the trained image scene recognition model is used to perform scene recognition on the image to be recognized, that is, extract the foreground region and the background region in the image to be recognized, input the foreground region into the foreground branch network, and input the background region into the background branch network. The foreground branch network extracts the self-attention background feature, and the background branch network extracts the self-attention foreground feature. The self-attention background feature and the self-attention foreground feature are fused to obtain a fused feature. Scene recognition is performed based on the fused feature to obtain the output image scene recognition result. Since the image scene recognition model simultaneously extracts the self-attention background feature and the self-attention foreground feature from the background region and the foreground region through the dual-branch network, and then performs the image scene recognition result through the self-attention background feature and the self-attention foreground feature, the accuracy of the image scene recognition result can be improved.
[0113] In one embodiment, the image scene recognition model includes an image feature extraction network; extracting the foreground region and the background region in the image to be recognized includes:
[0114] Input the image to be recognized into the image feature extraction network for feature extraction to obtain the image feature to be recognized; perform region division based on the image feature to be recognized to obtain the foreground region and the background region.
[0115] Among them, the image feature extraction network is used to extract features from the image to be recognized. The image feature extraction network can be a deep neural network, such as a convolutional neural network. The image feature to be recognized is used to represent the image to be recognized and is the high-dimensional image feature extracted by the deep neural network. This high-dimensional image feature is usually used to represent the foreground of the image to be recognized.
[0116] Specifically, the image scene recognition model inputs the image to be recognized into the image feature extraction network for feature extraction, obtains the image features to be recognized as the output, and uses the image features to be recognized for region division to obtain the foreground region and the background region. Among them, the region can be divided into the foreground region and the background region by the magnitude of the feature values in the image features to be recognized. The image to be recognized can also be divided by the image activation degree in the image features to be recognized to obtain the foreground region and the background region, and the image activation degree is used to characterize the importance of the corresponding image features to be recognized. In a specific embodiment, the image feature extraction network shown in Table 1 below is used to extract features from the image to be recognized, and the image features to be recognized as the output are obtained. The image feature extraction network is a resnet101 (residual network) network, and the image feature extraction network includes five convolutional layers. The input is the image to be recognized, and the output of the fifth convolutional layer is a 7*7*2048 feature map.
[0117] Table 1 Image Feature Extraction Network Structure Table (ResNet-101 Structure Table)
[0118]
[0119]
[0120] Among them, as Figure 4 shown, it is a schematic diagram of the structure of a block (module). The input of 256 dimensions is reduced to 64 dimensions through a 1X1 convolution, and finally restored through a 1X1 convolution. Among them, the Relu (Rectified Linear Unit) function is used as the activation function. This structure can alleviate the problems of model degradation and gradient disappearance to a certain extent.
[0121] In one embodiment, as Figure 5 shown, based on the image features to be recognized for region division, obtaining the foreground region and the background region includes:
[0122] Step 502, calculate the mean value corresponding to the feature values in the image features to be recognized.
[0123] Specifically, the server calculates the sum and the number of each eigenvalue in the image features to be recognized, then calculates the ratio of the sum of the eigenvalues to the number of eigenvalues to obtain the feature mean corresponding to the image features to be recognized. The feature mean is used as the threshold for dividing the foreground region and the background region of the image features to be recognized. In one embodiment, the median, mode, quantile, variance, or standard deviation, etc. corresponding to the eigenvalues in the image features to be recognized can also be calculated as the threshold for dividing the foreground region and the background region of the image features to be recognized. In one embodiment, the threshold corresponding to the image features to be recognized can also be calculated by the gray histogram algorithm. In one embodiment, the threshold corresponding to the image features to be recognized can be calculated by the maximum inter-class algorithm.
[0124] Step 504: Based on the mean value, perform binary division on the image features to be recognized to obtain a foreground mask.
[0125] Specifically, the server uses the mean value to perform binary division on the image to be recognized, that is, replaces the eigenvalues in the image features to be recognized that exceed the mean value with 1, and replaces the eigenvalues that do not exceed the mean value with 0, to obtain a foreground mask. Among them, exceeding the mean value represents the foreground region, and not exceeding the mean value represents the background region.
[0126] Step 506: Calculate the product of the foreground mask and the pixel values of the image to be recognized to obtain the foreground region in the image to be recognized.
[0127] Specifically, the server calculates the product of the foreground mask and the pixel values in the image to be recognized to obtain a binarized image, and then obtains the foreground region in the image to be recognized from the binarized image. In a specific embodiment, as Figure 6 shown, it is a schematic diagram of the foreground mask obtained by performing foreground mask extraction on three different pictures. Looking from top to bottom, for the first picture of a puppy, the foreground region extracted by performing foreground mask is the region of the puppy in the first picture. For the second picture of a street, the foreground region extracted by performing foreground mask is the region of the billboard in the second picture. For the third picture of a house, the foreground region extracted by performing foreground mask is the region of the house in the third picture.
[0128] Step 508: Invert the foreground mask to obtain a background mask, and calculate the product of the background mask and the pixel values of the image to be recognized to obtain the background region.
[0129] Specifically, the server inverts the foreground mask, that is, uses 1 to subtract the value of the foreground mask to obtain a background mask, and then calculates the product of the background mask and the pixel values of the image to be recognized to obtain a binarized image, and obtains the background region in the image to be recognized from the binarized image.
[0130] In the above embodiments, by using the features of the image to be recognized to divide the image to be recognized, a background region and a foreground region are obtained, which can improve the accuracy of the obtained background region and foreground region.
[0131] In one embodiment, the foreground branch network includes a foreground feature extraction network and a foreground attention feature extraction network; calculating self-attention weights based on the foreground features corresponding to the foreground region to obtain self-attention foreground weights, and adjusting the foreground features through the self-attention foreground weights to obtain self-attention foreground features, including:
[0132] Input the foreground region into the foreground feature extraction network for feature extraction to obtain the foreground features corresponding to the foreground region; input the foreground features into the foreground attention weight feature network for attention weight calculation to obtain self-attention foreground weights, and use the self-attention foreground weights to weight the foreground features to obtain self-attention foreground features.
[0133] Among them, the foreground feature extraction network refers to a network for extracting features from the foreground region, and the foreground attention feature extraction network refers to a network for extracting self-attention features from the foreground features.
[0134] Specifically, each branch network in the image scene recognition model has a corresponding feature extraction network and attention feature extraction network, that is, each branch network can extract features from the input image region and then perform attention feature extraction. That is, the foreground region can be input into the foreground feature extraction network for feature extraction to obtain the foreground feature map corresponding to the foreground region, and then the foreground features are input into the foreground attention weight feature network for attention weight calculation to obtain self-attention foreground weights, and the self-attention foreground weights are used to weight the foreground features to obtain self-attention foreground features. In a specific embodiment, the network structure of the foreground feature extraction network can be the network structure shown in Table 1, and the network parameters are obtained through training.
[0135] In one embodiment, as Figure 7 shown, input the foreground features into the foreground attention weight feature network for attention weight calculation to obtain self-attention foreground weights, and use the self-attention foreground weights to weight the foreground features to obtain self-attention foreground features, including:
[0136] Step 702, perform mean pooling on the foreground features through the mean pooling layer in the foreground attention feature extraction network to obtain foreground pooled features.
[0137] Among them, the mean pooling layer is used to perform mean pooling on the foreground attention features, that is, to reduce the dimension of the foreground attention features. The foreground pooled features refer to the features obtained after performing mean pooling on the foreground attention features.
[0138] Specifically, the server performs mean pooling on the foreground features through the mean pooling layer in the foreground attention feature extraction network to obtain foreground pooled features. For example, if the input foreground attention feature is a feature map with a dimension of 7*7*2048, the foreground pooled feature obtained through the mean pooling layer is a feature with a dimension of 1*1*2048, that is, a 1*2048 vector, which is used to represent the activation mean of 2048 different channels in the deep learning network layer in this foreground area. In one embodiment, max pooling can also be performed through the max pooling layer in the foreground attention feature extraction network to obtain foreground pooled features.
[0139] Step 704: Use the non-linear compression layer in the foreground attention feature extraction network to non-linearly compress the foreground pooled features to obtain foreground compressed features.
[0140] Among them, the non-linear compression layer is used for non-linear compression to achieve the extraction of important information in the foreground attention features. The foreground compressed features refer to the features obtained after non-linearly compressing the foreground pooled features.
[0141] Specifically, the server can non-linearly compress the foreground pooled features through the non-linear compression layer. For example, non-linearly compress the 1*2048-dimensional foreground vector to 64 dimensions to achieve the extraction of important information in the foreground features.
[0142] Step 706: Activate the foreground compressed features through the activation function layer in the foreground attention feature extraction network to obtain foreground activated features.
[0143] Specifically, the activation function layer is used to activate the foreground compressed features. The foreground activated features refer to the features obtained after activation through the activation function. Among them, activation can be performed through the RELU (Rectified Linear Unit) activation function, or through the S-shaped activation function, or through the Tanh (Hyperbolic Tangent) activation function. For example, the 64-dimensional foreground compressed features can be activated using the RELU function to obtain 64-dimensional foreground activated features.
[0144] Step 708: Based on the weight mapping layer in the foreground attention feature extraction network, perform weight mapping on the foreground activated features to obtain self-attention foreground weights.
[0145] Among them, the weight mapping layer is used for self-attention weight mapping, that is, mapping the refined features to weight vectors.
[0146] Specifically, the server inputs the foreground activation features into the weight mapping layer for weight mapping to obtain the self-attention foreground weights. For example, the server inputs the 64-dimensional foreground activation features into the weight mapping layer for weight mapping to obtain the output 2048-dimensional self-attention foreground weight vector, that is, the vector is used to represent the weight vectors corresponding to 2048 different channels of the deep learning network layer.
[0147] Step 710, use the self-attention foreground weights to weight the feature values in the foreground features to obtain weighted foreground features, and perform max pooling on the weighted foreground features through the max pooling layer in the foreground attention feature extraction network to obtain self-attention foreground features.
[0148] Specifically, the server uses the self-attention foreground weights to weight the feature values in the foreground features to obtain weighted foreground features. Input the weighted foreground features into the max pooling layer for max pooling to obtain self-attention foreground features. That is, the server uses the 2048-dimensional self-attention foreground weight vector to weight each channel in the foreground features to obtain the self-attention feature map. Then perform max pooling on the self-attention feature map to obtain the 2048-dimensional self-attention foreground feature vector. In a specific embodiment, the network structure of the foreground attention feature extraction network is shown in Table 2 below.
[0149] Table 2 Network structure of the self-attention feature extraction network
[0150]
[0151]
[0152] In the above embodiment, the self-attention background features extracted by the foreground feature extraction network and the foreground attention weight extraction network in the foreground branch network can make the extracted self-attention background features more accurate.
[0153] In one embodiment, the background branch network includes a background feature extraction network and a background attention weight extraction network;
[0154] Based on the background features corresponding to the background region, calculate the self-attention weights to obtain the self-attention background weights, and adjust the background features through the self-attention background weights to obtain the self-attention background features, including:
[0155] Input the background region into the background feature extraction network for feature extraction to obtain the background features corresponding to the background region; input the background features into the background attention weight feature network for attention weight calculation to obtain the self-attention background weights, and use the self-attention background weights to weight the background features to obtain the self-attention background features.
[0156] Among them, the background feature extraction network refers to a network that extracts features corresponding to the background region. The background attention weight feature network refers to a network that performs self-attention feature extraction on foreground features.
[0157] Specifically, the background branch network of the image scene recognition model inputs the background region into the background feature extraction network for feature extraction to obtain the background features corresponding to the background region. In a specific embodiment, the background feature extraction network trained using the network structure shown in Table 1 can be used and then applied. Then, the background features are input into the background attention weight feature network for attention weight calculation to obtain the self-attention background weight, and the product of the self-attention background weight and the background features is calculated to obtain the self-attention background features.
[0158] In one embodiment, as Figure 8 shown, inputting the background features into the background attention weight feature network for attention weight calculation to obtain the self-attention background weight, and using the self-attention background weight to weight the background features to obtain the self-attention background features, including:
[0159] Step 802, perform mean pooling on the background features through the mean pooling layer in the background attention feature extraction network to obtain the background pooled features.
[0160] Among them, the mean pooling layer in the background attention feature extraction network is used to perform mean pooling on the background features. The background pooled feature digitalization is the feature obtained after performing mean pooling on the background features.
[0161] Specifically, the server inputs the background features into the mean pooling layer in the background attention feature extraction network for mean pooling to obtain the background pooled features. For example, inputting the background features with a dimension of 7*7*2048 into the mean pooling layer in the background attention feature extraction network to obtain an output feature vector with a dimension of 1*2048. This vector is used to represent the activation mean of 2048 different channels in the deep learning network layer in this background region. In one embodiment, max pooling can also be performed through the max pooling layer in the background attention feature extraction network to obtain the background pooled features.
[0162] Step 804, perform non-linear compression on the background pooled features using the non-linear compression layer in the background attention feature extraction network to obtain the background compressed features.
[0163] Among them, the non-linear compression layer in the background attention feature extraction network is used for non-linear compression. The background compressed feature refers to the feature obtained after non-linear compression.
[0164] Specifically, the server inputs the background pooled features into the non-linear compression layer in the background attention feature extraction network to obtain the output background compressed features, that is, the 1*2048-dimensional background vector is compressed to 64 dimensions through non-linear compression to refine the important information in the background features.
[0165] Step 806: Activate the background compressed features through the activation function layer in the background attention feature extraction network to obtain the background activated features.
[0166] Among them, the activation function layer in the background attention feature extraction network is used to perform activation using an activation function. The activation function can be a RELU function, an S-shaped activation function, a tanh activation function, etc. The background activated features refer to the features obtained after activating the background compressed features.
[0167] Specifically, the server inputs the background compressed features into the activation function layer in the background attention feature extraction network for activation to obtain the background activated features. For example, the 64-dimensional background activated features are activated using the RELU activation function to obtain the background activated features.
[0168] Step 808: Perform weight mapping on the background activated features based on the weight mapping layer in the background attention feature extraction network to obtain the self-attention background weights.
[0169] Among them, the weight mapping layer in the background attention feature extraction network is used to perform weight mapping on the background activated features.
[0170] Specifically, the server inputs the background activated features into the weight mapping layer in the background attention feature extraction network for weight mapping to obtain the self-attention background weights. For example, the server inputs the 64-dimensional background activated features into the weight mapping layer for weight mapping to obtain the output 2048-dimensional self-attention background weight vector, that is, the background weight vector is used to represent the weight vectors corresponding to 2048 different channels in the deep learning network layer.
[0171] Step 810: Weight the feature values in the background features using the self-attention background weights to obtain the weighted background features, and perform max pooling on the weighted background features through the max pooling layer in the background attention feature extraction network to obtain the self-attention background features.
[0172] Specifically, the server calculates the product of the self-attention background weights and the background features to obtain weighted background features, and then performs max pooling on the weighted background features through the max pooling layer in the background attention feature extraction network to obtain self-attention background features. That is, the server weights each channel in the background features using a 2048-dimensional self-attention background weight vector to obtain a self-attention background feature map. Then, max pooling is performed on the self-attention background feature map to obtain a 2048-dimensional self-attention background feature vector. In a specific embodiment, the background attention feature extraction network can be trained and used with the network structure shown in Table 2.
[0173] In the above embodiment, the self-attention background features extracted by the background feature extraction network and the background attention weight extraction network in the background branch network can make the extracted self-attention background features more accurate.
[0174] In one embodiment, the image scene recognition model includes a fusion output network; the self-attention background features and the self-attention foreground features are fused to obtain fused features, and scene recognition is performed based on the fused features to obtain an image scene recognition result, including:
[0175] The self-attention background features and the self-attention foreground features are concatenated through the fusion layer in the fusion output network to obtain concatenated features; the concatenated features are input into the fully connected layer in the fusion output network for scene recognition to obtain an image scene recognition result.
[0176] Among them, the fusion output network is a network for fusing features and performing image scene recognition. The fusion layer in the fusion output network is used to fuse features, and the fully connected layer in the fusion output network is used to perform image scene recognition and output an image scene recognition result.
[0177] Specifically, the server concatenates the heads and tails of the self-attention background features and the self-attention foreground features through the fusion layer in the fusion output network to obtain concatenated features, and the server inputs the concatenated features into the fully connected layer in the fusion output network for multi-class scene recognition to obtain the probabilities of each image scene category, and the scene category with the highest probability is used as the output image scene recognition result according to the probabilities of the image scene categories.
[0178] In a specific embodiment, the network structure of the fusion output network is as shown in Table 3 below.
[0179] Table 3 Network structure table of the fusion output network
[0180]
[0181] Among them, N represents the number of categories of image scenes.
[0182] In the above embodiment, the self-attention background feature and the self-attention foreground feature are fused through the fusion output network, and the fused feature is used for multi-class scene recognition, thereby improving the accuracy of image scene recognition.
[0183] In one embodiment, as Figure 9 shown, a method for training an image scene recognition model is provided. Taking the server in Figure 1 as an example for illustration, it can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The method includes the following steps:
[0184] Step 902, obtain a training image and a corresponding training scene label, and input the training image into the initial image scene recognition model.
[0185] Among them, the training image refers to an image with a training scene label during training, and the training scene label refers to a label of a specific scene category corresponding to the training image. The initial image scene recognition model refers to an image scene recognition model with initialized model parameters.
[0186] Specifically, the server can directly obtain the training image and the corresponding training scene label from a database, can also collect the training image and the corresponding training scene label from the Internet, or can also obtain the training image and the corresponding training scene label from a service provider providing data services. Input the training image into the initial image scene recognition model for image scene recognition, and pre-establish an image scene recognition model with initialized model parameters. Among them, the initialization can be random initialization, zero initialization, Gaussian distribution initialization, etc. For example, the image scene recognition model can be initialized with a Gaussian distribution with a variance of 0.01 and a mean of 0. In one embodiment, the feature extraction parameters in the initial image scene recognition model can be pre-trained, and other parameters can be initialized using a Gaussian distribution.
[0187] Step 904, the initial image scene recognition model extracts the initial foreground training region and the initial background training region in the training image, inputs the initial foreground region into the initial foreground branch network, and inputs the initial background region into the initial background branch network.
[0188] Among them, the initial foreground training region refers to the foreground region in the training image extracted by the initial image scene recognition model. The initial background training region refers to the background region in the training image extracted by the initial image scene recognition model. The initial foreground branch network refers to a foreground branch network with initialized parameters. The initial background branch network refers to a background branch network with initialized parameters.
[0189] Specifically, the initial image scene recognition model in the server can extract the initial image features in the training images through the initial image feature extraction network, extract the initial foreground training region and the initial background training region in the training images according to the initial image features, and then input the initial foreground region into the initial foreground branch network, while inputting the initial background region into the initial background branch network. Among them, the initial image feature extraction network is an image feature extraction network with initialized parameters, which is used to extract features from images. The initialized parameters of this initial image feature extraction network can also be obtained through pre-training.
[0190] Step 906: The initial foreground branch network calculates the self-attention weight based on the initial foreground features corresponding to the initial foreground training region, obtains the initial self-attention foreground weight, and adjusts the initial foreground features through the initial self-attention foreground weight to obtain the initial self-attention foreground features.
[0191] Among them, the initial foreground features refer to the features corresponding to the extracted initial foreground training region. The initial self-attention foreground weight refers to the self-attention foreground weight corresponding to the initial foreground training region. The initial self-attention foreground features refer to the self-attention foreground features corresponding to the initial foreground training region.
[0192] Specifically, the initial foreground branch network inputs the initial foreground training region into the initial foreground feature extraction network for feature extraction to obtain the initial foreground features. The initialized parameters of this initial foreground feature extraction network can be obtained through pre-training. Then, the initial foreground features are input into the initial foreground attention feature extraction network for self-attention weight calculation to obtain the initial self-attention foreground weight, and the initial foreground features are weighted by the initial self-attention foreground weight to obtain the initial self-attention foreground features. Among them, the initialized parameters of the initial foreground attention feature extraction network can be the parameters obtained using the Gaussian distribution.
[0193] Step 908: The initial background branch network calculates the self-attention weight based on the initial background features corresponding to the initial background training region, obtains the initial self-attention background weight, and adjusts the initial background features through the initial self-attention background weight to obtain the initial self-attention background features.
[0194] Among them, the initial background features refer to the background features corresponding to the initial background training region. The initial self-attention background weight refers to the self-attention background weight corresponding to the initial background training region. The initial self-attention background features refer to the self-attention background features corresponding to the initial background training region.
[0195] Specifically, the initial background branch network inputs the initial background training region into the initial background region feature extraction network for feature extraction to obtain the initial background features. Here, the initial background region feature extraction network is a network for extracting features from the initial background training region, and the initialization parameters of this initial background region feature extraction network can be obtained through pre-training or initialization. The initial background features are input into the initial background attention feature extraction network for self-attention weight calculation to obtain the initial self-attention background weights, and the initial background features are weighted by the initial self-attention background weights to obtain the initial self-attention background features. Among them, the initialization parameters in the initial background attention feature extraction network can be parameters obtained using a Gaussian distribution.
[0196] Step 910, the initial image scene recognition model fuses the initial self-attention background features and the initial self-attention foreground features to obtain the initial fusion features, and performs scene recognition based on the initial fusion features to obtain the initial image scene recognition result.
[0197] Among them, the initial fusion features refer to the features obtained after the initial image scene recognition model performs feature fusion, and the initial image scene recognition result refers to the image scene recognition result output by the initial image scene recognition model.
[0198] Specifically, the initial image scene recognition model concatenates the initial self-attention background features and the initial self-attention foreground features end to end to obtain the initial fusion features, and inputs the initial fusion features into the initial fully connected network for scene recognition to obtain the initial image scene recognition result. Among them, the initial fully connected network is a multi-class fully connected network, and the initial parameters of this initial fully connected network are initialized through a Gaussian distribution.
[0199] Step 912, calculate the loss information between the initial image scene recognition result and the training scene label, and update the initial image scene recognition model based on the loss information.
[0200] Specifically, the server uses a loss function to calculate the error between the initial image scene recognition result and the training scene label to obtain the loss information.
[0201] Then use the loss information to calculate the gradient, and use the gradient descent algorithm to update the parameters in the initial image scene recognition model in reverse to obtain the updated image scene recognition model.
[0202] Step 914, determine whether the training completion condition is reached. When the training completion condition is reached, execute step 916. When the training completion condition is not reached, return to step 902 for iterative execution, that is, return to the step of inputting the training image into the initial image scene recognition model for iterative execution.
[0203] Step 916: Obtain the trained image scene recognition model.
[0204] The training completion condition refers to the condition for the completion of the training of the image scene recognition model, including at least one of the following: the loss information obtained from the training conforms to a preset loss threshold, the number of training iterations reaches the maximum number of iterations, and the model parameters do not change significantly.
[0205] Specifically, the server determines whether the model training has reached the training completion condition. When the training completion condition is not reached, it returns to the step of inputting the training images into the initial image scene recognition model and iteratively executes until the training completion condition is reached. The image scene recognition model that reaches the training completion condition is used as the trained image scene recognition model.
[0206] In the above image scene recognition model training method, by obtaining training images and corresponding training scene labels, the training images are input into the initial image scene recognition model. The initial image scene recognition model extracts self-attention foreground features through the initial foreground branch network and self-attention background features through the initial background branch network. Then, the initial self-attention background features and the initial self-attention foreground features are fused to obtain initial fusion features. Based on the initial fusion features, scene recognition is performed to obtain the initial image scene recognition result. The loss information between the initial image scene recognition result and the training scene label is calculated, and the initial image scene recognition model is updated based on the loss information until the training completion condition is reached, and the trained image scene recognition model is obtained. Since the self-attention foreground features and the self-attention background features are respectively extracted through the foreground branch network and the background branch network, and then the initial self-attention background features and the initial self-attention foreground features are fused, the image scene recognition result is obtained, so that the trained image scene recognition model can improve the accuracy of image scene recognition.
[0207] In one embodiment, as Figure 10 shown, calculating the loss information between the initial image scene recognition result and the training scene label, updating the initial image scene recognition model based on the loss information, and returning to the step of inputting the training images into the initial image scene recognition model and iteratively executing until the training completion condition is reached to obtain the trained image scene recognition model, includes:
[0208] Step 1002: Use the cross-entropy loss function to calculate the error between the initial image scene recognition result and the training scene label to obtain the loss information;
[0209] Specifically, the cross-entropy loss function can be used to calculate the error, that is, the loss information can be calculated using the following formula (1).
[0210]
[0211] Among them, L represents loss information, and y refers to the training scenario label, which is the true scenario category label of the image. refers to the initial image scene recognition result, which is the predicted scenario category.
[0212] Step 1004: When the loss information does not exceed the preset loss threshold, calculate the gradient based on the loss information, and use the gradient to update the initial image scene recognition model to obtain an updated image scene recognition model.
[0213] Among them, the preset loss threshold refers to the threshold of the pre-set loss information. The gradient is originally a vector, indicating that the directional derivative of a certain function at this point takes the maximum value along this direction, that is, the function changes fastest and has the largest change rate at this point along this direction (the direction of this gradient). The updated image scene recognition model is the image scene recognition model after parameter update.
[0214] Specifically, when the server determines that the loss information does not exceed the preset loss threshold, it calculates the gradient based on the loss information, and uses the gradient to update each parameter in the initial image scene recognition model, that is, update the parameters of the initial background branch network, update the parameters of the initial foreground branch network, update the network parameters of the initial fusion output network, and update the network parameters of the initial image feature extraction network, etc., to obtain an updated image scene recognition model.
[0215] Step 1006: Use the updated scene recognition model as the initial scene recognition model, and return to iterate the step of inputting the training image into the initial image scene recognition model until the loss information exceeds the preset loss threshold. Then, use the initial image scene recognition model that exceeds the preset loss threshold as the trained image scene recognition model.
[0216] Specifically, the server uses the updated scene recognition model as the initial scene recognition model, and returns to iterate the step of inputting the training image into the initial image scene recognition model until the loss information exceeds the preset loss threshold. Then, use the initial image scene recognition model that exceeds the preset loss threshold as the trained image scene recognition model. By using the cross-entropy loss function to train the image scene recognition model, the performance of the trained image scene recognition model can be better.
[0217] In one embodiment, the initial image scene recognition model includes an initial image feature extraction network, an initial foreground feature extraction network, and an initial background feature extraction network;
[0218] Such as Figure 11 shown, before step 902, that is, before obtaining the training image and the corresponding training scenario label, it further includes:
[0219] Step 1102: Obtain the pre-trained image and the pre-trained scene label.
[0220] The pre-trained image refers to the image used during pre-training. The pre-trained scene label refers to the scene category label corresponding to the pre-trained image during pre-training. Each pre-trained image has a corresponding pre-trained scene label, which is the true scene category label.
[0221] Specifically, the server can obtain the saved pre-trained image and pre-trained scene label from the database, or can collect the pre-trained image from the Internet and then obtain the pre-trained scene label corresponding to the pre-trained image. It can also obtain the training data, i.e., the pre-trained image and pre-trained scene label, from the service provider that provides data services.
[0222] Step 1104: Input the pre-trained image into the pre-trained scene recognition model. The pre-trained scene recognition model extracts features from the pre-trained image through the feature extraction network to obtain the pre-trained image features, and performs scene recognition based on the pre-trained image features to obtain the pre-trained image scene recognition result.
[0223] The pre-trained scene recognition model is the scene recognition model during pre-training. This pre-trained scene recognition model can be a model established using a deep neural network, and the model parameters of this pre-trained scene recognition model are randomly initialized. The pre-trained scene recognition model includes a feature extraction network, which is used to extract the features of the image. The pre-trained image features refer to the image features extracted from the pre-trained image. The pre-trained image scene recognition result refers to the predicted image scene category corresponding to the pre-trained image.
[0224] Specifically, the server inputs the pre-trained image into the pre-trained scene recognition model. The pre-trained scene recognition model extracts features from the pre-trained image through the feature extraction network to obtain the pre-trained image features, and performs scene recognition based on the pre-trained image features through the fully connected network in the pre-trained scene recognition model to obtain the pre-trained image scene recognition result.
[0225] Step 1106: Calculate the pre-training loss information based on the pre-trained scene recognition result and the pre-trained scene label, and update the pre-trained scene recognition model based on the pre-training loss information.
[0226] The pre-training loss information refers to the loss information obtained during pre-training.
[0227] Specifically, the server can calculate the pre-training loss information between the pre-trained scene recognition result and the pre-trained scene label using a classification loss function, and then use the pre-training loss information to reversely update the parameters in the pre-trained scene recognition model based on the gradient descent algorithm. Among them, the classification loss function can be a cross-entropy loss function, an exponential loss function, an S-shaped loss function, and so on.
[0228] Step 1108, determine whether the pre-training is completed. When the pre-training is not completed, return to the step of inputting the pre-training image into the pre-trained scene recognition model and execute it iteratively. When the pre-training is completed, execute step 1110.
[0229] Step 1110, obtain the initial image feature extraction network, the initial foreground feature extraction network, and the initial background feature extraction network in the initial image scene recognition model based on the pre-trained feature extraction network that has been completed.
[0230] Specifically, the server determines whether the pre-training is completed, that is, determines whether the pre-training completion condition is met. The pre-training completion condition includes at least one of the pre-training loss information reaching the preset pre-training loss threshold, the pre-training iteration times reaching the maximum iteration times, and the model parameters obtained by pre-training not changing significantly. When the pre-training is not completed, return to the step of inputting the pre-training image into the pre-trained scene recognition model and execute it iteratively. When the pre-training is completed, obtain the pre-trained pre-training image scene recognition model, and use the feature extraction network in the pre-trained pre-training image scene recognition model as the initial image feature extraction network, the initial foreground feature extraction network, and the initial background feature extraction network in the initial image scene recognition model. That is, the network structure and network parameters of the feature extraction network in the pre-trained image scene recognition model are the same as those of the initial image feature extraction network, the network structure and network parameters of the feature extraction network in the pre-trained image scene recognition model are the same as those of the initial foreground feature extraction network, and the network structure and network parameters of the feature extraction network in the pre-trained image scene recognition model are the same as those of the initial background feature extraction network. Use the initial image feature extraction network, the initial foreground feature extraction network, and the initial background feature extraction network to establish the initial image scene recognition model, and then train the established initial image scene recognition model to obtain the image scene recognition model.
[0231] In the above embodiments, an initial image feature extraction network, an initial foreground feature extraction network, and an initial background feature extraction network are obtained through pre-training, and then an initial image scene recognition model is established for training to obtain an image scene recognition model, which can improve the convergence speed when training the image scene recognition model and improve the training efficiency and accuracy. In a specific embodiment, the initial network parameters of the initial image feature extraction network, the initial foreground feature extraction network, and the initial background feature extraction network can use the parameters of ResNet101 pre-trained on the ImageNet dataset, and the parameters of newly added networks such as the self-attention feature extraction network and the fusion output network are initialized using a Gaussian distribution with a variance of 0.01 and a mean of 0.
[0232] In a specific embodiment, as Figure 12 shown, an image scene recognition method is provided, which is executed by a server and specifically includes the following steps:
[0233] Step 1202, the server obtains the image to be recognized from the terminal and inputs the image to be recognized into the image scene recognition model.
[0234] Step 1204, the image scene recognition model inputs the image to be recognized into the image feature extraction network for feature extraction to obtain the features of the image to be recognized.
[0235] Step 1206, the image scene recognition model calculates the mean value corresponding to the feature values in the features of the image to be recognized, and based on the mean value, the features of the image to be recognized are binary partitioned to obtain a foreground mask. Calculate the product of the foreground mask and the pixel values of the image to be recognized to obtain the foreground region in the image to be recognized.
[0236] Step 1208, the image scene recognition model takes the inverse of the foreground mask to obtain a background mask, calculates the product of the background mask and the pixel values of the image to be recognized to obtain the background region. Input the foreground region into the foreground branch network and input the background region into the background branch network.
[0237] Step 1210, the foreground branch network inputs the foreground region into the foreground feature extraction network for feature extraction to obtain the foreground features corresponding to the foreground region, inputs the foreground features into the foreground attention weight feature network for attention weight calculation to obtain the self-attention foreground weight, and uses the self-attention foreground weight to weight the foreground features to obtain the self-attention foreground features.
[0238] Step 1212, the background branch network inputs the background region into the background feature extraction network for feature extraction to obtain the background features corresponding to the background region, inputs the background features into the background attention weight feature network for attention weight calculation to obtain the self-attention background weight, and uses the self-attention background weight to weight the background features to obtain the self-attention background features.
[0239] In step 1214, the image scene recognition model splices the self-attention background feature and the self-attention foreground feature through the fusion layer in the fusion output network to obtain a spliced feature, inputs the spliced feature into the fully connected layer in the fusion output network for scene recognition to obtain the image scene recognition result, and the server returns the image scene recognition result to the terminal for display.
[0240] In a specific embodiment, as Figure 13 shown, a schematic diagram of the architecture of an image scene recognition model is provided. The image scene recognition model is a dual-branch self-attention recognition model. Specifically:
[0241] The image to be recognized is obtained, and the image is input into the image scene recognition model. The image scene recognition model extracts the image depth feature to obtain the image depth feature, and then uses the image depth feature to perform foreground extraction and background extraction to obtain a foreground image and a background image. Then, the foreground image feature and the background image feature are respectively extracted through a dual-branch network. And based on the self-attention weight 2 extracted from the foreground image feature through the self-attention network, the foreground image feature is weighted by the self-attention weight 2 to obtain the foreground classification feature. Based on the self-attention weight 1 extracted from the background image feature through the self-attention network, the background image feature is weighted by the self-attention weight 1 to obtain the background classification feature. Then, the background classification feature and the foreground classification feature are fused by connecting the head and the tail, and then the image scene is recognized through the fused feature to obtain the image scene recognition result, that is, the scene classification result.
[0242] The present application also provides an application scenario. The application of the image scene recognition method in this application scenario is as follows:
[0243] As Figure 14 shown, it is a schematic diagram of the application scenario of image scene recognition. Specifically, the user inputs a picture to be recognized through terminal A. For example, the picture can be as Figure 15 shown. Terminal A uploads the picture input by the user to the server. The image scene recognition model is deployed in the server. The image scene recognition model performs image scene recognition on the picture input by the user to obtain the image scene recognition result. For example, Figure 15 the recognition result can be a seaside scene, and then the image scene recognition result is sent to terminal B for display.
[0244] It should be understood that although Figure 2-12The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-12 At least some of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps.
[0245] In one embodiment, as Figure 16 shown, an image scene recognition device 1600 is provided. This device can be a software module, a hardware module, or a combination of both to form part of a computer device. Specifically, the device includes: an image acquisition module 1602, a region extraction module 1604, a foreground feature extraction module 1606, a background feature extraction module 1608, and a scene recognition module 1610, where:
[0246] The image acquisition module 1602 is configured to acquire an image to be recognized;
[0247] The region extraction module 1604 is configured to extract a foreground region and a background region from the image to be recognized;
[0248] The foreground feature extraction module 1606 is configured to calculate self-attention weights based on the foreground features corresponding to the foreground region to obtain self-attention foreground weights, and adjust the foreground features through the self-attention foreground weights to obtain self-attention foreground features;
[0249] The background feature extraction module 1608 is configured to calculate self-attention weights based on the background features corresponding to the background region to obtain self-attention background weights, and adjust the background features through the self-attention background weights to obtain self-attention background features;
[0250] The scene recognition module 1610 is configured to perform feature fusion on the self-attention background features and the self-attention foreground features to obtain fused features, and perform scene recognition based on the fused features to obtain an image scene recognition result corresponding to the image to be recognized.
[0251] In one embodiment, the image scene recognition device 1600 includes:
[0252] An image input module, configured to input the image to be recognized into the image scene recognition model;
[0253] A branch input module, which is used for an image scene recognition model to extract the foreground region and the background region in the image to be recognized, input the foreground region into the foreground branch network, and input the background region into the background branch network;
[0254] A foreground recognition module, which is used for the foreground branch network to extract the foreground features corresponding to the foreground region, calculate the self-attention weight using the foreground features to obtain the self-attention foreground weight, and weight the foreground features with the self-attention foreground weight to obtain the self-attention foreground features;
[0255] A background recognition module, which is used for the background branch network to extract the background features corresponding to the background region, calculate the self-attention weight using the background features to obtain the self-attention background weight, and weight the background features with the self-attention background weight to obtain the self-attention background features;
[0256] An image recognition module, which is used for the image scene recognition model to fuse the self-attention background features and the self-attention foreground features to obtain the fused features, and perform scene recognition based on the fused features to obtain the image scene recognition result.
[0257] In one embodiment, the image scene recognition model includes an image feature extraction network; the branch input module is further used to input the image to be recognized into the image feature extraction network for feature extraction to obtain the features of the image to be recognized; and perform region division based on the features of the image to be recognized to obtain the foreground region and the background region.
[0258] In one embodiment, the branch input module is further used to calculate the mean value corresponding to the feature values in the features of the image to be recognized; perform binary division on the features of the image to be recognized based on the mean value to obtain the foreground mask; calculate the product of the foreground mask and the pixel values of the image to be recognized to obtain the foreground region in the image to be recognized; invert the foreground mask to obtain the background mask, and calculate the product of the background mask and the pixel values of the image to be recognized to obtain the background region.
[0259] In one embodiment, the foreground branch network includes a foreground feature extraction network and a foreground attention feature extraction network; the foreground recognition module is further used to input the foreground region into the foreground feature extraction network for feature extraction to obtain the foreground features corresponding to the foreground region; input the foreground features into the foreground attention weight feature network for attention weight calculation to obtain the self-attention foreground weight, and weight the foreground features with the self-attention foreground weight to obtain the self-attention foreground features.
[0260] In one embodiment, the foreground recognition module is further configured to perform mean pooling on the foreground features through the mean pooling layer in the foreground attention feature extraction network to obtain foreground pooled features; perform non-linear compression on the foreground pooled features using the non-linear compression layer in the foreground attention feature extraction network to obtain foreground compressed features; activate the foreground compressed features through the activation function layer in the foreground attention feature extraction network to obtain foreground activated features; and perform weight mapping on the foreground activated features based on the weight mapping layer in the foreground attention feature extraction network to obtain self-attention foreground weights; use the self-attention foreground weights to weight the feature values in the foreground features to obtain weighted foreground features, and perform max pooling on the weighted foreground features through the max pooling layer in the foreground attention feature extraction network to obtain self-attention foreground features.
[0261] In one embodiment, the background branch network includes a background feature extraction network and a background attention weight extraction network; the background recognition module is further configured to input the background region into the background feature extraction network for feature extraction to obtain background features corresponding to the background region; input the background features into the background attention weight feature network for attention weight calculation to obtain self-attention background weights, and use the self-attention background weights to weight the background features to obtain self-attention background features.
[0262] In one embodiment, the background recognition module is further configured to perform mean pooling on the background features through the mean pooling layer in the background attention feature extraction network to obtain background pooled features; perform non-linear compression on the background pooled features using the non-linear compression layer in the background attention feature extraction network to obtain background compressed features; activate the background compressed features through the activation function layer in the background attention feature extraction network to obtain background activated features; and perform weight mapping on the background activated features based on the weight mapping layer in the background attention feature extraction network to obtain self-attention background weights; use the self-attention background weights to weight the feature values in the background features to obtain weighted background features, and perform max pooling on the weighted background features through the max pooling layer in the background attention feature extraction network to obtain self-attention background features.
[0263] In one embodiment, the image scene recognition model includes a fusion output network; the image recognition module is further configured to splice the self-attention background features and the self-attention foreground features through the fusion layer in the fusion output network to obtain spliced features; input the spliced features into the fully connected layer in the fusion output network for scene recognition to obtain the image scene recognition result.
[0264] In one embodiment, as Figure 17As shown, an image scene recognition model training device 1700 is provided. This device can be a software module, a hardware module, or a combination of both to form part of a computer device. Specifically, the device includes: a training data acquisition module 1702, a model processing module 1704, a foreground network processing module 1706, a background network processing module 1708, a model recognition module 1710, and an iteration module 1712, where:
[0265] The training data acquisition module 1702 is used to obtain training images and corresponding training scene labels, and input the training images into the initial image scene recognition model;
[0266] The model processing module 1704 is used to extract the initial foreground training region and the initial background training region from the training images by the initial image scene recognition model, input the initial foreground region into the initial foreground branch network, and input the initial background region into the initial background branch network;
[0267] The foreground network processing module 1706 is used to calculate the self-attention weight based on the initial foreground features corresponding to the initial foreground training region by the initial foreground branch network, obtain the initial self-attention foreground weight, and adjust the initial foreground features through the initial self-attention foreground weight to obtain the initial self-attention foreground features;
[0268] The background network processing module 1708 is used to calculate the self-attention weight based on the initial background features corresponding to the initial background training region by the initial background branch network, obtain the initial self-attention background weight, and adjust the initial background features through the initial self-attention background weight to obtain the initial self-attention background features;
[0269] The model recognition module 1710 is used to fuse the initial self-attention background features and the initial self-attention foreground features by the initial image scene recognition model to obtain the initial fusion features, and perform scene recognition based on the initial fusion features to obtain the initial image scene recognition result;
[0270] The iteration module 1712 is used to calculate the loss information between the initial image scene recognition result and the training scene label, update the initial image scene recognition model based on the loss information, and return to iteratively execute the step of inputting the training images into the initial image scene recognition model until the training completion condition is reached, and obtain the trained image scene recognition model.
[0271] In one embodiment, the iterative module 1712 is further configured to calculate the error between the initial image scene recognition result and the training scene label using a cross-entropy loss function to obtain loss information; when the loss information does not exceed a preset loss threshold, calculate the gradient based on the loss information, and update the initial image scene recognition model using the gradient to obtain an updated image scene recognition model; use the updated scene recognition model as the initial scene recognition model, and return to the step of inputting the training image into the initial image scene recognition model and iterate until the loss information exceeds the preset loss threshold, and use the initial image scene recognition model that exceeds the preset loss threshold as the trained image scene recognition model.
[0272] In one embodiment, the initial image scene recognition model includes an initial image feature extraction network, an initial foreground feature extraction network, and an initial background feature extraction network; the image scene recognition model training apparatus 1700 further includes:
[0273] A pre-training module, configured to obtain pre-training images and pre-training scene labels; input the pre-training images into a pre-training scene recognition model, the pre-training scene recognition model extracts features of the pre-training images through a feature extraction network to obtain pre-training image features, performs scene recognition based on the pre-training image features to obtain pre-training image scene recognition results; calculate pre-training loss information based on the pre-training scene recognition results and the pre-training scene labels, update the pre-training scene recognition model based on the pre-training loss information, and return to the step of inputting the pre-training images into the pre-training scene recognition model and iterate until pre-training is completed, and obtain the initial image feature extraction network, the initial foreground feature extraction network, and the initial background feature extraction network in the initial image scene recognition model based on the feature extraction network that has completed pre-training.
[0274] For the specific definitions of the image scene recognition apparatus and the image scene recognition model training apparatus, reference may be made to the definitions of the image scene recognition method and the image scene recognition model training method in the foregoing text, which will not be elaborated here. Each module in the above image scene recognition apparatus and the image scene recognition model apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above respective modules.
[0275] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 18As shown in the figure. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store training images and image data to be recognized. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an image scene recognition method and an image scene recognition model training method.
[0276] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 19 shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an image scene recognition method and an image scene recognition model training method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0277] Those skilled in the art can understand that Figure 18 and Figure 19 the structures shown in the figure are only block diagrams of some structures related to the solution of this application, and do not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0278] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0279] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor implements the steps in the above method embodiments.
[0280] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.
[0281] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0282] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0283] The above-described embodiments merely represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. An image scene recognition method, characterized in that, the method includes: Obtain the image to be recognized; Extract the foreground region and the background region in the image to be recognized; Calculate the self-attention weight based on the foreground features corresponding to the foreground region to obtain the self-attention foreground weight, and adjust the foreground features through the self-attention foreground weight to obtain the self-attention foreground features. The self-attention foreground weight is obtained by pooling the foreground features, compressing them, and then performing weight mapping on the compressed features; Calculate the self-attention weight based on the background features corresponding to the background region to obtain the self-attention background weight, and adjust the background features through the self-attention background weight to obtain the self-attention background features. The self-attention background weight is obtained by pooling the background features, compressing them, and then performing weight mapping on the compressed features; Fuse the self-attention background features and the self-attention foreground features to obtain fused features, and perform scene recognition based on the fused features to obtain the image scene recognition result corresponding to the image to be recognized.
2. The method according to claim 1, characterized in that, the method includes: Input the image to be recognized into an image scene recognition model; The image scene recognition model is used to extract the foreground region and the background region in the image to be recognized, input the foreground region into a foreground branch network, and input the background region into a background branch network; The foreground branch network is used to extract the foreground features corresponding to the foreground region, calculate the self-attention weight using the foreground features to obtain the self-attention foreground weight, and weight the foreground features through the self-attention foreground weight to obtain the self-attention foreground features; The background branch network is used to extract the background features corresponding to the background region, calculate the self-attention weight using the background features to obtain the self-attention background weight, and weight the background features through the self-attention background weight to obtain the self-attention background features; The image scene recognition model is further used to fuse the self-attention background features and the self-attention foreground features to obtain fused features, and perform scene recognition based on the fused features to obtain the image scene recognition result.
3. The method according to claim 2, characterized in that, the image scene recognition model includes an image feature extraction network; The extraction of the foreground region and the background region in the image to be recognized includes: Input the image to be recognized into the image feature extraction network for feature extraction to obtain the image features to be recognized; Perform region division based on the image features to be recognized to obtain the foreground region and the background region.
4. The method according to claim 3, characterized in that, the performing region division based on the image features to be recognized to obtain the foreground region and the background region includes: Calculate the mean value corresponding to the feature values in the image features to be recognized; Perform binary division on the image features to be recognized based on the mean value to obtain a foreground mask; Calculate the product of the foreground mask and the pixel values of the image to be recognized to obtain the foreground region in the image to be recognized; Invert the foreground mask to obtain a background mask, and calculate the product of the background mask and the pixel values of the image to be recognized to obtain the background region.
5. The method according to claim 2, wherein, the foreground branch network includes a foreground feature extraction network and a foreground attention feature extraction network; The self-attention foreground weights are calculated based on the foreground features corresponding to the foreground region, and the foreground features are adjusted by the self-attention foreground weights to obtain self-attention foreground features, including: Input the foreground region into the foreground feature extraction network for feature extraction to obtain the foreground features corresponding to the foreground region; Input the foreground features into the foreground attention feature extraction network for attention weight calculation to obtain the self-attention foreground weights, and use the self-attention foreground weights to weight the foreground features to obtain the self-attention foreground features.
6. The method according to claim 5, wherein, The step of inputting the foreground features into the foreground attention feature extraction network for attention weight calculation to obtain the self-attention foreground weights, and using the self-attention foreground weights to weight the foreground features to obtain the self-attention foreground features includes: Perform mean pooling on the foreground features through the mean pooling layer in the foreground attention feature extraction network to obtain foreground pooled features; Perform non-linear compression on the foreground pooled features through the non-linear compression layer in the foreground attention feature extraction network to obtain foreground compressed features; Activate the foreground compressed features through the activation function layer in the foreground attention feature extraction network to obtain foreground activated features; And perform weight mapping on the foreground activated features based on the weight mapping layer in the foreground attention feature extraction network to obtain the self-attention foreground weights; Use the self-attention foreground weights to weight the feature values in the foreground features, and perform max pooling on the weighted foreground features through the max pooling layer in the foreground attention feature extraction network to obtain the self-attention foreground features.
7. The method according to claim 2, wherein, the background branch network includes a background feature extraction network and a background attention weight extraction network; The self-attention background weights are calculated based on the background features corresponding to the background region, and the background features are adjusted by the self-attention background weights to obtain self-attention background features, including: Input the background region into the background feature extraction network for feature extraction to obtain the background features corresponding to the background region; Input the background features into the background attention weight extraction network for attention weight calculation to obtain the self-attention background weights, and use the self-attention background weights to weight the background features to obtain the self-attention background features.
8. The method according to claim 7, It is characterized in that inputting the background features into the background attention weight extraction network to calculate attention weights, obtaining the self-attention background weights, and using the self-attention background weights to weight the background features to obtain the self-attention background features, including: performing mean pooling on the background features through a mean pooling layer in the background attention weight extraction network to obtain background pooled features; performing non-linear compression on the background pooled features through a non-linear compression layer in the background attention weight extraction network to obtain background compressed features; activating the background compressed features through an activation function layer in the background attention weight extraction network to obtain background activated features; and performing weight mapping on the background activated features based on a weight mapping layer in the background attention weight extraction network to obtain the self-attention background weights; using the self-attention background weights to weight the feature values in the background features to obtain weighted background features, and performing max pooling on the weighted background features through a max pooling layer in the background attention weight extraction network to obtain the self-attention background features.
9. The method according to claim 2, It is characterized in that the image scene recognition model includes a fusion output network; fusing the self-attention background features and the self-attention foreground features to obtain fused features, and performing scene recognition based on the fused features to obtain the image scene recognition result corresponding to the image to be recognized, including: concatenating the self-attention background features and the self-attention foreground features through a fusion layer in the fusion output network to obtain concatenated features; inputting the concatenated features into a fully connected layer in the fusion output network for scene recognition to obtain the image scene recognition result.
10. A method for training an image scene recognition model, It is characterized in that the method includes: acquiring training images and corresponding training scene labels, and inputting the training images into an initial image scene recognition model; the initial image scene recognition model extracts an initial foreground training region and an initial background training region in the training images, inputs the initial foreground training region into an initial foreground branch network, and inputs the initial background training region into an initial background branch network; the initial foreground branch network calculates self-attention weights based on initial foreground features corresponding to the initial foreground training region to obtain initial self-attention foreground weights, and adjusts the initial foreground features through the initial self-attention foreground weights to obtain initial self-attention foreground features, where the initial self-attention foreground weights are obtained by pooling the initial foreground features, compressing the pooled features, and performing weight mapping on the compressed features; The initial background branch network calculates self-attention weights based on the initial background features corresponding to the initial background training region to obtain initial self-attention background weights, and adjusts the initial background features through the initial self-attention background weights to obtain initial self-attention background features. The initial self-attention background weights are obtained by pooling the initial background features, compressing them, and then performing weight mapping on the compressed features; The initial image scene recognition model fuses the initial self-attention background features and the initial self-attention foreground features to obtain initial fusion features, and performs scene recognition based on the initial fusion features to obtain an initial image scene recognition result; Calculate the loss information between the initial image scene recognition result and the training scene label, update the initial image scene recognition model based on the loss information, and return to the step of inputting the training image into the initial image scene recognition model and iterate until the training completion condition is reached, then obtain the trained image scene recognition model.
11. The method according to claim 10, wherein, The calculating the loss information between the initial image scene recognition result and the training scene label, updating the initial image scene recognition model based on the loss information, and returning to the step of inputting the training image into the initial image scene recognition model and iterating until the training completion condition is reached, then obtaining the trained image scene recognition model includes: Using a cross-entropy loss function to calculate the error between the initial image scene recognition result and the training scene label to obtain loss information; When the loss information does not exceed a preset loss threshold, calculate the gradient based on the loss information, and use the gradient to update the initial image scene recognition model to obtain an updated image scene recognition model; Take the updated image scene recognition model as the initial scene recognition model, and return to the step of inputting the training image into the initial image scene recognition model and iterate until the loss information exceeds the preset loss threshold, then take the initial image scene recognition model that exceeds the preset loss threshold as the trained image scene recognition model.
12. The method according to claim 10, wherein, The initial image scene recognition model includes an initial image feature extraction network, an initial foreground feature extraction network, and an initial background feature extraction network; Before obtaining the training image and the corresponding training scene label, it further includes: Obtain a pre-trained image and a pre-trained scene label; Input the pre-trained image into a pre-trained scene recognition model. The pre-trained scene recognition model extracts features from the pre-trained image through a pre-trained feature extraction network to obtain pre-trained image features, and performs scene recognition based on the pre-trained image features to obtain a pre-trained image scene recognition result; Calculate pre-training loss information based on the pre-training scenario recognition result and the pre-training scenario label, update the pre-training scenario recognition model based on the pre-training loss information, and return to iteratively execute the step of inputting the pre-training image into the pre-training scenario recognition model until pre-training is completed. Based on the pre-trained feature extraction network after pre-training is completed, obtain the initial image feature extraction network, the initial foreground feature extraction network, and the initial background feature extraction network in the initial image scenario recognition model.
13. An image scenario recognition device Characterized in that The device includes: An image acquisition module, configured to acquire an image to be recognized; A region extraction module, configured to extract a foreground region and a background region in the image to be recognized; A foreground feature extraction module, configured to calculate self-attention foreground weights based on foreground features corresponding to the foreground region, obtain self-attention foreground weights, and adjust the foreground features through the self-attention foreground weights to obtain self-attention foreground features. The self-attention foreground weights are obtained by pooling the foreground features and then compressing them, and mapping the weights of the compressed features; A background feature extraction module, configured to calculate self-attention background weights based on background features corresponding to the background region, obtain self-attention background weights, and adjust the background features through the self-attention background weights to obtain self-attention background features. The self-attention background weights are obtained by pooling the background features and then compressing them, and mapping the weights of the compressed features; A scenario recognition module, configured to fuse the self-attention background features and the self-attention foreground features to obtain a fused feature, and perform scenario recognition based on the fused feature to obtain an image scenario recognition result corresponding to the image to be recognized.
14. The device according to claim 13 Characterized in that The device includes: An image input module, configured to input the image to be recognized into an image scenario recognition model; A branch input module, configured to the image scenario recognition model is used to extract a foreground region and a background region in the image to be recognized, input the foreground region into a foreground branch network, and input the background region into a background branch network; A foreground recognition module, configured to the foreground branch network is used to extract foreground features corresponding to the foreground region, calculate self-attention foreground weights using the foreground features, obtain self-attention foreground weights, and weight the foreground features through the self-attention foreground weights to obtain self-attention foreground features; A background recognition module, configured to the background branch network is used to extract background features corresponding to the background region, calculate self-attention background weights using the background features, obtain self-attention background weights, and weight the background features through the self-attention background weights to obtain self-attention background features; An image recognition module, configured to the image scenario recognition model is further used to fuse the self-attention background features and the self-attention foreground features to obtain a fused feature, and perform scenario recognition based on the fused feature to obtain an image scenario recognition result.
15. The device according to claim 14, wherein, the image scene recognition model includes an image feature extraction network; the branch input module is further configured to input the image to be recognized into the image feature extraction network for feature extraction to obtain the features of the image to be recognized; and perform region division based on the features of the image to be recognized to obtain a foreground region and a background region.
16. The device according to claim 15, wherein, the branch input module is further configured to calculate the mean value corresponding to the feature values in the features of the image to be recognized; perform binary division on the features of the image to be recognized based on the mean value to obtain a foreground mask; calculate the product of the foreground mask and the pixel values of the image to be recognized to obtain the foreground region in the image to be recognized; invert the foreground mask to obtain a background mask, and calculate the product of the background mask and the pixel values of the image to be recognized to obtain the background region.
17. The device according to claim 14, wherein, the foreground branch network includes a foreground feature extraction network and a foreground attention feature extraction network; the foreground recognition module is further configured to input the foreground region into the foreground feature extraction network for feature extraction to obtain the foreground features corresponding to the foreground region; input the foreground features into the foreground attention feature extraction network for attention weight calculation to obtain the self-attention foreground weight, and use the self-attention foreground weight to weight the foreground features to obtain the self-attention foreground features.
18. The device according to claim 17, wherein, the foreground recognition module is further configured to perform mean pooling on the foreground features through the mean pooling layer in the foreground attention feature extraction network to obtain foreground pooled features; perform non-linear compression on the foreground pooled features through the non-linear compression layer in the foreground attention feature extraction network to obtain foreground compressed features; activate the foreground compressed features through the activation function layer in the foreground attention feature extraction network to obtain foreground activated features; and perform weight mapping on the foreground activated features based on the weight mapping layer in the foreground attention feature extraction network to obtain the self-attention foreground weight; use the self-attention foreground weight to weight the feature values in the foreground features to obtain weighted foreground features, and perform max pooling on the weighted foreground features through the max pooling layer in the foreground attention feature extraction network to obtain the self-attention foreground features.
19. The device according to claim 14, wherein, the background branch network includes a background feature extraction network and a background attention weight extraction network; the background recognition module is further configured to input the background region into the background feature extraction network for feature extraction to obtain the background features corresponding to the background region; input the background features into the background attention weight extraction network for attention weight calculation to obtain the self-attention background weight, and use the self-attention background weight to weight the background features to obtain the self-attention background features.
20. The device according to claim 19, wherein, the background recognition module is further configured to perform mean pooling on the background features through a mean pooling layer in the background attention weight extraction network to obtain background pooled features; perform non-linear compression on the background pooled features through a non-linear compression layer in the background attention weight extraction network to obtain background compressed features; activate the background compressed features through an activation function layer in the background attention weight extraction network to obtain background activated features; and perform weight mapping on the background activated features based on a weight mapping layer in the background attention weight extraction network to obtain the self-attention background weights; weight the feature values in the background features using the self-attention background weights to obtain weighted background features, and perform max pooling on the weighted background features through a max pooling layer in the background attention weight extraction network to obtain the self-attention background features.
21. The device according to claim 14, wherein, the image scene recognition model includes a fusion output network; the image recognition module is further configured to splice the self-attention background features and the self-attention foreground features through a fusion layer in the fusion output network to obtain spliced features; input the spliced features into a fully connected layer in the fusion output network for scene recognition to obtain the image scene recognition result.
22. An image scene recognition model training device, wherein, the device includes: a training data acquisition module, configured to acquire training images and corresponding training scene labels, and input the training images into an initial image scene recognition model; a model processing module, configured to extract an initial foreground training region and an initial background training region in the training images by the initial image scene recognition model, input the initial foreground training region into an initial foreground branch network, and input the initial background training region into an initial background branch network; a foreground network processing module, configured to calculate self-attention weights based on the initial foreground features corresponding to the initial foreground training region by the initial foreground branch network to obtain initial self-attention foreground weights, and adjust the initial foreground features through the initial self-attention foreground weights to obtain initial self-attention foreground features, where the initial self-attention foreground weights are obtained by pooling the initial foreground features, compressing the pooled features, and performing weight mapping on the compressed features; a background network processing module, configured to calculate self-attention weights based on the initial background features corresponding to the initial background training region by the initial background branch network to obtain initial self-attention background weights, and adjust the initial background features through the initial self-attention background weights to obtain initial self-attention background features, where the initial self-attention background weights are obtained by pooling the initial background features, compressing the pooled features, and performing weight mapping on the compressed features; A model recognition module, configured to enable the initial image scene recognition model to perform feature fusion on the initial self-attention background feature and the initial self-attention foreground feature to obtain an initial fusion feature, and perform scene recognition based on the initial fusion feature to obtain an initial image scene recognition result; An iteration module, configured to calculate loss information between the initial image scene recognition result and the training scene label, update the initial image scene recognition model based on the loss information, and return to iteratively execute the step of inputting the training image into the initial image scene recognition model until the training completion condition is reached, and obtain a trained image scene recognition model.
23. The apparatus according to claim 22, wherein, the iteration module is further configured to calculate an error between the initial image scene recognition result and the training scene label by using a cross-entropy loss function to obtain loss information; when the loss information does not exceed a preset loss threshold, calculate a gradient based on the loss information, use the gradient to update the initial image scene recognition model to obtain an updated image scene recognition model; use the updated image scene recognition model as the initial scene recognition model, and return to iteratively execute the step of inputting the training image into the initial image scene recognition model until the loss information exceeds the preset loss threshold, and use the initial image scene recognition model that exceeds the preset loss threshold as the trained image scene recognition model.
24. The apparatus according to claim 22, wherein, the initial image scene recognition model includes an initial image feature extraction network, an initial foreground feature extraction network, and an initial background feature extraction network; the apparatus further includes: A pre-training module, configured to obtain a pre-training image and a pre-training scene label; input the pre-training image into a pre-training scene recognition model, and the pre-training scene recognition model performs feature extraction on the pre-training image through a pre-training feature extraction network to obtain a pre-training image feature, perform scene recognition based on the pre-training image feature to obtain a pre-training image scene recognition result; calculate pre-training loss information based on the pre-training scene recognition result and the pre-training scene label, update the pre-training scene recognition model based on the pre-training loss information, and return to iteratively execute the step of inputting the pre-training image into the pre-training scene recognition model until pre-training is completed, and obtain the initial image feature extraction network, the initial foreground feature extraction network, and the initial background feature extraction network in the initial image scene recognition model based on the pre-trained pre-training feature extraction network.
25. A computer device, including a memory and a processor, where the memory stores a computer program, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
26. A computer-readable storage medium, storing a computer program, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
27. A computer program product, including a computer program, It is characterized in that when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Method and device for identifying scene integrated by multi-feature vision codebook
CN103366181A
Scene recognition method and device
CN105678267A
Medical image recognition method, device and equipment, and storage medium
CN111985574A
Scene recognition method and device, computer equipment and storage medium
CN112348117A