A method, device, equipment and medium for model training and scene recognition

By adopting multi-level feature extraction and decision-making layer models in machine audit technology, the problem of inaccurate scene recognition in the existing technology is solved, and higher accuracy of image scene recognition and effectiveness of machine audit are achieved.

CN114049584BActive Publication Date: 2025-06-17BIGO TECH PTE LTD

Patent Information

Application Number
CN202111174534.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-09
Publication Date
2025-06-17
Estimated Expiration
2041-10-09

AI Technical Summary

Technical Problem

When existing machine audit technology recognizes guns in images, it is easy to cause poor accuracy of audit results due to scene reasons, and it is impossible to effectively distinguish guns in animation or game scenes from actual violation images.

Method used

Using a model training and scene recognition method, the training scene recognition model can accurately identify scene information in the image through the core feature extraction layer, global information feature extraction layer, LCS modules at each level and fully connected decision-making layer.

Benefits of technology

Through this method, the accuracy of scene recognition is significantly improved, so that machine audit can more accurately determine whether the image is violated, and reduce misjudgment caused by inaccurate scene recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114049584B_ABST
    Figure CN114049584B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and medium for model training and scene recognition. When training a scene recognition model, first, the parameters of the core feature extraction layer and the global information feature extraction layer are trained through the first scene label of the sample image and the standard cross-entropy loss. Then, according to the feature maps output by the LCS modules at each level and the loss value calculated pixel by pixel with the first scene label of the sample image, the weight parameters of the LCS modules at each level are trained. Finally, the parameters of the fully connected decision layer of the scene recognition model are trained. This enables the scene recognition model to have the ability to extract features with high richness, and based on the scene recognition model for scene recognition, the accuracy of scene recognition is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a model training and scene recognition method, device, equipment and medium. Background Art

[0002] Machine review technology (referred to as machine review) is increasingly used in large-scale short video / image review. Illegal images identified by machine review are then pushed to staff for review (referred to as human review) to ultimately determine whether the image is illegal. The emergence of machine review has greatly improved the efficiency of image review. However, machine review tends to rely on the commonalities of image vision to make violation judgments, thereby ignoring changes in review results caused by changes in the macro environment. For example, in the review of gun violations, when the machine review recognizes that a gun appears in the image, it generally considers the image to be illegal, but the accuracy of such machine review results is poor. This is because, for example, if it is a gun in an animation or game scene, the picture is not an illegal picture. Therefore, scene recognition has a greater impact on the accuracy of the machine review results. At present, there is an urgent need for a scene recognition solution. Summary of the invention

[0003] The embodiments of the present invention provide a model training and scene recognition method, device, equipment and medium, so as to provide a scene recognition solution with high accuracy.

[0004] An embodiment of the present invention provides a scene recognition model training method, wherein the scene recognition model includes a core feature extraction layer and a global information feature extraction layer connected to the core feature extraction layer, LCS modules at each level, and a fully connected decision layer, and the method includes:

[0005] The parameters of the core feature extraction layer and the global information feature extraction layer are obtained by training through the first scene label of the sample image and the standard cross entropy loss;

[0006] Training weight parameters of the LCS modules at each level according to the feature maps output by the LCS modules at each level and the loss values ​​calculated pixel by pixel for the first scene label of the sample image;

[0007] The parameters of the fully connected decision layer are obtained by training through the first scene label of the sample image and the standard cross entropy loss.

[0008] On the other hand, an embodiment of the present invention provides a scene recognition method based on a scene recognition model trained by the above method, the method comprising:

[0009] Obtain an image to be recognized;

[0010] Input the image to be recognized into a pre-trained scene recognition model, and determine the scene information corresponding to the image to be recognized based on the scene recognition model.

[0011] On the other hand, an embodiment of the present invention provides a scene recognition model training device, which includes:

[0012] The first training unit is used to train the parameters of the core feature extraction layer and the global information feature extraction layer through the first scene label of the sample image and the standard cross-entropy loss.

[0013] The second training unit is used to train the weight parameters of each layer of the LCS module according to the loss value calculated pixel by pixel from the feature map output by each layer of the LCS module and the first scene label of the sample image.

[0014] The third training unit is used to train the parameters of the fully connected decision layer through the first scene label of the sample image and the standard cross-entropy loss.

[0015] On the other hand, an embodiment of the present invention provides a scene recognition device for a scene recognition model trained based on the above-mentioned device, which includes:

[0016] An acquisition module is used to acquire the image to be recognized;

[0017] The recognition module is used to input the image to be recognized into a pre-trained scene recognition model, and determine the scene information corresponding to the image to be recognized based on the scene recognition model.

[0018] On yet another aspect, an embodiment of the present invention provides an electronic device, which includes a processor, and the processor is used to implement the steps of the above-mentioned model training method or the steps of the above-mentioned scene recognition method when executing a computer program stored in a memory.

[0019] On yet another aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program implements the steps of the above-mentioned model training method or the steps of the above-mentioned scene recognition method when executed by a processor.

[0020] An embodiment of the present invention provides a model training and scene recognition method, device, equipment and medium. The scene recognition model includes a core feature extraction layer and a global information feature extraction layer connected to the core feature extraction layer, each layer of the LCS module, and a fully connected decision layer. The method includes:

[0021] The parameters of the core feature extraction layer and the global information feature extraction layer are trained through the first scene label of the sample image and the standard cross-entropy loss;

[0022] The weight parameters of the LCS modules at each level are trained according to the loss value calculated pixel by pixel from the feature map output by the LCS modules at each level and the first scene label of the sample image;

[0023] The parameters of the fully connected decision layer are trained through the first scene label of the sample image and the standard cross-entropy loss.

[0024] The above technical solution has the following advantages or beneficial effects:

[0025] The embodiment of the present invention provides a solution for image scene recognition based on a scene recognition model. When training the scene recognition model, first, the parameters of the core feature extraction layer and the global information feature extraction layer are trained through the first scene label of the sample image and the standard cross-entropy loss. Then, the weight parameters of the LCS modules at each level are trained according to the loss value calculated pixel by pixel from the feature map output by the LCS modules at each level and the first scene label of the sample image. Finally, the parameters of the fully connected decision layer of the scene recognition model are trained. This enables the scene recognition model to have the ability to extract features with high richness. Based on the scene recognition model for scene recognition, the accuracy of scene recognition is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0027] Figure 1 It is a schematic diagram of the training process of the scene recognition model provided by the embodiment of the present invention;

[0028] Figure 2 It is a schematic diagram of the application of the scene recognition method provided by the embodiment of the present invention;

[0029] Figure 3 It is a flowchart of the main model training stage provided by the embodiment of the present invention;

[0030] Figure 4 It is a flowchart of the model branch expansion stage provided by the embodiment of the present invention;

[0031] Figure 5 It is a schematic diagram of the structure of the core feature extraction part of the scene recognition model provided by the embodiment of the present invention;

[0032] Figure 6 Schematic diagram of the structure and execution principle of the global information feature extraction layer provided by an embodiment of the present invention;

[0033] Figure 7 Schematic diagram for detailed explanation of the principle of the local supervised learning module provided by an embodiment of the present invention;

[0034] Figure 8 Schematic diagram of the structure and the principle of the first round of training of the extended branch network of the scene recognition model provided by an embodiment of the present invention;

[0035] Figure 9 Schematic diagram of the structure and training of the branch extension stage of the scene recognition model provided by an embodiment of the present invention;

[0036] Figure 10 Schematic diagram of the scene recognition process provided by an embodiment of the present invention;

[0037] Figure 11 Schematic diagram of the structure of the scene recognition training device provided by an embodiment of the present invention;

[0038] Figure 12 Schematic diagram of the structure of the scene recognition device provided by an embodiment of the present invention;

[0039] Figure 13 Schematic diagram of the structure of the electronic device provided by an embodiment of the present invention;

[0040] Figure 14 Schematic diagram of the structure of another electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0041] The present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0042] Special abbreviations or custom nouns involved in the embodiments of the present invention are explained as follows:

[0043] Convolutional neural network: An end-to-end complex mapping for extracting features of images or videos and completing visual tasks such as classification and detection based on the extracted features, usually composed of multiple basic convolutional modules stacked together.

[0044] Convolutional layer: An operation layer that extracts features by weighted summation of an image using a kernel with a specific receptive field. Generally, this layer also combines a non-linear activation function to improve the mapping ability.

[0045] Pooling: An aggregation operation, such as aggregating pixel values within a specific range or dimension, usually including maximization, minimization, taking the average, etc.

[0046] Group convolution: The feature maps are grouped into several groups by channels, and the feature maps in each group perform the same or different convolution operations, which can be used to reduce the computational cost.

[0047] Feature pyramid: A method for extracting multi-scale features. Usually, feature maps are taken from different levels of the network, and then the feature maps are aligned through a certain upsampling scheme and fused to produce multi-scale features.

[0048] Residual block: A module composed of multiple convolutional layers with a cross-layer connection bypass. Using this module, a deeper convolutional neural network can be built, and the phenomenon of gradient disappearance can be avoided, accelerating the training of the network.

[0049] Heatmap: A feature map that can reflect the local importance of an image. Generally, the higher the importance, the greater the local heat value, or vice versa.

[0050] Local supervised learning: Directly connected labels and losses are used for certain parts of the model or a certain local area of the feature map to learn parameters or extraction capabilities.

[0051] Attention mechanism: A mechanism that forces the network to focus on important regions by fitting the importance degrees of different parts, and makes decisions based on the features of the important regions.

[0052] Sigmoid: An activation function that does not consider the mutual exclusivity of categories. Usually, the output values after activation will fall within the interval [0, 1] to complete normalization.

[0053] Deformable convolution: A convolution operation in which the convolution kernel is not in a standard geometric shape. The non-standard geometric shape is usually generated by adding offsets to the original shape.

[0054] Standard cross-entropy: A conventional loss evaluation function for simple classification problems, often used to train classification networks, including single-label classification and multi-label classification.

[0055] Focal Loss: A loss function for the class imbalance problem, which can give a larger penalty to the classes with less data volume to prevent the model from completely favoring the classes with more data volume.

[0056] It should be noted that the solution provided by the embodiment of the present invention is not directly applied to the machine review stage to directly generate the push review results; instead, it outputs the scene information required by the specific machine review model in the form of scene signals, and generates the push results together with the machine review model through appropriate strategies. The videos or pictures that are considered to be illegal by the final push results will be pushed to the human review stage for multiple rounds of review to obtain the punishment results; and the videos or pictures that are considered to be normal by the final push results will also be sampled and inspected in different areas at a certain sampling rate, or pushed to the human review stage for review according to the report results, to avoid missing videos / pictures that are seriously illegal.

[0057] Embodiment 1:

[0058] Figure 1 A schematic diagram of a scene recognition model training process provided by an embodiment of the present invention, the process includes the following steps:

[0059] S101: training the parameters of the core feature extraction layer and the global information feature extraction layer through the first scene label of the sample image and the standard cross entropy loss.

[0060] S102: training weight parameters of the LCS modules at each level according to the feature maps output by the LCS modules at each level and the loss values ​​calculated pixel by pixel based on the first scene label of the sample image.

[0061] S103: Training to obtain parameters of the fully connected decision layer through the first scene label of the sample image and standard cross entropy loss.

[0062] Among them, the scene recognition model includes a core feature extraction layer and a global information feature extraction layer connected to the core feature extraction layer, LCS modules at each level, and a fully connected decision layer.

[0063] The scene recognition method provided by the embodiment of the present invention is applied to an electronic device, which may be a smart device such as a PC, a tablet computer, or a server.

[0064] In order to meet the requirements of different fine-grained scene recognition, in an embodiment of the present invention, the scene recognition model further includes a branch expansion structure; the branch expansion structure includes a convolution layer and a local object association relationship module;

[0065] According to the loss value calculated pixel by pixel based on the feature map output by the convolution layer of the branch extension structure and the second scene label of the sample image, the weight parameters of the convolution layers at each level of the branch extension structure are trained; the parameters of the local object association relationship module are trained through the loss function with a scene confidence regularization term; wherein the granularity of the first scene label and the second scene label are different.

[0066] Generally, the first scene label is a coarse-grained scene label, and the second scene label is a fine-grained scene label.

[0067] As Figure 2 shown, the multi-level fine-grained scene recognition model proposed in the embodiments of the present invention will run in a synchronous manner. That is, for the images to be recognized in a certain video or picture set, it will be executed prior to other existing machine review models in the machine review process, and the scene information obtained from the execution will be stored in the cache queue. Then, a series of machine review models ( Figure 2 such as machine review model a, machine review model b, and machine review model c in) will be activated through a synchronization signal. Using the same images to be recognized in the video or picture set as the input, the machine review models will be run to obtain preliminary review results. Among them, the machine review model / tag configured with the scene policy will retrieve the corresponding scene signal from the cache queue, calculate it jointly with the preliminary machine review results, obtain the final review result, and decide whether to push it to the human review link; while the machine review model / tag without the configured scene policy will directly decide whether to push according to the results given by the machine review model.

[0068] If the scene information corresponding to the identified image to be recognized belongs to the illegal scene information, and the review result of the machine review is that the image to be recognized is an illegal image, then it is determined that the image to be recognized is an illegal image. And it is decided to push it to the human review link. If the scene information corresponding to the identified image to be recognized does not belong to the illegal scene information, or the review result of the machine review is that the image to be recognized is not an illegal image, then it is determined that the image to be recognized is not an illegal image. At this time, it is not pushed to the human review link, or it is sampled and inspected in different regions at a certain sampling rate, or it is pushed to the human review link for re-review according to the report result.

[0069] Among them, which scene information belongs to the illegal scene information can be pre-saved in the electronic device. After the scene information corresponding to the identified image to be recognized is determined, it is possible to determine whether the scene information corresponding to the image to be recognized belongs to the illegal scene information. The process of the machine review model reviewing whether the image to be recognized is an illegal image can adopt the solutions of the existing technology and will not be elaborated here.

[0070] Figure 3 is a schematic diagram of the main structure training process of the scene recognition model provided by the embodiments of the present invention, Figure 4 is a schematic diagram of the branch expansion structure training process of the scene recognition model provided by the embodiments of the present invention.

[0071] Figure 3 and Figure 4 show the overall training process of the multi-level fine-grained scene recognition model proposed in the embodiments of the present invention, which consists of two stages, namely the "model main body training stage" and the "model branch expansion stage". After the training is completed, a model such as Figure 2The scene recognition model with the structure shown on the left is used to improve the accuracy of the machine review model. For the scene recognition model, the feature extraction capability of the main structure is very important, because the various components of the main structure are generally used as the pre-components when the branch expansion structure is used, affecting the feature extraction capability of the specific fine-grained branch. In order to greatly improve the feature extraction capability of the main structure of the model, it is used to mine high-rich features; Figure 3 As shown in the figure, in the training stage of the main structure of the model, three rounds of training strategies with different targets are used to train each module of the main structure. Among them, the first round mainly optimizes the core feature extraction layer and the multi-scale global information feature extraction layer of the scene recognition model through the first scene label (generally the scene semantic label at the picture level) and the standard cross entropy loss after multiple iterations, and the weight parameters of the local supervised learning module are not optimized first. The main purpose of the first round of optimization is to enable the model to obtain the ability to extract abstract semantic features and global features at multiple scales. Then, the parameters (weight parameters) optimized in the first round are fixed, and a "local supervised learning module with attention mechanism" (Local Supervised module, LCS module) is connected to the convolutional feature map group after each pooling layer. Each level uses the convolutional feature map group after pooling as input. After the LCS module, a "focused" feature map is output. Through the feature map and the first scene label of the sample image (generally the label at the frame level), the weight parameters of the LCS modules at each level are optimized with a pixel-by-pixel binary sigmoid loss. After the second round of optimization, the model will be sensitive to the features of local objects. The local object features can be extracted through the LCS module and encoded through Fisher convolution features. While reducing feature redundancy, the loss of subtle features that have decision-making influences can be minimized. The third round of optimization focuses on the decision layer weight parameters for the fusion features. At this time, the fully connected output layer used in the first round of optimization will be removed and replaced with a fully connected decision layer corresponding to the fusion of three features. The first scene label (generally the scene semantic label at the picture level) and the standard cross entropy loss are also used for training optimization. The weight parameters outside the decision layer are fixed.

[0072] After the main training phase of the model is completed, the model can branch and expand on the main structure according to subsequent fine-grained scenario requirements. Figure 4An example of expanding a branch is given. Generally, a branch starts from the output of a certain convolutional layer in the main structure part, and accesses several convolutional pooling operation layers. During training, the weight parameters of the main structure part are fixed, and the training is completed after two rounds of optimization. In the first round of optimization, a per-pixel binary sigmoid loss is used to directly optimize the associated convolutional layer, and the LCS module is no longer used as a relay. The main purpose of the first round of optimization is also to enable the expanded branch to have the ability to learn features of local objects. In the second round of optimization, some "local object association relationship learning modules" composed of "deformable convolutions" are embedded in the expanded branch. Based on the fact that the branch already has the ability to extract local object features, the ability to mine local object association relationships is learned through the focus loss with a scene confidence regularization term. Using a similar method, multiple different branches can be expanded on the main network to handle different task requirements.

[0073] An embodiment of the present invention provides a solution for image scene recognition based on a scene recognition model. When training the scene recognition model, first, the parameters of the core feature extraction layer and the global information feature extraction layer are trained through the first scene label of the sample image and the standard cross-entropy loss. Then, according to the feature maps output by the LCS modules at each level and the loss values calculated pixel by pixel with the first scene label of the sample image, the weight parameters of the LCS modules at each level are trained. Finally, the parameters of the fully connected decision layer of the scene recognition model are trained. This enables the scene recognition model to have the ability to extract features with high richness. Based on the scene recognition model for scene recognition, the accuracy of scene recognition is greatly improved. Moreover, the scene recognition model also includes a branch expansion structure to adapt to the requirements of different fine-grained scene recognition.

[0074] From Figure 4 It can be seen that the core feature extraction layer in the scene recognition model is respectively connected to the global information feature extraction layer, the LCS modules at each level, the fully connected decision layer ( Figure 4 the FC layer in it) and the branch expansion structure. When performing scene recognition on the image to be recognized based on the scene recognition model, first, the core feature extraction layer extracts features from the image to be recognized, and then the results are respectively output to the global information feature extraction layer, the LCS modules at each level, the fully connected decision layer and the branch expansion structure. The global information feature extraction layer, the LCS modules at each level, the fully connected decision layer and the branch expansion structure process the received feature maps in parallel to obtain the final scene recognition result.

[0075] Figures 5 - 9 Shows the details of the scene recognition model and model training.

[0076] Embodiment 2:

[0077] The core feature extraction layer includes a first type of grouped multi-receptive field residual convolution module and a second type of grouped multi-receptive field residual convolution module;

[0078] The first type of grouped multi-receptive field residual convolution module includes a first group, a second group, and a third group. The convolution sizes of the first group, the second group, and the third group are different. The first group, the second group, and the third group include a residual calculation bypass structure; each group outputs a feature map through convolution operations and residual calculations. The feature maps output by each group are concatenated in the channel dimension and channel shuffling is performed, and after convolution fusion, the output is sent to the next module;

[0079] The second type of grouped multi-receptive field residual convolution module includes a fourth group, a fifth group, and a sixth group. The convolution sizes of the fourth group, the fifth group, and the sixth group are different. The fifth group and the sixth group respectively include a 1×1 convolution bypass structure and a residual calculation bypass structure; the feature maps output by each group are concatenated in the channel dimension and channel shuffling is performed, and after convolution fusion, the output is sent to the next module.

[0080] The scene recognition model proposed in the embodiment of the present invention has a core feature extraction layer structure as Figure 5 shown on the right, which is composed of a series of "grouped multi-receptive field residual convolution modules". Figure 5 The two types of grouped multi-receptive field residual convolution modules are described in detail on the left. They are composed of three convolution branches with different receptive fields. In order to save computational effort, the group of feature maps obtained by the previous module will be divided into three groups and transmitted to different convolution branches for convolution operations respectively to further extract features. The first type of grouped multi-receptive field residual convolution module is Figure 5 the "GM-Resblock" in. To cover different receptive fields, the convolution kernels of the three branches of the first group, the second group, and the third group use three different sizes of 1x1, 3x3, and 5x5 respectively. Among them, the 5x5 convolution operation is replaced by two layers of 3x3 convolution operations. In this way, while maintaining the same receptive field, the number of non-linear mappings can be increased, and the fitting ability can be improved. Each branch also adds a bypass structure for calculating residuals to avoid gradient disappearance while expanding the depth of the model. In the embodiment of the present invention, multi-receptive field convolution is used as the module mainly because scene recognition belongs to a complex visual problem, and local features at different scales may all affect the discrimination of the scene, and the multi-receptive field mechanism can try to capture more factors promoting decision-making. To ensure the regularity of the number of channels, the results output by the three convolution branches of GM-Resblock will be concatenated in the channel dimension and channel shuffling is performed, and finally fused with a 1x1 convolution and then sent to the next module. It should be noted that the second type of grouped multi-receptive field residual convolution module is Figure 5The "GM projection block" in it. The GM projection block includes a fourth group, a fifth group, and a sixth group. The convolutional kernels of the three branches of the fourth group, the fifth group, and the sixth group use three different sizes of 1x1, 3x3, and 5x5 respectively. Among them, the 5x5 convolution operation is replaced by two layers of 3x3 convolution operations. The GM projection block is also used for downsampling the feature map, so its structure is slightly modified. For example, the bypass of the 1x1 convolution branch is cancelled, and a 1x1 convolution is added to the bypasses of the 3x3 and 5x5 convolution branches to maintain the consistency of the feature map size and the number of channels. To ensure the regularity of the number of channels, the results output by the three convolution branches of the GM projection block are concatenated in the channel dimension and channel shuffling is performed, and finally fused through a 1x1 convolution and then passed to the next module.

[0081] Example 3:

[0082] The parameters of the core feature extraction layer and the global information feature extraction layer obtained by training through the first scene label of the sample image and the standard cross-entropy loss include:

[0083] Upsample the feature maps of different levels in the core feature extraction layer using deconvolution operations with different dilation factors, align the number of channels using bilinear interpolation algorithm in the channel dimension, add and merge the feature maps of each level channel by channel, perform convolution fusion on the merged feature map group, and obtain the global information feature vector through global average pooling of each channel. Concatenate the global information feature vector and the fully connected layer FC feature vector, and train the parameters of the core feature extraction layer and the global information feature extraction layer through the standard cross-entropy loss.

[0084] In the first round of the model main body training stage, the global information feature extraction module is trained together with the core feature extraction layer of the model. Figure 6Briefly shows its details and principles. To extract the global information of an image, the present invention starts from feature maps at multiple scales, fuses the global information at different scales, and obtains a high-quality global information feature vector. Compared with the global information feature at a single scale, the features at multiple scales can, on the one hand, reduce information loss, and on the other hand, make the model more sensitive to important regions at the global spatial level. The embodiments of the present invention draw on the idea of a feature pyramid, and perform upsampling on the feature map groups at different levels in the core feature extraction layer of the model using deconvolution operations with different dilation factors to ensure that the sizes of the feature maps are consistent. Here, deconvolution is used instead of ordinary padding upsampling mainly to alleviate the problem of image distortion caused by the upsampling operation. After the upsampling operation is completed, the number of channels of the feature maps at different levels is still inconsistent. Here, the bilinear interpolation algorithm is simply used in a loop in the channel dimension to supplement the insufficient channels and align the number of channels. Then, the feature maps at each level perform a channel-wise addition operation to complete the merging. The merged feature map group is fused using 1x1 convolution, and a global information feature vector is obtained through channel-wise global average pooling. According to Figure 3 , this vector will be concatenated with the FC feature vector that records the abstract features, and then connected to the standard cross-entropy loss for optimization.

[0085] Embodiment 4:

[0086] Training the weight parameters of each level of the LCS module according to the loss value calculated pixel by pixel from the feature map output by each level of the LCS module and the first scene label of the sample image includes:

[0087] Using an activation function through an attention mechanism in the channel dimension to obtain the importance weight of each channel, and performing weighted summation on the feature map of each channel according to the importance weight of each channel to obtain a summary heat map;

[0088] Calculating the loss value pixel by pixel according to the summary heat map, the object-scene association importance degree, and the area of the object, and training the weight parameters of each level of the LCS module according to the loss value.

[0089] In the main structure part of the model, another important step is to enable it to have better extraction ability for local object features. As Figure 7 shown, after the first round of the model main body training stage is completed, also at multiple levels, the present invention proposes to use a "local supervised learning module with an attention mechanism" and a "local object supervised loss" to enhance the extraction ability of this part of the model. Among them, the structure of the local supervised learning module with an attention mechanism (LCS module) is as Figure 7As shown in the lower left. First, the feature map groups of each layer are taken out and first mapped through a 3x3 convolution while keeping the number of channels unchanged. Considering that the importance of feature maps in different channels is not the same at the same pixel position, in the embodiments of the present invention, while downsampling in terms of size, the importance of feature maps in different channels is also controlled through an attention mechanism in the channel dimension to obtain a more accurate summary heat map indicating the pixel information at each position, so as to better guide the LCS module to learn the local object features. In addition, the LCS module uses ordinary 3x3 convolution to complete downsampling instead of pooling operation, so as to avoid too large deviation of local object activation when reducing redundancy. The attention mechanism in the channel dimension uses Sigmoid activation to obtain importance weights because the importance between channels is not a mutually exclusive relationship; finally, an importance weight vector will be output through a fully connected layer, and then the importance value is multiplied by the corresponding channel to complete the weighting of the channel feature map.

[0090] After the LCS module outputs the feature map group enhanced by attention, it will specifically connect to the "local object supervision loss" to supervisedly guide the module to learn the extraction ability of local object features. Specifically, first sum the feature map group enhanced by attention pixel by pixel across channels to obtain a heat map reflecting the activation situation at different pixel positions. Then, use this heat map and the mask map based on the block diagram object and the "object-scene association importance" label to calculate the loss and perform backpropagation. The mask map is a label obtained based on the scene semantic label at the image level according to the influence degree of the object in the scene image on the scene judgment. Among them, the object in the image will give a mask according to the block diagram range it occupies. The object that has a great influence on the scene judgment is marked as "important" (the mask value is given 1.0), the common object that has a small influence on the scene judgment and appears in multiple scenes is marked as "unimportant" (the mask value is given 0.5), and the background mask value is given 0.0. In order to achieve the effect of "local supervised learning", the loss uses the binary sigmoid loss for each pixel, and the penalty weight is selected according to the ratio of the area of the "important object" to the area of the "unimportant object". When the area of the "important object" is much smaller than that of the "unimportant object", the relative gap between the penalty weight of the "important object" and the penalty weight of the "unimportant object" will increase, so that the LCS module will increase the learning intensity for the "important object" when the "important object" is a small target, and avoid leaning towards the learning of the "unimportant object" or the "background". It should be noted that since the goal of the LCS module is to extract local object features, the penalty weight of the "background" will take a smaller value in both cases. The specific loss expression is as follows:

[0091]

[0092] Among them, p i,j represents the activation value of the pixel on the heat map, mask i,jThe representative pixel-level label, area represents the area, and in the present invention, λ im , λ unim , λ′ im , λ′ unim , λ back take the values of 0.8, 0.6, 1.0, 0.5, and 0.3 respectively. It should be noted that when training the LCS module in the present invention, the modules at each level are directly connected to the loss and backpropagate independently, and the mask map will be downsampled accordingly as needed.

[0093] After the LCS module is trained, the features directly extracted by it are still a group of feature maps with a size of HxWxC. Directly using them as features will still have too much redundancy, while using a non-linear fully connected layer to extract feature vectors will cause some subtle decisive features to be lost. Therefore, the embodiment of the present invention uses the Fisher convolution coding method to reduce the dimension of the feature map and uses the Fisher convolution feature coding technology to extract local object feature vectors, avoiding the influence of geometric transformation caused by redundant features while reducing the loss of subtle decisive features. The process of Fisher convolution feature coding is relatively simple. It mainly uses a variety of general Gaussian distributions to mix the vectors on different pixels and reduce the number of features in the size dimension. The specific steps are as follows:

[0094] Flatten the feature map in the size dimension so that it is represented as HxW C-dimensional vectors.

[0095] Use PCA to reduce the dimension of each C-dimensional vector to M dimensions.

[0096] Calculate K Gaussian mixture parameter values using K Gaussian distributions on the HxW M-dimensional vectors.

[0097] Evolve the HxW M-dimensional vectors into K M-dimensional Gaussian vectors.

[0098] Calculate the mean vector and variance vector of all Gaussian vectors, splice them and perform L2 regularization, and finally output a local object feature vector with a length of 2MK, and each level outputs one vector.

[0099] The difference here from the global information feature extraction is that in order to obtain some subtle local object characteristics, the features at different levels are not fused but output separately. As Figure 3 shown in step 3, after obtaining the local object features and global information features, these features will also be combined with the FC abstract features to reconstruct a main decision layer for high-abundance features, and use these features to complete high-precision decision-making.

[0100] Example 5:

[0101] The branch extension structure is constructed using depthwise separable convolutional residual blocks DW. In the main path of the residual block, a depthwise convolutional layer is used in the middle layer, and 1x1 convolutional layers are used before and after the depthwise convolution.

[0102] The local object association relationship learning module includes a deformable convolutional layer, a convolutional layer, and an average pooling layer;

[0103] The deformable convolutional layer obtains the convolutional kernel offset value at the current pixel position. The offset value is added to the current position of the convolutional kernel parameters as its actual effective position, and the pixel value of the feature image at the actual effective position is obtained. After convolutional operation and average pooling operation, a feature map is output.

[0104] After the training of the model main body is completed, the branch extension stage is entered. Usually, the branches are extended according to the new fine-grained scenario requirements, and an appropriate network structure can be adopted according to the requirements to design new branches. In the embodiments of the present invention, considering the multiple expandability of the branches, in order to control the overhead of each branch, depthwise separable convolutional residual blocks (Depth-Wise, DW) are used to construct the branches, as Figure 8 shown. In the main path of the residual block, a depthwise convolution is used in the middle layer to replace the ordinary convolutional layer, reducing the computational overhead by about two-thirds. 1x1 convolutions are used before and after the depthwise convolution to implement the inverse channel shrinking operation, and a linear activation is used for the output. This is mainly to avoid discarding too many features when Relu is activated for negative values. The present invention finally uses three modules (constituent part a of the branch module, constituent part b of the branch module, and constituent part c of the branch module) connected in series to form a fine-grained branch. Since the branch is extended from the core feature extraction layer of the scene model main body, this part is not specifically optimized for the local object feature learning ability. Therefore, the corresponding level of the extended branch network will directly access the previously proposed LCS loss for pre-training optimization. Here, it is different from the training of the main body part. No additional LCS module is added, but the convolutional layer parameters are shared with the extended branch network. On the one hand, this is to reduce the overhead, and on the other hand, it is to enable learning the association relationship of local objects on the basis of local object feature extraction during the second round of training in the branch extension stage, and realizing the recognition of fine-grained complex scenarios by combining local object features and the global spatial association of local objects.

[0105] In order to obtain the learning ability of the association relationship on the basis of the local object feature extraction ability, in the second round of the branch extension stage of the present invention, an "association relationship learning module" is embedded between the constituent parts of each branch module, and these modules are trained together with the constituent parts of the original branch network. As Figure 9As shown below, the association learning module consists of a deformable convolution layer, a 1x1 convolution layer, and an average pooling layer. The deformable convolution layer is the core of the module. It uses a deformed convolution kernel when performing convolution operations. This is mainly because the global spatial association of local objects is generally not a regular geometric shape, and its association logic can be more accurately modeled through a deformed convolution kernel. The execution process of the deformable convolution is very simple. Before performing the convolution operation, it needs to obtain the convolution kernel offset of the current pixel position through a branch. The offset includes X offset and Y offset (because the convolution kernel parameters usually only need to pay attention to the size dimension). Then the current position of the convolution kernel parameter plus the offset value is used as its actual effective position. Considering that the coordinates of the position may be floating point numbers, the feature map pixel value of the position corresponding to the convolution kernel parameter can be obtained using bilinear interpolation. After completing the deformable convolution operation, a 1x1 convolution operation and an average pooling operation (non-global average pooling, no size change) will be performed, which is mainly used to smooth the output results. It should be noted that the association relationship learning module is only a bypass of the branch expansion network connection position, and the original modules will still be directly connected. In this round of training, because it generally focuses on fine-grained scenes, it is easier to have data category imbalance and cross-category feature overlap. Therefore, this round of training uses focal loss as the main part of the loss function. This loss will give more training attention to categories with a smaller number, and it is also suitable as a multi-label training loss. In addition, the present invention also uses the confidence of each scene in the main part as a regular term to improve the efficiency of this round of training. The format of the loss function is as follows:

[0106]

[0107] Where L focus represents the standard focus loss, Represents the confidence score of the main part of the image for a certain scene category i, R is a regular term, and the present invention uses the L2 regular term as the penalty term for expansion. Branch expansion can be performed at any level of the main recognition feature extraction layer and expanded in a tree-like manner.

[0108] The technical solution of the present invention brings the following beneficial effects:

[0109] The present invention uses a three-stage training scheme to train the main feature extraction part of the model from three perspectives: abstract features, global information features, and local object features, so that the model has the ability to extract high-richness features and make scene discrimination based on them, greatly improving the scene recognition accuracy.

[0110] The present invention combines the idea of a feature pyramid to extract global information features from multiple scales, avoiding the loss of global spatial correlation information caused by excessive downsampling and non-linear transformation, providing high-quality global information features, and improving the recognition ability of background class scenes.

[0111] The present invention provides the ability to extract local object features for different levels through local supervised learning at multiple levels. Compared with the extraction of local object features at a single level, it reduces the loss of fine-grained scene decision-making information and enriches the local object features.

[0112] The present invention enhances the attention of the local supervised learning module to different channels through an attention mechanism, strengthens the activation of important local object features, and points the way for subsequent Fisher coding.

[0113] The present invention first proposes to optimize based on the aggregated heat map, combined with the importance of local objects at the block diagram level, using a new per-pixel binary Sigmoid loss, forcing the local supervised learning module to focus on the learning of "important local objects" and reducing the interference of "unimportant local objects" and "background" on decision-making.

[0114] The present invention uses Fisher convolution coding to extract feature vectors from the feature map, reducing redundancy while avoiding information loss due to over-abstraction.

[0115] In the main training stage, in order to increase the richness of features, the present invention uses multi-branch residual convolution as the basic module to ensure the feature extraction ability; while in the model branch expansion stage, the present invention uses strategies such as depthwise separable convolution and shared local learning modules to reduce overhead.

[0116] The present invention first proposes to use deformable convolution to build an association relationship learning module, and accurately model the association relationship of local objects using the geometric flexibility of deformable convolution.

[0117] The present invention also uses the scene confidence of the main body part as a regularization term, combined with focus loss, to well optimize the fine-grained scene recognition with class imbalance.

[0118] In the first round of the main body training stage of the model, focal loss can also be used to fully train only the core feature extraction layer, and then the global information feature extraction module is trained separately.

[0119] The global information feature extraction module can simply use two layers of transposed convolution to complete both size upsampling and channel expansion at the same time, but this will slow down the convergence speed.

[0120] The global information feature extraction module can also use a channel-level attention mechanism and a fully connected layer to complete feature fusion.

[0121] The local supervised learning module can be trained together with the fully connected layer by using the image-level semantic labels in combination with the auxiliary loss.

[0122] The fine-grained branch expansion network can also be extended on the existing branch expansion network, without necessarily starting from the main network as the expansion starting point.

[0123] The main part of the model can also use the basic module based on depthwise separable convolution to reduce the overhead. At the same time, the n×n convolution can be transformed into equivalent 1×n and n×1 convolutions to reduce the overhead.

[0124] For the learning of association relationships, a dedicated loss function can be designed and trained independently at multiple levels, without the need to be mixed in the branch expansion network for training.

[0125] Embodiment 6:

[0126] Figure 10 The following is a schematic diagram of the scene recognition process provided by the embodiment of the present invention. This process includes:

[0127] S201: Obtain the image to be recognized.

[0128] S202: Input the image to be recognized into the pre-trained scene recognition model, and determine the scene information corresponding to the image to be recognized based on the scene recognition model.

[0129] In the scene recognition method provided by the embodiment of the present invention, the electronic device to which the method is applied can be an intelligent device such as a PC or a tablet computer, or can also be a server. The electronic device for performing scene recognition can be the same as or different from the electronic device for performing model training in the above embodiment.

[0130] Since the process of model training is generally offline, the electronic device for performing model training can directly save the trained scene recognition model in the electronic device for performing scene recognition by using the method in the above embodiment, so that the subsequent electronic device for performing scene recognition can directly perform corresponding processing through the trained scene recognition model.

[0131] In the embodiment of the present invention, the image input to the scene recognition model is used as the image to be recognized. After obtaining the image to be recognized, the image to be recognized is input into the pre-trained scene recognition model, and the scene information corresponding to the image to be recognized is determined based on the scene recognition model.

[0132] Embodiment 7:

[0133] Figure 11 The following is a schematic diagram of the structure of the scene recognition model training device provided by the embodiment of the present invention. This device includes:

[0134] The first training unit 11 is configured to train the parameters of the core feature extraction layer and the global information feature extraction layer by using the first scene label of the sample image and the standard cross-entropy loss.

[0135] The second training unit 12 is configured to train the weight parameters of the LCS modules at each level according to the loss value calculated pixel by pixel from the feature maps output by the LCS modules at each level and the first scene label of the sample image.

[0136] The third training unit 13 is configured to train the parameters of the fully connected decision layer by using the first scene label of the sample image and the standard cross-entropy loss.

[0137] The device further includes:

[0138] The fourth training unit 14 is configured to train the weight parameters of the convolutional layers at each level of the branch expansion structure according to the loss value calculated pixel by pixel from the feature maps output by the convolutional layers of the branch expansion structure and the second scene label of the sample image; and train the parameters of the local object association relationship learning module by using a loss function with a scene confidence regularization term; wherein, the granularities of the first scene label and the second scene label are different.

[0139] The first training unit 11 is specifically configured to perform upsampling on the feature maps at different levels in the core feature extraction layer by using deconvolution operations with different dilation factors, align the number of channels in the channel dimension by using the bilinear interpolation algorithm, add and merge the feature maps at each level channel by channel, perform convolutional fusion on the merged feature map group, obtain a global information feature vector through global average pooling channel by channel, splice the global information feature vector and the fully connected layer FC feature vector, and train the parameters of the core feature extraction layer and the global information feature extraction layer by using the standard cross-entropy loss.

[0140] The second training unit 12 is specifically configured to obtain the importance weight of each channel by using an activation function through the attention mechanism in the channel dimension, perform weighted summation on the feature maps of each channel according to the importance weight of each channel to obtain a summary heat map; calculate the loss value pixel by pixel according to the summary heat map, the importance degree of object-scene association, and the area of the object, and train the weight parameters of the LCS modules at each level according to the loss value.

[0141] Embodiment 8:

[0142] Figure 12 The structural schematic diagram of the scene recognition device provided by the embodiment of the present invention, the device includes:

[0143] The acquisition unit 21 is configured to acquire an image to be recognized;

[0144] An identification unit 22, configured to input the image to be identified into a pre-trained scene recognition model, and determine scene information corresponding to the image to be identified based on the scene recognition model.

[0145] The apparatus further includes:

[0146] A determination unit 23, configured to determine that the image to be identified is a violation image if the scene information corresponding to the determined image to be identified belongs to violation scene information and the machine review result is that the image to be identified is a violation image.

[0147] Embodiment 9:

[0148] Based on the above embodiments, an electronic device is further provided in an embodiment of the present invention. As Figure 13 shown, it includes a processor 301, a communication interface 302, a memory 303, and a communication bus 304. Among them, the processor 301, the communication interface 302, and the memory 303 complete communication with each other through the communication bus 304;

[0149] A computer program is stored in the memory 303. When the program is executed by the processor 301, the processor 301 is caused to execute the following steps:

[0150] Training the parameters of the core feature extraction layer and the global information feature extraction layer through the first scene label of the sample image and the standard cross-entropy loss;

[0151] Training the weight parameters of each level of the LCS module according to the loss value calculated pixel by pixel from the feature map output by each level of the LCS module and the first scene label of the sample image;

[0152] Training the parameters of the fully connected decision layer through the first scene label of the sample image and the standard cross-entropy loss.

[0153] Based on the same inventive concept, an electronic device is further provided in an embodiment of the present invention. Since the principle of solving problems by the above electronic device is similar to the scene recognition model training method, the implementation of the above electronic device can refer to the implementation of the method, and the repeated parts will not be described again.

[0154] Embodiment 10:

[0155] Based on the above embodiments, an electronic device is further provided in an embodiment of the present invention. As Figure 14 shown, it includes a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete communication with each other through the communication bus 404;

[0156] The memory 403 stores a computer program, and when the program is executed by the processor 401, the processor 401 is caused to execute the following steps:

[0157] Obtain an image to be recognized;

[0158] Input the image to be recognized into a pre-trained scene recognition model, and determine scene information corresponding to the image to be recognized based on the scene recognition model.

[0159] Based on the same inventive concept, an electronic device is further provided in an embodiment of the present invention. Since the principle of the above electronic device for solving problems is similar to that of the scene recognition method, the implementation of the above electronic device can refer to the implementation of the method, and repeated parts will not be described again.

[0160] Embodiment 11:

[0161] Based on the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program executable by an electronic device. When the program runs on the electronic device, the electronic device is caused to execute the following steps when executed:

[0162] Train the parameters of the core feature extraction layer and the global information feature extraction layer through the first scene label of the sample image and the standard cross-entropy loss;

[0163] Train the weight parameters of the LCS modules at each level according to the loss value calculated pixel by pixel from the feature maps output by the LCS modules at each level and the first scene label of the sample image;

[0164] Train the parameters of the fully connected decision layer through the first scene label of the sample image and the standard cross-entropy loss.

[0165] Based on the same inventive concept, a computer-readable storage medium is further provided in an embodiment of the present invention. Since the principle of the processor for solving problems when executing the computer program stored on the above computer-readable storage medium is similar to that of the scene recognition model training method, the implementation of the processor for executing the computer program stored on the above computer-readable storage medium can refer to the implementation of the method, and repeated parts will not be described again.

[0166] Embodiment 12:

[0167] Based on the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program executable by an electronic device. When the program runs on the electronic device, the electronic device is caused to perform the following steps when executing:

[0168] Obtain an image to be recognized;

[0169] Input the image to be recognized into a pre-trained scene recognition model, and determine the scene information corresponding to the image to be recognized based on the scene recognition model.

[0170] Based on the same inventive concept, an embodiment of the present invention also provides a computer-readable storage medium. Since the principle of the processor solving problems when executing the computer program stored on the above computer-readable storage medium is similar to that of the scene recognition method, the implementation of the processor executing the computer program stored on the above computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be described again.

[0171] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the function specified in one block or multiple blocks.

[0172] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the function specified in one block or multiple blocks.

[0173] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the function specified in Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the function specified in one block or multiple blocks.

[0174] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0175] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for training a scene recognition model, characterized in that, The scene recognition model includes a core feature extraction layer, a global information feature extraction layer connected to the core feature extraction layer, LCS modules at each level, and a fully connected decision layer. The method includes: The parameters of the core feature extraction layer and the global information feature extraction layer are obtained by training through the first scene label of the sample image and the standard cross entropy loss; the sample image is used to indicate whether the image is illegal; Training weight parameters of the LCS modules at each level according to the feature maps output by the LCS modules at each level and the loss values ​​calculated pixel by pixel for the first scene label of the sample image; The parameters of the fully connected decision layer are obtained by training through the first scene label of the sample image and the standard cross entropy loss; The weight parameters of the LCS modules at each level are trained by calculating the loss values ​​pixel by pixel based on the feature maps output by the LCS modules at each level and the first scene labels of the sample images, including: The importance weight of each channel is obtained by using an activation function through the attention mechanism of the channel dimension, and the feature map of each channel is weighted summed according to the importance weight of each channel to obtain a summary heat map; Calculating the loss value pixel by pixel according to the summary heat map, the object scene association importance and the area of ​​the object, and training the weight parameters of the LCS modules of each level according to the loss value; The core feature extraction layer includes a first-class grouping multi-receptive field residual convolution module and a second-class grouping multi-receptive field residual convolution module.

2. The method according to claim 1, characterized in that, The scene recognition model also includes a branch expansion structure; the branch expansion structure includes a convolution layer and a local object association relationship learning module; According to the loss value calculated pixel by pixel based on the feature map output by the convolution layer of the branch extension structure and the second scene label of the sample image, the weight parameters of the convolution layers at each level of the branch extension structure are trained; the parameters of the local object association relationship learning module are trained through the loss function with a scene confidence regularization term; wherein the granularity of the first scene label and the second scene label are different.

3. The method according to claim 1, characterized in that, The first-class grouping multi-receptive field residual convolution module includes a first grouping, a second grouping and a third grouping, the first grouping, the second grouping and the third grouping have different convolution sizes, and the first grouping, the second grouping and the third grouping include a residual calculation bypass structure; each grouping outputs a feature map through a convolution operation and residual calculation, and the feature map output by each grouping is spliced ​​in the channel dimension and channel shuffled, and is output to the next module after convolution fusion; The second type of grouped multi-receptive field residual convolution module includes a fourth group, a fifth group and a sixth group. The convolution sizes of the fourth group, the fifth group and the sixth group are different. The fifth group and the sixth group respectively include a 1×1 convolution bypass structure and a residual calculation bypass structure; the feature map output by each group is spliced ​​in the channel dimension and channel shuffled, and is output to the next module after convolution fusion.

4. The method according to claim 1, characterized in that, The parameters of the core feature extraction layer and the global information feature extraction layer obtained by training through the first scene label of the sample image and the standard cross entropy loss include: Upsample the feature maps of different levels in the core feature extraction layer using deconvolution operations with different dilation factors, align the number of channels using bilinear interpolation algorithm in the channel dimension, add and merge the feature maps of each level channel by channel, perform convolutional fusion on the merged feature map group, and obtain the global information feature vector through global average pooling of each channel. Concatenate the global information feature vector and the fully connected layer FC feature vector, and train the parameters of the core feature extraction layer and the global information feature extraction layer through standard cross-entropy loss.

5. The method according to claim 2, characterized in that, The branch expansion structure is constructed using the depthwise separable convolutional residual block DW. In the main path of the residual block, a DW convolutional layer is used in the middle layer, and 1x1 convolutional layers are used before and after the DW convolution.

6. The method according to claim 2, characterized in that, The local object association relationship learning module includes a deformable convolutional layer, a convolutional layer, and an average pooling layer; The deformable convolutional layer obtains the offset value of the convolutional kernel at the current pixel position, adds the offset value to the current position of the convolutional kernel parameters as its actual effective position, obtains the pixel value of the feature image at the actual effective position, and outputs a feature map after convolutional operation and average pooling operation.

7. A scene recognition method for a scene recognition model trained by the method according to any one of claims 1-6, characterized in that, The method includes: Obtain an image to be recognized; Input the image to be recognized into a pre-trained scene recognition model, and determine the scene information corresponding to the image to be recognized based on the scene recognition model.

8. The method according to claim 7, characterized in that, The method further includes: If the scene information corresponding to the image to be recognized determined belongs to the illegal scene information, and the machine review result is that the image to be recognized is an illegal image, then determine that the image to be recognized is an illegal image.

9. A scene recognition model training device based on the scene recognition model training method according to claim 1, characterized in that, The device includes: The first training unit is used to train the parameters of the core feature extraction layer and the global information feature extraction layer through the first scene label of the sample image and standard cross-entropy loss; The second training unit is used to train the weight parameters of each level of the LCS module according to the loss value calculated pixel by pixel from the feature maps output by each level of the LCS module and the first scene label of the sample image; The third training unit is used to train the parameters of the fully connected decision layer through the first scene label of the sample image and standard cross-entropy loss; Specifically, the second training unit uses an activation function through an attention mechanism in the channel dimension to obtain the importance weight of each channel, performs weighted summation on the feature maps of each channel according to the importance weight of each channel to obtain a summary heat map; calculates the loss value pixel by pixel according to the summary heat map, the object-scene association importance degree, and the area of the object, and trains the weight parameters of each level of the LCS module according to the loss value.

10. A scene recognition device for a scene recognition model trained by the device according to claim 9, characterized in that, The device includes: An acquisition unit is used to acquire an image to be recognized; A recognition unit is used to input the image to be recognized into a pre-trained scene recognition model, and determine the scene information corresponding to the image to be recognized based on the scene recognition model.

11. An electronic device, characterized in that, The electronic device includes a processor, and the processor is used to implement the steps of the model training method as described in any one of claims 1-6, or implement the steps of the scene recognition method as described in any one of claims 7-8 when executing the computer program stored in the memory.

12. A computer-readable storage medium, characterized in that, It stores a computer program which, when executed by a processor, implements the steps of the model training method as described in any one of claims 1-6, or implements the steps of the scenario recognition method as described in any one of claims 7-8.

Citation Information

Patent Citations

  • Machine review model training and video machine review method, device, equipment and storage medium

    CN112926429A

Cited By

  • Model training and scene recognition method and apparatus, device, and medium

    WO2023056889A1