Scene recognition model training method, scene recognition method and model training device

Through the methods of global scene feature attention extraction and local feature merging, the problems of complex training and low accuracy in the existing scene recognition methods are solved, and the effect of simplifying the training process and improving the recognition accuracy is achieved.

CN113723159BActive Publication Date: 2025-08-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110222817.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-26
Publication Date
2025-08-22
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

Existing scene recognition methods require large-scale annotation of objects in the sample, and the training is complex and the recognition accuracy is low in scenarios without detection targets.

Method used

By extracting the global scene features of the training image, local features are obtained, and global and local features are merged to calculate local and fusion prediction loss values, simplifying the model training process and reducing the need for manual labeling.

Benefits of technology

This reduces the complexity of model training, improves the accuracy and applicability of scene recognition, and simplifies the end-to-end model with a one-stage model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113723159B_ABST
    Figure CN113723159B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a scene recognition model training method, a scene recognition method and a model training device. The scene recognition model training method obtains local features by performing attention extraction on the global scene features, obtains local prediction losses by using the local features, obtains fused features by merging the global scene features with the local features, obtains fused prediction losses by using the fused features, and then obtains a total prediction loss value based on the local prediction loss and the fused prediction loss to perform parameter correction on the scene recognition model. Since the local prediction loss value and the fused prediction loss value are calculated respectively by the scene category label of the training image in this embodiment, there is no need to label the local features of the training image, which can reduce the investment in manual labeling and the complexity of model training. The method can be widely used in the field of image recognition technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a scene recognition model training method, a scene recognition method, and a model training device. Background Art

[0002] Scene recognition is a hot topic in computer vision technology. It can be used to obtain scene information from images. It has a wide range of applications, including automated monitoring, human-computer interaction, video indexing, and image indexing. Scene recognition is more challenging than general object recognition because scene features often exist in the context of the scene being recognized. Conventional scene recognition methods typically focus on extracting features from specific objects or parts. This approach requires extensive annotation of objects in the sample when training scene recognition models, making training more complex. Summary of the Invention

[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0004] The embodiments of the present invention provide a scene recognition model training method, a scene recognition method, and a model training device, which can reduce the complexity of model training.

[0005] In one aspect, an embodiment of the present invention provides a scene recognition model training method, comprising the following steps:

[0006] Obtaining a training image and a scene category label of the training image;

[0007] Inputting the training image into a scene recognition model to obtain a first scene classification result and a target scene classification result;

[0008] Obtain a local prediction loss value based on the first scene classification result and the scene category label, obtain a fused prediction loss value based on the target scene classification result and the scene category label, and obtain a total prediction loss value based on the local prediction loss value and the fused prediction loss value;

[0009] Modifying parameters of the scene recognition model according to the total prediction loss value;

[0010] The step of inputting the training image into a scene recognition model to obtain a first scene classification result and a target scene classification result includes:

[0011] The global scene features of the training image are extracted through the scene recognition model, attention extraction is performed on the global scene features to obtain local features, and scene category prediction is performed on the local features to obtain a first scene classification result; the global scene features and the local features are merged to obtain a fusion feature, and scene category prediction is performed on the fusion feature to obtain a target scene classification result.

[0012] On the other hand, an embodiment of the present invention further provides a scene recognition method, comprising the following steps:

[0013] Obtain the image to be recognized;

[0014] Inputting the image to be recognized into a scene recognition model to obtain a target scene classification result;

[0015] The scene recognition model is trained using the above-mentioned scene recognition model training method.

[0016] On the other hand, an embodiment of the present invention further provides a scene recognition model training device, comprising:

[0017] A sample acquisition unit, configured to acquire a training image and a scene category label of the training image;

[0018] a recognition unit configured to input the training image into a scene recognition model, extract global scene features of the training image through the scene recognition model, perform attention extraction on the global scene features to obtain local features, perform scene category prediction on the local features to obtain a first scene classification result; merge the global scene features and the local features to obtain a fused feature, perform scene category prediction on the fused feature to obtain a target scene classification result;

[0019] a loss value calculation unit, configured to obtain a local prediction loss value based on the first scene classification result and the scene category label, obtain a fused prediction loss value based on the target scene classification result and the scene category label, and obtain a total prediction loss value based on the local prediction loss value and the fused prediction loss value;

[0020] A parameter correction unit is used to correct the parameters of the scene recognition model according to the total prediction loss value.

[0021] On the other hand, an embodiment of the present invention further provides a scene recognition device, comprising:

[0022] An image acquisition unit, configured to acquire an image to be recognized;

[0023] An image recognition unit, configured to input the image to be recognized into a scene recognition model to obtain a target scene classification result;

[0024] The scene recognition model is trained using the above-mentioned scene recognition model training method.

[0025] On the other hand, an embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned scene recognition model training method or scene recognition method when executing the computer program.

[0026] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, which stores a program, and the program is executed by a processor to implement the above-mentioned scene recognition model training method or scene recognition method.

[0027] In another aspect, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to implement the above-described scene recognition model training method or scene recognition method.

[0028] The embodiments of the present invention include at least the following beneficial effects: the embodiments of the present invention obtain local features by performing attention extraction on the global scene features, obtain local prediction losses using the local features, obtain fused features by merging the global scene features with the local features, obtain fused prediction losses using the fused features, and then obtain a total prediction loss value based on the local prediction loss and the fused prediction loss to perform parameter correction on the scene recognition model. Since the local prediction loss value and the fused prediction loss value are calculated separately using the scene category labels of the training images in this embodiment, there is no need to label the local features of the training images, which can reduce the investment in manual labeling and reduce the complexity of model training.

[0029] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purposes and other advantages of the present invention can be realized and obtained through the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation to the technical solution of the present invention.

[0031] Figure 1 This is a brief flowchart of the salient region extraction in the related art provided by an embodiment of the present invention;

[0032] Figure 2This is a schematic diagram of an optional architecture of a data processing system provided by an embodiment of the present invention;

[0033] Figure 3 is a flow chart of a scene recognition model training method provided by an embodiment of the present invention;

[0034] Figure 4 This is a schematic diagram of a structure of a scene recognition model provided by an embodiment of the present invention;

[0035] Figure 5 This is a schematic diagram of a model structure of a deep residual network layer 101 provided by an embodiment of the present invention;

[0036] Figure 6 This is a flowchart of the specific steps of extracting attention from global scene features to obtain local features, provided by an embodiment of the present invention;

[0037] Figure 7 This is a flowchart of the specific steps of extracting the original image frame corresponding to each candidate point from the training image provided by an embodiment of the present invention;

[0038] Figure 8 is a schematic diagram of a process for obtaining candidate points and corresponding magnified areas provided by an embodiment of the present invention;

[0039] Figure 9 This is a flowchart of the specific steps for selecting a target image frame from an original image frame provided by an embodiment of the present invention;

[0040] Figure 10 This is a flowchart of the specific steps of obtaining a vector corresponding to a target image frame and obtaining a second eigenvector provided by an embodiment of the present invention;

[0041] Figure 11 This is another specific step flow chart of obtaining a vector corresponding to a target image frame and obtaining a second eigenvector provided by an embodiment of the present invention;

[0042] Figure 12 This is a flowchart of the specific steps of merging global scene features and local features to obtain fused features, provided by an embodiment of the present invention;

[0043] Figure 13 This is a complete flowchart of the scene recognition model training method provided by an embodiment of the present invention;

[0044] Figure 14 This is another complete flow chart of the scene recognition model training method provided by an embodiment of the present invention;

[0045] Figure 15 is a flow chart of a scene recognition method provided by an embodiment of the present invention;

[0046] Figure 16 is a processing diagram of the scene recognition method provided by an embodiment of the present invention;

[0047] Figure 17 This is a flowchart of personalized recommendation in the scene recognition method provided by an embodiment of the present invention;

[0048] Figure 18 is a schematic diagram of applying the scene recognition method provided by an embodiment of the present invention to scene recognition;

[0049] Figure 19 is a schematic diagram of applying the scene recognition method provided by an embodiment of the present invention to personalized recommendation;

[0050] Figure 20 Schematic diagram of the structure of the scene recognition model training device provided by an embodiment of the present invention;

[0051] Figure 21 is a schematic structural diagram of a scene recognition device provided by an embodiment of the present invention;

[0052] Figure 22 is a block diagram of a portion of the structure of a mobile phone related to a terminal device provided by an embodiment of the present invention;

[0053] Figure 23 It is a structural diagram of a server provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0055] It should be understood that in the description of the embodiments of the present invention, "multiple" (or multiple) means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, and "above," "below," and "within" are understood to include the number itself. The terms "first," "second," and so on are used solely to distinguish technical features and are not to be construed as indicating or implying relative importance, or implicitly indicating the number of the indicated technical features, or implicitly indicating the order of the indicated technical features.

[0056] Before further explaining the embodiments of the present invention in detail, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are subject to the following interpretations:

[0057] Scene recognition: The goal of scene recognition is to determine the different types of scenes in an image. Unlike image classification, which classifies objects within an image, its goal is to classify local objects that occupy the main area of ​​the image. Image scene recognition requires a global consideration of multiple object categories in the image, rather than simply judging based on the category of local objects. For example, in order to determine whether an image belongs to a "beach" scene, it is necessary to analyze and determine whether objects of multiple categories such as "sand", "sea", and "blue sky" exist in the image at the same time. Conversely, if you simply judge based on whether there is a local object of the category "sand" in the image, you will not be able to correctly distinguish between the two different scene categories of "beach" and "desert".

[0058] Video understanding, as one of the applications of scene recognition, primarily involves identifying the scene in which the plot takes place within a video. Scene recognition is more difficult than general object recognition because scene features are often present in the background environment used for scene recognition. Conventional scene recognition methods typically focus on extracting features from specific objects or parts, which can easily lead to overfitting of the foreground in the target scene. This means that the scene recognition model memorizes the foreground in certain scenes (such as the clothing of foreground characters) rather than identifying the features of the surrounding background environment. Background environment features can occur in a variety of situations: one in which key environmental objects are concentrated in one place, and another in which key environmental objects are distributed across multiple locations. For example, a classroom study room has study tables and chairs, while a library study room has study tables and chairs plus multiple rows of bookshelves. The background of a classroom study room is concentrated tables and chairs, while the background of a library study room is distributed bookshelves. Therefore, ignoring the features of the background environment will reduce the accuracy of scene recognition.

[0059] In related technologies, scene recognition can be performed based on multi-scale salient region feature learning. Figure 1 , Figure 1 This is a brief flowchart of salient region extraction in related technologies. Specifically, the scene is first detected. One or more regions containing objects are obtained from the object detection results. Combined with the potential object density, the detected object regions are then captured at different scales to obtain salient regions at the optimal scale. Finally, the model is trained based on the obtained salient regions.

[0060] However, in the above-mentioned related technologies, whether it is the training process or the recognition process, it is necessary to first establish a target detection and positioning model, and then establish a scene recognition model, that is, the entire model is a two-stage model, which makes the model structure complicated; and, when training the target detection and positioning model, a large number of objects in the sample need to be labeled, and the training complexity is relatively high; in addition, not all scenes will have detection targets, such as the seaside, large forests, etc. At this time, the accuracy of scene recognition through the above-mentioned related technologies is not high.

[0061] Based on this, the embodiments of the present invention provide a scene recognition model training method, a scene recognition method and a model training device, which can reduce the complexity of model training and help improve the accuracy of recognition.

[0062] The scene recognition model training method and scene recognition method provided in the embodiments of the present invention can both be applied to artificial intelligence (AI). Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a manner similar to human intelligence. Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning, and decision-making.

[0063] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0064] Computer vision (CV) is the study of how machines can "see." Specifically, it refers to the use of cameras and computers to replace the human eye in identifying, detecting, and measuring objects. This is further processed by the computer to create images more suitable for human observation or transmission to instrumentation. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems capable of extracting information from images or multidimensional data.

[0065] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0066] The following describes exemplary applications of electronic devices implementing embodiments of the present invention. The electronic devices provided by embodiments of the present invention can be implemented as various types of user terminals, such as smartphones, tablet computers, laptop computers, and smart wearable devices. They can also be implemented as servers, where the server is a backend server that runs one or more applications including audio data processing, speech recognition, and text recognition. The following describes exemplary applications of electronic devices implemented as servers.

[0067] Reference Figure 2 , which is an optional architectural diagram of a data processing system 200 provided in an embodiment of the present invention. To support an exemplary application, terminals (terminal 210 and terminal 220 are shown as examples) are connected to a server 240 via a network 230. Network 230 can be a wide area network or a local area network, or a combination of the two, using a wireless link to achieve data transmission. It is understood that in other embodiments, the number of terminals is not limited to two. Figure 2 The number of terminals is for illustrative purposes only.

[0068] The scene recognition model training device provided in the embodiment of the present invention can be implemented as hardware or a combination of hardware and software. The following uses the scene recognition model training device implemented as server 240 to illustrate various exemplary implementations of the scene recognition model training method provided in the embodiment of the present invention.

[0069] Among them, the server 240 can be a background server corresponding to terminals such as mobile phones, computers, digital broadcast terminals, information transceiver equipment, game consoles, tablet devices, medical equipment, fitness equipment, personal digital assistants, etc. For example, it can be a background server corresponding to a terminal installed with a corresponding client. According to the structure of the server 240, it can be foreseen that the device is implemented as an exemplary structure of a terminal. Therefore, the structure described here should not be regarded as a limitation. For example, some components described below can be omitted, or components not described below can be added to meet the special needs of certain applications.

[0070] based on Figure 2 The data processing system 200 shown in FIG. Figure 3 An embodiment of the present invention provides a scene recognition model training method, which is described by taking the method applied in the server 240 as an example, wherein the method includes but is not limited to the following steps 301 to 304.

[0071] Step 301: Obtain training images and scene category labels of the training images;

[0072] The training images are used to train the scene recognition model. They can be obtained from the Internet or directly input locally. Specifically, the training images are selected from a candidate image set according to certain rules. For example, the candidate image set may contain 100,000 images. Random sampling can be performed, that is, randomly selecting images from the candidate image set as training sample images. Alternatively, weighted random sampling can be performed, where images are sampled based on the sampling weights of the images in the candidate sample image set. The larger the sampling weight, the greater the probability that the image will be used as a training sample image.

[0073] The scene category label is the scene category to which the training image belongs. It is used as a reference for calculating the loss value during model training. For example, the scene category label of image A is restaurant, the scene category label of image B is library, and so on.

[0074] Step 302: Input the training image into the scene recognition model to obtain a first scene classification result and a target scene classification result.

[0075] Specifically, after the training image is input into the scene recognition model, the global scene features of the training image are extracted through the scene recognition model, attention extraction is performed on the global scene features to obtain local features, and the scene category is predicted on the local features to obtain the first scene classification result; the global scene features and local features are merged to obtain fused features, and the scene category is predicted on the fused features to obtain the target scene classification result.

[0076] Among them, global scene features are used to characterize information about the characteristics of the image scene, and can express the characteristics of the entire image scene, that is, they can be used to describe the overall characteristics of the scene. Local features are local expressions of image features, reflecting the local particularity of the image. In an embodiment of the present invention, local features are obtained by performing attention extraction on global scene features, wherein attention extraction can be achieved through an attention network. By extracting local features with attention, significant features of the background can be mined, avoiding ignoring local features during the scene recognition process and reducing recognition accuracy. In related scene recognition schemes, local features are generally extracted after detecting specific objects or parts. However, not all scenes have corresponding detection targets, such as the seaside, large forests, etc. Therefore, the embodiment of the present invention obtains local features by performing attention extraction on global scene features, not only for specific objects or parts in the scene. Compared with the method of extracting local features by detecting specific objects or parts, it has a wider applicability in scene recognition and can avoid the problem of local parts of the image not being recognized.

[0077] It should be noted that the first scene classification result obtained by predicting the scene category of local features is different from the object classification result corresponding to the local features. The first scene classification result is for the scene category, such as restaurant, library, etc., while in related technologies, local features are used to obtain object classification results, such as people, animals, etc.

[0078] Among them, since the fusion features are obtained by merging global scene features and local features, the global scene features and local features of the image can be considered at the same time, so that the fusion features can more comprehensively represent the image scene, thereby making the scene recognition results more accurate.

[0079] Step 303: Obtain a local prediction loss value based on the first scene classification result and the scene category label, obtain a fusion prediction loss value based on the target scene classification result and the scene category label, and obtain a total prediction loss value based on the local prediction loss value and the fusion prediction loss value.

[0080] The loss value can be obtained based on a loss function, which is a function used to represent the "risk" or "loss" of an event. In one embodiment, a cross-entropy loss function can be used to calculate the local prediction loss value or the fusion prediction loss value. Specifically, the calculation formula is as follows:

[0081] L=-[ylogy , +(1-y)log(1-y , )]

[0082] Where L represents the loss value, y represents the scene category label, and y′ represents the first scene classification result or the target scene classification result. After obtaining the local prediction loss value and the fusion prediction loss value, the local prediction loss value and the fusion prediction loss value can be summed to obtain the total prediction loss value. In one embodiment, there can be multiple local loss values, and the total prediction loss value can be the sum of the multiple local prediction loss values ​​and the fusion prediction loss value.

[0083] Step 304: Modify the parameters of the scene recognition model according to the total prediction loss value.

[0084] In the related technologies that introduce local features for model training, the local feature extraction model is generally trained first, and then the recognition model is trained. That is, the local feature extraction loss value and the recognition loss value are optimized separately. The trained model is a two-stage model with high training complexity. In addition, this method requires the local features of the training image to be labeled in order to calculate the local feature extraction loss value, which increases the investment in manual labeling.

[0085] In an embodiment of the present invention, the local prediction loss value is calculated by using the scene category label. On the one hand, the attention extraction of global scene features can be located in the area associated with the image scene. On the other hand, the local prediction loss value and the fusion prediction loss value are both obtained based on the scene category label. Therefore, the scene recognition model training method provided by the embodiment of the present invention only needs to label the training image with the global scene category label, without labeling the local features, which can reduce the investment in manual labeling and reduce the complexity of model training.

[0086] On this basis, the total prediction loss value is obtained by using the local prediction loss value and the fusion prediction loss value, and the total prediction loss value is used to train the scene recognition model. This is beneficial for avoiding the problem of ignoring local features in the scene recognition process and improving the accuracy of scene recognition. It can also make the trained scene recognition model a one-stage end-to-end model, simplify the model structure, and reduce the complexity of model training.

[0087] In one embodiment, when training a scene recognition model, the scene recognition model can also be used to predict the scene category of the global scene features to obtain a second scene classification result, and a global prediction loss value can be obtained based on the second scene classification result and the scene category label. Based on this, in the above step 304, the parameters of the scene recognition model are corrected according to the total prediction loss value, specifically, the total prediction loss value can be obtained based on the local prediction loss value, the fusion prediction loss value and the global prediction loss value. Since the global prediction loss value is introduced as one of the bases for calculating the total prediction loss value, the extraction of global scene features is more accurate, which is conducive to improving the efficiency of scene recognition model training. Among them, the global prediction loss value can also be obtained using the calculation formula of the above-mentioned cross entropy loss function. In addition, since the global prediction loss value is also obtained based on the scene category label, it will not affect the complexity of the model training.

[0088] Among them, the total prediction loss value is obtained according to the local prediction loss value, the fusion prediction loss value and the global prediction loss value. Specifically, the total prediction loss value can be obtained by summing up. For example, loss represents the total prediction loss value, loss cr Represents the global prediction loss value, loss locate Represents the local prediction loss value, loss all Represents the fusion prediction loss value, then:

[0089] loss=loss cr +loss locate +loss all

[0090] On this basis, the weights of the global prediction loss value, the local prediction loss value, and the fusion prediction loss value can be introduced respectively, then:

[0091] loss = a * loss cr +b*loss locate +c*loss all

[0092] Among them, a, b, and c can be set according to actual needs, and the embodiment of the present invention does not limit them. By calculating the total prediction loss value in a weighted manner, the total prediction loss value is made more reasonable, which is conducive to improving the training effect of the scene recognition model.

[0093] Reference Figure 4 , Figure 4 A structural schematic diagram of one of the scene recognition models provided in an embodiment of the present invention, wherein the scene recognition model includes a basic recognition network 410, an attention extraction module 420, a local prediction module 430 and a fusion prediction module 440, and the basic recognition network 410 includes a feature extraction module 411 and a global prediction module 412.

[0094] In one embodiment, the basic recognition network 410 can be a deep neural network, and the feature extraction module 411 can use the parameters of ResNet101 (deep residual network 101 layer) pre-trained on the ImageNet dataset. A model structure of the deep residual network 101 layer is as follows: Figure 5 As shown, the deep residual network 101 layer can be a three-layer residual module for reducing the number of parameters. 3x3 represents the size of the convolution kernel, and 64 represents the number of channels. A plus sign in a circle represents addition, that is, identity mapping. ReLU (Rectified Linear Unit, linear rectification function) indicates activation using an activation function. 256-d represents that the input is 256-dimensional. Referring to Table 1, Table 1 is a structural table of ResNet101 in one embodiment, wherein x3, x4, and x23 represent 3 modules, 4 modules, and 23 modules, respectively. The volume ResNet101 includes multiple consecutive convolutional layers. There are 5 types of convolutional layers, and Conv5_x is the 5th convolutional layer.

[0095] Table 1

[0096]

[0097]

[0098] Pooling can be understood as compression, which aggregates features at different locations. For example, calculating the average value of a specific feature in a region of an image as the value of that region can reduce dimensionality while improving results and preventing overfitting. This aggregation operation is called pooling. Pooling includes average pooling and max pooling. The above method uses the average value of a specific feature in a region as the value of that region, which is called average pooling. The same method uses the maximum value of a specific feature in a region as the value of that region, which is called max pooling.

[0099] The global prediction module 412 may include a Max pool layer and a Full connetction (FC) layer. Referring to Table 2, which is a structural table of the global prediction module in one embodiment, Pool_cr and Fc_cr may be initialized using a Gaussian distribution with a variance of 0.01 and a mean of 0. Conv5 in the feature extraction module outputs the deep features of the global scene of the training image, obtains the first feature vector corresponding to the output global scene feature, and performs pooling processing through the Pool_cr layer. Then, the predicted probability distribution of N scene categories is obtained through the Fc_cr layer, and finally the second scene classification result is obtained. Among them, the output size of the Pool_cr layer can be 1x2048, and the output size of the Fc_cr layer can be 1xN. The second scene classification result is finally obtained based on the predicted probability distribution of N scene categories, which can be implemented using a softmax linear regression function.

[0100] Table 2

[0101] The name of the layer Output size layer Pool_cr 1x2048 Max Pooling Fc_cr 1xN Fully connected

[0102] In one embodiment, in addition to correcting the parameters of the scene recognition model according to the total prediction loss value, the parameters of the above-mentioned deep neural network can also be corrected separately according to the global prediction loss value. Among them, the second scene classification result obtained by the deep neural network can make the extraction of global scene features more accurate, which is conducive to improving the efficiency of scene recognition model training. Therefore, by correcting the parameters of the deep neural network separately, the accuracy of extracting global scene features can be further improved, and the overall training efficiency of subsequent scene recognition models can be further improved. Exemplarily, the convolution template parameter w and the bias parameter b of the deep neural network can be corrected by the gradient descent method based on SGD (Stochastic Gradient Descent).

[0103] Refer to Table 3, which is a structural table of the attention extraction module 420 in one embodiment, wherein the depth features of the global scene of the training image output by Conv5 are used as the input of the Down1_y layer. The function of the Down1_y layer is to perform spatial compression on the output of Conv5. The output of the Down1_y layer is used as the input of the Propost2_y layer. The function of the Down1_y layer is to perform channel compression on the output of the Down1_y layer. The output size of the Down1_y layer can be 19x31, and the output size of the Propost2_y layer can be 9x15. It is understandable that the output sizes of the Down1_y layer and the Propost2_y layer can be adjusted according to actual conditions, and the embodiments of the present invention do not limit this.

[0104] Finally, the topK second eigenvectors can be identified from the vector output by the Propost2_y layer to represent the topK local features.

[0105] Table 3

[0106]

[0107] Refer to Table 4, which is a structural table of the local prediction module 430 in one embodiment, wherein the input of the Fc_locate layer is the topK local features, and the output size of the Fc_locate layer is topKxN, that is, the output of the Fc_locate layer is the probability distribution of each local feature belonging to N scene categories, and finally the first scene classification result is obtained.

[0108] Table 4

[0109] The name of the layer Output size layer Fc_locate topKxN Fully connected

[0110] Refer to Table 5, which is a structural table of the fusion prediction module 440 in one embodiment, wherein the input of the Fc_al l layer is the fusion feature, and the output size of the Fc_all layer is 1xN, that is, the output of the Fc_all layer is the probability distribution of the fusion feature belonging to N scene categories, and finally the target scene classification result is obtained.

[0111] Table 5

[0112] The name of the layer Output size layer Fc_all 1xN Fully connected

[0113] Based on the example Figure 4 The scene recognition model shown, refer to Figure 6 In the above step 302, attention extraction is performed on the global scene features to obtain local features, which can further include steps 601 to 604, wherein steps 601 to 604 can be applied to the server 240.

[0114] Step 601: compress the first eigenvector to obtain a compressed eigenvector.

[0115] Among them, the compressed feature vector represents the attention intensity of each spatial coordinate in the compressed first feature vector. Taking the structure of the attention extraction module shown in Table 3 as an example, if m training images are input, the matrix size output by the Down1_y layer is mx128x19x31, where 128 represents the number of channels and 19x31 represents the length and width of the space after convolution. Then, after being processed by the Propost2_y layer, the output matrix size is mx6x9x15, where 6 represents the number of channels and 9x15 represents the length and width of the space after convolution. The point in 9x15 represents the attention intensity of the spatial coordinate where the point is located. At this time, the matrix of size mx6x9x15 is the compressed feature vector.

[0116] Step 602: Perform matrix transformation on the compressed feature vector to obtain candidate points corresponding to each attention intensity.

[0117] Among them, the reshape function can be used to perform matrix transformation processing on the compressed feature vector. The reshape function is a function that transforms the specified matrix into a matrix of a specific dimension, and the number of elements in the matrix after the transformation remains unchanged. Taking the structure of the attention extraction module shown in Table 3 as an example, the matrix of size mx6x9x15 is transformed, and finally 6x9x15=810 candidate points are obtained.

[0118] Step 603: extracting the original image frame corresponding to each candidate point from the training image, and filtering out the target image frame from the original image frame.

[0119] Since the compressed feature vector is obtained by compressing the first feature vector, each candidate point is actually compressed as well. Therefore, after amplifying each candidate point, it can be mapped back to the original image frame in the training image. Each original image frame must then be screened to identify the target image frame with high attention intensity before further local features are obtained. For example, if there are 810 original image frames, the number of target image frames obtained by screening may be 4.

[0120] Step 604: Obtain the vector corresponding to the target image frame to obtain a second eigenvector.

[0121] Among them, after the target image frame is obtained by screening from the original image frame, the vector corresponding to the target image frame is determined as the second eigenvector, which is used as the input of the Fc_locate layer in Table 4.

[0122] In one embodiment, referring to Figure 7In the above step 603 , extracting the original image frame corresponding to each candidate point from the training image may further include steps 701 to 702 , wherein steps 701 to 702 may be applied to the server 240 .

[0123] Step 701: performing a magnification process on each candidate point to obtain a magnified area corresponding to each candidate point, and determining the size of each magnified area according to the compression ratio of the compression process;

[0124] Among them, each candidate point has been compressed, for example, from 19x31 to 9x15, and the compression ratio can be determined. The enlargement of each candidate point can actually be regarded as the reverse process of the compression process. Therefore, the enlargement ratio and the compression ratio can be consistent. For example, from 9x15 to 19x31, the size of each enlarged area can be determined according to the compression ratio.

[0125] Step 702: Based on the position of each candidate point in the first eigenvector and the size of each magnified area, the plane coordinates of each magnified area in the training image are obtained, and the original image frame corresponding to each candidate point is extracted from the training image based on the plane coordinates of each magnified area.

[0126] The position of the magnified area in the training image can be determined based on the position of each candidate point in the first feature vector, and the plane coordinates of the magnified area in the training image can be determined based on the size of the magnified area. The original image frame corresponding to the candidate point can be extracted from the training image based on the plane coordinates of the magnified area. In one embodiment, the shape of the magnified area can be a rectangle, the candidate point can be the center of the magnified area, and the coordinates of the original image frame can be box(x1, y1, x2, y2). Figure 8 , Figure 8 This is a schematic diagram of the correspondence between candidate points and magnified areas. The first eigenvector is compressed by the Down1_y layer and the Propost2_y layer in sequence to obtain the candidate point. Correspondingly, the candidate point 801 is also magnified accordingly and finally corresponds back to the original image box 802 in the training image, that is, box(x1, y1, x2, y2).

[0127] In one embodiment, referring to Figure 9 In the above step 603 , filtering out the target image frame from the original image frame may further include steps 901 to 904 , wherein steps 901 to 904 may be applied to the server 240 .

[0128] Step 901: Obtain the confidence level corresponding to each original image frame.

[0129] Confidence refers to the probability that the true value falls within a certain range, centered around the measured value. A higher confidence level for a given original image frame indicates a more accurate positioning of that original image frame.

[0130] Step 902: Sort the confidences, and obtain candidate image frames according to the confidence ranking results.

[0131] The confidence levels may be sorted from highest to lowest or from lowest to highest, which is not limited in the present embodiment. The candidate image frame may be obtained based on the confidence level sorting result, and the original image frame with the highest confidence level may be retained as the candidate image frame. For example, if the confidence levels corresponding to original image frame A, original image frame B, original image frame C, and original image frame D are A1, A2, A3, and A4, respectively, and A1, A2, A3, and A4 are sorted from highest to lowest as A2, A1, A4, and A3, then the candidate image frame is original image frame B.

[0132] Step 903: Obtain the intersection-over-union ratios between the original image frames other than the candidate image frame and the candidate image frame.

[0133] The IoU refers to the overlap ratio between two image frames, that is, the ratio of their intersection to their union. The larger the IoU, the greater the overlap ratio of the two image frames. The ideal situation is that the two image frames completely overlap, that is, the ratio is 1. Based on the example of step 902, the IoU obtained in step 903 above includes: the IoU between the original image frame A and the original image frame B, the IoU between the original image frame C and the original image frame B, and the IoU between the original image frame D and the original image frame B.

[0134] Step 904: The candidate image frame and the original image frame whose IoU ratio is less than or equal to the threshold are used as the target image frame.

[0135] The threshold value can be set according to actual conditions, for example, 0.5, etc., and is not limited in the embodiment of the present invention. Based on the example of step 903, assuming that the intersection-and-union ratio between original image frame A and original image frame B is 0.6, the intersection-and-union ratio between original image frame C and original image frame B is 0.7, and the intersection-and-union ratio between original image frame D and original image frame B is 0.2, then the target image frames are original image frame B and original image frame D.

[0136] The specific principles of the above steps 901 to 904 are explained below with an actual example.

[0137] Now there are 5 original image frames box1, box2, box3, box4 and box5, with confidence levels of 0.8, 0.9, 0.7, 0.5 and 0.3 respectively. Then the order of these 5 original image frames from large to small according to confidence level is: box2>box1>box3>box4>box5. According to the confidence ranking results of these 5 original image frames, we can first determine that the target image frame is box2, and then calculate the intersection and union ratio between box1, box3, box4, box5 and box2 respectively. If the intersection and union ratio is greater than the preset threshold of 0.5, the corresponding box will be deleted. Specifically:

[0138] Intersection-over-union (box1, box2) = 0.1 < 0.5, keep box1;

[0139] Intersection-over-Union (box3, box2) = 0.7 > 0.5, delete box3;

[0140] Intersection-over-Union (box4, box2) = 0.6 > 0.5, delete box4;

[0141] Intersection-over-Union (box5, box2) = 0.8 > 0.5, delete box5;

[0142] The final target image boxes are box1 and box2.

[0143] In one embodiment, the above process may be performed iteratively, for example:

[0144] Intersection-over-union (box1, box2) = 0.1 < 0.5, keep box1;

[0145] Intersection-over-Union (box3, box2) = 0.7 > 0.5, delete box3;

[0146] Intersection-over-union (box4, box2) = 0.2 < 0.5, keep box4;

[0147] Intersection-over-union ratio (box5, box2) = 0.3 < 0.5, keep box5;

[0148] At this time, repeat the above process for box1, box4 and box5. First, sort box1, box4 and box5 by confidence: box1>box4>box5, and then calculate the intersection and union ratio between box4, box5 and box1 respectively. The final result is:

[0149] Intersection-over-Union (box4, box1) = 0.7 > 0.5, delete box4;

[0150] Intersection-over-Union (box5, box1) = 0.8 > 0.5, delete box5;

[0151] The final target image boxes are box1 and box2.

[0152] It is understandable that in other embodiments, the number of original image frames, the confidence ranking, and the intersection-and-union ratio may vary according to actual conditions. The above examples are merely schematic illustrations of the specific principles of steps 901 to 904. The condition for stopping iteration may be that the intersection-and-union ratio comparison can no longer be performed (for example, the number of original image frames remaining after confidence ranking and deletion based on the intersection-and-union ratio size is less than 2), or the number of retained original image frames reaches a preset threshold.

[0153] In one embodiment, referring to Figure 10 In the above step 604 , obtaining the vector corresponding to the target image frame to obtain the second feature vector may further include steps 1001 to 1003 , wherein steps 1001 to 1003 may be applied to the server 240 .

[0154] Step 1001: Obtain the plane coordinates corresponding to the target image frame in the training image;

[0155] Step 1002: extracting a target image block from the training image according to the plane coordinates;

[0156] Step 1003: extract features from the target image block to obtain a second feature vector.

[0157] The plane coordinates corresponding to the target image frame in the training image can determine the position of the target image frame, and the plane coordinates of the target image frame can be determined by the amplification method in the above steps 701 to 702. After the target image block is extracted from the training image according to the plane coordinates, the target image block can be obtained by Figure 4 The feature extraction module in the basic recognition network 410 performs feature extraction to obtain a second feature vector.

[0158] In one embodiment, referring to Figure 11 In the above step 604, the vector corresponding to the target image frame is obtained to obtain the second feature vector. In addition to the method of extracting the target image block from the training image to perform feature extraction in the above steps 1001 to 1003, the following steps may be further included:

[0159] Step 1101: Obtain the position of the candidate point corresponding to the target image frame in the first feature vector;

[0160] Step 1102: extract the vector corresponding to the candidate point from the first feature vector according to the position to obtain a second feature vector.

[0161] Among them, the target image frame has a corresponding candidate point, and the candidate point is obtained by the first eigenvector after compression processing and matrix transformation processing. Therefore, the position of the candidate point in the first eigenvector can be determined according to the corresponding matrix transformation processing and magnification processing, and then the part of the eigenvector in the first eigenvector corresponding to the candidate point can be determined to obtain the second eigenvector.

[0162] Figure 10 and Figure 11 Two methods of extracting the second eigenvector are shown respectively. Figure 10 The second eigenvector extraction method shown is to perform secondary feature extraction using the feature extraction module after determining the target image block. The advantage is that it can improve the accuracy of the second eigenvector, improve the training accuracy of the scene recognition model, and improve the scene recognition accuracy in subsequent model applications. Figure 11 The second feature vector extraction method shown is to extract it directly from the first feature vector through the position of the candidate point. This can avoid the time-consuming model operation problem caused by secondary feature extraction, and can improve the training efficiency of the scene recognition model and the scene recognition efficiency in subsequent model applications.

[0163] In one embodiment, referring to Figure 12 The fused feature is obtained by merging the global scene feature and the local feature, so the fused feature can be represented by a third feature vector. In the above step 302, the global scene feature and the local feature are merged to obtain the fused feature, which can further include steps 1201 to 1203, wherein steps 1201 to 1203 can be applied to the server 240.

[0164] Step 1201: performing pooling processing on the first eigenvector;

[0165] Step 1202: performing pooling processing on the second eigenvector;

[0166] Step 1203: Connect the first eigenvector and the second eigenvector after the pooling process end to end to obtain a third eigenvector.

[0167] Among them, since the first eigenvector and the second eigenvector may be multi-dimensional vectors, after the first eigenvector is pooled, a corresponding one-dimensional vector can be obtained, which is convenient for subsequent end-to-end connection processing, and the pooling process can be maximum pooling. By connecting the first and second eigenvectors after pooling, the third eigenvector can be obtained relatively simply and conveniently. For example, after the first eigenvector is pooled, a 1x2048 eigenvector can be obtained, and the second eigenvector can have topK eigenvectors, that is, the second eigenvector is pooled to obtain topK 1x2048 eigenvectors. Finally, the first and second eigenvectors after pooling are connected end to end, and the size of the obtained third eigenvector is (1+topK)x2048.

[0168] Reference Figure 13 , Figure 13 A complete flow chart of the scene recognition model training method provided in an embodiment of the present invention, wherein the scene recognition model training method includes steps 1301 to 1315, wherein steps 1301 to 1315 can be applied to the server 240.

[0169] Step 1301: Obtain a training image and a scene category label of the training image;

[0170] Step 1302: extract features from the training image using a deep neural network to obtain a first feature vector;

[0171] Step 1303: Predicting the scene category of the first feature vector using a deep neural network to obtain a first scene classification result;

[0172] Step 1304: Obtain a global prediction loss value based on the first scene classification result and the scene category label;

[0173] Step 1305: Modify the parameters of the deep neural network according to the global prediction loss value;

[0174] Step 1306: compress the first eigenvector to obtain a compressed eigenvector representing attention intensity;

[0175] Step 1307: Perform matrix transformation on the compressed feature vector to obtain multiple candidate points;

[0176] Step 1308: Obtain the original image frame corresponding to the candidate point from the training image, and select the topK target image frames from the original image frame;

[0177] Step 1309: extracting a target image block from the training image according to the target image frame, performing feature extraction on the target image block, and obtaining a second feature vector;

[0178] Step 1310: performing scene category prediction on the second feature vector to obtain a second scene classification result;

[0179] Step 1311: Obtain a local prediction loss value based on the second scene classification result and the scene category label;

[0180] Step 1312: Merge the first eigenvector and the second eigenvector to obtain a third eigenvector;

[0181] Step 1313: performing scene category prediction on the third eigenvector to obtain a target scene classification result;

[0182] Step 1314: Obtain a fusion prediction loss value based on the target scene classification result and the scene category label;

[0183] Step 1315: Sum the global prediction loss value, the local prediction loss value, and the fusion prediction loss value to obtain a total prediction loss value, and modify the parameters of the scene recognition model according to the total prediction loss value.

[0184] In step 1305, the parameters of the deep neural network are first modified based on the global prediction loss value, which can make the extraction of global scene features more accurate and help improve the efficiency of scene recognition model training. Of course, in other embodiments, the parameters of the deep neural network can also be modified without using the global prediction loss value.

[0185] Steps 1301 to 1315 above respectively calculate the global prediction loss value, the local prediction loss value, and the fused prediction loss value using the scene category labels of the training images, thereby eliminating the need to label the local features of the training images, reducing the investment in manual labeling, and lowering the complexity of model training. Furthermore, using the total prediction loss value to train the scene recognition model and the local prediction loss value as an auxiliary parameter not only helps avoid the problem of ignoring local features during the scene recognition process and improves the accuracy of scene recognition, but also allows the trained scene recognition model to be a one-stage end-to-end model, simplifying the model structure and reducing the complexity of model training.

[0186] Reference Figure 14 , Figure 14 Another complete flow chart of the scene recognition model training method provided in an embodiment of the present invention includes steps 1301 to 1415, wherein steps 1401 to 1415 can be applied to the server 240.

[0187] Step 1401: Obtain a training image and a scene category label of the training image;

[0188] Step 1402: extract features from the training image using a deep neural network to obtain a first feature vector;

[0189] Step 1403: Predicting the scene category of the first feature vector using a deep neural network to obtain a first scene classification result;

[0190] Step 1404: Obtain a global prediction loss value based on the first scene classification result and the scene category label;

[0191] Step 1405: Modify the parameters of the deep neural network according to the global prediction loss value;

[0192] Step 1406: compressing the first eigenvector to obtain a compressed eigenvector representing attention intensity;

[0193] Step 1407: Perform matrix transformation on the compressed feature vector to obtain multiple candidate points;

[0194] Step 1408: Obtain the original image frame corresponding to the candidate point from the training image, and select the topK target image frames from the original image frame;

[0195] Step 1409: extracting the corresponding vector from the first feature vector according to the position of the candidate point corresponding to the target image frame in the first feature vector to obtain a second feature vector;

[0196] Step 1410: performing scene category prediction on the second feature vector to obtain a second scene classification result;

[0197] Step 1411: Obtain a local prediction loss value based on the second scene classification result and the scene category label;

[0198] Step 1412: Merge the first eigenvector and the second eigenvector to obtain a third eigenvector;

[0199] Step 1413: performing scene category prediction on the third eigenvector to obtain a target scene classification result;

[0200] Step 1414: Obtain a fusion prediction loss value based on the target scene classification result and the scene category label;

[0201] Step 1415: Sum the global prediction loss value, the local prediction loss value, and the fusion prediction loss value to obtain a total prediction loss value, and modify the parameters of the scene recognition model according to the total prediction loss value.

[0202] In step 1405, the parameters of the deep neural network are first modified based on the global prediction loss value, which can make the extraction of global scene features more accurate and help improve the efficiency of scene recognition model training. Of course, in other embodiments, the parameters of the deep neural network can also be modified without using the global prediction loss value.

[0203] Steps 1401 to 1415 above respectively calculate the global prediction loss value, the local prediction loss value, and the fusion prediction loss value using the scene category labels of the training images, thereby eliminating the need to label the local features of the training images, reducing the investment in manual labeling, and lowering the complexity of model training. Furthermore, using the total prediction loss value to train the scene recognition model and the local prediction loss value as an auxiliary parameter not only helps avoid the problem of ignoring local features during the scene recognition process and improves the accuracy of scene recognition, but also allows the trained scene recognition model to be a one-stage end-to-end model, simplifying the model structure and reducing the complexity of model training.

[0204] The parameters of the scene recognition model are modified by modifying the parameters in Tables 1 to 5 in the above embodiments.

[0205] Reference Figure 15 Based on the scene recognition model obtained by the scene recognition model training method in the above embodiment, an embodiment of the present invention also provides a scene recognition method, including but not limited to the following steps 1501 to 1502, wherein steps 1501 to 1502 can be applied to the server 240.

[0206] Step 1501: Obtain an image to be recognized;

[0207] Step 1502: Input the image to be recognized into the scene recognition model to obtain the target scene classification result.

[0208] Among them, after the image to be recognized is input into the scene recognition model, the first eigenvector of the global scene feature is extracted by the scene recognition model, and the attention extraction is performed on the first eigenvector to obtain the second eigenvector of the local feature. The first eigenvector and the second eigenvector are then merged to obtain the third eigenvector of the fusion feature. The scene category prediction is performed on the third eigenvector to obtain the target scene classification result. The target scene classification result is the final scene recognition result. For example, referring to Figure 16, the input image to be identified is an image of a balcony, and feature extraction is performed on the image to be identified to obtain global scene features. After attention extraction is performed on the global scene features, local features 1601, local features 1602, local features 1603, and local features 1604 can be obtained. Then, the global scene features, local features 1601, local features 1602, local features 1603, and local features 1604 are merged to obtain fusion features. The fusion features are subjected to scene category prediction, and the scene of the image to be identified is a balcony. If the scene recognition model of the related technology is used to perform scene recognition on the image to be identified, the scene obtained may be a room. Since the scene recognition model of the embodiment of the present invention performs attention extraction on the image to be identified, for example Figure 16 The local feature 1602 in the image is a floor-to-ceiling window, which will affect the final scene recognition result, making the final scene recognition result a balcony, thereby avoiding the problem of ignoring local features in the scene recognition process and improving the accuracy of scene recognition.

[0209] The scene recognition method provided by the embodiment of the present invention can be exemplarily applied to personalized recommendation. Based on this, the image to be recognized can be obtained from the terminal. Figure 17 The above scene recognition method may further include steps 1701 to 1702 , wherein steps 1701 to 1702 may be applied to the server 240 .

[0210] Step 1701: Obtain recommended content for the corresponding terminal based on the target scene classification result;

[0211] Step 1702: Send recommended content to the terminal.

[0212] Among them, different users use different images on the terminal. For example, when a user uses the terminal to watch videos, the scene types involved in different videos may be different. By performing scene recognition on the terminal image, personalized recommended content can be provided to the user for easy viewing by the user, and the target scene classification result is obtained based on the scene recognition method provided by the embodiment of the present invention, so the accuracy is high, which can make the recommended content more accurate and more targeted.

[0213] It is understood that the steps in the above-mentioned various embodiments are only schematically applied to the server 240. In addition to being applicable to the server 240, the steps in the above-mentioned various embodiments can also be applied to the terminal (210, 220). Moreover, although the steps in the above-mentioned various flow charts are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in the present embodiment, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the above-mentioned flow charts may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0214] The following uses specific examples to illustrate application scenarios of the scene recognition method according to the embodiment of the present invention.

[0215] Reference Figure 18 , the scene recognition method of the embodiment of the present invention can be applied to the display of scene recognition results, wherein the user inputs the picture to be recognized through the front end A, and the picture to be recognized can be downloaded from the Internet or photographed using the camera component of the terminal. Front end A can be the interface of an application such as image recognition, and the back end performs scene recognition on the picture to be recognized through the scene recognition method provided by the embodiment of the present invention. The back end can run locally on the user's terminal or in a server. If the back end runs in the server, the picture input by the user can be transmitted to the server through a mobile network, a wireless network, Bluetooth or other communication methods. After the back end obtains the recognition result, it will eventually return the recognition result to the front end B for display to the user, wherein the front end A and the front end B can be the same interface or different interfaces. In this example, the scene recognition method can be called by the user operating the front end A.

[0216] based on Figure 18The processing flow shown is described below. Next, another example of applying the scene recognition method of an embodiment of the present invention to the display of scene recognition results is described, in which a user watches a video through front-end A. Front-end A can be the interface of an application such as a video player. Front-end A can be provided with a scene recognition button. When the user needs to know what scene the current plot is, the scene recognition method provided by the embodiment of the present invention can be called through the scene recognition button. The video player will then extract the image frames of the moments before and after the current playback moment and send them to the back-end. The back-end performs scene recognition on the image to be recognized through the scene recognition method provided by the embodiment of the present invention. Similarly, the back-end can run locally on the user's terminal or on the server. If the back-end runs on the server, the image frames extracted by the video player can be transmitted to the server via mobile networks, wireless networks, Bluetooth and other communication methods. After the back-end obtains the recognition result, it will eventually return the recognition result to the front-end B for display to the user.

[0217] Reference Figure 19 The scene recognition method of the embodiment of the present invention can be applied to personalized recommendations, wherein a user watches a video through front-end A, which can be the interface of an application such as a video player. The video player automatically calls the scene recognition method provided by the embodiment of the present invention, and then the video player extracts the image frames of the moments before and after the current playback moment and sends them to the back-end. The back-end performs scene recognition on the image to be recognized using the scene recognition method provided by the embodiment of the present invention. Similarly, the back-end can run locally on the user's terminal or on a server. If the back-end runs on the server, the image frames extracted by the video player can be transmitted to the server via a mobile network, a wireless network, Bluetooth, or other communication methods. After the back-end obtains the recognition result, it can obtain recommended content based on the recognition result and return the recommended content to front-end B for display to the user. The recommended content can be text content or video content. For example, based on the recognition result, it can predict the type of video that the user usually likes to watch and recommend the same type of video content to the user. It is understandable that front-end A and front-end B can belong to the same application or two different applications, such as two applications with associated accounts.

[0218] Reference Figure 20 , an embodiment of the present invention further provides a scene recognition model training device, comprising:

[0219] The sample acquisition unit 2001 is used to acquire training images and scene category labels of the training images;

[0220] Recognition unit 2002 is configured to input a training image into a scene recognition model, extract global scene features of the training image through the scene recognition model, perform attention extraction on the global scene features to obtain local features, perform scene category prediction on the local features to obtain a first scene classification result; merge the global scene features and the local features to obtain a fused feature, perform scene category prediction on the fused feature to obtain a target scene classification result;

[0221] A loss value calculation unit 2003 is configured to obtain a local prediction loss value based on the first scene classification result and the scene category label, obtain a fused prediction loss value based on the target scene classification result and the scene category label, and obtain a total prediction loss value based on the local prediction loss value and the fused prediction loss value;

[0222] The parameter correction unit 2004 is used to correct the parameters of the scene recognition model according to the total prediction loss value.

[0223] The above-mentioned scene recognition model training device and scene recognition model training method are based on the same inventive concept. Therefore, the scene recognition model training device does not need to label the local features of the training image, which can reduce the investment in manual labeling, reduce the complexity of model training, and improve the accuracy of the scene recognition model.

[0224] Reference Figure 21 , an embodiment of the present invention further provides a scene recognition device, comprising:

[0225] An image acquisition unit 2101 is used to acquire an image to be recognized;

[0226] The image recognition unit 2102 is used to input the image to be recognized into the scene recognition model to obtain the target scene classification result;

[0227] The scene recognition model is obtained by training using the scene recognition model training method in the above embodiment.

[0228] The scene recognition device and the scene recognition method are based on the same inventive concept, so the scene recognition device can avoid the problem of ignoring local features during the scene recognition process and improve the accuracy of scene recognition.

[0229] In addition, an embodiment of the present invention further provides an electronic device that can train a scene recognition model or perform scene recognition. The device is described below with reference to the accompanying drawings. Figure 21An embodiment of the present invention provides an electronic device, which may be a terminal device. The terminal device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), an in-vehicle computer, etc. Taking a mobile phone as an example:

[0230] Figure 22 FIG2 is a block diagram showing a partial structure of a mobile phone related to a terminal device provided by an embodiment of the present invention. Figure 22 The mobile phone includes components such as a radio frequency (RF) circuit 2210, a memory 2220, an input unit 2230, a display unit 2240, a sensor 2250, an audio circuit 2260, a wireless fidelity (WiFi) module 2270, a processor 2280, and a power supply 2290. It will be understood by those skilled in the art that Figure 22 The mobile phone structure shown in the figure does not constitute a limitation to the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0231] The following combination Figure 22 A detailed introduction to the various components of a mobile phone:

[0232] The RF circuit 2210 can be used to receive and send signals during information transmission or calls. In particular, after receiving downlink information from the base station, it is sent to the processor 2280 for processing. In addition, the designed uplink data is sent to the base station. Generally, the RF circuit 2210 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 2210 can also communicate with the network and other devices via wireless communication. The above-mentioned wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0233] The memory 2220 can be used to store software programs and modules. The processor 2280 executes the various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 2220. The memory 2220 may mainly include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 2220 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state memory device.

[0234] The input unit 2230 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile phone. Specifically, the input unit 2230 may include a touch panel 2231 and other input devices 2232. The touch panel 2231, also known as a touch screen, can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any other suitable object or accessory on or near the touch panel 2231) and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 2231 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch direction and detects the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch point coordinates, which are then sent to the processor 2280. It can also receive commands sent by the processor 2280 and execute them. In addition, the touch panel 2231 can be implemented using a variety of types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 2231, the input unit 2230 may further include other input devices 2232. Specifically, the other input devices 2232 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick.

[0235] The display unit 2240 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 2240 may include a display panel 2241. Optionally, the display panel 2241 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 2231 may cover the display panel 2241. When the touch panel 2231 detects a touch operation on or near it, it is transmitted to the processor 2280 to determine the category of the touch event. Subsequently, the processor 2280 provides corresponding visual output on the display panel 2241 according to the category of the touch event. Although in Figure 22 In the embodiment, the touch panel 2231 and the display panel 2241 are used as two independent components to realize the input and output functions of the mobile phone, but in some embodiments, the touch panel 2231 and the display panel 2241 can be integrated to realize the input and output functions of the mobile phone.

[0236] The mobile phone may also include at least one sensor 2250, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 2241 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 2241 and / or the backlight when the mobile phone is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described here.

[0237] Audio circuit 2260, speaker 2261, and microphone 2262 provide an audio interface between the user and the phone. Audio circuit 2260 converts received audio data into electrical signals and transmits them to speaker 2261, which then converts them into sound signals for output. Microphone 2262, on the other hand, converts collected sound signals into electrical signals, which are then received by audio circuit 2260 and converted into audio data. The audio data is then processed by processor 2280 and transmitted to, for example, another phone via RF circuit 2210, or stored in memory 2220 for further processing.

[0238] WiFi is a short-range wireless transmission technology. Mobile phones can help users send and receive emails, browse the web, and access streaming media through the WiFi module 2270. It provides users with wireless broadband Internet access. Figure 22 A WiFi module 2270 is shown, but it is understandable that it is not an essential component of the mobile phone and can be omitted as needed without changing the essence of the invention.

[0239] Processor 2280 is the control center of the phone, connecting all parts of the phone using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 2220 and accessing data stored in memory 2220, it performs various phone functions and processes data, thereby performing overall phone testing. Optionally, processor 2280 may include one or more processing units; preferably, processor 2280 may integrate an application processor and a modem processor, where the application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 2280.

[0240] The mobile phone also includes a power supply 2290 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 2280 through a power management system, thereby realizing functions such as charging, discharging, and power consumption management through the power management system.

[0241] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be described in detail here.

[0242] In this embodiment, the processor 2280 included in the terminal device is capable of executing the scene recognition model training method and the scene recognition method of the previous embodiment.

[0243] In the embodiment of the present invention, a server may also be used to execute the scene recognition model training method or the model training method, see Figure 23 As shown, Figure 23 This is a structural diagram of a server 2300 provided in an embodiment of the present invention. The server 2300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 2322 (for example, one or more processors) and a memory 2332, and one or more storage media 2330 (for example, one or more mass storage devices) for storing application programs 2342 or data 2344. Among them, the memory 2332 and the storage medium 2330 may be temporary storage or permanent storage. The program stored in the storage medium 2330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 2322 may be configured to communicate with the storage medium 2330 to execute a series of instruction operations in the storage medium 2330 on the server 2300.

[0244] The server 2300 may also include one or more power supplies 2326, one or more wired or wireless network interfaces 2350, one or more input and output interfaces 2358, and / or one or more operating systems 2341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0245] The processor in the server can be used to execute the scene recognition model training method or the scene recognition method.

[0246] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the scene recognition model training method or scene recognition method of each of the aforementioned embodiments.

[0247] Embodiments of the present invention further disclose a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the scene recognition model training method or scene recognition method of each of the aforementioned embodiments.

[0248] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0249] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0250] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0251] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0252] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0253] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0254] The step numbers in the above method embodiment are only provided for the convenience of explanation and do not limit the order of the steps. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0255] It should also be understood that the various implementations provided in the embodiments of the present invention can be arbitrarily combined to achieve different technical effects.

[0256] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above implementation. Those skilled in the art can also make various equivalent modifications or substitutions under the shared conditions that do not violate the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A scene recognition model training method, characterized in that: The following steps are involved: Obtaining a training image and a scene category label of the training image; Inputting the training image into a scene recognition model to obtain a first scene classification result and a target scene classification result; Obtain a local prediction loss value based on the first scene classification result and the scene category label, obtain a fused prediction loss value based on the target scene classification result and the scene category label, and obtain a total prediction loss value based on the local prediction loss value and the fused prediction loss value; Modifying parameters of the scene recognition model according to the total prediction loss value; The step of inputting the training image into a scene recognition model to obtain a first scene classification result and a target scene classification result includes: Extracting global scene features of the training image through the scene recognition model, performing attention extraction on the global scene features to obtain local features, performing scene category prediction on the local features to obtain a first scene classification result; merging the global scene features and the local features to obtain a fused feature, performing scene category prediction on the fused feature to obtain a target scene classification result, wherein the global scene features are represented by a first feature vector and the local features are represented by a second feature vector; performing attention extraction on the global scene features to obtain local features includes: The first eigenvector is compressed to obtain a compressed eigenvector, wherein the compressed eigenvector represents the attention intensity of each spatial coordinate in the compressed first eigenvector; the compressed eigenvector is matrix transformed to obtain candidate points corresponding to each attention intensity; the original image frame corresponding to each candidate point is extracted from the training image, and the target image frame is filtered out from the original image frame; the vector corresponding to the target image frame is obtained to obtain a second eigenvector.

2. The scene recognition model training method according to claim 1, characterized in that: The scene recognition model training method further includes: Performing scene category prediction on the global scene feature using the scene recognition model to obtain a second scene classification result, and obtaining a global prediction loss value based on the second scene classification result and the scene category label; Obtaining a total prediction loss value according to the local prediction loss value and the fusion prediction loss value includes: A total prediction loss value is obtained according to the local prediction loss value, the fusion prediction loss value and the global prediction loss value.

3. The scene recognition model training method according to claim 1, characterized in that: The extracting the original image frame corresponding to each candidate point from the training image includes: Performing an amplification process on each candidate point to obtain an amplified area corresponding to each candidate point, and determining a size of each amplified area according to a compression ratio of the compression process; Based on the position of each candidate point in the first feature vector and the size of each enlarged area, the plane coordinates of each enlarged area in the training image are obtained, and the original image frame corresponding to each candidate point is extracted from the training image according to the plane coordinates of each enlarged area.

4. The scene recognition model training method according to claim 1, characterized in that: The step of filtering out the target image frame from the original image frame includes: Obtaining the confidence level corresponding to each original image frame; Sorting the confidences, and obtaining candidate image frames according to the confidence ranking results; Obtaining an intersection-over-union ratio between the remaining original image frames except the candidate image frame and the candidate image frame; The candidate image frame and the original image frame whose intersection-over-union ratio is less than or equal to a threshold are used as target image frames.

5. The scene recognition model training method according to any one of claims 1 to 4, characterized in that: The acquiring of the vector corresponding to the target image frame to obtain the second eigenvector includes: Obtaining the plane coordinates corresponding to the target image frame in the training image; extracting a target image block from the training image according to the plane coordinates; Feature extraction is performed on the target image block to obtain a second feature vector.

6. The scene recognition model training method according to any one of claims 1 to 4, characterized in that: The acquiring of the vector corresponding to the target image frame to obtain the second eigenvector includes: Obtaining a position of a candidate point corresponding to the target image frame in the first feature vector; A vector corresponding to the candidate point is extracted from the first feature vector according to the position to obtain a second feature vector.

7. The scene recognition model training method according to claim 1, characterized in that: There are multiple local prediction loss values, and obtaining a total prediction loss value based on the local prediction loss values ​​and the fusion prediction loss value includes: The total prediction loss value is obtained by summing up the multiple local prediction loss values ​​and the fusion prediction loss value.

8. The scene recognition model training method according to claim 1, characterized in that: The global scene feature is represented by a first feature vector, the local feature is represented by a second feature vector, and the fused feature is represented by a third feature vector. The merging of the global scene feature and the local feature to obtain the fused feature includes: Performing pooling processing on the first feature vector; performing pooling processing on the second feature vector; The first eigenvector and the second eigenvector after the pooling process are connected end to end to obtain the third eigenvector.

9. The scene recognition model training method according to claim 1, characterized in that: The scene recognition model includes a plurality of consecutive convolutional layers, and extracting the global scene features of the training image through the scene recognition model includes: The training image is convolved through the multiple consecutive convolutional layers to obtain global scene features of the training image.

10. The scene recognition model training method according to claim 2, characterized in that: The scene recognition model includes a deep neural network, and performing scene category prediction on the global scene feature by the scene recognition model to obtain a second scene classification result includes: Performing scene category prediction on the global scene feature through the deep neural network to obtain a second scene classification result; The scene recognition model training method further includes: The parameters of the deep neural network are modified according to the global prediction loss value.

11. A scene recognition method, characterized in that: The following steps are involved: Obtain the image to be recognized; Inputting the image to be recognized into a scene recognition model to obtain a target scene classification result; Wherein, the scene recognition model is trained by the scene recognition model training method according to any one of claims 1 to 10.

12. The scene recognition method according to claim 11, characterized in that: The image to be recognized is obtained from a terminal, and the scene recognition method further includes: Obtaining recommended content corresponding to the terminal according to the target scene classification result; The recommended content is sent to the terminal.

13. A scene recognition model training device, characterized in that: include: A sample acquisition unit, configured to acquire a training image and a scene category label of the training image; A recognition unit is used to input the training image into a scene recognition model, extract the global scene features of the training image through the scene recognition model, perform attention extraction on the global scene features to obtain local features, perform scene category prediction on the local features to obtain a first scene classification result; merge the global scene features and the local features to obtain a fusion feature, perform scene category prediction on the fusion feature to obtain a target scene classification result, the global scene features are represented by a first feature vector, and the local features are represented by a second feature vector; the attention extraction on the global scene features to obtain local features includes: compressing the first feature vector to obtain a compressed feature vector, wherein the compressed feature vector represents the attention intensity of each spatial coordinate in the compressed first feature vector; performing matrix transformation on the compressed feature vector to obtain a candidate point corresponding to each attention intensity; extracting an original image frame corresponding to each candidate point from the training image, and filtering out a target image frame from the original image frame; obtaining a vector corresponding to the target image frame to obtain a second feature vector; a loss value calculation unit, configured to obtain a local prediction loss value based on the first scene classification result and the scene category label, obtain a fused prediction loss value based on the target scene classification result and the scene category label, and obtain a total prediction loss value based on the local prediction loss value and the fused prediction loss value; A parameter correction unit is used to correct the parameters of the scene recognition model according to the total prediction loss value.

14. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method for training a scene recognition model as described in any one of claims 1 to 10 is implemented, or the method for scene recognition as described in any one of claims 11 to 12 is implemented.

Citation Information

Patent Citations

  • Scene classification method and device based on self-supervision mechanism and regional suggestion network

    CN111062441A

  • Lightweight multi-branch pedestrian re-identification method and system based on attention mechanism

    CN111931624A