A scene recognition method, device, equipment, and storage medium
Through the multi-level image processing of the scene recognition model, the problem of judging the authenticity of scene images is solved, and accurate recognition of scene images and fraud prevention are achieved.
Patent Information
- Application Number
- CN202210698944.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-20
AI Technical Summary
The prior art is difficult to accurately determine whether the scene image is a real image, and there is a risk of fraud, such as the problem of synthetic photos spoofing surveillance cameras or cameras being moved illegally.
The scene recognition model is adopted, including the first sub-model, the second sub-model and the third sub-model. By acquiring multiple image frames, the significance prediction and attribution prediction models are used to output images, the feature map is fused and decisions are made, and the degree of matching between the scene image and the pre-stored scene map is judged.
Accurate recognition of scene images is realized, whether they match the preset scene image, prevent scene fraud, and improve the accuracy of the authenticity judgment of the surveillance image.
Smart Images

Figure CN115205778B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image technology, and particularly to a method, device, equipment and storage medium for scene recognition. Background Art
[0002] With the in-depth research on neural networks, neural networks are increasingly widely used in the field of image recognition.
[0003] For example, people can no longer use wallets, and when making electronic payments or conducting business in front of bank ATMs, they can no longer use passwords. Instead, they can use faces to replace passwords for payment operations or login operations; face recognition is used to determine whether the face in the current image matches the pre-stored customer face.
[0004] However, in some actual application scenarios, some problems are also faced. For example, criminals input synthetic photos into surveillance cameras to block the real scene; another example is that the cameras configured in bank ATMs are deliberately placed in other ATM locations, so that the original locations where the cameras are placed cannot be monitored.
[0005] In summary, a technical solution is needed to accurately determine whether a scene image is a real image. Summary of the Invention
[0006] Embodiments of this application provide a method, device, equipment and storage medium for customer recall. Banks can discover customers to be recalled with low stability, low activity and high churn risk among customer groups according to a screening mechanism, and send prompt messages to associated customers associated with the customers to be recalled. Through an incentive mechanism, the associated customers are urged to assist the bank in successfully recalling the customers to be recalled, thereby preventing customer churn and improving the overall customer operation efficiency of the bank.
[0007] Embodiments of this application provide a method for customer recall, which is applied to a scene recognition model. The scene recognition model includes a first sub-model, a second sub-model and a third sub-model. The method includes:
[0008] Obtain a scene image, where the scene image includes multiple image frames within a period of time;
[0009] Input the scene image into the trained scene recognition model;
[0010] Output a first image through the first sub-model, where the first image is used to highlight the recognition target object in the scene image;
[0011] Output a second image through the second sub-model, where the second image is used to represent the probability distribution of the contribution degree of each pixel in the scene image to the implementation of the prediction process, and the prediction process is the process of obtaining the second image from the scene image;
[0012] Fuse the first image and the second image to obtain a fused feature map;
[0013] Input the fused feature map into the third sub-model to obtain a classification score map, where the classification score map has a score and a category, the score is used to represent the matching degree between the scene image and the pre-stored scene map, and the category is used to represent the recognition result of the scene image.
[0014] In some embodiments, before obtaining the scene image, the method further includes:
[0015] Train the initial first sub-model and the initial second sub-model to obtain the trained first sub-model and second sub-model;
[0016] Use the trained first sub-model, the second sub-model, and the decision sub-model to obtain the trained scene recognition model.
[0017] Optionally, training the initial first sub-model and the initial second sub-model to obtain the trained first sub-model and second sub-model includes:
[0018] Obtain training samples, where the training samples include at least one first scene image with a score and at least one second scene image with a significant region, the original image corresponding to the first scene image is the same as the original image corresponding to the second scene image, and the score of the first scene image and the significant region of the second scene image are determined by the third sub-model;
[0019] Input the training samples into the initial scene recognition model to obtain the first loss function of the initial first sub-model and the second loss function of the initial second sub-model;
[0020] Repeat the above steps until the first loss function and the second loss function converge to obtain the trained scene recognition model.
[0021] Optionally, obtaining the first loss function of the initial first sub-model and the second loss function of the initial second sub-model includes:
[0022] Input the first scene image into the initial second sub-model to obtain the prediction attribution map of the first scene image;
[0023] Obtain the second loss function according to the predicted attribution map and the true attribution map of the first scenario image;
[0024] Input the second scenario image into the initial first sub-model to obtain the predicted saliency image of the second scenario image;
[0025] Obtain the first loss function of the initial first sub-model according to the predicted saliency image and the true saliency image of the first scenario image.
[0026] Optionally, the second loss function includes a region loss function and an area loss function. The region loss function is used to characterize the accuracy of the positions of the pixels included in the contribution degree probability distribution, and the area loss function is used to characterize the accuracy of the areas of the pixels included in the contribution degree probability distribution. Obtaining the second loss function includes:
[0027] Construct the region loss function according to the first prediction value in combination with the EM algorithm; the first prediction value is the prediction of the first scenario image of the known classification by the initial second model, and the first probability distribution of the positions of the pixels included in the first scenario image is obtained;
[0028] Construct the area loss function according to the second prediction value, and the area loss function is the sum value of all the probabilities included in the first probability distribution;
[0029] Calculate the sum of the region loss function and the area loss function, and the value of the sum is the value of the second loss function.
[0030] In some embodiments, before fusing the first image and the second image to obtain a fused feature map, the method further includes:
[0031] Calculate a first weight and a second weight;
[0032] The step of fusing the first image and the second image to obtain a fused feature map includes:
[0033] Fuse the first image and the second image according to the first weight and the second weight to obtain a fused feature map.
[0034] Optionally, calculating the first weight and the second weight includes:
[0035] Obtain a preset parameter and the score of a third scenario image; the third scenario image is any image frame in the scenario images;
[0036] Obtain the minimum value of the product of the preset parameter and the score of the third scenario image, and use the minimum value as the first weight for fusing the first image;
[0037] Obtain the difference between the preset value and the first weight, and use the difference as the second weight for fusing the second image.
[0038] An embodiment of the present application further provides a scene recognition device, including:
[0039] A scene image acquisition unit, configured to acquire a scene image, where the scene image includes multiple image frames within a period of time;
[0040] An input unit, configured to input the scene image into the trained scene recognition model;
[0041] A first sub-model unit, configured to output a first image through the first sub-model, where the first image is used to highlight a recognition target in the scene image;
[0042] A second sub-model unit, configured to output a second image through the second sub-model, where the second image is used to represent the probability distribution of the contribution degree of each pixel in the scene image to the realization of the prediction process, and the prediction process is the process of obtaining the second image from the scene image;
[0043] A fusion unit, configured to fuse the first image and the second image to obtain a fused feature map;
[0044] A third sub-model unit, configured to input the fused feature map into the third sub-model to obtain a scoring category map, where the scoring category map has a score and a category, the score is used to represent the matching degree between the scene image and a pre-stored scene image, and the category is used to represent the recognition result of the scene image.
[0045] An embodiment of the present application further provides a device, including a processor and a memory, where the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps in the above-mentioned scene recognition method.
[0046] An embodiment of the present application further provides a storage medium, where the storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the above-mentioned scene recognition method.
[0047] An embodiment of the present application can accurately recognize a scene image based on a scene recognition model, determine whether the scene image matches a preset scene image, so as to accurately determine whether the scene image is a real image, and thereby determine whether scene fraud occurs. Description of the Drawings
[0048] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0049] Figure 1a is an application scenario of a scene recognition method provided by an embodiment of the present application;
[0050] Figure 1b is a schematic flowchart of a scene recognition method provided by an embodiment of the present application;
[0051] Figure 2 is a schematic diagram of training a scene recognition model provided by an embodiment of the present application;
[0052] Figure 3 is a schematic structural diagram of a scene recognition device provided by an embodiment of the present application;
[0053] Figure 4 is a schematic structural diagram of a device provided by an embodiment of the present application. Detailed implementation manners
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0055] The terms "first", "second", etc. in the present application are used to distinguish different objects, rather than to describe a specific order. At the same time, the term "including" and any form of its transformation are intended to cover non-exclusive inclusion.
[0056] The embodiments of the present application provide a scene recognition method, device, device, and storage medium.
[0057] Among them, the scene recognition device can be specifically integrated in an electronic device, and the electronic device can be a terminal, a server, or other devices. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a notebook computer, or a personal computer (PC), etc.; the server can be a single server or a server cluster composed of multiple servers.
[0058] In some embodiments, the scene recognition device may also be integrated in multiple electronic devices. For example, the scene recognition device may be integrated in multiple servers, and the multiple servers are used to implement the scene recognition method of the present application.
[0059] In some embodiments, the server may also be implemented in the form of a terminal.
[0060] Reference Figure 1a , Figure 1a FIG. shows an application scenario of a scene recognition method according to an embodiment of the present application, that is, a user hopes to recognize the acquired monitoring image, and determine whether the monitoring camera is blocked, the lens angle is adjusted, or the camera is moved to another place according to the recognition result. For this purpose, the following steps are performed:
[0061] Obtain a monitoring image, where the monitoring image includes multiple image frames within a period of time.
[0062] Input the monitoring image into the trained scene recognition model.
[0063] Output a first image through the first sub-model, where the first image is used to highlight the recognition target object in the monitoring image.
[0064] Output a second image through the second sub-model, where the second image is used to represent the probability distribution of the contribution degree of each pixel in the monitoring image to the implementation of the prediction process, and the prediction process is the process of obtaining the second image from the monitoring image;
[0065] Fuse the first image and the second image to obtain a fused feature map;
[0066] Input the fused feature map into the third sub-model to obtain a scoring category map, where the scoring category map has a score and a category, the score is used to represent the matching degree between the monitoring image and a pre-stored monitoring scene map, and the category is used to represent the recognition result of the monitoring image.
[0067] According to the scoring category map, determine whether the monitoring image matches a pre-stored monitoring scene map, so as to determine whether scene fraud has occurred.
[0068] In summary, this embodiment can achieve accurate recognition of the monitoring image based on the scene recognition model, determine whether the monitoring image matches a preset monitoring scene map, so as to accurately determine whether the monitoring image is a real image, and thus determine whether scene fraud has occurred.
[0069] The following will be described in detail respectively. It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0070] In this embodiment, a scene recognition method related to machine learning is provided, which is applied to a scene recognition model. The scene recognition model includes a first sub-model, a second sub-model, and a third sub-model. As Figure 1b shown, the specific process of this method includes steps 110 to 160:
[0071] 110. Obtain a scene image, where the scene image includes multiple image frames within a period of time.
[0072] In some embodiments, the scene image generally refers to video images of any public or human occasion. The scene image can be obtained by a camera configured in a bank ATM, or by a surveillance camera in a certain factory building, or by an infrared camera outside a private residence.
[0073] In some embodiments, the scene recognition model includes a first sub-model, a second sub-model, and a third sub-model. The above three sub-models together constitute the scene recognition model. The above three sub-models and the scene recognition model can be neural network models, such as convolutional neural network models.
[0074] 120. Input the scene image into the trained scene recognition model.
[0075] Specifically, the scene recognition model may include a saliency prediction sub-model, an attribution prediction sub-model, and a decision-making sub-model, corresponding to the above-mentioned first sub-model, second sub-model, and third sub-model respectively.
[0076] Among them, the attribution prediction sub-model can be an F-CALM (Focus-Class Activation Latent Mapping) model improved from the CAM (Class Activation Mapping) model. The difference between the two is that the F-CALM model corresponding to the attribution prediction sub-model introduces the EM algorithm in the process of constructing the loss function on the basis of the CAM model, so as to give a deeper explanation of the contribution degree distribution of each pixel in the image by means of maximum likelihood estimation.
[0077] As Figure 2 shown, before step 120, the method of this application embodiment may further include:
[0078] Train the initial first sub-model and the initial second sub-model to obtain the trained first sub-model and second sub-model;
[0079] Use the trained first sub-model, the second sub-model, and the decision-making sub-model to obtain the trained scene recognition model.
[0080] As shown Figure 2 in Figure 2 the figure, it is a schematic diagram of a training scenario recognition model provided by an embodiment of the present application. In the embodiment of the present application, the scenario recognition model is trained by means of reinforcement learning. Among them, the third sub-model can be set as the agent of reinforcement learning to make a final decision on the recognition of the scenario image. The recognition processing of the first sub-model and the second sub-model on the scenario image can be set as the environment of reinforcement learning. It can be understood that since the first sub-model and the second sub-model are environmental factors of reinforcement learning, the intermediate prediction values of the first sub-model and the second sub-model will affect the final decision of the third sub-model. Specifically, one of the initial first sub-model and the initial second sub-model can be trained first, and after one of the models is completed, the other model can be trained to obtain the first sub-model and the second sub-model that are completed in training.
[0081] During the training process, the output of the reinforcement learning agent under the maximum reward mechanism can be used as a supervision signal. That is, the optimal output result of the third sub-model can be used as the label of the training sample for subsequent training of the first sub-model and the second sub-model.
[0082] Optionally, the step of "training the initial first sub-model and the initial second sub-model to obtain the trained first sub-model and the second sub-model" may include:
[0083] Obtain training samples;
[0084] Input the training samples into the initial scenario recognition model, obtain the first loss function of the initial first sub-model, and obtain the second loss function of the initial second sub-model;
[0085] Repeat the above steps until the first loss function and the second loss function converge to obtain the trained scenario recognition model.
[0086] Among them, the training samples include at least one first scenario image with a score and at least one second scenario image with a significant region. The original image corresponding to the first scenario image is the same as the original image corresponding to the second scenario image. The score of the first scenario image and the significant region of the second scenario image are determined by the third sub-model.
[0087] It can be understood that the first scene image and the second scene image can be the same original image with different labels. According to their respective label properties, the first scene image can be used to train the initial second sub-model, and the second scene image can be used to train the initial first sub-model. When the first loss function and the second loss function in the training process converge, it means that the training of the initial first sub-model and the initial second sub-model is completed. It should be noted that the labels of the first scene image and the second scene image can be generated by the output of the third sub-model or obtained by manual annotation, which is not limited in this embodiment.
[0088] Optionally, the step of "obtaining the first loss function of the initial first sub-model and obtaining the second loss function of the initial second sub-model" may include:
[0089] Input the first scene image into the initial second sub-model to obtain the predicted attribution map of the first scene image;
[0090] According to the predicted attribution map and the true attribution map of the first scene image, obtain the second loss function of the initial second sub-model;
[0091] Input the second scene image into the initial first sub-model to obtain the predicted saliency image of the second scene image;
[0092] According to the predicted saliency image and the true saliency image of the first scene image, obtain the first loss function of the initial first sub-model.
[0093] It can be understood that the second loss function is used to characterize the gap between the predicted attribution map and the true attribution map, and the first loss function is used to characterize the gap between the predicted saliency image and the true saliency image. Based on this, the performance of the detection model is detected, and the model is iteratively trained until the performance of the two models reaches the best.
[0094] Among them, the second loss function may include a region loss function and an area loss function. The region loss function is used to characterize the accuracy of the positions of the pixels included in the contribution degree probability distribution, and the area loss function is used to characterize the accuracy of the areas of the pixels included in the contribution degree probability distribution.
[0095] It can be understood that the attribution map is a heat map representing the degree of influence of each pixel point in the scene image on the scene image. That is, the heat on the attribution map can be understood as the contribution degree of each pixel of the input scene image to the realization of the model prediction process.
[0096] Among them, the second loss function may include a region loss function and an area loss function. The region loss function is used to characterize the accuracy of the positions of the pixels included in the contribution degree probability distribution, and the area loss function is used to characterize the accuracy of the areas of the pixels included in the contribution degree probability distribution. It can be understood that the second loss function for the second sub-model can be constructed from the above two aspects.
[0097] Optionally, the step of "obtaining the second loss function of the initial second sub-model" may include:
[0098] Construct the region loss function according to the first prediction value in combination with the EM algorithm; the first prediction value is the initial second sub-model's prediction of the first scene image with known classification, obtaining the first probability distribution of the positions of the pixels included in the first scene image;
[0099] Construct the area loss function according to the second prediction value; the second prediction value is the sum value of all the probabilities included in the first probability distribution;
[0100] Obtain the sum value of the region loss function and the area loss function, and take the sum value as the value of the second loss function.
[0101] Specifically, the latent variable used in the EM algorithm can be determined first. The position Z of each pixel in the input scene image is used as this dependent variable, and the EM algorithm is used to make the scene recognition model learn the conditional probability distribution of Z, that is, p(zly), on the premise of learning the known classification result Y. Among them, the classification result can be the category of the scoring attribution map.
[0102] Specifically, for an image X with a size of H×W, assume the number of categories of this image X is C, and it satisfies Z∈{1,2,…,HW}. Among them, H is the height of each pixel in the image X, and W is the width of each pixel.
[0103] Pass the image X through two convolutional layers respectively to change the number of channels of the feature map extracted by the classification network to C and 1. The feature map with a size of 1×H×W can be activated. For example, through L1 normalization activation, the activation value can be obtained as:
[0104] h z = p(z|x)
[0105] The feature map with a size of C×H×W can be activated along the channels. For example, through the softmax function activation, the activation value can be obtained as:
[0106] g yz = p(y|x,z)
[0107] h z= h after being processed by the broadcast mechanism of p(z|x) z and g yz Element-wise multiplication is performed to obtain the joint probability distribution p1 of predicting the class y and the position z respectively based on the input scene image x, and the probability distribution of deriving the correct attribution map from the initial second sub-model output according to this joint probability distribution, that is, the first prediction value p2:
[0108] p1 = p(y, z|x),
[0109] wherein, is the true class corresponding to the input scene image x.
[0110] Specifically, assume that p θ′ (z|x, y) is the probability distribution corresponding to each pixel position z under the condition that the classification result of the known input scene image x is y. Then, according to the EM algorithm, the regional loss function is obtained as:
[0111]
[0112] where θ is the parameter of the initial second sub-model, and θ′ is the distribution parameter of z.
[0113] Processing the formula L according to Bayes' formula EM gives:
[0114]
[0115] where l is the position of a certain pixel in the position set Z.
[0116] Therefore, the final regional loss function is obtained as:
[0117]
[0118] It can be understood that since the regional loss function can be used to characterize the accuracy of the positions of the pixels included in the contribution degree probability distribution, and the area loss function can be used to characterize the accuracy of the areas of the pixels included in the contribution degree probability distribution.
[0119] Therefore, on the basis of the regional loss function, the area loss function is further added, and the convergence of both loss functions is used as the loss function convergence condition of the initial second sub-model. It is possible to further control the accuracy of the areas formed by the above pixels while paying attention to the accuracy of the positions of the pixels in the attribution map, so that the positions of the pixels and the areas formed by the pixels in the heat map shown in the attribution map are more accurate, and the focus of the high-heat regions in the attribution map is stronger.
[0120] Among them, the second prediction value L can be used according toarea Construct an area loss function; the second prediction value L area can be the sum value of all probabilities included in the first probability distribution, and the second loss function can be the sum value of the regional loss function and the area loss function, which can be specifically expressed as:
[0121] L = L EM + λL area
[0122] where λ represents a preset hyperparameter used to control the proportion of the area loss function L area in the second loss function L.
[0123] In some embodiments, the first loss function of the initial first sub-model can be obtained according to the gap between the predicted saliency image and the true saliency image of the first scene image. For example, as the initial first sub-model is iteratively trained, for any stage in the training phase, the similarity distance between the above two images can be obtained as the first loss function. When the similarity distance between the two reaches the minimum or zero, it indicates that the first loss function converges, and the training of the initial first sub-model is completed.
[0124] 130. Obtain a first image through the output of the first sub-model, where the first image is used to highlight the recognition target in the scene image.
[0125] Among them, the first image can be used to highlight the recognition target in the scene image. Among them, the recognition target can refer to the target recognition object. Taking the surveillance image as an example, the target recognition object can be an object such as a strong background, a road, a door, etc. used to characterize the surveillance image information. In the first image, the position where the recognition target is located can be effectively displayed. Taking the recognition of the road in the surveillance image as an example, the saliency region can highlight the road part, while the brightness of the non-road part in the surveillance image remains unchanged or decreases, so as to clearly identify the corresponding contour and range of the road through the first image.
[0126] 140. Obtain a second image through the output of the second sub-model, where the second image is used to represent the probability distribution of the contribution degree of each pixel in the scene image to the realization of the prediction process, and the prediction process is the process of obtaining the second image from the scene image.
[0127] It can be understood that since the second sub-model in the embodiments of the present application can be the F-CALM model improved from the CAM model, the difference between the two is that the F-CALM model corresponding to the second sub-model introduces the EM algorithm in the process of constructing the loss function on the basis of the CAM model, so as to give a deeper explanation of the contribution degree distribution of each pixel in the image by means of maximum likelihood estimation. Therefore, the second image has a deeper interpretability compared with the traditional CAM map. It can not only reflect the contribution degree distribution of each pixel in the image to the display heat of the second image through the heat of each pixel in the heat map, but also has stronger focusing, making the heat areas of different depth regions in the second image more accurate.
[0128] 150. Fuse the first image and the second image to obtain a fused feature map.
[0129] 160. Input the fused feature map into the third sub-model to obtain a classification score map, where the classification score map has a score and a category. The score is used to represent the matching degree between the scene image and the pre-stored scene map, and the category is used to represent the recognition result of the scene image.
[0130] Among them, the classification score map has a score and a category. The score is used to represent the matching degree between the scene image and the pre-stored scene map, and the category is used to represent the recognition result of the scene image.
[0131] Among them, the matching degree refers to the comparison between the scene image and the pre-stored scene map. The matching degree can be presented by a percentage. The higher the percentage value, the more similar the scene image and the pre-stored scene map are.
[0132] Before step 150, the method of the embodiments of the present application may further include:
[0133] Calculate a first weight and a second weight;
[0134] The step of "fusing the first image and the second image to obtain a fused feature map" includes:
[0135] Fuse the first image and the second image according to the first weight and the second weight to obtain a fused feature map;
[0136] Input the fused feature map into the third sub-model.
[0137] Optionally, the step of "calculating the first weight and the second weight" may further include:
[0138] Obtain a preset parameter and the score of a third scene image; the third scene image is any image frame in the scene image;
[0139] Obtain the minimum value of the product of the preset parameter and the score of the third scenario image, and use the minimum value as the first weight for fusing the first image;
[0140] Obtain the difference between the preset value and the first weight, and use the difference as the second weight for fusing the second image.
[0141] Before the scenario image enters the third sub-model, it is necessary to fuse the first image and the second image, perform pooling processing, and then input them into the third sub-model. To ensure the rationality of the fused feature map as the fused feature, it is necessary to preset the weights for fusing the two images, including the first weight and the second weight. Optionally, in the embodiments of the present application, the weights of the top-down attention and bottom-up attention of the scenario image are controlled by the scores of the score category map. Among them, the first sub-model can output the first image through top-down attention, and the second sub-model can output the second image through the bottom-up attention mechanism.
[0142] Among them, the first weight can be set as where m is a preset hyperparameter, represents the score output by the third model under the maximum reward. It can be understood that since the first weight and the second weight are preset, the two weight values can be adjusted and optimized during the training process to obtain the optimal solutions of the first weight and the second weight.
[0143] It can be understood that for the fused image, the sum of the first weight and the second weight is 1. Therefore, the second weight is 1 - ρ, and the mapping value of the fused feature map can be expressed as:
[0144] S = (1 - ρ)S bu + ρS td
[0145] where bu represents bottom-up attention and td represents top-down attention.
[0146] As can be seen from the above, through the embodiments of the present application, accurate recognition of the scenario image is achieved through the scenario recognition model, and it is judged whether the scenario image matches the preset scenario image, so as to accurately judge whether the scenario image is a real image, and thus determine whether scenario fraud occurs.
[0147] To better implement the above method, an embodiment of the present application further provides a scene recognition device, which can be specifically integrated in an electronic device, and the electronic device can be devices such as a terminal, a server, etc. Among them, the terminal can be devices such as a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers.
[0148] For example, as Figure 3 shown, the scene recognition device may include:
[0149] A scene image acquisition unit 301, configured to acquire a scene image, where the scene image includes multiple image frames within a period of time;
[0150] An input unit 302, configured to input the scene image into the trained scene recognition model;
[0151] A first sub-model unit 303, configured to output a first image through the first sub-model, where the first image is used to highlight the recognition target in the scene image;
[0152] A second sub-model unit 304, configured to output a second image through the second sub-model, where the second image is used to represent the probability distribution of the contribution degree of each pixel in the scene image to the prediction process, and the prediction process is the process of obtaining the second image from the scene image;
[0153] A fusion unit 305, configured to fuse the first image and the second image to obtain a fusion feature map;
[0154] A third sub-model unit 306, configured to input the fusion feature map into the third sub-model to obtain a score classification map, where the score classification map has a score and a category, the score is used to represent the matching degree between the scene image and the pre-stored scene map, and the category is used to represent the recognition result of the scene image.
[0155] Specifically in implementation, the above units can be implemented as independent entities, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.
[0156] The embodiment of the present application provides a scene recognition device. Based on the foregoing method embodiment, this device not only solves the problem of time-consuming and laborious data processing in the implementation process of the method embodiment, but also solves the error problem existing in manual operations in the method embodiment through the setting of functional units. Through the systematic operation of functional units, while ensuring the accuracy and refinement of the implementation process, it saves the time of manual operations and improves the implementation efficiency of the technical solution protected by the present application.
[0157] The embodiment of the present application also provides a device, which can be a device such as a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.
[0158] In some embodiments, the scene recognition device can also be integrated in multiple devices. For example, the scene recognition device can be integrated in multiple servers, and the scene recognition method of the present application is implemented by multiple servers.
[0159] For example, as Figure 4 shown, it shows a schematic structural diagram of the device involved in the embodiment of the present application. Specifically:
[0160] The device may include a processor 401 with one or more processing cores, a memory 402 with one or more storage media, a power supply 403, an input module 404, a communication module 405 and other components. Those skilled in the art can understand that the device structure shown in the figure does not constitute a limitation on the device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:
[0161] The processor 401 is the control center of the device, connecting various parts of the entire device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, it executes various functions of the device and processes data. In some embodiments, the processor 401 may include one or more processing cores; in some embodiments, the processor 401 may integrate an application processor and a modulation and demodulation processor. Among them, the application processor mainly processes the operating system, user interface and application programs, etc., and the modulation and demodulation processor mainly processes wireless communication. It can be understood that the above modulation and demodulation processor may not be integrated into the processor 401.
[0162] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store data created according to the use of the device. In addition, the memory 402 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.
[0163] The device also includes a power supply 403 for powering each component. In some embodiments, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0164] The device may also include an input module 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0165] The device may also include a communication module 405. In some embodiments, the communication module 405 can include a wireless module. The device can perform short-distance wireless transmission through the wireless module of the communication module 405, thereby providing users with wireless broadband Internet access. For example, the communication module 405 can be used to help users send and receive emails, browse web pages, and access streaming media, etc.
[0166] Although not shown, the device may also include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to achieve various functions as follows:
[0167] Obtain a scene image, where the scene image includes multiple image frames within a period of time;
[0168] Input the scene image into the trained scene recognition model;
[0169] The first image is obtained by outputting through the first sub-model, where the first image is used to highlight the recognition target in the scene image;
[0170] The second image is obtained by outputting through the second sub-model, where the second image is used to represent the probability distribution of the contribution degree of each pixel in the scene image to the prediction process, and the prediction process is the process of obtaining the second image from the scene image;
[0171] The first image and the second image are fused to obtain a fused feature map;
[0172] The fused feature map is input into the third sub-model to obtain a scoring category map, where the scoring category map has a score and a category. The score is used to represent the matching degree between the scene image and the pre-stored scene map, and the category is used to represent the recognition result of the scene image.
[0173] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.
[0174] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. These instructions can be stored in a storage medium and loaded and executed by a processor.
[0175] Therefore, an embodiment of the present application provides a storage medium, in which multiple instructions are stored. These instructions can be loaded by a processor to execute the steps in any of the scene recognition methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:
[0176] Obtain a scene image, where the scene image includes multiple image frames within a period of time;
[0177] Input the scene image into the trained scene recognition model;
[0178] The first image is obtained by outputting through the first sub-model, where the first image is used to highlight the recognition target in the scene image;
[0179] The second image is obtained by outputting through the second sub-model, where the second image is used to represent the probability distribution of the contribution degree of each pixel in the scene image to the prediction process, and the prediction process is the process of obtaining the second image from the scene image;
[0180] The first image and the second image are fused to obtain a fused feature map;
[0181] Input the fused feature map into the third sub-model to obtain a classification score map, where the classification score map has a score and a category. The score is used to represent the matching degree between the scene image and the pre-stored scene map, and the category is used to represent the recognition result of the scene image.
[0182] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated herein.
[0183] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0184] Since the instructions stored in the storage medium can execute the steps in any of the scene recognition methods provided in the embodiments of the present application, the beneficial effects achievable by any of the scene recognition methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated herein.
[0185] The above has introduced in detail a scene recognition method, device, equipment and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A scene recognition method, characterized in that, Applied to a scene recognition model, the scene recognition model includes a first sub-model, a second sub-model, and a third sub-model, and the method includes: Training an initial first sub-model and an initial second sub-model to obtain a trained first sub-model and a second sub-model, including: obtaining training samples, where the training samples include at least one first scene image with a score and at least one second scene image with a saliency region, the original image corresponding to the first scene image is the same as the original image corresponding to the second scene image, and the score of the first scene image and the saliency region of the second scene image are determined by the third sub-model; inputting the training samples into the initial scene recognition model to obtain a first loss function of the initial first sub-model and a second loss function of the initial second sub-model; repeating the above steps until the first loss function and the second loss function converge to obtain the trained scene recognition model; Using the trained first sub-model and the second sub-model to obtain the trained scene recognition model; Obtaining a scene image, where the scene image includes multiple image frames within a period of time; Inputting the scene image into the trained scene recognition model; Outputting a first image through the first sub-model, where the first image is used to highlight the recognition target in the scene image; Outputting a second image through the second sub-model, where the second image is used to represent the probability distribution of the contribution degree of each pixel in the scene image to the realization of the prediction process, and the prediction process is the process of obtaining the second image from the scene image; Fusing the first image and the second image to obtain a fused feature map; Inputting the fused feature map into the third sub-model to obtain a score classification map, where the score classification map has a score and a category, the score is used to represent the matching degree between the scene image and a pre-stored scene map, and the category is used to represent the recognition result of the scene image; Wherein, the second loss function includes a region loss function and an area loss function, the region loss function is used to represent the accuracy of the positions of the pixels included in the contribution degree probability distribution, and the area loss function is used to represent the accuracy of the areas of the pixels included in the contribution degree probability distribution; the second loss is obtained in the following manner: constructing the region loss function according to a first prediction value in combination with the EM algorithm; the first prediction value is the initial second sub-model predicting the first scene image with a known classification to obtain a first probability distribution of the positions of the pixels included in the first scene image; constructing the area loss function according to a second prediction value, and the area loss function is the sum of all probabilities included in the first probability distribution; calculating the sum of the region loss function and the area loss function, and the sum value is the value of the second loss function.
2. The method according to claim 1, characterized in that, The obtaining of the first loss function of the initial first sub-model and the obtaining of the second loss function of the initial second sub-model include: Input the first scene image into the initial second sub-model to obtain the predicted attribution map of the first scene image; Obtain the second loss function according to the predicted attribution map and the true attribution map of the first scene image; Input the second scene image into the initial first sub-model to obtain the predicted saliency image of the second scene image; Obtain the first loss function of the initial first sub-model according to the predicted saliency image and the true saliency image of the first scene image.
3. The method according to claim 1, characterized in that, Before fusing the first image and the second image to obtain a fused feature map, the method further includes: Calculate a first weight and a second weight; The step of fusing the first image and the second image to obtain a fused feature map includes: Fuse the first image and the second image according to the first weight and the second weight to obtain a fused feature map.
4. The method according to claim 3, characterized in that The step of calculating the first weight and the second weight includes: Obtain a preset parameter and the score of a third scene image; the third scene image is any one image frame in the scene images; Obtain the minimum value of the product of the preset parameter and the score of the third scene image, and use the minimum value as the first weight for fusing the first image; Obtain the difference between a preset value and the first weight, and use the difference as the second weight for fusing the second image.
5. A scene recognition device, characterized in that, For implementing the method according to any one of claims 1 to 4, the apparatus includes: A scene image acquisition unit, configured to acquire scene images, where the scene images include multiple image frames within a period of time; An input unit, configured to input the scene images into the trained scene recognition model; A first sub-model unit, configured to output a first image through the first sub-model, where the first image is used to highlight the recognition target in the scene images; A second sub-model unit, configured to output a second image through the second sub-model, where the second image is used to represent the probability distribution of the contribution degree of each pixel in the scene images to the realization of the prediction process, and the prediction process is the process of obtaining the second image from the scene images; A fusion unit, configured to fuse the first image and the second image to obtain a fused feature map; A third sub-model unit, configured to input the fused feature map into the third sub-model to obtain a score classification map, where the score classification map has a score and a category, the score is used to represent the matching degree between the scene images and the pre-stored scene maps, and the category is used to represent the recognition result of the scene images.
6. A device, characterized in that, It includes a processor and a memory, and the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps in a scene recognition method according to any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in a scene recognition method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Character recognition method and device, computer readable medium and electronic equipment
CN111062389A
Image extraction method and device and storage medium
JP2001043376A