An artificial intelligence-based entity recognition method, related device, and storage medium

Through the subject recognition method based on artificial intelligence, the candidate area subject scores in the image are automatically identified and calculated, which solves the problem of manual operation dependence in the prior art, and realizes efficient and accurate image processing and cropping.

CN113723168BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110383066.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-09
Publication Date
2025-06-27
Estimated Expiration
2041-04-09

AI Technical Summary

Technical Problem

The prior art relies on manual operations in the process of image cutting, resulting in high time and labor costs, and the accuracy of image content diversity and scene complexity is difficult to guarantee, and cannot meet the needs of scale.

Method used

Using the subject recognition method based on artificial intelligence, the subject recognition and feature extraction are performed by obtaining the image to be identified, the subject scores of the candidate regions are calculated, and the target candidate regions and subjects are automatically determined.

Benefits of technology

It realizes automatic image processing, reduces time and labor costs, improves the accuracy of subject recognition, and can meet the needs of scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113723168B_ABST
    Figure CN113723168B_ABST
Patent Text Reader

Abstract

The present application discloses a method for entity recognition based on artificial intelligence, which can be used in the field of cloud security. The present application includes: obtaining an image to be recognized; performing region recognition processing on the image to be recognized to obtain N candidate regions, and performing feature extraction processing on the image to be recognized to obtain a target feature map; obtaining the entity score corresponding to each of the N candidate regions according to the target feature map and the N candidate regions; determining a target candidate region from the N candidate regions according to the entity score corresponding to each candidate region, and taking the candidate entity corresponding to the target candidate region as the target entity in the image to be recognized, where the target candidate region corresponds to the maximum entity score. The present application also provides related devices and storage media. The present application can achieve the purpose of automatically processing images, not only reducing the time cost and labor cost, but also the entity selected based on the entity score has high accuracy and can meet the large-scale requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology, and in particular, to a method for identifying a subject based on artificial intelligence, a related device, and a storage medium. Background Art

[0002] With the development of the information age, more and more people enjoy the convenience brought by the multimedia information age. For some multimedia information, secondary processing is also required. Among them, in the field of image processing, image cropping is an essential link, and image cropping usually refers to cropping out important regions from an image.

[0003] Considering the diversity of image sizes, for example, some images are 1:1 in size, some images are 4:3 in size, and some images are 3:4 in size. Currently, if it is necessary to crop out important regions from these images, it mainly depends on manual operation, that is, relevant personnel crop the images through professional software.

[0004] However, due to the diversity and scene complexity of image content, in the case of a large amount of image processing, the manual processing method not only consumes a large amount of time cost and labor cost, but also is prone to the situation of improper selection of the image subject, and it is difficult to meet the large-scale requirements. Summary of the Invention

[0005] Embodiments of this application provide a method for identifying a subject based on artificial intelligence, a related device, and a storage medium, which can achieve the purpose of automatically processing images, not only reducing the time cost and labor cost, but also having high accuracy for the subject selected based on the subject score, and can meet the large-scale requirements.

[0006] In view of this, on the one hand, this application provides a method for identifying a subject based on artificial intelligence, including:

[0007] Obtain an image to be recognized;

[0008] Perform region recognition processing on the image to be recognized to obtain N candidate regions, and perform feature extraction processing on the image to be recognized to obtain a target feature map, where each candidate region corresponds to a candidate subject, and the target feature map is obtained after splicing at least two feature maps, and N is an integer greater than or equal to 1;

[0009] According to the target feature map and the N candidate regions, obtain the subject score corresponding to each candidate region among the N candidate regions;

[0010] Determine a target candidate region from N candidate regions according to the subject score corresponding to each candidate region, and use the candidate subject corresponding to the target candidate region as the target subject in the image to be recognized, where the target candidate region corresponds to the largest subject score.

[0011] On the other hand, this application provides a subject recognition device, including:

[0012] An acquisition module, configured to acquire an image to be recognized;

[0013] A processing module, configured to perform region recognition processing on the image to be recognized to obtain N candidate regions, and perform feature extraction processing on the image to be recognized to obtain a target feature map, where each candidate region corresponds to a candidate subject, and the target feature map is obtained after splicing at least two feature maps, and N is an integer greater than or equal to 1;

[0014] The acquisition module is further configured to obtain the subject score corresponding to each candidate region in the N candidate regions according to the target feature map and the N candidate regions;

[0015] A determination module, configured to determine a target candidate region from the N candidate regions according to the subject score corresponding to each candidate region, and use the candidate subject corresponding to the target candidate region as the target subject in the image to be recognized, where the target candidate region corresponds to the largest subject score.

[0016] In a possible design, in the first implementation manner of the other aspect of the embodiment of this application,

[0017] The processing module is specifically configured to, based on the image to be recognized, obtain N candidate regions through a face detection network, where the candidate subject corresponding to each candidate region is a face;

[0018] Or,

[0019] Based on the image to be recognized, obtain N candidate regions through a human body detection network, where the candidate subject corresponding to each candidate region is a human body.

[0020] In a possible design, in the first implementation manner of the other aspect of the embodiment of this application,

[0021] The processing module is specifically configured to, based on the image to be recognized, obtain a saliency feature map through a first network included in the feature extraction network, where the saliency feature map corresponds to 1 channel;

[0022] Based on the image to be recognized, obtain a depth semantic embedding map through a second network included in the feature extraction network, where the depth semantic embedding map corresponds to C channels, and C is an integer greater than 1;

[0023] Perform splicing processing on the saliency feature map and the deep semantic embedding map to obtain a target feature map, where the target feature map includes (C + 1) channels.

[0024] In a possible design, in the first implementation manner of another aspect of the embodiments of the present application,

[0025] The processing module is specifically configured to, based on the image to be recognized, obtain a saliency feature map through a first network included in the feature extraction network, where the saliency feature map corresponds to 1 channel;

[0026] Based on the image to be recognized, obtain a deep semantic embedding map through a second network included in the feature extraction network, where the deep semantic embedding map corresponds to C channels, and C is an integer greater than 1;

[0027] Based on the image to be recognized, obtain a blurriness feature map through a third network included in the feature extraction network, where the blurriness feature map corresponds to 1 channel;

[0028] Perform splicing processing on the saliency feature map, the deep semantic embedding map, and the blurriness feature map to obtain a target feature map, where the target feature map includes (C + 2) channels.

[0029] In a possible design, in the first implementation manner of another aspect of the embodiments of the present application,

[0030] The obtaining module is specifically configured to perform matching processing on each of the N candidate regions with the target feature map to obtain N spatial feature maps, where there is a one-to-one correspondence between the spatial feature map and the candidate region;

[0031] For each of the N candidate regions, based on the spatial feature map corresponding to each candidate region, obtain a first image feature through a first convolutional network included in the main body selection network;

[0032] For each of the N candidate regions, based on the spatial feature map corresponding to each candidate region, obtain a second image feature through a second convolutional network included in the main body selection network;

[0033] For each of the N candidate regions, perform splicing processing on the first image feature and the second image feature corresponding to each candidate region to obtain a comprehensive image feature corresponding to each candidate region;

[0034] Based on the comprehensive image feature corresponding to each candidate region, obtain a main body score corresponding to each candidate region through a fully connected layer included in the main body selection network.

[0035] In a possible design, in the first implementation of another aspect of the embodiments of the present application, the subject recognition device further includes an update module;

[0036] The acquisition module is further configured to acquire a training image sample, where the training image sample includes a first annotation area and a second annotation area, the first annotation area corresponds to a first annotation score, and the second annotation area corresponds to a second annotation score;

[0037] The processing module is further configured to perform feature extraction processing on the training image sample to obtain a target feature map of the training image sample;

[0038] The processing module is further configured to perform matching processing between the first annotation area and the target feature map to obtain a first spatial feature map, and perform matching processing between the second annotation area and the target feature map to obtain a second spatial feature map;

[0039] The acquisition module is further configured to obtain a first predicted image feature through a first convolutional network included in the training subject selection network based on the first spatial feature map, and obtain a second predicted image feature through a second convolutional network included in the training subject selection network based on the second spatial feature map;

[0040] The acquisition module is further configured to obtain a third predicted image feature through the second convolutional network included in the training subject selection network based on the first annotation area, and obtain a fourth predicted image feature through the second convolutional network included in the training subject selection network based on the second annotation area;

[0041] The processing module is further configured to perform splicing processing on the first predicted image feature and the third predicted image feature corresponding to the first annotation area to obtain a first comprehensive image feature corresponding to the first annotation area, and perform splicing processing on the second predicted image feature and the fourth predicted image feature corresponding to the second annotation area to obtain a second comprehensive image feature corresponding to the second annotation area;

[0042] The acquisition module is further configured to obtain a first predicted subject score corresponding to the first annotation area through a fully connected layer included in the training subject selection network based on the first comprehensive image feature, and obtain a second predicted subject score corresponding to the second annotation area through a fully connected layer included in the training subject selection network based on the second comprehensive image feature;

[0043] The update module is configured to update the model parameters of the training subject selection network according to the first predicted subject score, the second predicted subject score, the first annotation score, and the second annotation score until the model training condition is met, and output the subject selection network.

[0044] In a possible design, in the first implementation of another aspect of the embodiments of the present application,

[0045] An acquisition module, specifically configured to, for each of the N candidate regions, crop the feature map corresponding to each candidate region from the target feature map;

[0046] Perform global average pooling on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1;

[0047] Determine the main score corresponding to each candidate region according to the region feature corresponding to each candidate region.

[0048] In a possible design, in the first implementation of another aspect of the embodiments of the present application,

[0049] An acquisition module, specifically configured to determine the feature average value corresponding to each candidate region according to the region feature corresponding to each candidate region, where the feature average value is the average value of the M feature values in the region feature;

[0050] For each of the N candidate regions, use the feature average value of the candidate region as the main score of the candidate region;

[0051] Or,

[0052] An acquisition module, specifically configured to determine the first feature average value corresponding to each candidate region according to the first region feature and the first region weight corresponding to each candidate region;

[0053] Determine the second feature average value corresponding to each candidate region according to the second region feature and the second region weight corresponding to each candidate region, where the second region weight is greater than the first region weight;

[0054] Determine the target average value corresponding to each candidate region according to the first feature average value and the second feature average value corresponding to each candidate region;

[0055] For each of the N candidate regions, use the target average value of the candidate region as the main score of the candidate region.

[0056] In a possible design, in the first implementation of another aspect of the embodiments of the present application,

[0057] An acquisition module, specifically configured to, for each of the N candidate regions, crop the feature map corresponding to each candidate region from the target feature map;

[0058] Perform global average pooling on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1;

[0059] Based on the region feature corresponding to each candidate region, obtain the main body score corresponding to each candidate region through a multi-layer perceptron.

[0060] In a possible design, in the first implementation manner of another aspect of the embodiments of the present application,

[0061] The acquisition module is further configured to acquire a training image sample, where the training image sample includes a first labeled region and a second labeled region, the first labeled region corresponds to a first labeled score, and the second labeled region corresponds to a second labeled score;

[0062] The processing module is further configured to perform feature extraction processing on the training image sample to obtain the target feature map of the training image sample;

[0063] The processing module is further configured to crop from the target feature map of the training image sample the feature map corresponding to the first labeled region and the feature map corresponding to the second labeled region;

[0064] The processing module is further configured to perform global average pooling on the feature map corresponding to the first labeled region to obtain the region feature corresponding to the first labeled region;

[0065] The processing module is further configured to perform global average pooling on the feature map corresponding to the second labeled region to obtain the region feature corresponding to the second labeled region;

[0066] The acquisition module is further configured to, based on the region feature corresponding to the first labeled region, obtain the first predicted main body score corresponding to the first labeled region through the training multi-layer perceptron;

[0067] The acquisition module is further configured to, based on the region feature corresponding to the second labeled region, obtain the second predicted main body score corresponding to the second labeled region through the training multi-layer perceptron;

[0068] The update module is further configured to update the model parameters of the training multi-layer perceptron according to the first predicted main body score, the second predicted main body score, the first labeled score, and the second labeled score until the model training condition is satisfied, and output the multi-layer perceptron.

[0069] In a possible design, in the first implementation manner of another aspect of the embodiments of the present application,

[0070] The acquisition module is specifically configured to perform frame splitting on the video to be processed to obtain K video frames, where K is an integer greater than 1;

[0071] For each of the K video frames, obtain the brightness, sharpness, and color uniformity.

[0072] Perform frame filtering on the K video frames according to the brightness and brightness threshold of each video frame, the sharpness and sharpness threshold of each video frame, and the color uniformity and color uniformity threshold of each video frame, to obtain L video frames, where L is an integer greater than or equal to 1 and less than K.

[0073] Select one video frame from the L video frames as the image to be recognized.

[0074] The obtaining module is further configured to determine a target candidate region from the N candidate regions according to the main body score corresponding to each candidate region, and use the candidate main body corresponding to the target candidate region as the target main body in the image to be recognized. Then, according to the target candidate region, crop the target image from the image to be recognized, where the target image corresponds to the target size and the target image includes the target candidate region.

[0075] The determining module is further configured to use the target image as the video cover of the video to be processed.

[0076] In a possible design, in the first implementation manner of the other aspect of the embodiments of the present application,

[0077] The obtaining module is specifically configured to determine a secondary candidate region from the N candidate regions according to the main body score corresponding to each candidate region, where the secondary candidate region corresponds to the second largest main body score.

[0078] If the main body score corresponding to the secondary candidate region is greater than or equal to the main body score threshold, obtain the first position coordinates of the center of the secondary candidate region in the image to be recognized, and obtain the second position coordinates of the center of the target candidate region in the image to be recognized.

[0079] Determine the region distance according to the first position coordinates and the second position coordinates.

[0080] If the region distance is less than or equal to the distance threshold corresponding to the target size, crop the target image from the image to be recognized according to the target candidate region and the secondary candidate region, where the target image further includes the secondary candidate region.

[0081] Another aspect of the present application provides a computer device, including: a memory, a processor, and a bus system;

[0082] Wherein, the memory is used to store programs;

[0083] The processor is used to execute a program in a memory, and the processor is used to execute the methods of the above aspects according to the instructions in the program code;

[0084] The bus system is used to connect the memory and the processor so that the memory and the processor can communicate.

[0085] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the methods of the above aspects.

[0086] Another aspect of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.

[0087] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:

[0088] In the embodiments of the present application, a method for identifying a subject based on artificial intelligence is provided. First, an image to be identified is obtained, and then the image to be identified is subjected to region recognition processing to obtain N candidate regions, and each candidate region corresponds to a candidate subject. In addition, it is also necessary to perform feature extraction processing on the image to be identified to obtain a target feature map, and then according to the target feature map and the N candidate regions, obtain the subject score corresponding to each candidate region in the N candidate regions. Finally, according to the subject score corresponding to each candidate region, determine the target candidate region from the N candidate regions, and use the candidate subject corresponding to the target candidate region as the target subject in the image to be identified. The target candidate region corresponds to the largest subject score. Through the above method, several candidate regions are first identified, and then the subject score of each candidate region is automatically calculated. Finally, the candidate region with the highest subject score is used as the target candidate region, that is, the candidate subject in the target candidate region is the subject in the image, and the candidate subjects in other candidate regions are the secondary subjects in the image. Thus, the purpose of automatically processing the image is achieved, which not only reduces the time cost and labor cost, but also the subject selected based on the subject score has high accuracy and can meet the large-scale requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 It is a schematic diagram of an architecture of a subject recognition system in an embodiment of the present application;

[0090] Figure 2 It is a schematic diagram of a comparison of the image screen sizes in an embodiment of the present application;

[0091] Figure 3 Schematic diagram of an embodiment of the main body recognition method in the embodiment of the present application;

[0092] Figure 4 Schematic diagram of recognizing a face region from an image to be recognized in the embodiment of the present application;

[0093] Figure 5 Schematic diagram of recognizing a human body region from an image to be recognized in the embodiment of the present application;

[0094] Figure 6 Schematic diagram of generating a target feature map based on two types of feature maps in the embodiment of the present application;

[0095] Figure 7 Another schematic diagram of generating a target feature map based on three types of feature maps in the embodiment of the present application;

[0096] Figure 8 Schematic diagram of a prediction process of the ranking main body selection model in the embodiment of the present application;

[0097] Figure 9 Schematic diagram of a training process of the ranking main body selection model in the embodiment of the present application;

[0098] Figure 10 Schematic diagram of a process for implementing intelligent image clipping in the embodiment of the present application;

[0099] Figure 11 Schematic diagram of the effect comparison obtained by sample analysis based on the main body recognition method in the embodiment of the present application;

[0100] Figure 12 Schematic diagram of a main body recognition device in the embodiment of the present application;

[0101] Figure 13 Schematic diagram of a structure of a terminal device in the embodiment of the present application;

[0102] Figure 14 Schematic diagram of a structure of a server in the embodiment of the present application. Specific embodiments

[0103] The embodiment of the present application provides an artificial intelligence-based main body recognition method, related device and storage medium, which can achieve the purpose of automatically processing images, not only reducing the time cost and labor cost, but also the main body selected based on the main body score has high accuracy and can meet the large-scale requirements.

[0104] In the description, claims, and the above-mentioned drawings of this application, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0105] The image main body recognition can detect the coordinate position of the main body in the image, and can be used to further crop out the main body area, and cooperate with the image recognition interface to improve the recognition accuracy. It is widely applicable to scenarios such as intelligent cropping, intelligent beauty photo, and intelligent assisted image recognition. The method of main body recognition will be introduced below in combination with the above scenarios.

[0106] Scenario 1, intelligent cropping;

[0107] With the rise of the short video consumption scenario on mobile devices, more and more vertical videos have occupied a large amount of the user's video content consumption time. Video platforms are also striving to build vertical video communities. However, the supply of vertical videos is still scarce, and there is a lack of high-quality existing materials. The supply source of vertical videos is an important bottleneck in the vertical video scenario of video platforms. In the business scenario of long videos, a large amount of high-quality copyrighted content is also one of the core competitiveness of video platforms. Therefore, using existing high-quality resources to produce vertical video content has great commercial value. The main body recognition method provided in this application can identify the main body and secondary main body from a certain key frame in the video, and preferentially use the main body as the basis for generating the video cover image. Thus, it can automatically generate cover images for videos of different sizes, provide materials for video display, and meet differentiated needs.

[0108] Scenario 2, intelligent beauty photo;

[0109] Detect the main body according to the photo uploaded by the user to implement functions such as image cropping or background blurring, which can be applied to application programs with beauty photo functions.

[0110] Scenario 3, intelligent assisted image recognition;

[0111] The image main body detection can be used to crop out the image main body area, and cooperate with image recognition to improve the recognition accuracy.

[0112] In order to improve the accuracy of main body recognition in the above scenarios, this application proposes an artificial intelligence-based main body recognition method, which is applied to Figure 1The subject recognition system shown, as shown in the figure, includes a server and terminal devices. The server involved in this application can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a laptop computer, a palm computer, a personal computer, a smart TV, a smart watch, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. The number of servers and terminal devices is also not restricted. Based on Figure 1 The subject recognition system shown supports offline subject recognition and online subject recognition. The methods of offline subject recognition and online subject recognition will be introduced separately below.

[0113] I. Offline subject recognition;

[0114] The subject recognition system includes a terminal device. First, the terminal device acquires the image to be recognized, then performs region recognition processing on the image to be recognized to obtain N candidate regions. In addition, it is also necessary to perform feature extraction processing on the image to be recognized to obtain a target feature map. Then, based on the target feature map and the N candidate regions, the subject score corresponding to each candidate region in the N candidate regions is obtained. Finally, the candidate region corresponding to the maximum subject score is used as the target candidate region. Among them, the candidate subject within the target candidate region is the selected target subject. The terminal device highlights the target candidate region on the image to be recognized.

[0115] II. Online subject recognition;

[0116] The subject recognition system includes a terminal device and a server. First, the terminal device acquires the image to be recognized, and then sends the image to be recognized to the server. The server performs region recognition processing on the image to be recognized to obtain N candidate regions. In addition, the server also needs to perform feature extraction processing on the image to be recognized to obtain a target feature map. Then, the server obtains the subject score corresponding to each candidate region in the N candidate regions based on the target feature map and the N candidate regions. Finally, the server uses the candidate region corresponding to the maximum subject score as the target candidate region. Among them, the candidate subject within the target candidate region is the selected target subject. The server sends the subject recognition result to the terminal device, and the terminal device highlights the target candidate region on the image to be recognized according to the subject recognition result.

[0117] It can be seen that this application needs to perform region recognition processing, feature extraction processing, and main body score calculation on the image to be recognized. These processes involve computer vision (CV) technology and machine learning (ML) technology based on artificial intelligence (AI), etc. Among them, CV is a science that studies how to make machines "see". Further speaking, it refers to using cameras and computers to replace human eyes for object recognition, trace tracing, measurement, etc. of machine vision, and further performing graphic processing to make the computer-processed image more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, CV studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. CV technology usually includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0118] ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0119] In summary, AI is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, AI is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0120] AI technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. AI basic technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. AI software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0121] As can be seen from the foregoing introduction, the subject recognition method can be applied to a variety of scenarios. In addition, it can also perform subject recognition on image frames of different sizes, and can achieve good results whether in image display or video playback. For ease of understanding, please refer to Figure 2 , Figure 2 which is a comparison schematic diagram of the image frame size in the embodiments of the present application. As shown in Figure 2 Figure (A) therein, taking the vertical screen video playback or image display as an example, the aspect ratio of the terminal device is 16:9. Assuming that the subject is determined to be a black person through subject recognition, therefore, the black person is displayed on the terminal device. As shown in Figure 2 Figure (B) therein, taking the horizontal screen video playback or image display as an example, the aspect ratio of the terminal device is 9:16. Assuming that the subject is determined to be a black person through subject recognition, therefore, the black person is displayed on the terminal device.

[0122] Combined with the above introduction, the subject recognition method based on artificial intelligence in the present application will be introduced below. Please refer to Figure 3 One embodiment of the subject recognition method based on artificial intelligence in the embodiments of the present application includes:

[0123] 101. Obtain the image to be recognized;

[0124] In this embodiment, the subject recognition device obtains the image to be recognized. Among them, the image to be recognized can be a photo, a painting, or a frame in a video, and no limitation is made here.

[0125] It should be noted that the subject recognition device involved in the present application can be deployed on a terminal device, or can be deployed on a server, or can be deployed on a subject recognition system composed of a terminal device and a server, and no limitation is made here.

[0126] 102. Perform region recognition processing on the image to be recognized to obtain N candidate regions, and perform feature extraction processing on the image to be recognized to obtain a target feature map, where each candidate region corresponds to a candidate subject, and the target feature map is obtained after splicing at least two feature maps, and N is an integer greater than or equal to 1;

[0127] In this embodiment, the subject recognition device can perform region recognition processing on the image to be recognized, thereby recognizing N candidate regions in the image to be recognized. Each candidate region includes a candidate subject, and the candidate subject includes but is not limited to a human face, a human body, an animal (such as a cat or a dog, etc.), a building, a plant, and other objects (such as a cup, glasses, or a teapot, etc.). In addition, the subject recognition device also needs to extract the image features of the image to be recognized, that is, obtain the target feature map, where the target feature map is obtained after splicing at least two feature maps, and different feature maps can reflect different image features. For example, they can reflect the color features, texture features, shape features, and spatial features of the image, etc.

[0128] Specifically, in practical applications, a trained network model can be used to perform region recognition processing on the image to be recognized. It can be understood that the network model includes but is not limited to the regions with convolutional neural network (R-CNN), Fast R-CNN, Faster R-CNN, the you only look once (YOLO) model, and the single shot multibox detector (SSD), etc.

[0129] 103. Obtain the subject score corresponding to each candidate region among the N candidate regions according to the target feature map and the N candidate regions;

[0130] In this embodiment, it is assumed that the size of the image to be recognized is 512*512*3, and it is assumed that the size of the target feature map obtained after feature extraction is 16*16*512. Since the image size has changed, it is necessary to map the candidate regions to the target feature map. It can be understood that this application can use methods such as region of interest pooling (ROI Pooling), or region of interest pooling align (ROIAlign), or region of interest warping layer (ROI Warping Layer), etc. to achieve the mapping of the candidate regions.

[0131] Specifically, the corresponding main body score of each candidate region can be further determined according to its mapping result. For example, the eigenvalues of each mapping result are summed, and the sum result is used as the main body score. For another example, the eigenvalues of each candidate region are input into a trained deep neural network, and the deep neural network predicts the corresponding main body score. Other methods can also be used to determine the main body score corresponding to the candidate region.

[0132] 104. According to the main body score corresponding to each candidate region, a target candidate region is determined from the N candidate regions, and the candidate main body corresponding to the target candidate region is used as the target main body in the image to be recognized, where the target candidate region corresponds to the maximum main body score.

[0133] In this embodiment, the main body recognition device sorts the main body scores of the N candidate regions in descending order, and takes the candidate region corresponding to the maximum main body score as the target candidate region. The candidate main body included in the target candidate region is the target main body in the image to be recognized, and the candidate main bodies included in other candidate regions are all secondary main bodies in the image to be recognized.

[0134] In the embodiment of the present application, an artificial intelligence-based main body recognition method is provided. Through the above method, several candidate regions are first recognized, then the main body scores of each candidate region are automatically calculated, and finally the candidate region with the highest main body score is used as the target candidate region, that is, the candidate main body in the target candidate region is the main body in the image, and the candidate main bodies in other candidate regions are the secondary main bodies in the image. Thereby, the purpose of automatically processing images is achieved, which not only reduces the time cost and labor cost, but also the main body selected based on the main body score has high accuracy and can meet the large-scale requirements.

[0135] Optionally, on the basis of the corresponding embodiment above, in another optional embodiment provided by the embodiment of the present application, region recognition processing is performed on the image to be recognized to obtain N candidate regions, which may specifically include: Figure 3 Based on the image to be recognized, N candidate regions are obtained through a face detection network, where the candidate main body corresponding to each candidate region is a face;

[0136] Or,

[0137] Based on the image to be recognized, N candidate regions are obtained through a human body detection network, where the candidate main body corresponding to each candidate region is a human body.

[0138]

[0139] ​In this embodiment, a detection method is introduced that uses the area where a human face or body is located as a candidate area. As can be seen from the foregoing embodiments, taking the image cropping scenario as an example, during the cropping process from a landscape image to a portrait image, some picture information is inevitably lost. In order to minimize this loss, during the intelligent image cropping process, it is necessary to identify the main subject in the picture and use the main subject as the center of the cropping area, so as to produce a meaningful portrait picture with most of the picture information. Since people are the main content shown in most images (or videos), therefore, in this application, person localization is used as the main target in the task. It can be understood that in practical applications, the type of the main subject can also be flexibly adjusted according to the situation. This is only an example here and should not be construed as a limitation to this application.

[0140] Exemplarily, taking face detection as an example, for ease of understanding, please refer to Figure 4 , Figure 4 is a schematic diagram of identifying a face area from the image to be recognized in an embodiment of this application. As shown in the figure, the area circled by the black frame in the figure is the recognized face area. This application can use the open-source dual shot face detector (DSFD) as the face detection network to discover the face area in the image to be recognized. DSFD has a Feature Enhance Module (FEM), which can learn more effective content and semantic learning in terms of width and depth, so as to enhance the discriminability and robustness of features. DSFD also uses the Progressive AnchorLoss (PAL) to assist in feature learning, forming more effective supervision for the entire model during the training process. DSFD also uses an improved anchor matching strategy to make the anchor match the real face as much as possible.

[0141] Exemplarily, taking human body detection as an example, for ease of understanding, please refer to Figure 5 , Figure 5 is a schematic diagram of identifying a human body area from the image to be recognized in an embodiment of this application. As shown in the figure, the area circled by the black frame in the figure is the recognized human body area. This application can use Free Anchor, and by modifying the loss function to remove the process of manually specifying the anchor point, the network model can autonomously learn which anchor point to match with the real object.

[0142] In a portrait image, it is easier and more important to keep the complete face than the complete human body. Therefore, in practical applications, it is more inclined to use the face detector as the main detector.

[0143] Secondly, in the embodiments of the present application, a detection method is provided in which the area where a human face or body is located is used as a candidate area. Through the above method, the trained detection network model can automatically identify objects such as human faces or bodies that appear in the image, thereby realizing the function of automatic detection without manual participation, improving the detection efficiency, and saving time and labor costs.

[0144] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, feature extraction processing is performed on the image to be recognized to obtain a target feature map, which may specifically include:

[0145] Based on the image to be recognized, a saliency feature map is obtained through the first network included in the feature extraction network, where the saliency feature map corresponds to 1 channel;

[0146] Based on the image to be recognized, a deep semantic embedding map is obtained through the second network included in the feature extraction network, where the deep semantic embedding map corresponds to C channels, and C is an integer greater than 1;

[0147] The saliency feature map and the deep semantic embedding map are concatenated to obtain a target feature map, where the target feature map includes (C + 1) channels.

[0148] In this embodiment, a method of extracting two types of feature maps as the target feature map is introduced. An integrated feature extraction network is used to extract two different feature maps of the image to be recognized, namely, the saliency feature map and the deep semantic embedding map.

[0149] Specifically, for ease of understanding, please refer to Figure 6 , Figure 6 is a schematic diagram of generating a target feature map based on two types of feature maps in the embodiments of the present application. As shown in the figure, the feature extraction network includes a first network and a second network. Among them, the first network can be a Cascaded Partial Decoder (CPD), and CPD discards the shallower features to ensure higher computational efficiency, and then refines the deeper features to improve their representation ability. The second network can be a Convolutional Neural Networks (CNN), Residual Network - 50 (ResNet - 50), or Residual Network - 18 (ResNet - 18) pre - trained using the ImageNet image dataset, etc. It should be noted that the first network and the second network can also be other networks, which are not limited here.

[0150] The image to be recognized passes through the first network to obtain a saliency feature map with 1 channel. The image to be recognized passes through the second network to obtain a depth semantic embedding map with C channels. Then, the two types of feature maps are concatenated in the channel dimension to obtain a target feature map with (C + 1) channels, that is, the target feature map can be expressed as F = [F sal , F e . Wherein, F represents the target feature map, F sal represents the saliency feature map, and F e represents the depth semantic embedding map.

[0151] Secondly, in the embodiments of the present application, a method for extracting two types of feature maps as the target feature map is provided. Through the above method, on the one hand, the target feature map containing the saliency feature map can reflect the contour of the candidate subject. On the other hand, the target feature map containing the depth semantic embedding map has better representation ability, that is, it can refine the image to be recognized into a better data representation. Therefore, based on this target feature map for subsequent tasks, the accuracy of subject selection can be improved.

[0152] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, the image to be recognized is subjected to feature extraction processing to obtain a target feature map, which may specifically include:

[0153] Based on the image to be recognized, a saliency feature map is obtained through the first network included in the feature extraction network, wherein the saliency feature map corresponds to 1 channel;

[0154] Based on the image to be recognized, a depth semantic embedding map is obtained through the second network included in the feature extraction network, wherein the depth semantic embedding map corresponds to C channels, and C is an integer greater than 1;

[0155] Based on the image to be recognized, a blur feature map is obtained through the third network included in the feature extraction network, wherein the blur feature map corresponds to 1 channel;

[0156] The saliency feature map, the depth semantic embedding map, and the blur feature map are subjected to concatenation processing to obtain a target feature map, wherein the target feature map includes (C + 2) channels.

[0157] In this embodiment, a method for extracting three types of feature maps as the target feature map is introduced. An integrated feature extraction network is used to extract three different feature maps of the image to be recognized, namely, the saliency feature (SalientFeature) map, the depth semantic embedding (Deep Semantic Embedding) map, and the blur feature (BlurFeature) map.

[0158] Specifically, for ease of understanding, please refer to Figure 7 , Figure 7 which is another schematic diagram of generating a target feature map based on three types of feature maps in the embodiments of the present application. As shown in the figure, the feature extraction network includes a first network, a second network, and a third network. Among them, the first network can be CPD. The second network can be a CNN pre-trained with ImageNet, resnet-50, or resnet-18, etc. The third network can adopt the Thresholded Gradient Magnitude Maximization (Tenengrad) detection algorithm. The blur detection aims to detect the Just Noticeable Blur (JNB) caused by defocusing, which only spans a small number of pixels in the image.

[0159] After the image to be recognized passes through the first network, a saliency feature map with 1 channel is obtained. After the image to be recognized passes through the second network, a depth semantic embedding map with C channels is obtained. After the image to be recognized passes through the third network, a blur degree feature map with 1 channel is obtained. Then, the three types of feature maps are concatenated (concat) in the channel dimension to obtain a target feature map with (C + 2) channels. That is, the target feature map can be expressed as F = [F sal , F b , F e . Among them, F represents the target feature map, F b represents the blur degree feature map, F sal represents the saliency feature map, and F e represents the depth semantic embedding map.

[0160] Secondly, in the embodiments of the present application, a method for extracting three types of feature maps as the target feature map is provided. Through the above method, on the one hand, the target feature map containing the saliency feature map can reflect the contour of the candidate subject. On the other hand, the target feature map containing the depth semantic embedding map has better representation ability, that is, it can refine the image to be recognized into a better data expression. In addition, the target feature map containing the blur degree feature map can blur the background part and clarify the foreground part, that is, improve the recognition degree of the image to be recognized. Thus, based on this target feature map for subsequent tasks, the accuracy of subject selection can be improved.

[0161] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, according to the target feature map and N candidate regions, obtaining the subject score corresponding to each candidate region in the N candidate regions may specifically include:

[0162] Match each of the N candidate regions with the target feature map to obtain N spatial feature maps, where there is a one-to-one correspondence between the spatial feature maps and the candidate regions;

[0163] For each of the N candidate regions, based on the spatial feature map corresponding to each candidate region, obtain the first image feature through the first convolutional network included in the main body selection network;

[0164] For each of the N candidate regions, based on the spatial feature map corresponding to each candidate region, obtain the second image feature through the second convolutional network included in the main body selection network;

[0165] For each of the N candidate regions, perform splicing processing on the first image feature and the second image feature corresponding to each candidate region to obtain the comprehensive image feature corresponding to each candidate region;

[0166] Based on the comprehensive image feature corresponding to each candidate region, obtain the main body score corresponding to each candidate region through the fully connected layer included in the main body selection network.

[0167] In this embodiment, a method for predicting the main body score based on the Rank-Subject Selection (Rank-SS) network is introduced. After the detector finds the candidate regions (i.e., candidate main bodies) in the image to be recognized, it is a highly experience- and subjective-dependent task to select the target candidate region (i.e., target main body) from these candidate regions (i.e., candidate main bodies). The criteria for selecting the target main body mainly include four criteria, specifically:

[0168] (1) Center criterion: The target main body often lies in the center of the scene, that is, the target main body will appear near the center of the picture.

[0169] (2) Focus criterion: The target main body should appear within the focal length, that is, the target main body is the focus of the picture rather than blurred.

[0170] (3) Proportion criterion: The target main body tends to occupy most of the scene, that is, the target main body is relatively large in the picture.

[0171] (4) Pose criterion: The target main body often shows a relatively prominent pose, that is, usually a frontal face rather than a side face or a back view.

[0172] Specifically, in combination with the business data distribution and the main body selection annotation, the present application proposes a Rank-SS model based on R-CNN, considering the candidate region main body selection problem as a region ranking problem. For ease of understanding, please refer to Figure 8 , Figure 8This is a schematic diagram of a prediction process of the sorting subject selection model in the embodiments of the present application. As shown in the figure, the image to be recognized is input into a detection network (for example, a face detection network or a human body detection network, etc.) and a feature extraction network (for example, three types of feature maps are extracted) respectively. N candidate regions are output through the detection network. Taking Figure 8 as an example, that is, 2 candidate regions are output. The feature extraction network extracts a depth semantic embedding map, a blurriness feature map, and a saliency feature map respectively. After concatenating the three types of feature maps, a target feature map is obtained. Based on ROI Align, each candidate region is matched with the target feature map to obtain N spatial feature maps (for example, 2 spatial feature maps). That is, the following method is used for processing:

[0173] F i =ROIAlign(F, c i );

[0174] where F represents the target feature map, c i represents the i-th candidate region, and F i represents the spatial feature map of the i-th candidate region.

[0175] Each spatial feature map is input into the first convolutional network and the second convolutional network included in the subject selection network respectively, and thus the first image feature and the second image feature are obtained. After concatenating the first image feature and the second image feature, the comprehensive image feature of each candidate region is obtained. Similar to the model of typical two-stage object detection, it can effectively share the features extracted between regions, and at the same time obtain relatively independent regional features. The spatial feature map passes through convolutional and fully connected layers to obtain the subject score of the i-th candidate region. That is, the subject score is determined in the following manner:

[0176] s i =RankSS(I, c i , w);

[0177] where s i represents the subject score of the i-th candidate region, c i represents the i-th candidate region, I represents the image to be recognized, and w represents the model parameters of Rank-SS.

[0178] Again, in the embodiments of the present application, a method for predicting the subject score based on the Rank-SS network is provided. Through the above method, the mutual relationship between candidate regions in the image can be concerned, and the relative subject order relationship within the image can be maintained. Therefore, the Rank-SS network can better identify the main subject and the secondary subject in the image, thereby improving the accuracy of subject recognition.

[0179] Optionally, in the aboveFigure 3 Based on the corresponding embodiment, in another alternative embodiment provided by the embodiments of the present application, it may further include:

[0180] Obtain a to-be-trained image sample, where the to-be-trained image sample includes a first annotation area and a second annotation area, the first annotation area corresponds to a first annotation score, and the second annotation area corresponds to a second annotation score;

[0181] Perform feature extraction processing on the to-be-trained image sample to obtain the target feature map of the to-be-trained image sample;

[0182] Match the first annotation area with the target feature map to obtain a first spatial feature map, and match the second annotation area with the target feature map to obtain a second spatial feature map;

[0183] Based on the first spatial feature map, obtain a first predicted image feature through the first convolutional network included in the to-be-trained subject selection network, and based on the second spatial feature map, obtain a second predicted image feature through the second convolutional network included in the to-be-trained subject selection network;

[0184] Based on the first annotation area, obtain a third predicted image feature through the second convolutional network included in the to-be-trained subject selection network, and based on the second annotation area, obtain a fourth predicted image feature through the second convolutional network included in the to-be-trained subject selection network;

[0185] Perform splicing processing on the first predicted image feature and the third predicted image feature corresponding to the first annotation area to obtain a first comprehensive image feature corresponding to the first annotation area, and perform splicing processing on the second predicted image feature and the fourth predicted image feature corresponding to the second annotation area to obtain a second comprehensive image feature corresponding to the second annotation area;

[0186] Based on the first comprehensive image feature, obtain a first predicted subject score corresponding to the first annotation area through the fully connected layer included in the to-be-trained subject selection network, and based on the second comprehensive image feature, obtain a second predicted subject score corresponding to the second annotation area through the fully connected layer included in the to-be-trained subject selection network;

[0187] Update the model parameters of the to-be-trained subject selection network according to the first predicted subject score, the second predicted subject score, the first annotation score, and the second annotation score until the model training condition is met, and output the subject selection network.

[0188] In this embodiment, a method for training the RANK-SS network is introduced. During the training process, the typical point-wise loss and margin-ranking of pair-wise loss in the ranking task can be adopted. The RANK-SS network includes two parts, namely the subject selection network and the feature extraction network. Among them, the feature extraction network can adopt a pre-trained network or be jointly trained with the subject selection network. Hereinafter, the case of using a pre-trained feature extraction network will be taken as an example for introduction.

[0189] Specifically, training the subject selection network usually requires a large number of training image samples. Taking any training image sample as an example, first, the first annotation area and the second annotation area in the training image sample need to be manually annotated. Among them, assuming that the first annotation area is the subject, its corresponding first annotation score can be set to 1, and assuming that the second annotation area is the secondary subject, its corresponding second annotation score can be set to 0.

[0190] For ease of understanding, please refer to Figure 9 , Figure 9 which is a schematic diagram of a training process of the ranking subject selection model in the embodiment of the present application. As shown in the figure, the pre-trained feature extraction network is used to perform feature extraction processing on the training image sample to obtain the target feature map of the training image sample. Then, the first annotation area is matched with the target feature map to obtain the first spatial feature map, and the second annotation area is matched with the target feature map to obtain the second spatial feature map. The first spatial feature map is input into the first convolutional network included in the training subject selection network to obtain the first predicted image feature, and the second spatial feature map is input into the second convolutional network included in the training subject selection network to obtain the second predicted image feature. At the same time, the first annotation area is input into the second convolutional network included in the training subject selection network to obtain the third predicted image feature, and the second annotation area is input into the second convolutional network included in the training subject selection network to obtain the fourth predicted image feature.

[0191] Based on this, the first predicted image feature and the third predicted image feature corresponding to the first annotation area are spliced to obtain the first comprehensive image feature corresponding to the first annotation area, and the second predicted image feature and the fourth predicted image feature corresponding to the second annotation area are spliced to obtain the second comprehensive image feature corresponding to the second annotation area.

[0192] Input the first comprehensive image feature into the fully connected layer included in the to-be-trained subject selection network to obtain the first predicted subject score corresponding to the first annotation region, and input the second comprehensive image feature into the fully connected layer included in the to-be-trained subject selection network to obtain the second predicted subject score corresponding to the second annotation region. Calculate the loss value among the first predicted subject score, the second predicted subject score, the first annotation score, and the second annotation score based on the loss function, and use the loss value to train the to-be-trained subject selection network based on gradient backpropagation until the model training condition is satisfied, that is, obtain the subject selection network.

[0193] It should be noted that the situation where the model training condition is satisfied can be reaching the number of iterations or the loss value has converged, which is not limited here.

[0194] Furthermore, in the embodiments of the present application, a method for training the RANK-SS network is provided. Through the above method, during the training process, only the ranking relationship between the main subject and the secondary subject needs to be considered, without classifying the candidate subjects. Instead, the subject score is determined based on the ordered relationship between the candidate subjects. Thus, the diversity and accuracy of subject recognition are improved.

[0195] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding embodiment, according to the target feature map and N candidate regions, obtain the subject score corresponding to each candidate region in the N candidate regions, which may specifically include:

[0196] For each candidate region among the N candidate regions, crop the feature map corresponding to each candidate region from the target feature map;

[0197] Perform global average pooling on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1;

[0198] Determine the subject score corresponding to each candidate region according to the region feature corresponding to each candidate region.

[0199] In this embodiment, a method for predicting the subject score based on the Naive Subject Selection ( -Subject Selection, NSS) network is introduced. First, crop (Crop) the feature map corresponding to each candidate region from the target feature map, and then perform global average pooling (Golbal Average Pooling, GAP) on the feature map of each candidate region to obtain its corresponding region feature. Based on the region feature of each candidate region, directly calculate its corresponding subject score.

[0200] Specifically, the main score of the i-th candidate region is calculated in the following manner:

[0201] f i = GAP(Crop(F, c i ));

[0202] s i = NSS(f i , c i ; w) = w T [f i , c i ;

[0203] Among them, f i represents the region feature of the i-th candidate region, F represents the target feature map, c i represents the i-th candidate region, s i represents the main score of the i-th candidate region, and w represents the weight matrix that can be manually debugged.

[0204] Again, in the embodiments of the present application, a method for predicting the main score based on NSS is provided. Through the above method, the main score can be directly predicted without data training, and the model has strong interpretability and can quickly adapt to scene pictures with different features.

[0205] Optionally, on the basis of the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiments of the present application, according to the region feature corresponding to each candidate region, the main score corresponding to each candidate region is determined, which may specifically include:

[0206] According to the region feature corresponding to each candidate region, the feature average value corresponding to each candidate region is determined, where the feature average value is the average value of M feature values in the region feature;

[0207] For each candidate region among the N candidate regions, the feature average value of the candidate region is used as the main score of the candidate region;

[0208] Or,

[0209] According to the region feature corresponding to each candidate region, determining the main score corresponding to each candidate region includes:

[0210] According to the first region feature and the first region weight corresponding to each candidate region, the first feature average value corresponding to each candidate region is determined;

[0211] According to the second region feature and the second region weight corresponding to each candidate region, the second feature average value corresponding to each candidate region is determined, where the second region weight is greater than the first region weight;

[0212] Determine the target average value corresponding to each candidate region according to the first feature average value and the second feature average value corresponding to each candidate region;

[0213] For each of the N candidate regions, use the target average value of the candidate region as the main score of the candidate region.

[0214] In this embodiment, two methods for predicting the main score based on the NSS network are introduced. For ease of understanding, the following 4*4 matrix will be used as the region feature of a candidate region for illustration:

[0215]

[0216] Combined with the above region feature including 16 feature values (i.e., M = 16). In one example, the average value of these 16 feature values can be calculated, that is, the feature average value is obtained. After calculation, the feature average value of the above region feature is 0.55, that is, the main score of this candidate region is also 0.55.

[0217] In another example, the boundary feature in the candidate region is used as the first region feature, and the remaining features are used as the second region feature. Among them, the first region feature corresponds to the first region weight, assuming the first region weight is 0.5, and the second region feature corresponds to the second region weight, assuming the second region weight is 1. After calculation, the first feature average value is 2.85, and the second feature average value is 3.1. Based on this, the first feature average value and the second feature average value are summed to obtain the face average value of 5.95. That is, the main score of this candidate region is also 5.95.

[0218] Furthermore, in the embodiments of the present application, two methods for predicting the main score based on NSS are provided. Through the above methods, the main score can be calculated in different ways, and there is no need for data training, with strong interpretability, and can quickly adapt to scene pictures with different features.

[0219] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to the above embodiment, according to the target feature map and N candidate regions, obtain the main score corresponding to each candidate region, which may specifically include:

[0220] For each of the N candidate regions, crop the feature map corresponding to each candidate region from the target feature map;

[0221] Perform global average pooling processing on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1;

[0222] Based on the region features corresponding to each candidate region, the subject score corresponding to each candidate region is obtained through a multi-layer perceptron.

[0223] In this embodiment, a method for predicting the subject score based on a multi-layer perceptron subject selection (MLP-SS) network is introduced. In a complex scenario with a large number of candidate subjects, sub-optimal solutions often occur. Therefore, an MLP-SS network trained with data can be used, which has higher accuracy in subject classification. After taking the region features as the input of the MLP-SS network, whether it is a subject is treated as a binary classification problem, and the one with the highest subject score during prediction is the final target subject.

[0224] Specifically, the following method is used to calculate the subject score of the i-th candidate region:

[0225] f i = GAP(Crop(F, c i ));

[0226] s i = MLP(f i , c i ; w);

[0227] Among them, f i represents the region features of the i-th candidate region, F represents the target feature map, c i represents the i-th candidate region, and s i represents the subject score of the i-th candidate region.

[0228] Furthermore, in the embodiment of the present application, a method for predicting the subject score based on the MLP-SS network is provided. Through the above method, since the multi-layer perceptron subject selection network involves data training, the relationship between image features and subject scores can be learned, thus effectively improving the accuracy of subject selection and being applicable to subject selection in a relatively simple multi-person scenario, thereby enhancing the feasibility and operability of the solution.

[0229] Optionally, based on the above Figure 3 corresponding embodiment, in another optional embodiment provided by the embodiment of the present application, it may further include:

[0230] Obtain the image sample to be trained, where the image sample to be trained includes a first labeled region and a second labeled region, the first labeled region corresponds to a first labeled score, and the second labeled region corresponds to a second labeled score;

[0231] Perform feature extraction processing on the image sample to be trained to obtain the target feature map of the image sample to be trained;

[0232] Crop the feature map corresponding to the first annotation region and the feature map corresponding to the second annotation region from the target feature map of the image sample to be trained;

[0233] Perform global average pooling on the feature map corresponding to the first annotation region to obtain the region feature corresponding to the first annotation region;

[0234] Perform global average pooling on the feature map corresponding to the second annotation region to obtain the region feature corresponding to the second annotation region;

[0235] Based on the region feature corresponding to the first annotation region, obtain the first predicted subject score corresponding to the first annotation region through the multi-layer perceptron to be trained;

[0236] Based on the region feature corresponding to the second annotation region, obtain the second predicted subject score corresponding to the second annotation region through the multi-layer perceptron to be trained;

[0237] According to the first predicted subject score, the second predicted subject score, the first annotation score, and the second annotation score, update the model parameters of the multi-layer perceptron to be trained until the model training conditions are met, and output the multi-layer perceptron.

[0238] In this embodiment, a method for training the MLP-SS network is introduced. The MLP-SS network includes two parts, namely MLP and a feature extraction network. Among them, the feature extraction network can use a pre-trained network or be jointly trained with the MLP. Hereinafter, the case of using a pre-trained feature extraction network will be taken as an example for introduction.

[0239] Specifically, training the MLP usually requires a large number of image samples to be trained. Taking any image sample to be trained as an example, first, the first annotation region and the second annotation region in the image sample to be trained need to be manually annotated. Among them, it is assumed that the first annotation region is the main body, that is, the corresponding first annotation score can be set to 1, and it is assumed that the second annotation region is the secondary main body, that is, the corresponding second annotation score can be set to 0. Use the trained feature extraction network to perform feature extraction processing on the image sample to be trained to obtain the target feature map of the image sample to be trained. Then, crop the feature map corresponding to the first annotation region and the feature map corresponding to the second annotation region from the target feature map. Perform GAP processing on the feature map corresponding to the first annotation region and the feature map corresponding to the second annotation region respectively to obtain the region feature corresponding to the first annotation region and the region feature corresponding to the second annotation region.

[0240] Input the region features corresponding to the first labeled region into the MLP to be trained, so as to obtain the first predicted subject score. Similarly, input the region features corresponding to the second labeled region into the MLP to be trained, so as to obtain the second predicted subject score. Calculate the loss value between the first predicted subject score and the first labeled score based on the loss function, and calculate the loss value between the second predicted subject score and the first labeled score. Add the two loss values to obtain the target loss value, and use the target loss value to train the MLP to be trained based on gradient backpropagation until the model training conditions are met, that is, the MLP is obtained.

[0241] It should be noted that the situation where the model training conditions are met can be reaching the number of iterations or the loss value has converged, which is not limited here.

[0242] Furthermore, in the embodiments of the present application, a method for training the MLP-SS network is provided. Through the above method, the process of training the multi-layer perceptron main body selection network based on supervised learning is relatively simple, can be widely applied to industrial problems. At the same time, during the training process, the amount of calculation required for classification is small, the speed is fast, and the storage resources are less.

[0243] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to the above embodiment, obtaining the image to be recognized may specifically include:

[0244] Perform frame processing on the video to be processed to obtain K video frames, where K is an integer greater than 1;

[0245] For each of the K video frames, obtain the brightness, clarity, and color simplicity;

[0246] According to the brightness of each video frame and the brightness threshold, the clarity of each video frame and the clarity threshold, and the color simplicity of each video frame and the color simplicity threshold, perform frame filtering processing on the K video frames to obtain L video frames, where L is an integer greater than or equal to 1 and less than K;

[0247] Select one video frame from the L video frames as the image to be recognized;

[0248] After determining the target candidate region from the N candidate regions according to the subject score corresponding to each candidate region, and taking the candidate subject corresponding to the target candidate region as the target subject in the image to be recognized, it may further include:

[0249] Crop the target image from the image to be recognized according to the target candidate region, where the target image corresponds to the target size and the target image includes the target candidate region;

[0250] Use the target image as the video cover of the video to be processed.

[0251] In this embodiment, a method for generating a video cover based on subject recognition is introduced. First, obtain the video to be processed, and then perform frame splitting on the video to be processed to generate K video frames. Then, obtain the brightness, clarity, and color uniformity of each video frame respectively. Among them, video frame filtering mainly filters out low-quality video frames and transition frames. Low-quality video frames often have low brightness, clarity, or color uniformity. Therefore, filter out video frames with brightness less than or equal to the brightness threshold, video frames with clarity less than or equal to the clarity threshold, and video frames with color uniformity less than or equal to the color uniformity threshold from the K video frames, and finally obtain L video frames.

[0252] Specifically, any one of the L video frames can be selected as the image to be recognized, or a video frame can be manually selected as the image to be recognized, or other strategies can be used to select the image to be recognized, which is not limited here. Based on this, after obtaining the subject scores corresponding to each candidate region in the image to be recognized, the target candidate region can be selected, and the target candidate region can be used as the basis for image cropping to obtain the target image. For example, if the size of the image to be recognized is 16:9 and the target size is 4:3, then the size of the output target image is 4:3 and includes the target subject (e.g., Person A) in the target candidate region. Finally, directly use the target image as the video cover of the video to be processed.

[0253] For ease of understanding, please refer to Figure 10 , Figure 10 is a schematic flow diagram for implementing intelligent image cropping in an embodiment of the present application. As shown in the figure, assume that the image to be recognized is a landscape image. After subject recognition, the position of the target subject in the image to be recognized is determined. Based on this, the target candidate region where the target subject is located is cropped to obtain the target image, and the target image is a portrait image.

[0254] Again, in the embodiment of the present application, a method for generating a video cover based on subject recognition is provided. Through the above method, for any-sized video to be processed, the purpose of intelligent cropping can be achieved, thereby providing materials for video display, meeting differentiated requirements, and saving a large amount of labor costs.

[0255] Optionally, based on the corresponding embodiment above, in another optional embodiment provided by the embodiment of the present application, cropping the target image from the image to be recognized according to the target candidate region may specifically include: Figure 3

[0256] ​Determine secondary candidate regions from N candidate regions according to the main scores corresponding to each candidate region, where the secondary candidate regions correspond to the second largest main scores;

[0257] If the main score corresponding to the secondary candidate region is greater than or equal to the main score threshold, obtain the first position coordinates of the center of the secondary candidate region in the image to be recognized, and obtain the second position coordinates of the center of the target candidate region in the image to be recognized;

[0258] Determine the regional distance according to the first position coordinates and the second position coordinates;

[0259] If the regional distance is less than or equal to the distance threshold corresponding to the target size, crop the target image from the image to be recognized according to the target candidate region and the secondary candidate region, where the target image also includes the secondary candidate region.

[0260] In this embodiment, an intelligent cropping method is introduced. As can be seen from the foregoing embodiments, according to the main scores corresponding to each candidate region, the candidate region with the largest main score can be selected as the target candidate region, and the candidate main body in the target subsequent region is the target main body, that is, the main body in the image to be recognized. The remaining candidate regions also have a main score respectively. Therefore, select the candidate region corresponding to the second largest main score from the remaining candidate regions as the secondary candidate region.

[0261] Specifically, if the main score corresponding to the secondary candidate region is greater than or equal to the main score threshold, it means that the candidate main body in the secondary candidate region can also be considered to be cropped out for display. However, since the distance between the secondary candidate region and the target candidate region may not allow them to appear in the target image at the same time, further determination is required. The determination method is as follows: First, obtain the first position coordinates of the center of the secondary candidate region in the image to be recognized, and obtain the second position coordinates of the center of the target candidate region in the image to be recognized. Based on the first position coordinates and the second position coordinates, calculate the straight-line distance, that is, obtain the regional distance. Based on this, first place the target candidate region at the center position of the target image, that is, the first position coordinates coincide with the center coordinates of the target image. Then, according to the target size, determine its corresponding distance threshold. If the regional distance is less than or equal to the distance threshold corresponding to the target size, the secondary candidate region and the target candidate region can be cropped into the same image (i.e., the target image) and output at the same time.

[0262] Combined with the above introduction, the following will be illustrated with actual examples. Please refer to Figure 11 , Figure 11This is a schematic diagram of the effect comparison obtained by analyzing the samples based on the subject recognition method in the embodiment of the present application. As shown in the figure, the first row belongs to the output results of intelligent cropping, and the second row belongs to the sample with poor effect. Among them, the white solid line is the subject automatically identified based on the method provided by the present application, the white dotted line is the subject manually marked, and the gray area is the target image (i.e., the cropping area) with a ratio of 9 to 16. It can be seen that the subject recognition method provided by the present application can effectively select a suitable subject (i.e., the target subject) in scenes such as crowds or side faces. At the same time, when a secondary subject appears around the main subject (i.e., the target subject), if the cropping area is sufficient, it will also be adaptively covered.

[0263] Furthermore, in an embodiment of the present application, a smart cropping method is provided. Through the above method, considering that the loss of picture information is inevitable during the cropping process, the main subject in the image is first found, and then the secondary subject is placed in the cropping area as much as possible, thereby reducing the amount of information loss and helping to improve the cropping effect.

[0264] The subject identification device in this application is described in detail below. Figure 12 , Figure 12 This is a schematic diagram of an embodiment of a subject identification device in an embodiment of the present application. The subject identification device 20 includes:

[0265] An acquisition module 201 is used to acquire an image to be recognized;

[0266] The processing module 202 is used to perform region recognition processing on the image to be recognized to obtain N candidate regions, and perform feature extraction processing on the image to be recognized to obtain a target feature map, wherein each candidate region corresponds to a candidate subject, and the target feature map is obtained by splicing at least two feature maps, and N is an integer greater than or equal to 1;

[0267] The acquisition module 201 is further used to acquire a subject score corresponding to each of the N candidate regions according to the target feature map and the N candidate regions;

[0268] The determination module 203 is used to determine a target candidate region from N candidate regions according to the subject score corresponding to each candidate region, and use the candidate subject corresponding to the target candidate region as the target subject in the image to be identified, wherein the target candidate region corresponds to the maximum subject score.

[0269] In the embodiments of the present application, a subject recognition device is provided, and the above device is adopted. First, several candidate regions are recognized, then the subject scores of each candidate region are automatically calculated, and finally, the candidate region with the highest subject score is used as the target candidate region, that is, the candidate subject in the target candidate region is the subject in the image, and the candidate subjects in other candidate regions are secondary subjects in the image. Thus, the purpose of automatically processing the image is achieved, which not only reduces the time cost and labor cost, but also the subject selected based on the subject score has high accuracy and can meet the large-scale requirements.

[0270] Optionally, on the basis of the above Figure 12 corresponding embodiment, in another embodiment of the subject recognition device 20 provided by the embodiments of the present application,

[0271] The processing module 202 is specifically configured to obtain N candidate regions through a face detection network based on the image to be recognized, where the candidate subject corresponding to each candidate region is a human face;

[0272] Or,

[0273] Based on the image to be recognized, N candidate regions are obtained through a human body detection network, where the candidate subject corresponding to each candidate region is a human body.

[0274] In the embodiments of the present application, a subject recognition device is provided, and the above device is adopted. The trained detection network model can automatically recognize objects such as human faces or human bodies that appear in the image, thereby realizing the function of automatic detection without manual participation, thus improving the detection efficiency and saving time and labor costs.

[0275] Optionally, on the basis of the above Figure 12 corresponding embodiment, in another embodiment of the subject recognition device 20 provided by the embodiments of the present application,

[0276] The processing module 202 is specifically configured to obtain a saliency feature map through a first network included in the feature extraction network based on the image to be recognized, where the saliency feature map corresponds to 1 channel;

[0277] Based on the image to be recognized, a depth semantic embedding map is obtained through a second network included in the feature extraction network, where the depth semantic embedding map corresponds to C channels, and C is an integer greater than 1;

[0278] The saliency feature map and the depth semantic embedding map are subjected to splicing processing to obtain a target feature map, where the target feature map includes (C + 1) channels.

[0279] In an embodiment of the present application, a subject recognition device is provided, and the above device is adopted. On the one hand, the target feature map containing the saliency feature map can reflect the contour of the candidate subject. On the other hand, the depth semantic embedding map included has better representation ability, that is, it can refine the image to be recognized into a better data representation. Therefore, performing subsequent tasks based on this target feature map can improve the accuracy of subject selection.

[0280] Optionally, based on the corresponding embodiment above, Figure 12 In another embodiment of the subject recognition device 20 provided in the embodiment of the present application,

[0281] The processing module 202 is specifically configured to obtain a saliency feature map based on the image to be recognized through a first network included in the feature extraction network, where the saliency feature map corresponds to 1 channel;

[0282] Based on the image to be recognized, obtain a depth semantic embedding map through a second network included in the feature extraction network, where the depth semantic embedding map corresponds to C channels, and C is an integer greater than 1;

[0283] Based on the image to be recognized, obtain a blurriness feature map through a third network included in the feature extraction network, where the blurriness feature map corresponds to 1 channel;

[0284] Perform splicing processing on the saliency feature map, the depth semantic embedding map, and the blurriness feature map to obtain a target feature map, where the target feature map includes (C + 2) channels.

[0285] In an embodiment of the present application, a subject recognition device is provided, and the above device is adopted. On the one hand, the target feature map containing the saliency feature map can reflect the contour of the candidate subject. On the other hand, the depth semantic embedding map included has better representation ability, that is, it can refine the image to be recognized into a better data representation. In addition, the target feature map containing the blurriness feature map can blur the background part and clarify the foreground part, that is, improve the recognition degree of the image to be recognized. Therefore, performing subsequent tasks based on this target feature map can improve the accuracy of subject selection.

[0286] Optionally, based on the corresponding embodiment above, Figure 12 In another embodiment of the subject recognition device 20 provided in the embodiment of the present application,

[0287] The acquisition module 201 is specifically configured to perform matching processing on each of the N candidate regions and the target feature map to obtain N spatial feature maps, where the spatial feature maps have a one-to-one correspondence with the candidate regions;

[0288] For each of the N candidate regions, based on the spatial feature map corresponding to each candidate region, a first image feature is obtained through the first convolutional network included in the main body selection network;

[0289] For each of the N candidate regions, based on the spatial feature map corresponding to each candidate region, a second image feature is obtained through the second convolutional network included in the main body selection network;

[0290] For each of the N candidate regions, the first image feature and the second image feature corresponding to each candidate region are concatenated to obtain the comprehensive image feature corresponding to each candidate region;

[0291] Based on the comprehensive image feature corresponding to each candidate region, the main body score corresponding to each candidate region is obtained through the fully connected layer included in the main body selection network.

[0292] In the embodiment of the present application, a main body recognition device is provided, which adopts the above device. It can pay attention to the mutual relationship between candidate regions in the image and maintain the relative main body order relationship in the image. Therefore, the Rank-SS network can better identify the main body and secondary main body in the image, thereby improving the accuracy of main body recognition.

[0293] Optionally, on the basis of the corresponding embodiment above, Figure 12 In another embodiment of the main body recognition device 20 provided in the embodiment of the present application, the main body recognition device 20 further includes an update module 204;

[0294] The acquisition module 201 is further configured to acquire a training image sample, where the training image sample includes a first labeled region and a second labeled region, the first labeled region corresponds to a first labeled score, and the second labeled region corresponds to a second labeled score;

[0295] The processing module 202 is further configured to perform feature extraction processing on the training image sample to obtain a target feature map of the training image sample;

[0296] The processing module 202 is further configured to match the first labeled region with the target feature map to obtain a first spatial feature map, and match the second labeled region with the target feature map to obtain a second spatial feature map;

[0297] The acquisition module 201 is further configured to obtain a first predicted image feature based on the first spatial feature map through the first convolutional network included in the training main body selection network, and obtain a second predicted image feature based on the second spatial feature map through the second convolutional network included in the training main body selection network;

[0298] The obtaining module 201 is further configured to obtain a third predicted image feature through a second convolutional network included in the to-be-trained subject selection network based on the first labeled region, and obtain a fourth predicted image feature through the second convolutional network included in the to-be-trained subject selection network based on the second labeled region;

[0299] The processing module 202 is further configured to perform splicing processing on the first predicted image feature corresponding to the first labeled region and the third predicted image feature to obtain a first comprehensive image feature corresponding to the first labeled region, and perform splicing processing on the second predicted image feature corresponding to the second labeled region and the fourth predicted image feature to obtain a second comprehensive image feature corresponding to the second labeled region;

[0300] The obtaining module 201 is further configured to obtain a first predicted subject score corresponding to the first labeled region through a fully connected layer included in the to-be-trained subject selection network based on the first comprehensive image feature, and obtain a second predicted subject score corresponding to the second labeled region through the fully connected layer included in the to-be-trained subject selection network based on the second comprehensive image feature;

[0301] The updating module 204 is configured to update the model parameters of the to-be-trained subject selection network according to the first predicted subject score, the second predicted subject score, the first labeled score, and the second labeled score until the model training condition is met, and output the subject selection network.

[0302] In the embodiment of the present application, a subject recognition device is provided, which adopts the above device. During training, only the sorting relationship between the main subject and the secondary subject needs to be considered, and there is no need to classify the candidate subjects. Instead, the subject scores are determined based on the ordered relationship between the candidate subjects. Thus, the diversity and accuracy of subject recognition are improved.

[0303] Optionally, on the basis of the above Figure 12 corresponding embodiment, in another embodiment of the subject recognition device 20 provided in the embodiment of the present application,

[0304] The obtaining module 201 is specifically configured to, for each of the N candidate regions, crop the feature map corresponding to each candidate region from the target feature map;

[0305] Perform global average pooling processing on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1;

[0306] Determine the subject score corresponding to each candidate region according to the region feature corresponding to each candidate region.

[0307] In an embodiment of the present application, a subject recognition device is provided. By using the above device, the subject score can be directly predicted without data training, and the model has strong interpretability, and can quickly adapt to scene pictures with different features.

[0308] Optionally, based on the corresponding embodiment above, in another embodiment of the subject recognition device 20 provided in the embodiment of the present application, Figure 12 the obtaining module 201 is specifically configured to determine the average feature value corresponding to each candidate region according to the region features corresponding to each candidate region, where the average feature value is the average of M feature values in the region features;

[0309] For each candidate region among the N candidate regions, use the average feature value of the candidate region as the subject score of the candidate region;

[0310] Or,

[0311] the obtaining module 201 is specifically configured to determine the first average feature value corresponding to each candidate region according to the first region feature and the first region weight corresponding to each candidate region;

[0312] Determine the second average feature value corresponding to each candidate region according to the second region feature and the second region weight corresponding to each candidate region, where the second region weight is greater than the first region weight;

[0313] Determine the target average value corresponding to each candidate region according to the first average feature value and the second average feature value corresponding to each candidate region;

[0314] For each candidate region among the N candidate regions, use the target average value of the candidate region as the subject score of the candidate region.

[0315] In an embodiment of the present application, a subject recognition device is provided. By using the above device, the subject score can be calculated in different ways, and it does not require data training, has strong interpretability, and can quickly adapt to scene pictures with different features.

[0316] Optionally, based on the corresponding embodiment above, in another embodiment of the subject recognition device 20 provided in the embodiment of the present application,

[0317] the obtaining module 201 is specifically configured to, for each candidate region among the N candidate regions, crop the feature map corresponding to each candidate region from the target feature map; Figure 12

[0318]

[0319] ​​Perform global average pooling on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1;

[0320] Based on the region feature corresponding to each candidate region, obtain the main body score corresponding to each candidate region through a multi-layer perceptron.

[0321] In the embodiment of the present application, a main body recognition device is provided and the above device is adopted. Since the main body selection network of the multi-layer perceptron involves data training, the relationship between the image feature and the main body score can be learned, thereby effectively improving the accuracy of main body selection and being applicable to main body selection in a relatively simple multi-person scenario, thus enhancing the feasibility and operability of the solution.

[0322] Optionally, on the basis of the corresponding embodiment above, in another embodiment of the main body recognition device 20 provided in the embodiment of the present application, Figure 12 The acquisition module 201 is further configured to acquire a to-be-trained image sample, where the to-be-trained image sample includes a first labeled region and a second labeled region, the first labeled region corresponds to a first labeled score, and the second labeled region corresponds to a second labeled score;

[0323] The processing module 202 is further configured to perform feature extraction processing on the to-be-trained image sample to obtain the target feature map of the to-be-trained image sample;

[0324] The processing module 202 is further configured to crop from the target feature map of the to-be-trained image sample the feature map corresponding to the first labeled region and the feature map corresponding to the second labeled region;

[0325] The processing module 202 is further configured to perform global average pooling on the feature map corresponding to the first labeled region to obtain the region feature corresponding to the first labeled region;

[0326] The processing module 202 is further configured to perform global average pooling on the feature map corresponding to the second labeled region to obtain the region feature corresponding to the second labeled region;

[0327] The processing module 202 is further configured to perform global average pooling on the feature map corresponding to the second labeled region to obtain the region feature corresponding to the second labeled region;

[0328] The acquisition module 201 is further configured to obtain the first predicted main body score corresponding to the first labeled region through the to-be-trained multi-layer perceptron based on the region feature corresponding to the first labeled region;

[0329] The acquisition module 201 is further configured to obtain the second predicted main body score corresponding to the second labeled region through the to-be-trained multi-layer perceptron based on the region feature corresponding to the second labeled region;

[0330] The update module 204 is further configured to update the model parameters of the multi-layer perceptron to be trained according to the first predicted subject score, the second predicted subject score, the first labeled score, and the second labeled score until the model training condition is met, and output the multi-layer perceptron.

[0331] In the embodiments of the present application, a subject recognition device is provided, which adopts the above device. The process of training the multi-layer perceptron subject selection network based on supervised learning is relatively simple, can be widely applied to industrial problems. At the same time, during the training process, the amount of calculation required for classification is small, the speed is fast, and the storage resources are less.

[0332] Optionally, based on the corresponding embodiments above, in another embodiment of the subject recognition device 20 provided in the embodiments of the present application, Figure 12

[0333] The acquisition module 201 is specifically configured to perform frame splitting on the video to be processed to obtain K video frames, where K is an integer greater than 1;

[0334] For each of the K video frames, obtain the brightness, clarity, and color uniformity;

[0335] Perform frame filtering on the K video frames according to the brightness of each video frame and the brightness threshold, the clarity of each video frame and the clarity threshold, and the color uniformity of each video frame and the color uniformity threshold, to obtain L video frames, where L is an integer greater than or equal to 1 and less than K;

[0336] Select one video frame from the L video frames as the image to be recognized;

[0337] The acquisition module 201 is further configured to determine a target candidate region from the N candidate regions according to the subject score corresponding to each candidate region, and use the candidate subject corresponding to the target candidate region as the target subject in the image to be recognized. Then, according to the target candidate region, crop the target image from the image to be recognized, where the target image corresponds to the target size and the target image includes the target candidate region;

[0338] The determination module 203 is further configured to use the target image as the video cover of the video to be processed.

[0339] In the embodiments of the present application, a subject recognition device is provided, which adopts the above device. For any size of the video to be processed, the purpose of intelligent cropping can be achieved, thereby providing materials for video display, meeting differentiated needs, and saving a large amount of labor costs.

[0340] Optionally, based on the above Figure 12 ​Based on the corresponding embodiments, in another embodiment of the subject recognition device 20 provided by the embodiments of the present application,

[0341] An acquisition module 201 is specifically configured to determine secondary candidate regions from N candidate regions according to the subject scores corresponding to each candidate region, where the secondary candidate regions correspond to the second largest subject scores;

[0342] If the subject score corresponding to the secondary candidate region is greater than or equal to the subject score threshold, obtain the first position coordinates of the center of the secondary candidate region in the image to be recognized, and obtain the second position coordinates of the center of the target candidate region in the image to be recognized;

[0343] Determine the regional distance according to the first position coordinates and the second position coordinates;

[0344] If the regional distance is less than or equal to the distance threshold corresponding to the target size, crop the target image from the image to be recognized according to the target candidate region and the secondary candidate region, where the target image further includes the secondary candidate region.

[0345] In the embodiments of the present application, a subject recognition device is provided, adopting the above device. Considering that the loss of picture information is inevitable during the cropping process, therefore, first find the subject in the image, and then try to put the secondary subject into the cropping area as much as possible, so as to reduce the amount of information loss and is beneficial to improving the cropping effect.

[0346] The embodiments of the present application further provide another subject recognition device, which can be deployed on a terminal device, such as Figure 13 As shown, for the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal device can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS) device, an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:

[0347] Figure 13 What is shown is a block diagram of a part of the structure of a mobile phone related to the terminal device provided by the embodiments of the present application. Referring to Figure 13 , the mobile phone includes: a radio frequency (RF) circuit 310, a memory 320, an input unit 330, a display unit 340, a sensor 350, an audio circuit 360, a wireless fidelity (WiFi) module 370, a processor 380, and a power supply 390, etc. Those skilled in the art can understand, Figure 13The mobile phone structure shown does not constitute a limitation on the mobile phone, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0348] The following will specifically introduce each component of the mobile phone in conjunction with Figure 13 :

[0349] The RF circuit 310 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is given to the processor 380 for processing; in addition, it sends the uplink data designed to the base station. Generally, the RF circuit 310 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 310 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0350] The memory 320 can be used to store software programs and modules. The processor 380 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 320. The memory 320 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 320 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0351] The input unit 330 can be used to receive input numeric or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 330 may include a touch panel 331 and other input devices 332. The touch panel 331, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 331), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 331 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 380, and can receive and execute commands sent by the processor 380. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 331. In addition to the touch panel 331, the input unit 330 may further include other input devices 332. Specifically, the other input devices 332 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0352] The display unit 340 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 340 may include a display panel 341. Optionally, the display panel 341 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 331 can cover the display panel 341. When the touch panel 331 detects a touch operation thereon or nearby, it is transmitted to the processor 380 to determine the type of touch event. Subsequently, the processor 380 provides corresponding visual output on the display panel 341 according to the type of touch event. Although in Figure 13 it, the touch panel 331 and the display panel 341 are implemented as two independent components to realize the input and input functions of the mobile phone, but in some embodiments, the touch panel 331 and the display panel 341 can be integrated to realize the input and output functions of the mobile phone.

[0353] The mobile phone may further include at least one sensor 350, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 341 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 341 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0354] The audio circuit 360, the speaker 361, and the microphone 362 can provide an audio interface between the user and the mobile phone. The audio circuit 360 can transmit the electrical signal converted from the received audio data to the speaker 361, and the speaker 361 converts it into a sound signal for output; on the other hand, the microphone 362 converts the collected sound signal into an electrical signal, which is received by the audio circuit 360 and then converted into audio data. After the audio data is output to the processor 380 for processing, it is sent through the RF circuit 310 to, for example, another mobile phone, or the audio data is output to the memory 320 for further processing.

[0355] WiFi belongs to short - range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 370, which provides users with wireless broadband Internet access. Although Figure 13 the WiFi module 370 is shown, it can be understood that it does not belong to an essential component of the mobile phone and can be omitted completely within the scope of not changing the essence of the invention according to needs.

[0356] The processor 380 is the control center of the mobile phone, connecting various parts of the entire mobile phone using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 320, and by calling data stored in the memory 320, it executes various functions of the mobile phone and processes data. Optionally, the processor 380 may include one or more processing units; optionally, the processor 380 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 380 either.

[0357] The mobile phone further includes a power supply 390 (such as a battery) for powering each component. Optionally, the power supply can be logically connected to the processor 380 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system.

[0358] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0359] In the above embodiments, the steps executed by the terminal device may be based on the Figure 13 structure of the terminal device shown.

[0360] The embodiment of the present application further provides another subject recognition device, which can be deployed on a server. Figure 14 FIG. is a schematic structural diagram of a server provided by the embodiment of the present application. The server 400 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 422 (for example, one or more processors) and a memory 432, and one or more storage media 430 (for example, one or more mass storage devices) for storing application programs 442 or data 444. Among them, the memory 432 and the storage media 430 may be transient storage or persistent storage. The program stored in the storage media 430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 422 may be configured to communicate with the storage media 430 and execute a series of instruction operations in the storage media 430 on the server 400.

[0361] The server 400 may further include one or more power supplies 426, one or more wired or wireless network interfaces 450, one or more input / output interfaces 458, and / or one or more operating systems 441, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0362] In the above embodiments, the steps executed by the server may be based on the Figure 14 structure of the server shown.

[0363] The embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, it causes the computer to execute the methods described in the foregoing embodiments.

[0364] In an embodiment of the present application, a computer program product including a program is further provided. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0365] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0366] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical or other forms.

[0367] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0368] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0369] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0370] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. An entity recognition method based on artificial intelligence, characterized in that Including: Obtain the image to be recognized; Perform region recognition processing on the image to be recognized to obtain N candidate regions, and perform feature extraction processing on the image to be recognized to obtain a saliency feature map, a depth semantic embedding map, and a blurriness feature map. Concatenate the saliency feature map, the depth semantic embedding map, and the blurriness feature map in the channel dimension to obtain a target feature map, where each candidate region corresponds to a candidate subject, and N is an integer greater than or equal to 1; According to the target feature map and the N candidate regions, obtain the subject score corresponding to each candidate region in the N candidate regions; According to the subject score corresponding to each candidate region, determine a target candidate region from the N candidate regions, and use the candidate subject corresponding to the target candidate region as the target subject in the image to be recognized, where the target candidate region corresponds to the largest subject score.

2. The subject recognition method according to claim 1, characterized in that, The performing region recognition processing on the image to be recognized to obtain N candidate regions includes: Based on the image to be recognized, obtain the N candidate regions through a face detection network, where the candidate subject corresponding to each candidate region is a face; Or, Based on the image to be recognized, obtain the N candidate regions through a human body detection network, where the candidate subject corresponding to each candidate region is a human body.

3. The subject recognition method according to claim 1, characterized in that, The performing feature extraction processing on the image to be recognized to obtain a saliency feature map, a depth semantic embedding map, and a blurriness feature map includes: Based on the image to be recognized, obtain a saliency feature map through a first network included in the feature extraction network, where the saliency feature map corresponds to 1 channel; Based on the image to be recognized, obtain a depth semantic embedding map through a second network included in the feature extraction network, where the depth semantic embedding map corresponds to C channels, and C is an integer greater than 1; Based on the image to be recognized, obtain a blurriness feature map through a third network included in the feature extraction network, where the blurriness feature map corresponds to 1 channel; Wherein, the target feature map includes C + 2 channels.

4. The subject recognition method according to any one of claims 1 to 3, characterized in that, The obtaining the subject score corresponding to each candidate region in the N candidate regions according to the target feature map and the N candidate regions includes: Perform matching processing on each candidate region in the N candidate regions with the target feature map to obtain N spatial feature maps, where there is a one-to-one correspondence between the spatial feature map and the candidate region; For each candidate region in the N candidate regions, based on the spatial feature map corresponding to each candidate region, obtain a first image feature through a first convolutional network included in the subject selection network; For each candidate region in the N candidate regions, based on the spatial feature map corresponding to each candidate region, obtain a second image feature through a second convolutional network included in the subject selection network; For each of the N candidate regions, perform splicing processing on the first image feature and the second image feature corresponding to each candidate region to obtain the comprehensive image feature corresponding to each candidate region; Based on the comprehensive image feature corresponding to each candidate region, through the fully connected layer included in the main body selection network, obtain the main body score corresponding to each candidate region.

5. The subject recognition method according to claim 4, characterized in that, The method further includes: Obtain a training image sample, where the training image sample includes a first labeled region and a second labeled region, the first labeled region corresponds to a first labeled score, and the second labeled region corresponds to a second labeled score; Perform feature extraction processing on the training image sample to obtain the target feature map of the training image sample; Match the first labeled region with the target feature map to obtain a first spatial feature map, and match the second labeled region with the target feature map to obtain a second spatial feature map; Based on the first spatial feature map, obtain a first predicted image feature through the first convolutional network included in the training main body selection network, and based on the second spatial feature map, obtain a second predicted image feature through the second convolutional network included in the training main body selection network; Based on the first labeled region, obtain a third predicted image feature through the second convolutional network included in the training main body selection network, and based on the second labeled region, obtain a fourth predicted image feature through the second convolutional network included in the training main body selection network; Perform splicing processing on the first predicted image feature and the third predicted image feature corresponding to the first labeled region to obtain a first comprehensive image feature corresponding to the first labeled region, and perform splicing processing on the second predicted image feature and the fourth predicted image feature corresponding to the second labeled region to obtain a second comprehensive image feature corresponding to the second labeled region; Based on the first comprehensive image feature, through the fully connected layer included in the training main body selection network, obtain a first predicted main body score corresponding to the first labeled region, and based on the second comprehensive image feature, through the fully connected layer included in the training main body selection network, obtain a second predicted main body score corresponding to the second labeled region; According to the first predicted main body score, the second predicted main body score, the first labeled score, and the second labeled score, update the model parameters of the training main body selection network until the model training condition is satisfied, and output the main body selection network.

6. The subject recognition method according to any one of claims 1 to 3, characterized in that The obtaining the main body score corresponding to each of the N candidate regions according to the target feature map and the N candidate regions includes: For each of the N candidate regions, crop the feature map corresponding to each candidate region from the target feature map; Perform global average pooling on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1; Determine the subject score corresponding to each candidate region according to the region feature corresponding to each candidate region.

7. The subject recognition method according to claim 6, characterized in that The determining the subject score corresponding to each candidate region according to the region feature corresponding to each candidate region includes: Determine the average feature value corresponding to each candidate region according to the region feature corresponding to each candidate region, where the average feature value is the average of the M feature values in the region feature; For each candidate region among the N candidate regions, use the average feature value of the candidate region as the subject score of the candidate region; Or, The determining the subject score corresponding to each candidate region according to the region feature corresponding to each candidate region includes: Determine the first average feature value corresponding to each candidate region according to the first region feature and the first region weight corresponding to each candidate region; Determine the second average feature value corresponding to each candidate region according to the second region feature and the second region weight corresponding to each candidate region, where the second region weight is greater than the first region weight; Determine the target average value corresponding to each candidate region according to the first average feature value and the second average feature value corresponding to each candidate region; For each candidate region among the N candidate regions, use the target average value of the candidate region as the subject score of the candidate region.

8. The subject recognition method according to any one of claims 1 to 3, characterized in that The obtaining the subject score corresponding to each candidate region according to the target feature map and the N candidate regions includes: For each candidate region among the N candidate regions, crop the feature map corresponding to each candidate region from the target feature map; Perform global average pooling on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1; Based on the region feature corresponding to each candidate region, obtain the subject score corresponding to each candidate region through a multi-layer perceptron.

9. The subject recognition method according to claim 8, characterized in that The method further includes: Obtain a training image sample, where the training image sample includes a first labeled region and a second labeled region, the first labeled region corresponds to a first labeled score, and the second labeled region corresponds to a second labeled score; Perform feature extraction on the training image sample to obtain the target feature map of the training image sample; Crop the feature map corresponding to the first labeled region and the feature map corresponding to the second labeled region from the target feature map of the training image sample; Perform global average pooling on the feature map corresponding to the first labeled region to obtain the region feature corresponding to the first labeled region; Perform global average pooling on the feature map corresponding to the second labeled region to obtain the region feature corresponding to the second labeled region; Based on the region feature corresponding to the first labeled region, obtain the first predicted subject score corresponding to the first labeled region through the multi-layer perceptron to be trained; Based on the region feature corresponding to the second labeled region, obtain the second predicted subject score corresponding to the second labeled region through the multi-layer perceptron to be trained; According to the first predicted subject score, the second predicted subject score, the first labeled score, and the second labeled score, update the model parameters of the multi-layer perceptron to be trained until the model training condition is satisfied, and output the multi-layer perceptron.

10. The subject recognition method according to any one of claims 1 to 3, characterized in that, The obtaining the image to be recognized includes: Perform frame splitting on the video to be processed to obtain K video frames, where K is an integer greater than 1; For each of the K video frames, obtain the brightness, clarity, and color simplicity; According to the brightness of each video frame and the brightness threshold, the clarity of each video frame and the clarity threshold, and the color simplicity of each video frame and the color simplicity threshold, perform frame filtering on the K video frames to obtain L video frames, where L is an integer greater than or equal to 1 and less than K; Select one video frame from the L video frames as the image to be recognized; After determining the target candidate region from the N candidate regions according to the subject score corresponding to each candidate region, and using the candidate subject corresponding to the target candidate region as the target subject in the image to be recognized, the method further includes: Crop a target image from the image to be recognized according to the target candidate region, where the target image corresponds to a target size and the target image includes the target candidate region; Use the target image as the video cover of the video to be processed.

11. The subject recognition method according to claim 10, characterized in that, The cropping the target image from the image to be recognized according to the target candidate region includes: Determine a secondary candidate region from the N candidate regions according to the subject score corresponding to each candidate region, where the secondary candidate region corresponds to the second largest subject score; If the subject score corresponding to the secondary candidate region is greater than or equal to the subject score threshold, obtain the first position coordinate of the center of the secondary candidate region in the image to be recognized, and obtain the second position coordinate of the center of the target candidate region in the image to be recognized; Determine the region distance according to the first position coordinate and the second position coordinate; If the region distance is less than or equal to the distance threshold corresponding to the target size, crop a target image from the image to be recognized according to the target candidate region and the secondary candidate region, where the target image further includes the secondary candidate region.

12. A main body recognition device, characterized in that, Includes: An acquisition module for acquiring an image to be recognized; A processing module, configured to perform region recognition processing on the image to be recognized to obtain N candidate regions, and perform feature extraction processing on the image to be recognized to obtain a saliency feature map, a depth semantic embedding map, and a blurriness feature map, and splice the saliency feature map, the depth semantic embedding map, and the blurriness feature map in the channel dimension to obtain a target feature map, where each candidate region corresponds to a candidate subject, and N is an integer greater than or equal to 1; The obtaining module is further configured to obtain a subject score corresponding to each of the N candidate regions in the N candidate regions according to the target feature map and the N candidate regions; A determining module, configured to determine a target candidate region from the N candidate regions according to the subject score corresponding to each candidate region, and use the candidate subject corresponding to the target candidate region as the target subject in the image to be recognized, where the target candidate region corresponds to the largest subject score.

13. The subject recognition device according to claim 12, wherein The processing module is specifically configured to: Based on the image to be recognized, obtain the N candidate regions through a face detection network, where the candidate subject corresponding to each candidate region is a face; Or, Based on the image to be recognized, obtain the N candidate regions through a human body detection network, where the candidate subject corresponding to each candidate region is a human body.

14. The subject recognition device according to claim 12, characterized in that, The processing module is specifically configured to: Based on the image to be recognized, obtain a saliency feature map through a first network included in the feature extraction network, where the saliency feature map corresponds to 1 channel; Based on the image to be recognized, obtain a depth semantic embedding map through a second network included in the feature extraction network, where the depth semantic embedding map corresponds to C channels, and C is an integer greater than 1; Based on the image to be recognized, obtain a blurriness feature map through a third network included in the feature extraction network, where the blurriness feature map corresponds to 1 channel; Wherein, the target feature map includes C + 2 channels.

15. The subject identification device according to any one of claims 12 to 14, characterized in that The obtaining module is specifically configured to: Perform matching processing on each of the N candidate regions and the target feature map to obtain N spatial feature maps, where there is a one-to-one correspondence between the spatial feature maps and the candidate regions; For each of the N candidate regions, based on the spatial feature map corresponding to each candidate region, obtain a first image feature through a first convolutional network included in the subject selection network; For each of the N candidate regions, based on the spatial feature map corresponding to each candidate region, obtain a second image feature through a second convolutional network included in the subject selection network; For each of the N candidate regions, perform splicing processing on the first image feature and the second image feature corresponding to each candidate region to obtain a comprehensive image feature corresponding to each candidate region; Based on the comprehensive image features corresponding to each candidate region, obtain the subject scores corresponding to each candidate region through the fully connected layer included in the subject selection network.

16. The subject recognition device according to claim 15, characterized in that, The device further includes: an update module; The obtaining module is further configured to obtain a training image sample, where the training image sample includes a first labeled region and a second labeled region, the first labeled region corresponds to a first labeled score, and the second labeled region corresponds to a second labeled score; The processing module is further configured to perform feature extraction processing on the training image sample to obtain a target feature map of the training image sample; The processing module is further configured to perform matching processing between the first labeled region and the target feature map to obtain a first spatial feature map, and perform matching processing between the second labeled region and the target feature map to obtain a second spatial feature map; The obtaining module is further configured to obtain a first predicted image feature based on the first spatial feature map through a first convolutional network included in the training subject selection network, and obtain a second predicted image feature based on the second spatial feature map through a second convolutional network included in the training subject selection network; The obtaining module is further configured to obtain a third predicted image feature based on the first labeled region through the second convolutional network included in the training subject selection network, and obtain a fourth predicted image feature based on the second labeled region through the second convolutional network included in the training subject selection network; The processing module is further configured to perform splicing processing on the first predicted image feature and the third predicted image feature corresponding to the first labeled region to obtain a first comprehensive image feature corresponding to the first labeled region, and perform splicing processing on the second predicted image feature and the fourth predicted image feature corresponding to the second labeled region to obtain a second comprehensive image feature corresponding to the second labeled region; The obtaining module is further configured to obtain a first predicted subject score corresponding to the first labeled region based on the first comprehensive image feature through the fully connected layer included in the training subject selection network, and obtain a second predicted subject score corresponding to the second labeled region based on the second comprehensive image feature through the fully connected layer included in the training subject selection network; The update module is configured to update the model parameters of the training subject selection network according to the first predicted subject score, the second predicted subject score, the first labeled score, and the second labeled score until the model training condition is satisfied, and output the subject selection network.

17. The subject recognition device according to any one of claims 12 to 14, characterized in that, The obtaining module is specifically configured to: For each of the N candidate regions, crop from the target feature map the feature map corresponding to each candidate region; Perform global average pooling processing on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1; Determine the main score corresponding to each candidate region according to the region features corresponding to each candidate region.

18. The subject recognition device according to claim 17, wherein The obtaining module is specifically configured to: Determine the average feature value corresponding to each candidate region according to the region features corresponding to each candidate region, where the average feature value is the average of M feature values in the region features; For each of the N candidate regions, use the average feature value of the candidate region as the main score of the candidate region; Or, The obtaining module is specifically configured to: Determine the first average feature value corresponding to each candidate region according to the first region feature and the first region weight corresponding to each candidate region; Determine the second average feature value corresponding to each candidate region according to the second region feature and the second region weight corresponding to each candidate region, where the second region weight is greater than the first region weight; Determine the target average value corresponding to each candidate region according to the first average feature value and the second average feature value corresponding to each candidate region; For each of the N candidate regions, use the target average value of the candidate region as the main score of the candidate region.

19. The subject recognition device according to any one of claims 12 to 14, characterized in that, The obtaining module is specifically configured to: For each of the N candidate regions, crop the feature map corresponding to each candidate region from the target feature map; Perform global average pooling processing on the feature map corresponding to each candidate region to obtain the region feature corresponding to each candidate region, where the region feature includes M feature values, and M is an integer greater than 1; Based on the region features corresponding to each candidate region, obtain the main score corresponding to each candidate region through a multi-layer perceptron.

20. The subject recognition device according to claim 19, characterized in that, The apparatus further includes: an updating module; The obtaining module is further configured to obtain a training image sample, where the training image sample includes a first labeled region and a second labeled region, the first labeled region corresponds to a first labeled score, and the second labeled region corresponds to a second labeled score; The processing module is further configured to perform feature extraction processing on the training image sample to obtain the target feature map of the training image sample; The processing module is further configured to crop the feature map corresponding to the first labeled region and the feature map corresponding to the second labeled region from the target feature map of the training image sample; The processing module is further configured to perform global average pooling processing on the feature map corresponding to the first labeled region to obtain the region feature corresponding to the first labeled region; The processing module is further configured to perform global average pooling processing on the feature map corresponding to the second labeled region to obtain the region feature corresponding to the second labeled region; The obtaining module is further configured to obtain the first predicted main score corresponding to the first labeled region through a multi-layer perceptron to be trained based on the region feature corresponding to the first labeled region; The obtaining module is further configured to obtain a second predicted subject score corresponding to the second labeled area through the multi-layer perceptron to be trained based on the area feature corresponding to the second labeled area. The updating module is configured to update model parameters of the multi-layer perceptron to be trained according to the first predicted subject score, the second predicted subject score, the first labeled score, and the second labeled score until a model training condition is met, and output the multi-layer perceptron.

21. The subject recognition device according to any one of claims 12 to 14, characterized in that, The obtaining module is specifically configured to: Perform frame splitting on the video to be processed to obtain K video frames, where K is an integer greater than 1. For each of the K video frames, obtain brightness, clarity, and color uniformity. Perform frame filtering on the K video frames according to the brightness of each video frame and a brightness threshold, the clarity of each video frame and a clarity threshold, and the color uniformity of each video frame and a color uniformity threshold to obtain L video frames, where L is an integer greater than or equal to 1 and less than K. Select one video frame from the L video frames as the image to be recognized. The obtaining module is further configured to: after determining a target candidate area from the N candidate areas according to the subject score corresponding to each candidate area, and taking the candidate subject corresponding to the target candidate area as the target subject in the image to be recognized, crop a target image from the image to be recognized according to the target candidate area, where the target image corresponds to a target size and the target image includes the target candidate area. The determining module is further configured to use the target image as a video cover of the video to be processed.

22. The subject recognition device according to claim 21, characterized in that, The obtaining module is specifically configured to: Determine a secondary candidate area from the N candidate areas according to the subject score corresponding to each candidate area, where the secondary candidate area corresponds to the second largest subject score. If the subject score corresponding to the secondary candidate area is greater than or equal to a subject score threshold, obtain a first position coordinate of the center of the secondary candidate area in the image to be recognized, and obtain a second position coordinate of the center of the target candidate area in the image to be recognized. Determine a regional distance according to the first position coordinate and the second position coordinate. If the regional distance is less than or equal to a distance threshold corresponding to the target size, crop a target image from the image to be recognized according to the target candidate area and the secondary candidate area, where the target image further includes the secondary candidate area.

23. A computer device, characterized in that, Including: A memory, a processor, and a bus system. Wherein, the memory is used for storing programs. The processor is configured to execute the programs in the memory, and the processor is configured to execute the subject recognition method according to any one of claims 1 to 11 based on instructions in the program code. The bus system is configured to connect the memory and the processor to enable communication between the memory and the processor.

24. A computer-readable storage medium includes instructions that, when run on a computer, cause the computer to execute the subject identification method according to any one of claims 1 to 11.

25. A computer program product, characterized in that, The computer program product includes computer instructions that are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the subject identification method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • A remote sensing image scene classification method based on the fusion of depth features and saliency features

    CN109165682A

  • Target detection method and device based on artificial intelligence

    CN112052837A

  • Fine-grained image recognition method, convolutional neural network and training method thereof

    CN112257758A