Image-based subject recognition method and device, storage medium and electronic equipment
By using an end-to-end subject recognition model, which combines feature extraction and classification subnetworks with graph convolutional subnetworks, the main objects in images can be identified. This solves the problem of poor generalization of the label + clustering method and achieves high-accuracy recognition in different scenarios.
Patent Information
- Application Number
- CN202210220533.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2042-03-08
AI Technical Summary
Existing label-based and clustering-based object recognition methods have poor generalization in different scenarios, require cumbersome hyperparameter tuning, and lack sufficient recognition accuracy.
An end-to-end subject recognition model is adopted. By combining a pre-trained feature extraction subnetwork and a classification subnetwork with a graph convolution subnetwork, the multimodal features of candidate detection boxes are identified to determine the main object in the image without the need to set hyperparameters for different scenes.
It improves the accuracy and generalization ability of object recognition, simplifies the hyperparameter adjustment process, and is suitable for image recognition in various scenarios.
Smart Images

Figure CN114581714B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of electronic information technology, and in particular, to an image-based subject identification method and device, a storage medium and an electronic device. BACKGROUND
[0002] For an object retrieval system with a search function, an object library needs to be accurately established. The object retrieval system queries information corresponding to a search request input by a user terminal from the object library and feeds back to the user terminal.
[0003] For a scenario of establishing an object library based on a picture, the picture may include both a subject object and a non-subject object, and the subject object needs to be accurately identified. The features of the subject object are written into the object library so as to accurately feed back search results to the user terminal based on the object library in the future. At present, the subject object in the picture is identified by a label + clustering method. However, the effect of the clustering algorithm is greatly affected by hyperparameters, and the overall generalization is poor. Different hyperparameters need to be set for different types of objects, but this process is relatively cumbersome. SUMMARY
[0004] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0005] In a first aspect, the present disclosure provides an image-based subject identification method, which comprises:
[0006] obtaining an image to be identified and object features used to describe subject objects in the image to be identified, the image to be identified including a plurality of candidate detection boxes, each candidate detection box including an object;
[0007] According to the image to be identified and the object features, a multi-modal feature of the candidate detection box is obtained through a feature extraction subnetwork in a pre-trained subject identification model.
[0008] According to the multi-modal feature of the candidate detection box, a classification result of the object in the candidate detection box is determined through a classification subnetwork in the subject identification model, the classification result being used to represent whether the object in the candidate detection box is a subject object in the image to be identified or not.
[0009] In a second aspect, the present disclosure provides an image-based subject identification device, which comprises:
[0010] The first obtaining module is configured to obtain a to-be-identified image and object features used for describing a subject object in the to-be-identified image, wherein the to-be-identified image comprises a plurality of candidate bounding boxes, and each candidate bounding box comprises an object.
[0011] The first determining module is configured to obtain, according to the to-be-identified image and the object features, multi-modal features of the candidate bounding boxes by using a feature extraction subnetwork in a pre-trained subject recognition model.
[0012] The second determining module is configured to determine, according to the multi-modal features of the candidate bounding boxes, a classification result of the object in the candidate bounding box by using a classification subnetwork in the subject recognition model, wherein the classification result is used to represent whether the object in the candidate bounding box is the subject object in the to-be-identified image or not.
[0013] In a third aspect, the present disclosure provides a computer readable medium, which stores a computer program, and the program is executed by a processing device to implement the steps of the subject recognition method in the first aspect.
[0014] In a fourth aspect, the present disclosure provides an electronic device, which comprises:
[0015] A storage device, which stores a computer program;
[0016] A processing device, which is configured to execute the computer program in the storage device to implement the steps of the subject recognition method in the first aspect.
[0017] According to the above technical solution, the to-be-identified image and the object features used for describing a subject object in the to-be-identified image are obtained, the to-be-identified image comprises a plurality of candidate bounding boxes, and each candidate bounding box comprises an object; according to the to-be-identified image and the object features, multi-modal features of the candidate bounding boxes are obtained by using a feature extraction subnetwork in a pre-trained subject recognition model; according to the multi-modal features of the candidate bounding boxes, a classification result of the object in the candidate bounding box is determined by using a classification subnetwork in the subject recognition model, and the classification result is used to represent whether the object in the candidate bounding box is the subject object in the to-be-identified image or not; and the subject object in the to-be-identified image is recognized based on the multi-modal features of the candidate bounding boxes by using an end-to-end subject recognition model, without the need to set different hyperparameters for different scenes.
[0018] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0019] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings. The same or similar components have the same or similar reference labels. It should be understood that the drawings are not necessarily to scale, with emphasis instead being placed upon illustrating the principles of the embodiments of the present disclosure. In the drawings:
[0020] Figure 1 is a flowchart of an image-based subject recognition method according to an exemplary embodiment of the present disclosure.
[0021] Figure 2 is a flowchart of determining multi-modal features of a candidate detection box according to an exemplary embodiment of the present disclosure.
[0022] Figure 3 is another flowchart of an image-based subject recognition method according to an exemplary embodiment of the present disclosure.
[0023] Figure 4 is a flowchart of determining fusion features of a candidate detection box according to an exemplary embodiment of the present disclosure.
[0024] Figure 5 is another flowchart of determining fusion features of a candidate detection box according to an exemplary embodiment of the present disclosure.
[0025] Figure 6 is another flowchart of determining fusion features of a candidate detection box according to an exemplary embodiment of the present disclosure.
[0026] Figure 7 is a block diagram of an image-based subject recognition apparatus according to an exemplary embodiment of the present disclosure.
[0027] Figure 8 is a structural schematic diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several embodiments of the present disclosure have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the present disclosure. It is to be understood that the drawings and descriptions are not to be taken literally, and that the present disclosure is not to be limited by the drawings and / or descriptions.
[0029] It should be understood that each step recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.
[0030] The term "comprising" and variations thereof as used herein are open-ended, that is, "comprising but not limited to." The term "based on" is "based, at least in part, on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related terms have analogous meanings.
[0031] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0032] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative rather than limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0033] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0034] As mentioned in the background, the subject object and the non-subject object can be included in the picture at the same time, so it is necessary to accurately identify the subject object of the picture, and then establish an object library based on the characteristics of the subject object, so as to accurately feed back the search result to the user terminal based on the object library in the subsequent. The label + clustering method is a non-end-to-end subject identification method, so the generalization of this method is poor. In order to improve the identification accuracy, for different scenes (for example, clothing scene, makeup scene), it is usually necessary to adjust the hyperparameters of different clustering algorithms, but the process of adjusting the hyperparameters of different clustering algorithms is relatively cumbersome.
[0035] Therefore, the embodiments of the present disclosure provide a subject identification method and device based on images, a storage medium and an electronic device, which identify the subject object in the image to be identified based on the multi-modal features of the candidate detection frame through an end-to-end subject identification model, without setting different hyperparameters for different scenes.
[0036] The present disclosure will be further explained and described in conjunction with the accompanying drawings.
[0037] Figure 1is a flowchart of an image-based subject recognition method according to an example embodiment of the present disclosure. The subject recognition method can be applied to an electronic device, with reference to Figure 1 The subject recognition method can include the following steps:
[0038] In step S101, an image to be recognized and an object feature used to describe a subject object in the image to be recognized are obtained. The image to be recognized includes a plurality of candidate detection boxes, and each candidate detection box includes an object.
[0039] In some embodiments, the object library can include features of the subject object and information (such as introduction links, purchase links) corresponding to the features. For example, taking the object as a commodity as an example, in a search scenario, the image to be recognized can be a frame image in a video played by a user terminal. The subject object in the image can be recognized based on the frame image in the video. In the case where the subject object in the image is recognized, the feature matching the feature of the subject object in the image can be queried from the object library, and the information corresponding to the feature can be fed back to the user terminal. The information can be an introduction link, a purchase link, or the like of the object. In the scenario of establishing the object library, the image to be recognized can be an image provided by a merchant. The subject object in the image provided by the merchant needs to be recognized so as to accurately establish the feature describing the subject object in the object library.
[0040] In some embodiments, the object feature can include a text feature extracted from a title text corresponding to the image to be recognized, and the object feature can also include a category feature describing the subject object in the image to be recognized. In this case, Figure 1 The step of obtaining the object feature used to describe the subject object in the image to be recognized in step S101 in the method can include: obtaining a title text corresponding to the image to be recognized; determining a first feature corresponding to the title text by a pre-trained text feature extraction model according to the title text; determining a second feature corresponding to the title text and the image to be recognized by a pre-trained category feature extraction model according to the title text and the image to be recognized; and determining the object feature used to describe the subject object in the image to be recognized according to the first feature and the second feature.
[0041] In the present embodiment, the text feature extraction model can be used to implement keyword extraction. It can be understood that the title text generally carries keyword information, which can be used to describe the subject object in the image to be recognized. Therefore, the keyword in the title text can be extracted as the first feature by using the text feature extraction model. For example, if the title carries a text such as clothing, it can be represented that the image to be recognized can be an image showing clothing, and the clothing object can be the subject object in the image to be recognized.
[0042] In this embodiment, the category feature extraction model belongs to a multi-modal model, which can predict the category of the subject object in the to-be-identified image according to the visual information of the to-be-identified image and the text information in the title text, and then obtain the category feature and take the category feature as the second feature. For example, when the category feature is a clothing category, it can be represented that the to-be-identified image is likely to be an image showing clothing, and the clothing object can be taken as the subject object in the to-be-identified image.
[0043] In some embodiments, the to-be-identified image can be detected by a detection model to obtain the position information of the bounding box, and the bounding box includes a target, and the bounding box corresponds to the candidate detection box in the present disclosure, and the target corresponds to the object in the present disclosure. Wherein, the detection model can refer to related technologies, and this embodiment will not be repeated here.
[0044] In step S102, the multi-modal feature of the candidate detection box is obtained by the feature extraction sub-network in the pre-trained subject recognition model according to the to-be-identified image and the object feature.
[0045] It should be noted that the to-be-identified image can reflect visual information. In the above examples of the object feature, the object feature can reflect visual information and text information, so the multi-modal information based on vision and text can be obtained based on the to-be-identified image and the object feature.
[0046] In some embodiments, the feature extraction sub-network can include a visual extraction layer, an ROI Pooling layer and a feature splicing layer, in which case, step S102 can include: extracting the visual feature of the to-be-identified image by the visual extraction layer in the feature extraction sub-network in the pre-trained subject recognition model according to the to-be-identified image; determining the visual feature corresponding to the candidate detection box by the ROI Pooling layer in the feature extraction sub-network according to the visual feature; and determining the multi-modal feature of the candidate detection box by the feature splicing layer in the feature extraction sub-network according to the visual feature corresponding to the candidate detection box and the object feature.
[0047] For example, referring to Figure 2 , the feature extraction sub-network can include a visual extraction layer, an ROI Pooling layer and a first feature splicing layer (to distinguish the feature splicing layer in the graph convolution sub-network in the following embodiments), specifically, the visual extraction layer extracts the overall visual feature of the to-be-identified image according to the to-be-identified image, extracts the visual feature corresponding to the candidate detection box according to the overall visual feature and the position information of the candidate detection box through the ROI Pooling layer, and splices the visual feature corresponding to the candidate detection box and the object feature through the first feature splicing layer to obtain the multi-modal feature of the candidate detection box output by the first feature splicing layer.
[0048] In this embodiment, the visual extraction layer can be ResNet-50, which is a neural network for visual tasks. The implementation principle of ResNet-50 can refer to related technologies, and will not be described here in this embodiment. The input of the ROI (Region of Interest, frame on the feature map) pooling layer includes a feature map and a candidate detection frame. The ROI pooling layer maps the feature map according to the position information of the candidate detection frame to obtain the feature map corresponding to the candidate detection frame. The implementation principle of ResNet-50 can refer to related technologies, and will not be described here in this embodiment. The first feature splicing layer can be a contat layer in a neural network. The implementation principle of the contat layer can refer to related technologies, and will not be described here in this embodiment.
[0049] Through the above-mentioned feature extraction sub-network, the multi-modal features of the candidate detection frame including visual features and text features can be obtained.
[0050] In step S103, according to the multi-modal features of the candidate detection frame, the classification result of the object in the candidate detection frame is determined through the classification sub-network in the subject recognition model. The classification result is used to represent whether the object in the candidate detection frame is a subject object in the to-be-recognized image or not.
[0051] In some embodiments, the classification sub-network can be a binary classification network, which is used to predict whether the object in the candidate frame is a subject object according to the multi-modal features of the candidate frame.
[0052] Specifically, the classification sub-network first outputs the probability that the object in the candidate detection frame belongs to the subject object in the to-be-recognized image. According to the size of the probability and the classification threshold, the classification result is further output. For example, if the classification threshold is 0.5, the probability that the object in the candidate detection frame belongs to the subject object in the to-be-recognized image is greater than or equal to 0.5, and the output classification result is that the object in the candidate detection frame belongs to the subject object in the to-be-recognized image. The probability that the object in the candidate detection frame belongs to the subject object in the to-be-recognized image is less than 0.5, and the output classification result is that the object in the candidate detection frame does not belong to the subject object in the to-be-recognized image.
[0053] In some embodiments, the classification sub-network can use a sigmoid activation function for processing. The sigmoid activation function converts the multi-modal feature activation into the probability that the object in the candidate detection frame belongs to the subject object in the to-be-recognized image.
[0054] By the above manner, the multi-modal features of the candidate detection box including visual features and text features are obtained by using the feature extraction sub-network, and whether the object in the candidate detection box belongs to the subject object of the to-be-identified image is predicted by the classification sub-network according to the multi-modal features of the candidate detection box, so that whether the object in the candidate detection box belongs to the subject object of the to-be-identified image is predicted in an end-to-end manner, without the need to adjust parameters for different object scenes.
[0055] Figure 3 is another flow chart of an image-based subject identification method according to an example embodiment of the present disclosure. Referring to Figure 3 , the method comprises the following steps:
[0056] In step S301, a to-be-identified image and an object feature used to describe a subject object in the to-be-identified image are obtained.
[0057] In step S302, multi-modal features of a candidate detection box are obtained by a feature extraction sub-network in a pre-trained subject identification model according to the to-be-identified image and the object feature.
[0058] The implementation process of step S301 can refer to the implementation process of step S101 in Figure 1 , and the implementation process of step S302 can refer to the implementation process of step S102 in Figure 1 . This embodiment will not be described here.
[0059] In step S303, a fusion feature of the candidate detection box is determined by a graph convolution sub-network according to the multi-modal features of the candidate detection box and the multi-modal features of other candidate detection boxes except the candidate detection box.
[0060] In some embodiments, for the graph convolution sub-network, graph data of each candidate detection box in the to-be-identified image can be established, in which each candidate detection box is taken as a node, and any two candidate detection boxes are connected by an edge. The graph convolution sub-network can determine other candidate detection boxes of a candidate detection box according to the graph data, so as to determine the fusion feature of the candidate detection box.
[0061] In some embodiments, the graph convolution sub-network can include a mean feature calculation layer and a feature splicing layer. In this case, step S303 can include: determining mean features of other candidate detection boxes by the mean feature calculation layer in the graph convolution sub-network according to the multi-modal features of the other candidate detection boxes; and determining the fusion feature of the candidate detection box by the feature splicing layer in the graph convolution sub-network according to the mean features of the other candidate detection boxes and the multi-modal features of the candidate detection box.
[0062] Refer to Figure 4, the graph convolution subnetwork can include a mean feature calculation layer and a second feature concatenation layer (to distinguish from the feature concatenation layer in the feature extraction subnetwork described above), and the to-be-identified image includes N candidate bounding boxes, wherein, Figure 4 The process for calculating the fusion feature of the candidate bounding box 1. Specifically, the mean feature calculation layer calculates the mean feature of the multi-modal features of the candidate bounding boxes 2 to N according to the multi-modal features of the candidate bounding boxes 2 to N, and concatenates the mean feature through the second feature concatenation layer, so as to obtain the multi-modal feature of the candidate bounding box 1.
[0063] In some embodiments, referring to Figure 5 , the graph convolution subnetwork can also include a fully connected layer, which is located after the feature concatenation layer in the graph convolution subnetwork, and can be used for dimension reduction operation on the feature output by the second feature concatenation layer, so as to output the multi-modal feature after dimension reduction. The classification subnetwork in the subject recognition model determines the classification result of the object in the candidate bounding box according to the multi-modal feature after dimension reduction.
[0064] Step S304, according to the fusion feature of the candidate bounding box, the classification subnetwork in the subject recognition model is used to determine the classification result of the object in the candidate bounding box.
[0065] It should be noted that the relevance feature of the objects between the candidate bounding boxes also affects the recognition of the subject object. Therefore, by fusing the features of multiple objects in the to-be-identified image through the graph convolution subnetwork to obtain the fusion feature, the classification result of the candidate bounding box output by the classification subnetwork depends not only on the feature of the candidate bounding box itself, but also on the features of other candidate bounding boxes except the candidate bounding box, so as to improve the accuracy of the subject object recognition.
[0066] In some embodiments, the graph convolution subnetwork can include multiple, multiple graph convolution subnetworks are connected in series, the output of the first graph convolution subnetwork in the adjacent two graph convolution subnetworks is taken as the multi-modal feature of the candidate bounding box input into the next graph convolution subnetwork, and the output of the last graph convolution subnetwork in the multiple graph convolution subnetworks is taken as the fusion feature of the candidate bounding box.
[0067] Referring to Figure 6For example, the graph convolution subnetworks include three graph convolution subnetworks 1, 2 and 3, and the fusion feature of the candidate bounding box 1 is determined by the multi-modal features of the candidate bounding boxes 1 to N through the graph convolution subnetworks 1, 2 and 3. Specifically, the candidate bounding boxes 1 to N are taken as the output of the graph convolution subnetwork 1 to obtain the fusion feature 1 output by the graph convolution subnetwork 1, the fusion feature 1 is taken as the new multi-modal feature of the candidate bounding box 1, and the new multi-modal feature of the candidate bounding box 1 and the candidate bounding boxes 2 to N are taken as the input of the graph convolution subnetwork 2 to obtain the fusion feature 2 output by the graph convolution subnetwork 2, the fusion feature 2 is taken as the new multi-modal feature of the candidate bounding box 1, and the new multi-modal feature of the candidate bounding box 1 and the candidate bounding boxes 2 to N are taken as the input of the graph convolution subnetwork 3 to obtain the fusion feature 3 output by the graph convolution subnetwork 3, and the fusion feature 3 output by the graph convolution subnetwork 3 can be taken as the input of the classification subnetwork to predict the classification result of the candidate bounding box 1.
[0068] wherein, Figure 6 The implementation process of each graph convolution subnetwork in the above formula can refer to the process of the above formula, and details are not described herein. Figure 4 or Figure 5 The implementation process of each graph convolution subnetwork in the above formula can refer to the process of the above formula, and details are not described herein.
[0069] In the above manner, multiple graph convolution networks are set to extract deeper feature information, so as to improve the classification accuracy of the classification subnetwork, and further improve the accuracy of the subject recognition model. The setting of three graph convolution subnetworks can well balance the model calculation amount and the accuracy, so that the model calculation amount and the accuracy can be optimized.
[0070] In some embodiments, the subject recognition model is obtained by the following method: obtaining a sample image and a sample object feature used to describe a sample subject object in the sample image, the sample image carrying a label corresponding to each sample candidate bounding box, the label of the sample candidate bounding box being used to represent whether the sample object in the sample candidate bounding box belongs to a subject sample object in the sample image; and training an initial model according to the sample image to obtain the subject recognition model.
[0071] It should be noted that, in the case of obtaining a subject recognition model including a graph convolution subnetwork, the model structure of the initial model includes a feature extraction subnetwork, a graph convolution subnetwork and a classification subnetwork. In the case of obtaining a subject recognition model not including a graph convolution subnetwork, the model structure of the initial model includes a feature extraction subnetwork and a classification subnetwork. In different cases, the subnetworks in the initial model can refer to the above related structures, and details are not described herein.
[0072] In some embodiments, the loss function of the initial model can be a cross-entropy loss function, which can represent the difference between the labels of each sample candidate detection box in the sample image and the predicted labels of each sample candidate detection box output by the initial model. The cross-entropy loss function can refer to related technologies, and this embodiment will not be described here.
[0073] In some embodiments, the initial model is iteratively trained according to the sample image, and the iteration stopping condition can be that the value of the cross-entropy loss function is less than a preset threshold or the number of iterations of the initial model is greater than a preset number.
[0074] Figure 7 is a block diagram of an image-based subject recognition device according to an exemplary embodiment of the present disclosure, which comprises:
[0075] The first acquisition module 701 is configured to acquire a to-be-recognized image and object features used to describe a subject object in the to-be-recognized image, wherein the to-be-recognized image comprises a plurality of candidate detection boxes, and each candidate detection box comprises an object.
[0076] The first determination module 702 is configured to obtain multi-modal features of the candidate detection boxes by a feature extraction subnetwork in a pre-trained subject recognition model according to the to-be-recognized image and the object features.
[0077] The second determination module 703 is configured to determine a classification result of the object in the candidate detection box by a classification subnetwork in the subject recognition model according to the multi-modal features of the candidate detection box, wherein the classification result is used to represent whether the object in the candidate detection box is a subject object in the to-be-recognized image or not.
[0078] Optionally, the subject recognition device 700 further comprises:
[0079] The fusion module is configured to determine the fusion features of the candidate detection box by the graph convolution subnetwork according to the multi-modal features of the candidate detection box and the multi-modal features of other candidate detection boxes except the candidate detection box.
[0080] The second determination module 703 is configured to determine the classification result of the object in the candidate detection box by the classification subnetwork in the subject recognition model according to the fusion features of the candidate detection box.
[0081] Optionally, the fusion module comprises:
[0082] a mean feature submodule configured to determine, according to the multi-modal features of other candidate bounding boxes except the candidate bounding box, mean features of the other candidate bounding boxes by a mean feature calculation layer in the graph convolutional subnetwork;
[0083] a fusion submodule configured to determine, according to the mean features of the other candidate bounding boxes and the multi-modal features of the candidate bounding box, fusion features of the candidate bounding box by a feature concatenation layer in the graph convolutional subnetwork.
[0084] Optionally, the first determination module 702 includes:
[0085] a first visual feature extraction submodule configured to extract, according to the to-be-recognized image, visual features of the to-be-recognized image by a visual extraction layer in a feature extraction subnetwork in a pre-trained subject recognition model;
[0086] a second visual feature extraction submodule configured to determine, according to the visual features, visual features corresponding to the candidate bounding box by an ROI Pooling layer in the feature extraction subnetwork;
[0087] a first determination submodule configured to determine, according to the visual features corresponding to the candidate bounding box and the object features, multi-modal features of the candidate bounding box by a feature concatenation layer in the feature extraction subnetwork.
[0088] Optionally, the graph convolutional subnetwork includes a plurality of graph convolutional subnetworks connected in series, the output of a previous graph convolutional subnetwork in two adjacent graph convolutional subnetworks is taken as input of the multi-modal features of the candidate bounding box in a next graph convolutional subnetwork, and the output of a last graph convolutional subnetwork in the plurality of graph convolutional subnetworks is taken as the fusion features of the candidate bounding box.
[0089] Optionally, the first acquisition module 701 includes:
[0090] a text acquisition submodule configured to acquire corresponding title text for describing the to-be-recognized image;
[0091] a first feature determination submodule configured to determine, according to the title text, first features corresponding to the title text by a pre-trained text feature extraction model;
[0092] a second feature determination submodule configured to determine, according to the title text and the to-be-recognized image, second features corresponding to the title text and the to-be-recognized image by a pre-trained category feature extraction model;
[0093] The object feature determination sub-module is configured to determine an object feature for describing the subject object in the image to be recognized according to the first feature and the second feature.
[0094] Optionally, the subject recognition apparatus 700 further comprises:
[0095] The second acquisition module is configured to acquire a sample image and a sample object feature for describing a sample subject object in the sample image, the sample image carrying a label corresponding to each sample candidate detection frame, the label of the sample candidate detection frame being used to represent whether the sample object in the sample candidate detection frame belongs to a subject sample object in the sample image.
[0096] The training module is configured to train an initial model according to the sample image to obtain the subject recognition model.
[0097] Reference will be made to the following description Figure 8 , which shows a structural schematic diagram of an electronic device 800 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. Figure 8 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0098] As shown in Figure 8 , the electronic device 800 can include a processing device (such as a central processor, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 802 or loaded into a random access memory (RAM) 803 from a storage device 808. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0099] Generally, the following devices can be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 808 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 809. The communication devices 809 can allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 8The electronic device 800 is illustrated with various means, but it is to be understood that not all of the illustrated means need be present in every embodiment. A greater or lesser number of means can alternatively be implemented.
[0100] In particular, in accordance with embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program comprising program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.
[0101] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is carried. Such a propagated data signal can take a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), etc., or any suitable combination thereof.
[0102] In some embodiments, the electronic device can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., a communications network) of any form or medium, such as the Internet. Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.
[0103] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled into the electronic device.
[0104] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire a to-be-recognized image and an object feature used to describe a subject object in the to-be-recognized image, the to-be-recognized image including a plurality of candidate detection boxes, each of the candidate detection boxes including an object; obtain, according to the to-be-recognized image and the object feature, a multi-modal feature of the candidate detection boxes through a feature extraction subnetwork in a pre-trained subject recognition model; and determine, according to the multi-modal feature of the candidate detection boxes, a classification result of the object in the candidate detection box through a classification subnetwork in the subject recognition model, the classification result being used to represent whether the object in the candidate detection box is the subject object in the to-be-recognized image or not.
[0105] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a number of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a computer system, includes electronic components, for example, a processor, that are configured to carry out the operations of the present disclosure.
[0106] The computer program product of the first aspect can include one or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform the operations of the first aspect. The one or more non-transitory computer-readable media can include, for example, magnetic media such as one or more magnetic disks, magnetic tapes or cassettes; optical media such as one or more compact discs, optical discs or Blu-ray discs; magneto-optical media such as one or more floptical discs; solid state media such as one or more solid state drives or other flash memory arrays; or any suitable combination of these. The one or more non-transitory computer-readable media can be encoded with instructions that, when executed, cause one or more processors to perform the operations of the first aspect.
[0107] The modules described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, the first obtaining module can also be described as a module that obtains a to-be-identified image and object features used to describe a subject object in the to-be-identified image.
[0108] The functions described in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0109] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0110] According to one or more embodiments of the present disclosure, example 1 provides a subject recognition method based on images, the subject recognition method comprising:
[0111] obtaining an image to be recognized and object features used to describe subject objects in the image to be recognized, the image to be recognized comprising a plurality of candidate bounding boxes, each of the candidate bounding boxes comprising an object;
[0112] According to the image to be recognized and the object features, obtaining multi-modal features of the candidate bounding boxes through a feature extraction subnetwork in a pre-trained subject recognition model;
[0113] According to the multi-modal features of the candidate bounding boxes, determining a classification result of the object in the candidate bounding box through a classification subnetwork in the subject recognition model, the classification result being used to represent whether the object in the candidate bounding box is a subject object in the image to be recognized or not.
[0114] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, the subject recognition model further comprising a graph convolution subnetwork, the method further comprising:
[0115] According to the multi-modal features of the candidate bounding box and multi-modal features of other candidate bounding boxes except the candidate bounding box, determining a fusion feature of the candidate bounding box through the graph convolution subnetwork;
[0116] The determining, according to the multi-modal features of the candidate bounding box, of the classification result of the object in the candidate bounding box through the classification subnetwork in the subject recognition model comprises:
[0117] The determining, according to the fusion feature of the candidate bounding box, of the classification result of the object in the candidate bounding box through the classification subnetwork in the subject recognition model.
[0118] According to one or more embodiments of the present disclosure, example 3 provides the method of example 2, the determining, according to the multi-modal features of the candidate bounding box and multi-modal features of other candidate bounding boxes except the candidate bounding box, of the fusion feature of the candidate bounding box through the graph convolution subnetwork comprises:
[0119] According to the multi-modal features of the other candidate bounding boxes except the candidate bounding box, determining mean features of the other candidate bounding boxes through a mean feature calculation layer in the graph convolution subnetwork;
[0120] According to the mean features of the other candidate bounding boxes and the multi-modal features of the candidate bounding box, determining the fusion feature of the candidate bounding box through a feature concatenation layer in the graph convolution subnetwork.
[0121] According to one or more embodiments of the present disclosure, example 4 provides the method of example 2, wherein the determining the multi-modal feature of the candidate bounding box from the image to be recognized and the object feature by a feature extraction sub-network in the pre-trained subject recognition model comprises:
[0122] extracting a visual feature of the image to be recognized by a visual extraction layer in the feature extraction sub-network in the pre-trained subject recognition model according to the image to be recognized;
[0123] determining a visual feature corresponding to the candidate bounding box by an ROI Pooling layer in the feature extraction sub-network according to the visual feature;
[0124] determining the multi-modal feature of the candidate bounding box by a feature concatenation layer in the feature extraction sub-network according to the visual feature corresponding to the candidate bounding box and the object feature.
[0125] According to one or more embodiments of the present disclosure, example 5 provides the method of example 2, wherein the graph convolution sub-network comprises a plurality of graph convolution sub-networks connected in series, an output of a previous graph convolution sub-network in two adjacent graph convolution sub-networks is taken as an input to the multi-modal feature of the candidate bounding box in a next graph convolution sub-network, and an output of a last graph convolution sub-network in the plurality of graph convolution sub-networks is taken as the fusion feature of the candidate bounding box.
[0126] According to one or more embodiments of the present disclosure, example 6 provides the method of example 1, wherein the obtaining the object feature for describing the subject object in the image to be recognized comprises:
[0127] obtaining a title text corresponding to the image to be recognized;
[0128] determining a first feature corresponding to the title text by a pre-trained text feature extraction model according to the title text;
[0129] determining a second feature corresponding to the title text and the image to be recognized by a pre-trained category feature extraction model according to the title text and the image to be recognized;
[0130] determining the object feature for describing the subject object in the image to be recognized according to the first feature and the second feature.
[0131] According to one or more embodiments of the present disclosure, example 7 provides the method of any one of examples 1-6, wherein the subject recognition model is trained by:
[0132] obtaining a sample image and a sample object feature used for describing a sample subject object in the sample image, the sample image carrying a label corresponding to each sample candidate bounding box, the label of the sample candidate bounding box being used to represent whether the sample object in the sample candidate bounding box belongs to a subject sample object in the sample image;
[0133] training an initial model according to the sample image to obtain the subject recognition model.
[0134] According to one or more embodiments of the present disclosure, example 8 provides a subject recognition device based on an image, the subject recognition device comprising:
[0135] a first obtaining module configured to obtain a to-be-recognized image and an object feature used for describing a subject object in the to-be-recognized image, the to-be-recognized image comprising a plurality of candidate bounding boxes, and each candidate bounding box comprising an object;
[0136] a first determining module configured to obtain a multi-modal feature of the candidate bounding box by a feature extraction sub-network in a pre-trained subject recognition model according to the to-be-recognized image and the object feature;
[0137] a second determining module configured to determine a classification result of the object in the candidate bounding box by a classification sub-network in the subject recognition model according to the multi-modal feature of the candidate bounding box, the classification result being used to represent whether the object in the candidate bounding box is the subject object in the to-be-recognized image or not.
[0138] According to one or more embodiments of the present disclosure, example 9 provides a computer readable medium having a computer program stored thereon, the program being executed by a processing device to implement the steps of the subject recognition method in any one of examples 1-7.
[0139] According to one or more embodiments of the present disclosure, example 10 provides an electronic device comprising:
[0140] a storage device having a computer program stored thereon;
[0141] a processing device configured to execute the computer program in the storage device to implement the steps of the subject recognition method in any one of examples 1-7.
[0142] The above description merely illustrates the preferred embodiment of the disclosure and a principle of applied technologies. It should be understood by those skilled in the art that the disclosed range of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.
[0143] Furthermore, although operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order illustrated or sequentially. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are contained in the above discussion, these should not be construed as limiting the scope of the disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0144] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely illustrative of specific forms of implementing the claims. With respect to the devices in the above-described embodiments, the specific manner in which the various modules perform operations has been described in detail in the embodiments related to the method, and will not be described here in detail.
Claims
1. An image-based subject identification method, characterized by, The subject recognition method comprises: obtaining a to-be-recognized image and object features used for describing a subject object in the to-be-recognized image, the to-be-recognized image comprising a plurality of candidate bounding boxes, each of the candidate bounding boxes comprising an object; obtaining, according to the to-be-recognized image and the object features, a multi-modal feature of the candidate bounding boxes by a feature extraction subnetwork in a pre-trained subject recognition model; determining, according to the multi-modal feature of the candidate bounding boxes, a classification result of the object in the candidate bounding box by a classification subnetwork in the subject recognition model, the classification result being used to represent whether the object in the candidate bounding box is a subject object in the to-be-recognized image or not; the obtaining of the to-be-recognized image and the object features used for describing the subject object in the to-be-recognized image comprises: obtaining corresponding title text used for describing the to-be-recognized image; determining, according to the title text, a first feature corresponding to the title text by a pre-trained text feature extraction model; determining, according to the title text and the to-be-recognized image, a second feature corresponding to the title text and the to-be-recognized image by a pre-trained category feature extraction model; and determining, according to the first feature and the second feature, the object features used for describing the subject object in the to-be-recognized image.
2. The subject identification method according to claim 1, characterized by, The subject recognition model further comprises a graph convolution subnetwork, and the method further comprises: determining, according to the multi-modal feature of the candidate bounding box and multi-modal features of other candidate bounding boxes except the candidate bounding box, a fusion feature of the candidate bounding box by the graph convolution subnetwork; the determining of the classification result of the object in the candidate bounding box by the classification subnetwork in the subject recognition model according to the multi-modal feature of the candidate bounding box comprises: determining, according to the fusion feature of the candidate bounding box, the classification result of the object in the candidate bounding box by the classification subnetwork in the subject recognition model.
3. The subject identification method according to claim 2, characterized by, the determining of the fusion feature of the candidate bounding box by the graph convolution subnetwork according to the multi-modal feature of the candidate bounding box and the multi-modal features of other candidate bounding boxes except the candidate bounding box comprises: determining, according to the multi-modal features of the other candidate bounding boxes except the candidate bounding box, mean features of the other candidate bounding boxes by a mean feature calculation layer in the graph convolution subnetwork; determining, according to the mean features of the other candidate bounding boxes and the multi-modal feature of the candidate bounding box, the fusion feature of the candidate bounding box by a feature splicing layer in the graph convolution subnetwork.
4. The subject identification method of claim 2, wherein the determining of the multi-modal feature of the candidate bounding box by the feature extraction subnetwork in the pre-trained subject recognition model according to the to-be-recognized image and the object features comprises: extracting, according to the to-be-recognized image, visual features of the to-be-recognized image by a visual extraction layer in the feature extraction subnetwork in the pre-trained subject recognition model; According to the visual feature, a ROI Pooling layer in the feature extraction sub-network is used to determine the visual feature corresponding to the candidate detection box; According to the visual feature corresponding to the candidate detection box and the object feature, a feature splicing layer in the feature extraction sub-network is used to determine the multi-modal feature of the candidate detection box.
5. The subject identification method according to claim 2, characterized by, The graph convolution sub-network includes multiple graph convolution sub-networks connected in series, the output of a previous graph convolution sub-network in adjacent two graph convolution sub-networks is taken as the input of the next graph convolution sub-network, the multi-modal feature of the candidate detection box, and the output of the last graph convolution sub-network in the multiple graph convolution sub-networks is taken as the fusion feature of the candidate detection box.
6. The subject identification method according to any one of claims 1 to 5, characterized by, The subject recognition model is obtained by training in the following manner: A sample image and a sample object feature used for describing a sample subject object in the sample image are obtained, the sample image carries a label corresponding to each sample candidate detection box, and the label of the sample candidate detection box is used to represent whether the sample object in the sample candidate detection box belongs to a subject sample object in the sample image. An initial model is trained according to the sample image to obtain the subject recognition model.
7. An image-based subject recognition apparatus, characterized by comprising: The subject recognition device includes: A first obtaining module is configured to obtain an image to be recognized and an object feature used for describing a subject object in the image to be recognized, the image to be recognized includes multiple candidate detection boxes, and each candidate detection box includes an object; A first determining module is configured to obtain a multi-modal feature of the candidate detection box by using a feature extraction sub-network in a pre-trained subject recognition model according to the image to be recognized and the object feature; A second determining module is configured to determine a classification result of the object in the candidate detection box by using a classification sub-network in the subject recognition model according to the multi-modal feature of the candidate detection box, and the classification result is used to represent whether the object in the candidate detection box is a subject object in the image to be recognized or not; The first obtaining module includes: A text obtaining sub-module is configured to obtain a title text used for describing the image to be recognized; A first feature determining sub-module is configured to determine a first feature corresponding to the title text by using a pre-trained text feature extraction model according to the title text; A second feature determining sub-module is configured to determine a second feature corresponding to the title text and the image to be recognized by using a pre-trained category feature extraction model according to the title text and the image to be recognized; An object feature determining sub-module is configured to determine an object feature used for describing the subject object in the image to be recognized according to the first feature and the second feature.
8. A computer readable medium having stored thereon a computer program, characterized in that The program is executed by a processing device to implement the steps of the subject recognition method in any one of claims 1-6.
9. An electronic device, comprising: The program is executed by a processing device to implement the steps of the subject recognition method in any one of claims 1-6. The program is executed by a processing device to implement the steps of the subject recognition method in any one of claims 1-6.
Citation Information
Patent Citations
Mixed-pasting bill image processing method, device, computer equipment and storage medium
CN111931664A
Region of interest determination method and image content identification method and device
CN112597997A