An Image Event Recognition Method Based on Multimodal Event Ontology
Through the image event recognition method based on multimodal event ontology, the multi-label classification and image feature matching technology is used to solve the problem of insufficient semantic information understanding in the prior art, and the accuracy of image event recognition is improved.
Patent Information
- Application Number
- CN202210690851.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-06-17
AI Technical Summary
In the prior art, semantic information is insufficient when identifying image events, resulting in low recognition accuracy.
The image event recognition method based on multimodal event ontology is adopted, image keywords are obtained through multi-label classification technology, and text matching is performed by combining the event class six-tuple representation structure to filter out event class sets with high matching degree, and the final event class is determined through image feature matching.
The accuracy of image event recognition is improved, and the probability of matching errors is reduced through the combination of structured information and multimodal technology, and the machine's ability to understand image semantic information is enhanced.
Smart Images

Figure CN114972884B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular, to an image event recognition method based on a multi-modal event ontology. Background Art
[0002] Images are an important auxiliary tool for humans to understand the world. With the rapid development of artificial intelligence technology, the processing of images by machines is no longer limited to simple classification tasks, but gradually focuses on the in-depth understanding and application of image information.
[0003] An event refers to a process that occurs at a specific time and in a specific environment, involves several roles, and exhibits specific actions or state changes. Representing an event in the form of a six-tuple of "object", "action", "time", "environment", "state", and "linguistic expression" can obtain a standardized description of the event.
[0004] Image event recognition mainly identifies the events occurring in an image through image processing technology. Its goal is to describe as detailed as possible the participants (people or objects), environmental information, and event categories in the event, which includes an intuitive judgment based on vision and an auxiliary reasoning process based on common sense. Therefore, in the recognition process, in addition to focusing on the visual features of the image, attention should also be paid to the understanding of its semantic information. It can be said that technologies such as object classification and recognition of images serve semantic understanding.
[0005] An event class refers to a set composed of events of the same or similar types, which is an abstract summary of multiple events. An event ontology refers to a knowledge base that can cover all scenarios obtained by screening and combining multiple related event classes for general or specific domain application scenarios, and combining event class relationships and certain reasoning rules. The event ontology can integrate a large amount of unstructured text events into a more structured form, making the representation form of events clearer.
[0006] Currently, the research community has begun to consider applying multi-modal information to the in-depth understanding process of images. Multi-modal technology is a technology that combines various types of information such as text, images, and speech. Each modality complements each other to improve the understanding ability of machines.
[0007] A multi-modal event ontology is to integrate the multi-modal idea into the event ontology model. Specifically, it uses the "multi-modal information" jointly composed of text and images as one of the elements for event (class) description. Therefore, when performing picture recognition, it can not only enhance the supplement of text semantic information, but also use visual features as an additional auxiliary for event judgment, thereby improving the accuracy of event recognition technology. Therefore, an image event recognition method based on a multi-modal event ontology is needed. Summary of the Invention
[0008] Based on the above problems, the present invention proposes an image event recognition method based on a multi-modal event ontology to solve the problem of insufficient semantic information understanding in the prior art when recognizing image events.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] An image event recognition method based on a multi-modal event ontology includes the following steps:
[0011] Image keyword acquisition: Using multi-label classification technology, obtain important keywords of the input image;
[0012] Screening event class set: Utilize the obtained keywords, and through text matching with the element information in the six-tuple representation structure of the event class, find the event class set with the highest matching degree in the event ontology model;
[0013] Image matching: For the images of all event classes in the screened event class set with high matching degree, match them with the input image, and select the corresponding event class with the highest score, which is the result of the final image event recognition.
[0014] Furthermore, the image keyword acquisition step further includes:
[0015] Image region extraction: Extract the key regions of the image to obtain several sub-images containing the key parts of the image, and these sub-images represent the main information of the image;
[0016] Multi-label classifier: Based on multi-label classification technology, process the sub-images generated in the region extraction technology respectively to obtain the keyword sets corresponding to each regional sub-image;
[0017] Keyword annotation: Perform part-of-speech annotation on the keyword sets of the regional sub-images, and make a new division of the keyword sets according to the part of speech.
[0018] Even further, in the region extraction part, adopt Selective Search or RPN (Region Proposal Network) technology to obtain the representative regions of the image, and make each representative regional sub-image retain only one key target as much as possible.
[0019] Even further, in the multi-label classification part, let the representative regional sub-images pass through the multi-label classification CNN model to obtain the keywords corresponding to the sub-images, and put the keywords generated by each sub-image into different sets to generate an image keyword sequence; in addition, according to the classification summary result, generate attributes such as the total number of objects.
[0020] The multi-label classifier here adopts the HCP (Hypotheses-CNN-Pooling) structure based on hypotheses.
[0021] Furthermore, the step of screening the corresponding event class set further includes:
[0022] Element matching: According to the existing multi-modal event ontology model, match the obtained image keywords with the corresponding event elements, and screen the required event class set;
[0023] External knowledge supplement: Further screen the results of element matching using external knowledge.
[0024] Furthermore, in the element matching part, text matching technologies such as semantic similarity are needed to complete the event element matching process, and generate an event class set with a higher matching degree.
[0025] Furthermore, in the external knowledge supplement part, according to the corpus, semantic dictionary or network resources, etc., calculate the semantic relevance between the text part of the "multi-modal information" element of the image keyword and the event class, and perform a secondary screening on the event class set according to the results.
[0026] Furthermore, the image matching step further includes:
[0027] Feature extraction: Extract the features of the input image to be recognized and all candidate images in the event class set after secondary screening;
[0028] Feature-based matching: Calculate the similarity between the input image and the features of all images to be screened respectively, use the similarity calculation as the scoring function for the final selection, sort according to the matching results, and the one with the highest score is the event class to which the image belongs.
[0029] Compared with the prior art, the beneficial effects of the present invention are:
[0030] Using the multi-modal event ontology model as supplementary information in the image event recognition process, the structured information therein makes the information matching process more standardized and structured; using the corpus, knowledge base, etc. as auxiliary tools for element matching reduces the probability of incorrect matching due to the lack of understanding ability of the machine; the multi-modal technology is introduced, fully combining the information covered by the image and text, and improving the accuracy of the image recognition process.
[0031] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following takes the preferred embodiments of the present invention and describes them in detail with the accompanying drawings. The specific implementation manners of the present invention are given in detail by the following embodiments and their accompanying drawings. Brief Description of the Drawings
[0032] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0033] Figure 1 is a flowchart of the steps of an image event recognition method based on a multi-modal event ontology in this application;
[0034] Figure 2 is a structural block diagram of an image event recognition method based on a multi-modal event ontology in this application. Detailed implementation manners
[0035] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention. In the following paragraphs, the present invention will be described more specifically by way of example with reference to the accompanying drawings. The advantages and features of the present invention will be clearer according to the following description and the claims. It should be noted that the accompanying drawings are all in very simplified forms and use non-precise scales, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.
[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0037] Please refer to Figures 1 - 2 , in an embodiment of the present invention, an image event recognition method based on a multi-modal event ontology Figure 1 is a flowchart of the steps shown according to the present invention, including: steps 101 to 103,
[0038] Step 101 is an image keyword acquisition step, that is, using multi-label classification technology to acquire important keywords in the input image information;
[0039] In this application, the step 101 may specifically include the following sub-steps:
[0040] Sub-step S11 is an image region extraction part, which uses Selective Search or RPN (Region Proposal Network) technology to acquire representative regions of the image, obtains several sub-images containing the key parts of the image, these sub-images represent the main information of the image, and makes each representative region sub-image retain only one key target as much as possible.
[0041] Among them, Selective Search is an improvement on the sliding window region extraction technology. It first segments the image, and then merges the segmented boxes based on the similarity of attributes such as color, texture, size, and shape compatibility, so as to obtain a set of sub-images of the most representative image regions; RPN integrates the region extraction function into the R-CNN network framework to realize the integration of R-CNN.
[0042] Sub-step S12, based on the multi-label classification technology, respectively pass the sub-images generated in the region extraction technology through the multi-label classification CNN model to obtain a sequence of keyword sets corresponding to each region sub-image. For example, if k sub-images are generated in the region extraction stage, k corresponding sets are generated: A 1 , A 2 , …, A k ; In addition, according to the classification summary results, attributes such as the total number of objects need to be generated.
[0043] The multi-label classifier here adopts the HCP (Hypotheses-CNN-Pooling) structure based on hypotheses, which is a multi-label classification model based on the region extraction technology, and obtains key information in the picture by proposing hypotheses for key regions.
[0044] When performing segmentation on the region extraction sub-images, the relevance of the things or descriptions existing in the sub-images is considered. Therefore, no additional calculation is required. By default, the label contents generated from the same sub-image have the highest relevance, so they are placed in the same set.
[0045] Sub-step S13, perform part-of-speech tagging on the keyword sets of the region sub-images, and make a new division of the keyword sets according to the part of speech. For example, the most important nouns, adjectives, and verbs in the keywords are respectively divided into: B 1 , B 2 , B 3 .
[0046] Step 102 is the step of screening event class sets. Using the obtained keywords, through text matching with the element information in the six-tuple representation structure of event classes, find the event class set with the highest matching degree in the event ontology model;
[0047] In this application, the step 102 may specifically include the following sub-steps:
[0048] Sub-step S21, element matching: According to the existing multi-modal event ontology model, with the help of text matching technologies such as semantic similarity, match the image keywords with the main event elements such as "action", "object", and "environment", and screen the required event class sets.
[0049] Our goal is to construct an event (class) structure from these words, and we need to use the method of "filling in the blanks" to build it:
[0050] Regard all the words in the noun set B 1 as objects, fill them into the object element part, and according to A i (i = 1,..., k) sets, regard the words in the adjective set B 2 as object attributes and fill them into the object element; then fill all the verbs in the verb set B 3 into the text part of the multi-modal information element; the image part of the multi-modal information element is the input image.
[0051] At this time, we have built an event (class) structure based on the content of the input image. The next step is to perform the element matching process:
[0052] For the newly built event (class), calculate the similarity between its text and the corresponding elements of the event class in the multi-modal event ontology, and obtain the sum as the scoring function in the first stage:
[0053] Among them, the sum represents the total value of the similarity values of the six elements, sim(·) is the similarity calculation function, and X and Y are the text sequences of the corresponding elements of the newly established event (class) and the existing event class respectively.
[0054] Sub-step S22, according to external knowledge such as the corpus, semantic dictionary or network resources, calculate the semantic relevance between the image keywords and the text part in the "multi-modal information" element of the event class, and perform a secondary screening on the event class set according to the result.
[0055] By learning external knowledge such as the corpus, semantic dictionary or network resources, calculate the total relevance between the newly established event (class) and the text sequence in the multi-modal information element of the existing event class, and obtain the scoring function in the second stage:
[0056] Score 2 = Score 1 + rel(M, N), where M and N are the text sequences of the language expression elements in the newly established event (class) and the existing event class respectively, and rel(·) is the relevance calculation function.
[0057] Step 103 is the image matching step. For the images of all event classes in the filtered high-matching event class set, match them with the input image, and select the corresponding event class with the highest score as the result of the final image event recognition.
[0058] In this application, the step 103 may specifically include the following sub-steps:
[0059] Sub-step S31: Extract the features of the input image to be recognized and all candidate images in the event class set after secondary screening.
[0060] The features of the image can be manually obtained features, convolutional features obtained by a neural network, or a combination of both.
[0061] Among them, the convolutional features can be obtained using pre-trained models such as the VGG classification network or the Faster R-CNN object detection network. The specific model can be modified or replaced according to the form selected for feature description.
[0062] Sub-step S32: Calculate the similarity between the input image and the features of all images to be screened respectively. Use the similarity calculation as the scoring function for the final selection, and sort according to the matching results. The one with the highest score is the event class to which the image belongs.
[0063] So far, the scoring function for the third stage can be obtained:
[0064] Score 3 = Score 2 + match(P,Q), where P and Q are the images in the multi-modal information elements of the newly established event (class) and the existing event classes respectively, and match(·) is the image matching function.
[0065] According to the scoring functions of the three stages, the event class corresponding to the input image can be screened.
[0066] If the multi-modal event ontology model contains an instance set of event classes, the method described in the present invention can be used for more detailed calculations to obtain more accurate results.
[0067] The above is only a preferred embodiment of the present invention, and it does not impose any form of limitation on the present invention. Any ordinary technician in the industry can smoothly implement the present invention according to the illustrations in the specification and the above description. However, any equivalent changes such as slight modifications, decorations, and evolutions made by those skilled in the art within the scope of the technical solution of the present invention using the technical content disclosed above are equivalent embodiments of the present invention. At the same time, any equivalent changes, modifications, and evolutions made to the above embodiments based on the essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. An image event recognition method based on a multi-modal event ontology, characterized in that, it includes the following steps: Image keyword acquisition: Using multi-label classification technology, obtain the important keywords of the input image; Screening event class set: Utilize the obtained keywords, and through text matching with the element information in the six-tuple representation structure of the event class, find the event class set with the highest matching degree in the event ontology model; Image matching: For the images of all event classes in the screened event class set with high matching degree, perform feature-based matching with the input image, and select the corresponding event class with the highest score as the result of the final image event recognition; The image keyword acquisition step further includes the following parts: Image region extraction: Extract the key regions of the image to obtain several sub-images containing the key parts of the image, and these sub-images represent the main information of the image; Multi-label classifier: Based on multi-label classification technology, process the sub-images generated in the region extraction technology respectively to obtain the keyword sets corresponding to each region sub-image; Keyword annotation: Perform part-of-speech annotation on the keyword sets of the region sub-images, and make a new division of the keyword sets according to the part of speech; The step of screening the corresponding event class set further includes: Element matching: According to the existing multi-modal event ontology model, perform corresponding event element matching between the obtained image keywords and it, and screen the required event class set; External knowledge supplementation: Use external knowledge to further screen the results of element matching; The image matching step further includes: Feature extraction: Extract the features of the input image to be recognized and all candidate images in the event class set after secondary screening; Feature-based matching: Calculate the similarity between the features of the input image and all images to be screened respectively, use the similarity calculation as the scoring function for the final selection, sort according to the matching results, and the one with the highest score is the event class to which the image belongs.
2. The method according to claim 1, characterized in that, In the region extraction part, use the SelectiveSearch or RPN technology to obtain the representative regions of the image, and make each representative region sub-image try to retain only one key target.
3. The method according to claim 2, characterized in that, In the multi-label classification part, let the representative region sub-images pass through the multi-label classification CNN model to obtain the keywords corresponding to the sub-images, put the keywords generated by each sub-image into different sets, and generate a sequence of image keyword sets; in addition, according to the classification summary results, generate the object total number attribute.
4. The method according to claim 1, characterized in that, In the element matching part, it is necessary to complete the event element matching process with the help of text matching technologies such as semantic similarity to generate an event class set with a relatively high matching degree.
5. The method according to claim 1, characterized in that, In the external knowledge supplementation part, it is necessary to calculate the semantic relevance between the image keywords and the text part of the "multi-modal information" element of the event class according to the corpus, semantic dictionary or network resources, and perform secondary screening on the event class set according to the results.
Citation Information
Patent Citations
Multi-modal event representation learning method based on event ontology
CN114201965A
Target image recognition method and device
CN114429649A