Method for generating image scene graph in the field of education based on multi-view information fusion

By constructing object views and semantic views in images in the field of education, and using graph convolution networks to fusion information to generate visual scene maps, the problem of sparse visual information at the bottom of visual objects in the image is solved, and more accurate image semantic understanding and information extraction effects are achieved.

CN115761036BActive Publication Date: 2025-05-30NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211156523.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2025-05-30
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

The underlying visual information of visual objects in images in educational fields is relatively sparse, making it difficult to accurately represent image semantics. Traditional methods are not effective when extracting image features from underlying visual information.

Method used

Using a multi-view information fusion method, by constructing object views and semantic views, using graph convolution networks to fusion information between multiple views, to generate visual scene maps based on semantic relationships and positional relationships.

Benefits of technology

Effectively capturing semantic information of images improves the information extraction effect, simplifies the network structure, reduces training costs, and ensures the effect of each part of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761036B_ABST
    Figure CN115761036B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for generating a scene graph of educational images based on multi-view information fusion, which captures the semantics of the image from multiple aspects such as the visual information of the objects in the image and the various interactive information between the objects. The contribution of the present invention is that a method for generating a scene graph Scene Graph based on multi-view information fusion for educational images is proposed. The method can use multiple different views of the image to mine the semantics of the objects in the image and the different types of associations between the objects, thereby generating a heterogeneous scene graph containing multiple types of nodes and edges. The scene graph represents the extracted and complex semantic information of educational images from the semantics of the visual objects and the different associations between the objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the research on computer application fields, multi-view information fusion, and visual scene graph generation methods in the education field, and particularly relates to a method based on convolutional neural networks and graph neural networks, using network topologies to enable a computer to generate a structured graph structure that can represent image features and information using the relationships between objects. Background Art

[0002] Scene graphs can represent the effective information in the original image in a concise, clear, and structured manner. This characteristic gives scene graphs high application value, and the information extracted through the scene graph generation task can be used as the input for other tasks. It has been proven that the information in the scene graph structure can help enhance the effects of various types of computer vision tasks. For example, image generation tasks can effectively generate high-quality images using the information provided by the scene graph, and Visual Question Answering (VQA) can directly use the semantic information contained in the scene graph to answer questions in the VQA task, enhancing the reasoning ability of the VQA model.

[0003] In the previous computer vision field, many tasks were carried out in one step. This means that each system needs to first extract information from the original image, then perform relevant processing on the information, and finally make predictions on the output. Such a network will at least need to include two modules: information feature extraction and information feature processing, resulting in a relatively complex network and high training costs. Moreover, due to the difficulty in distinguishing between key regions and non-key regions in the image, the network effect cannot be guaranteed. Using the scene graph as the carrier of information transmission can effectively enable downstream tasks to pay more attention to the improvement of their own network effects, handing over the problem of extracting effective structured information from the image to the scene graph generation task, reducing the complexity of the network, and the functional separation also ensures that the network effect of each part can be guaranteed. Summary of the Invention

[0004] Technical Problems to be Solved

[0005] The underlying visual information (such as color, texture, etc.) of visual objects in images in the education field is relatively sparse, and the understanding of image semantics also requires combining information such as numbers and interactions between objects that appear in the graph. This results in traditional image features based on underlying visual information being difficult to accurately represent image semantics. In view of the characteristics of images in the education field above, the present invention proposes a visual scene graph generation method based on multi-view information fusion, capturing the semantics of images from multiple aspects such as the visual appearance of objects in the graph and the interaction information between objects.

[0006] Technical Solutions

[0007] A method for generating an image scene graph in the field of education based on multi-view information fusion, characterized by the following steps:

[0008] Step 1: Construction of multi-views

[0009] Step 1.1: Construction of the object view: Use the Faster R-CNN object detector to identify the category and position coordinates of the objects in the graph, and use a convolutional neural network to construct the visual information encoding of the identified objects, and combine it with the objects in the graph to construct a fully connected object view as a node of the visual graph;

[0010] Step 1.2: Construction of the semantic view: First, based on OCR technology and unsupervised semantic segmentation technology, finely identify the text information in the graph, then perform weighted fusion on the context information of different categories to increase the reasoning ability of the network, and finally construct a fully connected semantic view;

[0011] Step 2: Construction of the multi-view information fusion module

[0012] Step 2.1: Fusion of the semantic view and the object view: Use the node information in the semantic view, and use a graph convolutional network to construct a fully connected network between the semantic view and the object view to update the node information in the object view, hoping to obtain the semantic information between objects in this way;

[0013] Step 2.2: Self-fusion of the object view: After fusing the semantic view, use the fusion method in Step 2.1 to construct a fully connected network between the object views to self-update the object views;

[0014] Step 3: Construction of the visual scene graph generation module

[0015] Step 3.1: Generation of the visual scene graph based on semantic relationships: After the node features of the object view are updated, generate a visual scene graph based on semantic relationships by calculating the probability distribution of semantic interactions between nodes; The nodes in the visual scene graph represent the regions of objects and their category labels, and the edges represent the semantic interaction categories between visual objects;

[0016] Step 3.2: Generation of the visual scene graph based on position relationships: First, use the Intersection over Union (IOU) of the bounding boxes of visual objects to determine whether there is an inclusion or overlap relationship between two objects; If the IOU < 0.5, then calculate the distance and angle between the center points of their bounding boxes to determine their fine-grained position interaction categories, and the fine-grained position interaction categories include eight categories: up, down, left, right, upper left, lower left, upper right, and lower right.

[0017] Further technical solution of the present invention: The fusion method of the semantic view and the object view described in step 2.1 is as follows:

[0018] Using the node information in the semantic view and the graph convolutional network to construct a fully connected network between the semantic view and the object view to update the node information in the object view ; for the node o i in the object view, its feature vector update formula is as follows:

[0019]

[0020] where N(i) is the set of neighbor nodes of node o i , b represents the offset of the model, W represents the parameters of the model, l represents the update iteration times of the node information in the object view, and c j i is calculated by the following formula, representing the standard deviation of the node degree;

[0021]

[0022] σ represents the Relu activation function, that is

[0023] f(x) = max(0, x), (2-3)

[0024] e ji represents the weight from node o j to node oi in N(i), which can be calculated by the following formula:

[0025]

[0026] ρ[s i , o j represents the similarity degree between s i and o j , which is jointly determined by the categories and visual features of s i and o j ; it is calculated by the following formula;

[0027] ρ[s i , o j = f s ((f l (s i )·f l (o j ))||(f fe (s i )·f fe (o j ))) (2-5)

[0028] fl (·) is a category encoder that obtains the word vectors representing the categories through Fasttext; f fe (·) represents a visual feature mapper that maps high-dimensional visual features into a low-dimensional space; f s (·) is a similarity encoder used to calculate the similarity between two nodes;

[0029] distance(s i ,o j ) represents the distance between two objects. The farther the distance between the objects, the weaker the degree of mutual influence between the two objects; "||" represents the concatenation operation, which concatenates the distance between the objects and the similarity degree; Normalize the influence coefficient of s i on o j at the node o j .

[0030] A further technical solution of the present invention: The self-fusion method of the object view described in step 2.2 is specifically as follows:

[0031] After completing the feature update of all nodes in the object view, use the fusion method in step 2.1 to construct a fully connected network between the object views and perform self-update on the object views; the feature vector update formula is as follows:

[0032]

[0033] Except for e ji , it is the same as step 2.1. e ji is obtained from the following formula

[0034]

[0035] The calculation methods of distance(o i ,o j ) and ρ[o i ,o j are the same as those in step 2.1.

[0036] A further technical solution of the present invention: The method for generating a visual scene graph based on semantic relationships described in step 3.1 is specifically as follows:

[0037] The probability distribution of semantic interaction between nodes in the visual scene graph can be calculated by the following formula:

[0038]

[0039] Among them, "||" represents the concatenation operation, which concatenates and respectively represent the nodes v in the object view iand v j The updated feature vector, v i,j is the visual feature corresponding to the region that covers both the i-th and j-th objects, f c is the semantic interaction classifier, and the interaction categories include contain, of, leftchild, rightchild, next, and hold;

[0040] After selecting the currently predicted edge type, a positive sample fully connected subgraph is constructed among the object types at both ends of the edge, and a negative sample subgraph with the same number of associations as in the subgraph is constructed in the whole graph. The relationship is scored and predicted in the two subgraphs, and the score can be obtained by the following formula:

[0041] score(o i , e ij , o j ) = f sc (o i ||o j *e ij ), (4 - 2)

[0042] where f sc represents the score predictor, and o i ||o j represents concatenating o i and o j as the input of f sc ;

[0043] The loss function of the score evaluation network for predicting semantic relationships between scene graph nodes is defined as:

[0044]

[0045] Due to the characteristics of images in the education field, the number of associations is related to the type of connections between nodes and the number of nodes. Therefore, a topk algorithm that adaptively adjusts and selects the maximum and minimum number according to node connections and the number of nodes is designed:

[0046]

[0047] where represents the numerical mapping from the number of nodes hums(o ij , o i , o j ) when the relationship type is e ij to the selected maximum and minimum number;

[0048] After obtaining all types of relationships, using the existing prior knowledge, pruning is performed by type, and finally a visual scene graph based on semantic relationships is obtained.

[0049] Further technical solution of the present invention: The method for generating a visual scene graph based on positional relationship described in step 3.2 is as follows:

[0050] Taking all the relationships in the visual scene graph given semantic relationships as the basis for position classification, for the position interaction between objects, first use the Intersection over Union (IOU) of the bounding boxes of visual objects to determine whether there is an inclusion or overlapping relationship between two objects; if IOU < 0.5, then by calculating the distance and angle between the center points of their bounding boxes, to determine the fine-grained position interaction category between them, including eight categories: up, down, left, right, upper left, lower left, upper right, and lower right; the calculation process is as follows:

[0051]

[0052] where f p ,f i is the position interaction category classifier.

[0053] A computer system, characterized in that it includes: one or more processors, and a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the above method.

[0054] A computer-readable storage medium, characterized in that it stores computer-executable instructions, and the instructions are used to implement the above method when executed.

[0055] Beneficial effects

[0056] An image scene graph generation method in the field of education based on multi-view information fusion provided by the present invention generates a visual scene graph based on semantic relationships and positional relationships by constructing and fusing the semantic view and object view of the image.

[0057] The scene graph generated by the present invention can represent the effective information in the original image in a concise, clear and structured manner, achieve the purpose of information feature extraction, and transmit the result to the downstream task, which can effectively enable it to pay more attention to the improvement of its own information feature processing effect, realize the functional separation, not only reduce the complexity of the network, but also ensure the effect of each part of the network.

[0058] Secondly, the present invention captures the semantics of images from multiple aspects such as the visual image of the objects in the figure and the interaction information between the objects, and fuses the semantic information and visual information in the scene graphs expressing different relationships, greatly improving the information extraction effect of images in the education field. In contrast, traditional methods often extract image features from low-level visual information, and the low-level visual information (such as color, texture, etc.) of visual objects in the images in this field is relatively sparse, resulting in the image features extracted by these methods being difficult to accurately represent the image semantics. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The drawings are only for the purpose of illustrating specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components.

[0060] Figure 1 is the overall task framework diagram;

[0061] Figure 2 is the structure diagram of the Faster R-CNN model;

[0062] Figure 3 is the types and quantities of triples in the dataset;

[0063] Figure 4 is the original image of the Array-list and the visual scene graph based on semantic and positional relationships;

[0064] Figure 5 is the original image of the Flowchart and the visual scene graph based on semantic and positional relationships;

[0065] Figure 6 is the comparison between Flowchart and Deadlock. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0067] The present method is an image scene graph generation method in the education field based on multi-view information fusion, which consists of three parts: the construction of multi-views, the construction of a multi-view information fusion module, and the construction of a visual scene graph generation module. The overall architecture schematic diagram is as Figure 1 shown and is described in detail as follows:

[0068] 1. Construction of multi-views

[0069] 1.1 Construction of the object view

[0070] The present invention uses Faster R-CNN to extract visual information from images. The model structure of Faster R-CNN is as Figure 2 shown.

[0071] The Resnet50 backbone network is used to extract the feature map from the image first, and this feature map is saved as the input shared by RPN and ROI pooling. The RPN network takes the feature sub-map wrapped by the anchor box in the global feature map as the input, uses the classifier to judge whether the feature sub-map can be determined as a target, and uses the regression function to output the coordinate position of its bounding box. At the same time, for each target, the ROI pooling is used to uniformly output the dimension, and its feature vector is obtained on the feature map, and finally the object view is obtained.

[0072] 1.2 Construction of semantic view

[0073] Compared with natural images, the semantic information of images in the education field has a relatively fixed sentence relationship structure. Making prior inferences on semantic information can effectively improve the reasoning ability of scene graphs.

[0074] In order to make the word vectors and sentence vectors contain as much syntactic information, sentence structure information and relationship reasoning information as possible, this method combines the image descriptions and problem descriptions in the dataset, plus a large number of data structure descriptions to form a rich corpus, and uses the Fasttext model to train the word vectors of the corpus. When using, first obtain the word vectors of each word, and then extract the relationship between the word vectors of a complete sentence through a bidirectional long short-term memory neural network (BiLSTM), so as to construct a sentence vector with generalization and complete semantic relationship, and finally construct a semantic view.

[0075] 2. Construction of multi-view information fusion module

[0076] Traditional scene graph generation methods based on low-level visual information for inference are difficult to accurately represent the image semantics of images in the education visual field. Therefore, a multi-view information fusion method is designed to fuse the objects in the graph and the interaction information between the numbers and objects that appear in the graph, and capture the image semantics from multiple aspects.

[0077] 2.1 Fusion of semantic view and object view

[0078] Using the node information in the semantic view , a fully connected network is constructed between the semantic view and the object view using a graph convolutional network to update the node information in the object view . For the node o j in the object view, its feature vector update formula is as follows:

[0079]

[0080] where N(i) is the set of neighbor nodes of node o i and b represents the offset of the model, W represents the parameters of the model, and c ji is calculated by the following formula and represents the standard deviation of the node degree.

[0081]

[0082] σ represents the Relu activation function, that is

[0083] f(x) = max(0, x), (2 - 3)

[0084] e ji represents the weight from node o j to node o i in N(i) and can be calculated by the following formula:

[0085]

[0086] ρ[s i , o j represents the similarity degree between s i and o j and is jointly determined by the categories and visual features of s i and o j and is calculated by the following formula.

[0087] ρ[s i , o j = f s ((f l (s i )·f l (o j ))||(f fe (s i )·f fe (o j ))), (2 - 5)

[0088] f l (·) is the category encoder that obtains the word vector representing the category through Fasttext. f fe (·) represents the visual feature mapper that maps high-dimensional visual features to a low-dimensional space. f s (·) is the similarity encoder used to calculate the similarity degree between two nodes.

[0089] distance(s i , o j) represents the distance between two objects. The farther the distance between the objects, the shallower the degree of mutual influence between the two objects. "||" represents the concatenation operation, which concatenates the distance between the objects and the similarity degree together. For s i the influence coefficient on o j is normalized at node o j .

[0090] 2.2 Self-Fusion of Object Views

[0091] After completing the feature update of all nodes in the object view, a fully connected network between the object views is constructed using a fusion method similar to that in step 2.1 to perform self-update on the object view. The feature vector update formula is as follows:

[0092]

[0093] Except for e ji it is the same as step 2.1, and e ji is obtained from the following formula

[0094]

[0095] distance(o i , o j ) and ρ[o i , o j are also calculated in the same way as step 2.1.

[0096] 3. Construction of the Visual Scene Graph Generation Module

[0097] 3.1 Generation of Visual Scene Graphs Based on Semantic Relationships

[0098] The present invention generates a visual scene graph based on semantic information from the object view after node feature update. The nodes in the visual scene graph represent the regions of the objects and their category labels, and the edges represent the semantic interactions between the visual objects. Specifically, the probability distribution of the semantic interactions between the nodes in the visual scene graph can be calculated by the following formula:

[0099]

[0100] Among them, "||" represents the concatenation operation, which and represent the updated feature vectors of nodes v i and v j in the object view respectively, v i,j is the visual feature corresponding to the region that simultaneously covers the i-th and j-th objects, and f cIt is a semantic interaction classifier, and the interaction categories include contain, of, leftchild, rightchild, next, and hold.

[0101] For images in the education field, multi-view fusion can effectively strengthen the implicit semantic connections between objects by adding semantic information to the object views, thereby better improving the model performance and can well solve the problem of insufficient reasoning ability caused by sparse visual information in the visual objects of images in the education field.

[0102] Among the objects in images in the education field, the objects with relationships and the relationships between objects are relatively fixed. Therefore, its <subject, predicate, object> triples can be set in advance as prior knowledge for scene graph generation and input into the model. When performing association prediction and classification, the object views can be divided into subgraphs according to the node types generated by object detection. Each time, one type of edge is selected and prediction is performed in the subgraph, which can greatly avoid redundant calculations on all nodes in the object views and significantly improve the efficiency. The types and quantities of triples in the dataset are shown in Figure 3 .

[0103] After selecting the current predicted edge type, a positive sample fully connected subgraph is constructed among the object types at both ends of the edge, and a negative sample subgraph with the same number of associations as in the subgraph is constructed in the full graph. The relationship scores are predicted respectively in the two subgraphs, and the scores can be obtained by the following formula:

[0104] score(o i , e ij , o j ) = f sc (o i ||o j * e ij ), (3 - 2)

[0105] where f sc represents the score predictor, and o i ||o j means concatenating o i and o j as the input of f sc .

[0106] The loss function of the score evaluation network for predicting semantic relationships between scene graph nodes is defined as:

[0107]

[0108] Due to the characteristics of images in the field of education, the number of associations is related to the number of types of connections between nodes and the number of nodes. Therefore, a topk algorithm that adaptively adjusts and selects the maximum and minimum number according to the node connections and the number of nodes is designed:

[0109]

[0110] Among them represents the numerical mapping from the number of nodes nums(o ij when the relationship type is e i , o j ) to the selected maximum and minimum number.

[0111] After obtaining all types of relationships, using the existing prior knowledge, pruning is performed according to the type (a unified analysis is carried out on the number of times and semantic importance of all relationships appearing in the samples, and relationships with few occurrences and low importance levels are pruned. For this standard, refer to the existing prior knowledge), and finally a visual scene graph based on semantic relationships is obtained.

[0112] 3.2 Generation of Visual Scene Graph Based on Position Relationships

[0113] Now assume that the visual scene graph based on semantic relationships can completely express the direct relationships in the image. In order to obtain the visual scene graph based on position relationships, the present invention proposes a method for generating a visual scene graph of position relationships based on the visual scene graph of semantic relationships. This method takes all the relationships in the visual scene graph given the semantic relationships as the basis for position classification. For the position interaction between objects, first, the intersection over union (IOU) of the bounding boxes of the visual objects is used to determine whether there is an inclusion or overlapping relationship between two objects. If IOU < 0.5, then by calculating the distance and angle between the center points of their bounding boxes, the fine-grained position interaction categories between them are determined, including eight categories: up, down, left, right, upper left, lower left, upper right, and lower right. The calculation process is as follows:

[0114]

[0115] where f p , f l is the position interaction category classifier.

[0116] 3.3 Implementation Details

[0117] The present invention uses Faster R-CNN with Resnet50 as the backbone network. The visual information obtained from Faster R-CNN is a 2048-dimensional vector. The word vector is a 50-dimensional vector, and the sentence vector is a 512-dimensional vector. When transmitting information, the bounding box, center point, length, width, and aspect ratio of the object are constructed into a 9-dimensional normalized vector, which is concatenated with the visual information carried by the object itself. A two-layer graph convolutional network is used for multi-view fusion. After obtaining the updated node representation, it is input into a two-layer MLP for prediction, and the output is a one-dimensional vector.

[0118] During the training process, the present invention uses two stages for separate training. First, the Faster R-CNN object detection network and the OCR network are trained on the dataset, and then the scene graph generation model is trained using the visual information and semantic information output by them. SGD is used as the optimization function, and its learning rate is set to 0.003.

[0119] 3.4 Quantitative evaluation and analysis

[0120] The present invention uses recall to quantitatively evaluate the model. Since the present invention only processes images in the education field, it is not compared with other methods. The quantitative evaluation is only to judge the feasibility of the model.

[0121] Recall (Recall@K, R@K) is the proportion of the top K classification results with the highest confidence found in the whole image in the relationship ground truth. In the quantitative analysis experiment, K is selected as 30, and the recall is calculated separately for images of different categories. The results are shown in Table 1.

[0122] Table 1 Recall of images of different categories

[0123] Category R@30 Array-list 77.8 Binary-tree 64.6 Deadlock 53.0 Directed-graph 60.1 Flowchart 51.7 Linked-List 73.8

[0124] Continued Table 1 Recall of images of different categories

[0125] Category R@30 Logic-circuit 44.8 Network-topology 48.5 Non-binary-tree 68.0 Queue 69.9 Stack 57.9 Undirected-graph 65.3

[0126] According to Table 1, it can be obtained that the recall rates of Array-list, Linked-List, and Queue are relatively high, while the recall rates of Logic-circuit and Network-topology are relatively low. The lowest recall rate of Logic-circuit is 44.8%, which is presumably related to the relatively sparse and clear image information in the dataset and the prior knowledge pre-input in the present invention. However, the prior knowledge also limits the generalization of the model to some extent. For example, in Network-topology, the target forms in the graph are diverse and it is relatively easy to generate detection errors. When the object category in the image does not conform to the expectation of the prior knowledge, the generation of relationships will be suppressed.

[0127] 3.5 Visualization Result Evaluation and Analysis

[0128] Select the relatively well-performing Array-list ( Figure 4 ) and the poorly-performing Flowchart ( Figure 5 ) as examples to analyze the visualization results of the two generated scene graphs.

[0129] It can be seen from Figure 5 that a relatively large error occurred during the object detection of the flowchart. The elliptical terminal was recognized as a process in the deadlock. Figure 6 This is a comparison of two types of pictures.

[0130] This phenomenon not only appears in these two types of pictures. For example, misidentifications will occur between the process in the Deadlock category and the node in the binary graph category, between the element in the Arraylist and the node in multiple categories, the nand gate and the and gate in Network topology, etc. This is due to the small difference in the underlying visual features of various types of nodes in the dataset. However, the impact on the scene graph generation task is acceptable. In the generated visual scene graph based on semantic relationships and the visual scene graph based on positional relationships, the relationships are relatively complete and the accuracy is guaranteed.

[0131] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention.

Claims

1. A method for generating an image scene graph in the field of education based on multi-view information fusion, characterized in that the steps are as follows: Step 1: Construction of multi-views Step 1.1: Construction of the object view: Using the Faster R-CNN object detector, identify the category and position coordinates of the objects in the graph, and use a convolutional neural network to construct the visual information encoding of the identified objects, which is combined with the objects in the graph to construct a fully connected object view as a node of the visual graph; Step 1.2: Construction of the semantic view: First, based on OCR technology and unsupervised semantic segmentation technology, finely identify the text information in the graph, then perform weighted fusion on the context information of different categories to enhance the inference ability of the network, and finally construct a fully connected semantic view; Step 2: Construction of the multi-view information fusion module Step 2.1: Fusion of the semantic view and the object view: Using the node information in the semantic view, use a graph convolutional network to construct a fully connected network between the semantic view and the object view to update the node information in the object view, expecting to obtain the semantic information between objects in this way; Step 2.2: Self-fusion of the object view: After fusing the semantic view, use the fusion method in Step 2.1 to construct a fully connected network between the object views to self-update the object view; Step 3: Construction of the visual scene graph generation module Step 3.1: Generation of the visual scene graph based on semantic relationships: After the node features of the object view are updated, generate a visual scene graph based on semantic relationships by calculating the probability distribution of semantic interactions between nodes; The nodes in the visual scene graph represent the regions of objects and their category labels, and the edges represent the semantic interaction categories between visual objects; Step 3.2: Generation of the visual scene graph based on positional relationships: First, use the Intersection over Union (IOU) of the bounding boxes of visual objects to determine whether there is an inclusion or overlapping relationship between two objects; If the IOU < 0.5, then by calculating the distance and angle between the center points of their bounding boxes, determine their fine-grained positional interaction categories, and the fine-grained positional interaction categories include eight categories: up, down, left, right, upper left, lower left, upper right, and lower right.

2. The method for generating an image scene graph in the field of education based on multi-view information fusion according to claim 1, characterized in that: The method for fusing the semantic view and the object view described in Step 2.1 is specifically as follows: Using the semantic view The node information in is used to construct a fully connected network between the semantic view and the object view using a graph convolutional network to update the node information in the object view For the node o in the object view i , its feature vector update formula is as follows: where N(i) is the set of neighbor nodes of node o i , b represents the offset of the model, W represents the parameters of the model, l represents the update iteration times of node information in the object view, c ji is calculated by the following formula and represents the standard deviation of the node degree; σ represents the Relu activation function, that is f(x) = max(0, x), (2 - 3) e ji represents the weight from node o to node o in N(i), which can be calculated by the following formula: j to node o i The weight of can be calculated by the following formula: ρ[s i ,o j represents the similarity degree of s i and o j , which is jointly determined by the categories and visual features of s i and o j ; it is calculated by the following formula; ρ[s i ,o j = f s ((f l (s i )·f l (o j )) || (f fe (s i )·f fe (o j ))), (2 - 5) f l (·) is a category encoder that obtains the word vectors representing the categories through Fasttext; f fe (·) represents a visual feature mapper that maps high-dimensional visual features into a low-dimensional space; f s (·) is a similarity encoder used to calculate the similarity between two nodes; distance(s i ,o j ) represents the distance between two objects. The farther the distance between the objects, the shallower the degree of mutual influence between the two objects; "||" represents the splicing operation, splicing the distance between the objects and the degree of similarity together; The influence coefficient of s i on o j is normalized at the node o j .

3. The method for generating an image scene graph in the field of education based on multi-view information fusion according to claim 2, characterized in that: The self-fusion method of the object view described in Step 2.2 is specifically that: After completing the feature update of all nodes in the object view, use the fusion method in Step 2.1 to construct a fully connected network between the object views to self-update the object view; The update formula of its feature vector is as follows: Except for e ji it is the same as step 2.1, and e ji is obtained from the following formula distance(o i ,o j ) and ρ[o i ,o j are calculated in the same way as in Step 2.

1.

4. The method for generating an image scene graph in the field of education based on multi-view information fusion according to claim 3, characterized in that: The method for generating a visual scene graph based on semantic relationships described in step 3.1 is as follows: The probability distribution of semantic interaction between nodes in the visual scene graph can be calculated by the following formula: Among them, "||" represents the splicing operation, which combines and respectively represent the updated feature vectors of nodes v i and v j in the object view. v i,j is the visual feature corresponding to the region that simultaneously covers the i-th and j-th objects, and f c is a semantic interaction classifier, and the interaction categories include contain, of, leftchild, rightchild, next, and hold; After selecting the currently predicted edge type, a positive sample fully connected subgraph is constructed among the object types at both ends of the edge, and a negative sample subgraph with the same number of associations as in the subgraph is constructed in the whole graph. The relationship scores are predicted respectively in the two subgraphs, and the scores can be obtained by the following formula: score(o i ,e ij ,o j ) = f sc (o i ||o j *e ij ), (4 - 2) where f sc represents a score predictor, o i ||o j represents concatenating o i and o j as the input to f sc ; The loss function of the scoring evaluation network for predicting semantic relationships between scene graph nodes is defined as: Due to the characteristics of images in the education field, the number of associations is related to the types of connections between nodes and the number of nodes. Therefore, a topk algorithm is designed to adaptively adjust and select the maximum and minimum number according to node connections and the number of nodes: Among them represents the numerical mapping of the number of nodes nums(o ij when the relationship type is e i , o j ) to the selected maximum value quantity; After obtaining all types of relationships, using the existing prior knowledge, pruning is performed by type, and finally a visual scene graph based on semantic relationships is obtained.

5. The method for generating an image scene graph in the education field based on multi-view information fusion according to claim 4, characterized in that: The method for generating a visual scene graph based on positional relationships described in step 3.2 is as follows: Taking all the relationships in the visual scene graph given semantic relationships as the basis for position classification, for the positional interaction between objects, first use the Intersection over Union (IOU) of the bounding boxes of visual objects to determine whether there is an inclusion or overlapping relationship between two objects; if IOU < 0.5, then by calculating the distance and angle between the center points of their bounding boxes, to determine the fine-grained position interaction category between them, including eight categories: above, below, left, right, upper left, lower left, upper right, and lower right; the calculation process is as follows: where f p , f i is a location interaction category classifier.

6. A computer system, characterized in that it includes: One or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to claim 1.

7. A computer-readable storage medium, characterized in that it stores computer-executable instructions that are used to implement the method according to claim 1 when executed.

Citation Information

Patent Citations

  • Image-text bothway retrieval method based on multi-view unite embedded space

    CN107330100A

  • Semantic-based image recognition system and semantic-based image recognition method

    CN112528705A