Scene graph generation method based on collaborative learning and IoST data

Through collaborative learning and scene graph generation method of IoST data, the object detection and Transformer architecture are used to calculate the difference-guided prompt vector for feature fusion, and a more accurate scene graph is generated, which solves the problem of scene graph generation of dynamic social interaction in IoST scenarios and improves the accuracy of relationship prediction.

CN120374960AActive Publication Date: 2025-07-25湖南工商大学
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510839242.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-25
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The existing technology scenario graph generation method in social intelligence (IoST) scenarios is difficult to adapt to dynamic social interactions, insufficient training data, and lack of predefined relationship templates, resulting in complex model construction and low accuracy of relation prediction.

Method used

The scene graph generation method based on collaborative learning and IoST data is adopted, and the visual features of the subject and object target are obtained through object detection, the difference-guided prompt vector is calculated, and the feature fusion is used for multi-layer Transformer and S²P_MSA mechanisms are used to generate scene graphs with a relationship decoder.

Benefits of technology

This improves the recall and accuracy of scene graph generation, realizes unbiased classification of target relationships in IoST scenarios, and reduces workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374960A_ABST
    Figure CN120374960A_ABST
Patent Text Reader

Abstract

The invention relates to a scene graph generation method based on collaborative learning and IoST data, and the method comprises the steps: carrying out the target detection of a picture in an IoST scene, and obtaining the visual features of a subject target / object target in the picture; for the visual features of the subject target, calculating a difference guide prompt vector between the subject target and the object target; enabling the initial visual block features obtained by adding the position codes and the difference guide prompt vectors to pass through multiple layers of first Transformers, fusing any visual block feature input in each layer with the corresponding difference guide prompt vector based on an attention mechanism in each layer, and outputting a plurality of subject visual feature blocks in the last layer; performing same processing on the visual features of the object target to obtain a plurality of object visual feature blocks; obtaining a relation classification result based on each subject visual feature block and each object visual feature block; and finally, constructing a scene graph based on a relationship classification result, the subject target and the object target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of scene graph generation, and in particular to a scene graph generation method based on collaborative learning and IoST data. Background Art

[0002] In the scenario of Intelligence of Social Things (IoST), the interactions of social entities (such as human-device, device-environment) have the characteristics of dynamic evolution. In the field of scene graph generation in the IoST scenario, traditional scene graph generation methods face problems such as existing tools being difficult to adapt to dynamic social interactions, insufficient training data, and lack of predefined relationship templates, resulting in complex model construction and low relationship prediction accuracy. Moreover, existing scene graph generation methods focus more on feature extraction or single classification stages, rather than a complete scene graph generation system. Summary of the Invention

[0003] Based on this, it is necessary to provide a scene graph generation method based on collaborative learning and IoST data, and this method includes: S1: Obtain the image to be processed in the IoST scenario, and use the target detector to perform target detection on the image to obtain the category, bounding box, and visual features of each target in the image, where the targets include subject targets / object targets; S2: Add position encoding to the visual features of the subject target to obtain the initial visual block features; calculate the differences in category and bounding box between the subject target and the object target to obtain the difference-guided prompt vector; pass the initial visual block features through multiple layers of the first Transformer, and starting from the second layer, the input of each layer of the first Transformer is the visual block features output by the previous layer; based on the S²P_MSA mechanism in each layer of the first Transformer, fuse the initial visual block features / visual block features input in each layer with the difference-guided prompt vector, and the last layer of the first Transformer outputs several subject visual feature blocks; S3: Replace the visual features of the subject target in step S2 with the visual features of the object target, and execute step S2 again to obtain several object visual feature blocks; S4: Connect each subject visual feature block and each object visual feature block along the channel direction and then pass them through a fully connected layer respectively to obtain several subject-object semantic visual features; pass the subject-object semantic visual features through a relationship decoder, and map the decoding result to a relationship classification result, and construct a scene graph based on the relationship classification result, subject target, and object target.

[0004] Preferably, the target detector includes a Faster R-CNN model.

[0005] Preferably, the process of obtaining the initial visual features includes: The visual features of the subject target are divided into several feature blocks, and position encoding vectors are added after the several feature blocks to obtain the initial visual features.

[0006] Preferably, the process of obtaining the difference-guided prompt vector includes: Based on the semantic vector difference between the subject target and the object target, the subject-object semantic difference feature is obtained; Obtain the upper left coordinate, lower right coordinate, and center point coordinate of the bounding box of the subject target; Obtain the upper left coordinate, lower right coordinate, and center point coordinate of the bounding box of the object target; Based on the upper left abscissa, lower right abscissa, and center point abscissa of the two bounding boxes, calculate the relative position difference of the corresponding abscissas; Based on the upper left ordinate, lower right ordinate, and center point ordinate of the two bounding boxes, calculate the relative position difference of the corresponding ordinates; Based on the widths and heights of the two bounding boxes, calculate the relative size difference; Based on the upper left coordinates and lower right coordinates of the two bounding boxes, calculate the ratio difference between the intersection area and the bounding box of the subject target; Input the relative position differences of the abscissas, the relative position differences of the ordinates, the relative size difference, and the ratio difference between the intersection area and the bounding box of the subject target between the two bounding boxes into the fully connected layer, and map them into the subject-object space difference feature; Concatenate the subject-object semantic difference feature and the subject-object space difference feature, pass them through the ReLU activation function, and regularize the obtained activation result through the Dropout function to obtain the spatial prompt vector; Multiply the normalized spatial prompt vector by any normalized initial visual block feature / visual block feature to obtain a mask vector, and perform matrix element multiplication on the mask vector and the initial visual block feature / visual block feature to obtain the corresponding spatially prompted visual block feature; Map the subject-object space difference feature through the fully connected layer to obtain the channel prompt vector; Multiply the corresponding spatially prompted visual block feature by the channel prompt vector, and pass the obtained product through the multi-layer perceptron to obtain the difference-guided prompt vector corresponding to the initial visual block feature / visual block feature.

[0007] Preferably, the process of obtaining the subject-object semantic difference feature includes: Pass the category of the subject target and the category of the object target through the Glove model respectively to obtain the subject semantic vector and the object semantic vector; Calculate the difference between the subject semantic vector and the object semantic vector, and pass the obtained difference through a fully connected layer to obtain the subject-object semantic difference feature.

[0008] Preferably, the calculation process of the relative position difference includes: Calculate the difference in the abscissa of the upper left corner, the difference in the ordinate of the upper left corner, the difference in the abscissa of the lower right corner, the difference in the ordinate of the lower right corner, the difference in the abscissa of the center point, and the difference in the ordinate of the center point between the two bounding boxes respectively; Divide the difference in the abscissa of the upper left corner, the difference in the abscissa of the lower right corner, and the difference in the abscissa of the center point by the width of the picture respectively to obtain the relative position differences of the corresponding abscissas; Divide the difference in the ordinate of the upper left corner, the difference in the ordinate of the lower right corner, and the difference in the ordinate of the center point by the height of the picture respectively to obtain the relative position differences of the corresponding ordinates.

[0009] Preferably, the calculation process of the relative size difference includes: The relative size difference includes the relative width difference, the relative height difference, and the relative ratio difference; Calculate the width difference between the two bounding boxes, and divide the obtained width difference by the width of the picture to obtain the relative width difference; Calculate the height difference between the two bounding boxes, and divide the obtained height difference by the height of the picture to obtain the relative height difference; Divide the width of the bounding box of the object target by the width of the picture to obtain the first quotient; divide the height of the bounding box of the object target by the height of the picture to obtain the second quotient; multiply the first quotient and the second quotient to obtain the first product; Divide the width of the bounding box of the subject target by the width of the picture to obtain the third quotient; divide the height of the bounding box of the subject target by the height of the picture to obtain the fourth quotient; multiply the third quotient and the fourth quotient to obtain the second product; Subtract the second product from the first product to obtain the relative ratio difference.

[0010] Preferably, the calculation process of the ratio difference between the intersection area and the bounding box of the subject target is: Judge whether there is an intersection area between the bounding box of the subject target and the bounding box of the object target, If there is no intersection area, the ratio difference is 0; if there is an intersection area, then: Pass the abscissa of the lower right corner of the two bounding boxes through the minimum value function to obtain the first minimum value; pass the abscissa of the upper left corner of the two bounding boxes through the maximum value function to obtain the first maximum value; subtract the first maximum value from the first minimum value to obtain the abscissa intersection; Pass the ordinate of the lower right corner of the two bounding boxes through the minimum value function to obtain the second minimum value; pass the ordinate of the upper left corner of the two bounding boxes through the maximum value function to obtain the second maximum value; subtract the second maximum value from the second minimum value to obtain the ordinate intersection; Divide the abscissa intersection by the width of the picture to obtain the fifth quotient; divide the ordinate intersection by the height of the picture; multiply the fifth quotient by the sixth quotient to obtain the third product; Subtract the second product from the third product to obtain the ratio difference between the intersection region and the bounding box of the subject target.

[0011] Preferably, fusing the initial visual patch feature / visual patch feature of each layer input with the difference-guided prompt vector through the S²P_MSA mechanism includes: In any first Transformer layer, Calculate the sum of the initial visual patch feature / the visual patch feature output from the previous layer and the difference-guided prompt vector, multiply the obtained result by the query projection matrix to obtain the query vector; Multiply the initial visual patch feature / the visual patch feature output from the previous layer by the key projection matrix to obtain the key vector; Multiply the initial visual patch feature / the visual patch feature output from the previous layer by the value projection matrix to obtain the value vector; Calculate the attention scores based on the query vector, key vector, and value vector, splice the attention scores in all attention heads, and fuse the splicing result through a fully connected layer into the visual patch feature / several subject visual feature patches output from the corresponding layer.

[0012] Preferably, the relationship decoder includes multiple layers of second Transformers, and the number of layers of the second Transformers is 3 layers more than the number of layers of the first Transformers; fuse the classification patch feature, semantic patch feature, spatial patch feature, and several subject-object semantic visual features through the relationship decoder, output the classification patch feature in the last layer of the second Transformers, and map the output classification patch feature to a relationship classification result through a linear classifier; The classification patch feature is a learning parameter; Connect the subject semantic vector and the object semantic vector and then pass through a fully connected layer to output the semantic patch feature; Connect the bounding box feature of the subject target and the bounding box feature of the object target and then pass through a fully connected layer to output the spatial patch feature; Pass the length, width, and center point coordinates of the bounding box of the subject target / object target through a fully connected layer to obtain the bounding box feature of the subject target / object target.

[0013] Beneficial effects: By performing object detection on images in the IoST scenario, the visual features of the subject object / object object in the image are obtained; for the visual features of the subject object, the difference-guided hint vector between the subject object and the object object is calculated; the initial visual patch features obtained by adding positional encoding and the difference-guided hint vector pass through multiple first Transformers, and based on the attention mechanism in each layer of the first Transformers, any one of the visual patch features input in each layer is fused with the corresponding difference-guided hint vector, and the last layer outputs several subject visual feature patches; the same processing is performed on the visual features of the object object to obtain several object visual feature patches; the relationship classification result is obtained based on each subject visual feature patch and each object visual feature patch; finally, a scene graph is constructed based on the relationship classification result, the subject object, and the object object. Description of the Drawings

[0014] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0015] Figure 1 It is a flowchart of the method for generating a scene graph based on collaborative learning and IoST data in the embodiments of the present application. Detailed Embodiments

[0016] To make the above objects, features, and advantages of the present application more obvious and understandable, the following will give a detailed description of the specific embodiments of the present application in conjunction with the drawings. Many specific details are set forth in the following description to fully understand the present application. However, the present application can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0017] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0018] As Figure 1 shown, this embodiment provides a method for generating a scene graph based on collaborative learning and IoST data, and the method includes: S1: Obtain the image to be processed in the IoST scenario, perform object detection on the image using an object detector, and obtain the category, bounding box, and visual features of each object in the image. The objects include subject objects / object objects.

[0019] In this embodiment, the object detector includes a Faster R-CNN model. In this embodiment, the Faster R-CNN object detector is used to detect the objects in the image, and the bounding box, category, and feature information of the objects are obtained.

[0020] In this embodiment, the subject features and object features of the object interaction prompt are fully utilized to strengthen the collaborative learning strategy of the object detector, and the training effect is improved by combining a small amount of labeled data and automatically generated pseudo-labels, strengthening the learning of feature extraction, thereby improving the encoding ability of the image and subsequent classification ability.

[0021] S2: Add positional encoding to the visual features of the subject object to obtain initial visual block features; calculate the differences in category and bounding box between the subject object and the object object to obtain a difference-guided prompt vector; pass the initial visual block features through multiple layers of the first Transformer, and starting from the second layer, the input of each layer of the first Transformer is the visual block features output by the previous layer; based on the S²P_MSA mechanism in each layer of the first Transformer, fuse the initial visual block features / visual block features input in each layer with the difference-guided prompt vector, and the last layer of the first Transformer outputs several subject visual feature blocks.

[0022] In this step, the process of obtaining the initial visual features includes: The visual features of the subject object are divided into several feature blocks, and positional encoding vectors are added after the several feature blocks to obtain the initial visual features.

[0023] Further, the process of obtaining the difference-guided prompt vector includes: Based on the semantic vector difference in category between the subject object and the object object, obtain the subject-object semantic difference feature. In this embodiment, the process of obtaining the subject-object semantic difference feature includes: Pass the category of the subject object and the category of the object object through the Glove model respectively to obtain the subject semantic vector and the object semantic vector; Calculate the difference between the subject semantic vector and the object semantic vector, and pass the obtained difference through a fully connected layer to obtain the subject-object semantic difference feature.

[0024] Obtain the upper left coordinate, lower right coordinate, and center point coordinate of the bounding box of the subject object (the bounding box information of the subject object); Obtain the upper left coordinates, lower right coordinates, and center point coordinates of the bounding box of the object target (the bounding box information of the object target); When calculating the difference, subtract the bounding box information of the subject target from the bounding box information of the object target; Based on the upper left abscissa, lower right abscissa, and center point abscissa of the two bounding boxes, calculate the relative position difference of the corresponding abscissa; Based on the upper left ordinate, lower right ordinate, and center point ordinate of the two bounding boxes, calculate the relative position difference of the corresponding ordinate.

[0025] In this embodiment, the calculation process of the relative position difference includes: Calculate the upper left abscissa difference, upper left ordinate difference, lower right abscissa difference, lower right ordinate difference, center point abscissa difference, and center point ordinate difference between the two bounding boxes respectively; Divide the upper left abscissa difference, lower right abscissa difference, and center point abscissa difference by the width of the picture respectively to obtain the relative position difference of the corresponding abscissa; Divide the upper left ordinate difference, lower right ordinate difference, and center point ordinate difference by the height of the picture respectively to obtain the relative position difference of the corresponding ordinate.

[0026] Based on the widths and heights of the two bounding boxes, calculate the relative size difference.

[0027] In this embodiment, the calculation process of the relative size difference includes: The relative size difference includes relative width difference, relative height difference, and relative ratio difference; Calculate the width difference between the two bounding boxes, and divide the obtained width difference by the width of the picture to obtain the relative width difference; Calculate the height difference between the two bounding boxes, and divide the obtained height difference by the height of the picture to obtain the relative height difference; Divide the width of the bounding box of the object target by the width of the picture to obtain the first quotient; divide the height of the bounding box of the object target by the height of the picture to obtain the second quotient; multiply the first quotient and the second quotient to obtain the first product; Divide the width of the bounding box of the subject target by the width of the picture to obtain the third quotient; divide the height of the bounding box of the subject target by the height of the picture to obtain the fourth quotient; multiply the third quotient and the fourth quotient to obtain the second product; Subtract the second product from the first product to obtain the relative ratio difference.

[0028] Based on the upper left coordinates and lower right coordinates of the two bounding boxes, calculate the ratio difference between the intersection area and the bounding box of the subject target.

[0029] In this embodiment, the calculation process of the ratio difference between the intersection region and the bounding box of the subject target is as follows: Determine whether there is an intersection region between the bounding box of the subject target and the bounding box of the object target. If there is no intersection region, the ratio difference is 0; if there is an intersection region, then: Pass the abscissa of the lower right corner of the two bounding boxes through the minimum function to obtain the first minimum value; pass the abscissa of the upper left corner of the two bounding boxes through the maximum function to obtain the first maximum value; subtract the first maximum value from the first minimum value to obtain the abscissa intersection. Pass the ordinate of the lower right corner of the two bounding boxes through the minimum function to obtain the second minimum value; pass the ordinate of the upper left corner of the two bounding boxes through the maximum function to obtain the second maximum value; subtract the second maximum value from the second minimum value to obtain the ordinate intersection. Divide the abscissa intersection by the width of the image to obtain the fifth quotient; divide the ordinate intersection by the height of the image; multiply the fifth quotient by the sixth quotient to obtain the third product. Subtract the second product from the third product to obtain the ratio difference between the intersection region and the bounding box of the subject target.

[0030] Input the relative position differences of the abscissas, the relative position differences of the ordinates, the relative size differences, and the ratio difference between the intersection region and the bounding box of the subject target between the two bounding boxes into the fully connected layer, and map them into the subject-object space difference features. Concatenate the subject-object semantic difference features and the subject-object space difference features, pass them through the ReLU activation function, and regularize the obtained activation result through the Dropout function to obtain the spatial cue vector. Multiply the normalized spatial cue vector by any normalized initial visual block feature / visual block feature to obtain the mask vector, and perform matrix element product operation on the mask vector and the initial visual block feature / visual block feature to obtain the corresponding spatially cued visual block feature. Map the subject-object space difference features through the fully connected layer to obtain the channel cue vector. Multiply the corresponding spatially cued visual block feature by the channel cue vector, and pass the obtained product through the multi-layer perceptron to obtain the difference-guided cue vector corresponding to the initial visual block feature / visual block feature.

[0031] In this embodiment, the process of obtaining the mask vector includes: Multiply each element in the normalized spatial cue vector by each element in any normalized initial visual block feature / visual block feature respectively to obtain a number of fourth products. Determine whether each fourth product is greater than a preset threshold. If so, set the value at the corresponding position of the fourth product to 1; otherwise, set the value at the corresponding position of the fourth product to 0.

[0032] In this embodiment, the preset threshold is 0.5 and can be set according to actual requirements.

[0033] The visual features adopted in this embodiment are the target-related features obtained by the target detection module. However, there is information in the visual features that is irrelevant to the relationship classification, resulting in a scattered feature space for the same type of relationships and a low classification accuracy. Since the relationship classification task is a multi-object interaction task, the visual attention area of the subject is related to the object. Therefore, this embodiment designs a subject-object semantic difference calculation module and a subject-object spatial difference calculation module to obtain the subject-object semantic difference features and the subject-object spatial difference features respectively; designs a difference-guided prompt generation module to update the query part with the difference-guided prompt vector, thereby guiding the model to focus on the relationship-related target visual features.

[0034] In this embodiment, the S²P_MSA mechanism is used to replace the multi-head self-attention mechanism MSA in the conventional Transformer. Specifically, the fusion of the initial visual patch features / visual patch features input to each layer and the difference-guided prompt vector through the S²P_MSA mechanism includes: In any first Transformer layer, Calculate the sum of the initial visual patch features / the visual patch features output from the previous layer and the difference-guided prompt vector, multiply the obtained result by the query projection matrix to obtain the query vector; Multiply the initial visual patch features / the visual patch features output from the previous layer by the key projection matrix to obtain the key vector; Multiply the initial visual patch features / the visual patch features output from the previous layer by the value projection matrix to obtain the value vector; Calculate the attention scores based on the query vector, key vector, and value vector, concatenate the attention scores in all attention heads, and fuse the concatenated result through a fully connected layer into the visual patch features / several subject visual feature patches output from the corresponding layer.

[0035] Use the difference-guided prompt vector to update the query part in the S²P_MSA mechanism, guiding the model to focus on the visual area related to the relationship classification rather than the entire target visual area, so as to obtain the relationship inherent features (subject visual feature patches / object visual feature patches) for relationship classification.

[0036] S3: Replace the visual features of the subject target in step S2 with the visual features of the object target, and execute step S2 again to obtain several object visual feature patches.

[0037] Considering the different characteristics of the target when it is used as the subject and the object, in this embodiment, a dual feature enhancement module is designed based on the Transformer architecture, which is respectively used to refine the visual features of the subject and the visual features of the object; fully exploit the semantic and spatial position feature differences between the subject and the object, and obtain the subject visual feature block and the object visual feature block for relation classification, providing a basis for subsequent relation classification.

[0038] S4: Connect each subject visual feature block and each object visual feature block along the channel direction and then pass them through a fully connected layer respectively to obtain a number of subject-object semantic visual features; pass the subject-object semantic visual features through a relation decoder, and map the decoding result to a relation classification result, and construct a scene graph based on the relation classification result, the subject target, and the object target.

[0039] Furthermore, the relation decoder includes multiple layers of a second Transformer, and the number of layers of the second Transformer is 3 layers more than the number of layers of the first Transformer; fuse the classification patch feature, the semantic patch feature, the spatial patch feature, and a number of subject-object semantic visual features through the relation decoder, and output the classification patch feature in the last layer of the second Transformer, and map the output classification patch feature to a relation classification result through a linear classifier; The classification patch feature is a learning parameter; Connect the subject semantic vector and the object semantic vector and then pass them through a fully connected layer to output the semantic patch feature; Connect the bounding box features of the subject target and the bounding box features of the object target and then pass them through a fully connected layer to output the spatial patch feature; Pass the length (calculated from the abscissa of the upper left corner and the abscissa of the lower right corner in the bounding box), width (calculated from the ordinate of the upper left corner and the ordinate of the lower right corner in the bounding box), and the center point coordinates of the bounding box of the subject target / object target through a fully connected layer to obtain the bounding box features of the subject target / object target.

[0040] Even further, construct a triple based on the subject target, the relation classification result, and the object target, and the set of all triples in the picture is the scene graph.

[0041] In this embodiment, a relation classification module based on the Transformer architecture is designed, and by exploring the associations among the subject-object semantic visual features, the semantic patch features, and the spatial patch features, the classification of relations is realized.

[0042] The scene graph generation method based on collaborative learning and IoST data provided in this embodiment has the following beneficial effects: 1. This method performs object detection on images in the IoST scenario to obtain the visual features of the subject object / object object in the image. For the visual features of the subject object, the difference-guided cue vector between the subject object and the object object is calculated. The initial visual patch features obtained by adding positional encoding and the difference-guided cue vector pass through multiple first Transformers. Based on the attention mechanism in each layer of the first Transformers, any one of the visual patch features input in each layer is fused with the corresponding difference-guided cue vector, and the last layer outputs several subject visual feature patches. The same processing is performed on the visual features of the object object to obtain several object visual feature patches. The relationship classification result is obtained based on each subject visual feature patch and each object visual feature patch. Finally, a scene graph is constructed based on the relationship classification result, the subject object, and the object object. This method realizes unbiased classification of the relationships between objects in images in the IoST scenario with a long-tail distribution, improves the recall rate and accuracy of scene graph generation, and greatly reduces the workload.

[0043] 2. After experiments, this method achieved recall rate indicators of 44.5% for mR@100 and 50.0% for F@100 respectively after training.

[0044] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0045] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for generating a scene graph based on collaborative learning and IoST data, characterized in that, Including: S1: Obtain the image to be processed in the IoST scenario, and use the target detector to perform target detection on the image to obtain the category, bounding box, and visual features of each target in the image. The targets include subject targets / object targets; S2: Add positional encoding to the visual features of the subject target to obtain the initial visual block features; calculate the differences in category and bounding box between the subject target and the object target to obtain the difference-guided prompt vector; pass the initial visual block features through multiple layers of the first Transformer. Starting from the second layer, the input of each layer of the first Transformer is the visual block features output by the previous layer; based on the S²P_MSA mechanism in each layer of the first Transformer, fuse the initial visual block features / visual block features input in each layer with the difference-guided prompt vector. The last layer of the first Transformer outputs several subject visual feature blocks; S3: Replace the visual features of the subject target in step S2 with the visual features of the object target, and execute step S2 again to obtain several object visual feature blocks; S4: Connect each subject visual feature block and each object visual feature block along the channel direction and then pass them through the fully connected layer respectively to obtain several subject-object semantic visual features; Pass the subject-object semantic visual features through the relationship decoder, and map the decoding result to the relationship classification result. Based on the relationship classification result, subject target, and object target, construct a scene graph.

2. The method for generating a scene graph based on collaborative learning and IoST data according to claim 1, wherein The target detector includes a Faster R-CNN model.

3. The method for generating a scene graph based on collaborative learning and IoST data according to claim 1, characterized in that, The process of obtaining the initial visual features includes: Divide the visual features of the subject target into several feature blocks, and add positional encoding vectors after the several feature blocks to obtain the initial visual features.

4. The method for generating a scene graph based on collaborative learning and IoST data according to claim 1, wherein The process of obtaining the difference-guided prompt vector includes: Based on the semantic vector difference in category between the subject target and the object target, obtain the subject-object semantic difference feature; Obtain the upper left coordinate, lower right coordinate, and center point coordinate of the bounding box of the subject target; Obtain the upper left coordinate, lower right coordinate, and center point coordinate of the bounding box of the object target; Based on the upper left abscissa, lower right abscissa, and center point abscissa of the two bounding boxes, calculate the relative position difference of the corresponding abscissas; Based on the upper left ordinate, lower right ordinate, and center point ordinate of the two bounding boxes, calculate the relative position difference of the corresponding ordinates; Based on the widths and heights of the two bounding boxes, calculate the relative size difference; Based on the upper left coordinate and lower right coordinate of the two bounding boxes, calculate the ratio difference between the intersection area and the bounding box of the subject target; Input the relative position differences of each abscissa, relative position differences of each ordinate, relative size difference, and ratio difference between the intersection area and the bounding box of the subject target between the two bounding boxes into the fully connected layer, and map them to the subject-object spatial difference feature; Concatenate the subject-object semantic difference feature and the subject-object spatial difference feature, pass them through the ReLU activation function, and regularize the obtained activation result through the Dropout function to obtain the spatial prompt vector; Multiply the normalized spatial cue vector by any one of the normalized initial visual patch features / visual patch features to obtain a mask vector, and perform an element-wise matrix multiplication of the mask vector and the initial visual patch features / visual patch features to obtain the corresponding spatially cued visual patch features; Map the subject-object spatial difference features through a fully connected layer to obtain a channel cue vector; Multiply the corresponding spatially cued visual patch features by the channel cue vector, and pass the obtained product through a multi-layer perceptron to obtain the difference-guided cue vector corresponding to the initial visual patch features / visual patch features.

5. The method for generating a scene graph based on collaborative learning and IoST data according to claim 4, wherein The process of obtaining the subject-object semantic difference features includes: Pass the categories of the subject target and the object target through the Glove model respectively to obtain the subject semantic vector and the object semantic vector; Calculate the difference between the subject semantic vector and the object semantic vector, and pass the obtained difference through a fully connected layer to obtain the subject-object semantic difference features.

6. The method for generating a scene graph based on collaborative learning and IoST data according to claim 4, wherein The calculation process of the relative position difference includes: Calculate the differences in the upper-left abscissa, upper-left ordinate, lower-right abscissa, lower-right ordinate, center abscissa, and center ordinate between the two bounding boxes respectively; Divide the differences in the upper-left abscissa, lower-right abscissa, and center abscissa by the width of the image respectively to obtain the relative position differences of the corresponding abscissas; Divide the differences in the upper-left ordinate, lower-right ordinate, and center ordinate by the height of the image respectively to obtain the relative position differences of the corresponding ordinates.

7. The method for generating a scene graph based on collaborative learning and IoST data according to claim 4, wherein The calculation process of the relative size difference includes: The relative size difference includes the relative width difference, relative height difference, and relative ratio difference; Calculate the width difference between the two bounding boxes, and divide the obtained width difference by the width of the image to obtain the relative width difference; Calculate the height difference between the two bounding boxes, and divide the obtained height difference by the height of the image to obtain the relative height difference; Divide the width of the bounding box of the object target by the width of the image to obtain the first quotient; divide the height of the bounding box of the object target by the height of the image to obtain the second quotient; multiply the first quotient and the second quotient to obtain the first product; Divide the width of the bounding box of the subject target by the width of the image to obtain the third quotient; divide the height of the bounding box of the subject target by the height of the image to obtain the fourth quotient; multiply the third quotient and the fourth quotient to obtain the second product; Subtract the second product from the first product to obtain the relative ratio difference.

8. The method for generating a scene graph based on collaborative learning and IoST data according to claim 7, wherein The calculation process of the ratio difference between the intersection area and the bounding box of the subject target is: Judge whether there is an intersection area between the bounding box of the subject target and the bounding box of the object target, If there is no intersection area, the ratio difference is 0; if there is an intersection area, then: Pass the lower-right abscissas of the two bounding boxes through the minimum function to obtain the first minimum value; pass the upper-left abscissas of the two bounding boxes through the maximum function to obtain the first maximum value; subtract the first maximum value from the first minimum value to obtain the abscissa intersection; Pass the lower-right ordinates of the two bounding boxes through the minimum function to obtain the second minimum value; pass the upper-left ordinates of the two bounding boxes through the maximum function to obtain the second maximum value; subtract the second maximum value from the second minimum value to obtain the ordinate intersection; Divide the abscissa intersection by the width of the picture to obtain the fifth quotient; divide the ordinate intersection by the height of the picture; multiply the fifth quotient by the sixth quotient to obtain the third product; Subtract the second product from the third product to obtain the ratio difference between the intersection region and the bounding box of the subject target.

9. The method for generating a scene graph based on collaborative learning and IoST data according to claim 1, wherein Fusing the initial visual patch feature / output visual patch feature of each layer input with the difference-guided prompt vector through the S²P_MSA mechanism includes: In any first Transformer layer, Calculate the sum of the initial visual patch feature / output visual patch feature of the previous layer and the difference-guided prompt vector, multiply the obtained result by the query projection matrix to obtain the query vector; Multiply the initial visual patch feature / output visual patch feature of the previous layer by the key projection matrix to obtain the key vector; Multiply the initial visual patch feature / output visual patch feature of the previous layer by the value projection matrix to obtain the value vector; Calculate the attention scores based on the query vector, key vector, and value vector, concatenate the attention scores in all attention heads, and fuse the concatenated result through a fully connected layer into the output visual patch feature / several subject visual feature patches of the corresponding layer.

10. The method for generating a scene graph based on collaborative learning and IoST data according to claim 5, wherein, The relationship decoder includes multiple second Transformers, and the number of layers of the second Transformer is 3 more than that of the first Transformer; fuse the classification patch feature, semantic patch feature, spatial patch feature, and several subject-object semantic visual features through the relationship decoder, output the classification patch feature in the last second Transformer layer, and map the output classification patch feature to a relationship classification result through a linear classifier; The classification patch feature is a learning parameter; Concatenate the subject semantic vector and the object semantic vector and then pass through a fully connected layer to output the semantic patch feature; Concatenate the bounding box feature of the subject target and the bounding box feature of the object target and then pass through a fully connected layer to output the spatial patch feature; Pass the length, width, and center point coordinates of the bounding box of the subject target / object target through a fully connected layer to obtain the bounding box feature of the subject target / object target.

Citation Information

Patent Citations

  • Visual relation detection method, device, visual relation detection training method and device

    CN108229272A

  • Relation visual attention mechanism-based scene graph generation method

    CN110991532A

  • Scene graph generation method based on global information and position embedding

    CN113836339A

  • Unbiased scene graph generation method based on effective feature representation

    CN115861779A

  • Scene graph generation method based on unified decoder

    CN119359904A