Scene Graph Generation Method and System Based on Deep Evidence Active Learning

By building a hybrid active learning model and deep evidence learning, we can solve the problem of data set annotation difficulties and high cost in the scene graph generation task, realize high-precision model training and stable training process, and reduce the annotation cost.

CN116501903BActive Publication Date: 2025-07-25NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310544482.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-15
Publication Date
2025-07-25
Estimated Expiration
2043-05-15

AI Technical Summary

Technical Problem

There are difficulties in the data set annotation and high cost in the scene graph generation task, and long-tail distribution and error annotation problems, resulting in difficulty in model training. It is difficult for existing methods to effectively use a small amount of annotation data to train excellent models.

Method used

Build a hybrid active learning model, combine deep evidence learning to perform uncertainty estimation, remove similar contexts and redundant pictures, explore background relationships, and provide reliable uncertainty sampling through relationship recommendation diagrams to form the final sampling result.

Benefits of technology

Improve model accuracy under a fixed annotation budget, stabilize the training process, reduce the annotation cost, and effectively solve the challenges in the scene graph generation task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116501903B_ABST
    Figure CN116501903B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for scene graph generation based on deep evidence active learning, including: S1, constructing a hybrid active learning model for the scene graph generation task, where the hybrid active learning model includes an active learning model based on uncertainty sampling and density sampling; S2, performing uncertainty estimation on the relationships of the constructed hybrid active learning model based on deep evidence learning; S3, removing similar contexts in the scene graph generation task to avoid context-level bias, and removing redundant pictures to eliminate picture-level bias; S4, mining and judging the background relationships in the pictures, judging the probability of the existence of foreground relationships between any two objects in the pictures, and combining the uncertainty estimation of deep evidence learning to provide reliable uncertainty sampling for the hybrid active learning framework to obtain the final sampling result. The technical solution of the present invention has higher accuracy and a more stable training process under a fixed annotation budget.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and particularly relates to a method and system for generating a scene graph based on deep evidence active learning. Background Art

[0002] The task of scene graph generation aims to generate a structured representation for a visual scene to describe objects, object attributes, and the relationships between objects. The task of scene graph generation has received much attention because it can provide rich semantic information and has good application potential in many other visual tasks, such as object detection, image search, visual question answering, etc. However, as far as we know, as a task that can communicate computer vision and natural language processing, the task of scene graph generation has still not been fully studied, and its main challenges include the following.

[0003] On the one hand, there are many intractable problems in the current scene graph generation datasets, such as long-tailed distribution, missing annotations, and incorrect annotations, etc., which makes it difficult to train a satisfactory model. On the other hand, the data annotation cost of the scene graph generation task is very expensive. The label form of the scene graph generation task is a triple, such as <person, rides, bicycle>, and it is very difficult and time-consuming to annotate the relationships of all object pairs in the picture. There is an urgent need to train an excellent model with a small annotation cost in many tasks. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and system for generating a scene graph based on deep evidence active learning to solve at least one of the above problems existing in the prior art.

[0005] Based on the above purpose, one or more embodiments in the present application provide a method for generating a scene graph based on deep evidence active learning, and the content includes the following steps:

[0006] S1. Construct a hybrid active learning model for the scene graph generation task, and the hybrid active learning model includes an active learning model based on uncertainty sampling and density sampling;

[0007] S2. Estimate the uncertainty of the relationship of the constructed hybrid active learning model based on deep evidence learning;

[0008] S3. Remove similar contexts in the scene graph generation task to avoid context-level bias, and remove redundant pictures to eliminate picture-level bias;

[0009] S4. Mine and judge the background relationship in the picture, judge the probability that there is a foreground relationship between any two objects in the picture, and combine the uncertainty estimation of deep evidence learning to provide reliable uncertainty sampling for the hybrid active learning framework to obtain the final sampling result.

[0010] Based on the above technical solution of the present invention, the following improvements can also be made:

[0011] Optionally, step S2 includes:

[0012] Set the input object pair and its relationship annotation to be x i and y ij , where j ∈ {1, 2, …, k} is the index of the k-th relationship classification. If x i belongs to the j-th classification, then y ij = 1, otherwise y ij = 0;

[0013] Then the loss function of deep evidence learning is expressed as follows:

[0014]

[0015] where e ij is the output of the input x i , α ij = e ij + 1, is the Dirichlet intensity, p ij = α ij / S i is the probability that the input x i belongs to the i-th classification. Among them, the uncertainty estimate of the input x i is u i = k / S i , and its maximum value is 1.

[0016] Optionally, the gradient of the loss function of deep evidence learning is where σ is the ReLU activation function, and when f(x i ) < 0, σ′ = 0;

[0017] Use cross-entropy loss to jointly train the mixed active learning model. The cross-entropy loss function is as follows:

[0018] L = L EDL + L ce ;

[0019] Among them

[0020] Optionally, the process of eliminating similar contexts in the scene graph generation task in step S3 includes the following steps: Set that there are M object pairs in picture I, and the features are Among them represents the features extracted by the model obtained in the (t - 1)-th round of active learning. Then calculate the context-level density, denoted as D I = {d1, d2, …, d M}, where Finally, according to the calculated context-level density, the contexts with larger density are removed, and the removal ratio is a preset value.

[0021] Optionally, the process of eliminating redundant pictures in step S3 includes the following steps: Denote the unlabeled data and annotation budget in the t-th round as l t and First, calculate the picture-level density, denoted as where is the feature extracted by the model Then, remove pictures with the largest density to obtain the final active learning sampling result, where l t is the actual annotation budget, and is a hyperparameter to be set.

[0022] Optionally, increases as active learning progresses, and set

[0023] Optionally, step S4 includes:

[0024] S41. Generate relationship clusters based on existing foreground annotations;

[0025] S42. Extract statistical graphs from the relationship clusters generated in step S41;

[0026] S43. Establish a probability graph;

[0027] S44. Combine the statistical graph and the probability graph to obtain a relationship recommendation graph, and the relationship recommendation graph judges the probability of the existence of a foreground relationship between any two objects in the picture in the form of probability;

[0028] S45. Combine the uncertainty estimation of the relationship of the constructed hybrid active learning framework based on deep evidence learning to obtain the final sampling result.

[0029] Optionally, in step S41, the relationship cluster is represented as where o p and o q are objects belonging to class p and class q respectively, r represents the relationship between the objects, and the cluster size is determined by the existing relationship annotations in the dataset.

[0030] Optionally, in step S43, when establishing the probability graph, define where o p 、o m 、oq Objects belonging to classification p, classification m, and classification q respectively, and r1 and r2 are two foreground relationships. Then the probability expression that there is a foreground relationship between object classification p and object classification q is:

[0031]

[0032] where k is the number of object classifications in the dataset.

[0033] According to a second aspect of the present invention, a scene graph generation system based on deep evidence active learning is provided. The system includes a memory, a processor, and a communication circuit. The memory and the communication circuit are respectively coupled to the processor. Among them, the communication circuit is connected to the processor, and the communication circuit performs data interaction with an external terminal device under the control of the processor; the memory includes local storage and stores a computer program; the processor is used to run the computer program to execute the above-mentioned scene graph generation method based on deep evidence active learning.

[0034] The beneficial effects of the present invention are as follows. The present invention provides a scene graph generation method and system based on deep evidence active learning, which has at least the following advantages: The hybrid active learning framework proposed by the present invention is superior to the baseline method and has higher accuracy under a fixed annotation budget. The standard deviation of the results shows that the hybrid active learning framework proposed by the present invention has a more stable training process, and the hybrid active learning framework can obtain the performance of full supervision only using about 10% of the annotation cost. At the same time, a large number of ablation experiments also show that introducing active learning into the scene graph generation task will face many challenges not observed in other tasks, and the technical solutions adopted by the present invention can effectively solve these problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of the hybrid active learning framework of a scene graph generation method and system based on deep evidence active learning according to an embodiment of the present invention.

[0036] Figure 2 It is a schematic diagram of the motivation for removing similar contexts in the scene graph generation task and removing redundant pictures in a scene graph generation method and system based on deep evidence active learning according to an embodiment of the present invention.

[0037] Figure 3 It is a performance comparison diagram of the hybrid active learning framework and the baseline method of a scene graph generation method and system based on deep evidence active learning according to an embodiment of the present invention under different scene graph generation models.

[0038] Figure 4A scene graph generated by the hybrid active learning framework of a scene graph generation method and system based on deep evidence active learning according to an embodiment of the present invention.

[0039] Figure 5 A relationship heat map of a scene graph generation method and system based on deep evidence active learning according to an embodiment of the present invention.

[0040] Figure 6 Four schematic diagrams with similar uncertainties and their respective relationship features of a scene graph generation method and system based on deep evidence active learning according to an embodiment of the present invention.

[0041] Figure 7 A schematic diagram showing the performance of different duplicate removal strategies of a scene graph generation method and system based on deep evidence active learning according to an embodiment of the present invention.

[0042] Figure 8 A schematic diagram showing the performance after removing different modules of a scene graph generation method and system based on deep evidence active learning according to an embodiment of the present invention.

[0043] Figure 9 A schematic diagram for open relationship recognition of a scene graph generation method and system based on deep evidence active learning according to an embodiment of the present invention. Detailed implementation manners

[0044] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further describes the present disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0045] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second", and similar terms used in one or more embodiments of the present application do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0046] A scene graph generation method and system based on deep evidence active learning in one or more embodiments of the present application, wherein, as Figure 1 shown, the method includes the following steps:

[0047] S1. Construct a mixed-initiative learning model for the scene graph generation task. The mixed-initiative learning model includes an active learning model based on uncertainty sampling and density sampling;

[0048] S2. Estimate the uncertainty of the constructed mixed-initiative learning model relationship based on deep evidence learning. Due to the serious long-tail distribution problem of relationships in the scene graph generation task, traditional uncertainty estimation methods perform poorly. Figure 2 (a) shows the distribution of 50 relationships in the VG150 dataset and the probabilities of each relationship appearing at different sampling rates. It can be found that at a given sampling rate, a considerable part of the relationships will not appear in the sampling results. In this embodiment, such relationships are called open relationships, and existing uncertainty estimation methods often cannot provide reliable uncertainty estimates for such open relationships. On the contrary, deep evidence learning models the probability and uncertainty of classification through the Dirichlet distribution. Even in the face of open objects, it can still provide reliable uncertainty estimates. For the open relationships in the scene graph generation task, in theory, deep evidence learning will not collect any evidence for them as known classifications. Therefore, evidence learning will assign high uncertainty to these open relationships.

[0049] Therefore, this embodiment uses deep evidence learning as the uncertainty estimation technology in the scene graph generation task. Let the input object pair and its relationship annotation be x i and y ij respectively, where j ∈ {1, 2,..., k} is the index of the k-th relationship classification. If x i belongs to the j-th classification, then y ij = 1, otherwise y ij = 0;

[0050] Then the loss function of deep evidence learning is expressed as follows:

[0051]

[0052] where e ij is the output of the input x i , α ij = e ij + 1, is the Dirichlet strength, p ij = α ij / S i is the probability that the input x i belongs to the i-th classification. Among them, the uncertainty estimate of the input x i is u i = k / S i , and its maximum value is 1.

[0053] In the actual training process, it is found that the above loss function may have problems of gradient disappearance and overfitting, which will damage the performance of the scene graph generation model. The gradient of the loss function of deep evidence learning is where σ is the ReLU activation function. When f(x i ) < 0, σ′ = 0;

[0054] The cross-entropy loss is used to jointly train the hybrid active learning model. The cross-entropy loss function is as follows:

[0055] L = L EDL + L ce ;

[0056] where

[0057] S3. Remove similar contexts in the scene graph generation task to avoid context-level bias, and remove redundant images to eliminate image-level bias; propose CBM to eliminate similar contexts: Suppose there are M object pairs in image I, and their features are where represents the features extracted by the model obtained in the (t - 1)-th round of active learning. Then calculate the context-level density, denoted as D I = {d1, d2, …, d M}}, where Finally, remove those contexts with larger density according to the calculated density, and the removal ratio is a preset value.

[0058] Design IBM to eliminate similar images in the sampling results: Denote the unlabeled data and annotation budget in the t-th round as and respectively. First, calculate the image-level density, denoted as where is the feature extracted by the model . Then remove images with the largest density to obtain the final active learning sampling result, where l t is the actual annotation budget, and is a hyperparameter that needs to be set. Increases as active learning progresses. Set

[0059] As shown in Figure 2 (c), a scene may include many identical context relationships (such as <person, sitting, chair>). The embodiment takes into account that identical context relationships can only provide similar feature spaces. Therefore, if the same relationships in a single-image class are treated equally, it will inevitably provide unreliable value estimates. The embodiment calls this problem the context-level bias in scene graph generation. And the image-level bias is asFigure 2 As shown in (d), it means that the uncertainty estimation method will be seriously biased towards the classification with poor performance of the current model during the sampling process.

[0060] S4. Mining and judging the background relationship in the image. The relationship annotations in the scene graph generation dataset are very sparse, so there is a serious foreground-background imbalance problem, where the foreground refers to the object pairs that have been annotated with specific relationships, and the background refers to the object pairs that do not have any relationship. For example, Figure 2 The five objects in (b) theoretically have 20 potential relationships, but in reality there are only four specific relationships that have been annotated (such as <woman, holding, racket>, etc.). The large number of background relationships will disturb or even dominate the uncertainty estimation of the relationship, so a natural idea is to use existing methods to distinguish whether the relationship is foreground or background, such as Directed Semantic Action Graph, Confusion matrix, Statistical prior, etc. However, these methods all require large-scale annotated data for training, so they are not applicable in active learning scenarios.

[0061] In order to solve this problem, this embodiment proposes a relationship recommendation graph for analysis, which, as a relationship mining method, can determine whether the relationship between a pair of objects is foreground or background in the form of probability:

[0062] The probability of a foreground relationship between any two objects in the image is determined, and combined with the uncertainty estimation of deep evidence learning, reliable uncertainty sampling is provided for the hybrid active learning framework to obtain the final sampling result.

[0063] S4 includes the following steps:

[0064] S41. Generate relation clusters based on existing foreground annotations; the relation clusters are expressed as Among them p and q They are objects belonging to category p and category q respectively, r represents the relationship between objects, and the cluster size Determined by the existing relationship annotations in the dataset.

[0065] S42, extracting a statistical graph from the relationship clusters generated in step S41; wherein the statistical graph models whether there is a relationship between any pair of objects, 1 indicates existence and 0 indicates non-existence. Obviously, since there are only a very small number of relationship annotations in the active learning scenario, the statistical graph will be very sparse and thus produce unreliable foreground-background relationship estimates.

[0066] S43. Establish a probability graph; in this embodiment, a probability graph is established, which has the same nodes as the scene graph, and the difference is that the probability graph models the relationship between object pairs as background or foreground in the form of probability.

[0067] When establishing the probability graph, define where o p 、o m 、o q are objects belonging to classification p, classification m, and classification q respectively, and r1 and r2 are two foreground relationships. Then the probability expression that there is a foreground relationship between object classification p and object classification q is:

[0068]

[0069] where k is the number of object classifications in the dataset.

[0070] S44. Combine the statistical graph and the probability graph to obtain a relationship recommendation graph, and the relationship recommendation graph judges the probability that there is a foreground relationship between any two objects in the picture in the form of probability;

[0071] S45. Combine the uncertainty estimation of the relationship of the constructed mixed active learning framework based on deep evidence learning to obtain the final sampling result.

[0072] Experimental settings

[0073] Experimental details. In this embodiment, experiments on the proposed active learning framework are carried out on the VG150 dataset, that is, all relationship annotations in the training set are removed to create a dataset U consisting only of target box annotations. The proposed framework first randomly samples from U at a sampling rate τ1 as the training data for the first round of active learning. In the subsequent i-th round, the proposed framework samples at a sampling rate τ i until the annotation budget is exhausted. The experimental settings in this paper are as follows: τ1 = 0.01, τ i = 0.01, i = 2, 3,..., the parameter λ = 10 in IBM and η = 4, and the deduplication rate in CBM is 0.2. -5 、η = 4, and the deduplication rate in CBM is 0.2.

[0074] Baseline methods and models. The active learning baseline methods compared in this embodiment include the following:

[0075] (1) Random sampling.

[0076] (2) Classical uncertainty sampling methods, including minimum confidence sampling, entropy sampling, distance sampling, etc.

[0077] (3) The latest methods include Integer Programming Approach (IPA), Temporal Output Discrepancy (TOD), Fisher Kernel Self Supervision (FKSS), The Core-set approach, etc. The scene graph generation model used in this embodiment includes Neural Motifs backbone, VCTree, Transformer-based backbone, etc.

[0078] Evaluation metrics and modes. This embodiment samples the most commonly used active learning evaluation mode, that is, the comparison of model performance under a fixed annotation budget. This embodiment uses three of the most commonly used scene graph evaluation modes, including Predicate classification (PredCls), Scene Graph Classification (SGCls), and Scene Graph Detection (SGDet).

[0079] Results and Analysis

[0080] From Figure 3 the quantitative results shown in Figure 4 and the qualitative results in

[0081] (1) The active learning framework proposed in this embodiment outperforms all baseline methods, that is, the proposed method has higher accuracy under a fixed annotation budget.

[0082] (2) The standard deviation of the results is as shown in Figure 3 the shaded part, indicating that the framework proposed in this embodiment has a more stable training process, which mainly benefits from the proposed text removal module and image removal module.

[0083] (3) The uncertainty sampling method used for comparison performs well in the early stage of active learning, but gradually lags behind other methods in the middle and late stages. That is, the data bias problem becomes more and more serious as the model is trained.

[0084] (4) The density sampling method used for comparison performs poorly in the scene graph generation task. This embodiment believes that this is mainly because the relational features of the scene graph generation task are relatively complex and cannot provide reliable feature information required by traditional density sampling methods.

[0085] Ablation Learning

[0086] The active learning framework proposed in this embodiment includes four main modules, namely evidence uncertainty estimation, relational recommendation graph, context removal module, and image removal module. The experimental results of removing the above key modules in turn are asFigure 8 As shown. It is worth noting that in order to keep the whole framework running properly, the minimum confidence sampling is used for the uncertainty estimation method of the framework when removing the evidence uncertainty estimation. Figure 8 The results in [reference] show that the four modules proposed in this embodiment all have gains for the scene graph generation task. Next, the effects and functions of each module will be specifically explored.

[0087] Ablation study on evidence uncertainty. The VG150 used in the experiments of this embodiment includes 150 object classes and 50 relation classes. In fact, the relation annotations in this dataset far exceed 50 classes. Therefore, the relations outside the 50 classes can be called open relations. In order to verify the recognition performance of the improved evidence uncertainty of this embodiment for open relations, open relation recognition is carried out in this embodiment, and the results are as Figure 9 shown. The experimental results show that the improved evidence uncertainty method of this embodiment has obvious advantages compared with other uncertainty methods.

[0088] Ablation study on the relation recommendation graph. As Figure 5 shown, simple foreground-background binary classification statistics can only cover a small number of true relations, especially in the early stage of active learning. For example, the binary classification statistics have a recall rate of only 24.07% in the first round of active learning. Therefore, this embodiment believes that in the setting of active learning, simple binary classification statistics cannot complete the task of filtering out background relations. On the contrary, the method proposed in this embodiment not only uses binary classification statistics, but also uses probability graphs to mine potential relations. As Figure 5 shown, even in the early stage of active learning, the relation recommendation graph proposed in this embodiment can still obtain satisfactory foreground-background classification effects.

[0089] Ablation study on the context removal module. Figure 6 shows four pictures with similar uncertainties, and visualizes their relation feature distributions through t-SNE. It can be found that even with similar uncertainties, the relation spaces of the four pictures are still very different. This phenomenon confirms that uncertainty sampling is prone to generate relation-level data bias problems in the scene graph generation task. And the relation removal module designed in this embodiment can remove similar features, and this has almost no impact on the entire feature space, as Figure 6 shown. Therefore, it can be considered that the proposed relation removal module can alleviate the relation-level data bias problem, and then perform reliable value estimation on the data.

[0090] Ablation study on the picture removal module. Figure 7Shows the performance of the model under different deduplication strategies, where DeCv represents the constant deduplication method, DeLu represents the exponential deduplication method, and DeTu represents the deduplication method proposed in this embodiment. The results show that in the early stage of active learning, the performance of the exponential deduplication method is comparable to that of the method proposed in this embodiment. However, as active learning progresses, its performance begins to gradually lag behind. This phenomenon fully confirms that different deduplication strategies should be used in different active learning stages.

[0091] Conclusion

[0092] Compared with simple tasks, the active learning framework for scene graph generation tasks will face a large number of open relationships, background relationships will perturb uncertainty estimation, image-level data bias and context-level data bias and other problems. For these challenges, this embodiment proposes a hybrid active learning framework to alleviate the expensive annotation cost in scene graph generation tasks. This embodiment experiments with a variety of scene graph generation models, and the framework proposed in this embodiment exceeds the existing methods under three scene graph evaluation models. A large number of ablation experiments confirm the effectiveness of each of the above modules.

[0093] In another possible embodiment, a scene graph generation system based on deep evidence active learning is provided. The system includes a memory, a processor, and a communication circuit. The memory and the communication circuit are respectively coupled to the processor. Among them, the communication circuit is connected to the processor, and the communication circuit performs data interaction with an external terminal device under the control of the processor; the memory includes local storage and stores a computer program; the processor is used to run the computer program to execute any one of the above-mentioned scene graph generation methods based on deep evidence active learning.

[0094] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0095] The present invention will be described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded computers, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate a system for realizing the functions specified in one or more flows Figure 1 or more flows and / or blocks Figure 1 or more blocks.

[0096] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction system that realizes the functions specified in one or more flows Figure 1 or more flows and / or blocks Figure 1 or more blocks.

[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one or more flows Figure 1 or more flows and / or blocks Figure 1 or more blocks.

[0098] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0099] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A method for generating a scene graph based on deep evidence active learning, characterized in that, It includes the following steps: S1. Construct a mixed-initiative learning model for the scene graph generation task. The mixed-initiative learning model includes an active learning model based on uncertainty sampling and density sampling; S2. Estimate the uncertainty of the relationship of the constructed mixed-initiative learning model based on deep evidence learning; Including: Set the input object pair and its relationship annotation to be x i and y ij , where j ∈ {1, 2, …, k} is the index of the k-th relationship classification. If x i belongs to the j-th classification, then y ij = 1, otherwise y ij = 0; The loss function of deep evidence learning is expressed as follows: where e ij is the output of the input x i , α ij = e ij + 1, is the Dirichlet intensity, p ij = α ij / S i is the probability that the input x i belongs to the i-th category, where the uncertainty estimate of the input x i is u i = k / S i , and its maximum value is 1; S3. Remove similar contexts in the scene graph generation task to avoid context-level bias, and remove redundant pictures to eliminate picture-level bias; S4. Mine and judge the background relationship in the picture, judge the probability that there is a foreground relationship between any two objects in the picture, and combine the uncertainty estimation of deep evidence learning to provide reliable uncertainty sampling for the mixed-initiative learning framework to obtain the final sampling result.

2. The method for generating a scene graph based on deep evidence active learning according to claim 1, characterized in that, The gradient of the loss function for deep evidence learning is where σ is the ReLU activation function, and when f(x i ) < 0, σ′ = 0; Use cross-entropy loss to jointly train the mixed-initiative learning model. The cross-entropy loss function is as follows: L = L EDL + L ce ; Among them 3. The method for generating a scene graph based on deep evidence active learning according to claim 2, wherein In the process of eliminating similar contexts in the scene graph generation task in step S3, the following steps are included: It is assumed that there are M object pairs in picture I, and the features are where represents the features extracted by the model obtained in the (t - 1)-th round of active learning. Then, the context-level density is calculated and denoted as D I ={d1, d2, …, d M}, where Finally, according to the calculated context-level density, the contexts with larger density are removed, and the removal ratio is a preset value.

4. The method for generating a scene graph based on deep evidence active learning according to claim 3, wherein The process of eliminating redundant images in step S3 includes the following steps: Denote the unlabeled data and annotation budget in the t-th round as l t and First, calculate the image-level density, denoted as where is the feature extracted by the model Then, remove images with the highest density to obtain the final active learning sampling result, where l t is the actual annotation budget, and is a hyperparameter that needs to be set.

5. The method for generating a scene graph based on deep evidence active learning according to claim 4, wherein, Increase as active learning progresses, set 6. The method for generating a scene graph based on deep evidence active learning according to claim 5, wherein The step S4 includes: S41. Generate relationship clusters based on existing foreground annotations; S42. Extract a statistical graph from the relationship clusters generated in step S41; S43. Establish a probability graph; S44. Combine the statistical graph and the probability graph to obtain a relationship recommendation graph. The relationship recommendation graph judges the probability that there is a foreground relationship between any two objects in the picture in the form of probability; S45. Combine the uncertainty estimation of the relationship of the constructed mixed-initiative learning framework based on deep evidence learning to obtain the final sampling result.

7. The method for generating a scene graph based on deep evidence active learning according to claim 6, wherein In the step S41, the relationship clustering is expressed as where o p and o q are objects belonging to the categories p and q respectively, r represents the relationship between the objects, and the clustering size is determined by the relationship annotations existing in the dataset.

8. The method for generating a scene graph based on deep evidence active learning according to claim 7, characterized in that, In step S43, when establishing a probability graph, it is defined that where o p , o m , o q are objects belonging to class p, class m, and class q respectively, and r1 and r2 are two foreground relationships. Then the probability expression that there is a foreground relationship between object class p and object class q is: Where k is the number of object classifications in the dataset.

9. A scene graph generation system based on deep evidence active learning, characterized in that, It includes a memory, a processor, and a communication circuit. The memory and the communication circuit are respectively coupled to the processor. The communication circuit is connected to the processor. The communication circuit performs data interaction with an external terminal device under the control of the processor; the memory includes local storage and stores a computer program; the processor is used to run the computer program to execute the scene graph generation method based on deep evidence active learning according to any one of claims 1-8.

Citation Information

Patent Citations

  • Scene graph generation method based on transformer model and category association

    CN114782791A

  • Generating scene graphs from digital images using external knowledge and image reconstruction

    US20200401835A1