Method and apparatus for generating image scene information
By detecting and encoding the target object features in the image, and generating the relationship features between the target object groups, the problem of insufficient modeling of scene graph generation context in the prior art is solved, and the accuracy of image scene information and the interpretability of the generation process are improved.
Patent Information
- Application Number
- CN202111323685.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-08
AI Technical Summary
The existing scene graph generation method does not model the context information sufficiently, resulting in insufficient accuracy of the generated image scene information.
By detecting the target object in the image, obtaining its detection results and feature information, using the attention network for context encoding, generating the relationship characteristics between the target object groups, building a scene map, and using the training sample set to adjust the relationship deviation to improve the accuracy of the model.
The accuracy of image scene information and interpretability of the generation process are improved, the problem of insufficient context modeling in the prior art is solved, and the robustness of the model is enhanced.
Smart Images

Figure CN114067196B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technologies, and in particular, to a method and apparatus for generating image scene information. Background Art
[0002] Scene graph generation technology is an important technology for computers to understand image information and is mainly applied to multimedia information analysis. Specifically, scene graph generation technology analyzes the input image data, obtains the target objects in the image, and analyzes the relationships between the target objects, so as to abstract the image into a directed graph. This graph structure describing the image scene information is called a scene graph. The existing scene graph generation methods do not model the context information sufficiently and there is room for improvement. Summary of the Invention
[0003] The embodiments of the present application propose a method and apparatus for generating image scene information.
[0004] In a first aspect, the embodiments of the present application provide a method for generating image scene information, including: detecting target objects in the acquired image to be processed, and obtaining the detection results and feature information of each target object; obtaining the context-aware feature of each target object according to the detection results and feature information of each target object; obtaining a relationship feature representing the relationships between the target objects in the target object group in the image to be processed according to the context-aware features of the target objects in the target object group in the image to be processed; and generating the image scene information corresponding to the image to be processed according to the relationship feature corresponding to each target object group.
[0005] In some embodiments, the detection result includes location information and classification information; and the obtaining the context-aware feature of each target object according to the detection results and feature information of each target object includes: for each target object, performing the following operations: splicing the location information, classification information, and feature information of the target object to obtain the object feature of the target object; performing a linear transformation on the object feature to obtain the transformed object feature; and performing context encoding on the transformed object feature through a first attention network to obtain the context-aware feature of the target object.
[0006] In some embodiments, the obtaining a relationship feature representing the relationships between the target objects in the target object group in the image to be processed according to the context-aware features of the target objects in the target object group in the image to be processed includes: for each target object group, performing the following operations: splicing the context-aware features of the target objects in the target object group to obtain the object group feature of the target object group; performing a linear transformation on the object group feature to obtain the transformed object group feature; and performing context encoding on the transformed object group feature through a second attention network to obtain a relationship feature representing the relationships between the target objects in the target object group.
[0007] In some embodiments, the context-aware features of the target objects in the target object group are concatenated to obtain the object group features of the target object group, including: determining the bounding box information of the target objects in the target object group included in the to-be-processed image; concatenating the context-aware features and the bounding box information of the target objects in the target object group to obtain the object group features of the target object group.
[0008] In some embodiments, generating the image scene information corresponding to the to-be-processed image according to the relationship features corresponding to each target object group includes: determining, by a relationship classification network, the relationships between the target objects in each target object group according to the relationship features of each target object group; generating a scene graph representing the scene information in the to-be-processed image according to the relationships between the target objects in each target object group.
[0009] In a second aspect, an embodiment of the present application provides a method for generating image scene information, including: obtaining a training sample set, where the training samples in the training sample set include sample images, object labels representing the target objects in the sample images, and relationship labels representing the relationships between the target objects in the target object groups in the sample images; detecting the target objects in the sample images to obtain the detection results and feature information of each target object; obtaining the context-aware features of each target object through a first attention network according to the detection results and feature information of each target object; obtaining the relationship features representing the relationships between the target objects in each target object group through a second attention network according to the context-aware features of the target objects in each target object group; using the relationship features as the input of the relationship classification network, and using the object labels and relationship labels corresponding to the input training samples as the expected outputs of the relationship classification network, and training to obtain a scene information generation network including the first attention network, the second attention network, and the relationship classification network.
[0010] In some embodiments, using the relationship features as the input of the relationship classification network, and using the object labels and relationship labels corresponding to the input training samples as the expected outputs of the relationship classification network, and training to obtain a scene information generation network including the first attention network, the second attention network, and the relationship classification network includes: determining the resistance deviation of each relationship included in the predicted training sample according to the distribution information of the training samples in the training sample set; using the relationship features as the input of the relationship classification network to obtain a classification result; correcting the classification result corresponding to each relationship through the resistance deviation corresponding to each relationship to obtain a corrected classification result; and training to obtain a scene information generation network based on the loss between the object labels, relationship labels corresponding to the input sample images, and the corrected classification result.
[0011] In some embodiments, determining the resistance deviation of each relationship included in the predicted training samples according to the distribution information of the training samples in the training sample set includes: combining the preset hyperparameters for adjusting the resistance magnitude, and determining the resistance deviation of each relationship included in the predicted training samples according to the distribution information of the training samples in the training sample set.
[0012] In some embodiments, determining the resistance deviation of each relationship included in the predicted training samples according to the distribution information of the training samples in the training sample set includes: determining the resistance deviation corresponding to each relationship according to the proportion of the training samples corresponding to each relationship in the training sample set; or determining the resistance deviation corresponding to each relationship according to the normalization result of the number of target object groups corresponding to each relationship; or for each relationship corresponding to a target object group, determining the resistance deviation corresponding to each relationship involved in the target object group according to the proportion of the training samples involved in the target object group belonging to the relationship in all the training samples involved in the target object group; or for each relationship corresponding to a target object group, determining the resistance deviation corresponding to each relationship involved in the target object group according to the proportion of the estimated number of the training samples involved in the target object group belonging to the relationship in the total estimated number of the training samples involved in the target object group, where the estimated number represents the general distribution information of the training samples involved in the target object group belonging to the relationship.
[0013] In some embodiments, the estimated number is determined in the following manner: determining the estimated number of the training samples involved in the target object group under the relationship according to the number of the training samples involved by the subject object in the target object group under the relationship and the number of the training samples involved by the object object in the target object group under the relationship.
[0014] In a third aspect, an embodiment of the present application provides a device for generating image scene information, including: a first detection unit configured to detect target objects in the acquired image to be processed, and obtain the detection results and feature information of each target object; a first feature processing unit configured to obtain the context-aware feature of each target object according to the detection results and feature information of each target object; a second feature processing unit configured to obtain the relationship feature representing the relationship between the target objects in the target object group in the image to be processed according to the context-aware features of the target objects in the target object group; a generation unit configured to generate the image scene information corresponding to the image to be processed according to the relationship feature corresponding to each target object group.
[0015] In some embodiments, the detection result includes location information and classification information; and a first feature processing unit, which is further configured to: for each target object, perform the following operations: splice the location information, classification information, and feature information of the target object to obtain an object feature of the target object; perform a linear transformation on the object feature to obtain a transformed object feature; perform context encoding on the transformed object feature through a first attention network to obtain a context-aware feature of the target object.
[0016] In some embodiments, a second feature processing unit, which is further configured to: for each target object group, perform the following operations: splice the context-aware features of the target objects in the target object group to obtain an object group feature of the target object group; perform a linear transformation on the object group feature to obtain a transformed object group feature; perform context encoding on the transformed object group feature through a second attention network to obtain a relationship feature representing the relationships between the target objects in the target object group.
[0017] In some embodiments, a second feature processing unit, which is further configured to: determine bounding box information of the target objects in the target object group included in the image to be processed; splice the context-aware features of the target objects in the target object group and the bounding box information to obtain an object group feature of the target object group.
[0018] In some embodiments, a generating unit, which is further configured to: through a relationship classification network, determine the relationships between the target objects in each target object group according to the relationship features of each target object group; generate a scene graph representing the scene information in the image to be processed according to the relationships between the target objects in each target object group.
[0019] Fourthly, an embodiment of the present application provides a device for generating image scene information, including: an acquisition unit configured to acquire a training sample set, where the training samples in the training sample set include sample images, object labels representing target objects in the sample images, and relationship labels representing the relationships between target objects in the target object groups in the sample images; a second detection unit configured to detect target objects in the sample images to obtain detection results and feature information of each target object; a first attention unit configured to obtain context-aware features of each target object through a first attention network according to the detection results and feature information of each target object; a second attention unit configured to obtain relationship features representing the relationships between target objects in each target object group through a second attention network according to the context-aware features of the target objects in each target object group; and a training unit configured to use the relationship features as the input of a relationship classification network, and use the object labels and relationship labels corresponding to the input training samples as the expected outputs of the relationship classification network, and train to obtain a scene information generation network including a first attention network, a second attention network, and a relationship classification network.
[0020] In some embodiments, the training unit is further configured to: determine a resistance deviation for predicting each relationship included in the training samples according to the distribution information of the training samples in the training sample set; use the relationship features as the input of the relationship classification network to obtain a classification result; correct the classification result corresponding to each relationship through the resistance deviation corresponding to each relationship to obtain a corrected classification result; and train to obtain the scene information generation network based on the loss between the object labels, relationship labels corresponding to the input sample images, and the corrected classification result.
[0021] In some embodiments, the training unit is further configured to: combine a preset hyperparameter for adjusting the resistance magnitude and determine a resistance deviation for predicting each relationship included in the training samples according to the distribution information of the training samples in the training sample set.
[0022] In some embodiments, the training unit is further configured to: determine the resistance deviation corresponding to each relationship according to the proportion of the training samples corresponding to each relationship in the training sample set; or determine the resistance deviation corresponding to each relationship according to the normalization result of the number of target object groups corresponding to each relationship; or for each relationship corresponding to a target object group, determine the resistance deviation corresponding to each relationship involved in the target object group according to the proportion of the training samples involved in the target object group belonging to the relationship in all the training samples involved in the target object group; or for each relationship corresponding to a target object group, determine the resistance deviation corresponding to each relationship involved in the target object group according to the proportion of the estimated number of the training samples involved in the target object group belonging to the relationship in the total estimated number of the training samples involved in the target object group, where the estimated number represents the general distribution information of the training samples involved in the target object group belonging to the relationship.
[0023] In some embodiments, the estimated number is determined in the following manner: determine the estimated number of the training samples involved in the target object group under the relationship according to the number of the training samples involved in the subject object in the target object group under the relationship and the number of the training samples involved in the object object in the target object group under the relationship.
[0024] In a fifth aspect, an embodiment of the present application provides a computer-readable medium, on which a computer program is stored, where when the program is executed by a processor, the method described in any implementation manner of the first aspect or the second aspect is implemented.
[0025] In a sixth aspect, an embodiment of the present application provides an electronic device, including: one or more processors; a storage device, on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect or the second aspect.
[0026] The method and apparatus for generating image scene information provided by the embodiments of the present application detect target objects in the acquired image to be processed, and obtain the detection results and feature information of each target object; according to the detection results and feature information of each target object, obtain the context-aware features of each target object; according to the context-aware features of the target objects in the target object group in the image to be processed, obtain the relationship features representing the relationships between the target objects in each target object group; and generate the image scene information corresponding to the image to be processed according to the relationship features corresponding to each target object group. Thus, a method is provided for first determining the context features of the target objects in the image to be processed, and then determining the relationship features representing the relationships between the target objects in the target object group to generate scene graph information, which fully models the target objects and the relationships between the target objects, and improves the accuracy of the image scene information determined by the relationship features. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Other features, objects, and advantages of the present application will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0028] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present application can be applied;
[0029] Figure 2 is a flowchart of an embodiment of the method for generating image scene information according to the present application;
[0030] Figure 3 is a schematic diagram of an application scenario of the method for generating image scene information according to this embodiment;
[0031] Figure 4 is a flowchart of an embodiment of the method for generating image scene information according to the present application;
[0032] Figure 5 is a schematic structural diagram of the scene information generation model according to the present application;
[0033] Figure 6 is a structural diagram of an embodiment of the apparatus for generating image scene information according to the present application;
[0034] Figure 7 is a structural diagram of an embodiment of the apparatus for generating image scene information according to the present application;
[0035] Figure 8 is a schematic structural diagram of a computer system suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the relevant invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only the parts related to the relevant invention are shown in the drawings.
[0037] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0038] Figure 1 An exemplary architecture 100 of a method and apparatus for generating image scene information to which the present application can be applied, and a method and apparatus for generating image scene information is shown.
[0039] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The terminal devices 101, 102, 103 are communicatively connected to form a topology network, and the network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0040] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 may be hardware devices or software that support network connections for data interaction and data processing. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices that support network connections, information acquisition, interaction, display, processing, etc., including but not limited to smart phones, tablet computers, e-book readers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they can be installed in the above-listed electronic devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or can also be implemented as a single software or software module. No specific limitation is made here.
[0041] Server 105 can be a server that provides various services. For example, for the to-be-processed images provided by terminal devices 101, 102, and 103, it first determines the context features of the target objects in the to-be-processed images, and then determines the relationship features that characterize the relationships between the target objects in the target object group to generate image scene information. The server can also train a scene information generation network for generating image scene information based on a training sample set. The scene information generation network includes a first attention network for determining the context features of the target objects in the to-be-processed images, a second attention network for determining the relationship features that characterize the relationships between the target objects in the target object group, and a relationship classification network for determining the relationships based on the relationship features. As an example, server 105 can be a cloud server.
[0042] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules (such as software or software modules for providing distributed services) or as a single software or software module. Specific limitations are not made here.
[0043] It also should be noted that the method for generating image scene information provided in the embodiments of the present application can be executed by the server, can be executed by the terminal device, or can be executed by the server and the terminal device in cooperation with each other. Correspondingly, the device for generating image scene information and each part (such as each unit) included in the device for generating image scene information can be all set in the server, can be all set in the terminal device, or can be respectively set in the server and the terminal device.
[0044] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0045] Continuing to refer to Figure 2 , a flowchart 200 of an embodiment of the method for generating image scene information is shown, including the following steps:
[0046] Step 201, detect the target objects in the acquired to-be-processed image to obtain the detection results and feature information of each target object.
[0047] In this embodiment, the execution subject of the method for generating image scene information (such asFigure 1 The terminal device or server in ) can obtain the image to be processed from a remote or local location based on a wired or wireless connection method, and detect the target objects in the obtained image to be processed, obtaining the detection results and feature information of each target object.
[0048] The image to be processed can be an image in various scenarios and including any content. For example, the image to be processed can be an image in specific application scenarios such as intelligent robots, autonomous driving, and visual impairment assistance. The target object can be an object such as a person or an object included in the image to be processed.
[0049] In this embodiment, the above-mentioned execution entity can input the image to be processed into a pre-trained target detection network. The feature extraction network in the target detection network extracts the feature information of the target objects in the image to be processed, and the result output network determines the detection results of the target objects according to the feature information.
[0050] Among them, the target detection network can be any network model with the detection function of target objects. For example, the target detection network can be a CNN (Convolutional Neural Networks, convolutional neural network), an RNN (Recurrent Neural Network, recurrent neural network), or a Faster RCNN (fast recurrent convolutional neural network).
[0051] Step 202: Obtain the context-aware features of each target object according to the detection results and feature information of each target object.
[0052] In this embodiment, the above-mentioned execution entity can obtain the context-aware features of each target object according to the detection results and feature information of each target object. The context-aware features represent the features of the context awareness corresponding to the target object (obtaining the context information of the target object).
[0053] As an example, the above-mentioned execution entity can perform context encoding on each target object based on the detection results and feature information of each target object to obtain the context-aware features of each target object.
[0054] In some optional implementation manners of this embodiment, the detection results include the location information and classification information of the target object. The above-mentioned execution entity can execute the above step 202 in the following manner:
[0055] For each target object, perform the following operations:
[0056] First, splice the location information, classification information, and feature information of the target object to obtain the object feature of the target object.
[0057] Then, perform a linear transformation on the object features to obtain the transformed object features.
[0058] As an example, denote the set of all target objects detected in the image to be processed as V. For any target object v ∈ V, the transformed object features are obtained through the following formula:
[0059] e v = W o [pos(b v ),g v ,ebd(c v )]
[0060] where pos(b v ) represents the position encoding of the target object, g v represents the feature information of the target object, ebd(c v ) represents the word embedding vector generated according to the classification result of the target object, [·, ·] represents the concatenation operation, and W o is the linear transformation layer.
[0061] Finally, perform context encoding on the transformed object features through the first attention network to obtain the context-aware features of the target object.
[0062] Based on the transformed object features {e v , v ∈ V} of all target objects, use a stacked Transformer (attention network) to perform context encoding on the target objects to obtain the context-aware features corresponding to each target object
[0063] Step 203, according to the context-aware features of the target objects in the target object group in the image to be processed, obtain relationship features representing the relationships between the target objects in each target object group.
[0064] In this embodiment, the above execution subject can obtain relationship features representing the relationships between the target objects in each target object group according to the context-aware features of the target objects in the target object group in the image to be processed.
[0065] The image to be processed includes at least one target object group, and each target object group includes two target objects in the image to be processed. There is a certain relationship between the two target objects. For example, the relationship between the two target objects in the target image group is "on", "has" relationship, or a more precise and information-rich relationship, such as "parked on", "carrying".
[0066] As an example, the above-mentioned execution entity can perform context encoding on each target object group based on the context-aware features of two target objects in the target object group, so as to obtain the relationship features of each target object group.
[0067] In some alternative implementation manners of this embodiment, the above-mentioned execution entity can execute the above step 203 in the following manner:
[0068] For each target object group, perform the following operations:
[0069] First, splice the context-aware features of the target objects in the target object group to obtain the object group features of the target object group.
[0070] Then, perform a linear transformation on the object group features to obtain the transformed object group features.
[0071] Finally, perform context encoding on the transformed object group features through a second attention network to obtain relationship features representing the relationships between the target objects in the target object group.
[0072] In some alternative implementation manners of this embodiment, the above-mentioned execution entity can obtain the object group features in the following manner:
[0073] First, determine the bounding box information of the target objects in the target object group included in the image to be processed; then, splice the context-aware features of the target objects in the target object group and the bounding box information to obtain the object group features of the target object group. As an example, the bounding box information may be the information of the minimum bounding box including the target objects in the target object group.
[0074] Specifically, for any target object group (s, o) composed of a subject object s ∈ V and an object object o ∈ V, the corresponding transformed object group features are:
[0075]
[0076] Among them, g U(s,o) represents the bounding box information, represents the context-aware feature of the subject object, represents the context-aware feature of the object object, [·, ·] represents the splicing operation, and W r represents the linear transformation layer.
[0077] Based on the transformed object group features {e s,o , s ∈ V, o ∈ V} of all target object groups, use another set of stacked Transformers to perform context encoding on the transformed object group features to obtain the context-aware relationship features of each target object group
[0078] Step 204: Generate image scene information corresponding to the image to be processed according to the relationship features corresponding to each target object group.
[0079] In this embodiment, generate image scene information corresponding to the image to be processed according to the relationship features corresponding to each target object group.
[0080] As an example, the above-mentioned execution entity can map the relationship features to the corresponding relationship categories through a mapping network to obtain the relationships corresponding to each target object group; count the relationships corresponding to each target object group to obtain the image scene information corresponding to the image to be processed.
[0081] In some alternative implementation manners of this embodiment, the above-mentioned execution entity can execute the above-mentioned step 204 in the following manner:
[0082] First, determine the relationships between the target objects in each target object group according to the relationship features of each target object group through a relationship classification network; then, generate a scene graph representing the scene information in the image to be processed according to the relationships between the target objects in each target object group.
[0083] As an example, the above-mentioned execution entity can determine the relationships corresponding to each target object group through the following formula:
[0084]
[0085] where represents the relationship classification network, represents the relationship features corresponding to the target object group, σ represents the softmax function, and p s,o represents the predicted probability distribution of the relationship between the finally obtained target objects s and o.
[0086] As an example, the above-mentioned execution entity can determine the relationship with the highest probability in the probability distribution as the relationship corresponding to the target object group.
[0087] In this embodiment, the above-mentioned execution entity can construct a scene graph in a manner that nodes represent the target objects in the image to be processed and the directed edges between the nodes represent the relationships between the target objects.
[0088] It should be noted that the information processing process shown in the above steps 201-204 can be executed by the scene information generation model obtained in the subsequent Embodiment 400.
[0089] Continue to refer to Figure 3 , Figure 3 is a schematic diagram 300 of an application scenario of the method for generating image scene information according to this embodiment. In Figure 3In an application scenario, the server 301 obtains the image to be processed 303 from the terminal device 302. After obtaining the image to be processed 303, first, the target objects in the obtained image to be processed are detected to obtain the detection results and feature information of each target object. Then, according to the detection results and feature information of each target object, the context-aware features of each target object are obtained. Then, according to the context-aware features of the target objects in the target object group in the image to be processed, the relationship features representing the relationships between the target objects in each target object group are obtained. Finally, according to the relationship features corresponding to each target object group, the image scene information corresponding to the image to be processed is generated.
[0090] The method provided by the above embodiment of the present application detects the target objects in the obtained image to be processed to obtain the detection results and feature information of each target object; according to the detection results and feature information of each target object, the context-aware features of each target object are obtained; according to the context-aware features of the target objects in the target object group in the image to be processed, the relationship features representing the relationships between the target objects in each target object group are obtained; according to the relationship features corresponding to each target object group, the image scene information corresponding to the image to be processed is generated, thereby providing a method for first determining the context features of the target objects in the image to be processed, and then determining the relationship features representing the relationships between the target objects in the target object group to generate scene graph information, fully modeling the target objects and the relationships between the target objects, and improving the accuracy of the image scene information determined by the relationship features.
[0091] Moreover, since the context features of the target objects in the image to be processed are first determined, and then the relationship features representing the relationships between the target objects in the target object group are determined to generate scene graph information, based on each operation step, the interpretability and robustness of the generation process of the image scene information are improved.
[0092] Continuing to refer to Figure 4 , a schematic flowchart 400 of an embodiment of a method for generating image scene information according to the present application is shown, including the following steps:
[0093] Step 401, obtain a training sample set.
[0094] In this embodiment, the execution subject of the method for generating image scene information (for example, Figure 1 the server or terminal device in
[0095] Step 402: Detect the target objects in the sample image to obtain the detection results and feature information of each target object.
[0096] In this embodiment, the above-mentioned execution entity can detect the target objects in the sample image to obtain the detection results and feature information of each target object.
[0097] As an example, the above-mentioned execution entity can input the sample image into a pre-trained object detection network. The feature extraction network in the object detection network extracts the feature information of the target objects in the image to be processed, and the result output network determines the detection results of the target objects according to the feature information.
[0098] Step 403: Through the first attention network, obtain the context-aware features of each target object according to the detection results and feature information of each target object.
[0099] In this embodiment, the above-mentioned execution entity can obtain the context-aware features of each target object through the first attention network according to the detection results and feature information of each target object.
[0100] As an example, the first attention network can be a multi-head and stacked attention module.
[0101] Specifically, for each target object, the following operations are performed:
[0102] First, splice the location information, classification information, and feature information of the target object to obtain the object feature of the target object. Then, perform a linear transformation on the object feature to obtain the transformed object feature. Finally, perform context encoding on the transformed object feature through the first attention network to obtain the context-aware feature of the target object.
[0103] Step 404: Through the second attention network, obtain the relationship features representing the relationships between the target objects in each target object group according to the context-aware features of the target objects in each target object group.
[0104] In this embodiment, the above-mentioned execution entity can obtain the relationship features representing the relationships between the target objects in each target object group through the second attention network according to the context-aware features of the target objects in each target object group.
[0105] As an example, the second attention network can be a multi-head and stacked attention module.
[0106] Specifically, for each target object group, the above-mentioned execution entity performs the following operations: First, splice the context-aware features of the target objects in the target object group and the bounding box information of the target objects in the target object group included in the to-be-processed image to obtain the object group features of the target object group. Then, perform a linear transformation on the object group features to obtain the transformed object group features. Finally, perform context encoding on the transformed object group features through the second attention network to obtain the relationship features representing the relationships between the target objects in the target object group.
[0107] Step 405: Use the relationship features as the input of the relationship classification network, and use the object label and relationship label corresponding to the input training sample as the expected output of the relationship classification network, and train to obtain a scene information generation network including a first attention network, a second attention network, and a relationship classification network.
[0108] In this embodiment, the above-mentioned execution entity can use the relationship features as the input of the relationship classification network, and use the object label and relationship label corresponding to the input training sample as the expected output of the relationship classification network, and train to obtain a scene information generation network including a first attention network, a second attention network, and a relationship classification network.
[0109] As an example, the above-mentioned execution entity can use the relationship features as the input of the relationship classification network to obtain a classification result; furthermore, based on the loss between the object label, relationship label corresponding to the input sample image and the corrected classification result, update the first attention network, the second attention network, and the relationship classification network until a preset end condition is reached, and train to obtain a scene information generation network.
[0110] Among them, the preset end condition can be, for example, that the training time exceeds the time threshold, the number of training times exceeds the number threshold, and the loss tends to converge.
[0111] In some optional implementation manners of this embodiment, in order to make the entire training process more coordinated and make the finally obtained model have high accuracy, the above-mentioned execution entity can also incorporate the target detection network for determining the detection result and feature information into the model update process, so as to update the scene information generation network including the target detection network, the first attention network, the second attention network, and the relationship classification network according to the loss.
[0112] In some optional implementation manners of this embodiment, the above-mentioned execution entity can execute the above step 405 in the following manner:
[0113] First, according to the distribution information of the training samples in the training sample set, determine the resistance deviation of each relationship included in the predicted training sample.
[0114] Each training sample may include multiple target object groups, and there may be various relationships among the target objects in each target object group. For example, the sample image includes person A and horse B, and the various relationships between person A and horse B can be: the person leads the horse, and the horse is in front of the person.
[0115] In this implementation, the resistance deviation is used to provide resistance for the prediction process of each relationship included in the training sample, so as to improve the recognition accuracy of the relationship by the finally obtained scene information generation model.
[0116] As an example, the above-mentioned execution entity can determine the resistance deviation for predicting each relationship included in the training sample based on the principle that the distribution information is negatively correlated with the resistance deviation.
[0117] Second, use the relationship features as the input of the relationship classification network to obtain the classification result.
[0118] Third, correct the classification result corresponding to each relationship through the resistance deviation corresponding to each relationship to obtain the corrected classification result.
[0119] As an example, for each relationship, subtract the resistance deviation corresponding to the relationship from the classification result of the relationship to obtain the corrected classification result.
[0120] Fourth, train to obtain the scene information generation network based on the loss between the object label, relationship label corresponding to the input sample image and the corrected classification result.
[0121] Specifically, the above-mentioned execution entity can obtain gradient information according to the loss, and then update the scene information generation network according to the gradient information.
[0122] In the prior art, mainly the labeled training sample set is used to train the scene information generation network, but there is a problem of unbalanced sample distribution in the labeled training sample set, and the prominent manifestation is the long-tail distribution problem of the relationships between target objects. Specifically speaking, due to the differences in the difficulty of data collection and the labeling tendencies of labelers, a small number of common relationships or coarse-grained relationship descriptions appear in large quantities in the data set. This part of the relationships is called the head relationships, and they dominate the sample distribution of the entire training sample set, while the data volume of other relationships is relatively small and is called the tail relationships.
[0123] For example, in the labeled data of the widely used Visual Genome (VG) database, there are a large number of "on" and "has" relationships, while more precise and information-rich relationships such as "parked on" and "carrying" appear relatively less in the labeled data. This unbalanced data distribution leads to a serious skewness problem in the model trained based on this data, thereby reducing the quality of the scene graph generated by the model and making it unable to meet the requirements of practical applications.
[0124] In this implementation, the use of resistance deviation avoids the skewness problem of the scene information generation model obtained by training caused by the unbalanced data distribution in the training sample set, and improves the accuracy of the model.
[0125] In some optional implementation manners of this embodiment, the above execution subject may execute the above first step in the following manner:
[0126] Combined with the preset hyperparameters for adjusting the resistance magnitude, according to the distribution information of the training samples in the training sample set, determine the resistance deviation of each relationship included in the predicted training samples.
[0127] By presetting the hyperparameters for adjusting the resistance magnitude, the resistance deviation of each relationship can be adjusted more flexibly based on actual requirements, improving the flexibility of the training process.
[0128] In some optional implementation manners of this embodiment, the above execution subject may determine the resistance deviation of each relationship in any of the following manners:
[0129] Method 1: Determine the resistance deviation corresponding to each relationship according to the proportion of the training samples corresponding to each relationship in the training sample set.
[0130] Method 2: Determine the resistance deviation corresponding to each relationship according to the normalization result of the number of target object groups corresponding to each relationship.
[0131] Method 3: For each relationship corresponding to each target object group, according to the proportion of the training samples involved in the target object group belonging to this relationship in all the training samples involved in the target object group, determine the resistance deviation of each relationship involved in the target object group.
[0132] Method 4: For each type of relationship corresponding to each target object group, determine the resistance deviation corresponding to each type of relationship involved in the target object group according to the proportion of the estimated number of training samples involved in the target object group belonging to this type of relationship in the total estimated number of training samples involved in the target object group. Among them, the estimated number characterizes the general distribution information of the training samples involved in the target object group belonging to this type of relationship.
[0133] Specifically, count the training samples in the training sample set and generate the resistance deviation according to the statistical data. The basic calculation form of the resistance deviation is as follows:
[0134]
[0135] Among them, C r represents the set of relationship categories, ω i represents the weight of relationship i in the training sample set, and α and ∈ represent preset hyperparameters used to adjust the resistance deviation.
[0136] By adjusting α and ∈, the debiasing effect of resistance training can be conveniently adjusted. For example, by default, α = 1 and ∈ = 0.001 can be taken. When α decreases from 1 to 0, or ∈ increases from 0.001 to 1, the intensity of the resistance gradually decreases, and the debiasing effect of resistance training gradually weakens.
[0137] In order to describe the distribution of the training sample set from different perspectives, multiple methods are designed to describe the weight of each relationship in the training sample set:
[0138] Corresponding to the above Method 1, the above execution entity considers the influence of the proportion of each relationship in the data on model training. The relationship with a smaller proportion is more difficult to identify. Initialize the weight of each relationship with the proportion of the number of samples of each relationship category in the total training sample set to obtain the resistance deviation based on the number of samples.
[0139] Corresponding to the above Method 2, in addition to the difference in the number of samples of relationships caused by the difference in sample acquisition difficulty, the long-tail distribution is also reflected in that the coarse-grained description annotations are much more than the finer ones. In order to reflect the differences between annotations of different granularities in the long-tail distribution, the resistance deviation based on the target object group corresponding to each relationship is designed. If there is a relationship I between the target objects in the target object group (s, o), then (s, o) is a target object group corresponding to relationship I. The result of normalizing the number of target object groups of each relationship is used as the weight of each relationship to obtain the resistance deviation based on the number of valid combinations.
[0140] Considering that relationship prediction is not only related to the relationship category but also closely related to the target object group, the resistance deviation is further refined here, and different resistance deviations are assigned to different relationships of different target object groups. The resistance deviation can be calculated by the following formula:
[0141]
[0142] where ω s,o,i is the weight of the triple (s, o, i) in the training sample set, where s, o, and i correspond to the label of the subject object, the label of the object object, and the relationship label respectively. Based on the above formula, the calculation methods of the resistance deviation in the above Method 3 and Method 4 are provided.
[0143] Corresponding to the above Method 3, for each given target object group (s, o), the proportion of the number of samples of each relationship category i involved in the target object group (s, o) to the total number of samples of the target object group is used as the weight of the corresponding triple (s, o, i), and the resistance deviation based on the number of samples of the target object group is obtained.
[0144] Corresponding to the above Method 4, since many target object groups have only very few samples in the training sample set, the distribution of each relationship of these target object groups cannot reflect the general distribution of the relationships. For each given target object group (s, o), the proportion of the estimated number of samples of each relationship category i to the total estimated number of the target object group is used as the weight of the corresponding triple (s, o, i), and the resistance deviation based on the number of samples of the target object group is obtained.
[0145] As an example, the estimated number can be the estimated number obtained based on a large amount of data statistics to reflect the general distribution of the relationships, so that its weight is more in line with the actual scenario.
[0146] In some optional time modes of this embodiment, the estimated number is determined by the following method: According to the number of training samples involved in the subject object in the target object group under this relationship and the number of training samples involved in the object object in the target object group under this relationship, the estimated number of training samples involved in the target object group under this relationship is determined.
[0147] Specifically, the general distribution of the relationships between target objects is estimated by the following subject-predicate - predicate-object method:
[0148]
[0149] where C e represents the target object set, and n s,o,i is the number of samples of the triple (s, o, i) in the training sample set. n s,o′,iThe number of training samples involved in the subject object s in the target object group under relationship i, n s,o′,i Only the subject object is limited, and the object object is not limited; n s′,o,i The number of training samples involved in the object object o in the target object group under relationship i, n s′,o,i Only the object object is limited, and the subject object is not limited.
[0150] Furthermore, the weight of each triple is determined through the following formula to obtain the resistance deviation of the distribution of the target object group based on quantity estimation:
[0151]
[0152] In this implementation manner, multiple ways to determine the resistance deviation are provided. During the specific training process, it can be specifically selected according to the actual situation, further improving the flexibility and accuracy of the training process.
[0153] Continue to refer to Figure 5 , which shows a specific scene information generation model 500. The scene information generation model 500 includes a target detection network 501, a first attention network 502, a second attention network 503, and a relationship classification network 504.
[0154] The method provided by the above embodiments of the present application, by obtaining a training sample set, wherein the training samples in the training sample set include sample images, object labels representing the target objects in the sample images, and relationship labels representing the relationships between the target objects in the target object groups in the sample images; detecting the target objects in the sample images to obtain the detection results and feature information of each target object; through the first attention network, according to the detection results and feature information of each target object, obtaining the context-aware features of each target object; through the second attention network, according to the context-aware features of the target objects in each target object group, obtaining the relationship features representing the relationships between the target objects in each target object group; using the relationship features as the input of the relationship classification network, and using the object labels and relationship labels corresponding to the input training samples as the expected outputs of the relationship classification network, training to obtain a scene information generation network including the first attention network, the second attention network, and the relationship classification network, provides a training method for a scene information generation model, fully models the target objects and the relationships between the target objects, and improves the accuracy of the model.
[0155] Moreover, since the context features of the target objects in the image are first determined, and then the relationship features representing the relationships between the target objects in the target object groups are determined to generate scene graph information, based on each operation step, the interpretability and robustness of the model's generation process of image scene information are improved.
[0156] Continue to refer to Figure 6 , as an implementation of the methods shown in the above figures, an embodiment of an apparatus for generating image scene information is provided in the present application. This apparatus embodiment corresponds to Figure 2 the method embodiment shown, and this apparatus can be specifically applied to various electronic devices.
[0157] As Figure 6 shown, the apparatus for generating image scene information includes: a first detection unit 601, configured to detect target objects in the acquired image to be processed, and obtain the detection results and feature information of each target object; a first feature processing unit 602, configured to obtain the context-aware feature of each target object according to the detection results and feature information of each target object; a second feature processing unit 603, configured to obtain a relationship feature representing the relationship between the target objects in each target object group according to the context-aware features of the target objects in the target object group in the image to be processed; a generation unit 604, configured to generate the image scene information corresponding to the image to be processed according to the relationship feature corresponding to each target object group.
[0158] In some optional implementation manners of this embodiment, the detection result includes position information and classification information; and the first feature processing unit 602 is further configured to: for each target object, perform the following operations: splice the position information, classification information, and feature information of the target object to obtain the object feature of the target object; perform a linear transformation on the object feature to obtain the transformed object feature; perform context encoding on the transformed object feature through a first attention network to obtain the context-aware feature of the target object.
[0159] In some optional implementation manners of this embodiment, the second feature processing unit 603 is further configured to: for each target object group, perform the following operations: splice the context-aware features of the target objects in the target object group to obtain the object group feature of the target object group; perform a linear transformation on the object group feature to obtain the transformed object group feature; perform context encoding on the transformed object group feature through a second attention network to obtain the relationship feature representing the relationship between the target objects in the target object group.
[0160] In some optional implementation manners of this embodiment, the second feature processing unit 603 is further configured to: determine the bounding box information of the target objects in the image to be processed that include the target objects in the target object group; splice the context-aware features of the target objects in the target object group and the bounding box information to obtain the object group feature of the target object group.
[0161] In some alternative implementation manners of this embodiment, the generating unit 604 is further configured to: determine the relationships between the target objects in each target object group according to the relationship features of each target object group through a relationship classification network; generate a scene graph representing the scene information in the image to be processed according to the relationships between the target objects in each target object group.
[0162] In this embodiment, the first detection unit in the apparatus for generating image scene information detects the target objects in the acquired image to be processed, and obtains the detection results and feature information of each target object; the first feature processing unit obtains the context-aware feature of each target object according to the detection results and feature information of each target object; the second feature processing unit obtains the relationship feature representing the relationships between the target objects in the target object group in the image to be processed according to the context-aware features of the target objects in the target object group; the generating unit generates the image scene information corresponding to the image to be processed according to the relationship features corresponding to each target object group, thereby providing an apparatus that first determines the context features of the target objects in the image to be processed, and then determines the relationship features representing the relationships between the target objects in the target object group to generate scene information, which can fully model the target objects and the relationships between the target objects, and improves the accuracy of the image scene information determined by the relationship features.
[0163] Continue to refer to Figure 7 , as an implementation of the methods shown in the above figures, this application provides an embodiment of an apparatus for generating image scene information. This apparatus embodiment corresponds to Figure 4 the method embodiment shown, and this apparatus can be specifically applied to various electronic devices.
[0164] As Figure 7As shown in the figure, the apparatus for generating image scene information includes: an acquisition unit 701 configured to acquire a training sample set, where the training samples in the training sample set include sample images, object labels representing target objects in the sample images, and relationship labels representing the relationships between the target objects in the target object groups in the sample images; a second detection unit 702 configured to detect the target objects in the sample images to obtain the detection results and feature information of each target object; a first attention unit 703 configured to obtain the context-aware features of each target object through a first attention network according to the detection results and feature information of each target object; a second attention unit 704 configured to obtain the relationship features representing the relationships between the target objects in each target object group through a second attention network according to the context-aware features of the target objects in each target object group; and a training unit 705 configured to use the relationship features as the input of a relationship classification network, and use the object labels and relationship labels corresponding to the input training samples as the expected output of the relationship classification network, and train to obtain a scene information generation network including a first attention network, a second attention network, and a relationship classification network.
[0165] In some optional implementation manners of this embodiment, the training unit 705 is further configured to: determine the resistance deviation of each relationship included in the predicted training sample according to the distribution information of the training samples in the training sample set; use the relationship features as the input of the relationship classification network to obtain a classification result; correct the classification result corresponding to each relationship through the resistance deviation corresponding to each relationship to obtain a corrected classification result; and train to obtain a scene information generation network based on the loss between the object labels, relationship labels corresponding to the input sample images, and the corrected classification result.
[0166] In some optional implementation manners of this embodiment, the training unit 705 is further configured to: combine a preset hyperparameter for adjusting the resistance magnitude, and determine the resistance deviation of each relationship included in the predicted training sample according to the distribution information of the training samples in the training sample set.
[0167] In some alternative implementation manners of this embodiment, the training unit 705 is further configured to: determine the resistance deviation corresponding to each relationship according to the proportion of the training samples corresponding to each relationship in the training sample set; or determine the resistance deviation corresponding to each relationship according to the normalization result of the number of target object groups corresponding to each relationship; or for each relationship corresponding to each target object group, determine the resistance deviation corresponding to each relationship involved in the target object group according to the proportion of the training samples involved in the target object group belonging to the relationship in all the training samples involved in the target object group; or for each relationship corresponding to each target object group, determine the resistance deviation corresponding to each relationship involved in the target object group according to the proportion of the estimated number of the training samples involved in the target object group belonging to the relationship in the total estimated number of the training samples involved in the target object group, where the estimated number represents the general distribution information of the training samples involved in the target object group belonging to the relationship.
[0168] In some alternative implementation manners of this embodiment, the estimated number is determined in the following manner: determine the estimated number of the training samples involved in the target object group under the relationship according to the number of the training samples involved in the subject object in the target object group under the relationship and the number of the training samples involved in the object object in the target object group under the relationship.
[0169] In this embodiment, the acquisition unit in the device for generating image scene information acquires a training sample set, where the training samples in the training sample set include sample images, object labels representing the target objects in the sample images, and relationship labels representing the relationships between the target objects in the target object groups in the sample images; the second detection unit detects the target objects in the sample images to obtain the detection results and feature information of each target object; the first attention unit obtains the context-aware features of each target object through a first attention network according to the detection results and feature information of each target object; the second attention unit obtains the relationship features representing the relationships between the target objects in each target object group through a second attention network according to the context-aware features of the target objects in each target object group; the training unit uses the relationship features as the input of the relationship classification network, and uses the object labels and relationship labels corresponding to the input training samples as the expected outputs of the relationship classification network, and trains to obtain a scene information generation network including the first attention network, the second attention network, and the relationship classification network, providing a training method for a scene information generation model, fully modeling the target objects and the relationships between the target objects, and improving the accuracy of the model.
[0170] The following refers to Figure 8 , which shows a device suitable for implementing the embodiments of the present application (such as Figure 1Schematic diagram of the structure of computer system 800 of devices 101, 102, 103, 105 shown. Figure 8 The devices shown are merely examples and should not impose any limitations on the functions and scope of use of the embodiments of this application.
[0171] As Figure 8 shown, computer system 800 includes a processor (e.g., CPU, central processing unit) 801, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 802 or a program loaded from storage section 808 into random access memory (RAM) 803. In RAM 803, various programs and data required for the operation of system 800 are also stored. Processor 801, ROM 802, and RAM 803 are connected to each other via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0172] The following components are connected to I / O interface 805: input section 806 including a keyboard, a mouse, etc.; output section 807 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and speakers, etc.; storage section 808 including a hard disk, etc.; and communication section 809 including a network interface card such as a LAN card, a modem, etc. Communication section 809 performs communication processing via a network such as the Internet. Drive 810 is also connected to I / O interface 805 as required. Removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on drive 810 as required so that a computer program read from it can be installed into storage section 808 as required.
[0173] Specifically, according to the embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of this application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by processor 801, the above functions defined in the methods of this application are executed.
[0174] It should be noted that the computer-readable medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0175] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the client computer, partially on the client computer, executed as a stand-alone software package, partially on the client computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the client computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the boxes may occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0177] The units involved in the embodiments described in the present application can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor including a first detection unit, a first feature processing unit, a second feature processing unit, and a generation unit, or it can also be described as: a processor including an acquisition unit, a second detection unit, a first attention unit, a second attention unit, and a training unit. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself. For example, the training unit can also be described as "a unit that uses the relationship feature as the input of the relationship classification network, uses the object label and relationship label corresponding to the input training sample as the expected output of the relationship classification network, and trains to obtain a scene information generation network including a first attention network, a second attention network, and a relationship classification network".
[0178] As another aspect, the present application also provides a computer-readable medium, which may be included in the device described in the above embodiments; or may exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the computer device is caused to: detect target objects in the acquired image to be processed, and obtain the detection results and feature information of each target object; obtain the context-aware features of each target object according to the detection results and feature information of each target object; obtain relationship features representing the relationships between the target objects in each target object group according to the context-aware features of the target objects in the target object group in the image to be processed; generate image scene information corresponding to the image to be processed according to the relationship features corresponding to each target object group. The computer device is also caused to: obtain a training sample set, wherein the training samples in the training sample set include sample images, object labels representing the target objects in the sample images, and relationship labels representing the relationships between the target objects in the target object group in the sample images; detect the target objects in the sample images, and obtain the detection results and feature information of each target object; obtain the context-aware features of each target object according to the detection results and feature information of each target object through a first attention network; obtain relationship features representing the relationships between the target objects in each target object group according to the context-aware features of the target objects in each target object group through a second attention network; use the relationship features as the input of a relationship classification network, and use the object labels and relationship labels corresponding to the input training samples as the expected outputs of the relationship classification network, and train to obtain a scene information generation network including a first attention network, a second attention network, and a relationship classification network.
[0179] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present application.
Claims
1. A method for generating image scene information, comprising: Detecting target objects in the acquired image to be processed, and obtaining the detection results and feature information of each target object; Obtaining the context-aware feature of each target object through a first attention network according to the detection results and feature information of each target object; Obtaining the relationship feature representing the relationship between the target objects in each target object group in the image to be processed through a second attention network according to the context-aware features of the target objects in each target object group; Determining the relationship between the target objects in each target object group through a relationship classification network according to the relationship feature corresponding to each target object group; Generating the image scene information corresponding to the image to be processed according to the relationship between the target objects in each target object group; Wherein, the scene information generation network including the first attention network, the second attention network and the relationship classification network is trained in the following manner: Obtaining a training sample set, wherein the training samples in the training sample set include sample images, object labels representing the target objects in the sample images, and relationship labels representing the relationship between the target objects in the target object groups in the sample images; Detecting the target objects in the sample images, and obtaining the detection results and feature information of each target object; Obtaining the context-aware feature of each target object through the first attention network according to the detection results and feature information of each target object; Obtaining the relationship feature representing the relationship between the target objects in each target object group through the second attention network according to the context-aware features of the target objects in each target object group; Determining the resistance deviation of each relationship included in the predicted training sample according to the distribution information of the training samples in the training sample set, wherein the resistance deviation is used to provide resistance for the prediction process of each relationship included in the training sample; Taking the relationship feature as the input of the relationship classification network to obtain a classification result; Correcting the classification result corresponding to each relationship through the resistance deviation corresponding to each relationship to obtain a corrected classification result; Training the scene information generation network based on the loss between the object labels, relationship labels corresponding to the input sample images and the corrected classification results.
2. The method according to claim 1, wherein, The detection result includes position information and classification information; And The obtaining the context-aware feature of each target object through the first attention network according to the detection results and feature information of each target object includes: For each target object, perform the following operations: Concatenating the position information, classification information and feature information of the target object to obtain the object feature of the target object; Performing a linear transformation on the object feature to obtain a transformed object feature; Performing context encoding on the transformed object feature through the first attention network to obtain the context-aware feature of the target object.
3. The method according to claim 1, wherein, The obtaining the relationship feature representing the relationship between the target objects in each target object group through the second attention network according to the context-aware features of the target objects in each target object group in the image to be processed includes: For each target object group, perform the following operations: Concatenate the context-aware features of the target objects in the target object group to obtain the object group feature of the target object group; Perform a linear transformation on the object group feature to obtain the transformed object group feature; Perform context encoding on the transformed object group feature through the second attention network to obtain a relationship feature representing the relationships between the target objects in the target object group.
4. The method according to claim 3, wherein The operation of concatenating the context-aware features of the target objects in the target object group to obtain the object group feature of the target object group includes: Determine the bounding box information of the target objects in the target object group included in the to-be-processed image; Concatenate the context-aware features of the target objects in the target object group and the bounding box information to obtain the object group feature of the target object group.
5. A method for generating image scene information, including: Obtain a training sample set, where the training samples in the training sample set include sample images, object labels representing the target objects in the sample images, and relationship labels representing the relationships between the target objects in the target object groups in the sample images; Detect the target objects in the sample images to obtain the detection results and feature information of each target object; Through a first attention network, obtain the context-aware feature of each target object according to the detection results and feature information of each target object; Through a second attention network, obtain a relationship feature representing the relationships between the target objects in each target object group according to the context-aware features of the target objects in each target object group; Use the relationship feature as the input of a relationship classification network, and use the object label and relationship label corresponding to the input training sample as the expected output of the relationship classification network, and train to obtain a scene information generation network including the first attention network, the second attention network, and the relationship classification network, including: According to the distribution information of the training samples in the training sample set, determine the resistance deviation of each relationship included in the predicted training sample, where the resistance deviation is used to provide resistance for the prediction process of each relationship included in the training sample; Use the relationship feature as the input of the relationship classification network to obtain a classification result; Correct the classification result corresponding to each relationship through the resistance deviation corresponding to each relationship to obtain the corrected classification result; Based on the loss between the object label, relationship label corresponding to the input sample image and the corrected classification result, train to obtain the scene information generation network.
6. The method according to claim 5, wherein, The operation of determining the resistance deviation of each relationship included in the predicted training sample according to the distribution information of the training samples in the training sample set includes: Combine a preset hyperparameter for adjusting the resistance magnitude, and determine the resistance deviation of each relationship included in the predicted training sample according to the distribution information of the training samples in the training sample set.
7. The method according to claim 5 or 6, wherein The operation of determining the resistance deviation of each relationship included in the predicted training sample according to the distribution information of the training samples in the training sample set includes: Determine the resistance deviation corresponding to each relationship according to the proportion of the training samples corresponding to each relationship in the training sample set; or Determine the resistance deviation corresponding to each relationship according to the normalization result of the number of target object groups corresponding to each relationship; or For each relationship corresponding to each target object group, determine the resistance deviation corresponding to each relationship involved in the target object group according to the proportion of the training samples involved in the target object group belonging to the relationship in all the training samples involved in the target object group; or For each relationship corresponding to each target object group, determine the resistance deviation corresponding to each relationship involved in the target object group according to the proportion of the estimated number of the training samples involved in the target object group belonging to the relationship in the total estimated number of the training samples involved in the target object group, where the estimated number represents the general distribution information of the training samples involved in the target object group belonging to the relationship.
8. The method according to claim 7, wherein The estimated number is determined by the following method: Determine the estimated number of the training samples involved in the target object group under the relationship according to the number of the training samples involved in the subject object in the target object group under the relationship and the number of the training samples involved in the object object in the target object group under the relationship.
9. An apparatus for generating image scene information, comprising: A first detection unit configured to detect target objects in the acquired image to be processed, and obtain the detection result and feature information of each target object; A first feature processing unit configured to obtain the context-aware feature of each target object through a first attention network according to the detection result and feature information of each target object; A second feature processing unit configured to obtain the relationship feature representing the relationship between the target objects in each target object group through a second attention network according to the context-aware features of the target objects in the target object group in the image to be processed; A generating unit configured to determine the relationship between the target objects in each target object group through a relationship classification network according to the relationship feature corresponding to each target object group; Generate the image scene information corresponding to the image to be processed according to the relationship between the target objects in each target object group; Wherein, the scene information generation network including the first attention network, the second attention network and the relationship classification network is trained by the following method: Obtain a training sample set, where the training samples in the training sample set include sample images, object labels representing target objects in the sample images, and relationship labels representing the relationships between target objects in the target object groups in the sample images; Detect target objects in the sample images to obtain the detection results and feature information of each target object; Through the first attention network, according to the detection results and feature information of each target object, obtain the context-aware feature of each target object; Through the second attention network, according to the context-aware features of the target objects in each target object group, obtain relationship features representing the relationships between the target objects in each target object group; According to the distribution information of the training samples in the training sample set, determine the resistance deviation for each relationship included in the predicted training sample, where the resistance deviation is used to provide resistance for the prediction process of each relationship included in the training sample; Use the relationship features as the input of the relationship classification network to obtain a classification result; Correct the classification result corresponding to each relationship through the resistance deviation corresponding to each relationship to obtain a corrected classification result; Based on the loss between the object labels, relationship labels corresponding to the input sample images and the corrected classification results, train to obtain the scene information generation network.
10. An apparatus for generating image scene information, comprising: An acquisition unit configured to acquire a training sample set, where the training samples in the training sample set include sample images, object labels representing target objects in the sample images, and relationship labels representing the relationships between target objects in the target object groups in the sample images; A second detection unit configured to detect target objects in the sample images to obtain the detection results and feature information of each target object; A first attention unit configured to, through a first attention network, according to the detection results and feature information of each target object, obtain the context-aware feature of each target object; A second attention unit configured to, through a second attention network, according to the context-aware features of the target objects in each target object group, obtain relationship features representing the relationships between the target objects in each target object group; A training unit, configured to use the relationship feature as the input of a relationship classification network, and use the object label and relationship label corresponding to the input training sample as the expected output of the relationship classification network, and train to obtain a scene information generation network including the first attention network, the second attention network, and the relationship classification network, including: determining a resistance deviation for each relationship included in the predicted training sample according to the distribution information of the training samples in the training sample set, where the resistance deviation is used to provide resistance for the prediction process of each relationship included in the training sample; using the relationship feature as the input of the relationship classification network to obtain a classification result; correcting the classification result corresponding to each relationship through the resistance deviation corresponding to each relationship to obtain a corrected classification result; and training to obtain the scene information generation network based on the loss between the object label, relationship label corresponding to the input sample image, and the corrected classification result.
11. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by a processor, it implements the method according to any one of claims 1-8.
12. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-8.
Citation Information
Patent Citations
Scene map generation method, device and equipment
CN111931928A