Scene graph generation method, device and intelligent agent

By performing object detection on the target scene image, the interaction behavior relationship and relative position relationship of the target object are obtained, and the target scene map is generated, the problem of inability to fully characterize the object relationship in the existing technology is solved, and the accuracy of the execution of the agent task is improved.

CN119379777BActive Publication Date: 2025-05-16BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411985858.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The prior art is difficult to fully characterize the interaction behavioral relationship and relative position relationship between multiple target objects, resulting in insufficient accuracy of the agent in task execution.

Method used

By acquiring the first acquired image of the target scene, object detection is performed to obtain the interactive behavior relationship and relative position relationship of the multiple target objects, and a target scene map is generated based on the first detection result and the second detection result.

Benefits of technology

The generated target scene map can simultaneously characterize the interactive behavioral relationship and relative position relationship between multiple target objects, improving the accuracy of the agent to complete tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119379777B_ABST
    Figure CN119379777B_ABST
Patent Text Reader

Abstract

The present application discloses a scene graph generation method, device and intelligent agent, which belongs to the field of intelligent agent technology. The method is applied to an intelligent agent, and the method includes: acquiring a first captured image of a target scene; performing target detection on the first captured image to obtain multiple target objects in the target scene, and acquiring a first detection result and a second detection result of the multiple target objects; wherein the first detection result is used to characterize the interactive behavior relationship between the multiple target objects, and the second detection result is used to characterize the relative position relationship between the multiple target objects; based on the first detection result and the second detection result, a target scene graph corresponding to the target scene is generated. The obtained target scene graph can simultaneously characterize the interactive behavior relationship and the relative position relationship between multiple target objects, and the characterized object relationship is more comprehensive, which can improve the accuracy of the intelligent agent in completing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of intelligent agent technology, and in particular, relates to a scene graph generation method, device and intelligent agent. Background Art

[0002] With the development of artificial intelligence technology, intelligent agents are widely used in many fields such as medicine, finance, manufacturing, and entertainment. Intelligent agents can automatically handle complex and repetitive tasks instead of humans according to set algorithm rules. Intelligent agents can also interact with users to provide personalized services to users. When intelligent agents perform tasks in specific scenes or based on objects in specific scenes, comprehensive determination of the relationship between objects in the scene helps to improve the accuracy of the intelligent agent in completing tasks. At present, the object relationships determined by the methods for determining the relationship between objects in the scene are relatively simple. Summary of the invention

[0003] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a scene graph generation method, device and intelligent agent, and the target scene graph obtained can simultaneously represent the interactive behavior relationship and relative position relationship between multiple target objects, and the represented object relationship is more comprehensive, which can improve the accuracy of the intelligent agent in completing the task.

[0004] In a first aspect, the present application provides a scene graph generation method, the method being applied to an intelligent agent, the method comprising:

[0005] Acquire a first captured image of a target scene;

[0006] Performing target detection on the first collected image to obtain a plurality of target objects in the target scene, and acquiring a first detection result and a second detection result of the plurality of target objects, wherein the first detection result is used to characterize an interactive behavior relationship between the plurality of target objects, and the second detection result is used to characterize a relative position relationship between the plurality of target objects;

[0007] Based on the first detection result and the second detection result, generating a target scene graph corresponding to the target scene;

[0008] The first detection result is detected by the following steps:

[0009] Acquire a content description text corresponding to the first acquired image, wherein the content description text includes a plurality of description objects and interaction behavior relationships between the plurality of description objects;

[0010] Performing matching detection on the target object and the description object;

[0011] When it is determined that the target object matches the description object, the target object is substituted into the description object in the content description text to obtain the first detection result.

[0012] According to the scene graph generation method of the present application, target detection is performed on a first captured image of a target scene to obtain multiple target objects in the target scene, the interaction behavior between the multiple target objects is detected to obtain a first detection result, and the relative positions between the multiple target objects are detected to obtain a second detection result. The target scene graph is obtained by combining the first detection result and the second detection result. The obtained target scene graph can simultaneously characterize the interaction behavior relationship and relative position relationship between the multiple target objects, and the characterized object relationship is more comprehensive, which can improve the accuracy of the intelligent agent in completing tasks.

[0013] According to an embodiment of the present application, obtaining a content description text corresponding to the first acquired image includes:

[0014] The first collected image is input into an image description model, the description object is detected by the image description model, and the interactive behavior relationship between multiple description objects is detected to obtain the content description text output by the image description model.

[0015] According to an embodiment of the present application, the matching detection of the target object and the description object includes:

[0016] Inputting the content description text and the first collected image into an image positioning model, and locating the area where the described object is located in the first collected image based on the content description text by the image positioning model;

[0017] When the degree of overlap between the region where the description object is located and the region where the target object is located is greater than a degree of overlap threshold, it is determined that the target object matches the description object.

[0018] According to an embodiment of the present application, the matching detection of the target object and the description object includes:

[0019] Inputting the content description text and the first collected image into an image positioning model, and locating the area where the described object is located in the first collected image based on the content description text by the image positioning model;

[0020] The object corresponding to the area where the described object is located is used as the target object matching the described object.

[0021] According to an embodiment of the present application, the matching detection of the target object and the description object includes:

[0022] Calculating the text similarity between the object label corresponding to the target object and the description text corresponding to the description object;

[0023] When the text similarity is greater than a text similarity threshold, it is determined that the target object matches the description object.

[0024] According to one embodiment of the present application, the target scene graph is expressed in the form of a three-dimensional probability matrix, the first dimension of the three-dimensional probability matrix represents the target object, the second dimension of the three-dimensional probability matrix represents the target object, and the third dimension of the three-dimensional probability matrix represents the interaction behavior and relative position.

[0025] According to one embodiment of the present application, the second detection result is detected by the following steps:

[0026] Inputting the first collected image into a target neural network model, detecting the relative position relationship between the plurality of target objects through the target neural network model, and obtaining the second detection result output by the target neural network model;

[0027] The target neural network model is trained by the following steps:

[0028] Acquire three-dimensional image training data and two-dimensional image training data, wherein the three-dimensional image training data and the two-dimensional image training data are annotated with a plurality of sample objects and relative positional relationships between the plurality of sample objects;

[0029] Training a first neural network model based on the three-dimensional image training data, and training a second neural network model based on the two-dimensional image training data, wherein the second neural network model acquires features in the first neural network model for learning, and based on the learning results, updates the training parameters of the first neural network model;

[0030] The trained first neural network model is determined as the target neural network model.

[0031] In a second aspect, the present application provides a scene graph generation device, the device is arranged in an intelligent agent, and the device comprises:

[0032] An acquisition module, used for acquiring a first captured image of a target scene;

[0033] A first processing module is used to perform target detection on the first collected image to obtain multiple target objects in the target scene, and obtain first detection results and second detection results of the multiple target objects, wherein the first detection result is used to characterize the interactive behavior relationship between the multiple target objects, and the second detection result is used to characterize the relative position relationship between the multiple target objects;

[0034] A second processing module, configured to generate a target scene graph corresponding to the target scene based on the first detection result and the second detection result;

[0035] The first detection result is detected by the following steps:

[0036] Acquire a content description text corresponding to the first acquired image, wherein the content description text includes a plurality of description objects and interaction behavior relationships between the plurality of description objects;

[0037] Performing matching detection on the target object and the description object;

[0038] When it is determined that the target object matches the description object, the target object is substituted into the description object in the content description text to obtain the first detection result.

[0039] According to the scene graph generation device of the present application, target detection is performed on a first captured image of the target scene to obtain multiple target objects in the target scene, the interactive behavior between the multiple target objects is detected to obtain a first detection result, and the relative positions between the multiple target objects are detected to obtain a second detection result. The target scene graph is obtained by combining the first detection result and the second detection result. The obtained target scene graph can simultaneously characterize the interactive behavior relationship and relative position relationship between the multiple target objects, and the characterized object relationship is more comprehensive, which can improve the accuracy of the intelligent agent in completing the task.

[0040] In a third aspect, the present application provides an intelligent agent, the intelligent agent comprising:

[0041] A scene graph generating device as described in the second aspect above.

[0042] According to the intelligent agent of the present application, target detection is performed on a first captured image of a target scene to obtain multiple target objects in the target scene, interactive behaviors between the multiple target objects are detected to obtain a first detection result, and relative positions between the multiple target objects are detected to obtain a second detection result. The first detection result and the second detection result are combined to obtain a target scene graph. The obtained target scene graph can simultaneously characterize the interactive behavior relationship and relative position relationship between the multiple target objects, and the characterized object relationship is more comprehensive, which can improve the accuracy of the intelligent agent in completing tasks.

[0043] In a fourth aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the scene graph generation method as described in the first aspect above is implemented.

[0044] In a fifth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the scene graph generation method as described in the first aspect above is implemented.

[0045] In a sixth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the scene graph generation method as described in the first aspect above.

[0046] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0048] Figure 1 It is one of the flow diagrams of the scene graph generation method provided in the embodiment of the present application;

[0049] Figure 2 This is the second flow chart of the scene graph generation method provided in the embodiment of the present application;

[0050] Figure 3 This is the third flow chart of the scene graph generation method provided in the embodiment of the present application;

[0051] Figure 4 is a structural schematic diagram of a scene graph generating device provided in an embodiment of the present application;

[0052] Figure 5 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0054] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0055] In the following, in combination with the accompanying drawings, the scene graph generation method, scene graph generation device, electronic device and readable storage medium provided in the embodiments of the present application are described in detail through specific embodiments and their application scenarios.

[0056] Among them, the scene graph generation method can be applied to intelligent agents.

[0057] An intelligent agent is a system that can perceive its environment and perform actions to affect the environment. An intelligent agent can be a software program, a hardware device, or a system that combines software and hardware.

[0058] For example, intelligent agents can be industrial robots, service robots, self-driving cars, drones, virtual assistants, and electronic pets.

[0059] The scene graph generation method provided in the embodiment of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the scene graph generation method. The electronic devices mentioned in the embodiment of the present application include but are not limited to mobile phones, computers, cameras, wearable devices, etc. The scene graph generation method provided in the embodiment of the present application is described below using the electronic device as an example of the execution subject.

[0060] like Figure 1 As shown, the scene graph generation method includes: step 110, step 120 and step 130.

[0061] Step 110: Acquire a first captured image of the target scene.

[0062] The target scene is the scene for which the scene graph is to be generated. The target scene may be the scene corresponding to the perspective of the camera of the intelligent agent, or may be any scene specified by the user.

[0063] For example, the target scene can be a family life scene, an industrial scene, a medical scene, a natural landscape, an emergency rescue scene, and a game scene.

[0064] The first acquired image is an image obtained by capturing the target scene through an image acquisition device such as a camera or a camera. The first acquired image can be a static picture or a dynamic video frame.

[0065] In this step, the intelligent agent can be provided with an image acquisition device, the intelligent agent can move to the target scene, and acquire the first acquisition image through its own image acquisition device, or acquire the first acquisition image through an image acquisition device external to the intelligent agent, and transmit the acquired first acquisition image to the intelligent agent.

[0066] Step 120: Perform target detection on the first collected image to obtain multiple target objects in the target scene, and obtain first detection results and second detection results of the multiple target objects.

[0067] The first detection result is used to characterize the interactive behavior relationship between the multiple target objects, and the second detection result is used to characterize the relative position relationship between the multiple target objects.

[0068] In this embodiment, target detection is a process of identifying a target object in a first acquired image and determining a category and position of the target object. Target detection is performed on the first acquired image to identify each target object and determine the category of the target object. A bounding box is drawn for each identified target object to indicate the position and size of the object.

[0069] Among them, the target objects can be people and objects in the target scene. For example, in a home scene, the target objects can include sofas, televisions, dining tables, water cups, pets, and people, etc. In an autonomous driving scene, the target objects can include other vehicles, pedestrians, traffic lights, and road signs, etc.

[0070] In actual implementation, target detection may be performed on the first acquired image through a target detection model.

[0071] In this embodiment, after obtaining a plurality of target objects, the interaction behaviors between the plurality of target objects are detected to obtain a first detection result, and the relative positions between the plurality of target objects are detected to obtain a second detection result.

[0072] Among them, interactive behaviors are interactive actions or interactive events between target objects. Interactive behaviors may include physical contact, spatial proximity, object transfer, and visual interaction. Interactive behavior relationships may characterize the dynamic interactive relationships between multiple target objects in the target scene.

[0073] The relative position may be the distance between target objects, the relative direction between target objects, or a combination of distance and direction. The relative position relationship may characterize the static spatial position relationship between multiple target objects in the target scene.

[0074] In this step, the first detection result may be detected by an interactive behavior detection model, and the second detection result may be detected by a relative position detection model.

[0075] Step 130: Generate a target scene graph corresponding to the target scene based on the first detection result and the second detection result.

[0076] Among them, the scene graph is a semantic graph structure that can be represented by nodes and edges. Objects correspond to nodes, and the relationships between objects correspond to edges.

[0077] The target scene graph can be used to describe the target objects in the target scene, as well as the interactive behavior relationships and relative position relationships between the target objects.

[0078] In this step, the first detection result and the second detection result may be in the form of a data structure of a scene graph, and the first detection result and the second detection result may be merged to obtain a target scene graph.

[0079] According to the scene graph generation method provided in the embodiment of the present application, target detection is performed on a first captured image of the target scene to obtain multiple target objects in the target scene, the interaction behavior between the multiple target objects is detected to obtain a first detection result, and the relative positions between the multiple target objects are detected to obtain a second detection result. The target scene graph is obtained by combining the first detection result and the second detection result. The obtained target scene graph can simultaneously characterize the interaction behavior relationship and relative position relationship between the multiple target objects, and the characterized object relationship is more comprehensive, which can improve the accuracy of the intelligent agent in completing the task.

[0080] In some embodiments, the first detection result is detected by the following steps:

[0081] Obtaining a content description text corresponding to the first captured image;

[0082] Match and detect the target object with the described object;

[0083] When it is determined that the target object matches the description object, the target object is substituted into the description object in the content description text to obtain a first detection result.

[0084] The content description text includes multiple description objects and interactive behavior relationships between the multiple description objects. The content description text may be a text in a natural language form. The content description text may express interactive behaviors between the description objects. The description objects are people and objects.

[0085] For example, the content description text may be a girl holding a cartoon cup and a man pointing at a black TV, wherein the description objects include the girl, the cartoon cup, the man and the black TV. A girl holding a cartoon cup indicates the interaction between the girl and the cartoon cup, and a man pointing at a black TV indicates the interaction between the man and the black TV.

[0086] In this embodiment, the first collected image may be analyzed by performing image recognition, template matching, and combining context information to identify the object types and interaction behaviors between objects in the first collected image and generate content description text.

[0087] In this embodiment, the description object in the content description text can be a target object, and the interactive behavior relationship between the description objects can be an interactive behavior relationship between the target objects. The description object and the target object are matched and detected. When the description object and the target object match, the description object can be considered to be the target object. The target object is substituted into the description object in the content description text to obtain the interactive behavior relationship between the target objects, thereby determining the first detection result.

[0088] For example, the content description text is a girl holding a cartoon cup, the description objects include the girl and the cartoon cup, the target objects include the cup, the child and the kitten, and the description object and the target object are matched and detected to determine that the cup in the target object matches the cartoon cup in the description object, and the child in the target object matches the girl in the description object. The cup is substituted into the cartoon cup in the content description text, and the child is substituted into the girl in the content description text, then the first detection result is that the child is holding a cup.

[0089] In this embodiment, the matching detection between the target object and the description object can be achieved by comparing the features of the target object and the description object, or by pattern matching.

[0090] In some embodiments, obtaining a content description text corresponding to the first collected image includes:

[0091] The first captured image is input into the image description model, the description object is detected by the image description model, and the interactive behavior relationship between multiple description objects is detected to obtain the content description text output by the image description model.

[0092] Among them, the image caption model is a trained model that can convert visual information into natural language text and imitate humans' understanding and description of image content. The image caption model can capture detailed information in the image and generate more accurate description text.

[0093] In this embodiment, the first captured image is input into the image description model. The image description model can identify and locate objects in the first captured image through a built-in object detection algorithm, assign a category label to each object, and determine which objects interact and how they interact through human-object interaction detection (HOI), generate triple-tuple data of (person, action, object), and the image description model outputs content description text based on the triple-tuple data.

[0094] In some embodiments, matching and detecting the target object with the description object includes:

[0095] Inputting the content description text and the first collected image into an image positioning model, and locating the area where the description object is located in the first collected image based on the content description text by the image positioning model;

[0096] When the degree of overlap between the region where the description object is located and the region where the target object is located is greater than a threshold value of the degree of overlap, it is determined that the target object matches the description object.

[0097] Among them, the Image Grounding model is a model that associates sentences or phrases in natural language with objects in an image. The Image Grounding model can be used to understand which part of an image a sentence or phrase describes.

[0098] In this embodiment, the content description text and the first acquired image are input into the image positioning model, the image positioning model performs feature extraction on the content description text to obtain text features, performs feature extraction on the first acquired image to obtain image features, matches the text features with the image features, obtains a bounding box corresponding to the description object, and thereby locates the area where the description object is located in the first acquired image.

[0099] In this embodiment, the target object may correspond to a bounding box for indicating the area where the target object is located. When the degree of overlap between the bounding box of the target object and the bounding box of the description object is greater than a threshold of the degree of overlap, it is determined that the target object matches the description object.

[0100] The overlap threshold is a preset value, for example, 95%.

[0101] In some embodiments, matching and detecting the target object with the description object includes:

[0102] Inputting the content description text and the first collected image into an image positioning model, and locating the area where the description object is located in the first collected image based on the content description text by the image positioning model;

[0103] The object corresponding to the area where the described object is located is taken as the target object matching the described object.

[0104] In this embodiment, a bounding box corresponding to the described object can be obtained through an image positioning model, and the bounding box corresponding to the described object is directly used as the bounding box of the matching target object, and the object in the bounding box is used as the matching target object.

[0105] In some embodiments, matching and detecting the target object with the description object includes:

[0106] Calculate the text similarity between the object label corresponding to the target object and the description text corresponding to the description object;

[0107] When the text similarity is greater than the text similarity threshold, it is determined that the target object matches the description object.

[0108] The object label is used to indicate the category and attribute of the target object, etc., and target detection is performed on the first collected image to identify multiple target objects and annotate the target objects with object labels.

[0109] Text similarity is used to characterize the similarity between the object label corresponding to the target object and the description text corresponding to the description object.

[0110] In this embodiment, feature extraction can be performed on the object label corresponding to the target object and the description text corresponding to the description object respectively, the similarity between the features of the object label and the features of the description text can be calculated, the text similarity can be determined, and the text similarity can be compared with a preset text similarity threshold. When the text similarity is greater than the text similarity threshold, it is determined that the target object matches the description object.

[0111] For example, the object label corresponding to the target object is a water cup, and the description text corresponding to the description object is a cartoon cup. If the text similarity between the water cup and the cartoon cup is greater than the text similarity threshold, it is determined that the target object matches the description object.

[0112] In some embodiments, the target scene graph is expressed in the form of a three-dimensional probability matrix, the first dimension of the three-dimensional probability matrix represents the target object, the second dimension of the three-dimensional probability matrix represents the target object, and the third dimension of the three-dimensional probability matrix represents the interaction behavior and relative position.

[0113] In this embodiment, the target scene graph is expressed in the form of a three-dimensional probability matrix, and the three-dimensional probability matrix may be a data structure for representing a target object and the relationship between multiple target objects.

[0114] The first dimension of the three-dimensional probability matrix can be a row, and the first dimension represents each target object, and each row vector represents a target object. The second dimension of the three-dimensional probability matrix can be a column, and the second dimension can also represent each target object, and each column vector represents a target object. The third dimension can represent the interaction behaviors and relative positions between target objects. The third dimension can include interaction behaviors such as approaching, moving away, touching, and looking, as well as relative positions such as left, right, front, and back.

[0115] The data in the three-dimensional probability matrix may be probability values ​​of corresponding interaction behaviors occurring between target objects, or presenting corresponding relative positions.

[0116] In this embodiment, the output of the target scene graph can be implemented using a logical function (sigmoid) to obtain n (n-1) K, where n is the number of target objects in the target scene and K represents the number of defined types of relationships between objects.

[0117] The target scene graph is represented as a three-dimensional probability matrix, which can provide greater flexibility for the agent to perform downstream task planning.

[0118] For example, when user a interacts with the agent by voice, user a asks the agent: How many books are there on the bookshelf? The agent parses the three-dimensional probability matrix and determines that (book 1, isIn, bookshelf 0) is 0.93, i.e. the confidence of book 1 being in bookshelf 0, and determines that (book 1, isOn, bookshelf 0) is 0.91, i.e. the confidence of book 1 being on bookshelf 0. The agent generates the answer: There is a book on the bookshelf. User b asks the agent: How many books are there on the bookshelf? The agent generates the answer: There is a book on the bookshelf.

[0119] For another example, when user c asks the agent: Is there food on the plate? The agent analyzes the three-dimensional probability matrix and determines that the confidence of (apple, isIn, plate) is 0.6, that is, the apple is in the plate, and the confidence of (apple, isOn, plate) is 0.75, that is, the apple is on the plate. The agent believes that there may be a piece of food on the plate that is an apple, but it is not very sure. It plans an autonomous task to walk closer and take a closer look. By walking closer to observe, the three-dimensional probability matrix is ​​updated, and the information that the confidence of (apple, isIn, plate) is 0.9, that is, the apple is in the plate, and the confidence of (apple, isOn, plate) is 0.93, that is, the apple is on the plate, is generated. The reply: There is food on the plate, it is an apple.

[0120] In some embodiments, the second detection result is detected by the following steps:

[0121] The first collected image is input into the target neural network model, and the relative position relationship between multiple target objects is detected by the target neural network model to obtain a second detection result output by the target neural network model.

[0122] In this embodiment, the target neural network model is a trained model. The target neural network model can detect the relative position relationship between multiple target objects according to the first collected image. The target neural network model is trained by the following steps:

[0123] Acquire three-dimensional image training data and two-dimensional image training data;

[0124] Training a first neural network model based on three-dimensional image training data, and training a second neural network model based on two-dimensional image training data, wherein the second neural network model acquires features in the first neural network model for learning, and based on the learning results, updates the training parameters of the first neural network model;

[0125] The trained first neural network model is determined as the target neural network model.

[0126] Among them, the three-dimensional image training data can be data in the form of three-dimensional point cloud, and the two-dimensional image training data can be image data including two-dimensional features of width and height. The three-dimensional image training data and the two-dimensional image training data are annotated with multiple sample objects and the relative position relationship between the multiple sample objects.

[0127] The sample object is an object in a sample scene, the relative position relationship between multiple sample objects is the relative position relationship between multiple sample objects in the sample scene, and the sample scene may be the same scene as the target scene, or a scene similar to the target scene.

[0128] In this embodiment, color (RGBD) video data with depth information can be collected in the sample scene by using a depth camera external to the intelligent body or a depth camera set in the intelligent body, and a series of video frames and corresponding depth maps can be obtained. The sample scene point cloud corresponding to the sample scene is restored by using the intrinsic and extrinsic parameters of the depth camera, so as to obtain three-dimensional image training data.

[0129] The sample scene point cloud is annotated with object instance-level mask information and object relationship information, that is, the scene graph corresponding to the sample scene is used as the ground truth. Each edge in the scene graph is represented by a triple of (object a, relationship label label, object b). For each annotated triple, the set template is used to generate the corresponding natural language text description sentence to obtain annotated 3D image training data.

[0130] Projecting the three-dimensional image training data into a two-dimensional space and other operations are performed to obtain two-dimensional image training data.

[0131] The matching rules between images and texts are set. Based on each triplet and the corresponding description sentence, the sample objects that best match the description sentence are screened by visual language alignment models such as Contrastive Language-Image Pre-training (CLIP) and Bootstrapped Language-Image Pre-training 2 (BLIP2). For example, the sample object whose image features are most similar to the text description features and whose projected bounding box accounts for a larger proportion in the sample scene can be selected as the most matching sample object.

[0132] For the matched sample objects and description sentences, the corresponding language model is used to process the natural language text description sentences, making the description more natural and diverse.

[0133] In this embodiment, the second neural network model is a model for assisting training of the first neural network model. The input of the second neural network model is two-dimensional image training data. The second neural network model includes two encoder modules. The two encoder modules encode nodes and edges respectively, where nodes are objects and edges are relationships. The encoded objects and relationships are passed through the graph neural network (GNN) in the second neural network model to output two features, one for object feature and the other for relationship feature. The two features are passed through the corresponding classifiers in the second neural network model to obtain object classification results and relationship classification results.

[0134] In addition, object features and relationship features can be transformed into triplet features through the triplet-level regularization module. The triplet features can be aligned with the text features input by the pre-trained text encoder during training, which can improve the accuracy and generalization of prediction.

[0135] Among them, the pre-trained Text Encoder can be implemented using the Text Encoder module in visual language alignment models such as CLIP and BLIP2. It is frozen during training and no weight update is performed. The input text is a natural language text description sentence and a processed description sentence.

[0136] The first neural network model also includes two encoder modules, which encode nodes and edges respectively, where nodes are objects and edges are relationships. The encoded objects and relationships are output through the graph neural network (GNN) of the first neural network model, one for object feature and the other for relationship feature. The two features are passed through the corresponding classifiers in the first neural network model to obtain the classification results of the object and the classification results of the relationship.

[0137] The second neural network model obtains features from the first neural network model for learning, and the node features output by the node encoder of the first neural network model and the node features output by the node encoder of the second neural network model are fused once. The fused features are passed through the GNN structure of the second neural network model to generate object features.

[0138] The second neural network model updates the training parameters of the first neural network model according to the learning results of the features obtained in the first neural network model, and the relationfeature output by the GNN module of the first neural network model is fused with the relationfeature output by the GNN module of the second neural network model, and the fused features are then used for triple feature generation and classification prediction.

[0139] After the training is completed, the second neural network model can be removed and the trained first neural network model can be determined as the target neural network model.

[0140] In the related art, the use of two-dimensional data lacks depth information, and in some special viewing angles, it is easy to make misjudgments when judging the front and back position relationship or other spatial position relationship of the object.

[0141] The use of three-dimensional data is limited by the quality of three-dimensional data acquisition. When the point cloud of an object is incomplete, it has a greater impact on relationship recognition.

[0142] The customized rule method based on the three-dimensional bounding box requires manual design of more thresholds, and it is difficult to properly adjust some fuzzy relationships to a robust threshold.

[0143] In an embodiment of the present application, a first neural network model is trained using three-dimensional image training data, and a second neural network model is trained using two-dimensional image training data. The second neural network model acquires features in the first neural network model for learning, and based on the learning results, the training parameters of the first neural network model are updated. The trained first neural network model is determined as a target neural network model. The target neural network model can combine the advantages of two-dimensional data training and three-dimensional data training without setting too many thresholds, and can more accurately identify the spatial position relationship between objects.

[0144] The scene graph generation method provided in the embodiment of the present application can be executed by a scene graph generation device. In the embodiment of the present application, the scene graph generation device executing the scene graph generation method is taken as an example to illustrate the scene graph generation device provided in the embodiment of the present application.

[0145] An embodiment of the present application also provides a scene graph generation device, which is disposed in an intelligent agent.

[0146] like Figure 4 As shown, the scene graph generating device comprises:

[0147] An acquisition module 410 is used to acquire a first captured image of a target scene;

[0148] The first processing module 420 is used to perform target detection on the first collected image to obtain multiple target objects in the target scene, and obtain first detection results and second detection results of the multiple target objects;

[0149] The first detection result is used to characterize the interactive behavior relationship between the multiple target objects, and the second detection result is used to characterize the relative position relationship between the multiple target objects;

[0150] The second processing module 430 is used to generate a target scene graph corresponding to the target scene based on the first detection result and the second detection result.

[0151] According to the scene graph generation device provided in the embodiment of the present application, target detection is performed on a first captured image of the target scene to obtain multiple target objects in the target scene, the interaction behavior between the multiple target objects is detected to obtain a first detection result, and the relative positions between the multiple target objects are detected to obtain a second detection result. The target scene graph is obtained by combining the first detection result and the second detection result. The obtained target scene graph can simultaneously characterize the interaction behavior relationship and relative position relationship between the multiple target objects, and the characterized object relationship is more comprehensive, which can improve the accuracy of the intelligent agent in completing the task.

[0152] In some embodiments, the first processing module 420 is used to obtain a content description text corresponding to the first collected image, where the content description text includes a plurality of description objects and an interactive behavior relationship between the plurality of description objects;

[0153] Match and detect the target object with the described object;

[0154] When it is determined that the target object matches the description object, the target object is substituted into the description object in the content description text to obtain a first detection result.

[0155] In some embodiments, the first processing module 420 is used to input the first captured image into an image description model, detect the description object through the image description model, and detect the interactive behavior relationship between multiple description objects to obtain the content description text output by the image description model.

[0156] In some embodiments, the first processing module 420 is used to input the content description text and the first collected image into an image positioning model, and locate the area where the description object is located in the first collected image based on the content description text through the image positioning model;

[0157] When the degree of overlap between the region where the description object is located and the region where the target object is located is greater than a threshold value of the degree of overlap, it is determined that the target object matches the description object.

[0158] In some embodiments, the first processing module 420 is used to input the content description text and the first collected image into an image positioning model, and locate the area where the description object is located in the first collected image based on the content description text through the image positioning model;

[0159] The object corresponding to the area where the described object is located is taken as the target object matching the described object.

[0160] In some embodiments, the first processing module 420 is used to calculate the text similarity between the object label corresponding to the target object and the description text corresponding to the description object;

[0161] When the text similarity is greater than the text similarity threshold, it is determined that the target object matches the description object.

[0162] In some embodiments, the target scene graph is expressed in the form of a three-dimensional probability matrix, the first dimension of the three-dimensional probability matrix represents the target object, the second dimension of the three-dimensional probability matrix represents the target object, and the third dimension of the three-dimensional probability matrix represents the interaction behavior and relative position.

[0163] In some embodiments, the first processing module 420 is used to input the first collected image into the target neural network model, detect the relative position relationship between multiple target objects through the target neural network model, and obtain a second detection result output by the target neural network model;

[0164] Among them, the target neural network model is trained through the following steps:

[0165] Acquire three-dimensional image training data and two-dimensional image training data, wherein the three-dimensional image training data and the two-dimensional image training data are annotated with a plurality of sample objects and relative positional relationships between the plurality of sample objects;

[0166] Training a first neural network model based on three-dimensional image training data, and training a second neural network model based on two-dimensional image training data, wherein the second neural network model acquires features in the first neural network model for learning, and based on the learning results, updates the training parameters of the first neural network model;

[0167] The trained first neural network model is determined as the target neural network model.

[0168] The scene graph generating device in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or it can be other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc. It can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0169] The scene graph generation device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0170] The scene graph generation device provided in the embodiment of the present application can achieve Figures 1 to 3 To avoid repetition, the various processes implemented by the method embodiment are not described here.

[0171] The embodiment of the present application also provides an intelligent agent.

[0172] The intelligent agent includes the above-mentioned scene graph generating device.

[0173] It should be noted that the acquisition module can be connected to an image acquisition device such as a camera, and processes the color image or depth image etc. acquired by the image acquisition device to obtain a first acquired image.

[0174] The first processing module may include an object detection and segmentation module and a scene point cloud fusion module. The object detection and segmentation module may be integrated with a target detection model and a segmentation model (Segment Anything Model, SAM). The object detection and segmentation module is used to perform target detection on the first acquired image. The scene point cloud fusion module may fuse multi-frame single-view point clouds into a full-scene instance-level point cloud based on the intrinsic parameters, extrinsic parameters and depth information of the depth camera and the result of target detection, using a fusion strategy.

[0175] The first processing module may also include an interactive behavior relationship detection module and a relative position relationship detection module. The interactive behavior relationship detection module may be integrated with an image description model, and the interactive behavior relationship detection module may detect the interactive behavior relationship between target objects. The relative position relationship detection module may be integrated with a target neural network model, and the relative position relationship detection module may detect the relative position relationship between target objects.

[0176] According to the intelligent agent provided in the embodiment of the present application, target detection is performed on a first captured image of a target scene to obtain multiple target objects in the target scene, interaction behaviors between the multiple target objects are detected to obtain a first detection result, and relative positions between the multiple target objects are detected to obtain a second detection result. The target scene graph is obtained by combining the first detection result and the second detection result. The obtained target scene graph can simultaneously characterize the interaction behavior relationship and relative position relationship between the multiple target objects, and the characterized object relationship is more comprehensive, which can improve the accuracy of the intelligent agent in completing tasks.

[0177] The following introduces a specific implementation example of a method for generating a scene graph executed by an intelligent agent.

[0178] like Figure 2 As shown, step 1, obtain an RGBD video frame sequence of the target scene to obtain a first captured image.

[0179] Step 2: Input the two-dimensional RGB data corresponding to the first collected image into the object detection and segmentation module to perform target detection on the first collected image.

[0180] Step 3: The first collected image and the target detection result are input into an interactive behavior relationship detection module, and the interactive behavior relationship detection module detects the interactive behavior relationship between the target objects and outputs a first detection result.

[0181] Step 4: The target detection result and depth information are input into the scene point cloud fusion module. The scene point cloud fusion module can obtain the three-dimensional point cloud corresponding to the first acquired image according to the intrinsic parameters, extrinsic parameters and depth information of the depth camera in combination with the target detection result.

[0182] The three-dimensional point cloud corresponding to the first collected image is input into the relative position relationship detection module, the relative position relationship between the target objects is detected, and a second detection result is output.

[0183] Step 5: Combine the first detection result and the second detection result to obtain a target scene graph.

[0184] Among them, Figure 3 As shown, the execution process of the interactive behavior relationship detection module is as follows:

[0185] Step 1: Input the two-dimensional RGB data corresponding to the first acquired image into the target detection model, wherein the target detection model may be a model in an object detection and segmentation module. Inputting the two-dimensional RGB data corresponding to the first acquired image into the target detection model may correspond to inputting the two-dimensional RGB data corresponding to the first acquired image into the object detection and segmentation module, and inputting the two-dimensional RGB data corresponding to the first acquired image into the image description model. The image description model outputs content description text. The image description model may be an enhanced multimodal pre-training model (Bootstrap Language-Image Pre-training 2, BLIP-2).

[0186] Step 2: Perform a matching test on the target object and the description object in the matching test model. When it is determined that the target object and the description object match, substitute the target object into the description object in the content description text to obtain a first test result.

[0187] In some embodiments, Figure 5 As shown, an embodiment of the present application also provides an electronic device 500, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the program is executed by the processor 501, each process of the above-mentioned scene graph generation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0188] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0189] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned scene graph generation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0190] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0191] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned scene graph generation method when executed by a processor.

[0192] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0193] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned scene graph generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0194] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0195] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0196] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0197] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

[0198] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0199] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.

Claims

1. A scene graph generation method, characterized in that: The method is applied to an intelligent agent, and the method comprises: Acquire a first captured image of a target scene; Performing target detection on the first collected image to obtain a plurality of target objects in the target scene, and acquiring a first detection result and a second detection result of the plurality of target objects, wherein the first detection result is used to characterize an interactive behavior relationship between the plurality of target objects, and the second detection result is used to characterize a relative position relationship between the plurality of target objects; Based on the first detection result and the second detection result, generating a target scene graph corresponding to the target scene; The first detection result is detected by the following steps: Acquire a content description text corresponding to the first acquired image, wherein the content description text includes a plurality of description objects and interaction behavior relationships between the plurality of description objects; Performing matching detection on the target object and the description object; When it is determined that the target object matches the description object, the target object is substituted into the description object in the content description text to obtain the first detection result; The second detection result is detected by the following steps: Inputting the first collected image into a target neural network model, detecting the relative position relationship between the plurality of target objects through the target neural network model, and obtaining the second detection result output by the target neural network model; The target neural network model is trained by the following steps: Acquire three-dimensional image training data and two-dimensional image training data, wherein the three-dimensional image training data and the two-dimensional image training data are annotated with a plurality of sample objects and relative positional relationships between the plurality of sample objects; Training a first neural network model based on the three-dimensional image training data, and training a second neural network model based on the two-dimensional image training data, wherein the second neural network model acquires features in the first neural network model for learning, and based on the learning results, updates the training parameters of the first neural network model; The trained first neural network model is determined as the target neural network model.

2. The scene graph generation method according to claim 1, characterized in that: The obtaining of the content description text corresponding to the first collected image includes: The first collected image is input into an image description model, the description object is detected by the image description model, and the interactive behavior relationship between multiple description objects is detected to obtain the content description text output by the image description model.

3. The scene graph generation method according to claim 1, characterized in that: The matching detection of the target object and the description object includes: Inputting the content description text and the first collected image into an image positioning model, and locating the area where the described object is located in the first collected image based on the content description text by the image positioning model; When the degree of overlap between the region where the description object is located and the region where the target object is located is greater than a degree of overlap threshold, it is determined that the target object matches the description object.

4. The scene graph generation method according to claim 1, characterized in that: The matching detection of the target object and the description object includes: Inputting the content description text and the first collected image into an image positioning model, and locating the area where the described object is located in the first collected image based on the content description text by the image positioning model; The object corresponding to the area where the described object is located is used as the target object matching the described object.

5. The scene graph generation method according to claim 1, characterized in that: The matching detection of the target object and the description object includes: Calculating the text similarity between the object label corresponding to the target object and the description text corresponding to the description object; When the text similarity is greater than a text similarity threshold, it is determined that the target object matches the description object.

6. The scene graph generation method according to any one of claims 1 to 5, characterized in that: The target scene graph is expressed in the form of a three-dimensional probability matrix, wherein the first dimension of the three-dimensional probability matrix represents the target object, the second dimension of the three-dimensional probability matrix represents the target object, and the third dimension of the three-dimensional probability matrix represents the interaction behavior and relative position.

7. A scene graph generating device, characterized in that: The device is arranged on an intelligent body, and comprises: An acquisition module, used for acquiring a first captured image of a target scene; A first processing module is used to perform target detection on the first collected image to obtain multiple target objects in the target scene, and obtain first detection results and second detection results of the multiple target objects, wherein the first detection result is used to characterize the interactive behavior relationship between the multiple target objects, and the second detection result is used to characterize the relative position relationship between the multiple target objects; A second processing module, configured to generate a target scene graph corresponding to the target scene based on the first detection result and the second detection result; The first detection result is detected by the following steps: Acquire a content description text corresponding to the first acquired image, wherein the content description text includes a plurality of description objects and interaction behavior relationships between the plurality of description objects; Performing matching detection on the target object and the description object; When it is determined that the target object matches the description object, the target object is substituted into the description object in the content description text to obtain the first detection result; The second detection result is detected by the following steps: Inputting the first collected image into a target neural network model, detecting the relative position relationship between the plurality of target objects through the target neural network model, and obtaining the second detection result output by the target neural network model; The target neural network model is trained by the following steps: Acquire three-dimensional image training data and two-dimensional image training data, wherein the three-dimensional image training data and the two-dimensional image training data are annotated with a plurality of sample objects and relative positional relationships between the plurality of sample objects; Training a first neural network model based on the three-dimensional image training data, and training a second neural network model based on the two-dimensional image training data, wherein the second neural network model acquires features in the first neural network model for learning, and based on the learning results, updates the training parameters of the first neural network model; The trained first neural network model is determined as the target neural network model.

8. An intelligent agent, characterized in that: include: The scene graph generating device as claimed in claim 7.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the scene graph generation method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Image content retrieval method based on scene graph

    CN115952306A

  • Open word list scene graph generation method, system and equipment and storage medium

    CN116524513A