Method, apparatus, device, and storage medium for generating rail transit scene map

By using attention mechanism to mark the target area and generate models in rail transit scenarios, the problem of incomplete and inaccurate scene diagrams in the prior art is solved, and a higher-level scenario understanding is achieved.

CN114463462BActive Publication Date: 2025-07-01CHINA ACADEMY OF RAILWAY SCI CORP LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210032568.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-12
Publication Date
2025-07-01
Estimated Expiration
2042-01-12

AI Technical Summary

Technical Problem

The scene maps generated by the prior art in rail transit scenarios are incomplete and inaccurate, and fail to effectively utilize the correlation properties between objects.

Method used

The attention mechanism is used to mark the target area in the rail transit image, and a multi-layer directed rail transit scene map is constructed based on the visual relationship between the marked nodes and nodes through the generation model.

Benefits of technology

The generated rail transit scene map is more comprehensive and accurate, which can better represent the visual situation of rail transit scenes and improve the level of scene understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463462B_ABST
    Figure CN114463462B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, device, and storage medium for generating a rail transit scene map. The method includes: obtaining a rail transit image, and annotating a target area in the rail transit image according to an attention mechanism; inputting the rail transit image after annotating the target area into a generation model, and obtaining a rail transit scene map output by the generation model. The method adopted in the present invention achieves a higher-level scene understanding through the visual perception of the rail transit scene and the interaction between objects in the image, and uses the scene map to represent the visual situation of the rail transit scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rail transit systems, and particularly to a method, device, equipment and storage medium for generating a rail transit scene graph. Background Art

[0002] With the improvement of computing power and model capabilities, people's cognitive understanding of scenes is no longer limited to object detection. In recent years, some studies have focused on the visual knowledge description of the entire image, which is not only a collection of objects, object attributes, and image attributes, but also a representation of the relationships between objects. Visual cognition of rail transit scenes is one of the core technologies in the field of rail transit systems. For rail transit scenes, a higher level of scene understanding can be achieved if the interactions between objects in the image are understood.

[0003] In the prior art, some characteristics of the objects themselves in the rail transit scene are ignored, and the relevant properties existing between the objects themselves are not utilized, resulting in the problems of incomplete and inaccurate generated scene graphs. Summary of the Invention

[0004] The present invention provides a method, device, equipment and storage medium for generating a rail transit scene graph, so as to solve the defects of incomplete and inaccurate scene graphs in the prior art in the rail transit scene, and obtain a rail transit scene graph representing the visual situation of the rail transit scene.

[0005] The present invention provides a method for generating a rail transit scene graph, including:

[0006] Obtain a rail transit image, and label the target area in the rail transit image according to the attention mechanism;

[0007] Input the rail transit image after labeling the target area into a generation model, and obtain the rail transit scene graph output by the generation model;

[0008] Wherein, the generation model is used to construct the rail transit scene graph according to the nodes and the visual relationships between the nodes in the labeled target area. The rail transit scene graph includes a multi-layer directed graph, and each layer stores multiple nodes. There are edges connecting the nodes within or between the layers. The nodes are used to describe the rail transit scene information of different semantic categories, and the edges are used to describe the visual relationships between the nodes.

[0009] According to the method for generating a rail transit scene graph provided by the present invention, the labeling of the target area in the rail transit image according to the attention mechanism includes:

[0010] Obtain the area where there are entity objects in the rail transit image, and obtain the track-visual area in the rail transit image according to the attention mechanism;

[0011] Label the area where the entity object exists and the track-visual area as the target area.

[0012] According to a method for generating a rail transit scene graph provided by the present invention, the track-visual area includes an upper view area of the track, a left view area of the track, and a right view area of the track.

[0013] According to a method for generating a rail transit scene graph provided by the present invention, the visual relationships between the nodes include track topological relationships, spatial relationships, occlusion relationships, and subordination relationships.

[0014] According to a method for generating a rail transit scene graph provided by the present invention, the rail transit scene graph includes a source image layer, a track layer, a foreground layer, a background layer, and a scene layer.

[0015] According to a method for generating a rail transit scene graph provided by the present invention, the nodes of the source image layer are used to store rail transit images;

[0016] The nodes of the track layer are used to store the track subgrade, bridge and tunnel buildings under the track, and semantic information on the track embankment;

[0017] The nodes of the foreground layer are used to store the semantic information of foreground objects;

[0018] The nodes of the background layer are used to store the semantic information of the ground and structures;

[0019] The nodes of the scene layer are used to store the overall attributes of the scene.

[0020] The present invention also provides a rail transit scene graph generation device, including:

[0021] An acquisition module, configured to acquire a rail transit image and label the target area in the rail transit image according to the attention mechanism;

[0022] A generation module, configured to input the rail transit image with the labeled target area into a generation model to obtain the rail transit scene graph output by the generation model;

[0023] Wherein, the generation model is used to construct the rail transit scene graph according to the nodes and the visual relationships between the nodes in the labeled target area. The rail transit scene graph includes multiple layers and each layer stores multiple nodes. There are edges connecting the nodes within or between the layers. The nodes are used to describe rail transit scene information of different semantic categories, and the edges are used to describe the visual relationships between the nodes.

[0024] A device for generating a rail transit scene map provided by the present invention, the acquisition module includes an annotation sub-module, and the annotation sub-module is used to obtain the area with entity objects in the rail transit image, and obtain the track-visual area in the rail transit image according to the attention mechanism;

[0025] The annotation sub-module is further used to label the area with entity objects and the track-visual area as target areas.

[0026] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the rail transit scene map generation method as described in any one of the above are implemented.

[0027] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the rail transit scene map generation method as described in any one of the above are implemented.

[0028] A method, device, equipment and storage medium for generating a rail transit scene map provided by the present invention. First, compared with the general visual rail transit scene, the present invention uses fewer categories of rail transit scene elements. Secondly, the present invention proposes an effective generation model, which is suitable for the hierarchical multi-mode rail transit scene framework. Finally, based on the structure of the rail transit scene and the attention characteristics of people, the present invention improves the existing scene dataset annotation method to better describe the visual knowledge of this area and provides a more effective solution for the identification of the rail transit scene. Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 It is a flowchart of the method for generating a rail transit scene map provided by an embodiment of the present invention;

[0031] Figure 2 It is a schematic diagram of the scene hierarchy framework provided by an embodiment of the present invention;

[0032] Figure 3 It is a schematic diagram of the target area positioning provided by an embodiment of the present invention;

[0033] Figure 4 It is a rail transit scene map provided by an embodiment of the present invention;

[0034] Figure 5 It is a schematic structural diagram of a rail transit scenario map generation device provided by an embodiment of the present invention;

[0035] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments

[0036] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0037] The following combines Figures 1 - 4 to describe the rail transit scenario map generation method of the present invention, including:

[0038] Step 101: Obtain a rail transit image, and label the target area in the rail transit image according to the attention mechanism;

[0039] Step 102: Input the rail transit image after labeling the target area into the generation model, and obtain the rail transit scenario map output by the generation model;

[0040] Wherein, the generation model is used to construct the rail transit scenario map according to the nodes and the visual relationships between the nodes in the labeled target area. The rail transit scenario map includes a multi-layer directed graph, and each layer stores multiple nodes. There are edges between the nodes within or between the layers. The nodes are used to describe rail transit scenario information of different semantic categories, and the edges are used to describe the visual relationships between the nodes.

[0041] It should be noted that the generation model is a pre-constructed standard for generating a rail transit scenario map. In this embodiment, the standard is a Figure 2 scene hierarchy framework as shown. The scene hierarchy framework consists of multi-layer directed graphs. The nodes in the directed graph represent scene information of different semantic categories, and the edges represent the visual relationships between the nodes.

[0042] The rail transit scenario map generation method of the embodiment of the present invention uses fewer categories of rail transit scenario elements and the method is simple. Through the generation model, the present invention is suitable for application in a hierarchical multi-mode rail transit scenario framework. Based on the structure of the rail transit scenario and the attention characteristics of people, the present invention improves the existing scene dataset annotation method to better describe the visual knowledge of this area, providing a more effective solution for the identification of rail transit scenarios.

[0043] In at least one embodiment of the present invention, the method for annotating a target area in a rail transit image according to an attention mechanism includes:

[0044] Step 201: Obtain the area where there are entity objects in the rail transit image, and obtain the track-visual area in the rail transit image according to the attention mechanism;

[0045] Step 202: Label the area where there are entity objects and the track-visual area as the target area.

[0046] It should be noted that step 201 specifically includes the following sub-steps:

[0047] 1.1) For several categories such as railway tracks, buildings, sky, ground, etc., a "joint track-visual area positioning method" is proposed as in step 1.2.

[0048] 1.2) Starting from the most prominent part of the image, describe the image one by one to obtain area descriptions, and draw bounding boxes covering all the objects mentioned in the descriptions. In order not to lose the regional relationship, based on the track area positioning, different track numbers are positioned as track 1, track 2, etc. For the human visual attention mechanism in the rail transit scenario, the upper visual area of the track (mainly referring to the area where there are entities above the entire track area), the left visual area of the track (mainly referring to the area where there are entities at the leftmost side of the entire track area), and the right visual area of the track (mainly referring to the area where there are entities at the rightmost side of the entire track area) are added for area positioning.

[0049] 1.3) By describing the relationships in the located area and the encoded information of the relationships between objects, attributes, and data, it is used in the subsequent generation of the scene graph.

[0050] In at least one embodiment of the present invention, the track-visual area includes the upper visual area of the track, the left visual area of the track, and the right visual area of the track.

[0051] In at least one embodiment of the present invention, the visual relationships between the nodes include track topological relationships, spatial relationships, occlusion relationships, and subordination relationships.

[0052] It should be noted that the relationship of the edge is established through the following steps:

[0053] 3.1) Key classes are selected in the scene graph to describe the rail transit scene. According to the occurrence frequency of each category in various rail transit datasets, the category with the highest occurrence frequency is selected as the semantic object category for this work, and it is grouped, such as: locomotives, tracks, nature, buildings, ground, sky, etc., and stored in different layers. The key class is the category with the highest semantic frequency.

[0054] 3.2) Connect the nodes in the underlying layer to the nodes in the foreground layer and the background layer; since the track layer is the basis for other layers, connect the nodes of the track layer to the foreground layer and the background layer respectively to represent the spatial relationship; connect the structural nodes in the background layer to the ground nodes to represent the adjacency relationship, and also connect them to the nodes in the foreground layer to represent the contact relationship. Achieve the mutual connection of nodes between different layers or across different layers to achieve the unification of multi-modal data.

[0055] 3.3) Construct a relationship list suitable for the rail transit scene graph and a relationship sub-list applicable to each layer: Remove the relationships irrelevant to the rail transit scene from the 50 most common relationships in Visual Genome, and supplement and define the common relationships in the rail transit scene based on the rail semantic graph. Five main relationships are defined between different tracks: parallel relationship, intersection relationship, parallel and intersecting relationship from visual proximal to visual distal, intersecting and parallel relationship from visual proximal to visual distal, and vertical relationship in space.

[0056] 3.4) Form the scene graph by connecting each other within the layer or across layers through relationships.

[0057] In at least one embodiment of the present invention, the rail transit scene graph includes a source image layer, a track layer, a foreground layer, a background layer, and a scene layer.

[0058] In at least one embodiment of the present invention, the nodes of the source image layer are used to store rail transit images;

[0059] It should be noted that the semantic information of the underlying layer is stored on the source image layer, and the nodes of this layer are used to represent the entire source image;

[0060] The nodes of the track layer are used to store the semantic information of the track subgrade, the bridge and tunnel buildings under the track, and the embankment on the track;

[0061] It should be noted that this layer is the most basic and important layer in the scene graph;

[0062] The nodes of the foreground layer are used to store the semantic information of foreground objects;

[0063] It should be noted that the nodes of the foreground layer are the parts that need to be noted in the rail transit scene, which are called "foreground objects", mainly referring to some dynamic objects in the track scene, such as locomotives, people, etc. running on the track;

[0064] The nodes of the background layer are used to store the semantic information of the ground and structures;

[0065] It should be noted that the background layer contains two types of nodes: ground and structure, representing the background structure in the rail transit scenario, mainly referring to some static objects in the track scenario, such as the ground where the track is located, buildings around the track, nature, and the sky, etc.;

[0066] The nodes of the scene layer are used to store the overall attributes of the scene;

[0067] It should be noted that the nodes of the scene layer, such as cities, mountainous areas, bridges, tunnels, etc., reflect their geographical location attributes. When it is necessary to establish visual scene connections between different images, the scene reflected by the entire image can be regarded as a node.

[0068] In at least one embodiment of the present invention, a generation process of a complete rail transit scene graph is proposed. As Figure 3 shown, first, the selected picture is subjected to image segmentation to obtain a fine-grained instance segmentation image, then the two image regions are located, and then the relationships in different regions are labeled. Finally, the encoded data is input into the generation model to obtain the rail transit scene graph, that is, the nodes are obtained through semantic segmentation recognition, the target regions are obtained through manual screening and attention mechanisms, and the relationships are obtained through manual labeling, as Figure 4 shown.

[0069] It should be noted that the edge relationships in the finally obtained rail transit scene graph are divided into: "instances within the foreground layer", "instances within the track layer", "instances between the foreground layer and the track layer", "instances between the foreground layer and the background layer", "instances between the track layer and the background layer", and "instances within the background layer". All regions are described by each layer.

[0070] Next, the rail transit scene graph generation device provided by the present invention will be described. The rail transit scene graph generation device described below can be mutually referred to the rail transit scene graph generation method described above. As Figure 5 shown, the rail transit scene graph generation device includes:

[0071] An acquisition module 501, configured to acquire a rail transit image and label the target region in the rail transit image according to the attention mechanism;

[0072] A generation module 502, configured to input the rail transit image after labeling the target region into the generation model to obtain the rail transit scene graph output by the generation model;

[0073] Among them, the generation model is used to construct the rail transit scene graph according to the nodes and the visual relationships between the nodes in the labeled target area. The rail transit scene graph includes multiple layers, and each layer stores multiple nodes. There are edges connecting the nodes within or between the layers. The nodes are used to describe the rail transit scene information of different semantic categories, and the edges are used to describe the visual relationships between the nodes.

[0074] In at least one embodiment of the present invention, the acquisition module includes an annotation sub-module. The annotation sub-module is used to obtain the area with entity objects in the rail transit image and obtain the track-visual area in the rail transit image according to the attention mechanism.

[0075] The annotation sub-module is further used to label the area with entity objects and the track-visual area as the target area.

[0076] In at least one embodiment of the present invention, the track-visual area includes an upper view area of the track, a left view area of the track, and a right view area of the track.

[0077] In at least one embodiment of the present invention, the visual relationships between the nodes include track topological relationships, spatial relationships, occlusion relationships, and subordination relationships.

[0078] In at least one embodiment of the present invention, the rail transit scene graph includes a source image layer, a track layer, a foreground layer, a background layer, and a scene layer.

[0079] In at least one embodiment of the present invention, the nodes of the source image layer are used to store the rail transit image.

[0080] The nodes of the track layer are used to store the semantic information of the track subgrade, the bridge and tunnel buildings under the track, and the embankment on the track.

[0081] The nodes of the foreground layer are used to store the semantic information of the foreground objects.

[0082] The nodes of the background layer are used to store the semantic information of the ground and the structure.

[0083] The nodes of the scene layer are used to store the overall attributes of the scene.

[0084] Figure 6 An example of the schematic physical structure of an electronic device is shown in Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute the method for generating a rail transit scenario graph. The method includes:

[0085] Obtain a rail transit image, and label the target area in the rail transit image according to the attention mechanism;

[0086] Input the rail transit image with the target area labeled into the generation model to obtain the rail transit scenario graph output by the generation model;

[0087] Among them, the generation model is used to construct the rail transit scenario graph according to the nodes and the visual relationships between the nodes in the labeled target area. The rail transit scenario graph includes a multi-layer directed graph, and each layer stores multiple nodes. There are edges connecting the nodes within or between the layers. The nodes are used to describe the rail transit scenario information of different semantic categories, and the edges are used to describe the visual relationships between the nodes.

[0088] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0089] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for generating a rail transit scenario graph provided by the above-mentioned various methods. The method includes:

[0090] Obtain a rail transit image, and label the target area in the rail transit image according to the attention mechanism;

[0091] Input the rail transit image after marking the target area into a generation model to obtain the rail transit scene map output by the generation model;

[0092] Among them, the generation model is used to construct the rail transit scene map according to the nodes and the visual relationships between the nodes in the marked target area. The rail transit scene map includes a multi-layer directed graph, and each layer stores multiple nodes. There are edges connecting the nodes within or between the layers. The nodes are used to describe the rail transit scene information of different semantic categories, and the edges are used to describe the visual relationships between the nodes.

[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a rail transit scenario map, characterized in that Including: Obtain a rail transit image and label the target area in the rail transit image according to the attention mechanism; Input the rail transit image after labeling the target area into a generation model to obtain the rail transit scene graph output by the generation model; Wherein, the generation model is used to construct the rail transit scene graph according to the nodes and the visual relationships between the nodes in the labeled target area. The rail transit scene graph includes a multi-layer directed graph, and each layer stores multiple nodes. There are edges connecting the nodes within or between the layers. The nodes are used to describe the rail transit scene information of different semantic categories, and the edges are used to describe the visual relationships between the nodes; Wherein, the labeling of the target area in the rail transit image according to the attention mechanism includes: Obtain the area with entity objects in the rail transit image, and obtain the rail-visual area in the rail transit image according to the attention mechanism; Label the area with entity objects and the rail-visual area as the target area; The rail-visual area includes the on-rail view area, the left-rail view area, and the right-rail view area; The obtaining the area with entity objects in the rail transit image and obtaining the rail-visual area in the rail transit image according to the attention mechanism includes: Starting from the most prominent part of the rail transit image, describe the rail transit image one by one to obtain area descriptions, and draw bounding boxes covering all the entity objects mentioned in the area descriptions; On the basis of rail area positioning, locate different rails according to the number of different rails, and for the human visual attention mechanism of the rail transit scene, add the on-rail view area, the left-rail view area, and the right-rail view area for rail-visual area positioning; Describe the relationships in the located rail-visual area and the encoded information of the relationships between objects, attributes, and data, and use it in the subsequent generation of the rail transit scene graph; The rail transit scene graph includes a source image layer, a rail layer, a foreground layer, a background layer, and a scene layer; The relationship of the edges is established through the following steps: According to the occurrence frequency of each category in various rail transit datasets, select the category with the highest occurrence frequency as the semantic object category, group it, and store it in different layers to describe the rail transit scene; Connect the nodes in the bottom layer with the nodes in the foreground layer and the background layer; Connect the nodes in the rail layer with the foreground layer and the background layer respectively to represent the spatial relationship; Connect the structural nodes in the background layer to the ground nodes to represent the adjacency relationship, and connect the structural nodes in the background layer with the nodes in the foreground layer to represent the contact relationship; Construct a relationship list suitable for the rail transit scene graph and a relationship sub-list applicable to each layer: Remove the relationships irrelevant to the rail transit scene from the 50 most common relationships in the Visual Genome, and supplement and define the common relationships in the rail transit scene based on the rail semantic graph; The relationships between different rails are defined in five types: parallel relationship, intersection relationship, parallel and intersection relationship from visual proximal to visual distal, intersection and parallel relationship from visual proximal to visual distal, and up-down relationship in space; Form the rail transit scene graph by connecting with each other within or across layers through relationships.

2. The method for generating a rail transit scene diagram according to claim 1, wherein The visual relationships between the nodes include rail topological relationships, spatial relationships, occlusion relationships, and subordination relationships.

3. The method for generating a rail transit scene map according to claim 1 or 2, characterized in that The nodes in the source image layer are used to store rail transit images; The nodes in the rail layer are used to store the track subgrade, bridge and tunnel buildings under the track, and semantic information on the track embankment; The nodes in the foreground layer are used to store the semantic information of foreground objects; The nodes in the background layer are used to store the semantic information of the ground and structures; The nodes in the scene layer are used to store the overall attributes of the scene.

4. A device for generating a rail transit scenario diagram, characterized in that, Include: An acquisition module, configured to acquire rail transit images and label target regions in the rail transit images according to the attention mechanism; A generation module, configured to input the rail transit images with labeled target regions into a generation model to obtain the rail transit scene graph output by the generation model; Wherein, the generation model is used to construct the rail transit scene graph according to the nodes and visual relationships between the nodes in the labeled target regions. The rail transit scene graph includes a multi-layer directed graph, and each layer stores multiple nodes. There are edges connecting the nodes within or between the layers. The nodes are used to describe rail transit scene information of different semantic categories, and the edges are used to describe the visual relationships between the nodes; Wherein, the acquisition module includes a labeling sub-module, and the labeling sub-module is configured to obtain the regions with entity objects in the rail transit images and obtain the rail-visual regions in the rail transit images according to the attention mechanism; The labeling sub-module is further configured to label the regions with entity objects and the rail-visual regions as target regions; The rail-visual regions include the on-rail viewing area, the left-rail viewing area, and the right-rail viewing area; The labeling sub-module, specifically: Start from the most prominent part of the rail transit image, describe the rail transit image one by one to obtain region descriptions, and draw bounding boxes covering all the entity objects mentioned in the region descriptions; on the basis of rail region positioning, locate different tracks according to the number of different tracks, and add the on-rail viewing area, the left-rail viewing area, and the right-rail viewing area for rail-visual region positioning according to the human visual attention mechanism of the rail transit scene; Describe the relationships in the located rail-visual regions and the encoding information of the relationships between objects, attributes, and data, and use it in the subsequent generation of the rail transit scene graph; The rail transit scene graph includes a source image layer, a rail layer, a foreground layer, a background layer, and a scene layer; The generation module is specifically configured to establish the relationship of the edges in the following way: According to the occurrence frequency of each category in various rail datasets, select the category with the highest occurrence frequency as the semantic object category, group it, and store it in different layers to describe the rail transit scene; Connect the nodes in the bottom layer with the nodes in the foreground layer and the background layer; connect the nodes in the rail layer with the foreground layer and the background layer respectively to represent spatial relationships; connect the structure nodes in the background layer to the ground nodes to represent adjacency relationships, and connect the structure nodes in the background layer with the nodes in the foreground layer to represent contact relationships; Construct a relationship list suitable for the rail transit scenario graph and relationship sub-lists applicable to each layer: Remove the relationships irrelevant to the rail transit scenario from the 50 most common relationships in Visual Genome, and supplement and define the common relationships in the rail transit scenario based on the rail semantic graph; The relationships between different rails are defined in five types: parallel relationship, intersecting relationship, parallel and intersecting relationship from visual proximal to visual distal, intersecting and parallel relationship from visual proximal to visual distal, and vertical relationship in space. Form the rail transit scenario graph by connecting each other within or across layers through the relationships.

5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the rail transit scenario graph generation method according to any one of claims 1 to 3.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the rail transit scenario graph generation method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Relation visual attention mechanism-based scene graph generation method

    CN110991532A

  • Scene Graph Generation apparatus

    KR102254768B1