Semantic change detection method and system supporting block scene
By constructing a single-stage scene graph generation model based on Transformer and combining visual and spatial features, semantic change detection in street scenes is performed, solving the problems of speed and accuracy in semantic change detection in street scenes, and realizing real-time, accurate, and comprehensive monitoring and analysis of the street environment.
Patent Information
- Application Number
- CN202510912976.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-21
AI Technical Summary
Existing technologies struggle to quickly and accurately detect semantic changes in street-level scenarios, and lack in-depth analysis of changes in spatial layout and semantic relationships between elements, thus failing to effectively support urban managers in making decisions and optimizing resource allocation at a fine scale.
A single-stage scene graph generation model based on Transformer is constructed. By extracting visual, semantic, and spatial features, and combining image annotation tools and time-series datasets, the scene graph generation model is trained to perform object detection and relationship modeling, construct an image spatial knowledge graph, and perform semantic change detection.
It enables real-time, accurate, and comprehensive monitoring and analysis of street scenes, providing intuitive visualization results of semantic changes and improving the monitoring and analysis capabilities of the street environment.
Smart Images

Figure CN120823503A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence image understanding, and specifically relates to a semantic change detection method and system supporting block scenes. Background Art
[0002] Urban blocks are crucial spaces for human production, life, and social interaction. Changes in their semantic information directly reflect the progress of urban development and the effectiveness of governance. Block scene change detection technology has important applications in areas such as digital urban management, infrastructure maintenance, and disaster assessment. However, how to quickly, accurately, and comprehensively detect semantic changes at the fine-scale of blocks and express these changes in a more intuitive manner remains a pressing technical challenge in urban governance and planning.
[0003] Currently, scene analysis methods based on street view imagery mostly employ feature extraction, object detection, semantic segmentation, and image classification techniques. While these methods can, to a certain extent, identify and represent the basic elements and their classification labels within street scenes, they remain significantly deficient in revealing the complex semantic relationships between elements within the scene and the underlying semantic variations. Furthermore, the results produced by these methods typically represent the identification or category labeling of a single element, lacking in-depth analysis of changes in the spatial layout and semantic relationships between elements. This makes them ineffective in supporting urban managers' fine-scale decision-making and resource optimization.
[0004] With the rise of scene graph technology, scene graphs can intuitively represent objects and their relationships in an image as a graph structure, thereby abstractly representing the high-level semantic information of the image. However, scene graph technology is currently mainly used for semantic understanding of general-purpose images, and there has been no systematic research on the spatial semantic variation characteristics of elements in block scenes, especially complex block environments. Therefore, how to build scene semantic change detection methods and systems suitable for block-level scenes to improve the real-time, accurate, and comprehensive monitoring and analysis capabilities of semantic changes in block environments is one of the key technical issues that need to be addressed. Summary of the Invention
[0005] The present invention aims to solve the technical problems existing in the background technology and provides a semantic change detection method and system for street scene, including: analyzing the geographical elements and correlation relationships of street scene, constructing a temporal street scene change detection dataset; using the Visual Genome dataset as training data, training a single-stage scene graph generation model based on Transformer, and extracting the visual features, semantic features and spatial features of the target object; The Genome dataset is combined with a time-series dataset of annotated street view images. Using the best pre-trained scene graph generation model as the training basis, the scene graph generation model is retrained to adapt to street scenes. Taking street view image pairs from different time periods as input, the trained scene graph generation model is used to perform object detection and relationship modeling inference on street scene entities. Considering the spatial location of target entities in street view images, an inference module for the scene graph generation model is designed. During the inference process of subject-object triples, the pixel coordinates of the entity target detection anchor boxes are calculated, and a JSON file containing the inference results is generated for each street view image. Based on the subject-verb-object triples extracted from the scene graph generation model and the spatial anchor box positions of the entities, an image spatial knowledge graph corresponding to each street view image is constructed. Semantic change detection for street view images of different time periods is achieved based on the image spatial knowledge graph generated for a single street view image. Through a series of steps such as node alignment and fusion, graph differencing, similarity calculation, and graph visualization, the semantic changes in street views are presented from a graph perspective, thus achieving street view semantic change detection.
[0006] In order to solve the technical problem, the technical solution of the present invention is:
[0007] A method for detecting semantic changes in a supporting block scene, the method comprising:
[0008] S1: Based on the preset dataset and the data characteristics of street view image change detection, we collect pairs of urban block street view images at the same location in different time periods and construct a street view image time series dataset using image annotation tools.
[0009] S2: Using the preset dataset, training a Transformer-based single-stage scene graph generation model, and extracting visual features, semantic features, and spatial features of the target object to obtain a trained optimal pre-trained scene graph generation model;
[0010] S3: merging the preset dataset and the street view image time series dataset as input, using the optimal pre-trained scene graph generation model as a training basis, retraining the scene graph generation model to adapt the model to the street scene, and obtaining a trained street-adapted scene graph generation model;
[0011] S4: Using the trained block-adapted scene graph generation model and pairs of street view images at different times as input, the trained block-adapted scene graph generation model is used to perform object detection and relationship modeling reasoning on the street view images. During the reasoning process, the spatial position of the target entity is considered, and the pixel coordinates of the target detection anchor box are calculated. The reasoning result includes subject-verb-object triple relationship data, and a JSON file is output for each street view image, which contains object detection information, triple relationship reasoning results, and the coordinates of the entity target.
[0012] S5: Based on the subject-verb-object triples extracted from the JSON file and the spatial anchor box position of each target entity, an image spatial knowledge graph corresponding to each street view image is constructed;
[0013] S6: Based on the image spatial knowledge graphs of different time phases, semantic change detection is performed. Through node alignment and graph structure difference calculation, newly added or deleted nodes and edges are identified. The graph differences between the two time phases are compared, the similarity scores are calculated, and a visualization of the changes is generated. Finally, the node alignment fusion results, graph structure difference results, similarity measurement scores, and graph change visualizations are obtained.
[0014] It can be understood that this method realizes semantic change detection of street view images from a graph perspective. It can not only effectively handle the differences in multi-temporal street view images, but also present the semantic change content in the form of graphs, provide intuitive visualization results, and enhance the real-time, accurate, and comprehensive monitoring and analysis capabilities of semantic changes in block environments.
[0015] Furthermore, the step S1 includes:
[0016] S101, Street View Image Collection and Collation:
[0017] We selected urban blocks as the research objects and used panoramic cameras to continuously collect street view images of the same location at different times. We then organized the image data and constructed a time-series street view image dataset.
[0018] S102. Use of image annotation tools:
[0019] Use professional image annotation tools, using the COCO dataset annotation format as a reference, to annotate the objects in the street view image in the form of target boxes, marking the name, specific location, and shape of the object;
[0020] S103, target relation triple annotation:
[0021] According to the positional relationship of the target annotation boxes, the target relationship in the street view image is described in the form of <subject, relationship, object> triples. The target annotation boxes and relationships are numbered respectively, and the relationship between the targets is annotated in the form of triples according to the sequence numbers.
[0022] S104. Target selection and marking:
[0023] In terms of targets, new street-specific target objects have been added to the original VG dataset target categories, including: road traffic lights, signboards, green spaces, garbage, drains, street lights, and roads;
[0024] S105, scene graph construction:
[0025] Based on the annotated target boxes and relationship triplet sets, instantiate the composition to form a street view image scene graph;
[0026] In the scene graph, graph nodes represent the ground objects in the image, edges represent the relationships between objects, and there is a one-to-one correspondence between nodes and target box labels. The relationships in the scene graph cover spatial relationships of topology, direction, distance, and semantic relationships.
[0027] Further, the step S2 includes:
[0028] S201, input representation:
[0029] Extract image features through convolutional neural networks and inject position information using position encoding to preserve spatial structure;
[0030] The feature representation of a given image is Where N is the number of regions in the input image, D is the feature dimension of each region, and the position encoding is added to the input embedding:
[0031] X input =X+P
[0032] Among them, X input It is the superposition of input feature map and position encoding;
[0033] S202, select encoder:
[0034] The encoder uses a standard Transformer architecture to extract relationship information between input features, which mainly consists of a multi-layer attention mechanism and a feedforward neural network.
[0035] Self-Attention Mechanism:
[0036] Calculate the similarity between different positions in the input sequence and update the representation of each position according to the weighted similarity; given input X input , computing the representation of each region through multi-head self-attention;
[0037] For single-head self-attention, the calculation formula is as follows:
[0038]
[0039] Among them, Q, K, V are query, key and value matrices respectively, d k Is the dimension of the key. For multi-head attention, it is divided into multiple heads, calculated in parallel, and finally the results of all heads are spliced together:
[0040] MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O
[0041] The calculation of each head is as follows:
[0042] headh i =Attention(QW i Q ,KW i K ,VW i V )
[0043] Among them, W i Q ,W i K ,W i V is the weight matrix of each head, W O is the output weight matrix;
[0044] Feedforward Neural Network:
[0045] The self-attention output of each layer passes through a feed-forward neural network consisting of two fully connected layers and activation functions:
[0046] FFN(x)=max(0,xW1+b1)W2+b2
[0047] The feedforward network enhances the representation ability of the network through nonlinear transformation, and the encoder output is It contains the global relationship information of the input features;
[0048] S203, Relationship Modeling:
[0049] By calculating the relationship matrix between targets Represents the relationship between each pair of targets; the matrix R can be constructed by calculating the similarity or interaction between targets, including similarity metrics based on spatial position, visual features or other contextual information;
[0050] By combining the relationship matrix with the encoder output H encoded Combined, the model enhances the interactive features between targets, thereby optimizing the relational reasoning between targets;
[0051] S204, decoder:
[0052] The feature map output by the encoder is converted into the final prediction information of the target. The decoder combines the context information of the target detection through the self-attention mechanism and the cross-attention mechanism to generate the final output;
[0053] Cross-Attention Mechanism:
[0054] The cross-attention mechanism is used to associate the encoder output with the decoder target to generate the final target category and position prediction; the formula is as follows:
[0055]
[0056] Among them, Q dec is the query vector of the decoder, K enc and V enc are the key and value of the encoder respectively;
[0057] Object Detection:
[0058] For the target detection task, the output of the decoder will pass through a classification layer and a regression layer to predict the location and category of the target. The output of the classification layer is the category probability distribution of each target, while the output of the regression layer is the coordinate prediction of the target bounding box.
[0059] S205, output layer:
[0060] For each target area, the output layer will give the target category and location prediction. The final target detection result is output by the following formula:
[0061] y class =softmax(W class H decoded +b class )
[0062] y bbox =W bbox H decoded +b bbox
[0063] Among them, W class and W bbox are the weight matrices for categories and bounding boxes, H decoded is the output of the decoder.
[0064] Further, the step S3 includes:
[0065] S301, Image Feature Embedding:
[0066] The feature embedding module embeds visual features and combines them with position coding information. It uses two position coding methods: sinusoidal position coding and learned position coding. The sinusoidal position coding formula is as follows:
[0067]
[0068] PE (pos,2i+1) =cos(pos / 10000 2i / d )
[0069] Among them, pos is the position, i is the dimension index, and d is the feature dimension;
[0070] S302, encoder and decoder parts:
[0071] The image features are processed by the Transformer encoder and decoder to generate the final scene graph;
[0072] During the training process, the cross entropy loss function is used to optimize the model so that the generated target objects and relationships are as close to the real annotations as possible; the specific loss function is as follows:
[0073]
[0074] in, is the cross entropy loss of the target class, is the bounding box regression loss, is the generalized intersection-union loss, is the loss of relationship prediction;
[0075] S303, relational reasoning module:
[0076] In order to enhance the relational reasoning ability of the model, the relationship between objects is further explored through graph neural networks or self-attention mechanisms.
[0077] Further, the step S4 includes:
[0078] S401, input data processing:
[0079] The input is a pair of street view images at different times. Given an image, we first preprocess the image to obtain a standardized Tensor format data, which is convenient for use as the input of the model. The calculation formula for this processing is:
[0080]
[0081] Where μ and σ are the mean and standard deviation of the image data respectively;
[0082] S402, relational reasoning module:
[0083] The model extracts features from the input image and combines them with position encoding to obtain the category prediction, detection box coordinates, and relationships between entities for each target entity, and generates triplets between target entities.
[0084] The coordinates of the detection box of each target entity are calculated using the following formula:
[0085]
[0086] Among them, (x center ,y center ) is the center coordinate of the target box, w and h are the width and height of the target box.
[0087] Further, the step S5 includes:
[0088] S501. Construction of graph structure:
[0089] Combined with the JSON file output from step S4 containing the subject-verb-object triples and the coordinates of the entity target, use the trained street adaptation scene graph generation model in step S3 and use the Networkx library in Python to build the graph structure;
[0090] S502. Calculate node spatial position attributes:
[0091] In order to enhance the distinguishability of nodes, especially when there are multiple entities with the same name in street view images, a spatial location attribute is introduced in the graph node to store the spatial location coordinates of the entity in the image;
[0092] The spatial position of each entity is represented by the coordinates of the center point of its anchor box. The two-dimensional coordinates of the center point of the anchor box can be calculated by the following formula:
[0093]
[0094] Among them, x min ,y min is the coordinate of the upper left corner of the anchor box, x max ,y max is the coordinate of the lower right corner of the anchor box, (cx,cy) is the coordinate of the center point of the anchor box;
[0095] S503, Intra-graph Node Fusion Strategy:
[0096] Design a node merging strategy within the graph to merge nodes with the same name and close location based on the Euclidean distance threshold;
[0097] Assume that node N i and N j The position coordinates are (x i ,y i ) and (x j,y j ), then the Euclidean distance d between them is calculated by the following formula:
[0098]
[0099] When d is less than the preset threshold, the two nodes are considered to be the same entity and merged;
[0100] S504, coordinate scaling and normalization:
[0101] In order to align the positions of knowledge graph nodes with those of the original image, the coordinates of the nodes need to be properly scaled and normalized;
[0102] The Y coordinate of the image is inverted by the image height to fit the visualization coordinate system;
[0103] The node coordinates will be normalized by dividing the original image coordinates by the width and height of the image, thereby converting each coordinate point from the pixel coordinate system to the standardized unit coordinate system; the normalized coordinates (x ′ ,y ′ ) is calculated using the following formula:
[0104]
[0105] Where width is the width of the image, height is the height of the image, and x and y are the original coordinates of the node.
[0106] Further, the step S6 includes:
[0107] S601, node alignment and fusion:
[0108] Align and fuse the spatial position differences of the same entity at different time points, and design a two-stage node merging strategy. The specific steps are as follows:
[0109] Merge nodes within the graph:
[0110] In the node merging phase within the graph, the algorithm merges nodes with similar locations and the same name by calculating the Euclidean distance between nodes; the formula is as follows:
[0111] For node N i and N j , when:
[0112] name(x i )=name(x j )
[0113]
[0114] The two nodes are merged into the same node, where θpos is the preset threshold;
[0115] Inter-graph node alignment:
[0116] For the graph structure nodes at different time points T1 and T2, θ align is the preset threshold, and the node pair (N t1 ,N t2 )’s similarity:
[0117]
[0118] When S align >θ align When , take the average position of two similar nodes:
[0119]
[0120] S602, Image Difference:
[0121] Detect changes between two graph structures, including the addition or deletion of nodes and edges, divided into node set differencing and edge-level differencing;
[0122] Node differential:
[0123] By calculating G t2 G t1 Nodes that are not in G, identify new nodes, and calculate G t1 G t2 Identify and delete nodes that are not in the list;
[0124] new_nodes=G t2 .nides()-G t1 .nides()
[0125] deleted_nodes=G t1 .nides()-G t2 .nodes()
[0126] Edge Difference:
[0127] By calculating G t2 G t1 Identify the edges that are not in G and add them; by calculating G t1 G t2 Identify and delete edges that are not present in ;
[0128] new_edges=G t2 .edges()-G t1 .edges()
[0129] deleted_edges=G t1 .edges()-G t2 .edges()
[0130] S603, similarity calculation:
[0131] By combining the two graph embedding methods, node2vec and graph2vec, we generate node embedding vectors and graph embedding vectors from the node-level and graph-level similarity calculation methods respectively, and then calculate the similarity of the two graphs;
[0132] Node2Vec node-level similarity calculation:
[0133] Similarity score S between nodes node It can be measured by calculating the cosine similarity between node embedding vectors; the formula is:
[0134]
[0135] Among them, e u and e v are the embedding vectors of node u and node v respectively;
[0136] Graph2Vec graph-level similarity calculation:
[0137] Similarity score S between graphs graph It is measured by calculating the cosine similarity between graph embedding vectors; the formula is:
[0138]
[0139] in, and are the embedding vectors of graph G1 and graph G2 respectively;
[0140] Combining the results of node2vec and graph2vec, we calculate node-level similarity and graph-level similarity respectively. The final similarity metric can be obtained by a weighted average fusion method. The comprehensive similarity score formula is as follows:
[0141] S total =α·S node +β·S graph
[0142] Among them, α and β are weight parameters that control the contribution of node-level similarity and graph-level similarity to the final similarity score;
[0143] S604, Graph Visualization
[0144] Enter G t1 and G t2The graph structure and its node positions show the changes in the graph.
[0145] A semantic change detection system supporting street scenes, the system being applied to any of the above methods, the system comprising:
[0146] Data collection and annotation module: Based on the preset dataset and the data characteristics of street view image change detection, it collects pairs of urban block street view images from the same location at different time periods and constructs a street view image time series dataset using image annotation tools;
[0147] Pre-trained model generation module: using the preset data set, training a single-stage scene graph generation model based on Transformer, and extracting the visual features, semantic features and spatial features of the target object to obtain the trained optimal pre-trained scene graph generation model;
[0148] The block adaptation training module combines the preset dataset and the street view image time series dataset as input, uses the optimal pre-trained scene graph generation model as a training basis, retrains the scene graph generation model to adapt the model to the block scene, and obtains a trained block adaptation scene graph generation model;
[0149] Target Detection and Relationship Reasoning Module: This module uses the trained block-adapted scene graph generation model and pairs of street view images from different time periods as input, and performs target detection and relationship modeling reasoning on the street view images using the trained block-adapted scene graph generation model. During the reasoning process, the spatial position of the target entity is considered, and the pixel coordinates of the target detection anchor box are calculated. The reasoning results include subject-verb-object triple relationship data, and a JSON file is output for each street view image, containing target detection information, triple relationship reasoning results, and the coordinates of the entity target.
[0150] Knowledge graph construction module: Based on the subject-verb-object triples extracted from the JSON file and the spatial anchor box position of each target entity, it constructs the image spatial knowledge graph corresponding to each street view image;
[0151] Semantic change detection and visualization module: Based on the image space knowledge graphs in different time phases, semantic change detection is performed. Through node alignment and graph structure difference calculation, new or deleted nodes and edges are identified, the graph differences between the two time phases are compared, the similarity scores are calculated and a visualization of the changes is generated. Finally, the node alignment fusion results, graph structure difference results, similarity measurement scores and graph change visualization graphs are obtained.
[0152] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, any of the above-mentioned methods for detecting semantic changes in supporting block scenes is implemented.
[0153] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements any of the above-mentioned semantic change detection methods for supporting block scenes.
[0154] Compared with the prior art, the advantages of the present invention are:
[0155] Construction of a temporal street view dataset: By collecting street view images of the same location at different time points, a temporal street view dataset is constructed, which helps to better detect changes and updates of various elements in the block environment.
[0156] Block scene graph generation model: By combining the Visual Genome dataset with a time-series dataset of annotated street view images, and using the best pre-trained scene graph generation model as the training basis, the scene graph generation model is retrained to better adapt to block scenes, laying a solid foundation for subsequent semantic change detection.
[0157] Relationship modeling reasoning: By designing a scene graph generation model inference module, the spatial position of the target entity in the street view image is taken into account. During the inference process, the pixel coordinates of the entity target detection anchor box are calculated and the corresponding JSON file is generated. This effectively improves the recognition of target entities and relationship modeling in street view imagery at different times.
[0158] Image Spatial Knowledge Graph Construction: By combining the subject-verb-object triples extracted by the scene graph generation model with the spatial anchor box positions of the target entities, we construct an image spatial knowledge graph corresponding to each street view image. This ensures that the semantic information of each street view image includes not only visual features but also spatial structure information, providing more comprehensive and accurate data support for subsequent change detection.
[0159] Street View Image Semantic Change Detection Module: This module proposes a street view semantic change detection module, which includes node alignment and fusion, graph differencing, similarity calculation, and graph visualization. Compared to traditional pixel-level change detection methods, this module not only effectively handles differences in multi-temporal street view imagery but also presents semantic changes through graphs, providing intuitive visualization results. BRIEF DESCRIPTION OF THE DRAWINGS
[0160] Figure 1 , the main technical roadmap of the semantic change detection method supporting block scenes in the present invention. DETAILED DESCRIPTION
[0161] The specific implementation of the present invention is described below in conjunction with examples:
[0162] It should be noted that the structures, proportions, sizes, etc. shown in this specification are only used to match the contents disclosed in the specification for people familiar with this technology to understand and read, and are not used to limit the conditions under which the present invention can be implemented. Any structural modification, change in proportional relationship or adjustment of size should still fall within the scope of the technical content disclosed in the present invention without affecting the efficacy and purpose that can be achieved by the present invention.
[0163] At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" quoted in this specification are only for the convenience of description and are not used to limit the scope of implementation of the present invention. Changes or adjustments to their relative relationships should be regarded as the scope of implementation of the present invention without substantially changing the technical content.
[0164] Example 1:
[0165] like Figure 1 As shown, the present invention provides a semantic change detection method for supporting block scenes, comprising the following steps:
[0166] S1. Analyze the geographical elements and related relationships of street scenes and construct a time series street scene change detection dataset;
[0167] S2. Using the Visual Genome dataset as training data, train a Transformer-based single-stage scene graph generation model and extract the visual, semantic, and spatial features of the target object;
[0168] S3. Combine the VisualGenome dataset with the annotated street view image time series dataset, use the best pre-trained scene graph generation model as the training basis, and retrain the scene graph generation model to adapt to the street scene.
[0169] S4. Using street view image pairs from different time periods as input, the trained scene graph generation model is used to perform target detection and relationship modeling reasoning on street scene entities. Considering the spatial position of the target entity in the street view image, a scene graph generation model inference module is designed. In the process of inferring the subject-object triples, the pixel coordinates of the entity target detection anchor box are calculated, and a JSON file containing the inference results is obtained for each street view image.
[0170] S5. Based on the subject-verb-object triples extracted from the scene graph generation model and combined with the spatial anchor frame positions of the entities, an image space knowledge graph corresponding to each street view image is constructed.
[0171] S6. Semantic change detection of street view images at different times is achieved based on the image space knowledge graph generated corresponding to a single street view image. Through a series of steps such as node alignment and fusion, graph differentiation, similarity calculation, and graph visualization, the semantic change content is presented from a graph perspective.
[0172] In this embodiment, S1 constructs a semantic change detection dataset for street scenes. Considering the data characteristics of street view image change detection, this embodiment references existing Visual Genome and Open Images datasets, incorporates the temporal and spatial characteristics of urban street scenes, and collects pairs of street view images of urban blocks at the same location at different time periods. These images are annotated using professional image annotation tools. During dataset construction, representative urban blocks are first selected, including commercial areas, residential areas, transportation hubs, and other areas. The street view images of each area are ensured to reflect distinct change characteristics over different time periods. When collecting street view images, fixed time intervals are set to capture images of the same location at different time periods to ensure the temporal and spatial continuity of the data. After image acquisition, professional image annotation tools are used to annotate the captured street view images with target boxes, accurately marking the specific location and shape of each target object. Target objects include buildings, shops, billboards, road signs, vehicles, etc., and each target object annotation includes information such as the target box's location and size. This process generates a corresponding target box for each street view image. Based on the annotated target boxes, the relationships between each target object are further annotated. Through manual or intelligently assisted annotation, <subject, relationship, object> triples are constructed between target objects. Based on this annotation information, a scene graph for instantiated street view images can be generated. In the scene graph, nodes represent target objects in the image, while edges represent the relationships between these objects. There is a one-to-one correspondence between nodes and target box annotations. This graph structure effectively represents the target objects and their relationships in street view imagery, providing an accurate data foundation for subsequent semantic change detection.
[0173] In this embodiment, S3 trains a block scene graph generation model;
[0174] To further adapt to the complex features of block street view images and maximize the use of existing data during training, we combined the Visual Genome dataset and the street view image time series dataset constructed by annotation, and based on the pre-trained scene graph generation model trained with the VG dataset, we fine-tuned the model and retrained it to adapt to block scenes.
[0175] The scene graph generation model is based on the Transformer architecture. Its basic structure includes the following key components: an image encoding module, an image feature embedding module, a Transformer encoder and decoder module, and a relational reasoning module. In the image encoding module, a pre-trained ResNet50 is used as the image encoder, which effectively extracts rich visual features from street view imagery. For each street view image, the ResNet50 output is a set of high-dimensional visual features, which serve as the basis for the Transformer input.
[0176] In the image feature embedding module, the feature embedding module embeds the visual features extracted by ResNet50 and combines them with position encoding information. Two position encoding methods are used: sinusoidal position encoding and learned position encoding. The sinusoidal position encoding formula is as follows:
[0177]
[0178] PE (pos,2i+1) =cos(pos / 10000 2i / d )
[0179] Where pos is the position, i is the dimension index, and d is the feature dimension. In this way, the network can use the spatial position information in the image to guide the learning of objects and their relationships.
[0180] In the encoder and decoder, image features are processed by the Transformer encoder. The encoder applies a multi-level attention mechanism to the extracted features to capture the dependencies between various positions in the image. The decoder receives contextual information from the encoder and generates the final scene graph.
[0181] During the training process, the cross entropy loss function is used to optimize the model so that the generated target objects and relationships are as close to the real annotations as possible. The specific loss function is as follows:
[0182]
[0183] in, is the cross entropy loss of the target class, is the bounding box regression loss, It is the generalized intersection-over-union loss, which is used to measure the overlap between the predicted box and the true box. It is the loss of relationship prediction, which is used to optimize the relationship inference between objects in the scene.
[0184] In the relational reasoning module, in order to enhance the relational reasoning ability of the model, graph neural networks or self-attention mechanisms are used to further explore the relationships between objects.
[0185] When fine-tuning the model using annotated street view imagery, we started with the pre-trained model weights and froze some network layers, particularly the image encoding layer, leaving only the Transformer decoder and relational reasoning modules for training. A low learning rate was used to preserve the basic structure and visual features of the pre-trained model, while the network's reasoning was tuned to be more suitable for street view data. Training parameters were set to decay the learning rate every 100 epochs, with a batch size of 2 and a model iteration count of 200. By fine-tuning the pre-trained scene graph generation model, we were able to effectively adapt the model to the new street view imagery dataset, achieving both object detection and relational reasoning tasks.
[0186] In this embodiment, S4 designs a scene graph generation model inference module:
[0187] In this module, a scene graph generation model inference method based on street view image pairs at different times is designed. The trained street scene graph generation model is used to perform target detection and relationship modeling reasoning on the target entities in the street scene, fully considering the spatial position of the target entity in the street view image. By inferring the subject-object triplet and the pixel coordinates of the target detection anchor frame, the inference results corresponding to each street view image are finally obtained and saved in JSON file format.
[0188] The input data consists of pairs of street view images taken at different times. Each street view image represents a specific viewpoint or time point, containing information about various objects, buildings, and human activities in the neighborhood. Each image is first normalized by resizing it to 800 pixels and normalizing it using mean and standard deviation.
[0189] Given an image, we first preprocess the image to obtain a standardized Tensor format data, which is convenient for use as the input of the model. The calculation formula for this processing is:
[0190]
[0191] Where μ and σ are the mean and standard deviation of the image data, respectively.
[0192] In the relational reasoning module, the model extracts features from the input image, combines them with position encoding, and processes them through the Transformer network to obtain the category prediction, detection box coordinates, and relationship between entities for each target entity, and generates triplets between the target entities.
[0193] The coordinates of the detection box of each target entity are calculated using the following formula:
[0194]
[0195] Among them, (x center ,y center ) is the center coordinate of the target box, w and h are the width and height of the target box. The detection box corresponding to each target entity is given by the output layer of the model and converted into pixel coordinates in the image after decoding.
[0196] In order to convert the normalized coordinates of the detection box into the pixel coordinates of the image, the following formula is used to denormalize the coordinates:
[0197] scaled box=bbox×[img width,img height,img width,img height]
[0198] Where bbox represents the normalized detection box coordinates, img width and img height are the width and height of the image respectively. This formula can be used to convert the normalized detection box to pixel coordinates in the actual image.
[0199] By reasoning about each target entity and its relationship, the corresponding triples are finally generated and saved as JSON files. The main contents include subject, relationship and object categories, confidence and their respective bounding box coordinates.
[0200] In this embodiment, S5 constructs an image space knowledge graph corresponding to each street view image based on the subject-verb-object triple data extracted from the scene graph generation model and the spatial anchor frame position of the entity.
[0201] S501. Construction of graph structure:
[0202] Combined with the JSON file output from S4 containing subject-verb-object triples and target anchor box locations, the trained model in S3 is used to process street view images, and the Networkx library in Python is used to build the graph structure.
[0203] S502. Calculate node spatial position attributes:
[0204] In order to enhance the distinguishability of nodes, especially when there are multiple entities with the same name (such as multiple people or multiple vehicles) in street view images, a spatial location attribute is introduced in the graph node to store the spatial location coordinates of the entity in the image.
[0205] The spatial position of each entity is represented by the coordinates of the center point of its anchor box. The two-dimensional coordinates of the center point of the anchor box can be calculated by the following formula:
[0206]
[0207] Among them, x min ,y minis the coordinate of the upper left corner of the anchor box, x max ,y max is the coordinate of the lower right corner of the anchor box, and (cx,cy) is the coordinate of the center point of the anchor box.
[0208] S503, Intra-graph Node Fusion Strategy:
[0209] Due to the uncertainty of target detection, the same semantic entity may be identified as multiple similar but independent nodes in the graph structure, resulting in semantic redundancy and structural duplication in the graph.
[0210] Design a node merging strategy within the graph to merge nodes with the same name and close location based on the Euclidean distance threshold.
[0211] Assume that node N i and N j The position coordinates are (x i ,y i ) and (x j ,y j ), then the Euclidean distance d between them can be calculated by the following formula:
[0212]
[0213] When d is less than a preset threshold, the two nodes are considered to be the same entity and are merged.
[0214] S504, coordinate scaling and normalization:
[0215] In order to align the positions of knowledge graph nodes with the original image positions, the coordinates of the nodes need to be properly scaled and normalized.
[0216] The image's Y coordinate is inverted by the image's height to fit the visualization coordinate system.
[0217] The node coordinates will be normalized by dividing the original image coordinates by the width and height of the image, thereby converting each coordinate point from the pixel coordinate system to the standardized unit coordinate system. The normalized coordinates (x ′ ,y ′ ) is calculated using the following formula:
[0218]
[0219] Where width is the width of the image, height is the height of the image, and x and y are the original coordinates of the node.
[0220] In this embodiment, S6 realizes semantic change detection of street view images at different phases based on the image space knowledge graph generated corresponding to a single street view image. Through a series of steps such as node alignment and fusion, graph difference, similarity calculation and graph visualization, the semantic change content is presented from the perspective of the graph.
[0221] S601, node alignment and fusion:
[0222] Align and fuse the spatial position differences of the same entity at different time points. A two-stage node merging strategy is designed. The specific steps are as follows:
[0223] Merge nodes within the graph:
[0224] In the node merging phase within the graph, the algorithm calculates the Euclidean distance between nodes and merges nodes with similar locations and the same name. The formula is as follows:
[0225] For node N i and N j , when:
[0226] name(x i )=name(x j )
[0227]
[0228] The two nodes are merged into the same node, where θ pos is the preset threshold.
[0229] Inter-graph node alignment:
[0230] For the graph structure nodes at different time points T1 and T2, θ align is the preset threshold, and the node pair (N t1 ,N t2 )’s similarity:
[0231]
[0232] When S align >θ align When , take the average position of two similar nodes:
[0233]
[0234] S602, Image Difference:
[0235] Detect changes between two graph structures, including the addition or deletion of nodes and edges, mainly divided into node set difference and edge level difference
[0236] Node differential:
[0237] By calculating G t2G t1 Identify the nodes that are not in G. t1 G t2 Identify and delete nodes that are not in the list.
[0238] new_nodes=G t2 .nodes()-G t1 .nodes()
[0239] deleted_nodes=G t1 .nodes()-G t2 .nodes()
[0240] Edge Difference:
[0241] By calculating G t2 G t1 Identify the edges that are not in G. t1 G t2 Identify and delete edges that are not present in .
[0242] new_edges=G t2 .edges()-G t1 .edges()
[0243] deleted_edges=G t1 .edges()-G t2 .nodes()
[0244] S603, similarity calculation:
[0245] By combining the two graph embedding methods, node2vec and graph2vec, we can perform a comprehensive evaluation from the local similarity between nodes to the global similarity of the entire graph, thereby improving the accuracy and robustness of graph structure similarity calculation.
[0246] Node2Vec node-level similarity calculation:
[0247] A graph embedding algorithm based on biased random walk. The core idea is to generate node sequences by performing biased random walks in the graph, thereby capturing the similarity between nodes. The similarity score S between nodes node It can be measured by calculating the cosine similarity between node embedding vectors. The formula is:
[0248]
[0249] Among them, e u and e vare the embedding vectors of node u and node v, respectively. This similarity measure can capture the first-order similarity (homogeneity) and second-order similarity (structural equivalence) between nodes.
[0250] Graph2Vec graph-level similarity calculation:
[0251] Unlike node2vec which focuses on similarity calculation at the node level, Graph2Vec focuses on the structure of the entire graph. Its main idea is to extract the subgraph features of the graph by generating the root node subgraph through the Weisfeiler-Lehman kernel and learn the graph-level embedding using the Skip-Gram model. The similarity score S between graphs is graph It can be measured by calculating the cosine similarity between graph embedding vectors. The formula is:
[0252]
[0253] in, and are the embedding vectors of graph G1 and graph G2 respectively.
[0254] To more comprehensively evaluate the similarity of street scenes at different times, we combined the results of node2vec and graph2vec to calculate node-level similarity and graph-level similarity, respectively. The final similarity metric can be obtained through a weighted average fusion method. The comprehensive similarity score formula is as follows:
[0255] S total =β·S node +β·S graph
[0256] Here, α and β are weight parameters that control the contribution of node-level and graph-level similarity to the final similarity score. By adjusting the weight parameters, we can balance the influence of node-level and graph-level similarity and evaluate the similarity of street scenes at different times from different perspectives.
[0257] S604, Graph Visualization
[0258] Enter G t1 and G t2 The graph structure and its node positions show the changes in the graph (added and deleted nodes and edges).
[0259] Example 2:
[0260] This embodiment provides a semantic change detection system for supporting street scenes. The system can be used to implement the above-mentioned semantic change detection method for street scenes. Specifically, the system includes:
[0261] Dataset construction module: Referring to existing datasets, such as the Visual Genome dataset and the OpenImages dataset, and taking into account the data characteristics of street view image change detection, we collect pairs of urban block street view images at the same location but different time periods, and use image annotation tools to construct a street view image time series dataset.
[0262] Single-stage scene graph generation module: Using the VisualGenome dataset as training data, it extracts the visual, semantic, and spatial features of the target object.
[0263] Block scene graph generation module:
[0264] Taking the labeled street view image time series dataset as input and the best pre-trained scene graph generation model as the training basis, the scene graph generation model is retrained to adapt the model to the block scene.
[0265] Scene graph generation model inference module:
[0266] A scene graph generation model inference module is designed to calculate the pixel coordinates of the entity target detection anchor box during the process of inferring the subject-object triplet, and obtain a JSON file containing the inference results corresponding to each street view image.
[0267] Image space knowledge graph building module:
[0268] Based on the subject-verb-object triples extracted from the scene graph generation model and combined with the spatial anchor frame positions of the entities, an image space knowledge graph corresponding to each street view image is constructed.
[0269] The street view semantic change detection module includes the following submodules:
[0270] Node alignment and fusion module: Design a two-stage node merging strategy to align and fuse the spatial position differences of the same entity at different time points.
[0271] Graph difference module: detects changes between two graph structures, including the addition or deletion of nodes and edges. It is mainly divided into node set difference and edge level difference, and detects changes in entities and relationships.
[0272] Similarity calculation module: It uses node-level and graph-level similarity calculation methods to generate node embedding vectors and graph embedding vectors, and then calculates the similarity of the two graphs.
[0273] Graph visualization module: input G t1 and G t2 The graph structure and its node positions show the changes in the graph (added and deleted nodes and edges).
[0274] Example 3:
[0275] In one embodiment of the present invention, a terminal device is provided, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used to support the operation of the semantic change detection method of the block scene, including the following steps:
[0276] S1. With reference to existing datasets, such as the Visual Genome dataset and the Open Images dataset, and taking into account the data characteristics of street view image change detection, we collected pairs of urban block street view images at the same location but at different time periods. Using image annotation tools, we constructed a street view image time series dataset.
[0277] S2. Using the Visual Genome dataset as training data, train a Transformer-based single-stage scene graph generation model and extract the visual, semantic, and spatial features of the target object;
[0278] S3: The Visual Genome dataset and the street view image time series dataset constructed by annotation are merged as input. The best pre-trained scene graph generation model is used as the training basis, and the scene graph generation model is retrained to adapt the model to the street scene.
[0279] S4. Using street view image pairs from different time periods as input, the trained scene graph generation model is used to perform target detection and relationship modeling reasoning on street scene entities. Considering the spatial position of the target entity in the street view image, a scene graph generation model inference module is designed. In the process of inferring the subject-object triples, the pixel coordinates of the entity target detection anchor box are calculated, and a JSON file containing the inference results is obtained for each street view image.
[0280] S5. Based on the subject-verb-object triples extracted from the scene graph generation model and combined with the spatial anchor frame positions of the entities, an image space knowledge graph corresponding to each street view image is constructed.
[0281] S6. Detect semantic changes in street view images at different times based on the image space knowledge graph generated from a single street view image. The street view semantic change detection module includes: node alignment and fusion, graph differencing, similarity calculation, and graph visualization.
[0282] Example 4:
[0283] In one embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory.
[0284] The processor may load and execute one or more instructions stored in a computer-readable storage medium to implement the corresponding steps of the semantic change detection method for supporting a street scene in the above embodiment. The processor may load and execute the following steps:
[0285] S1. With reference to existing datasets, such as the Visual Genome dataset and the Open Images dataset, and taking into account the data characteristics of street view image change detection, we collected pairs of urban block street view images at the same location but at different time periods. Using image annotation tools, we constructed a street view image time series dataset.
[0286] S2. Using the Visual Genome dataset as training data, train a Transformer-based single-stage scene graph generation model and extract the visual, semantic, and spatial features of the target object;
[0287] S3: The Visual Genome dataset and the street view image time series dataset constructed by annotation are merged as input. The best pre-trained scene graph generation model is used as the training basis, and the scene graph generation model is retrained to adapt the model to the street scene.
[0288] S4. Using street view image pairs from different time periods as input, the trained scene graph generation model is used to perform target detection and relationship modeling reasoning on street scene entities. Considering the spatial position of the target entity in the street view image, a scene graph generation model inference module is designed. In the process of inferring the subject-object triples, the pixel coordinates of the entity target detection anchor box are calculated, and a JSON file containing the inference results is obtained for each street view image.
[0289] S5. Based on the subject-verb-object triples extracted from the scene graph generation model and combined with the spatial anchor frame positions of the entities, an image space knowledge graph corresponding to each street view image is constructed.
[0290] S6. Detect semantic changes in street view images at different times based on the image space knowledge graph generated from a single street view image. The street view semantic change detection module includes: node alignment and fusion, graph differencing, similarity calculation, and graph visualization.
[0291] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0292] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0293] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0294] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0295] The preferred embodiments of the present invention are described in detail above, but the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.
[0296] Many other changes and modifications can be made without departing from the spirit and scope of the present invention. It should be understood that the present invention is not limited to the specific embodiments, and the scope of the present invention is defined by the appended claims.
Claims
1. A semantic change detection method for supporting street scenes, characterized in that: The method comprises: S1: Based on the preset dataset and the data characteristics of street view image change detection, we collect pairs of urban block street view images at the same location in different time periods and construct a street view image time series dataset using image annotation tools. S2: Using the preset dataset, training a Transformer-based single-stage scene graph generation model, and extracting visual features, semantic features, and spatial features of the target object to obtain a trained optimal pre-trained scene graph generation model; S3: merging the preset dataset and the street view image time series dataset as input, using the optimal pre-trained scene graph generation model as a training basis, retraining the scene graph generation model to adapt the model to the street scene, and obtaining a trained street-adapted scene graph generation model; S4: Using the trained block-adapted scene graph generation model and pairs of street view images at different times as input, the trained block-adapted scene graph generation model is used to perform object detection and relationship modeling reasoning on the street view images. During the reasoning process, the spatial position of the target entity is considered, and the pixel coordinates of the target detection anchor box are calculated. The reasoning result includes subject-verb-object triple relationship data, and a JSON file is output for each street view image, which contains object detection information, triple relationship reasoning results, and the coordinates of the entity target. S5: Based on the subject-verb-object triples extracted from the JSON file and the spatial anchor box position of each target entity, an image spatial knowledge graph corresponding to each street view image is constructed; S6: Based on the image spatial knowledge graphs of different time phases, semantic change detection is performed. Through node alignment and graph structure difference calculation, newly added or deleted nodes and edges are identified. The graph differences between the two time phases are compared, the similarity scores are calculated, and a visualization of the changes is generated. Finally, the node alignment fusion results, graph structure difference results, similarity measurement scores, and graph change visualizations are obtained.
2. The semantic change detection method for supporting block scenes according to claim 1 is characterized in that: The step S1 comprises: S101, Street View Image Collection and Collation: We selected urban blocks as the research objects and used panoramic cameras to continuously collect street view images of the same location at different times. We then organized the image data and constructed a time-series street view image dataset. S102. Use of image annotation tools: Use professional image annotation tools, using the COCO dataset annotation format as a reference, to annotate the objects in the street view image in the form of target boxes, marking the name, specific location, and shape of the object; S103, target relationship triple annotation: According to the positional relationship of the target annotation boxes, the target relationship in the street view image is described in the form of <subject, relationship, object> triples. The target annotation boxes and relationships are numbered respectively, and the relationship between the targets is annotated in the form of triples according to the sequence numbers. S104. Target selection and marking: In terms of targets, new street-specific target objects have been added to the original VG dataset target categories, including: road traffic lights, signboards, green spaces, garbage, drains, street lights, and roads; S105, scene graph construction: Based on the annotated target boxes and relationship triplet sets, instantiate the composition to form a street view image scene graph; In the scene graph, graph nodes represent the ground objects in the image, edges represent the relationships between objects, and there is a one-to-one correspondence between nodes and target box labels. The relationships in the scene graph cover spatial relationships of topology, direction, distance, and semantic relationships.
3. The method for detecting semantic changes in a supporting block scene according to claim 1, characterized in that: The step S2 comprises: S201, input representation: Extract image features through convolutional neural networks and inject position information using position encoding to preserve spatial structure; The feature representation of a given image is Where N is the number of regions in the input image, D is the feature dimension of each region, and the position encoding is added to the input embedding: X input =X+P Among them, X input It is the superposition of input feature map and position encoding; S202, select encoder: The encoder uses a standard Transformer architecture to extract relationship information between input features, which mainly consists of a multi-layer attention mechanism and a feedforward neural network. Self-Attention Mechanism: Calculate the similarity between different positions in the input sequence and update the representation of each position according to the weighted similarity; given input X input , computing the representation of each region through multi-head self-attention; For single-head self-attention, the calculation formula is as follows: Among them, Q, K, V are query, key and value matrices respectively, d k Is the dimension of the key. For multi-head attention, it is divided into multiple heads, calculated in parallel, and finally the results of all heads are spliced together: MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O The calculation of each head is as follows: Among them, W i Q , is the weight matrix of each head, W O is the output weight matrix; Feedforward Neural Network: The self-attention output of each layer passes through a feed-forward neural network consisting of two fully connected layers and activation functions: FFN(x)=max(0,xW1+b1)W2+b2 The feedforward network enhances the representation ability of the network through nonlinear transformation, and the encoder output is It contains the global relationship information of the input features; S203, Relational Modeling: By calculating the relationship matrix between targets Represents the relationship between each pair of targets; the matrix R can be constructed by calculating the similarity or interaction between targets, including similarity metrics based on spatial position, visual features or other contextual information; By combining the relationship matrix with the encoder output H encoded Combined, the model enhances the interactive features between targets, thereby optimizing the relational reasoning between targets; S204, decoder: The feature map output by the encoder is converted into the final prediction information of the target. The decoder combines the context information of the target detection through the self-attention mechanism and the cross-attention mechanism to generate the final output; Cross-Attention Mechanism: The cross-attention mechanism is used to associate the encoder output with the decoder target to generate the final target category and position prediction; the formula is as follows: Among them, Q dec is the query vector of the decoder, K enc and V enc are the key and value of the encoder respectively; Object Detection: For the target detection task, the output of the decoder will pass through a classification layer and a regression layer to predict the location and category of the target. The output of the classification layer is the category probability distribution of each target, while the output of the regression layer is the coordinate prediction of the target bounding box. S205, output layer: For each target area, the output layer will give the target category and location prediction. The final target detection result is output by the following formula: y class =softmax(W class H decoded +b class ) y bbox =W bbox H decoded +b bbox Among them, W class and W bbox are the weight matrices for categories and bounding boxes, respectively, H decoded is the output of the decoder.
4. The method for detecting semantic changes in a supporting block scene according to claim 1, characterized in that: The step S3 comprises: S301, Image Feature Embedding: The feature embedding module embeds visual features and combines them with position coding information. It uses two position coding methods: sinusoidal position coding and learned position coding. The sinusoidal position coding formula is as follows: ON (pos,2i+1) =cos(pos / 10000 2i / d ) Among them, pos is the position, i is the dimension index, and d is the feature dimension; S302, encoder and decoder parts: The image features are processed by the Transformer encoder and decoder to generate the final scene graph; During the training process, the cross entropy loss function is used to optimize the model so that the generated target objects and relationships are as close to the real annotations as possible; the specific loss function is as follows: in, is the cross entropy loss of the target class, is the bounding box regression loss, is the generalized intersection-union loss, is the loss of relationship prediction; S303, relational reasoning module: In order to enhance the relational reasoning ability of the model, the relationship between objects is further explored through graph neural networks or self-attention mechanisms.
5. The method for detecting semantic changes in a supporting block scene according to claim 1, characterized in that: The step S4 comprises: S401, input data processing: The input is a pair of street view images at different times. Given an image, we first preprocess the image to obtain a standardized Tensor format data, which is convenient for use as the input of the model. The calculation formula for this processing is: Where μ and σ are the mean and standard deviation of the image data respectively; S402, relational reasoning module: The model extracts features from the input image and combines them with position encoding to obtain the category prediction, detection box coordinates, and relationships between entities for each target entity, and generates triplets between target entities. The coordinates of the detection box of each target entity are calculated using the following formula: Among them, (x center ,y center ) is the center coordinate of the target box, w and h are the width and height of the target box.
6. The method for detecting semantic changes in a supporting block scene according to claim 1, characterized in that: The step S5 comprises: S501. Construction of graph structure: Combined with the JSON file output from step S4 containing the subject-verb-object triples and the coordinates of the entity target, use the trained street adaptation scene graph generation model in step S3 and use the Networkx library in Python to build the graph structure; S502. Calculate node spatial position attributes: To enhance the distinguishability of nodes, especially when there are multiple entities with the same name in street view images, a spatial location attribute is introduced in the graph node to store the spatial location coordinates of the entity in the image; The spatial position of each entity is represented by the coordinates of the center point of its anchor box. The two-dimensional coordinates of the center point of the anchor box can be calculated by the following formula: Among them, x min ,y min is the coordinate of the upper left corner of the anchor box, x max ,y max is the coordinate of the lower right corner of the anchor box, (cx,cy) is the coordinate of the center point of the anchor box; S503, Intra-graph Node Fusion Strategy: Design a node merging strategy within the graph to merge nodes with the same name and close location based on the Euclidean distance threshold; Assume that node N i and N j The position coordinates are (x i ,y i ) and (x j ,y j ), then the Euclidean distance d between them is calculated by the following formula: When d is less than the preset threshold, the two nodes are considered to be the same entity and merged; S504, coordinate scaling and normalization: In order to align the positions of knowledge graph nodes with those of the original image, the coordinates of the nodes need to be properly scaled and normalized; The Y coordinate of the image is inverted by the image height to fit the visualization coordinate system; The node coordinates are normalized by dividing the original image coordinates by the image width and height, thereby converting each coordinate point from the pixel coordinate system to the standardized unit coordinate system; the normalized coordinates (x', y') are calculated using the following formula: Where width is the width of the image, height is the height of the image, and x and y are the original coordinates of the node.
7. The method for detecting semantic changes in a supporting block scene according to claim 1, characterized in that: The step S6 comprises: S601, node alignment and fusion: Align and fuse the spatial position differences of the same entity at different time points, and design a two-stage node merging strategy. The specific steps are as follows: Merge nodes within the graph: In the node merging phase within the graph, the algorithm merges nodes with similar locations and the same name by calculating the Euclidean distance between nodes; the formula is as follows: For node N i and N j , when: name(x i )=name(x j ) The two nodes are merged into the same node, where θ pos is the preset threshold; Inter-graph node alignment: For the graph structure nodes at different time points T1 and T2, θ align is the preset threshold, and the node pair (N t1 ,N t2 )’s similarity: When S align >θ align When , take the average position of two similar nodes: S602, Image Difference: Detect changes between two graph structures, including the addition or deletion of nodes and edges, divided into node set differencing and edge-level differencing; Node differential: By calculating G t2 G t1 Nodes that are not in G, identify new nodes, and calculate G t1 G t2 Identify and delete nodes that are not in the list; new_nodes=G t2 .nodes()-G t1 .nodes() deleted_nodes=G t1 .nodes()-G t2 .nodes() Edge Difference: By calculating G t2 G t1 Identify the edges that are not in G and add them; by calculating G t1 G t2 Identify and delete edges that are not present in ; new_edges=G t2 .edges()-G t1 .edges() deleted_edges=G t1 .edges()-G t2 .edges() S603, similarity calculation: By combining the two graph embedding methods, node2vec and graph2vec, we generate node embedding vectors and graph embedding vectors from the node-level and graph-level similarity calculation methods respectively, and then calculate the similarity of the two graphs; Node2Vec node-level similarity calculation: Similarity score S between nodes node It can be measured by calculating the cosine similarity between node embedding vectors; the formula is: Among them, e u and e v are the embedding vectors of node u and node v respectively; Graph2Vec graph-level similarity calculation: Similarity score S between graphs graph It is measured by calculating the cosine similarity between graph embedding vectors; the formula is: in, and are the embedding vectors of graph G1 and graph G2 respectively; Combining the results of node2vec and graph2vec, we calculate node-level similarity and graph-level similarity respectively. The final similarity metric can be obtained by a weighted average fusion method. The comprehensive similarity score formula is as follows: S total =α·S node +β·S graph Among them, α and β are weight parameters that control the contribution of node-level similarity and graph-level similarity to the final similarity score; S604, Graph Visualization Enter G t1 and G t2 The graph structure and its node positions show the changes in the graph.
8. A semantic change detection system supporting street scenes, characterized by: The system is applied to the method according to any one of claims 1 to 7, and the system includes: Data collection and annotation module: Based on the preset dataset and the data characteristics of street view image change detection, it collects pairs of urban block street view images from the same location at different time periods and constructs a street view image time series dataset using image annotation tools. Pre-trained model generation module: using the preset data set, training a single-stage scene graph generation model based on Transformer, and extracting the visual features, semantic features and spatial features of the target object to obtain the trained optimal pre-trained scene graph generation model; The block adaptation training module combines the preset dataset and the street view image time series dataset as input, uses the optimal pre-trained scene graph generation model as a training basis, retrains the scene graph generation model to adapt the model to the block scene, and obtains a trained block adaptation scene graph generation model; Target Detection and Relationship Reasoning Module: This module uses the trained block-adapted scene graph generation model and pairs of street view images from different time periods as input, and performs target detection and relationship modeling reasoning on the street view images using the trained block-adapted scene graph generation model. During the reasoning process, the spatial position of the target entity is considered, and the pixel coordinates of the target detection anchor box are calculated. The reasoning results include subject-verb-object triple relationship data, and a JSON file is output for each street view image, containing target detection information, triple relationship reasoning results, and the coordinates of the entity target. Knowledge graph construction module: Based on the subject-verb-object triples extracted from the JSON file and the spatial anchor box position of each target entity, it constructs the image spatial knowledge graph corresponding to each street view image; Semantic change detection and visualization module: Based on the image space knowledge graphs in different time phases, semantic change detection is performed. Through node alignment and graph structure difference calculation, new or deleted nodes and edges are identified, the graph differences between the two time phases are compared, the similarity scores are calculated and a visualization of the changes is generated. Finally, the node alignment fusion results, graph structure difference results, similarity measurement scores and graph change visualization graphs are obtained.
9. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for detecting semantic changes in a supporting block scene as claimed in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, which, when executed by a processor, implements the semantic change detection method for supporting block scenes described in any one of claims 1-7.
Citation Information
Cited By
Method for generating long-time behavior of intelligent robot with body based on thinking chain strategy decomposition
CN122165442A