Method for judging image semantic misrecognition based on triple confidence detection

By building a triple confidence detection network, combined with the HG-GRU model and other evaluators, the problem of misidentification of image semantics in autonomous driving systems is solved, the perception accuracy and reliability are improved, the risk of traffic accidents is reduced, and the commercialization and safe operation of autonomous driving systems is supported.

CN120107725BActive Publication Date: 2025-07-01XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510602414.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-01
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

There is a problem of misrecognition of image semantics in autonomous driving systems, which is difficult to effectively solve in the existing technology, resulting in possible major traffic accidents.

Method used

Using the image semantic misidentification and judgment method based on triple confidence detection, a triple confidence detection network is constructed by constructing an HG-GRU model, introducing a GC feature extractor and SEM module, and combining ResourceRank, RelTrans and HGReasoner evaluator, a triple confidence detection network is constructed, each element and its attributes in the road scene is captured, triple information is extracted, and image semantic misidentification and judgment is achieved through confidence detection.

Benefits of technology

It improves the accuracy and reliability of the autonomous driving perception process, reduces the risk of traffic accidents caused by identification errors, provides quantitative judgment basis, enhances the understanding and decision-making ability of complex scenarios, and supports the large-scale commercialization and safe operation of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107725B_ABST
    Figure CN120107725B_ABST
Patent Text Reader

Abstract

The present invention discloses an image semantic misrecognition judgment method based on triple confidence detection, including: adding a historical dynamic gate module on the basis of GRU to construct an HG-GRU model, and introducing GC and SEM modules at its input and output ends respectively; constructing an HGReasoner evaluator based on the above three models, and combining the ResourceRank and RelTrans evaluators to construct a triple confidence detection network; capturing and fusing each element and its attributes in the road scene to construct a road scene graph; extracting triple information of each target in the road according to the scene graph; using the triple confidence detection network to solve the confidence results of each triple; combining a preset threshold to judge the triple confidence, so as to realize the judgment of image semantic misrecognition. The method of the present invention solves the problem that the prior art cannot autonomously judge image semantic misrecognition, and provides a quantitative judgment basis for improving the accuracy of the perception process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving perception methods, and relates to an image semantic misrecognition judgment method based on triple confidence detection. Background Art

[0002] In an autonomous driving system, correct target recognition is a prerequisite for performing all downstream tasks. Although with the development of deep learning, the accuracy of target detection has been greatly improved, due to problems such as environmental noise interference, training data deviation, and similar target features, the problem of misrecognition still cannot be avoided. Moreover, more importantly, the requirements for target detection tasks in the field of autonomous driving are not exactly the same as those in the field of computer vision. Target detection in the field of autonomous driving has more stringent requirements for context information. To distinguish different types of misrecognition problems, the present invention defines misrecognition as image feature misrecognition and image semantic misrecognition. Among them, image feature misrecognition refers to an error caused by improper classification of target image features, such as misclassifying the type of object A as object B. And image semantic misrecognition refers to an error caused by ignoring the context semantics of the target although the classification of target image features is correct. For example, misclassifying a traffic light in transit placed in the cargo box of a pickup truck as a normally working traffic light. This type of misrecognition may cause major traffic accidents in autonomous driving. Therefore, timely detecting and correcting image semantic misrecognition problems is crucial for improving the safety performance of autonomous driving technology. Currently, the method adopted in the field of autonomous driving is to follow the target detection method in the field of computer vision, such as constructing a neural network with a deeper structure and more complex parameters. Although this paradigm can effectively reduce the occurrence of image feature misrecognition, it is powerless for image semantic misrecognition problems.

[0003] Although adopting the architecture of a large language model can effectively solve the problem of image semantic misrecognition, due to the huge number of parameters and deep network structure of the large language model itself, its reasoning process shows more complexity and energy consumption than traditional driving perception or decision-making schemes. If it is applied in the autonomous driving system in the full time domain and unconditionally, it will inevitably cause huge waste of energy and computing power. Therefore, how to construct an image semantic misrecognition judgment method using a neural network to enable the large language model to be triggered and called only in the case of image semantic misrecognition, avoid the waste of computing power, communication, and energy caused by calling the large language model in the full time domain, improve the accuracy and reliability of the environmental perception link of the autonomous driving system, and ensure the large-scale commercial implementation and safe operation of the autonomous driving system is of great significance. Summary of the Invention

[0004] The objective of the present invention is to provide a method for judging image semantic misrecognition based on triple confidence detection, which solves the problem that the prior art cannot autonomously judge image semantic misrecognition and provides a quantitative judgment basis for improving the accuracy of the perception process.

[0005] The technical solution adopted by the present invention is as follows:

[0006] The method for judging image semantic misrecognition based on triple confidence detection is specifically implemented according to the following steps:

[0007] Step 1: Add a historical dynamic gate module on the basis of GRU to construct an HG-GRU model, introduce a GC feature extractor at its input end, and introduce a SEM module at its output end;

[0008] Step 2: Based on HG-GRU and its affiliated models GC and SEM, construct an HGReasoner evaluator, and combine the ResourceRank and RelTrans evaluators to construct a triple confidence detection network;

[0009] Step 3: Capture and fuse each element and its attributes in the road scene to construct a road scene graph;

[0010] Step 4: Extract the triple information of each target in the road according to the scene graph;

[0011] Step 5: Use the triple confidence detection network to solve the confidence results of each triple;

[0012] Step 6: Combine the preset threshold to judge the triple confidence, so as to realize the judgment of image semantic misrecognition.

[0013] The beneficial effects of the present invention are:

[0014] (1) The method of the present invention constructs an HG-GRU model on top of the traditional gated recurrent unit (GRU) network architecture, enhancing the modeling ability for time series data. By introducing a dynamic history gate (DHG) to GRU, the ability to capture sequence information is optimized, enabling the model to effectively maintain and utilize historical information over long time spans, improving the learning effect for multi-step dependencies. This mechanism ensures that the model can accurately identify and respond to key temporal semantic features in complex dynamic environments. At the same time, by combining a path-level graph attention network (PGAT) with a convolutional neural network (CNN), the processing ability for unsequenced data and graph-structured data is enhanced. This integrated method makes the model more flexible and accurate in capturing global and local features, thus enhancing the understanding and decision-making ability for complex scenarios. Meanwhile, the introduced stable enhancement module (SEM), through techniques such as layer normalization, residual connection, and global pooling, significantly improves the training stability and gradient propagation efficiency of the model. Such a design effectively reduces the vanishing gradient phenomenon commonly seen in long sequence training, ensuring the model is more stable during the learning process. Compared with traditional GRU, the method of the present invention, through the introduction of a more efficient structural design, enables it to better capture dynamic changes in complex scenarios, improving the recognition ability of traffic elements in different temporal and spatial contexts. This improvement not only enhances the accuracy of the model but also its robustness in complex environments, providing a more reliable decision-making basis for autonomous driving systems;

[0015] (2) The method of the present invention constructs an HGReasoner evaluator based on the HG-GRU model, and then combines the ResourceRank and RelTrans evaluators to construct a triple confidence detection network. This network effectively discriminates the confidence by evaluating the relationships between traffic elements and their context information, realizing the semantic judgment of triples, and providing a quantitative basis for the semantic understanding of each triple in the scene graph. This method solves the problem of insufficient semantic information processing in current autonomous driving systems when facing complex environments. Whether in urban complex traffic, rural roads, or highways, this method can flexibly adjust response strategies, enhancing the system's understanding ability of traffic scenes, laying a solid foundation for subsequent decision-making and operation guidance, and helping to ensure the efficient and safe operation of autonomous driving systems. In addition, by continuously training and optimizing the model, it can quickly adapt to new situations, improving the flexibility and response speed of the overall system, ensuring that autonomous driving vehicles make accurate judgments and decisions under various conditions;

[0016] (3) The method of the present invention detects the scene Figure 3The confidence of the tuple is used to judge the problem of image semantic misrecognition in the field of autonomous driving, which can significantly reduce the risk of traffic accidents caused by recognition errors, and provide a reliable and effective judgment basis for downstream decision-making of autonomous driving or the intervention and correction of large language models. The research of this comprehensive perception model aims to improve the perception ability of autonomous vehicles in complex scenarios, which will provide new ideas and solutions for the further development of autonomous driving technology. Description of the Drawings

[0017] Figure 1 is the flowchart of the method of the present invention;

[0018] Figure 2 is the structural schematic diagram of the existing GRU model;

[0019] Figure 3 is the structural schematic diagram of the HG-GRU model constructed in the method of the present invention;

[0020] Figure 4 is the structural diagram of the dynamic history gate (DHG) constructed in the method of the present invention;

[0021] Figure 5 is the structural unfolded schematic diagram of the GC and SEM networks constructed in the method of the present invention;

[0022] Figure 6 is the structural schematic diagram of the triple confidence detection network constructed in the method of the present invention;

[0023] Figure 7 is the example diagram of image semantic misrecognition in Embodiment 1 of the present invention;

[0024] Figure 8 is the schematic diagram of converting road scene image information into a scene graph in Embodiment 1 of the present invention;

[0025] Figure 9 is the schematic diagram of extracting triples from partial information of the scene graph in Embodiment 1 of the present invention. Detailed Description of the Invention

[0026] The present invention will be described in detail below with reference to the drawings and specific embodiments.

[0027] The method for judging image semantic misrecognition based on triple confidence detection of the present invention first constructs a Hybrid-Gated GRU (HG-GRU) model and designs a triple confidence detection network; secondly, captures and fuses different traffic elements and their features in the road scene and constructs a Scene Graph; finally, uses the triple confidence detection network to solve the confidence of each triple in the scene graph, and combines the preset threshold to complete the semantic confidence judgment of the triple, so as to realize the judgment of image semantic misrecognition.

[0028] The image semantic misrecognition judgment method based on triple confidence detection of the present invention is as follows Figure 1 shown, and is specifically implemented according to the following steps:

[0029] Step 1, construct an HG-GRU model:

[0030] The construction process of HG-GRU is as follows: as Figure 2 , Figure 3 shown, by adding a Dynamic History Gate (DHG) module on the basis of the traditional GRU module, the path semantic features can be efficiently captured.

[0031] Specifically, the core innovation in the construction process of HG-GRU is reflected in the following aspects:

[0032] (1) Add a Dynamic History Gate (DHG). As Figure 4 shown, the method of the present invention uses a multi-head attention mechanism in the DHG, so that the gating can adapt to different input sequences and dynamically adjust according to different time steps of the sequence. The DHG is realized through three core steps: First, maintain the hidden states of the previous time steps for multi-step historical state modeling; Second, use the multi-head attention mechanism to dynamically weight the historical states to generate a context vector; Finally, further adjust the importance of the historical states on the basis of the context vector, so that the key information in the long-distance and short-distance dependencies can be flexibly captured, especially in the tasks of multi-step dependencies, and the semantic value of multi-historical states can be better mined.

[0033] The specific calculation process of the DHG module is as follows:

[0034] First, use the multi-head attention mechanism to extract context information features from the historical hidden states :

[0035]

[0036]

[0037] In the formula:

[0038] —— represents the dimensions of the key, query, and value vectors of each attention head, and the projection of the query vector is ( is the query projection matrix);

[0039] —— are the key and value respectively ( is the linear projection matrix);

[0040] —— Represents the historical context, i.e., the past hidden state at a time step;

[0041] —— Represents the attention weights calculated through the multi - head attention mechanism.

[0042] The multi - head attention mechanism uses independent attention heads and concatenates the outputs of each head to generate the final context vector :

[0043]

[0044] In the formula:

[0045] —— Represents h the context vector of the u -th head among

[0046] vectors;

[0047] —— Represents the output projection matrix;

[0048] DHG calculates the dynamic weight using the context vector for adjusting the contribution of context features to the hidden state update:

[0049]

[0050] In the formula:

[0051] —— Represent the DHG weight matrix and bias term respectively;

[0052] —— Represents the Sigmoid activation function.

[0053] The output of DHG is calculated through the weight and the context vector :

[0054]

[0055] In the formula:

[0056] —— Hadamard product (element - wise product) operator;

[0057] —— represents a hyperparameter used to adjust the overall output amplitude of the DHG.

[0058] Combine the DHG output with the candidate hidden state of the GRU to update the hidden state :

[0059]

[0060] In the formula:

[0061] —— represents the value of the new candidate state;

[0062] —— represents the forgetting coefficient.

[0063] (2) Additionally, as Figure 5 shown, the method of the present invention introduces a GC feature extractor (Graph attention networks and CNN, GC) in the input part of the HG-GRU model to comprehensively enhance its ability to process non-sequential data and graph-structured data. Specifically, the GC feature extractor integrates a path-level graph attention mechanism module (Path-level Graph Attention Network, PGAT) and a convolutional neural network (CNN), which are respectively responsible for global relationship modeling and local feature extraction, thereby significantly improving the overall performance of the model.

[0064] First, in order to optimize the model's ability to process graph-structured data and non-sequential data, as Figure 5 shown in the GC module, the method of the present invention designs a PGAT module. This module captures the semantic relationships between node pairs by performing a linear transformation on the features of the nodes in the path and combining the attention mechanism, thereby learning high-order feature representations in the path context. After integrating the PGAT into the HG-GRU framework, the model can efficiently capture the global dependencies between sequences. Especially when processing data with complex topological structures, it can more accurately model the interaction relationships between nodes, making up for the deficiencies of traditional GRU in handling such tasks. Second, in order to enhance the traditional GRU's ability to capture local features, the method of the present invention introduces the CNN algorithm.

[0065] The specific calculation process of the GC module is as follows:

[0066] First, through Figure 5 the graph neural network PGAT in the GC feature extractor in extracts the graph structure information of the nodes in the path, dynamically aggregates the neighbor features of the nodes, and assigns different importance weights to different neighbors to extract context features. Second, for the node vectors Perform a linear transformation and map it to a new feature space , and obtain the output matrix :

[0067]

[0068]

[0069] In the formula:

[0070] —— represents the weight matrix, which maps the input features to the hidden layer space;

[0071] —— represents the bias vector;

[0072] N —— represents the number of feature spaces .

[0073] After that, perform row normalization on the output matrix to obtain the attention weights :

[0074]

[0075]

[0076] In the formula:

[0077] —— represents the attention score;

[0078] , —— respectively represent the feature spaces corresponding to nodes in the matrix i and j ;

[0079] —— represents the concatenation operation;

[0080] —— represents the parameter vector of the attention mechanism.

[0081] Aggregate the features of neighbor nodes through the attention weights : :

[0082]

[0083] —— represents the Swish activation function;

[0084] —— represents the neighbor node set of node .

[0085] After that, to enhance the expressive ability, the multi-head attention mechanism is introduced to concatenate the results of independent attention heads to obtain the final output path feature matrix .

[0086] To further capture the semantic information of short sequence data and extract the local continuous dependence features in the path, therefore, a one-dimensional CNN operation is used for each time step in the path. By using its local receptive field mechanism, the features between adjacent nodes are extracted, focusing on the local features between consecutive nodes, and the result after one-dimensional convolution (Conv1D) and pooling is obtained :

[0087]

[0088] In the formula:

[0089] —— represents the MaxPooling1D pooling operation.

[0090] (3) Finally, the Stability Enhancement Module (SEM) is introduced. As shown in the SEM module in Figure 5 , it improves the stability of the model and the ability of gradient propagation by combining layer normalization, residual connection, pooling layer and fully connected layer. Specifically, SEM first performs layer normalization on the candidate hidden states of HG-GRU to stabilize the feature distribution and accelerate training; then enhances the information flow through residual connection to alleviate the problem of gradient vanishing. The global average pooling layer further extracts the global context features of the time series, generates a compact representation with a fixed dimension, and reduces the risk of overfitting; the fully connected layer combines normalization, Dropout and non-linear transformation to strengthen the global feature expression ability and complete the final output. Specifically:

[0091] The specific calculation process of the SEM module is as follows:

[0092] After the calculation of HG-GRU, to enhance the training stability and performance of the model, the hidden state needs to be processed by the SEM module. First, it undergoes residual network and average pooling processing to obtain the output , and then the global feature representation of the sequence is performed, and the scores of each path are calculated through the fully connected layer. For the th path, its score is:

[0093]

[0094]

[0095] Where:

[0096] ——represents a learnable weight matrix used to transform the input Dimensions mapped to hidden states;

[0097] ——Representation layer normalization process;

[0098] ——represents the weight of the fully connected layer;

[0099] —— represents the bias term.

[0100] Figure 5 The structural expansion diagram of the GC and SEM networks is shown, which clearly presents the interaction mechanism and information flow process of the three core modules. By introducing the above modules, local semantic information can be captured more accurately in short sequence tasks, and the modeling ability of multi-granularity and long-term dependencies can be significantly improved.

[0101] Step 2: Build a triplet confidence detection network

[0102] The triplet confidence detection network, namely the TRG-trust model, is used to judge the semantic misrecognition of the autonomous driving system. Its structure is as follows: Figure 6 As shown in the figure, a method for solving the accuracy of autonomous judgment of image semantic misrecognition through a triplet confidence detection network is used to solve the problem of misrecognition of autonomous vehicle targets in the field of autonomous driving. Specifically, the core of the TRG-Trust model is based on a cross neural network structure, and its design is carried out from two dimensions: vertical level and horizontal module. Vertical level: The upper layer is composed of multiple credibility evaluation units, and their outputs converge to the fusion unit of the lower layer. The fusion unit is a multilayer perceptron (MLP) used to generate the final credibility score for each triplet. Horizontal module: For a given triplet The credibility is analyzed progressively from three evaluation units: ResourceRank, RelTrans (RT), and HGReasoner. Three algorithms are used to answer three questions: Is there a relationship between entity pairs (h, t)? Is there a relationship r between entity pairs (h, t)? From a global perspective, is the triple (h, r, t) credible?

[0103] Specifically, the construction process of the triple confidence detection network is as follows:

[0104] (1) Construct a ResourceRank evaluator to determine whether there is a relationship between the head and tail entities;

[0105] (2) Determine whether a certain relationship exists between the head and tail entities through RelTrans;

[0106] (3) Construct the HGReasoner evaluator to evaluate whether the triples extracted in the subsequent steps are credible through social relationships.

[0107] Finally, the confidence results of these three are weighted and integrated into a multi-layer perceptron (MLP) to generate the final triple confidence score, and the resulting score reflects the likelihood that a relationship exists between the triples.

[0108] (1) Construction of the ResourceRank evaluator

[0109] The ResourceRank algorithm is constructed by characterizing the association strength between entity pairs through the idea of resource allocation. This algorithm assumes that the stronger the association between the entity pair (h, t), the more resources flow from the head entity h to the tail entity t through all relevant paths, and the amount of resources finally flowing into t cleverly reflects the association strength between h and t. The ResourceRank (RR) algorithm mainly includes three steps: First, construct a directed graph centered on the head entity h, then iterate the flow of resources in the graph until the resource distribution converges, and calculate the resource retention value of the tail entity t. Finally, integrate other features and output the likelihood that the triple (h,?, t) holds.

[0110] Specifically, each entity is abstracted as a node. If there is a relationship between the entity and , then there will be a directed edge between the nodes and . Therefore, the knowledge graph can be mapped to a weakly connected directed graph, and other nodes in the graph can be reached from the head entity h. In the initial state, the resource amount of h is 1, the resources of other nodes are 0, and the total resources of all nodes is 1. If a certain node does not exist in the graph, its resources are always 0. In addition, there may be multiple relationships between the entity pair ( , ), but there will only be one directed edge from and in the graph. According to the number of these relationships, the bandwidth of each edge will be different. The larger the bandwidth, the more resources flow through the edge. The resources owned by the node h will flow to other nodes through all relevant paths in the graph until the resource distribution is stable, and the resource flow is simulated based on the PageRank algorithm until the distribution is stable. The resource value of the tail entity t can be calculated by the formula:

[0111]

[0112] In the formula:

[0113] —— represents the set of all nodes pointing to the link of the tail entity t;

[0114] —— represents the node 's out-degree;

[0115] —— represents the bandwidth from the node to t;

[0116] —— represents each node in 's resources to t;

[0117] —— represents the total number of resource flows to the node;

[0118] —— Since there may be noise in the knowledge graph, to improve the fault tolerance of the model, it is assumed that the resource flow of each node can jump to any node with the same probability and the probability of this part of the resource flowing to t is .

[0119] After that, by constructing a feature vector to characterize the association strength between the nodes and , as the formula:

[0120]

[0121] In the formula:

[0122] —— represents the resource value from the head node h to the tail node t;

[0123] —— represents the in-degree of the head node h, that is, how many other nodes point to h;

[0124] —— represents the out-degree of the head node h, that is, how many other nodes h points to;

[0125] —— represents the in-degree of the tail node t, that is, how many other nodes point to t;

[0126] —— represents the out-degree of the tail node t, that is, how many other nodes t points to;

[0127] ——Indicates the path depth between the head node h and the tail node t, that is, how many intermediate nodes need to be passed from h to t in the graph.

[0128] Finally, the feature vector needs to pass through a non-linear activation function and a linear transformation to output the final probability value , whose value range is between [0, 1]. The closer the value is to 1, the greater the possibility that there is a relationship between them. As the formula:

[0129]

[0130] In the formula:

[0131] and ——respectively represent the learnable weight matrix and bias during model training.

[0132] (2)Construction of the RelTrans (RT) evaluator

[0133] Although TransE has a significant modeling effect on simple relationships, its modeling effect on complex relationships is very unsatisfactory. Complex relationships such as 1-N, N-1, and N-N are difficult to model. At the same time, since most of the data in this study are one-to-many and many-to-many relationships, therefore, the method of the present invention constructs the RelTrans (RT) algorithm based on TransH. As shown in the RelTrans module in Figure 6 , in the vector space, the same relationship vector can be mapped to the same hyperplane and can be freely translated within the hyperplane without changing its properties. Since it is assumed in TransH that each relationship has a corresponding hyperplane, the translation operation of the relationship does not occur in a single space, but occurs within the hyperplane defined by each relationship. Specifically, first project the head entity h and the tail entity t onto the hyperplane of the relationship r:

[0134]

[0135]

[0136] In the formula:

[0137] ——represents the unit vector of the hyperplane defined by the relationship r;

[0138] and ——respectively represent the projections of the head entity h and the tail entity t on the hyperplane.

[0139] After that, the head entity h will be translated within the hyperplane by the relationship r so that . The triples (car, speed, 100km / h) and (person, speed, 10km / h) are correct. However, according to the translational invariance of the relation vector, (person, speed, 100km / h) must be wrong. Therefore, a credible triple should satisfy , such that after the head entity is translated by the relation, it can approach the tail entity. Thus, the energy function is defined as . When the energy value is smaller, it is considered that the probability of the existence of the relation r between the entity pair (h, t) is greater, and the credibility of (h, r, t) is higher, and vice versa.

[0140] In the RT algorithm, first, the word embedding vector method is used to implement the low-dimensional distributed representation of entities or relations, and the energy value of each triple is calculated , where represents different relations, such as speed, color, and distance, etc. Then, based on the sigmoid function, is converted into the probability that the entity pair (h, t) constitutes the relation , as the formula:

[0141]

[0142] In the formula:

[0143] ——represents the hyperparameter used for model training;

[0144] ——represents the relation threshold. When , the probability value p = 0.5. If , then p > 0.5.

[0145] (3) Based on the HG-GRU model and its affiliated models GC and SEM constructed in step 1, construct the HGReasoner evaluator

[0146] As Figure 6 shown, for the HGReasoner module, each path consists of multiple triple embeddings, and the node vectors after mapping of all paths are , and the whole path is represented as a matrix :

[0147]

[0148] As Figure 6 shown, the last step of HGReasoner is the path filter, and its calculation process is as follows:

[0149] Filter and score the Top- The features of the paths are spliced ​​together, and the Rating of paths Concatenate to form a vector, and pass the vector through a fully connected layer And Sigmoid activation function , get the overall path confidence value of the HGReasoner module :

[0150]

[0151] (4) Evaluator Fusion

[0152] The fusion device based on multi-layer perceptron (MLP) is used to output the final triple credibility value. , and Concatenate into feature vectors:

[0153]

[0154] Then the feature vector Through multiple hidden layers, the final semantic confidence is output. The output layer is a binary classifier that assigns the label y=1 to the true tuple and the label y=0 to the false tuple:

[0155]

[0156] Where:

[0157] ——Indicates the Hidden layer output;

[0158] and ——Respectively represent The parameter matrix and bias term of the hidden layer;

[0159] and ——represent the parameter matrix and bias term of the output layer respectively.

[0160] Step 3: Capture and integrate the elements and attributes in the road scene to construct a road scene graph:

[0161] Obtain real-time time-series scene photos, which are acquired by on-vehicle cameras on autonomous vehicles. Use a convolutional neural network to extract and encode features of the scene graph, and then use a long short-term memory neural network decoding method to capture the text information of traffic scene elements; at the same time, use YOLO V10 to identify traffic signs and traffic lights, and obtain the motion state information of scene elements through binocular vision and tracking algorithms; apply GPS and map software to perceive the GPS information of the vehicle itself in the scene, and obtain the required macroscopic position and time information.

[0162] Step 4: According to the scene graph obtained in Step 3, extract the triple information of each target in the road;

[0163] Suppose the obtained scene graph is G, then the triples extracted from it include multiple detected relationships, such as (h1, r1, t1), (h2, r2, t2), (h3, r3, t3), etc. Among them, h represents the head entity, r represents the relationship, and t represents the tail entity. These triples are used to represent the relationships between elements in the scene, providing an information basis for subsequent semantic understanding and decision-making.

[0164] Step 5: Use the triple confidence detection network constructed in Step 2 to solve the confidence results of each triple in Step 4:

[0165] Take the triples extracted in Step 4 as the input of the triple confidence detection network, and solve their accuracies through a neural network respectively.

[0166] To solve the confidence results of each triple, use the TRG-trust model to progressively analyze the credibility of the triples (h, r, t) generated in Step 4 using the following three evaluation units:

[0167] (1) ResourceRank evaluator: Calculate the association strength between the entity pair (h, t);

[0168] (2) RelTrans evaluator: Judge whether there is a relationship r between the head entity and the tail entity, and evaluate it through an energy value;

[0169] (3) HGReasoner evaluator: Analyze the credibility of the triple (h, r, t) from a global perspective, and verify it by integrating historical data and context information.

[0170] Finally, weight and integrate the confidence results of these three into a multi-layer perceptron (MLP) to generate the final triple confidence score, and the obtained result reflects the possibility of the existence of a relationship between the triples.

[0171] Step 6: Combine a preset threshold to judge the triple confidence obtained in Step 5, so as to realize the judgment of image semantic misrecognition.

[0172] By setting a confidence threshold in semantic detection in advance, when the confidence of a triple is lower than this threshold, the system will identify it as a misidentification situation, the system will issue an alarm, and hand over the error result to the downstream task for processing. If it is not lower than the confidence threshold, it indicates that no misidentification has occurred, and then the next task will be executed.

[0173] Embodiment 1:

[0174] The specific implementation method of this embodiment is as follows:

[0175] As Figure 7 shown, in reality, a car is driving normally on the highway, but mistakenly regards the vehicle on the road billboard as a real vehicle, resulting in an emergency brake. This is a misidentification. For this misidentification, applying the method of the present invention, the specific steps are as follows:

[0176] Step 1-2: Construct a triple confidence detection network;

[0177] Step 3: Obtain real-time sequential scene photos. The scene photos are obtained according to the on-vehicle camera on the autonomous vehicle. Use an algorithm to extract the information and attributes of each element in the road, and the finally generated scene graph is as Figure 8 shown;

[0178] Step 4: Perform triple extraction on the scene graph obtained in Step 3 to obtain multiple triples (the guardrail on the left / right of the autonomous vehicle, is, no entry), (the autonomous vehicle, driving on, the lane), (the autonomous vehicle, of, driving speed), (the driving speed, is, 60 km / h), (there is a vehicle in front of the autonomous vehicle), (the vehicle, driving in, the air). As Figure 9 After extracting a part of the content in the scene graph, a triple (the vehicle, driving in, the air) is obtained, where "the vehicle" is the head entity (h), "driving in" is the relationship (r), and "the air" is the tail entity (t), so as to facilitate understanding and analyzing the relationship between the characteristics of each element in the scene graph.

[0179] Step 5: Load the triples obtained in Step 4 into the triple confidence detection network constructed in Step 2.

[0180] The confidence of the triples in Figure 9 is detected through the triple confidence detection network. For (the vehicle, driving in, the air), its credibility is analyzed step by step from three evaluation units: ResourceRank, RelTrans, and HGReasoner. Specifically, it is to evaluate whether there is a relationship between "the vehicle" and "the air"? Whether there is a relationship "driving in"? From a global perspective, it is judged whether (the vehicle, driving in, the air) in the scene is credible?

[0181] Step 6: Combine with a preset threshold to judge the confidence of the triples obtained in Step 5, so as to realize the judgment of image semantic misdetection.

[0182] After solving, the confidence of the triple (vehicle, driving on, air) is 0.27, which is lower than the set threshold of 0.5, indicating that there is a semantic deviation in this triple. In addition, judged by common sense in reality, vehicles on the road do not drive in the air. Therefore, after this evaluation, the association strength between "vehicle" and "air" is weak, so the relationship "driving on" between the two does not hold. The conclusion obtained by using the method of the present invention conforms to common sense. Therefore, it can be concluded that there is no real "driving on" relationship between the head entity "vehicle" and the tail entity "air", indicating that the triple is misrecognized and handed over to the downstream task for processing.

[0183] Embodiment 2:

[0184] The method for judging image semantic misrecognition based on triple confidence detection in this embodiment is specifically implemented according to the following steps:

[0185] Step 1: Add a historical dynamic gate module on the basis of GRU to construct an HG-GRU model, introduce a GC feature extractor at its input end, and introduce a SEM module at the output end;

[0186] Step 2: Based on HG-GRU and its affiliated models GC and SEM, construct an HGReasoner evaluator, and combine the ResourceRank and RelTrans evaluators to construct a triple confidence detection network;

[0187] Step 3: Capture and fuse the elements and their attributes in the road scene to construct a road scene graph;

[0188] Step 4: Extract the triple information of each target in the road according to the scene graph;

[0189] Step 5: Use the triple confidence detection network to solve the confidence results of each triple;

[0190] Step 6: Combine with a preset threshold to judge the triple confidence, so as to realize the judgment of image semantic misrecognition.

[0191] Embodiment 3:

[0192] On the basis of Embodiment 2, in Step 1, the specific calculation process of the historical dynamic gate module is as follows:

[0193] First, use the multi-head attention mechanism to extract the context information features from the historical hidden state :

[0194]

[0195]

[0196] In the formula:

[0197] —— represents the dimensions of the key, query, and value vectors for each attention head. The projection of the query vector is , is the query projection matrix;

[0198] —— are the key and value respectively, where is the linear projection matrix;

[0199] —— represents the historical context, that is, the hidden states of the past time steps;

[0200] —— represents the attention weights calculated through the multi - head attention mechanism;

[0201] The multi - head attention mechanism uses independent attention heads and concatenates the outputs of each head to generate the final context vector :

[0202]

[0203] In the formula:

[0204] —— represents h the context vector of the u - th head among

[0205] —— represents the output projection matrix;

[0206] —— represents the vector concatenation operation;

[0207] Secondly, the dynamic weight is calculated through the context vector for adjusting the contribution of the context features to the update of the hidden state:

[0208]

[0209] In the formula:

[0210] —— represent the DHG weight matrix and the bias term respectively;

[0211] —— represents the Sigmoid activation function;

[0212] Through the weight and the context vector Calculate the output of the dynamic history gate :

[0213]

[0214] In the formula:

[0215] —— Hadamard product (element-wise product) operator;

[0216] —— Represents a hyperparameter used to adjust the overall output amplitude of the DHG;

[0217] Finally, combine the output with the candidate hidden state of the gated recurrent unit module to update the hidden state :

[0218]

[0219] In the formula:

[0220] —— Represents the value of the new candidate state;

[0221] —— Represents the forgetting coefficient.

[0222] Example 4:

[0223] Based on Example 3, in step 1, the GC feature extractor integrates the path-level graph attention mechanism module PGAT and the convolutional neural network CNN, which are responsible for global relationship modeling and local feature extraction respectively;

[0224] The specific calculation process of the GC feature extractor is as follows:

[0225] First, extract the graph structure information of the nodes in the path through PGAT, dynamically aggregate the neighbor features of the nodes, assign different importance weights to different neighbors at the same time, extract the context features, and perform a linear transformation on the node vectors after mapping for all paths to map to a new feature space to obtain the output matrix :

[0226]

[0227]

[0228] In the formula:

[0229] —— represents the weight matrix, which maps the input features to the hidden layer space;

[0230] —— represents the bias vector;

[0231] N —— represents the number of feature spaces;

[0232] Secondly, after normalizing the output matrix H, the attention weights are obtained :

[0233]

[0234]

[0235] In the formula:

[0236] —— represents the attention score;

[0237] , —— respectively represent the feature spaces corresponding to nodes in the matrix i and j ;

[0238] —— represents the concatenation operation;

[0239] —— represents the parameter vector of the attention mechanism;

[0240] Through the attention weights aggregate the features of neighbor nodes :

[0241]

[0242] —— represents the Swish activation function;

[0243] —— represents the set of neighbor nodes of node ;

[0244] Again, introduce the multi - head attention mechanism, and concatenate the results of independent attention heads to obtain the final output path feature matrix ;

[0245] Finally, for At each time step in the path, a one-dimensional CNN operation is used. Utilizing its local receptive field mechanism, it extracts the features between adjacent nodes, focuses on the local features between consecutive nodes, and obtains the result after one-dimensional convolution Conv1D and pooling. :

[0246]

[0247] In the formula:

[0248] —— represents the MaxPooling1D pooling operation.

[0249] Example 5:

[0250] Based on Example 4, in step 1, the specific calculation process of the SEM module is as follows:

[0251] After the calculation of the HG-GRU model, first, the hidden state is processed by a residual network and average pooling to obtain the output . Then, the global feature representation of the sequence is performed, and the score of each path is calculated through a fully connected layer. For the th path, its score is:

[0252]

[0253]

[0254] In the formula:

[0255] —— represents the Sigmoid activation function;

[0256] —— represents a learnable weight matrix used to map the input to the dimension of the hidden state;

[0257] —— represents the layer normalization process;

[0258] —— represents the weight of the fully connected layer;

[0259] —— represents the bias term.

[0260] Example 6:

[0261] Based on Example 5, in the triple confidence detection network of step 2:

[0262] The ResourceRank evaluator is used to determine whether there is a relationship between the head and tail entities;

[0263] The RelTrans evaluator is used to determine whether a certain relationship exists between the head and tail entities;

[0264] The HGReasoner evaluator is used to evaluate whether a triple is credible through social relationships;

[0265] Finally, the confidence results of these three evaluators are weighted and integrated into a multi-layer perceptron to generate the final triple confidence score, and the result reflects the possibility that there is a relationship between triples.

Claims

1. An image semantic misrecognition judgment method based on triple confidence detection, characterized in that: Follow the steps below to implement it: Step 1: Add a historical dynamic gate module based on GRU to build the HG-GRU model, introduce the GC feature extractor at its input end, and introduce the SEM module at its output end; Step 2: Based on HG-GRU and its subsidiary models GC and SEM, HGReasoner evaluator is constructed, and triple confidence detection network is constructed by combining ResourceRank and RelTrans evaluators; Step 3: Capture and integrate the elements and attributes in the road scene to construct a road scene graph; Step 4: Extract triplet information of each target on the road according to the scene graph; Step 5: Utilize the triple confidence detection network to solve the confidence results of each triple; Step 6: Combined with the preset threshold, the confidence of the triplet is judged to achieve image semantic misrecognition judgment; In step 1, the specific calculation process of the historical dynamic gate module is as follows: First, a multi-head attention mechanism is used to extract contextual information features from historical hidden states. : Where: ——denotes the dimensions of the key, query, and value vectors of each attention head. The projection of the query vector is , is the query projection matrix; ——are keys and values, respectively, where is the linear projection matrix; ——Indicates historical context, i.e. the past The hidden state of time steps; ——represents the attention weight calculated by the multi-head attention mechanism; Use of multi-head attention mechanism The output of each head is concatenated to generate the final context vector : Where: --express h In the vector u The context vector of each head; ——represents the output projection matrix; ——Indicates vector concatenation operation; Secondly, through the context vector Calculating dynamic weights , which is used to adjust the contribution of context features to the hidden state update: Where: ——represent the DHG weight matrix and bias term respectively; ——Indicates the Sigmoid activation function; By weight and context vector Calculate the output of the dynamic history gate : Where: ——Hadamard product operator; ——represents a hyperparameter used to adjust the overall output amplitude of DHG; Finally, the output Combined with the candidate hidden state of the gated recurrent meta-module, update the hidden state : Where: ——Indicates the new candidate state value; ——Indicates the forgetting coefficient.

2. The image semantic misrecognition judgment method based on triple confidence detection according to claim 1 is characterized in that: In step 1, the GC feature extractor integrates the path-level graph attention mechanism module PGAT and the convolutional neural network CNN, which are responsible for global relationship modeling and local feature extraction respectively; The specific calculation process of the GC feature extractor is as follows: First, PGAT is used to extract the graph structure information of the nodes in the path, dynamically aggregate the neighbor features of the nodes, assign different importance weights to different neighbors, extract context features, and map the node vectors of all paths. Perform linear transformation and map it to the new feature space , and get the output matrix : Where: ——represents the weight matrix, which maps the input features to the hidden space; —— represents the bias vector; N——represents the feature space The number of Secondly, the output matrix H is normalized to obtain the attention weight : Where: ——Indicates the attention score; , ——represents matrices respectively Midpoint i and j The corresponding feature space; ——Indicates splicing operation; ——Parameter vector representing the attention mechanism; By attention weight Aggregate neighbor node characteristics : —— represents the Swish activation function; ——Indicates the node The set of neighbor nodes of Again, the multi-head attention mechanism is introduced to Results of independent attention heads Splice to get the final output path feature matrix ; Finally, Each time step in the path uses a one-dimensional CNN operation, using its local receptive field mechanism to extract features between adjacent nodes, focusing on local features between consecutive nodes, and obtaining the result after one-dimensional convolution Conv1D and pooling. : Where: ——Indicates MaxPooling1D pooling operation.

3. The image semantic misrecognition judgment method based on triple confidence detection according to claim 1 is characterized in that: In step 1, the specific calculation process of the SEM module is as follows: After the calculation of the HG-GRU model, the hidden state is first Perform residual network and average pooling to get the output , then the global feature representation of the sequence is performed, and the score of each path is calculated through the fully connected layer. Paths, their scores for: Where: ——Indicates the Sigmoid activation function; ——represents a learnable weight matrix used to transform the input Dimensions mapped to hidden states; ——Representation layer normalization process; ——represents the weight of the fully connected layer; —— represents the bias term.

4. The image semantic misrecognition judgment method based on triple confidence detection according to claim 1, characterized in that: In the triple confidence detection network of step 2: The ResourceRank evaluator is used to determine whether there is a relationship between the head and tail entities; The RelTrans evaluator is used to determine whether a relationship exists between the head and tail entities; The HGReasoner evaluator is used to evaluate whether a triple is credible through social relations; Finally, the confidence results of these three evaluators are weighted and integrated into the multi-layer perceptron to generate the final triple confidence score, which reflects the possibility of the existence of a relationship between the triplets.

5. The image semantic misrecognition judgment method based on triple confidence detection according to claim 1 or 4, characterized in that: In step 2, the ResourceRank algorithm includes the following steps: First, a directed graph is constructed with the head entity h as the center. Then, the flow of resources is iterated in the graph until the resource distribution converges. The resource retention value of the tail entity t is calculated. Finally, other features are integrated to output the possibility of the triple (h, ?, t) being true. Among them: the resource retention value of the tail entity t Calculated by the following formula: Where: ——represents the set of all nodes that point to the tail entity t link; ——Indicates the node The out-degree of ——Indicates slave node The bandwidth to t; --express Each node in to t's resources; ——Indicates the total number of resources flowing to the node; ——In order to improve the fault tolerance of the model, it is assumed that the resource flow of each node has the same probability Jump to any node, and the probability that this part of the resources flows to t is ; Then, by constructing a feature vector Characterization Node and The strength of the association between them is as follows: Where: ——Indicates the resource value from the head node h to the tail node t; ——Indicates the in-degree of the head node h, that is, how many other nodes point to h; ——Indicates the out-degree of the head node h, that is, how many other nodes h points to; ——Indicates the in-degree of the tail node t, that is, how many other nodes point to t; ——Indicates the out-degree of the tail node t, that is, how many other nodes t points to; ——Indicates the path depth from the head node h to the tail node t, that is, how many intermediate nodes are needed to go from h to t in the graph; Finally, the eigenvector It is necessary to output the final probability value through nonlinear activation function and linear transformation , whose value range is between [0,1]. The closer the value is to 1, the greater the possibility that there is a relationship between them, as shown in the formula: Where: and ——represent the weight matrix and bias that can be learned during model training.

6. The image semantic misrecognition judgment method based on triple confidence detection according to claim 1 or 4, characterized in that: In step 2, the RelTrans algorithm is built based on TransH, and its energy function is defined as: Where: ——represent the head entity, relationship and tail entity respectively; and ——represent the projections of the head entity h and the tail entity t on the hyperplane respectively; When the energy value The smaller it is, the greater the probability that there is a relationship r between the entity pair (h, t), and the higher the credibility of (h, r, t), otherwise the lower the credibility; In the RelTrans algorithm, the word embedding vector method is first used to implement a low-dimensional distributed representation of entities or relations, and the energy value of each triple is calculated. ,in Represents different relationships; then based on the sigmoid function Convert to entity pairs (h, t) to form a relationship The probability is as follows: Where: ——represents the hyperparameters used for model training; ——Indicates the relationship threshold, when When p =0.5, if ,but p >0.

5.

7. The image semantic misrecognition judgment method based on triple confidence detection according to claim 1 or 4, characterized in that: In step 2, each path of the HGReasoner evaluator consists of multiple triple embeddings, and the node vectors of all paths after mapping are , the entire path is represented as a matrix : The last step of HGReasoner is the path filter, which is calculated as follows: Filter and score the Top- The features of the paths are spliced ​​together, and the Rating of paths Concatenate to form a vector, and pass the vector through a fully connected layer And Sigmoid activation function , get the overall path confidence value of the HGReasoner module : 。 8. The image semantic misrecognition judgment method based on triple confidence detection according to claim 4 is characterized in that: In step 2, the multilayer perceptron first combines the results of the HGReasoner evaluator, ResourceRank evaluator, and RelTrans evaluator , and Concatenate into feature vectors: Then the feature vector Transformed through multiple hidden layers, the final semantic confidence is output; the output layer is a binary classifier that assigns the label y=1 to the true tuple and the label y=0 to the false tuple: Where: ——indicates the head entity; ——Indicates the Sigmoid activation function; ——represents hyperparameters; ——Indicates the Hidden layer output; and ——Respectively represent the The parameter matrix and bias term of the hidden layer; and ——represent the parameter matrix and bias term of the output layer respectively.

Citation Information

Patent Citations

  • Unbiased scene graph construction method, system and equipment based on generative template

    CN117709454A

  • Driving pressure interpretable prediction method based on layered road environment scene graph

    CN119380305A