Target detection method and device based on Gaussian mixture model, vehicle and medium

By using Gaussian hybrid model in the Transformer model to model key sequences and calculate attention scores, the problem of high memory consumption and complexity in object detection is solved, and a more efficient and interpretable object detection effect is achieved.

CN120014576APending Publication Date: 2025-05-16CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510085112.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The calculation memory consumption of the Transformer model in object detection is an exponential increase, which increases the time for object detection, resulting in high computational complexity and high parameters, affecting the performance and safety of the autonomous driving system.

Method used

Using the object detection method based on the Gaussian mixed model, the calculation cost of the Transformer model is reduced by modeling each key in the key sequence into a Gaussian mixed model containing a preset number of Gaussian distributions.

Benefits of technology

By reducing computational costs and improving the interpretability of model output, better target detection results are achieved and the performance and safety of the autonomous driving system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014576A_ABST
    Figure CN120014576A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a target detection method and device based on a Gaussian mixture model, a vehicle and a medium, and the method comprises the steps: carrying out the image feature extraction of a driving scene image, obtaining a feature sequence, carrying out the extraction through a linear projection layer, and obtaining a query, a key and a value; each key is modeled as a preset number of Gaussian mixture models of Gaussian distribution, the probabilities of the queries in the Gaussian mixture models are calculated, and the attention scores of the current key and all the queries are calculated based on the Gaussian mixture models corresponding to the keys and the probabilities of the queries in the Gaussian mixture models; then attention scores of all keys and all queries are obtained, dot product multiplication is carried out on the attention scores and a value sequence to obtain a target feature sequence, and finally a detection result of a target is determined. Through a self-attention mechanism based on Gaussian mixture distribution, the calculation cost of a self-attention model is reduced, the explanatability of model output is improved, and the detection accuracy of the target is improved. And a better target detection result is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a target detection method, device, vehicle and medium based on a Gaussian mixture model. Background Art

[0002] At present, Transform-based artificial intelligence algorithms have surpassed traditional convolutional neural network (CNN) algorithms in many fields and achieved breakthrough progress. For example, the ViT (Vision Transformer) algorithm proposed by JANAI J et al. broke the dominance of CNN in image classification tasks for the first time and promoted the integration of vision and text.

[0003] The self-attention mechanism is a powerful attention model that can help better understand the structure and characteristics of input data and improve the accuracy and speed of target detection. Therefore, the Transformer algorithm can be put into practical applications such as autonomous driving, and target detection and positioning can be performed on environmental images collected in autonomous driving to overcome the complex environment, diverse data types, and multi-target interactions in open world scenarios. However, in order to improve the accuracy of Transfoemer's target detection, a large amount of labeled data and powerful computing resources are required to optimize the Transformer model. The three main problems faced include: Since the Transformer architecture requires interactive calculation of all position information between feature sequences, when facing high-precision images or videos, as the feature sequence input increases, the consumption of computing memory grows exponentially, and since all position information of the feature sequence needs to be interactively calculated, the time for target detection will be increased, making the model calculation complex and the number of parameters high. The real-time and interpretability of the model are facing challenges, affecting the performance and safety of the autonomous driving system. Summary of the invention

[0004] In view of this, the present invention provides a target detection method, device, vehicle and medium based on a Gaussian mixture model to solve the problem that with the increase of feature sequence input, the consumption of Transformer computing memory grows exponentially, and the time of target detection is increased, which makes the model calculation complex and the number of parameters high, and the real-time and interpretability of the model face challenges, affecting the performance and safety of the autonomous driving system.

[0005] In a first aspect, the present invention provides a target detection method based on a Gaussian mixture model, the method comprising: obtaining a driving scene image of a vehicle during driving, and performing image feature extraction on the driving scene image to obtain a first feature sequence; extracting the first feature sequence through a multi-layer linear projection layer to obtain a query sequence, a key sequence and a value sequence corresponding to the first feature sequence, and modeling each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions; based on the Gaussian mixture models corresponding to all keys in the key sequence, calculating the probabilities of all queries in the query sequence in the Gaussian mixture model respectively; based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture model respectively, calculating the attention scores of the current key and all queries in the key sequence, traversing all keys in the key sequence to obtain the attention scores of all keys and all queries; performing dot product multiplication of the attention scores of all keys and all queries with the value sequence to obtain a target feature sequence, and determining the detection result of the target in the driving scene image based on the target feature sequence.

[0006] The target detection method based on Gaussian mixture model provided by the present invention extracts image features from a driving scene image to obtain a first feature sequence, then extracts the first feature sequence through a multi-layer linear projection layer to obtain a query sequence, a key sequence and a value sequence, and models each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions, calculates the probabilities of all queries in the query sequence in the Gaussian mixture model, calculates the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture model, traverses all keys to obtain the attention scores of all keys and all queries, then dot-products the attention scores of all keys and all queries with the value sequence to obtain a target feature sequence, and finally determines the detection result of the target in the driving scene image based on the target feature sequence, reduces the computational cost of the Transformer model through a self-attention mechanism based on a mixed Gaussian distribution, improves the interpretability of the model output, and achieves better target detection results.

[0007] In an optional embodiment, the target feature sequence is a feature sequence processed by a single attention head in a self-attention network, and determining the detection result of the target in the driving scene image further includes: traversing all attention heads in the self-attention network to obtain target feature sequences corresponding to all attention heads; transforming the target feature sequences corresponding to all attention heads through a total transformation matrix, and then performing a residual connection with the first feature sequence to obtain an output sequence; and determining the detection result of the target in the driving scene image based on the output sequence.

[0008] In an optional implementation, the Gaussian mixture model corresponding to the key is The calculating the probabilities of all queries in the query sequence in the Gaussian mixture model based on the Gaussian mixture model corresponding to all keys in the key sequence includes: calculating the probabilities of all queries in the query sequence in the Gaussian mixture model using the following total probability formula:

[0009]

[0010] Among them, p(q i ) represents the probability of the i-th query in the Gaussian mixture model, q i represents the i-th query in the query sequence; π jr =p(z r =1|t j =1) is represented by the rth Gaussian distribution of the jth key, z r represents the rth Gaussian distribution in the Gaussian mixture model, t j represents a T-dimensional binary random variable, indicating the position of the key; M represents the number of Gaussian distributions; μ jr represents the mean of the rth Gaussian distribution of the jth key; represents the covariance matrix of the r-th Gaussian distribution of the j-th key; I represents the identity matrix.

[0011] The present invention can improve the interpretation ability of each key and reduce the chance of learning redundant heads by modeling each key as a Gaussian mixture model of M Gaussian distributions.

[0012] In an optional implementation, the calculating the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture models respectively includes: calculating the posterior probability of the current key and all queries using the following Bayesian formula, wherein the posterior probability represents the attention score of the current key and all queries in the key sequence:

[0013]

[0014] Among them, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; k jr represents the r-th Gaussian distribution of the j-th key in the key sequence.

[0015] In an optional implementation, the calculating the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture models respectively includes: calculating the attention scores of the current key and all queries in the key sequence using the following Gaussian kernel function between the current key and all queries:

[0016]

[0017] Among them, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; k jr represents the r-th Gaussian distribution of the j-th key in the key sequence.

[0018] The present invention calculates the attention scores of the current key and all queries by utilizing the Gaussian kernel function between the current key and all queries, which can meet practical applications and further improve the accuracy of target detection.

[0019] In an optional implementation, the calculating the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture models respectively includes: selecting a Gaussian distribution with the largest probability density function from the Gaussian mixture model corresponding to the current key as the Gaussian distribution for calculating the attention score; and calculating the attention scores of the current key and all queries in the key sequence using the following Gaussian kernel function between the current key and all queries:

[0020]

[0021] in, Represents the Gaussian distribution with the largest probability density function, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; k jr represents the r-th Gaussian distribution of the j-th key in the key sequence.

[0022] The present invention makes an approximate comparison between the query and only one of the most critical Gaussian distributions, thereby reducing the number of parameters and the amount of calculation.

[0023] In a second aspect, the present invention provides a target detection device based on a Gaussian mixture model, the device comprising: a feature sequence acquisition module, used to acquire a driving scene image of a vehicle during driving, and perform image feature extraction on the driving scene image to obtain a first feature sequence; a Gaussian modeling module, used to extract the first feature sequence through a multi-layer linear projection layer, obtain a query sequence, a key sequence and a value sequence corresponding to the first feature sequence, and model each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions; a query probability calculation module, used to calculate the probabilities of all queries in the query sequence in the Gaussian mixture model based on the Gaussian mixture models corresponding to all keys in the key sequence; an attention score calculation module, used to calculate the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture model, traverse all keys in the key sequence, and obtain the attention scores of all keys and all queries; a target detection module, used to perform dot product multiplication of the attention scores of all keys and all queries with the value sequence to obtain a target feature sequence, and determine the detection result of the target in the driving scene image based on the target feature sequence.

[0024] In a third aspect, the present invention provides a vehicle, comprising a controller, the controller comprising a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the Gaussian mixture model-based target detection method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0025] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the Gaussian mixture model-based target detection method of the first aspect or any corresponding embodiment thereof.

[0026] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the target detection method based on a Gaussian mixture model of the above-mentioned first aspect or any corresponding embodiment thereof.

[0027] The target detection method based on Gaussian mixture model provided by the present invention extracts image features from a driving scene image to obtain a first feature sequence, then extracts the first feature sequence through a multi-layer linear projection layer to obtain a query sequence, a key sequence and a value sequence, and models each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions, calculates the probabilities of all queries in the query sequence in the Gaussian mixture model, calculates the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture model, traverses all keys to obtain the attention scores of all keys and all queries, then dot-products the attention scores of all keys and all queries with the value sequence to obtain a target feature sequence, and finally determines the detection result of the target in the driving scene image based on the target feature sequence, reduces the computational cost of the Transformer model through a self-attention mechanism based on a mixed Gaussian distribution, improves the interpretability of the model output, and achieves better target detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0029] Figure 1 is a schematic diagram of a process of a target detection method based on a Gaussian mixture model according to an embodiment of the present invention;

[0030] Figure 2 is an example diagram of the architecture of the self-attention mechanism according to an embodiment of the present invention;

[0031] Figure 3 is a flowchart of another target detection method based on a Gaussian mixture model according to an embodiment of the present invention;

[0032] Figure 4 is a flowchart of a target detection method based on a Gaussian mixture model according to an embodiment of the present invention;

[0033] Figure 5a is the initial environment image;

[0034] Figure 5b It is an image processed using the traditional attention mechanism;

[0035] Figure 5c is an image processed by the object detection method based on the Gaussian mixture model according to an embodiment of the present invention;

[0036] Figure 6a is an attention distribution diagram of a traditional attention mechanism according to an embodiment of the present invention;

[0037] Figure 6b is an attention distribution diagram of a self-attention mechanism in a target detection method based on a Gaussian mixture model according to an embodiment of the present invention;

[0038] Figure 7 is a structural block diagram of a target detection device based on a Gaussian mixture model according to an embodiment of the present invention;

[0039] Figure 8 Schematic diagram of the hardware structure of the controller according to the embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0041] At present, the research on target detection and positioning applied to autonomous driving by Transformer mainly faces the following problems: First, there is the problem of high memory consumption. Since the Transformer architecture needs to interactively calculate all the position information between feature sequences, when facing high-precision images or videos, as the feature sequence input increases, the consumption of computing memory grows exponentially. In addition, since all the position information between feature sequences needs to be interactively calculated, the target detection time will be increased, which will affect the performance and safety of the autonomous driving system. Secondly, there is the problem of interpretability of model calculation and design, including the fact that the internal calculation of Transformer is relatively complex, and the results are difficult to explain their effective meaning. There are also some studies on solving the problem of difficult model convergence. Ordinary DETR algorithms generally require more than three hundred iterations to converge, which is a huge consumption of computing resources.

[0042] The multi-head attention mechanism module in the Transformer architecture has the problems of large computational complexity and computational redundancy, which makes the model computational complexity and parameter amount high, and the real-time and interpretability of the model face challenges, which is not conducive to the application in related fields such as the field of autonomous driving. In order to solve the above problems, the present invention optimizes the design of the multi-head attention mechanism module based on the image recognition algorithm and target detection algorithm of Transformer, and proposes a target detection method based on the self-attention mechanism of mixed Gaussian distribution, which reduces the computational cost of the Transformer model, improves the interpretability of the model output, and achieves better detection effect.

[0043] According to an embodiment of the present invention, an embodiment of a target detection method based on a Gaussian mixture model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0044] In this embodiment, a target detection method based on a Gaussian mixture model is provided, which can be used in the above-mentioned controller. Figure 1 is a flow chart of a target detection method based on a Gaussian mixture model according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0045] Step S101, obtaining a driving scene image of a vehicle during driving, and extracting image features from the driving scene image to obtain a first feature sequence.

[0046] The embodiment of the present invention obtains a driving scene image of a vehicle during driving, such as Figure 2 As shown, it can be passed through a backbone network and a feature pyramid to extract image features to obtain feature maps of different sizes, and then the features of all feature maps can be converted into sequences of the same feature dimension through a specific linear projection layer, where the sequence is in the form of a sequence to better input into the Transformer for calculation. The sequence of the same feature dimension is added to the initialized position code as the first feature sequence that is finally input into the self-attention network.

[0047] Step S102: extract the first feature sequence through a multi-layer linear projection layer to obtain a query sequence, a key sequence and a value sequence corresponding to the first feature sequence, and model each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions.

[0048] The self-attention network of the embodiment of the present invention may include single-head attention or multi-head attention, wherein the number of attention heads may be set according to the actual accuracy and real-time performance of target detection. In a single attention head, the first feature sequence is extracted through three linear projection layers to obtain the query, key and value of each feature respectively, and the query sequence Q, key sequence K and value sequence V corresponding to the first feature are obtained, and the query sequence q is defined as i ∈Q and key sequence k j ∈K, and then we can set the key data distribution r Gaussian distributions to get a mixed Gaussian distribution, that is, the key k at position j in the key sequence j The model is a Gaussian mixture model containing M Gaussian distributions, which is only used as an example.

[0049] Step S103, based on the Gaussian mixture models corresponding to all the keys in the key sequence, the probabilities of all the queries in the query sequence in the Gaussian mixture models are calculated.

[0050] In the embodiment of the present invention, the distribution of the query sequence Q may be considered to be a mixed Gaussian distribution of all key sequences in K, and p(q i ) distribution is modeled as follows:

[0051]

[0052] Among them, π j For a given k j Prior distribution, let t be a T-dimensional binary random variable, where t j is equal to 1, and all other elements are equal to 0, which can represent key k j The position of j represents the mean of the Gaussian distribution of the jth key; Represents the covariance matrix of the Gaussian distribution of the j-th bond.

[0053] Step S104, based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture models, calculate the attention scores of the current key and all queries in the key sequence, traverse all keys in the key sequence, and obtain the attention scores of all keys and all queries.

[0054] The embodiment of the present invention can calculate the posterior probability according to the Bayesian formula, that is, given k j Next, q i k with respect to the relevant position j The similarity is calculated as follows:

[0055]

[0056] Where T represents the dimension of the feature sequence. To reduce the consumption of computational memory, both Q and K can be regarded as standardized distributions, and the corresponding variances are the same. Therefore, it can be simplified to:

[0057]

[0058] The embodiment of the present invention can calculate the attention scores of the current key and all queries in the key sequence based on formula (3), and then traverse all the keys in the key sequence to obtain the attention scores of all keys and all queries.

[0059] Step S105, dot-product the attention scores and value sequences of all keys and all queries to obtain a target feature sequence, and determine the detection result of the target in the driving scene image based on the target feature sequence.

[0060] First, the standard single-head self-attention module uses the dot product multiplication of the feature sequence to calculate the cosine similarity, and calculates the similarity between the query Q and the key K through softmax probability normalization. The formula is as follows:

[0061]

[0062] Among them, Attention(Q,K,V)=[h1,K,h N ] represents the result of feature sequence purification after attention mechanism feature extraction, A represents the attention score between Q and K, D k Represents the dimension of the feature sequence.

[0063] The embodiment of the present invention can perform dot product multiplication of the attention scores and value sequences of all keys and all queries through the following formula to obtain a target feature sequence, and can generate a series of candidate regions based on the target feature sequence, and then classify and locate the candidate regions to determine the target category and bounding box position. In order to remove duplicate detection results, a non-maximum suppression algorithm is usually used to screen the candidate boxes, retain the best target detection results, and finally post-process the target detection results to determine the detection results of the targets in the driving scene image, as an example only.

[0064]

[0065] It can be seen that the attention scores calculated by formula (5), formula (3) and formula (4) are similar, which proves the correctness of formula (5). In the single-head attention unit, the query q i and key k j The attention score between is the posterior probability p(t j =1|q i ), indicating the key k j Explain query q iThe weight it occupies can be said to be the similarity, which in turn determines how much attention the query at position i in the first feature sequence should pay to the sequence information at position j.

[0066] The target detection method based on Gaussian mixture model provided in this embodiment extracts image features from a driving scene image to obtain a first feature sequence, extracts the first feature sequence through multiple linear projection layers to obtain a query sequence, a key sequence and a value sequence, and models each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions, calculates the probabilities of all queries in the query sequence in the Gaussian mixture model, calculates the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture model, traverses all keys to obtain the attention scores of all keys and all queries, and then performs dot product multiplication of the attention scores of all keys and all queries with the value sequence to obtain a target feature sequence, and finally determines the detection result of the target in the driving scene image based on the target feature sequence, reduces the computational cost of the Transformer model through a self-attention mechanism based on a mixed Gaussian distribution, improves the interpretability of the model output, and achieves better target detection results.

[0067] In this embodiment, a target detection method based on a Gaussian mixture model is provided, which can be used in the above-mentioned controller. Figure 3 is a flow chart of a target detection method based on a Gaussian mixture model according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0068] Step S301, obtaining a driving scene image of the vehicle during driving, and extracting image features from the driving scene image to obtain a first feature sequence. Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.

[0069] Step S302: extract the first feature sequence through a multi-layer linear projection layer to obtain the query sequence, key sequence and value sequence corresponding to the first feature sequence, and model each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions. Figure 1 Step S102 of the illustrated embodiment will not be described in detail here.

[0070] Step S303, based on the Gaussian mixture models corresponding to all the keys in the key sequence, the probabilities of all the queries in the query sequence in the Gaussian mixture models are calculated.

[0071] Specifically, the Gaussian mixture model corresponding to the key is The following total probability formula is used to calculate the probability of all queries in the query sequence in the Gaussian mixture model:

[0072]

[0073] Among them, p(q i ) represents the probability of the i-th query in the Gaussian mixture model, q i represents the i-th query in the query sequence; π jr =p(z r =1|t j =1) is represented by the rth Gaussian distribution of the jth key, z r represents the rth Gaussian distribution in the Gaussian mixture model, t j represents a T-dimensional binary random variable, indicating the position of the key; M represents the number of Gaussian distributions; μ jr represents the mean of the rth Gaussian distribution of the jth key; represents the covariance matrix of the r-th Gaussian distribution of the j-th key; I represents the identity matrix.

[0074] In order to increase the intrinsic representation of each attention head, the embodiment of the present invention needs to increase the value of each key k j To improve the interpretability of the key and reduce the chance of learning redundant heads, each key k at position j can be j Modeled as a Gaussian mixture model containing M Gaussian distributions Where r = 0, 1, 2, ... M, and z can be defined as an M-dimensional binary random variable. r represents the rth Gaussian distribution in the mixed distribution, and defines π jr =p(z r =1|t j =1) is the rth Gaussian distribution of the jth key, and then the total probability formula can be used to calculate the probability of all queries in the query sequence in the Gaussian mixture model:

[0075]

[0076] The present invention can improve the interpretation ability of each key and reduce the chance of learning redundant heads by modeling each key as a Gaussian mixture model of M Gaussian distributions.

[0077] Step S304, based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture models, calculate the attention scores of the current key and all queries in the key sequence, traverse all keys in the key sequence, and obtain the attention scores of all keys and all queries.

[0078] Specifically, the following Bayesian formula is used to calculate the posterior probability of the current key and all queries, where the posterior probability represents the attention score of the current key and all queries in the key sequence:

[0079]

[0080] Among them, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; k jr represents the r-th Gaussian distribution of the j-th key in the key sequence.

[0081] The embodiment of the present invention takes into account p(t j =1|q i ) can be approximated by a finite number of Gaussian mixture distributions. Based on the above setting analysis, the formula (2) for p(t j =1|q i ) can be calculated with richer approximations. Similar to the above deduction, the posterior probability in the Gaussian mixture model can be written as formula (7).

[0082] In an optional implementation, the attention scores of the current key and all queries in the key sequence are calculated using the following Gaussian kernel function between the current key and all queries:

[0083]

[0084] The embodiment of the present invention takes into account that in the Gaussian mixture model, assuming that the query q i and key k jr The value of is the normalized result, so when actually calculating the attention score, it can be calculated by querying q i and key k jr The Gaussian kernel function between them is used to measure the distance between vectors instead of their dot product results. The attention scores of the current key and all queries can be calculated by formula (8).

[0085] The present invention calculates the attention scores of the current key and all queries by utilizing the Gaussian kernel function between the current key and all queries, which can meet practical applications and further improve the accuracy of target detection.

[0086] Furthermore, the Gaussian distribution with the largest probability density function is selected from the Gaussian mixture model corresponding to the current key as the Gaussian distribution for calculating the attention score; the attention scores of the current key and all queries in the key sequence are calculated using the following Gaussian kernel function between the current key and all queries:

[0087]

[0088] in, Represents the Gaussian distribution with the maximum probability density function.

[0089] In the embodiment of the present invention, the assumption of the Gaussian mixture model is taken into account. Each query sequence needs to perform Gaussian kernel distance calculation with the keys of multiple Gaussian mixture distributions, which has a high computational complexity. Therefore, a simplified process can be performed, and the query is only approximately compared with one of the most critical Gaussian distributions, which can greatly reduce the number of parameters and the amount of calculation. The more M is set, the more the amount of calculation is reduced, and the attention score can be effectively calculated. At the same time, the prior distribution π jr It no longer plays a role in the algorithm. Therefore, the Gaussian distribution with the largest probability density function can be selected from the Gaussian mixture model corresponding to the current key to be used as the Gaussian distribution for calculating the attention score. Then, the attention scores of the current key and all queries in the key sequence can be calculated using formula (9).

[0090] Step S305, dot product the attention scores and value sequences of all keys and all queries to obtain a target feature sequence, and determine the detection result of the target in the driving scene image based on the target feature sequence.

[0091] In the embodiment of the present invention, the attention score calculated according to formula (9) is dot-producted with the value sequence to obtain the target feature sequence:

[0092]

[0093] Specifically, the target feature sequence is a feature sequence processed by a single attention head in the self-attention network, and the above step S305 includes:

[0094] Step S3051, traverse all attention heads in the self-attention network to obtain the target feature sequences corresponding to all attention heads.

[0095] In step S3052, the target feature sequences corresponding to all attention heads are transformed through the total transformation matrix, and then residually connected with the first feature sequence to obtain an output sequence.

[0096] Step S3053: Determine the detection result of the target in the driving scene image based on the output sequence.

[0097] It can be seen from formula (3) that the key k jr With query q i The similarity between them can be explained by the posterior probability. By this simple connection, each query q i is considered as a mixed sample from T bonds, i.e. Considering that the distribution of each key may be asymmetric, skewed, or even multimodal, using a Gaussian distribution for each key may limit the explanatory power and diversity of each key. Therefore, the use of multi-head attention can be designed to make up for the above shortcomings.

[0098] like Figure 2 As shown, the self-attention network may include S attention heads. After the target feature sequence is calculated in each attention head according to the above formula (10), the target feature sequence processed by all attention heads can be transformed into a total transformation matrix, and then residually connected with the first feature sequence to obtain the attention mechanism result. Finally, the initially processed output sequence can be obtained through layer normalization, feedforward neural network, layer normalization and residual connection in the standard encoder module, and then the detection result of the target in the driving scene image is determined based on the output sequence.

[0099] In a specific embodiment, Figure 4 As shown, image features are extracted from the driving scene images obtained during the vehicle driving process to obtain a first feature sequence X. Among them, n x represents the length of the feature sequence, d x represents the dimension of the first feature sequence; W O Represents the total transformation projection matrix; W Q W K W V Represents the linear projection matrix respectively; can be obtained from W Q W K W V Extract a single attention head W from i Q W i K W i V ,in, d k is the dimension of the query and key vectors, d v represents the dimension of the value vector; h represents the number of attention heads, d q =d k =d v =d m / h, where d m =d x , the feature sequence X can be obtained by W i Q W i K W i V After obtaining Q, K, and V of each head, the posterior distribution p(t j =1|q i ), and then use formula (10) to get h for each attention head. Finally, we can use X'=MultiHead(Q,K,V)=Concat(h1L h s )W O Get the final output sequence Please refer to the above embodiments for detailed description, which will not be repeated here.

[0100] The embodiments of the present invention perform two types of visualization analysis to verify the effectiveness of the present invention: (1) using the attention distribution diagram to verify the present invention's ability to capture information that is strongly correlated with the attention heads; and (2) using the attention diagrams of different attention heads to verify the present invention's effect in mitigating the redundancy of the attention heads.

[0101] (1) The data distribution after the two multi-head attention mechanisms were purified and analyzed was visualized, and the attention distribution after the same data input was compared; the attention distribution diagrams of four pictures after being processed by the two attention mechanisms were given. Figure 5a For the original picture, Figure 5b is the attention distribution after the original method performs feature processing on the query at this position, Figure 5c The attention distribution map after being processed by the method proposed in this study. The red box indicates the position of a query position encoding corresponding to the position in the original image.

[0102] Figure 5a A simulated high-speed driving scene was selected, and the query was located at the position of the vehicle sunroof. The original method's attention distribution was very scattered, and the focus was on the door position, which was not strongly related to the original window features. This method can accurately focus on the area of ​​​​the query, which is a more ideal processing result; the second picture selected a busy traffic intersection, and the query was on a black vehicle. It can be seen that the original method's attention distribution is still very scattered. Although most of the attention points can be placed on the vehicle, there are still many attention points scattered to unrelated areas, while this method can still accurately focus on the area near the query; Figure 5b The same picture as the second one is selected, but the query area is different. The query is placed on the delivery man. It can be seen that the original method successfully places the focus area near the delivery man, but there is still a small area on the zebra crossing. This method not only successfully focuses on the area, but also has a larger distribution and can better aggregate similar information around it. Figure 5c The robustness of the driving scene verification method in foggy weather was selected. The query was on the black car on the left. The original method could focus on the car, but the most focused area was mistakenly placed on another black car. The attention distribution of this method accurately placed the focus area on the car, and was not limited to the query position, and was able to detect similar areas around it.

[0103] According to the visual analysis and comparison of the above three pictures, it can be seen that this method can accurately purify the key attention areas and query nearby areas with high similarity, which enhances the ability of data aggregation, thereby making the final output results of the Transformer module more ideal and improving all aspects of the model's performance.

[0104] (2) The two methods are compared under the same setting conditions, and the results of each attention head of each method are visualized to compare and analyze the redundancy; Figure 6a As shown in the figure, it represents the analysis results of the six attention heads of the original algorithm, and also includes the final average result, such as Figure 6b The following figure shows the analysis results of the 6 attention heads of the AMG algorithm and the final average result. Note that this experiment uses the same layer for visualization extraction. As can be seen from the figure, the results of the analysis of the 2nd, 3rd, 4th and 5th attention heads of the original algorithm are roughly the same, with only slight differences, indicating that these heads all focus on similar attention areas, which makes the calculation of the 6 attention heads too redundant and does not meet the original purpose of using multiple attention heads to enable the model to learn from multiple aspects. The results obtained by analyzing different attention heads of this algorithm are more different, indicating that the query can pay attention to multiple keys and is equivalent to other tags at different positions in the input sequence. This diversity of attention patterns helps to significantly reduce the possibility of the model learning similar and redundant attention matrices in different heads.

[0105] In the present embodiment, a target detection device based on a Gaussian mixture model is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0106] This embodiment provides a target detection device based on a Gaussian mixture model. Figure 7As shown, it includes: a feature sequence acquisition module 701, which is used to acquire a driving scene image of a vehicle during driving, and extract image features of the driving scene image to obtain a first feature sequence; a Gaussian modeling module 702, which is used to extract the first feature sequence through a multi-layer linear projection layer to obtain a query sequence, a key sequence and a value sequence corresponding to the first feature sequence, and model each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions; a query probability calculation module 703, which is used to calculate the probability of all queries in the query sequence in the Gaussian mixture model based on the Gaussian mixture model corresponding to all keys in the key sequence; an attention score calculation module 704, which is used to calculate the attention score of the current key in the key sequence and all queries based on the Gaussian mixture model corresponding to all keys and the probability of all queries in the Gaussian mixture model, traverse all keys in the key sequence, and obtain the attention scores of all keys and all queries; a target detection module 705, which is used to perform dot product multiplication of the attention scores of all keys and all queries with the value sequence to obtain a target feature sequence, and determine the detection result of the target in the driving scene image based on the target feature sequence.

[0107] In some optional embodiments, the target feature sequence is a feature sequence processed by a single attention head in the self-attention network, and the target detection module 705 also includes: an attention head traversal unit, which is used to traverse all attention heads in the self-attention network to obtain target feature sequences corresponding to all attention heads; a sequence conversion unit, which is used to convert the target feature sequences corresponding to all attention heads through a total conversion matrix, and then perform a residual connection with the first feature sequence to obtain an output sequence; a target detection unit, which is used to determine the detection result of the target in the driving scene image based on the output sequence.

[0108] In some optional implementations, the Gaussian mixture model corresponding to the key is The query probability calculation module 703 includes: using the following total probability formula to calculate the probabilities of all queries in the query sequence in the Gaussian mixture model:

[0109]

[0110] Among them, p(q i ) represents the probability of the i-th query in the Gaussian mixture model, q i represents the i-th query in the query sequence; π jr =p(z r =1|t j =1) is represented by the rth Gaussian distribution of the jth key, z r represents the rth Gaussian distribution in the Gaussian mixture model, t j represents a T-dimensional binary random variable, indicating the position of the key; M represents the number of Gaussian distributions; μjr represents the mean of the rth Gaussian distribution of the jth key; represents the covariance matrix of the r-th Gaussian distribution of the j-th key; I represents the identity matrix.

[0111] In some optional implementations, the attention score calculation module 704 includes:

[0112] The following Bayesian formula is used to calculate the posterior probability of the current key and all queries. The posterior probability represents the attention score of the current key and all queries in the key sequence:

[0113]

[0114] Among them, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; k jr represents the r-th Gaussian distribution of the j-th key in the key sequence.

[0115] In some optional implementations, the attention score calculation module 704 includes:

[0116] The attention score of the current key and all queries in the key sequence is calculated using the Gaussian kernel function between the current key and all queries as follows:

[0117]

[0118] Among them, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; k jr represents the r-th Gaussian distribution of the j-th key in the key sequence.

[0119] In some optional implementations, the attention score calculation module 704 includes:

[0120] Filter out the Gaussian distribution with the largest probability density function from the Gaussian mixture model corresponding to the current key as the Gaussian distribution for calculating the attention score;

[0121] The attention score of the current key and all queries in the key sequence is calculated using the Gaussian kernel function between the current key and all queries as follows:

[0122]

[0123] in, Represents the Gaussian distribution with the largest probability density function, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; kjr represents the r-th Gaussian distribution of the j-th key in the key sequence.

[0124] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0125] The target detection device based on the Gaussian mixture model in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0126] The embodiment of the present invention also provides a vehicle having Figure 8 Controller shown.

[0127] See also Figure 8 , Figure 8 is a schematic diagram of the structure of a controller provided by an optional embodiment of the present invention, such as Figure 8 As shown, the controller includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the controller, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple controllers can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 8 A processor 10 is taken as an example.

[0128] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0129] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0130] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created according to the use of the controller, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the controller via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0131] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0132] The controller also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 8 The example of connecting through bus is taken in the following.

[0133] The input device 30 can receive input digital or character information and generate key signal input related to the user settings and function control of the controller, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0134] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0135] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0136] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A target detection method based on Gaussian mixture model, characterized in that: The method comprises: Acquire a driving scene image of the vehicle during driving, and extract image features from the driving scene image to obtain a first feature sequence; Extracting the first feature sequence through a multi-layer linear projection layer to obtain a query sequence, a key sequence, and a value sequence corresponding to the first feature sequence, and modeling each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions; Based on the Gaussian mixture models corresponding to all the keys in the key sequence, calculating the probabilities of all the queries in the query sequence in the Gaussian mixture models respectively; Based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture models, the attention scores of the current key and all queries in the key sequence are calculated, and all keys in the key sequence are traversed to obtain the attention scores of all keys and all queries; The attention scores of all the keys and all the queries are dot-producted with the value sequence to obtain a target feature sequence, and based on the target feature sequence, a detection result of the target in the driving scene image is determined.

2. The method according to claim 1, characterized in that The target feature sequence is a feature sequence processed by a single attention head in a self-attention network, and the determining of the detection result of the target in the driving scene image further includes: Traverse all the attention heads in the self-attention network to obtain the target feature sequences corresponding to all the attention heads; The target feature sequences corresponding to all the attention heads are transformed by a total transformation matrix, and then residually connected with the first feature sequence to obtain an output sequence; Based on the output sequence, a detection result of the target in the driving scene image is determined.

3. The method according to claim 1, characterized in that The Gaussian mixture model corresponding to the key is The calculating, based on the Gaussian mixture model corresponding to all the keys in the key sequence, the probabilities of all the queries in the query sequence in the Gaussian mixture model respectively comprises: The following total probability formula is used to calculate the probabilities of all queries in the query sequence in the Gaussian mixture model: Among them, p(q i ) represents the probability of the i-th query in the Gaussian mixture model, q i represents the i-th query in the query sequence; π jr =p(z r =1|t j =1) is represented by the rth Gaussian distribution of the jth key, z r represents the rth Gaussian distribution in the Gaussian mixture model, t j represents a T-dimensional binary random variable, indicating the position of the key; M represents the number of Gaussian distributions; μ jr represents the mean of the rth Gaussian distribution of the jth key; represents the covariance matrix of the r-th Gaussian distribution of the j-th key; I represents the identity matrix.

4. The method according to claim 3, characterized in that The calculating the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all the keys and the probabilities of all the queries in the Gaussian mixture models respectively includes: The following Bayesian formula is used to calculate the posterior probability of the current key and all queries, where the posterior probability represents the attention score of the current key and all queries in the key sequence: Among them, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; k jr represents the r-th Gaussian distribution of the j-th key in the key sequence.

5. The method according to claim 3, characterized in that: The calculating the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all the keys and the probabilities of all the queries in the Gaussian mixture models respectively includes: The attention score of the current key and all queries in the key sequence is calculated using the Gaussian kernel function between the current key and all queries as follows: Among them, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; k jr represents the r-th Gaussian distribution of the j-th key in the key sequence.

6. The method according to claim 3, characterized in that The calculating the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all the keys and the probabilities of all the queries in the Gaussian mixture models respectively includes: Filter out the Gaussian distribution with the largest probability density function from the Gaussian mixture model corresponding to the current key as the Gaussian distribution for calculating the attention score; The attention score of the current key and all queries in the key sequence is calculated using the Gaussian kernel function between the current key and all queries as follows: in, Represents the Gaussian distribution with the largest probability density function, p(t j =1|q i ) represents the attention score of the current key and all queries; T represents the dimension of the key sequence; k jr represents the r-th Gaussian distribution of the j-th key in the key sequence.

7. A target detection device based on a Gaussian mixture model, characterized in that: The device comprises: A feature sequence acquisition module is used to acquire a driving scene image of the vehicle during driving, and perform image feature extraction on the driving scene image to obtain a first feature sequence; A Gaussian modeling module, configured to extract the first feature sequence through a multi-layer linear projection layer to obtain a query sequence, a key sequence, and a value sequence corresponding to the first feature sequence, and model each key in the key sequence as a Gaussian mixture model containing a preset number of Gaussian distributions; A query probability calculation module, used to calculate the probabilities of all queries in the query sequence in the Gaussian mixture model based on the Gaussian mixture models corresponding to all keys in the key sequence; An attention score calculation module, used to calculate the attention scores of the current key and all queries in the key sequence based on the Gaussian mixture models corresponding to all keys and the probabilities of all queries in the Gaussian mixture models, traverse all keys in the key sequence, and obtain the attention scores of all keys and all queries; The target detection module is used to perform dot product multiplication of the attention scores of all the keys and all the queries with the value sequence to obtain a target feature sequence, and determine the detection result of the target in the driving scene image based on the target feature sequence.

8. A vehicle, characterized in that: The vehicle includes a controller, the controller includes a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the target detection method based on the Gaussian mixture model described in any one of claims 1 to 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the target detection method based on a Gaussian mixture model according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to enable a computer to execute the target detection method based on a Gaussian mixture model according to any one of claims 1 to 6.

Citation Information

Cited By

  • Method and system for calculating posterior attention under long sequence condition

    CN122472104A