A method for evaluating the risk level of a scenario

Through the Transformer Encoder-Classification algorithm assisted by object detection, image feature processing is optimized, and the problems of high and slow computing resources in the prior art are solved, and fast and accurate scenario hazard level evaluation and classification are achieved.

CN116109913BActive Publication Date: 2025-07-08NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211262767.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-15
Publication Date
2025-07-08
Estimated Expiration
2042-10-15

AI Technical Summary

Technical Problem

The existing scenario understanding technology has problems such as high computing resource requirements, slow processing speed and poor classification results in the evaluation of hazardous scenario level, especially the Transformer-based model does not perform well in complex scenario understanding tasks.

Method used

The Transformer Encoder-Classification scene classification algorithm assisted by object detection is used to identify the dangerous object areas in the image and perform masking operations through the Pyramid Pooling Transformer model. The Encoder and image classification detection head are designed in combination with the Transformer E-Cls model, and the multi-head self-attention mechanism and position coding are used to optimize image feature processing.

Benefits of technology

It realizes a fast and accurate scenario hazard level evaluation. The model can process 130 pictures per second, with an evaluation accuracy of 90%, and is universal and robust, and can be applied to other scenario classification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116109913B_ABST
    Figure CN116109913B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for evaluating the risk level of a scene, including: preprocessing of the target detection image: before the image is input into the scene classification model, the Pyramid Pooling Transformer target detection model is used to identify the areas related to people and dangerous items in the image; after masking other irrelevant areas according to the target detection recognition result, the image matrix is input into the scene classification model; processing using the Transformer E-Cls scene classification model: including two parts, namely, encoder Encoder processing and image classification detection head processing. The Transformer E-Cls model of the present invention has a higher recognition accuracy in the task of evaluating the risk level of dangerous scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to scene understanding technology, and particularly to a method for evaluating the danger level of a scene. Background Art

[0002] At present, it has very important application value to realize the automatic judgment of the danger level of a picture scene through related scene understanding technologies. When dangerous events such as armed injury and terrorist attacks occur in some crowded public places such as railway stations, waiting halls, and city squares, if relevant security personnel cannot discover dangerous situations in time and formulate corresponding measures, irreparable consequences will be caused. Therefore, using technologies such as image scene understanding to realize the automatic and rapid judgment of dangerous scenes and provide relevant information can effectively improve security efficiency and minimize personnel and property losses.

[0003] Scene understanding is a relatively complex computer vision problem. The concept of scene understanding mainly includes: the overall scene semantic understanding, identifying each object category, and annotating multiple attribute values for the picture. That is, scene understanding mainly includes the following several tasks: multiple complex problems such as object recognition and detection, image semantic segmentation, and scene classification. Classifying the above several tasks, the image scene understanding task can be further divided into two tasks based on global features and based on local objects. The task based on global features focuses on the whole image, mainly for the annotation and classification of the scene. By analyzing the whole image, the category to which the image scene belongs is identified. The task based on local objects focuses on local targets in the image. By analyzing specific objects in the image, the category and location information of the targets in the image are determined, and finally the scene category is determined. The local scene understanding task mainly includes semantic segmentation and target detection.

[0004] Regarding related problems of scene understanding, the research contents at home and abroad mainly focus on finding a high-performance model to solve prediction tasks such as image classification, target detection, and semantic segmentation under the problem of scene understanding. At present, in the related field of scene understanding, the mainstream algorithm models can be divided into two categories: convolutional neural network models and network models based on Transformer.

[0005] Many studies have been conducted on tasks such as image scene classification, target detection, and semantic segmentation at home and abroad.

[0006] Inspired by the success of scaling in Transformers in NLP, Vision Transformer (abbreviated as ViT) basically realizes the direct application of standard Transformers to image processing tasks with minimal modifications. Vision Transformer divides an image into patches and takes the sequence of linear embeddings of these patches as the input of the Transformer. The image patches are processed in the same way as tokens (words) in NLP applications.

[0007] The Pyramid Vision Transformer (abbreviated as PVT) model is mainly based on Vision Transformer and has been optimized and modified. PVT breaks the conventional model of Transformer by introducing a progressive shrinking pyramid. It can generate multi-scale feature maps like traditional convolutional network models, thus solving the difficulty of applying Transformer to high-resolution images. The main modifications of the PVT model to the Vision Transformer model include the following three aspects: (1) Using fine-grained image patches (i.e., 4×4 pixels per patch) as input to learn high-resolution representations, which is crucial for dense prediction tasks; (2) As the network deepens, reducing the sequence length of the Transformer, significantly reducing the computational cost; (3) Adopting a spatial reduction attention layer (SRA) when learning high-resolution features, further reducing resource consumption.

[0008] The authors of the Transformer based on Pyramid Pooling (abbreviated as P2T) model believe that there are mainly the following two problems in Vision Transformer and some Transformer models based on Vision Transformer:

[0009] (1) The calculation of the multi-head self-attention mechanism has a high computational / space complexity;

[0010] (2) Most Vision Transformer models are overly optimized for image classification, while ignoring the differences between image classification (simple scenes, more similar to natural semantic processing tasks) and downstream scene understanding tasks (complex scenes, with rich structural and contextual features).

[0011] To solve the above two problems, P2T refers to the pyramid pooling algorithm that has been widely used in convolutional neural networks and applies it to the multi-head attention mechanism, thereby constructing a Transformer network for downstream tasks such as object detection and semantic segmentation. Many experiments show that when P2T is used as the backbone network, compared with previous CNN- and Vision Transformer-based networks, it has significant advantages in various downstream scene understanding tasks such as semantic segmentation, object detection, instance segmentation, and visual saliency detection.

[0012] The assessment of dangerous scene levels can mainly be transformed into a scene classification problem. Scene classification, as an image classification problem, is different from ordinary image classification in that: the objects of image classification (such as cats, dogs, or people) are basically located in the middle of the image; while scene classification is mainly classified according to the relationships of some items in the scene.

[0013] The existing scene classification technologies mainly include two categories: convolutional neural network models and Transformer-based network models. Convolutional neural network models mainly focus on local features of images, while scene classification requires judgments based on the global relationships of images. Therefore, convolutional neural network models do not perform well on scene classification datasets. On the other hand, Transformer-based network models can use the self-attention mechanism to capture global relationships, but most Transformer models have more parameters, require higher computing and memory resources, and are slower in processing images. Summary of the Invention

[0014] The main purpose of the present invention is to provide a method for assessing the dangerous level of a scene,

[0015] The technical solution adopted by the present invention is: a method for assessing the dangerous level of a scene, including:

[0016] Preprocessing of object detection images: Before the image is input into the scene classification model, the Pyramid Pooling Transformer object detection model is used to identify the regions related to people and dangerous items in the image; after masking other irrelevant regions according to the object detection results, the image matrix is input into the scene classification model;

[0017] Processing using the Transformer E-Cls scene classification model: It includes two parts: encoder Encoder processing and image classification detection head processing;

[0018] The encoder module is designed based on the Transformer model and is used to process image features to generate a feature map;

[0019] The image classification detection head consists of a simple linear transformation for making classification predictions based on the feature maps generated by the encoder module to obtain corresponding results;

[0020] Before inputting the image into the encoder, the Transformer E-Cls scene classification model adds two modules, patch embedding and position encoding, for processing.

[0021] Furthermore, the encoder Encoder is stacked by 3 or 6 identical layers; each layer of the encoder includes two sub-layers;

[0022] The first layer is the multi-head self-attention mechanism, and the second layer is a simple, fully position-connected feed-forward network; before the two sub-layers of the Transformer E-D encoder, a residual connection is made first, and then layer normalization is performed; the output of each sub-layer can be expressed as:

[0023] LayerNorm(x + Sublayer(x)) (1)

[0024] where Sublayer(x) represents the function implemented by the sub-layer itself, and the output dimensions of all sub-layers and the embedding layer in the model are the same.

[0025] Furthermore, the patch embedding includes: the Transformer E-Cls scene classification model divides the image into 16×16 small patches, and after flattening this small patch, it can be regarded as a word vector with a length of 768; each divided small patch is integrated as the input of the encoder Encoder.

[0026] Furthermore, the position encoding includes: after dividing the image into small patches, each small patch can correspond to a Token or word vector in the input of the Transformer E-Cls scene classification model, and the position of the small patch in the image is determined by two two-dimensional coordinates;

[0027] As shown in Equation (2), using sine and cosine functions with different frequencies, two position encodings with dimensions equal to half of the input dimension are generated, and the two position encodings are concatenated to generate an encoding for identifying the position information of the divided small patches in the picture;

[0028]

[0029] In Equation (2), pos represents the position and i represents the dimension; each dimension of the position encoding corresponds to a sine curve; its wavelength forms a geometric progression from 2π to 10000·2π.

[0030] Furthermore, the multi-head self-attention mechanism includes:

[0031] The multi-head self-attention mechanism is designed using the self-attention mechanism based on pyramid pooling, P-MHSA, in the Pyramid Pooling Transformer algorithm model. P-MHSA reshapes the input X into a two-dimensional space and applies multiple average pooling layers with different ratios on the reshaped X to generate pyramid feature maps:

[0032]

[0033] In Equation (3), {P1, P2,..., P n} represents the generated pyramid feature maps, and n is the number of pooling layers; P-MHSA inputs the pyramid feature maps into a depth convolution with relative position encoding:

[0034]

[0035] In Equation (4), DWConv(·) represents a depth convolution with a kernel size of 3×3 or 5×5, represents P with relative position encoding i ; P-MHSA flattens and concatenates these pyramid feature maps:

[0036]

[0037] Assume that the query, key, and value matrices in MHSA are Q, K, and V respectively. Compared with the traditional multi-head self-attention mechanism:

[0038]

[0039] P-MHSA uses:

[0040]

[0041] In Equation (7), W q 、W k and W v represent the weight matrices of the linear transformations used to generate the query, key, and value matrices respectively; Input Q, K, and V into the convolutional neural network attention module to calculate the attention matrix A, and its formula is as follows:

[0042]

[0043] In Equation (8), d K is the channel dimension of K, can be used as an approximate normalization; The Softmax function is applied along the rows of the matrix.

[0044] Furthermore, the feedforward network includes: two linear transformations, with a ReLU activation function connecting the two linear transformation functions in the middle;

[0045] .

[0046] Advantages of the present invention:

[0047] After being trained with relevant datasets, the present invention can achieve a rapid assessment of the danger level of scene images;

[0048] The model algorithm is modular, and the model effect can be improved by optimizing the multi-head self-attention mechanism and combining convolutional feature maps without redesigning the overall architecture of the model and the remaining modules;

[0049] After combining with the object detection model, it can simultaneously identify and label dangerous items in the scene, and at the same time improve the generalization ability of the scene classification model.

[0050] The number of Transformer encoder layers is 3 or 6 layers. Compared with other Transformer-based models, it can process more pictures per second.

[0051] Compared with other Transformer-based scene classification models, the Transformer E-Cls model proposed by the present invention has a higher recognition accuracy in the dangerous scene level assessment task.

[0052] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The present invention will be further described in detail below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0054] Figure 1 is the image input of the present invention;

[0055] Figure 2 is the overall flowchart of the algorithm of the present invention;

[0056] Figure 3 is the preprocessing result diagram of the object detection image of the present invention ((a) is the object detection result; (b) is the masking processing result);

[0057] Figure 4It is the overall architecture diagram of the Transformer E-Cls of the present invention. (In the figure, Image represents the input image; Patch Embedding represents the patch embedding module; Positional Encoding represents the positional encoding module; MHSA represents the multi-head self-attention mechanism; Add & Norm represents the residual connection and normalization function; Feed Forward represents the feed-forward neural network; Encoder represents the encoder model; Classification Head represents the linear classification head; Output is the final output of the model). Detailed implementation manners

[0058] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.

[0059] Aiming at the problems existing in the existing scene understanding technology and the characteristics of the dangerous scene level assessment task, a Transformer Encoder-Classification scene classification algorithm assisted by object detection is proposed.

[0060] It can evaluate the scene danger level in real time and accurately. When performing the scene danger level assessment task, the scene danger level assessment model proposed by the present invention can process 130 pictures per second, and the assessment accuracy is about 90%.

[0061] The model has universality and robustness. In addition to being applicable to the scene danger level assessment task, the model can be well applied to other scene classification tasks through training with other data sets.

[0062] Some algorithms of the model can be easily replaced and can be used in combination with convolutional networks.

[0063] The Transformer Encoder-Classification scene classification algorithm assisted by object detection is divided into two parts: object detection image preprocessing and Transformer Encoder-Classification model (abbreviated as Transformer E-Cls) scene classification.

[0064] The input of this algorithm is relevant scene images, such as Figure 1 shown. After object detection preprocessing and scene danger level assessment, the algorithm outputs the scene danger level. The scene danger level is divided into the following six levels according to the degree of danger:

[0065] (1) Level 5: Someone in the scene holds a firearm;

[0066] (2) Level 4: Someone in the scene is holding a knife;

[0067] (3) Level 3: Someone in the scene is holding a club;

[0068] (4) Level 2: Someone in the scene is holding an implement such as a racket;

[0069] (5) Level 1: There are dangerous items in the scene but no one is holding them;

[0070] (6) Level 0: There is no abnormal situation in the scene.

[0071] The overall process design of this algorithm is as Figure 2 shown.

[0072] (1) Preprocessing of the target detection image: Before the image is input into the scene classification model, the Pyramid Pooling Transformer target detection model is used to identify the regions related to people and dangerous items in the image. Subsequently, according to the target detection recognition results, other irrelevant regions are masked (here, the pixel values of the irrelevant regions are selected to be modified to 0), and then the image matrix is input into the scene classification model. The target detection recognition and masking results are as Figure 3 shown.

[0073] (2) Transformer Encoder-Classification scene classification model: The overall architecture of the Transformer E-Cls model is as Figure 4 shown. The Transformer E-Cls is mainly composed of two parts: an encoder and an image classification detection head. The encoder module is designed based on the Transformer model and is responsible for processing the image features to generate a feature map; the image classification detection head consists of a simple linear transformation and is responsible for classifying and predicting according to the feature map generated by the encoder module to obtain the corresponding results. Before inputting the image into the encoder, the Transformer E-Cls model adds two modules: patch embedding and positional encoding.

[0074] The following is a detailed introduction to the three parts of the encoder, patch embedding, and positional encoding in the Transformer E-Cls model:

[0075] 1. Encoder: Stacked by 3 or 6 identical layers. Each layer of the encoder has two sub-layers. The first layer is the multi-head self-attention mechanism, and the second layer is a simple, position-wise fully connected feed-forward network. Transformer E-D first performs a residual connection between the two sub-layers of the encoder and then layer normalization. That is, the output of each sub-layer can be expressed as:

[0076] LayerNorm(x + Sublayer(x)) (1)

[0077] where Sublayer(x) represents the function implemented by the sub-layer itself (for example, the multi-head self-attention mechanism or the feed-forward network will be introduced later). To facilitate the application of residual connections, the output dimensions of all sub-layers in the model and the embedding layer are the same.

[0078] 2. Patch Embedding: To enable the input of images into the Transformer encoder, referring to the relevant design in Vision Transformer, the Transformer E-D model divides the image into 16×16 small patches. After flattening this small patch, it can be regarded as a word vector with a length of 768. Integrating each divided small patch can be used as the input of the Transformer encoder.

[0079] 3. Position Encoding: After dividing the image into small patches, each small patch can correspond to a Token or word vector in the input of the Transformer model. The position of the small patch in the image is determined by two two-dimensional coordinates, which is quite different from the one-dimensional coordinate of the word position in the translation task. Transformer E-D designs a position encoding that can uniquely identify the two coordinates in the picture. This position encoding is achieved by referring to the position encoding scheme in the basic Transformer model and making simple modifications: using sine and cosine functions with different frequencies (as shown in Equation (2)), generating two position encodings with dimensions equal to half of the input dimension, and concatenating the two position encodings to generate the encoding for identifying the position information of the divided small patches in the picture.

[0080]

[0081] In Equation (2), pos represents the position and i represents the dimension. That is, each dimension of the position encoding corresponds to a sine curve. Its wavelength forms a geometric progression from 2π to 10000·2π. This function is chosen because it is assumed that it allows the model to easily learn relative positions, because for any fixed offset k, PEpos + k can be expressed as a linear function of PEpos.

[0082] Among them, the multi-head self-attention mechanism and the feed-forward network are designed as follows:

[0083] 1. Multi-head self-attention mechanism: The multi-head self-attention mechanism adopts the design of the self-attention mechanism based on pyramid pooling (abbreviated as P-MHSA) in the Pyramid Pooling Transformer algorithm model. First, P-MHSA reshapes the input X into a two-dimensional space. Then, multiple average pooling layers with different ratios are applied to the reshaped X to generate pyramid feature maps, such as:

[0084]

[0085] In Equation (3), {P1, P2,..., P n} represents the generated pyramid feature maps, and n is the number of pooling layers. Then, P-MHSA inputs the pyramid feature maps into the depth convolution of relative position encoding:

[0086]

[0087] In Equation (4), DWConv(·) represents the depth convolution with a kernel size of 3×3 (or 5×5), represents P with relative position encoding i . Since P i is the merged feature, the computational cost of such an operation in the above formula is very small. Next, P-MHSA flattens and concatenates these pyramid feature maps:

[0088]

[0089] For simplicity, the flattening operation is omitted here. In this way, if the pooling ratio is large enough, then P can be a sequence shorter than the input X. In addition, P contains the global context abstraction of the input, so it can be used as a strong substitute for the input when calculating MHSA.

[0090] Suppose the query, key, and value matrices in MHSA are Q, K, and V respectively. Compared with the traditional multi-head self-attention mechanism:

[0091]

[0092] P-MHSA tends to use:

[0093]

[0094] In Equation (7), W q 、W k and W vThey respectively represent the weight matrices of the linear transformations used to generate the query, key, and value matrices. Then, Q, K, and V are input into the convolutional neural network attention module to calculate the attention matrix A, and its formula is as follows:

[0095]

[0096] In Equation (8), d K is the channel dimension of K, which can be used as an approximate normalization. The Softmax function is applied along the rows of the matrix. The concept of multi-heads is omitted in the above formula for simplicity.

[0097] Since the lengths of K and V are less than that of X, the proposed P-MHSA is more efficient than the traditional MHSA. In addition, since K and V contain highly abstract multi-scale information, P-MHSA still has strong capabilities in global context dependence modeling, which helps in scene understanding.

[0098] 2. Feed-Forward Network: In the Transformer E-Cls model, the feed-forward network consists of two linear transformations, and the two linear transformation functions are connected by a ReLU activation function in the middle.

[0099]

[0100] In the Transformer E-Cls model of the present invention, the multi-head self-attention mechanism can be replaced by the traditional multi-head self-attention mechanism; in the Transformer E-Cls model, the input image matrix can be replaced by a convolutional feature map; the object detection model can be replaced by models such as PVT and Swin Transformer.

[0101] The present invention can evaluate the scene danger level in real time and accurately. When performing the scene danger level evaluation task, the scene danger level evaluation model proposed by the present invention can process 130 pictures per second, and the evaluation accuracy is about 90%;

[0102] The model has universality and robustness. In addition to being applicable to the scene danger level evaluation task, by training with other datasets, it can be well applied to other scene classification tasks.

[0103] Some algorithms of the model can be easily replaced and can be used in combination with convolutional networks.

[0104] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for evaluating the risk level of a scenario, characterized in that, Including: Target detection image preprocessing: Before the image is input into the scene classification model, the Pyramid Pooling Transformer target detection model is used to identify the regions related to people and dangerous items in the image; after masking other irrelevant regions according to the target detection results, the image matrix is input into the scene classification model; Processing using the Transformer E-Cls scene classification model: It includes two parts: encoder Encoder processing and image classification detection head processing; The encoder module is designed based on the Transformer model and is used to process image features to generate a feature map; The image classification detection head consists of a simple linear transformation and is used to perform classification prediction based on the feature map generated by the encoder module to obtain the corresponding results; Before the Transformer E-Cls scene classification model inputs the image into the encoder, two modules, patch embedding and position encoding, are added for processing; The encoder Encoder is stacked by 3 or 6 identical layers; each layer of the encoder includes two sub-layers; The first layer is the multi-head self-attention mechanism, and the second layer is a simple, fully connected feed-forward network in terms of position; before the two sub-layers of the Transformer E-D encoder, a residual connection is first made, and then layer normalization is performed; the output of each sub-layer can be expressed as: LayerNorm(x + Sublayer(x)) (1) where Sublayer(x) represents the function implemented by the sub-layer itself, and the output dimensions of all sub-layers and the embedding layer in the model are the same; The feed-forward network includes: consisting of two linear transformations, and a ReLU activation function is connected in the middle of the two linear transformation functions; FFN(x) = max(0, xW1 + b1)W2 + b2 (9).

2. The method for evaluating the risk level of a scenario according to claim 1, wherein, The patch embedding includes: The Transformer E-Cls scene classification model divides the image into 16×16 small blocks, and after flattening this small block, it can be regarded as a word vector with a length of 768; each divided small block is integrated and used as the input of the encoder Encoder.

3. The method for evaluating the risk level of a scenario according to claim 1, wherein The position encoding includes: After the image is divided into small blocks, each small block can correspond to a Token or word vector in the input of the Transformer E-Cls scene classification model, and the position of the small block in the image is determined by two two-dimensional coordinates; As shown in Equation (2), using sine and cosine functions with different frequencies, two position encodings with dimensions equal to half of the input dimension are generated, and the two position encodings are concatenated to generate an encoding for identifying the position information of the divided small blocks in the picture; , (2) In Equation (2), pos represents the position and i represents the dimension; each dimension of the position encoding corresponds to a sine curve; its wavelength forms a geometric progression from 2π to 10000·2π.

4. The method for evaluating the risk level of a scenario according to claim 2, wherein, The multi-head self-attention mechanism includes: The multi-head self-attention mechanism adopts the self-attention mechanism P-MHSA based on pyramid pooling in the Pyramid Pooling Transformer algorithm model. P-MHSA reshapes the input X into a two-dimensional space, and applies multiple average pooling layers with different ratios on the reshaped X to generate pyramid feature maps: P1 = AvgPool1(X), P2 = AvgPool2(X), · · · , P n = AvgPool n (X), (3) In formula (3), {P1, P2,..., P n} represents the generated pyramid feature map, and n is the number of pooling layers; P-MHSA inputs the pyramid feature map into the depth convolution of relative position encoding: DWConv(P i ) + P i , i = 1, 2, · · · , n (4) In Equation (4), DWConv(·) represents depthwise convolution with a kernel size of 3×3 or 5×5, represents P with relative position encoding i ; P-MHSA flattens and concatenates these pyramid feature maps: P = LayerNorm(Concat( , ,..., )) (5) Assume that the query, key, and value matrices in MHSA are Q, K, and V respectively. Compared with the traditional multi-head self-attention mechanism: (Q, K, V) = (XW q , XW k , XW v ) (6) P-MHSA uses: (Q, K, V) = (XW q , PW k , PW v ) (7) In formula (7), W q , W k and W v respectively represent the weight matrices of the linear transformations used to generate the query, key, and value matrices; Q, K, and V are input into the convolutional neural network attention module to calculate the attention matrix A, and its formula is as follows: A = Softmax((Q×K T ) / )×V (8) d in Equation (8) K is the channel dimension of K, which can be used as an approximate normalization; the Softmax function is applied along the rows of the matrix.

Citation Information

Patent Citations

  • Time sequence attention mechanism scene image recognition method

    CN113688822A

  • High-efficiency insulator detection system in complex space environment

    CN114399628A