Multi-modal Feature Fusion Image Classification Method and Its Application in Humanoid Robots
Through the multimodal feature fusion image classification method, combined with depth images and radar images, the Transformer network is used for feature fusion and multi-scale processing, which solves the problem that humanoid robots cannot effectively perceive objects of different scales in complex scenarios, and achieves better environmental adaptability and navigation capabilities.
Patent Information
- Application Number
- CN202410712100.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-06-04
AI Technical Summary
In the prior art, humanoid robots can only rely on a single modal information to perceive the environment, and cannot effectively perceive objects of different scales in the environment in complex scenarios.
The multimodal feature fusion image classification method is adopted, and depth images and radar images are collected, block processing and data alignment are performed, and feature fusion is performed using different convolutional neural networks, combined with the Transformer network based on cross attention, and further processing is performed through multi-scale Transformer to realize image classification.
This method can better perceive objects at different scales in complex environments, improve the environmental adaptability and navigation capabilities of humanoid robots, and enhance their perception capabilities in complex scenarios.
Smart Images

Figure CN118628802B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition and understanding, and particularly relates to a multi-modal feature fusion image classification method and its application in humanoid robots. Background Art
[0002] The environmental perception ability of a humanoid robot refers to the robot's ability to obtain information in the environment through sensors and process it to understand the surrounding situation and conditions. This ability is crucial for the robot to perform tasks, interact with humans, and execute complex actions in the real world.
[0003] In actual application scenarios, a humanoid robot needs to accurately understand the surrounding environment in order to perform effective path planning, obstacle avoidance, target tracking and other tasks. Image classification is a basic and key cognitive ability for humanoid robots, enabling the robot to have a high-level understanding of the surrounding environment based on visual information, and then adapt to various application scenarios, effectively interact with elements in the environment, execute specified tasks, and perform self-monitoring and maintenance. Through image classification, the robot can identify key elements in the environment, such as furniture, pedestrians, obstacles, road signs, doors, stairs, etc., and thus make corresponding decisions.
[0004] Current robot environmental perception technologies based on computer vision mostly rely on convolutional neural networks (CNNs) and recurrent neural networks (RNNs). For the environmental perception technology based on the CNN network, through convolutional layers and pooling layers, it can effectively capture local features of image visual information, such as edges, textures, etc., and parameter sharing and sparse connections result in a relatively small number of parameters, which helps to reduce overfitting. However, it is difficult to process long inputs and cannot handle environmental understanding tasks in complex environments. The environmental perception technology based on the RNN network is suitable for processing sequential data and can capture temporal dependencies. However, traditional RNNs have problems of gradient vanishing or gradient explosion, making it difficult to model long-term dependencies, resulting in difficulty in capturing long-distance context information, and low computational efficiency. Due to the recurrent structure, it is difficult to parallelize, which limits the training speed. In addition to the methods based on the above two neural networks, in recent years, a batch of environmental perception technologies based on Transformers have emerged. The Transformer network introduces a self-attention mechanism, which can better capture the relationships between different positions, is suitable for sequential tasks and local and global relationships in images, has strong parallelization ability and high computational efficiency, which is beneficial to the training speed. However, the methods adopted have a single scale, and in complex scenarios, it is impossible to analyze the category features of objects of multiple scales separately. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention provides a multi-modal feature fusion image classification method and its application in humanoid robots. The purpose of the present invention is to solve the problems in the existing technology that robots can only rely on single-modal information to perceive the environment and cannot effectively perceive objects of different scales in complex scenarios.
[0006] The first aspect of the present invention is to provide a multi-modal feature fusion image classification method, including the following steps:
[0007] Step S101: Collect depth images and radar images, and perform block processing and data alignment operations on the obtained depth images and radar images;
[0008] Step S102: For the depth image blocks and radar image blocks after block processing, use different convolutional neural networks for preliminary feature extraction respectively, capture the spatial depth feature information of the depth image blocks and the abstract feature information of the radar image blocks, and align the data dimensions of the preliminary features of the depth image blocks and the radar image blocks;
[0009] Step S103: Fuse the preliminarily extracted depth image features and radar image features through a Transformer network based on cross-attention to achieve adaptive weight allocation and interaction of information between different modalities, and capture cross-modal semantic relationships;
[0010] Step S104: Process the fused features using a multi-scale Transformer to further capture abstract and refined feature information;
[0011] Step S105: Input the feature information after the above multi-scale feature fusion processing into the classification task head to complete image classification.
[0012] Further, in step S101 of the above multi-modal feature fusion image classification method, block processing is performed on the obtained depth images and radar images. According to the processing capabilities of the edge computing device, the images are segmented into image blocks of a preset size, and boundary padding processing is performed on the segmented image blocks so that the widths and heights of the filled depth image blocks and radar image blocks are the same.
[0013] Furthermore, in step S102 of the above multi-modal feature fusion image classification method, a feature extraction module combining three-dimensional convolution (Conv3D) and heterogeneous kernel convolution (HetConv2D) is used to process the depth image patches. The depth image patches are input into a three-dimensional convolutional layer (Conv3D) for convolutional processing, and on the basis of extracting spatial features, the height and width of the output image features are kept the same as those of the input image. Subsequently, the output of the three-dimensional convolutional layer is further processed by HetConv2D. The HetConv2D module uses parallel groupwise convolution and pointwise convolution, and integrates features by adding the outputs of the two convolutional layers to enhance the abstract representation of the depth image patches. In this process, batch normalization and ReLU activation functions are applied to optimize feature extraction and keep the feature dimensions.
[0014] Furthermore, in step S102 of the above multi-modal feature fusion image classification method, a feature extraction module including two-dimensional convolution (Conv2D), batch normalization (Batch Norm) and ReLU activation layer is used to extract features and align dimensions for the radar image patches. The two-dimensional convolutional layer is used to perform preliminary feature extraction on the radar image patches. The subsequent batch normalization layer is used to reduce the internal covariate shift and promote training stability. The subsequent ReLU activation function layer introduces non-linear activation for backpropagation in the output feature map.
[0015] Furthermore, in step S103 of the above multi-modal feature fusion image classification method, the cross-attention network adopts the Multi-Head Attention structure in Transformer, including two Cross Attention networks. The Query of the first Cross Attention network is generated from the abstract feature tensor of the radar image, while the Key and Value are generated from the abstract feature tensor of the depth image. The Query of the second Cross Attention network is generated from the abstract feature tensor of the depth image, and the Key and Value are generated from the abstract feature tensor of the radar image. Weights are assigned through the Softmax function to adaptively emphasize the correlation and importance of different modal data, and to achieve information interaction and semantic relationship capture between modalities.
[0016] Furthermore, step S104 of the above multi-modal feature fusion image classification method includes:
[0017] Step S1041: Use two-dimensional convolution on the fused features to generate Query / Key / Value (Q / K / V) token information, and retain the local spatial context structure.
[0018] Step S1042: By introducing a multi-scale multi-head self-attention mechanism (MSMHSA), features are extracted and fused at different scales to meet the recognition requirements of objects of different sizes.
[0019] Further, in step S1041 of the above multi-modal feature fusion image classification method, two-dimensional convolutional projection is used on the fused features to generate Query / Key / Value, and three independent two-dimensional convolutions W Q , W K , W V are used to generate token information with the same dimension as the original features, providing a basic representation for the subsequent multi-scale self-attention mechanism.
[0020] Further, in step S1042 of the above multi-modal feature fusion image classification method, a multi-scale multi-head self-attention mechanism is used to construct a pyramid structure in the self-attention module. The feature map is dynamically segmented according to a preset scale list to form feature representations of multiple scales. Each scale corresponds to a computing head, and each computing head independently performs self-attention calculation. Weighted fusion is performed on features of different scales. Within each computing head, the feature tensor is segmented by scale and then attention calculation is performed, and then the results of all computing heads are integrated through a concatenation operation to form a multi-scale feature representation.
[0021] The second aspect of the present invention proposes an application of the multi-modal feature fusion image classification method in a humanoid robot, and the multi-modal feature fusion image classification method is the multi-modal feature fusion image classification method introduced above.
[0022] Further, in the application of the multi-modal feature fusion image classification method in a humanoid robot, the multi-modal feature fusion image classification method is integrated into the interaction system of the humanoid robot as a program. The humanoid robot uses the sensor devices equipped on itself to collect depth images and radar images of the surrounding environment, identifies and classifies the types of objects in the surrounding environment to identify the passable areas, and plans a walking path within the passable areas.
[0023] Beneficial effects
[0024] The multi-modal feature fusion image classification method provided by the present invention can be applied to the usage scenarios of humanoid robots. In this method, a multi-scale Transfomer network is combined with a CNN neural network, which can capture information at different scales, including local details and global context. This enables it to integrate the advantages of CNN in capturing local image features and Transformer in capturing global relationships; it can generate multi-scale representations, thereby better modeling multi-scale information in images or sequences. This has advantages in dealing with objects of different sizes or events at different time scales; the multi-scale Transformer can more flexibly adapt to different scale requirements in different tasks and scenarios. In the image classification task, it can better handle different object sizes and image details; it performs excellently in capturing long-range dependencies and can better understand complex environments; similar to the standard Transformer, the multi-scale Transformer also has strong parallelization capabilities, which can accelerate model operations during training and inference; it can also adapt to multi-modal inputs, fully learn the complementary features between different modal data, and enhance the perception ability of humanoid robots in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a flowchart of the multi-modal feature fusion image classification method of the present invention.
[0026] Figure 2 is Figure 1 a flowchart of the multi-head self-attention mechanism part in
[0027] Figure 3 a processing flowchart for extracting depth map information.
[0028] Figure 4 a processing flowchart for extracting radar map information.
[0029] Figure 5 is a schematic diagram of the feature fusion module based on the cross-attention network.
[0030] Figure 6 is a schematic diagram of the structure of the cross-attention network.
[0031] Figure 7 is a schematic diagram of the convolutional projection process. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The present invention will be further clarified below through specific embodiments. These embodiments are exemplary and are intended to illustrate the problem and explain the present invention, rather than a limitation.
[0033] Embodiment 1
[0034] A multi-modal feature fusion image classification method, which can be applied to visual recognition in scenarios such as humanoid robots, and its process is as follows Figure 1 and Figure 2 shown, including steps S101 to S105.
[0035] S101: Perform tiling and padding processing on the sampled depth image and radar image.
[0036] According to the computing power of the edge computing devices of different humanoid robots, the collected images can be segmented into small image blocks of shapes such as 11×11 pixels, 16×16 pixels, 64×64 pixels, etc. Hardware parameters such as the processor performance, memory capacity, and storage bandwidth of the edge computing devices will affect their ability to process image data. The collected images are segmented into small image blocks of allowed sizes according to the device processing performance.
[0037] After that, the boundary filling method is used to fill the small image blocks to make the images easier to process. The specific filling method is shown in formula (1).
[0038]
[0039] Where: N is the number of samples, C is the number of channels, H in is the height of the input image, W in is the width of the input image, H out is the height of the output image, W out is the width of the output image.
[0040] S102: Use different CNN networks for preliminary feature extraction for different modalities of data.
[0041] The ability of automatic context modeling between features helps CNN make strong inferences and has good performance. For depth images, the depth information they carry enables CNN to make full use of its advantages in extracting hierarchical features and controlling the depth of output features. CNN can extract high-level abstract features. In the model proposed by the present invention, CNN is used to extract preliminary abstract features used as the input of Transformer.
[0042] Such as Figure 3As shown in the figure, for the extraction of depth features and spatial feature information of depth image blocks, the present invention designs a feature extraction module that combines three-dimensional convolution (Conv3D) and heterogeneous kernel-based convolutions (HetConv2D). This module can not only make full use of the context information of HIS for feature induction, but also make the depth control of the depth image feature map more convenient.
[0043] The depth image block X after edge padding H With a size of (16×16×C, where C is the depth information of the original depth image) is first decompressed into a shape of (1×16×16×C). Through a three-dimensional convolutional layer Conv3D with a convolutional kernel size of (3×3×9) and edge padding of (1×1×0), while extracting spatial features, the height and width of the output image features are kept equal to those of the input image. The relevant attributes of Conv3D are shown in Equation (2).
[0044] X Conv3D = Conv3D(X in , k=(3, 3, 9), p=(1, 1, 0)) Equation (2)
[0045] After that, the abstract features extracted by Conv3D are input into HetConv2D. HetConv2D parallelly uses two Conv2D layers, one of which performs groupwise convolution and the other performs pointwise convolution. The outputs of these two convolutions are integrated into the output of HetConv2D in the way of element-wise numerical addition. The relevant attributes and feature extraction process of HetConv2D are shown in Equation (3).
[0046]
[0047] The outputs of both the Conv3D layer and the HetConv2D layer pass through their respective batch normalization (Batch Norm, BN) and ReLU activation layers. The BN layer is used to solve the overfitting problem caused by a small number of training samples and accelerate the training performance. The ReLU layer helps activate backpropagation by introducing non-linearity in the output feature map. After passing through the three-dimensional convolution and heterogeneous kernel convolution modules, the shapes of the depth image blocks are (8×16×16×(C - 8)) and (16×16×64) respectively.
[0048] As Figure 4As shown in the figure, the feature extraction module of the radar image consists of a two-dimensional convolution, followed by a batch normalization (Batch Norm, BN) and a ReLU activation layer. In addition to performing preliminary abstract feature extraction, it also aligns the abstract feature dimensions of the depth image and the radar image, facilitating subsequent feature fusion. The shape of the output features is (16×16×64). The process is shown in Equation (4).
[0049] X out = Conv2D(X in , k=(3,3), p=(1,1)) Equation (4)
[0050] S103: Fuse the initially obtained feature information.
[0051] Use a Transformer network based on cross-attention for multi-modal data fusion: Using cross-attention can allocate weights according to the correlation between modalities, adaptively select the importance of each modality. Cross-attention allows effective interaction between modalities, which helps to capture cross-modal semantic relationships. However, using a cross-attention network may increase computational complexity and requires a large amount of data and a long training time to learn the relationships between modalities. The feature fusion module based on the cross-attention network is as Figure 5 shown.
[0052] The cross-attention network mainly adopts the Multi-Head Attention structure in Transformer. Its detailed structure is as Figure 6 shown. The feature fusion module contains two Cross Attention networks. Taking the part of extracting the abstract features of the depth image as an example, Q in the Attention mechanism is generated from the LiDAR abstract feature matrix, and K and V are generated from the depth image abstract feature tensor; the sources of Q, K, and V of the Cross Attention for extracting LiDAR abstract features are opposite to those of the part for extracting the abstract features of the depth image. The Cross Attention process is shown in Equation (5). The specific process of the feature fusion module for obtaining abstract features is shown in Equation (6), and the shape of the feature tensor does not change during the feature extraction process and is all (16×16×64).
[0053]
[0054] S104: Use a multi-scale Transformer for further feature extraction and fusion, and repeat this process E x times to obtain better inference results.
[0055] S1041: Generate the token information of Query / Key / Value (Q / K / V) using convolution. If linear projection is used in ViT to generate Q, K, and V, the dimensionality of the feature tensor needs to be collapsed into one-dimensional data first, which will destroy the spatial context structure of the original feature tensor. However, the Q, K, and V generated by the present invention based on the two-dimensional convolution projection method can retain the local spatial context information in the abstract features to the greatest extent. The convolution projection process is as Figure 7 shown. Moreover, the shape of the feature tensor generated by the convolution projection is the same as that of the input matrix, and it can be directly used for residual and scale division operations with the original features without deformation, facilitating the calculation of multi-scale attention.
[0056] Input the features extracted in S103 into three different two-dimensional convolutions W Q , W K , W V to generate their respective token information, preparing for further extraction of abstract feature information using the multi-scale multi-head self-attention mechanism. The generation process of the token information is shown in Equation (7).
[0057] Q, K, V = Conv2D(X in , k=(1, 1)) Equation (7)
[0058] The detailed parameters of the corresponding two-dimensional convolution are as follows:
[0059] X out = Conv2D(X in , k=(3, 3), p=(1, 1))
[0060] X = LeakyReLU(X otu , 0.2)
[0061] S1042: Use the multi-scale multi-head self-attention mechanism and introduce a pyramid structure in the self-attention module. This structure helps generate multi-scale feature maps, enabling the fusion of pixel-level features and complementary features. The proposed multi-scale multi-head self-attention mechanism is applied to different heads of each layer of the transformer. All heads follow a unified method to calculate self-attention. Its specific structure is as Figure 2 shown.
[0062] To achieve better classification performance on different datasets, a hyperparameter depth E x is added to the feature encoder, and the MSMHSA module is repeated different times on different datasets. The segmentation scales scale_list in MSMHSA are a set of hyperparameters. Here, the shape of the i-th scale is denoted as p i = H i × W i。There is a series of computing heads in the MSMHSA. The number of computing heads is determined by the number N of segmentation scales. Attention calculations for feature tensors of corresponding scales are performed in each computing head. The feature tensor generated from convolutional projection can be expressed as The feature map is in C A dimensions. According to the number N of segmentation scales, it is evenly divided into N equal parts. Then the input feature tensor in the i-th computing head is expressed as At this time, the feature map in the i-th computing head is segmented according to the segmentation scale p i into n i parts, and the values are shown in formula (8). The segmented feature tensor is denoted as q i , k i , v i , and its shape information is shown in formula (9).
[0063]
[0064] The output attention feature tensor in the i-th computing head is denoted as h i , and its shape information is shown in formula (10). Feature tensors h i from different computing heads are deformed and concatenated to form the feature output of the entire MSMHSA module whose definition is shown in formula (11).
[0065]
[0066] H = Concatenate(Reshape(h 1 ), Reshape(h 2 ),..., Reshape(h n ))
[0067] Formula (11)
[0068] The definition of the attention feature tensor h i is shown in formula (12):
[0069]
[0070] where C m = H A / H i , C n = W A / W i , q i,m represents the m-th part segmented from the input feature tensor Q i in the i-th computing head, k i,n , v i,n represent the feature tensors K i, V i The nth part segmented from
[0071] In the present invention, different from adding extra tokens in ViT for image classification, the present invention uses a sequence-based representation method to learn from all tokens in the input. The cross-entropy loss is used to obtain better performance in the classification task of humanoid robot environmental perception.
[0072] S105: Input the feature information obtained in the above process into the classification task head to obtain the classification result.
[0073] Embodiment 2
[0074] Based on the multi-modal feature fusion image classification method in Embodiment 1, an application of the multi-modal feature fusion image classification method in a humanoid robot is provided.
[0075] Specifically, integrate the multi-modal feature fusion image classification method into the interaction system of the humanoid robot in the form of a program. The humanoid robot uses the sensor devices equipped on itself to collect the depth image and radar image of the surrounding environment, identify and classify the types of objects in the surrounding environment to identify the passable area, and plan the walking path within the passable area. Thus, this application method not only improves the environmental adaptability and navigation ability of the humanoid robot, but also provides a more intelligent and stable basic condition for human-computer interaction, and has significant benefits in improving the autonomy of the robot.
[0076] Technical term explanations:
[0077] Padding: Padding;
[0078] Upsampling: Upsampling;
[0079] Fusion: Fusion (features);
[0080] Conv3D: 3D Convolution;
[0081] Conv2D: 2D Convolution;
[0082] Norm: Normalization layer;
[0083] Convolutional Projection: Convolutional Projection;
[0084] Concatenate: Concatenate;
[0085] Feed forward: Feed forward neural network;
[0086] MLP: Multi-Layer Perceptron;
[0087] Multi-Scale MHSA: Multi-Scale multi-head self attention, which is multi-scale multi-head self-attention.
[0088] The above embodiments are exemplary, aiming to illustrate the technical concept and features of the present invention, so that those skilled in this field can understand the content of the present invention and implement it accordingly. However, it should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A multimodal feature fusion image classification method, characterized by: The following steps are involved: Step S101: collecting a depth image and a radar image, and performing block processing and data alignment operations on the acquired depth image and radar image; Step S102: for the depth image block and the radar image block after the block processing, different convolutional neural networks are used to perform preliminary feature extraction respectively, to capture the spatial depth feature information of the depth image block and the abstract feature information of the radar image block, and to align the data dimensions of the preliminary features of the depth image block and the radar image block; Step S103: The initially extracted depth image features and radar image features are fused through a Transformer network based on cross attention to achieve adaptive weight allocation and effective interaction of information between different modalities and capture cross-modal semantic relationships; Step S104: applying a multi-scale Transformer to process the fused features to further capture abstract and refined feature information; Step S105: inputting the feature information processed by multi-scale feature fusion into the classification task head to complete image classification; In step S103, the cross attention network adopts the Multi-Head Attention structure in Transformer, including two Cross Attention networks; the query of the first Cross Attention network is generated by the abstract feature tensor of the radar image, and the key and value are generated by the abstract feature tensor of the depth image; the query of the second Cross Attention network is generated by the abstract feature tensor of the depth image, and the key and value are generated by the abstract feature tensor of the radar image; the weights are assigned by the Softmax function, and the relevance and importance of different modal data are adaptively emphasized to achieve information interaction and semantic relationship capture between modalities; Step S104 includes: Step S1041: Generate Query / Key / Value tag information using two-dimensional convolution on the fused features, retaining the local spatial context structure; Step S1042: By introducing a multi-scale multi-head self-attention mechanism, features are extracted and fused at different scales to meet the recognition requirements of objects of different sizes; In step S1041, two-dimensional convolution projection is used to generate Query / Key / Value for the fused features, using three independent two-dimensional convolutions W Q , W K , W V Generates label information with the same dimension as the original feature, providing a basic representation for the subsequent multi-scale self-attention mechanism; In step S1042, a multi-scale multi-head self-attention mechanism is used to construct a pyramid structure in the self-attention module, and the feature map is dynamically segmented according to a preset scale list to form feature representations of multiple scales. Each scale corresponds to a computing head, and each computing head independently performs self-attention calculations, and performs weighted fusion on features of different scales. In each computing head, the feature tensor is segmented by scale and then attention calculations are performed. Then, the results of all computing heads are integrated through cascade operations to form a multi-scale feature expression.
2. The multimodal feature fusion image classification method according to claim 1, characterized in that: In step S101, the acquired depth image and radar image are subjected to block processing. According to the processing capability of the edge computing device, the image is divided into image blocks of a preset size, and the divided image blocks are subjected to boundary filling processing so that the filled depth image blocks and radar image blocks have the same width and height.
3. The multimodal feature fusion image classification method according to claim 1, characterized in that: In step S102, the depth image block is processed by a feature extraction module that combines three-dimensional convolution and heterogeneous kernel convolution; the depth image block is input into the three-dimensional convolution layer for convolution processing, and the height and width of the output image features are kept consistent with the input image on the basis of extracting spatial features; the output of the three-dimensional convolution layer is then further processed by heterogeneous kernel convolution. The heterogeneous kernel convolution module uses parallel group convolution and point-by-point convolution to integrate features by adding the outputs of the two convolution layers to enhance the abstract representation of the depth image block, and batch normalization and ReLU activation function are applied in this process to optimize feature extraction and maintain feature dimension.
4. The multimodal feature fusion image classification method according to claim 1, characterized in that: In step S102, the radar image block is processed using a feature extraction module including two-dimensional convolution, batch normalization and ReLU activation layers for feature extraction and dimension alignment; a two-dimensional convolution layer is used to perform preliminary feature extraction on the radar image block, and the subsequent batch normalization layer is used to reduce internal covariate shift and promote training stability, and the subsequent ReLU activation function layer introduces nonlinear activation back propagation in the output feature map.
5. An application of a multimodal feature fusion image classification method in a humanoid robot, characterized in that: The multimodal feature fusion image classification method is the multimodal feature fusion image classification method according to any one of claims 1 to 4.
6. The application of the multimodal feature fusion image classification method in a humanoid robot according to claim 5, characterized in that: The multimodal feature fusion image classification method is integrated into the interactive system of the humanoid robot through a program. The humanoid robot uses its own sensor devices to collect depth images and radar images of the surrounding environment, identifies the passable area by identifying and classifying the types of objects in the surrounding environment, and plans the walking path within the passable area.
Citation Information
Patent Citations
Image classification method
CN115222998A
Image recognition method and system based on cross-modal feature fusion
CN117036891A