A device based on an image recognition architecture
By using GhostNet network to replace convolutional calculation in the MTCNN model, the problem of high resource consumption of image recognition technology in edge devices is solved, and high-precision image recognition and parameter saving effect is achieved, which is suitable for monitoring equipment of space stations.
Patent Information
- Application Number
- CN202410200659.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-02-23
AI Technical Summary
When existing image recognition technology is deployed within edge devices, convolutional computing consumes high resources, resulting in large size of equipment and cannot be installed on a large scale in the space station.
The GhostNet network is used to replace the convolution calculation in the MTCNN model, and feature extraction is performed through the Bottleneck block of the GhostNet network, reducing the number of convolution kernels and saving parameters.
It improves the recognition accuracy of the image recognition architecture, reduces the required parameters, and is convenient for deployment on edge devices, and is suitable for monitoring equipment for space stations.
Smart Images

Figure CN117787349B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly relates to a device based on an image recognition architecture. Background Art
[0002] A space station, also known as a spacecraft, a space station, or an orbital station, is a manned spacecraft that can operate in a low-Earth orbit for a long time and can be visited by multiple astronauts for long-term work and life. Space stations are divided into two types: single-module and combined. A single-module space station can be launched into orbit at one time by a space vehicle, while a combined space station requires components to be sent into orbit in batches by a space vehicle and then assembled in space.
[0003] After the space station is sent into orbit, it needs to operate independently and for a long time in orbit. Due to the increasing space activities, there are more and more abandoned artificial objects around the low-Earth orbit. Coupled with the frequent visits of space meteorites, problems such as small cracks, holes, or firmware detachment are likely to occur on the space station, which will affect the safety of the space station and astronauts.
[0004] With the increasing level of space technology, the launch frequency of space stations is becoming more and more frequent. To improve the safety of space stations, monitoring devices capable of automatic monitoring, such as intelligent image recognition architectures, need to be equipped on space stations. However, ordinary convolutional operations consume a lot of resources when deployed inside edge devices, have high requirements for the computing power and energy consumption of hardware, increase the volume of monitoring devices, and the monitoring devices cannot be installed on a large scale inside space stations. Summary of the Invention
[0005] In view of this, the problem to be solved by the present invention is to provide a device based on an image recognition architecture. Compared with the original MTCNN model, only a small number of parameters are added, which can greatly improve the recognition accuracy of the modified MTCNN model, and the modified MTCNN model can operate normally inside edge devices.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is:
[0007] A device based on an image recognition architecture, including an MTCNN model, where the MTCNN model includes a P-Network, an R-Network, and an O-Network, and convolutional calculations inside the P-Network, R-Network, and O-Network are replaced by a GhostNet network to save calculation parameters.
[0008] Further, the GhostNet network includes a Bottleneck block with a stride of 1 and a Bottleneck block with a stride of 2;
[0009] The Bottleneck block with a stride of 1 is used to perform a small amount of convolution calculations to generate a first feature map. The first feature map is subjected to simple linear calculations to obtain a first new feature map, and the first feature map and the first new feature map are jointly output.
[0010] The Bottleneck block with a stride of 2 is used to perform a small amount of convolution calculations to generate a second feature map. The second feature map is subjected to simple linear calculations to obtain a second new feature map, and the second feature map and the second new feature map are compressed and jointly output.
[0011] Furthermore, the P-Network includes a first P convolutional layer, a first GhostNet network, and a second P convolutional layer. The first GhostNet network is composed of a stacked 1-stride Bottleneck block and a 2-stride Bottleneck block.
[0012] Furthermore, the R-Network includes an R convolutional layer, a second GhostNet network, and an R fully connected layer. The second GhostNet network is composed of a stacked 1-stride Bottleneck block and two 2-stride Bottleneck blocks.
[0013] Furthermore, the O-Network sequentially includes an O convolutional layer, a third GhostNet network, and an O fully connected layer. The third GhostNet network is composed of a stacked 1-stride Bottleneck block and three 2-stride Bottleneck blocks.
[0014] Furthermore, it includes an image acquisition module, and the image acquisition module communicates with the image recognition module.
[0015] Furthermore, the image recognition module includes an FPGA chip.
[0016] The advantages and positive effects of the present invention are:
[0017] By replacing the second convolutional layer and subsequent convolutional layers of the P-Network, R-Network, and O-Network with the GhostNet network, the number of convolutional kernels in the GhostNet network is less than that of a conventional convolutional layer, and the feature extraction ability is high, which can improve the recognition accuracy of the image recognition architecture. Compared with the recognition network under the same recognition accuracy, fewer parameters are required, facilitating the deployment of the image recognition architecture on edge devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0019] Figure 1 It is a framework diagram of a device based on an image recognition architecture of the present invention;
[0020] Figure 2 It is a framework flow chart of a device based on image recognition architecture of the present invention. DETAILED DESCRIPTION
[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0023] The present invention provides a device based on an image recognition architecture, such as Figure 1 As shown, the MTCNN model includes a P-Net network, an R-Net network and an O-Net network. Taking hole recognition as an example: the P-Net network receives and processes images to be recognized of different sizes to generate a large number of candidate frame images; based on the candidate frame images, the original sample image is cropped or scaled to a size of 24x24 pixels, and sent to the R-Net network for processing to screen out hole frame images containing holes; based on the hole frame images, the original image to be recognized is cropped or scaled to a size of 48x48 pixels, and sent to the O-Net network. The O-Net network recognizes and gives the final judgment result, which includes the location and number of holes.
[0024] like Figure 2 As shown in the figure, in order to save parameters in the MTCNN model, the second convolution layer and subsequent convolution layers of the P-Net network, R-Net network and O-Net network are replaced by the GhostNet network. The operation process of the GhostNet network is as follows: receiving the feature map output by the first layer of convolution, the feature map generates a partial feature map through a small number of convolution kernels, and the partial feature map generates a new feature map through a series of linear transformations; after repeating the above steps several times, the height and width of all features are compressed to complete the feature extraction.
[0025] When the number of convolutional kernels remains unchanged, more feature information can be obtained, improving the feature extraction performance of the MTCNN model and thus enhancing the recognition accuracy. When reducing the number of convolutional kernels, the same feature information can still be obtained, ensuring the recognition accuracy of the MTCNN model.
[0026] When the MTCNN model is used on edge devices, by reducing the number of convolutional kernels (required parameters) in the MTCNN model, it can operate normally at high speed. At the same time, due to the reduced storage space and storage area of the MTCNN model, compared with the picture recognition architecture, it has stronger anti-single particle interference ability and is more suitable for the aviation field.
[0027] The GhostNet network is composed of several Bottleneck blocks with different strides stacked together. The Bottleneck block includes a Bottleneck block with a stride of 1 and a Bottleneck block with a stride of 2. The Bottleneck block with a stride of 1 includes two Ghost blocks and a residual edge group. The Bottleneck block with a stride of 1 is used to perform a small amount of convolutional calculations to generate the first feature map. The first feature map is obtained through simple linear calculations to acquire the first new feature map, and the first feature map and the first new feature map are jointly output. The first Ghost block serves as an expansion layer, increasing the number of channels, and the second Ghost module reduces the number of channels.
[0028] The difference between the Bottleneck block with a stride of 2 and the Bottleneck block with a stride of 1 is that a [2×2] depthwise separable convolution is added between the two Ghost modules to complete the width and height compression operation. The Bottleneck block with a stride of 2 is used to perform a small amount of convolutional calculations to generate the second feature map. The second feature map is obtained through simple linear calculations to acquire the second new feature map, and the second feature map and the second new feature map are compressed and jointly output.
[0029] Taking the recognition of holes as an example:
[0030] The Net network includes a first P convolutional layer, a first GhostNet network, and a second P convolutional layer. The first GhostNet network is composed of a Bottleneck block with a stride of 1 and a Bottleneck block with a stride of 2 stacked together.
[0031] The processing process of the P-Network is as follows: The first P-layer convolution receives the sample image, and after convolution processing, outputs several first P-feature maps. The Bottleneck block with a stride of 1 receives the first P-feature maps, performs a small amount of convolution calculations, and then outputs several second P-feature maps through simple linear calculations. The Bottleneck block with a stride of 2 performs convolution calculations on the second P-feature maps and then simple linear calculations to generate several third P-feature maps. The length and width of the third P-feature maps are compressed to form the total feature map and output. The second P-convolution layer receives the total feature map and performs feature extraction to output several first candidate maps of different sizes.
[0032] The R-Network includes an R-convolution layer, a second GhostNet network, and an R-fully connected layer. The second GhostNet network is composed of the stacking of a Bottleneck block with a stride of 1 and two Bottleneck blocks with a stride of 2.
[0033] The processing process of the R-Network is as follows: The R-convolution layer receives the first candidate maps, and after convolution processing, outputs several second candidate maps. The Bottleneck block with a stride of 1 receives the second candidate maps, performs a small amount of convolution calculations, and then outputs several third candidate maps through simple linear calculations. The first Bottleneck block with a stride of 2 performs convolution calculations on the third candidate maps and then simple linear transformation to generate several fourth candidate maps. After compressing the length and width of the fourth candidate maps, they are output. The second Bottleneck block with a stride of 2 performs convolution calculations on the fourth candidate maps and then simple linear transformation to generate several fifth candidate maps. After compressing the length and width of the fifth candidate maps, they are output. The R-fully connected layer selects several first hole block diagrams from the fifth candidate maps and deletes the non-hole block diagrams.
[0034] The O-Network sequentially includes an O-convolution layer, a third GhostNet network, and an O-fully connected layer. The third GhostNet network is composed of the stacking of a Bottleneck block with a stride of 1 and three Bottleneck blocks with a stride of 2.
[0035] The processing process of the O-Network is as follows: The O convolutional layer receives the first hole block diagram, and after convolutional processing, outputs several second hole block diagrams. The Bottleneck block with a stride of 1 receives the second hole block diagram, performs a small amount of convolutional calculations, and then outputs several third hole block diagrams through simple linear calculations. The first Bottleneck with a stride of 2 performs convolutional calculations on the third hole block diagram, and then performs a simple linear transformation to generate several fourth hole block diagrams, compresses the length and width of the fourth hole block diagram and then outputs. The second Bottleneck with a stride of 2 performs convolutional calculations on the fourth hole block diagram, and then performs a simple linear transformation to generate several fifth hole block diagrams, compresses the length and width of the fifth hole block diagram and then outputs. The third Bottleneck with a stride of 2 performs convolutional calculations on the fifth hole block diagram, and then performs a simple linear transformation to generate several sixth hole block diagrams, compresses the length and width of the sixth hole block diagram and then outputs. The O fully connected layer receives all the sixth hole block diagrams to output the final hole judgment result.
[0036] It saves the calculation parameters of the image recognition architecture, improves the feature extraction efficiency at the same time, and then improves the operation efficiency and image recognition accuracy of the picture recognition architecture in the edge device.
[0037] A monitoring device based on a picture recognition architecture, including a picture acquisition module and a picture recognition module, and the picture acquisition module is directly interconnected with the picture recognition module. An embodiment of the application is that the picture recognition module includes an FPGA chip.
[0038] The picture acquisition module is used to continuously acquire pictures in the space station. The picture recognition module directly receives the pictures, judges whether there are problems such as cracks and holes on the pictures, and gives a timely reminder when a problem is detected. The picture recognition module is directly connected to the picture acquisition module, reducing the problem of multiple caches during the picture data transmission process, and reducing the impact of single-event upsets on the accuracy of the picture data during transfer.
[0039] The above has described the embodiments of the present invention in detail, but the above content is only the preferred embodiments of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equal changes and improvements made according to the scope of the present invention should still fall within the scope covered by this patent.
Claims
1. A device based on an image recognition architecture, characterized in that: Including an MTCNN model, wherein the MTCNN model includes a P-Net network, an R-Net network, and an O-Net network, wherein the second convolutional layer and subsequent convolutional layers in the P-Net network, the R-Net network, and the O-Net network are replaced by a GhostNet network to save parameters; The GhostNet network includes a 1-step Bottleneck block and a 2-step Bottleneck block; the 1-step Bottleneck block includes two Ghost blocks and a residual edge group; The Bottleneck block with a step size of 1 is used to perform convolution calculation to generate a first feature map, the first feature map is linearly calculated to obtain a first new feature map, and the first feature map and the first new feature map are jointly output; the first Ghost block is used as an expansion layer to increase the number of channels, and the second Ghost module reduces the number of channels; the difference between the Bottleneck block with a step size of 2 and the Bottleneck block with a step size of 1 is that a [2×2] depthwise separable convolution is added between the two Ghost modules to complete the width and height compression operation; The 2-step Bottleneck block is used to perform convolution calculation to generate a second feature map, the second feature map is linearly calculated to obtain a second new feature map, the second feature map and the second new feature map are compressed and output together; The P-Net network receives and processes the images to be identified of different sizes and generates candidate frame images; based on the candidate frame images, the original sample image is cropped or scaled to a size of 24×24 pixels and sent to the R-Net network for processing to filter out hole frame images containing holes; based on the hole frame images, the original image to be identified is cropped or scaled to a size of 48×48 pixels and sent to the O-Net network, which identifies and gives the final judgment result, which includes the location and number of holes; The operation process of the GhostNet network is as follows: receiving the feature map output by the first layer of convolution, generating partial feature maps through the convolution kernel, and generating new feature maps through a series of linear transformations of the partial feature maps; after repeating the above steps several times, compressing the height and width of all features to complete feature extraction, while reducing the number of convolution kernels, it can also ensure that the same characteristic information is obtained, ensuring the recognition accuracy of the MTCNN model; The P-Net network includes a first P convolutional layer, a first GhostNet network and a second P convolutional layer, the R-Net network includes an R convolutional layer, a second GhostNet network and an R fully connected layer, and the O-Net network includes an O convolutional layer, a third GhostNet network and an O fully connected layer in sequence; The processing process of the P-Net network is as follows: the first P layer convolution receives the sample image, and outputs several first P feature maps after convolution processing. The 1-step Bottleneck block receives the first P feature map, performs convolution calculation, and then outputs several second P feature maps through linear calculation. The 2-step Bottleneck block performs convolution calculation on the second P feature map, and then performs linear calculation to generate several third P feature maps. The length and width of the third P feature map are compressed to form a total feature map and output. The second P convolution layer receives the total feature map, performs feature extraction, and outputs several first candidate images of different sizes. It includes a picture acquisition module, the picture acquisition module and the picture recognition module are data-interoperable, the picture recognition module includes an FPGA chip, and the picture recognition module is directly connected to the picture acquisition module, so as to reduce the problem of multiple caches during the picture data transmission process and reduce the accuracy of the picture data affected by the transfer due to single particle failure; The R-Net network includes R convolutional layers, the second GhostNet network, and R fully connected layers. The second GhostNet network is composed of a 1-step Bottleneck block and a 2-step Bottleneck block stacked together. The processing process of the R-Net network is as follows: the R convolution layer receives the first candidate image, outputs several second candidate images after convolution processing, the 1-step Bottleneck block receives the second candidate image, performs convolution calculation, and then outputs several third candidate images through linear calculation, the first 2-step Bottleneck performs convolution calculation on the third candidate image, and then linearly transforms it to generate several fourth candidate images, compresses the length and width of the fourth candidate image and then outputs it, the second 2-step Bottleneck performs convolution calculation on the third candidate image, and then linearly transforms it to generate several fifth candidate images, compresses the length and width of the fifth candidate image and then outputs it, the R fully connected layer selects several first hole frame images in the fifth candidate image, and deletes non-hole frame images; The O-Net network includes O convolutional layers, the third GhostNet network, and O fully connected layers in sequence. The third GhostNet network is composed of a 1-step Bottleneck block and three 2-step Bottleneck blocks stacked together; The processing process of the O-Net network is as follows: the O convolutional layer receives the first hole frame diagram, outputs several second hole frame diagrams after convolution processing, the 1-step Bottleneck block receives the second hole frame diagram, performs convolution calculation, and then outputs several third hole frame diagrams through linear calculation, the first 2-step Bottleneck performs convolution calculation on the third hole frame diagram, and then linearly transforms to generate several fourth hole frame diagrams, compresses the length and width of the fourth hole frame diagram and then outputs, the second 2-step Bottleneck performs convolution calculation on the fourth hole frame diagram, and then linearly transforms to generate several fifth hole frame diagrams, compresses the length and width of the fifth hole frame diagram and then outputs; The third 2-step Bottleneck performs convolution calculation on the fifth hole frame diagram, and then linearly transforms it to generate several sixth hole frame diagrams, compresses the length and width of the sixth hole frame diagram and outputs them; the O fully connected layer receives all the sixth hole frame diagrams to output the final hole judgment result.
2. The device based on the image recognition architecture according to claim 1, characterized in that: The first GhostNet network is composed of a 1-step Bottleneck block and a 2-step Bottleneck block stacked together.
3. The device based on the image recognition architecture according to claim 1, characterized in that: The second GhostNet network consists of a 1-step Bottleneck block and two 2-step Bottleneck blocks stacked together.
4. The device based on the image recognition architecture according to claim 1, characterized in that: The third GhostNet network is composed of a 1-step Bottleneck block and three 2-step Bottleneck blocks stacked together.
Citation Information
Patent Citations
Edge processing data de-identification
CN115443466A
Method and device for lightening target detection model, electronic equipment and medium
CN116863419A