Dual-path Residual Structure Neural Network Model and Image Target Recognition System
By introducing a dual-path residual structure into the deep learning model, the gradient diffusion problem is solved, and the feature extraction and classification performance of the model is improved.
Patent Information
- Application Number
- CN202111540075.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-15
AI Technical Summary
When the existing deep learning model reaches a certain depth, due to the gradient diffusion problem, it cannot effectively promote weight adjustment, which affects the improvement of the classification performance of the model.
Using a dual-channel residual structure neural network model, the forward propagation of image features and the back propagation of gradients are promoted by setting the first residual branch and the second residual branch.
It effectively alleviates the problem of gradient diffusion and improves the feature extraction performance and target classification performance of deep neural networks.
Smart Images

Figure CN114399023B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision image recognition, and particularly to a dual-path residual structure neural network model and an image target recognition system. Background Art
[0002] Deep learning technology is an important method for image classification. A deep neural network constructed by stacking single-layer neural networks layer by layer has good feature extraction and image classification performance.
[0003] However, there is a significant problem in existing deep learning models for image classification, that is, the gradient vanishing problem. Restricted by this problem, when the deep learning model reaches a certain depth, due to the weak backpropagation gradient, it is unable to promote the adjustment of weights, thus affecting the improvement of the model's classification performance.
[0004] Deep Residual Learning is an effective method to alleviate the gradient vanishing of deep neural networks. This method designs cross-layer connections in the deep neural network to promote the forward propagation of features and the backward propagation of gradients. However, the existing Deep Residual Learning model (ResNet, Kaiming He, CVPR 2015) has only one cross-layer connection, which is a single-path residual structure and still cannot more fully promote the forward propagation of features and the backward propagation of gradients. Summary of the Invention
[0005] To solve the above problems in the prior art, the present invention provides a dual-path residual structure neural network model and an image target recognition system.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A dual-path residual structure neural network model, comprising: at least two cascaded network units in sequence;
[0008] Each network unit includes a plurality of cascaded convolutional neural network modules in sequence;
[0009] A first residual branch is provided between the input end and the output end of the network unit; a second residual branch is provided between two network units; the input and output of the first residual branch are both the image to be recognized; the input and output of the second residual branch are both the output of any one of the convolutional neural network modules in the previous network unit; the second residual branch is used to superimpose the output of any one of the convolutional neural network modules in the previous network unit onto any convolutional neural network module in the subsequent network unit for convolutional processing to obtain the target recognition result.
[0010] Preferably, each network unit includes a first convolutional neural network module, a second convolutional neural network module, and a third convolutional neural network module that are cascaded in sequence;
[0011] A first residual branch is provided between the input end of the first convolutional neural network module in the current network unit and the output end of the third convolutional neural network module; a second residual branch is provided between the input end of the second convolutional neural network module in the current network unit and the output end of the first convolutional neural network module in the next network unit cascaded with the current network unit; the input of the first convolutional neural network module in the current network unit is the image to be recognized; the superimposed result of the output of the first residual branch and the output of the third convolutional neural network module in the current network unit is used as the input of the first convolutional neural network module in the next network unit.
[0012] Preferably, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module are all composed of a single convolutional layer.
[0013] Preferably, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module all include a single convolutional layer and a non-linear activation layer; the non-linear activation layer is added after the single convolutional layer.
[0014] Preferably, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module are all composed of a double convolutional layer.
[0015] Preferably, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module all include a double convolutional layer and a non-linear activation layer; the non-linear activation layer is added after the double convolutional layer.
[0016] Preferably, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module are all composed of a triple convolutional layer.
[0017] Preferably, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module all include a triple convolutional layer and a non-linear activation layer; the non-linear activation layer is added after the triple convolutional layer.
[0018] According to the specific embodiments provided by the present invention, the following technical effects of the present invention are disclosed:
[0019] The dual-path residual structure neural network model provided by the present invention obtains a dual-path residual structure by setting a first residual branch and a second residual branch, which can promote the forward propagation of image features and the backward propagation of gradients, so as to make up for the problem that the prior art cannot fully solve the gradient dispersion problem, and further improve the feature extraction performance and target classification performance of the deep neural network.
[0020] Corresponding to the above-provided dual-path residual structure neural network model, the present invention also provides an implementation manner as follows:
[0021] An image target recognition system includes: an image acquisition unit and a processing unit; the image acquisition unit is connected to the processing unit; an image target recognition model is implanted in the processing unit; the image target recognition model is the above-provided dual-path residual structure neural network model;
[0022] The image acquisition unit is used to acquire an image to be recognized; the processing unit is used to use the image target recognition model implanted therein to obtain a target recognition result based on the image to be recognized acquired by the image acquisition unit.
[0023] The image target recognition method and system provided by the present invention can further improve the accuracy and timeliness of image target recognition based on the above-provided dual-path residual structure neural network model.
[0024] The above general description and the following description are only exemplary and explanatory, and are not used to limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and among them:
[0026] Figure 1 is a schematic structural diagram of the dual-path residual structure neural network model provided by the present invention;
[0027] Figure 2 is a schematic structural diagram of the image target recognition system provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In order to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only, and are not intended to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, numerous details are provided to give a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be shown in a simplified manner to simplify the drawings.
[0029] The dual-path residual structure neural network model provided by the present invention includes: at least two cascaded network units in sequence;
[0030] Each network unit includes a plurality of cascaded convolutional neural network modules in sequence;
[0031] A first residual branch is provided between the input end and the output end of the network unit; a second residual branch is provided between two network units; the input and output of the first residual branch are both the image to be recognized; the input and output of the second residual branch are both the output of any one of the convolutional neural network modules in the previous network unit; the second residual branch is used to superimpose the output of any one of the convolutional neural network modules in the previous network unit onto any convolutional neural network module in the subsequent network unit for convolutional processing to obtain the target recognition result.
[0032] Taking each network unit including three cascaded convolutional neural network modules (the first convolutional neural network module A, the second convolutional neural network module B, and the third convolutional neural network module C) as an example, the data transmission process of the above-mentioned dual-path residual structure neural network model will be described below.
[0033] As Figure 1 shown, the above-mentioned dual-path residual structure neural network model is named DoubleResBlock. Each network unit includes three stacked convolutional processing modules, that is, the first convolutional neural network module A, the second convolutional neural network module B, and the third convolutional neural network module C cascaded in sequence. The input information will sequentially pass through the first convolutional neural network module A, the second convolutional neural network module B, and the third convolutional neural network module C to form the main path of information processing. The inputs and outputs of the first convolutional neural network module A, the second convolutional neural network module B, and the third convolutional neural network module C on the main path are called the main inputs and main outputs of the first convolutional neural network module A, the second convolutional neural network module B, and the third convolutional neural network module C.
[0034] A first residual branch is provided between the input end of the first convolutional neural network module A in the current network unit and the output end of the third convolutional neural network module C (that is, Figure 1The residual branch in (1). There is a second residual branch (i.e., Figure 1 the residual branch in (2).
[0035] Based on the above structural data transmission, the processing process can be as follows: The main input information of the first convolutional neural network module A (i.e., the image to be target-recognized) will be cross-connected to the third convolutional neural network module C via a newly added branch, and then added to the main output information of the third convolutional neural network module C to form the final output information of the third convolutional neural network module C. The main output information of the first convolutional neural network module A will be cross-connected to the second convolutional neural network module B of the next network unit via the second residual branch, and added to the main input of the second convolutional neural network module B to form the final input information of the second convolutional neural network module B (i.e., the target recognition result).
[0036] In order to enable the information passing through this branch to be properly transformed and thus conform to the relevant matrix data addition rules, the above first residual branch and second residual branch can both selectively add convolutional layers.
[0037] Based on the above settings, the present invention can promote the forward propagation of image features and the backward propagation of gradients through a dual-path residual structure (i.e., the first residual branch and the second residual branch), making up for the problem that the prior art cannot fully solve the gradient dispersion problem, thereby improving the feature extraction performance and target classification performance of the deep neural network.
[0038] In the specific design of the deep learning model, multiple network units need to be cascaded in series, and the first residual branch and the second residual branch with cross-layer connections are set between adjacent modules, so as to link multiple cascaded modules with the main information path and the residual branch to achieve collaborative work.
[0039] Furthermore, in order to increase the diversity of the formed dual-path residual structure neural network model, the convolutional neural network modules A, B, and C embedded in each network unit can be composed of single-layer, double-layer, or triple-layer convolutional units. A non-linear activation layer can be added after each convolution, or not added.
[0040] Therefore, the neural network model based on the dual-path residual structure constituted by the present invention can construct an image classification model and an image target detection and recognition model based on deep learning technology, and has a wide range of uses.
[0041] Corresponding to the above-provided dual-path residual structure neural network model, the present invention also provides an image target recognition system, such as Figure 2As shown in the figure, the system includes an image acquisition unit 200 and a processing unit 201. The image acquisition unit 200 is connected to the processing unit 201. An image target recognition model is implanted in the processing unit 201. The image target recognition model is the dual-path residual structure neural network model provided above.
[0042] The image acquisition unit 200 is used to acquire the image to be recognized. The processing unit 201 is used to obtain the target recognition result based on the image to be recognized acquired by the image acquisition unit 200 by using the image target recognition model implanted therein.
[0043] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. Embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. The scope of the embodiments of the present disclosure includes the entire scope of the claims and all available equivalents of the claims. When used in this application, although terms such as "first", "second", etc. may be used in this application to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without changing the meaning of the description, the first element may be called the second element, and similarly, the second element may be called the first element, as long as all occurrences of the "first element" are consistently renamed and all occurrences of the "second element" are consistently renamed. The first element and the second element are both elements, but they may not be the same element. Moreover, the terms used in this application are only used to describe the embodiments and do not limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising", etc. refer to the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups of these. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, or device comprising the element. Herein, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts may refer to the description of the method part.
[0044] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0045] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. In addition, the functional units in the embodiments of the present disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0046] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion thereof that contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A method for constructing a dual-path residual structure neural network model, characterized in that, the dual-path residual structure neural network model includes: at least two cascaded network units in sequence; each network unit includes a plurality of cascaded convolutional neural network modules in sequence; a first residual branch is provided between the input end and the output end of the network unit; a second residual branch is provided between two network units; the input and output of the first residual branch are both the image to be recognized; the input and output of the second residual branch are both the output of any one of the convolutional neural network modules in the previous network unit; the second residual branch is used to superimpose the output of any one of the convolutional neural network modules in the previous network unit onto any convolutional neural network module in the subsequent network unit for convolutional processing to obtain the target recognition result; wherein, each network unit includes a first convolutional neural network module, a second convolutional neural network module, and a third convolutional neural network module cascaded in sequence; a first residual branch is provided between the input end of the first convolutional neural network module in the current network unit and the output end of the third convolutional neural network module; a second residual branch is provided between the input end of the second convolutional neural network module in the current network unit and the output end of the first convolutional neural network module in the next network unit cascaded with the current network unit; the input of the first convolutional neural network module in the current network unit is the image to be recognized; the superimposed result of the output of the first residual branch and the output of the third convolutional neural network module in the current network unit is used as the input of the first convolutional neural network module in the next network unit.
2. The method for constructing a dual-path residual structure neural network model according to claim 1, characterized in that, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module are all composed of a single convolutional layer.
3. The method for constructing a dual-path residual structure neural network model according to claim 1, characterized in that, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module all include a single convolutional layer and a non-linear activation layer; the non-linear activation layer is added after the single convolutional layer.
4. The method for constructing a dual-path residual structure neural network model according to claim 1, characterized in that, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module are all composed of a double convolutional layer.
5. The method for constructing a dual-path residual structure neural network model according to claim 1, characterized in that, the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module all include a double convolutional layer and a non-linear activation layer; the non-linear activation layer is added after the double convolutional layer.
6. The method for constructing a dual-path residual structure neural network model according to claim 1, characterized in that, The first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module are all composed of three convolutional layers.
7. The method for constructing a dual-path residual structure neural network model according to claim 1, characterized in that the first convolutional neural network module, the second convolutional neural network module, and the third convolutional neural network module all include three convolutional layers and a non-linear activation layer; the non-linear activation layer is added after the three convolutional layers.
8. An image target recognition system, characterized in that it includes: an image acquisition unit and a processing unit; the image acquisition unit is connected to the processing unit; an image target recognition model is implanted in the processing unit; the image target recognition model is a dual-path residual structure neural network model constructed by the method for constructing a dual-path residual structure neural network model according to any one of claims 1-7; the image acquisition unit is used to acquire an image to be recognized; the processing unit is used to obtain a target recognition result based on the image to be recognized acquired by the image acquisition unit by using the image target recognition model implanted therein.
Citation Information
Patent Citations
Hybrid degraded image enhancement method based on convolutional neural network
CN112801904A
Meta-architecture design for a CNN network
DE102017128082A1