Image recognition method and device
By using the neural network model of rotation invariance, the problem of low accuracy in picture book recognition of picture book robots is solved, and efficient image recognition is achieved at different placement positions and angles.
Patent Information
- Application Number
- CN202010761239.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-07-31
AI Technical Summary
Existing picture book robots have high requirements for the placement of picture books when identifying picture books, which makes it difficult for young children to accurately place picture books, resulting in a decrease in recognition accuracy or even unrecognition.
Using a neural network model with rotation invariance, feature extraction is performed on the image to be recognized through the first neural network, and point multiplication operation of the feature map is combined with the second neural network to build a rotation invariance feature network to improve the accuracy of image recognition.
It improves the accuracy of image recognition, reduces the requirements for the placement position and angle of picture books, and enhances the efficiency and accuracy of the image recognition process.
Smart Images

Figure CN112084849B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of neural networks, and particularly to an image recognition method and apparatus. Background Art
[0002] With the development of technology, some early childhood education product robots with picture book reading functions (referred to as picture book robots for short) have emerged on the market. Before reading a picture book, the picture book robot needs to accurately recognize the picture book. Specifically, the robot first captures an image of a certain page of the picture book through a camera, then performs local feature detection on the image, and then matches the detection result with a pre-stored picture book image template in the database to obtain the image with the highest matching degree with the detection result, and uses the image with the highest matching degree with the detection result as the image to be read. Subsequently, the picture book robot reads the image to be read.
[0003] The above method for recognizing a picture book has relatively high requirements for the placement position of the picture book. For example, it is required that the picture book is laid flat on a horizontal plane, which is the same as the horizontal plane where the picture book robot is located; it is also required that the distance and angle between the picture book and the picture book robot meet certain requirements. In addition, it is also required that the picture book robot stands upright without falling.
[0004] However, in actual applications, young children usually have difficulty placing the picture book according to the above requirements, which will cause the accuracy rate of the picture book robot for recognizing the picture book to be greatly reduced, or even unable to recognize. Summary of the Invention
[0005] The embodiments of this application provide an image recognition method and apparatus, which helps to improve the accuracy rate of image recognition.
[0006] To achieve the above object, this application provides the following technical solutions:
[0007] In a first aspect, an image recognition method is provided, including: First, obtain an image to be recognized. Then, use a first neural network to extract features from the image to be recognized to obtain a first feature map. Next, use a second neural network to extract features from the first feature map to obtain a second feature map, and perform a dot product of the second feature map and the first feature map to obtain a third feature map; wherein, the third feature map represents the feature map obtained after transforming the features of the image to be recognized to the main direction. And, obtain a first score map of the image to be recognized based on the third feature map. Finally, recognize the image to be recognized based on the third feature map and the first score map. In this technical solution, using the second neural network to extract features from the first feature map to obtain a second feature map, and performing a dot product of the second feature map and the first feature map to obtain a third feature map helps to construct a network with rotation-invariant features, and performing image recognition based on this network helps to improve the accuracy of image recognition.
[0008] In a possible design, a first neural network is used to extract features from the image to be recognized, obtaining a first feature map, including: using the first neural network to perform at least one convolutional operation on the image to be recognized to obtain the first feature map.
[0009] In a possible design, a second neural network is used to extract features from the first feature map, obtaining a second feature map, including: using the second neural network to perform at least one convolutional operation on the first feature map to obtain the second feature map. In this possible design, the second feature map is obtained by performing at least one convolutional operation on the first feature map for feature extraction, and the operation is simple.
[0010] In a possible design, a second neural network is used to extract features from the first feature map, obtaining a second feature map, including: using the second neural network to perform at least one convolutional operation on the first feature map; performing at least one pooling operation and / or fully connected operation on the first feature map after the convolutional operation to obtain the second feature map. In this possible design, the first feature map is subjected to feature extraction through at least one convolutional operation, and at least one pooling operation and / or fully connected operation, which helps to achieve more complex feature extraction, thereby helping to make the result of feature extraction more accurate, and further helping to improve the accuracy of image recognition.
[0011] In a possible design, the size of the third feature map is M1*N1*P1, the size of the first score map is M1*N1, P1 is the size of the feature direction dimension, M1*N1 is the size perpendicular to the feature direction dimension, and M1, N1, and P1 are all positive integers. In this possible design, the third feature map and the first score map are directly used to recognize the image to be recognized. This solution is simple to implement.
[0012] In a possible design, the size of the third feature map is M2 * N2 * P2, and the size of the first score map is M1 * N1. P2 is the size of the feature direction dimension, and M1, N1, P1, M2, N2, and P2 are all positive integers. Based on the third feature map and the first score map, the image to be recognized is recognized, including: extracting features from the third feature map to obtain a fourth feature map; wherein, the size of the fourth feature map is M1 * N1 * P1; P1 is the size of the feature direction dimension, and P1 is a positive integer; based on the fourth feature map and the first score map, the image to be recognized is recognized; wherein, the size of the first score map is M1 * N1. Based on this optional implementation, using the first score map and the feature map obtained by extracting features from the third feature map to recognize the image to be recognized helps to change the size of the feature map. Since generally, the larger the size of the feature map, the lower the efficiency of the image recognition process, and the larger the size of the feature map, the more accurately the feature map can represent the image to be recognized; therefore, changing the size of the feature map helps to balance the efficiency and accuracy of the image recognition process, thereby improving the overall performance of the image recognition process.
[0013] In a possible design, M1 * N1 < M2 * N2. In this way, it helps to reduce the size of the feature map used in the image recognition process, thereby reducing the processing complexity of the image recognition process to improve the processing efficiency of the image recognition process.
[0014] In a possible design, obtaining the first score map of the image to be recognized based on the third feature map includes: performing a convolution operation on the third feature map using a 1-channel convolution kernel to obtain X fifth feature maps; wherein, the size of the feature direction of the fifth feature map is smaller than the size of the feature direction of the third feature map; X is an integer greater than 2; weighted summing the elements of the X fifth feature maps to obtain a sixth feature map; extracting features from the sixth feature map to obtain the first score map. In this possible design, during the process of obtaining the score map, only the size of the feature direction of the third feature map is compressed, so the implementation is simple.
[0015] In a possible design, obtaining the first score map of the image to be recognized based on the third feature map includes: extracting features from the third feature map to obtain a seventh feature map; wherein, the dimension size perpendicular to the feature direction of the third feature map is greater than the dimension size perpendicular to the feature direction of the seventh feature map; X is an integer greater than 2; performing a convolution operation on the seventh feature map using a 1-channel convolution kernel to obtain X fifth feature maps; weighted summing the elements of the X fifth feature maps to obtain a sixth feature map; extracting features from the sixth feature map to obtain the first score map. In this possible design, during the process of obtaining the score map, both the size of the feature direction and the size perpendicular to the feature direction of the third feature map are compressed, so it helps to reduce the complexity of the image processing process, thereby improving the processing efficiency of the image recognition process.
[0016] In a possible design, the size of the image to be recognized is larger than the size of the first score map. Since the size of the first score map (assumed to be a*b) used in the image recognition process represents the number of features in the feature map used in this process, in this possible design, if the features of the image to be recognized are dense features, then the feature map corresponding to the first score map is a sparse feature. Using sparse features for image recognition helps reduce the complexity of the image processing process, thereby improving the processing efficiency of the image recognition process.
[0017] In a second aspect, the present application provides an image recognition device.
[0018] In a possible design, the image recognition device is used to execute any one of the methods provided in the first aspect above. The present application can divide the functional modules of the image recognition device according to any one of the methods provided in the first aspect above. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. Exemplarily, the present application can divide the image recognition device into an acquisition unit, a feature extraction unit, a recognition unit, etc. according to functions. The descriptions of the possible technical solutions and beneficial effects executed by each of the above-divided functional modules can refer to the technical solutions provided in the first aspect above or its corresponding possible design, which will not be elaborated here.
[0019] In another possible design, the image recognition device includes: a memory and one or more processors, and the memory and the processor are coupled. The memory is used to store computer instructions, and the processor is used to call the computer instructions to execute any one of the methods provided in the first aspect and any of its possible design manners.
[0020] In a third aspect, the present application provides a computer-readable storage medium, such as a non-transitory computer-readable storage medium. Computer programs (or instructions) are stored thereon. When the computer programs (or instructions) run on the image recognition device, the image recognition device is enabled to execute any one of the methods provided in any one of the possible implementation manners in the first aspect above.
[0021] In a fourth aspect, the present application provides a computer program product, which when running on a computer, enables any one of the methods provided in any one of the possible implementation manners in the first aspect to be executed.
[0022] In a fifth aspect, the present application provides a chip system, including: a processor, and the processor is used to call and run the computer program stored in the memory and execute any one of the methods provided in the implementation manner in the first aspect.
[0023] It can be understood that any of the above-provided image recognition devices, computer storage media, computer program products, chip systems, etc. can be applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods, which will not be elaborated here.
[0024] In this application, the names of the above image recognition devices do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear under other names. As long as the functions of each device or functional module are similar to those of this application and fall within the scope of the claims of this application and their equivalent technologies.
[0025] These aspects or other aspects of this application will be more clearly understood in the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 Schematic diagram of the hardware structure of a computer device applicable to an embodiment of this application;
[0027] Figure 2a Schematic diagram of a deep learning network model provided by an embodiment of this application;
[0028] Figure 2b Schematic diagram of another deep learning network model provided by an embodiment of this application;
[0029] Figure 3 Schematic diagram of the logical structure of a first neural network provided by an embodiment of this application;
[0030] Figure 4 Schematic diagram of the logical result of a second neural network provided by an embodiment of this application;
[0031] Figure 5 Schematic diagram of each dimension of a feature map provided by an embodiment of this application;
[0032] Figure 6 Schematic diagram of the flowchart of a method for obtaining training data provided by an embodiment of this application;
[0033] Figure 7 Schematic diagram of a reference image applicable to an embodiment of this application and a sample image obtained after performing a homography transformation on the reference image;
[0034] Figure 8 Schematic diagram of the relationship between reference data and training data provided by an embodiment of this application;
[0035] Figure 9 Schematic diagram of the connection relationship between a front-end network, an adversarial network, and a siamese network provided by an embodiment of this application;
[0036] Figure 10 Schematic flowchart of the method for training the front - end network provided by the embodiment of the present application;
[0037] Figure 11 Schematic diagram of the logical structure of an adversarial network provided by the embodiment of the present application;
[0038] Figure 12 Schematic diagram of the logical structure of an extraction network provided by the embodiment of the present application;
[0039] Figure 13 Schematic diagram of the logical structure of a representation network provided by the embodiment of the present application;
[0040] Figure 14 Schematic flowchart of an image recognition process provided by the embodiment of the present application;
[0041] Figure 15 Schematic flowchart of another image recognition process provided by the embodiment of the present application;
[0042] Figure 16 Schematic diagram of the structure of an image recognition device provided by the embodiment of the present application;
[0043] Figure 17 Schematic diagram of the structure of a chip system provided by the embodiment of the present application;
[0044] Figure 18 Conceptual partial view of a computer program product provided by the embodiment of the present application. Detailed implementation manners
[0045] First, some terms and technologies involved in the present application are described:
[0046] Feature: That is, an image feature, which may include color features, texture features, shape features, local feature points, etc.
[0047] Global feature / Local feature: A global feature refers to the overall attribute of an image. Common global features include color features, texture features, and shape features, etc. A global feature uses all the features of an image to represent the image, and such features have a large amount of redundant information. A local feature refers to the local attribute of an image. A local feature uses local feature points of an image to represent the image. Each local feature point only contains the information of the image block where it is located and is not aware of the global information of the image.
[0048] Feature point (i.e., local feature point): In image processing, for the same object or scene, if multiple images are collected from different angles and the same parts of the object or scene can be recognized with the same result, then these parts are said to have scale invariance. The pixel points or pixel blocks (i.e., pixel blocks composed of multiple pixel points) with "scale invariance" are feature points. In one example, if a pixel point in an image is an extreme point in its neighborhood (such as the point with the maximum or minimum value), then it is determined that this pixel point is a feature point.
[0049] Image patch: A local square region in an image, such as an image region of 4*4 pixels or 8*8 pixels. Here, a*a pixels represents a square region with a width and height of a pixels respectively, and a is an integer greater than or equal to 1.
[0050] Homograph: Also known as projective transformation. It maps points (three-dimensional homogeneous vectors) on one projective plane to another projective plane. It satisfies Y = H*X, where H is a 3*3 matrix (also called the homography matrix), X is the position coordinate of a pixel point in the source image, and Y is the position coordinate of the corresponding pixel point on the mapped target image. In picture book recognition, a picture book can be regarded as a plane, and the subset of its corresponding geometric transformations is homograph. The homography matrix that determines the transformation is a matrix composed of properties such as rotation, translation, and scaling (such as a 3*3 matrix). If one image is obtained from another image through homograph transformation, then it is considered that there is a homograph transformation relationship between these two images.
[0051] Histogram of Oriented Gradient (HOG): A histogram, also known as a quality distribution diagram, is a statistical report diagram, which is represented by a series of vertical stripes or line segments with different heights to show the data distribution. Generally, the horizontal axis represents the data type, and the vertical axis represents the distribution. The histogram of oriented gradient is a statistical value used to calculate the direction information of local image gradients.
[0052] Principal direction: In an image / image patch, by calculating the gradient directions between adjacent pixels (i.e., the unit vector of the vector difference between adjacent pixels), a histogram of oriented gradient is established. The gradient at the peak in the histogram of oriented gradient is the principal direction of the image / image patch.
[0053] Convolutional Neural Network (CNN): A type of feedforward neural network, whose artificial neurons can respond to surrounding units within a certain coverage range and has excellent performance in large-scale image processing.
[0054] Max pooling: The most direct purpose of the pooling layer is to reduce the amount of data to be processed in the next layer. In max pooling, for a certain filter that extracts several feature values, only the largest of these feature values is retained as the reserved value, and all other feature values are discarded. The largest value represents retaining only the strongest of these features and discarding other weaker such features.
[0055] Rotation invariance: In physics, if the properties of a physical system are independent of its orientation in space, then the system has rotational invariance. In image processing, if, under any arbitrary rotation angle in the plane, the features extracted by a feature extractor for an image change very little, then the feature extractor is said to have rotational invariance. Among them, the feature extractor can be a picture book robot or a functional module in a picture book robot, such as a neural network.
[0056] Loss function: The loss function is used to measure the degree of inconsistency between the predicted value f(x) of a model and the true value Y. It is a non-negative real-valued function, usually denoted by L(Y, f(x)). The smaller the loss function, the better the robustness of the model. The goal of an optimization problem is to minimize the loss function. An objective function is usually either the loss function itself or its negative value. When an objective function is the negative value of the loss function, the value of the objective function seeks to be maximized.
[0057] Sparse features, dense features: In local feature detection, if the position index of each pixel in the image is recorded, and each index should correspond to a feature, then sparse features refer to the situation where, in the index set, most of the indexes are empty, or most of the indexes have no corresponding features. Dense features refer to the situation where most of the indexes are not empty, that is, most of the indexes have their corresponding feature descriptions.
[0058] Local feature detection algorithm: The local feature detection algorithm consists of two parts: "extraction" and "representation". The purpose of "extraction" is to determine whether each pixel point (or image patch) in the image is a feature point. "Representation" means that for all detected feature points, according to their neighborhoods, they are represented as feature values in the same dimension. By calculating the distance between the feature values of two feature points, it is possible to determine whether the two feature points are similar, and then, based on the number or ratio of similar feature points in two images, the similarity degree of the two images can be judged. Therefore, the evaluation criterion for the local feature detection algorithm is: the matching accuracy rate at which feature points are successfully matched for two images with the same / similar regions.
[0059] A high homography transformation scenario refers to a scenario where the feature representations before and after the transformation are very different (i.e., the feature points determined before and after the transformation are very different), such as a picture book recognition scenario.
[0060] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0061] In the embodiments of the present application, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0062] In the present application, the meaning of the term "at least one" is one or more, and the meaning of the term "a plurality" is two or more. For example, a plurality of second messages means two or more second messages. In this document, the terms "system" and "network" are often used interchangeably.
[0063] It should be understood that the terms used in the description of various examples herein are only for the purpose of describing specific examples and are not intended to be limiting. As used in the description of various examples and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0064] It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" is a description of the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the associated objects before and after.
[0065] It should also be understood that in the various embodiments of the present application, the magnitudes of the sequence numbers of the various processes do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0066] It should be understood that determining B based on A does not mean determining B solely based on A. B can also be determined based on A and / or other information.
[0067] It should also be understood that the term "comprising" (also referred to as "includes", "including", "comprises" and / or "comprising") when used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0068] It should also be understood that the term "if" can be interpreted to mean "when" ("when" or "upon") or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined..." or "if [the stated condition or event] is detected" can be interpreted to mean "when it is determined..." or "in response to determining..." or "when [the stated condition or event] is detected" or "in response to detecting [the stated condition or event]".
[0069] It should be understood that the "one embodiment", "an embodiment", "a possible implementation" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment or implementation are included in at least one embodiment of this application. Therefore, the "in one embodiment" or "in an embodiment", "a possible implementation" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.
[0070] Currently, the following local detection algorithms are usually adopted for picture book recognition:
[0071] The first one: a local detection algorithm based on the manual feature method, that is, the extraction and representation of local feature points are both rule-based. For example, in the judgment of extreme points, it is necessary to compare the pixel value of each pixel point with the pixel values of its surrounding neighborhood pixel points one by one. In the judgment of the main direction, it is necessary to construct a gradient direction histogram one by one, etc. When representing features, complex steps such as normalization and direction correction are required. And fixed parameters need to be set through experiments for each of these steps.
[0072] When determining whether two images are locally similar, if the geometric shape of the similar region (i.e., the positions of the same physical region under different shooting angles) becomes smaller, in the manual-based local feature detection, the corresponding features change less, and it is relatively easier to match correctly. In the picture book recognition scenario, for the same page in a picture book, when the position of the picture book is different, the feature changes of the image of this page scanned by the picture book robot are relatively large. When the similar region undergoes a large geometric deformation, the distribution of its extreme points changes greatly. When the local region shrinks or the geometric deformation is large, the pixel points that were originally extreme points may no longer be represented as extreme points under the manual rules, and thus cannot be determined as feature points, which will lead to deviations in the representation of some feature points and make it impossible to match correctly.
[0073] Second: The local detection algorithm of the deep learning-based method, that is, the input of the neural network is an image, and the output is a score map (that is, the probability value from 0 to 1 corresponding to the possibility that each pixel point (or pixel block) in the image is considered a feature point) of each pixel point (or pixel block) in the image being considered a feature point, and a feature map of each pixel point (or pixel block) corresponding to a feature value. This method is a non-end-to-end method. On the one hand, the extraction of features in this method still depends on manual feature extraction. Therefore, the above problems will also exist. On the other hand, this neural network is usually a convolutional neural network, and the convolutional neural network only has rotational invariance to a certain extent, and will not perform rotation and normalization on feature points like the first method above. Therefore, in the high homography transformation scenario, the difference in feature representations before and after the transformation is very large, resulting in a very low matching accuracy.
[0074] Based on this, the embodiments of the present application provide a neural network model training method and an image recognition method, which are applied to the high homography transformation scenario (such as the picture book recognition scenario). Specifically: In the model training stage, a neural network with rotational invariance is trained based on multiple images. More precisely, a neural network with a higher degree of rotational invariance than the convolutional neural network in the prior art is trained. Among them, the multiple images include images with a homography transformation relationship. In the image recognition stage, the image is recognized based on the neural network with rotational invariance. In this way, compared with the prior art, it helps to make the difference in feature representations before and after the transformation small, thereby improving the matching accuracy.
[0075] The neural network model training method and the image recognition method provided by the embodiments of the present application can be applied to the same or different computer devices respectively. For example, the neural network model training method can be executed by a computer device such as a server or a terminal. The image recognition method can be executed by a terminal (such as a picture book robot, etc.). The embodiments of the present application do not limit this.
[0076] Such as Figure 1As shown, it is a schematic diagram of the hardware structure of a computer device 10 applicable to the embodiments of the present application.
[0077] Referring to Figure 1 , the computer device 10 includes a processor 101, a memory 102, an input / output device 103, and a bus 104. Among them, the processor 101, the memory 102, and the input / output device 103 can be connected through the bus 104.
[0078] The processor 101 is the control center of the computer device 10, which can be a general-purpose central processing unit (CPU) or other general-purpose processors. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0079] As an example, the processor 101 can include one or more CPUs, such as Figure 1 the CPUs 0 and 1 shown in
[0080] The memory 102 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0081] In a possible implementation, the memory 101 can exist independently of the processor 101. The memory 102 can be connected to the processor 101 through the bus 104 and is used to store data, instructions, or program code. When the processor 101 calls and executes the instructions or program code stored in the memory 102, it can implement the neural network model training method and / or the image recognition method provided by the embodiments of the present application.
[0082] In another possible implementation, the memory 102 can also be integrated with the processor 101.
[0083] An input / output device 103 is configured to input parameter information such as a sample image and an image to be recognized, so that the processor 101 executes instructions in the memory 102 according to the input parameter information to perform the neural network model training method and / or the image recognition method provided in the embodiments of the present application. Generally, the input / output device 103 may be an operation panel or a touch screen, or any other device capable of inputting parameter information, which is not limited in the embodiments of the present application.
[0084] The bus 104 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 1 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.
[0085] It should be noted that, Figure 1 the structure shown in the figure does not constitute a limitation on the computer device 10. Except Figure 1 for the components shown, the computer device 10 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0086] Hereinafter, the technical solutions provided in the embodiments of the present application will be described with reference to the accompanying drawings:
[0087] The model adopted in the embodiments of the present application is a deep learning network model (or neural network model, hereinafter simply referred to as network model). As Figure 2a and Figure 2b shown, it is a schematic diagram of two deep learning network models provided in the embodiments of the present application.
[0088] Figure 2a The network model shown in includes: a front-end network 41 and a representation network 42.
[0089] Figure 2b The network model shown in includes: a front-end network 41, a representation network 42, and an extraction network 43.
[0090] The input of the front-end network 41 is an image, and the output is the third feature map of the image. Among them, the third feature map represents the feature map obtained by transforming the features (such as texture features, etc.) of the image input to the front-end 41 to the main direction. In the training stage, the input of the front-end network 41 is a sample image. In the image recognition stage, the input of the front-end network 41 is an image to be recognized.
[0091] Optionally, the front-end network 41 may include a first neural network 411 and a second neural network 412.
[0092] The first neural network 411 is used to extract features from the input image (i.e., the input image). For example, at least one convolutional operation is performed on the input image to obtain a first feature map. The first feature map can be a three-dimensional tensor, and an element in the tensor corresponds to a region in the input image, and this region can also be referred to as the receptive field of the convolutional neural network.
[0093] Exemplarily, as Figure 3 shown, it is a schematic diagram of the logical structure of a first neural network 411 provided by an embodiment of the present application. Among them, the size of the input image of the first neural network 411 is H*W*3, and the size of the output first feature map is H / 4*W / 4*64. The first neural network 411 includes 4 convolutional layers (labeled conv1-1, conv1-2, conv2-1, conv2-2) respectively.
[0094] The second neural network 412 is used to correct the first feature map to obtain a third feature map. Optionally, the second neural network 412 is used to extract features from the first feature map to obtain a second feature map, and the second feature map is multiplied point by point with the first feature map to obtain a third feature map.
[0095] In one implementation, the second neural network 412 is specifically used to perform at least one convolutional operation on the first feature map to obtain a second feature map.
[0096] In another implementation, the second neural network 412 is specifically used to perform at least one convolutional operation on the first feature map, and then perform at least one pooling operation and / or fully connected operation on the first feature map after the convolutional operation to obtain a second feature map. Exemplarily, as Figure 4 shown, it is a schematic diagram of the logical result of a second neural network 412 provided by an embodiment of the present application. Figure 4 It is drawn based on Figure 3 The size of the first feature map input to the second neural network 412 is H / 4*W / 4*64, and the size of the output third feature map is H / 4*W / 4*64. Figure 4 The second neural network 412 shown includes 2 convolutional layers, 1 fully connected layer and 1 point multiplication layer. This implementation can achieve more complex feature extraction, which helps to make the result of feature extraction more accurate, and thus helps to improve the accuracy of image recognition.
[0097] Based on Figure 2a the network model shown:
[0098] A network 42 is represented, which is used to obtain a first score map based on a third feature map. The first score map is the score map of the image input to the front-end network. Among them, the size of the third feature map is M1*N1*P1, the size of the first score map is M1*N1, P1 is the size of the feature direction dimension, M1*N1 is the size perpendicular to the feature direction dimension, and M1, N1, and P1 are all positive integers.
[0099] As Figure 5 shown, it is a schematic diagram of each dimension of a feature map provided by an embodiment of the present application. Figure 5 In it, a feature map with a size of H / 4*H / 4*64 is used as an example for illustration. In the embodiment of the present application, the dimension size of the feature direction of this feature map is 64, and the dimension perpendicular to the feature direction is H / 4*H / 4. The descriptions of the dimensions of other feature maps are similar, and will not be elaborated here one by one.
[0100] Based on Figure 2a the network model shown, the third feature map output by the second neural network 42 is used as the feature map used in the process of image recognition. When applied to the image recognition stage, the first score map and the third feature map are used to recognize the image to be recognized.
[0101] Combined with Figure 2a the network model shown and Figure 4 the size of the third feature map output by the second neural network 412 shown, it can be known that M1*N1*P1 is equivalent to H / 4*H / 4*64. Specifically, M1 = H / 4, N1 = H / 4, and P1 = 64. In this case, the size of the first score map is H / 4*H / 4.
[0102] Based on Figure 2b the network model shown:
[0103] A network 42 is represented, which is used to obtain a first score map based on a third feature map. The size of the third feature map is M2*N2*P2, the size of the first score map is M1*N1, P2 is the size of the feature direction dimension, and M1, N1, P1, M2, N2, and P2 are all positive integers. Optionally, M1*N1 < M2*N2.
[0104] An extraction network 43 is used to perform feature extraction on the third feature map to obtain a fourth feature map. Among them, the size of the fourth feature map is M1*N1*P1; P1 is the size of the feature direction dimension, and P1 is a positive integer. That is to say, the feature extraction here is to further reduce the size of the dimension perpendicular to the feature direction of the feature map. In this way, it helps to reduce the computational complexity of the image recognition process when using the fourth feature map for image recognition later, thereby improving the recognition efficiency.
[0105] Based on Figure 2bWhen the network model shown is applied to the image recognition stage, the first score map and the fourth feature map obtained based on the image to be recognized are used to recognize the image to be recognized. All the specific examples in the following text are illustrated by taking the network model shown in 2b as an example, which is hereby uniformly explained and will not be elaborated hereinafter.
[0106] The technical solution provided by the embodiments of the present application includes a training stage and an image recognition stage, which are described separately as follows:
[0107] Training stage
[0108] The training stage includes a training data acquisition stage and a model training stage, which are described separately as follows:
[0109] a) Training data acquisition stage
[0110] As Figure 6 shown, it is a schematic flowchart of a method for acquiring training data provided by the embodiments of the present application. The execution subject of this method can be a computer device, and this method may include the following steps:
[0111] S101: Acquire a reference image set, where the reference image set includes multiple reference images; then, acquire the score map of each reference image in the multiple reference images.
[0112] The embodiments of the present application do not limit the reference image set. For example, the reference image set can be an existing data set, such as the HPatches data set, specifically, a three-dimensional reconstruction data set, etc.
[0113] Since local feature detection methods are mostly used in fields such as three-dimensional modeling, simultaneous localization and mapping (SLAM), etc., and there are few cases of high homography transformation like the picture book recognition scenario in these fields, the training data sets used in these fields usually do not contain sample images in the case of high homography transformation. Due to the high difficulty and cost of constructing the data set, in some embodiments of the present application, the existing data set is enhanced to obtain sample images applicable to the case of high homography transformation. Among them, the sample images in the case of high homography transformation include: images with a homography transformation relationship. The enhancement process can refer to S102 - S103.
[0114] The score map of an image can be characterized by a matrix. For example, the value of the element in the i-th row and j-th column of the matrix represents the probability that the pixel (or pixel block) in the i-th row and j-th column of the image is a feature point. Here, both i and j are integers greater than or equal to 0. In one example, if the reference image set is an existing data set such as the HPatches data set, the score map of the reference image in the reference image set can be the score map of the corresponding image in an existing data set such as the HPatches data set. In this way, the score map of the image in the prior art can be directly used without further calculation, which helps to reduce the computational complexity.
[0115] In some other embodiments of the present application, other methods can be used to obtain sample images suitable for the case of high homography transformation, rather than enhancing based on the existing training data set. Correspondingly, the score map of each reference image in the reference image set can also be obtained by other methods, which are not limited in the embodiments of the present application.
[0116] S102: Use multiple reference images (such as each reference image) as sample images respectively, and use the score maps of the multiple reference images as the score maps of the corresponding sample images respectively. Moreover, perform homography transformation on multiple reference images (such as each reference image) respectively to obtain multiple sample images.
[0117] Performing homography transformation on the reference image specifically includes: multiplying the reference image by a transformation matrix to obtain a sample image. Among them, the homography transformation matrix can be predefined or randomly generated. In S102, for any one reference image, multiply the reference image by one or different multiple transformation matrices to obtain one or more sample images. For one reference image, the transformation matrix corresponds one-to-one with the sample image obtained based on the reference image.
[0118] As Figure 7 shown, it is a schematic diagram of a reference image applicable to the embodiments of the present application and a sample image obtained by performing homography transformation on the reference image. Among them, Figure 7 the H in represents the transformation matrix used during homography transformation.
[0119] S103: For each sample image obtained based on homography transformation, based on the score of the reference image corresponding to the sample image (i.e., the reference image used when obtaining the sample image), and the transformation matrix corresponding to the sample image (i.e., the transformation matrix used when obtaining the sample image), obtain the score map of the sample image.
[0120] Hereinafter, taking the example of obtaining a sample image by performing homography transformation on a reference image, the method for obtaining the score map of the sample image is described:
[0121] First, mark the reference image as D, mark the transformation matrix used for the homography transformation of the reference image as H, and mark the score of the pixel point d ij (i.e., the pixel point in the i-th row and j-th column of the reference image, where both i and j are integers) as s ij . Multiply the pixel point d ij in the reference image by the homography transformation coefficient H i , and mark the obtained pixel point as . Mark the score of
[0122] as ij is mapped from d, so 's score (i.e., ) is affected by the homography transformation matrix H. When local deformation occurs in the image and the image loss is large (this loss has a correlation with the spatial rotation parameter and scaling parameter of the homography matrix, and this correlation can be calculated in the expression in b)), it causes the possibility that some feature points in the image before transformation are considered as feature points to decrease when mapped to the image after transformation. And if the scores of the pixel points with a mapping relationship remain unchanged before and after mapping, it will cause serious distortion of the sample, resulting in difficult convergence of subsequent network training. For this reason, the embodiments of the present application provide a method for estimating the score in the image obtained after transformation, which may specifically include the following steps:
[0123] Step A): Expand s ij into a matrix [s ij , 1, 1]. In order to perform normalization processing on the data, perform a normalization operation on the matrix [s ij , 1, 1] to obtain S = [a, b, c].
[0124] Step B): According to the reference image and the score map of the reference image, calculate the score map transformation matrix T = [λ1, λ2, λ3] used when the reference image is transformed into the sample image.
[0125] Specifically: According to the matching correspondence relationship between the image patches in the reference image and the image patches in the sample image, using the scores of multiple image patches in the reference image before deformation and the homography transformation matrix H as inputs, and using the scores of the image patches obtained after deformation of the multiple images as outputs, perform least squares fitting to obtain the score map transformation matrix T. Among them, if the image patches in the reference image and the image patches in the sample image physically represent the same object, then there is a matching correspondence relationship between these two image patches.
[0126] Step C): Based on the score map transformation matrix T, obtain Specifically:
[0127] If, during the transformation process, the pixel point P' on the sample image is obtained by transforming the pixel point P on the reference image, the pixel point Q' on the sample image is obtained by transforming the pixel point Q on the reference image, and P coincides with Q, then the following formula is satisfied:
[0128] where n is the number of coincident points.
[0129] If, during the transformation process, the pixel point P' on the sample image is obtained by transforming the pixel point P on the reference image, and there exists a pixel point Q in the reference image, where the pixel point Q is a pixel point within the neighborhood of the pixel point P and Q is a fitted estimated point, then the following formula is satisfied:
[0130] where n is the number of neighbors, that is, the number of pixel points within the neighborhood.
[0131] The neighborhood of the pixel point P can be predefined. The embodiments of the present application do not limit the size and position of the neighborhood of the pixel point P.
[0132] It should be noted that each score of the feature point score map is jointly constrained by the points within its neighborhood, which increases the receptive field in the constraint and at the same time compensates for the sample distortion problem during the data augmentation process, reducing the contingency of feature point selection.
[0133] Thus, the training data is obtained. The training data includes: the sample images in the sample image set and the score map of each sample image. Among them, the sample image set includes the reference image and the image obtained by performing a homography transformation on the reference image.
[0134] As Figure 8 shown, it is a schematic diagram of the relationship between the reference data and the training data provided by the embodiments of the present application. Among them, the reference data includes the reference image set and the score map of each reference image in the reference image set, Figure 8 which schematically shows that the reference image set includes reference image 1 and reference image 2. The training data includes the sample image set and the score map of each sample image in the sample image set, Figure 8 which schematically shows that the sample image set includes: sample image 10 (i.e., reference image 1), sample image 11 (i.e., the image obtained by multiplying reference image 1 by transformation matrix 11), sample image 12 (i.e., the image obtained by multiplying reference image 1 by transformation matrix 12), sample image 20 (i.e., reference image 2), and sample image 21 (i.e., the image obtained by multiplying reference image 2 by transformation matrix 21), etc. Figure 8 The double-headed arrows in indicate the corresponding relationship between the image and its score map.
[0135] For training data, two pairs of input sample images of size H*W are used, the image blocks corresponding to each image, and the feature matrix corresponding to the image blocks (1*W dimensions); the value range of the score corresponding to each image block is ([0,1]). Construct the triple tri=(D i ,D j ,D k ), where D i ,D j ,D k are all image blocks, (D i ,D j ) is a matching pair of similar image patches, (D i ,D k ) are matching pairs of dissimilar image patches, where dissimilar image matching patches are randomly selected from the same image or different images at the same scale.
[0136] In order to make the image robust at multiple scales, images of multiple scales are used as training inputs. In the embodiment of the present application, the size of the image can be adjusted to three sizes: (H*2)*(W*2), H*W, and (H / 2)*(W / 2) according to the training data. In the corresponding score map, the score map corresponding to the image of size (H*2)*(W*2) is obtained by interpolation, and the score map corresponding to the image of size (H / 2)*(W / 2) is obtained by downsampling (max pooling).
[0137] It should be noted that the training data is based on the annotation information of natural scenes to estimate the score map, which makes the estimated samples more inclined to real scenes, thus helping to improve the accuracy of image recognition.
[0138] b) Model training stage
[0139] based on Figure 2a In the network model shown, the computer device can first train the front-end network 41 and then train the representation network 42 separately.
[0140] based on Figure 2b In the network model shown, the computer device can first train the front-end network 41, and then respectively train the representation network 42 and the extraction network 43. The training of the representation network 42 and the extraction network 43 can be performed in parallel, and the training order between the two can be in no particular order.
[0141] The process of training each network (including the front-end network 41, the representation network 42, and the extraction network 43, etc.) can be considered as the process of obtaining the actual value of the parameters of the network (such as the value of each element in the convolution kernel, etc.). The actual value here refers to the value of the parameter used by the network when applied to the image recognition stage.
[0142] Train the front-end network 41
[0143] Before training the front-end network 41, the following information can be pre-configured:
[0144] The operation layers respectively included in the first neural network 411 and the second neural network 412 in the front-end network 41, the size of the input of the operation layer, the size of the parameters of the operation layer, the size of the output of the operation layer, and the association relationship between the operation layers. Among them, the operation layer may include: one or more of a convolutional layer, a pooling layer, a fully connected layer, or a dot product layer, etc. The parameters of the operation layer include the parameters used when performing the operation of this layer. For example, the parameters of the convolutional layer include the number of convolutional layers and the size of the convolutional kernel used for each convolutional layer. The association relationship between the operation layers can also be referred to as the connection relationship between the operation layers. For example, which operation layer's output is used as the input of which operation layer, etc.
[0145] It can be understood that the input of the first operation layer in the front-end network 41 is the input of the front-end network 41, and the output of the last operation layer of the front-end network 41 is the output of the front-end network 41.
[0146] The input of the front-end network 41 is an image. In one example, the size of the input of the front-end network 41 is marked as H*W*3. Among them, H represents the height size of the input image, W represents the width size of the input image, and 3 represents the number of channels. The values of H and W can be predefined.
[0147] The output of the front-end network 41 is the third feature map. The third feature map refers to the feature map obtained by rotating the features of the image input to the front-end network 41 to the main direction.
[0148] Thus, the pre-configuration process of the front-end network is completed.
[0149] After pre-configuring the front-end network 41, the input sizes, parameter sizes, and output sizes of different operation layers are adapted. Here, "adapted" means the sizes that satisfy the operation relationship between matrices / tensors in mathematics. For example, the principle for matrix A and matrix B to satisfy the dot product is that the number of columns of matrix A is equal to the number of rows of matrix B. Other examples are not listed one by one.
[0150] After the pre-configuration process ends, the computer device can configure initial values for the parameters in the front-end network 41 (for example, the parameters of each operation layer in the front-end network 41). For example, each convolutional kernel used in each convolutional layer has an initial value. The embodiments of the present application do not limit the initial values of the parameters. For example, they can be randomly generated.
[0151] The basic principle of training the front-end network 41 is as follows: Based on the images in the sample image set and the initial values of the parameters in the front-end network 41, training is performed under the constraints of the adversarial network 44 and the siamese network 45 in the front-end network 41 to achieve that "the third feature map output by the front-end network 41 is the feature map obtained by transforming the features of the input image to the main direction". And the parameters of the front-end network 41 used when achieving this purpose are used as the training results. Among them, the connection relationships among the front-end network 41, the adversarial network 44, and the siamese network 45 can be as Figure 9 shown.
[0152] The result of the training process is used as the value (or actual value) of the parameters of the front-end network during the image recognition process using the front-end network 41.
[0153] The following describes the method for training the front-end network 41 provided in the embodiments of the present application. The execution subject of this method can be a computer device. As Figure 10 shown, this method may include the following steps:
[0154] S201: Input any image in the sample image set as the input image into the first neural network 411. The first neural network 411 extracts features from the input image to obtain the first feature map of the input image.
[0155] For example, the first neural network 411 uses the initial values of the parameters of the first neural network to extract features from the input image to obtain the first feature map of the input image.
[0156] Optionally, the first neural network performs a convolution operation with a preset number of layers on the input image to obtain the first feature map of the input image. By way of example, based on Figure 3 , when executing S201, the first neural network 411 performs 4-layer convolution operations on the input image to obtain the first feature map of the input image.
[0157] It should be noted that this is only an example. In actual implementation, the first neural network 411 may also perform other operations on the input image to obtain the first feature map, and the embodiments of the present application do not limit this.
[0158] S202: Input the first feature map of the input image into the second neural network 412 to extract features from the first feature map to obtain a third feature map. The third feature map can be understood as the feature map obtained by processing the input image features (such as texture features, etc.) by the second neural network 412 and transforming them to the main direction.
[0159] For example, the second neural network 412 sequentially performs a convolution operation and a fully connected operation on the first feature map of the input image, and performs a dot product operation on the result of the fully connected operation and the first feature map of the input image to obtain a third feature map.
[0160] Exemplarily, based on Figure 4 , when executing S202, the second neural network 412 sequentially Figure 3 performs 2-layer convolution operations and 1-layer fully connected operation on the obtained first feature map, and performs a dot product operation on the result of the fully connected operation and the first feature map to obtain a third feature map. For example, based on Figure 4 , the second neural network 412 sequentially Figure 3 After performing 2-layer convolution operations and 1-layer fully connected operation on the obtained first feature map, h*w matrices of 2*2 can be obtained. Taking the kernel of 2*2*(h*w) as the direction matrix of the main direction of the corresponding channel feature, through the dot product method, the features cycled to the main direction are obtained.
[0161] Since the dot product is differentiable, the second neural network 412 can perform backpropagation during the training process of the front-end network 41. Specifically, it is constrained by the adversarial network 44 and the siamese network 45 of the front-end network 41 to train the front-end network 41 to obtain the actual values of the parameters of the front-end network 41. The working principle of the adversarial network 44 is described in step S203 below, and the working principle of the siamese network 45 is described in step S204.
[0162] As an example, the second neural network 412 can be called a local spatial transformation network (LSTN). The design of LSTN, under the learning of the generative adversarial network, enables the local area to be corrected to its main direction, so that the network can converge when training high homography transformation samples.
[0163] S203: Use the third feature map as the input of the adversarial network 44. The adversarial network 44 performs a deconvolution operation on the third feature map to obtain a fifth feature map. Among them, the size of the fifth feature map is the same as the size of the input image of the front-end network 41, such as both being H*W*3. Then, the adversarial network 44 divides the fifth feature map into multiple data blocks.
[0164] Optionally, the adversarial network 44 performs two-layer deconvolution operations on the third feature map to obtain a fifth feature map.
[0165] Optionally, the adversarial network 44 divides the fifth feature map into multiple data blocks, which may include: the adversarial network 44 evenly divides the fifth feature map into multiple data blocks. The embodiments of the present application do not limit the size of each data block.
[0166] Such as Figure 11As shown, it is a schematic diagram of the logical structure of a confrontation network 44 provided by an embodiment of the present application. Figure 11 It is drawn based on Figure 4 Specifically, based on Figure 4 The obtained third feature map (with a size of H / 4 * H / 4 * 64), the confrontation network 44 performs two-layer deconvolution operations on the third feature map to obtain a feature map with a size of H / 2 * H / 2 * 32 and a feature map with a size of H * W * 3 (i.e., the fifth feature map). Then, the elements in each matrix with a size of H * W in the feature map with a size of H * W * 3 are evenly divided into data blocks of 16 * 16.
[0167] S204: Input the multiple data blocks generated by the confrontation network 44 into the siamese network 45. The siamese network 45 uses a loss function for constraint to determine whether the third feature map is the feature map after being rotated to the main direction. Among them, the basic idea of the siamese network 45 is to minimize the feature distance between matching pairs of similar data blocks while maximizing the feature distance between pairs of dissimilar data blocks.
[0168] If so, that is, the judgment result is that the third feature map is the feature map after being rotated to the main direction, then the training process of the previous network 41 ends. Subsequently, the values of the parameters used when executing S201 and S202 this time can be used as the parameter values of the previous network during the recognition stage.
[0169] If not, that is, the judgment result is that the third feature map is not the feature map after being rotated to the main direction, then the previous network 41 can feedback relevant information to the previous network 41 to assist in adjusting the values of the parameters of the previous network 41. After the parameters of the previous network 41 are adjusted, S201 is executed again, and so on in a loop until the judgment result is that the third feature map is the feature map after being rotated to the main direction after one or more executions of S204.
[0170] The embodiment of the present application does not limit the specific implementation manner of the confrontation network 44 and the siamese network 45 assisting in adjusting the previous network 41. For example, the feedback adjustment process during the training process of the values of the parameters of the previous network 41 in other application scenarios in the prior art can be referred to, and details are not described here.
[0171] Optionally, by constructing a loss function constraint of the triplet tri, it is determined whether the third feature map is correct. The loss function of the triplet is as shown in the following formula 1, and its idea is: minimizing the feature distance between matching pairs of similar image blocks while maximizing the feature distance between pairs of dissimilar image blocks, where M is a bias value to ensure model convergence.
[0172] Formula 1: L tri (D i , D j , D k ) = ∑i,j,k∈P max(0, dist(D i , D j ) - dist(D i , D k ) + M).
[0173] It should be noted that the loss function of the triple is used simultaneously in the feature representation and the adversarial network of the main direction, which enlarges the distribution of dissimilar feature points and enables the subsequent matching to obtain the nearest neighbor features more accurately.
[0174] Training the extraction network 43
[0175] Before training the extraction network 43, the following information can be pre-configured:
[0176] The operation layers included in the extraction network 43, the size of the input of the operation layer, the size of the parameters of the operation layer, the size of the output of the operation layer, and the association relationship between the operation layers (i.e., which operation layer's output is used as the input of which operation layer, etc.). Among them, the operation layer can include: a convolutional layer, or a grouped weighted layer, etc. The parameters of the convolutional layer include the number of convolutional layers and the size of the convolutional kernel used in each convolutional layer.
[0177] In one implementation, the extraction network 43 includes a grouped weighted layer 432.
[0178] The grouped weighted layer 432 is used to: perform a convolution operation on the third feature map using a 1-channel convolutional kernel to obtain X fifth feature maps; where the size of the feature direction of the fifth feature map is smaller than the size of the feature direction of the third feature map; X is an integer greater than 2; perform weighted summation of the elements of the X fifth feature maps to obtain a sixth feature map; perform feature extraction on the sixth feature map to obtain a first score map. For the specific description of the grouped weighted layer 432 in this implementation, it can be inferred based on the following implementation, and will not be elaborated here.
[0179] In another implementation, the extraction network 43 includes a convolutional layer 431 and a grouped weighted layer 432.
[0180] The convolutional layer 431 is used to: perform a convolution operation on the third feature map to obtain a seventh feature map. In the embodiments of the present application, the number of convolutional layers and the size of the convolutional kernel for the convolution operation are not limited. Optionally, the purpose of "performing feature extraction on the third feature map to obtain a seventh feature map" is to reduce the dimensional size perpendicular to the feature direction.
[0181] The grouped weighted layer 432 is used to: perform a convolution operation on the seventh feature map using a 1-channel convolution kernel to obtain X fifth feature maps. The size of the feature direction of the fifth feature map is smaller than that of the third feature map. X is an integer greater than 2. Weighted sum the elements of the X fifth feature maps to obtain a sixth feature map. Perform feature extraction on the sixth feature map to obtain a first score map.
[0182] The 1-channel convolution kernel can be understood as a convolution kernel with a dimension size of 1 perpendicular to the feature direction and a dimension size of X in the feature direction. The size of the sixth feature map is the same as that of the fifth feature map. The purpose of performing feature extraction on the sixth feature map is to compress the dimension size of the feature direction of the sixth feature map to 1. Specifically, the grouped weighted layer 432 can perform one or more layers of convolution operations on the sixth feature map to obtain the first score map. The first score map is a two-dimensional matrix, that is, the dimension size of its feature direction is 1.
[0183] It should be noted that the design of the grouped weighted layer (or called grouped weighted network) enables the calculation of the score map to utilize both local and global information to find local features.
[0184] As Figure 12 shown, it is a schematic diagram of the logical structure of an extraction network 43 provided by an embodiment of the present application. Figure 12 is drawn based on Figure 4 Specifically: Based on Figure 4 the third feature map with a size of H / 4 * H / 4 * 64 obtained, the convolution layer 431 in the extraction network 43 is used to perform a convolution operation on the third feature map with a size of H / 4 * H / 4 * 64 to obtain a seventh feature map with a size of H / 8 * H / 8 * 256. The dimension of the feature direction of this seventh feature map is 256, and the dimension size perpendicular to the feature direction is H / 8 * H / 8. The grouped weighted layer 432 in the extraction network 43 is used to perform a convolution operation on the seventh feature map with a size of H / 8 * H / 8 * 256 using a 1*1*16 convolution kernel to obtain 16 feature maps with a size of H / 8 * H / 8 * 16 respectively. Then, weighted sum the elements of these 16 feature maps with a size of H / 8 * H / 8 * 16 to obtain a sixth feature map with a size of H / 8 * H / 8 * 16. Among them, weighted sum the elements with the same coordinate positions in different H / 8 * H / 8 * 16 feature maps to obtain the element at this coordinate position in the sixth feature map. Then, perform a convolution operation on the sixth feature map to obtain a first score map with a size of H / 8 * H / 8.
[0185] The formula for weighted summing the elements of 16 feature maps with a size of H / 8 * H / 8 * 16 is as shown in Formula 2:
[0186] Formula 2: s k= ∑ ij exp(a ij *p ij ) / ∑ k ∑ ij exp(a k,ij *p k,ij )
[0187] where s k represents the score representation for each element in each channel. Among them, in one channel, i represents the i-th group, j represents the j-th element in the i-th group, and the maximum value of k is the number of elements in a single channel; a ij represents the weight corresponding to the j-th element in the i-th group in the first channel (this weight is learned through backpropagation); a k,ij represents the weight corresponding to the j-th element in the i-th group in the k-th channel; p is the value of the corresponding element.
[0188] It should be noted that in actual implementation, during the training of the extraction network 43, feedback constraints are required by the loss function ( Figure 12 not shown in the figure), for example, the loss function of local feature extraction (i.e., the extraction network) is shown in Formula 3:
[0189] Formula 3: L score (sx, sy) = log(∑ h,w exp(l(sx hw , sy hw )))
[0190] where sy is the label, and its value is not directly obtained from the corresponding pixel position in the dataset. Instead, the scores on the score map in the corresponding n*n region (n is user-defined, and the recommended value is 9*9) are calculated, and the scores of each pixel point are obtained through Formula 2, and then the maximum value among the scores in the n*n region is taken as the score of the current point. is the score corresponding to each pixel point in the local region (the pixel points without corresponding scores are filled with a score of 0.0). sx is the score obtained through forward propagation, and sy is the reference score given by the dataset. sx hw is the score calculated for the h-th row and w-th column of the image, and sy hw is the score for the h-th row and w-th column of the image in the reference given by the dataset. The representation of Formula 3 is in the general form of a neural network loss function, using the reference data in the dataset to constrain the data calculated by the neural network, and updating each parameter in the neural network through backpropagation.
[0191] Training the representation network 42
[0192] Before training the representation network 42, the following information can be pre-configured:
[0193] The representation network 42 includes an operation layer, the size of the input of the operation layer, the size of the parameters of the operation layer, the size of the output of the operation layer, and the association relationship between the operation layers (that is, which operation layer's output is used as the input of which operation layer, etc.). Among them, the operation layer may include a convolutional layer, etc. The parameters of the convolutional layer include the number of layers of the convolutional layer and the size of the convolutional kernel used in each convolutional layer.
[0194] As Figure 13 shown, it is a schematic diagram of the logical structure of a representation network 42 provided by an embodiment of the present application. Figure 13 It is drawn based on Figure 4 Specifically: Based on Figure 4 the third feature map with a size of H / 4*H / 4*64 obtained, a convolutional layer in the representation network 42 performs a convolutional operation on the third feature map, and then outputs the output result to another convolutional layer for convolutional operation to obtain a fourth feature map. Among them, the size of the fourth feature map may be H / 8*H / 8*128. That is to say, after being processed by the representation network 42, the size in the direction perpendicular to the feature dimension is reduced. In this way, it helps to reduce the computational complexity in the subsequent image recognition process, thereby improving the image recognition efficiency.
[0195] It should be noted that in actual implementation, during the training process of the representation network 42, it is necessary to be feedback-constrained by a loss function ( Figure 13 not shown in the figure). For example, the loss function in the local feature representation stage (that is, the representation network) uses a triplet loss function, that is, constructing similar matching pairs and dissimilar matching pairs, and then using formula 1 to minimize the distance of the similar matching pairs and maximize the distance of the dissimilar matching pairs. In this stage, all channels of a single element of the feature map are extracted as features, that is, a matrix of 1*128 dimensions.
[0196] In addition, it should be noted that in actual implementation, a loss function of the overall network can be established as shown in formula 4:
[0197] Formula 4:
[0198] Among them, P represents the set of all matching image points, and p and q are points in P, and the two may be similar points or dissimilar points. The overall loss function is the sum of the loss functions of the local feature score (that is, the extraction network) and the feature representation (that is, the representation network). and They respectively represent the scores of two points A and B, and A and B are respectively taken from two images with a homography transformation relationship. The loss function of Formula 4 is a global loss calculation, which is different from the simple weighted addition of loss functions. Its purpose is to jointly act on similar and dissimilar matching pairs, cross-multiply the scores of similar pairs with similar pairs, and calculate their proportion in the global scope, thereby strengthening the constraint and enabling a better impact on the overall loss.
[0199] Image recognition stage
[0200] In the image recognition stage, the forward inference of the network includes Figure 2a or Figure 2b the network structure shown, and does not include adversarial networks and siamese networks, etc.
[0201] As Figure 14 shown, it is a schematic flowchart of an image recognition method provided by an embodiment of the present application. Figure 14 The method shown includes the following steps:
[0202] S301: The image recognition device acquires the image to be recognized. For example, the picture book robot takes pictures of the picture book to obtain the image to be recognized.
[0203] S302: The image recognition device uses the first neural network to extract features from the image to be recognized, and obtains the first feature map.
[0204] S303: The image recognition device uses the second neural network to extract features from the first feature map, obtains the second feature map, and multiplies the second feature map with the first feature map point by point to obtain the third feature map. Among them, the third feature map represents the feature map obtained after transforming the features of the image to be recognized to the main direction.
[0205] The first neural network here can be any of the trained first neural networks 411 provided above, and the second neural network can be any of the trained second neural networks 412 provided above.
[0206] In one example, the image recognition device uses the second neural network to perform at least one layer of convolution operation on the first feature map to obtain the second feature map. The specific implementation process can refer to the relevant steps executed by the computer device above.
[0207] In another example, the image recognition device uses the second neural network to perform at least one layer of convolution operation on the first feature map; performs at least one layer of pooling operation and / or fully connected operation on the first feature map after the convolution operation to obtain the second feature map. The specific implementation process can refer to the relevant steps executed by the computer device above.
[0208] S304: The image recognition device obtains a first score map of the image to be recognized based on the third feature map.
[0209] In one example, the image recognition device uses a 1-channel convolution kernel to perform a convolution operation on the third feature map, obtaining X fifth feature maps; wherein, the size of the feature direction of the fifth feature map is smaller than that of the third feature map; X is an integer greater than 2; the elements of the X fifth feature maps are weighted and summed to obtain a sixth feature map; feature extraction is performed on the sixth feature map to obtain the first score map. The specific implementation process can refer to the relevant steps executed by the computer device above.
[0210] In one example, the image recognition device performs feature extraction on the third feature map to obtain a seventh feature map; wherein, the dimension size perpendicular to the feature direction of the third feature map is greater than that of the seventh feature map; X is an integer greater than 2; a 1-channel convolution kernel is used to perform a convolution operation on the seventh feature map, obtaining X fifth feature maps; the elements of the X fifth feature maps are weighted and summed to obtain a sixth feature map; feature extraction is performed on the sixth feature map to obtain the first score map. The specific implementation process can refer to the relevant steps executed by the computer device above.
[0211] Optionally, the size of the image to be recognized is greater than that of the first score map.
[0212] S305: The image recognition device recognizes the image to be recognized based on the third feature map and the first score map.
[0213] In one example, the size of the third feature map is M1*N1*P1, and the size of the first score map is M1*N1, where P1 is the size of the feature direction dimension, M1*N1 is the size perpendicular to the feature direction dimension, and M1, N1, and P1 are all positive integers. In this case, the image recognition device directly recognizes the image to be recognized based on the third feature map and the first score map.
[0214] In another example, the size of the third feature map is M2*N2*P2, and the size of the first score map is M1*N1, where P2 is the size of the feature direction dimension, and M1, N1, P1, M2, N2, and P2 are all positive integers. In this case, the image recognition device performs feature extraction on the third feature map to obtain a fourth feature map; wherein, the size of the fourth feature map is M1*N1*P1; P1 is the size of the feature direction dimension and is a positive integer; then, the image to be recognized is recognized based on the fourth feature map and the first score map; wherein, the size of the first score map is M1*N1. Optionally, M1*N1 < M2*N2.
[0215] Examples of the specific implementation manner of S305 can refer to the specific examples in the following step S405.
[0216] The image recognition method provided by the embodiment of the present application utilizes the network described above. Since the network has rotational invariance, in image processing for an image with rotational invariance, when the image is rotated by any angle in the plane, the features extracted by the image recognition device hardly change. Therefore, the requirements for the placement and shooting of the image to be recognized are not high. In addition, compared with the prior art technical solution using a network without rotational invariance for image recognition, it helps to improve the accuracy of image recognition.
[0217] The following uses a specific example to illustrate the image recognition process provided by the embodiment of the present application.
[0218] As Figure 15 shown, it is a schematic flowchart of another image recognition method provided by the embodiment of the present application. Figure 15 The method shown includes the following steps:
[0219] S401: The image recognition device acquires two images to be matched. One of the two images is the image to be recognized, and the other image is the sample image. For example, the image recognition device can be a picture book robot.
[0220] For example, when applied to the picture book recognition process, the image to be recognized is the image captured by the picture book robot, and the sample image is a certain page of the picture book stored in the predefined picture book database.
[0221] S402: The image recognition device scales the image to be recognized to three scales, such as (0.5, 1, 2), and inputs the scaled images into the network respectively. At the same time, the sample image scaled to the same size is input into the network. Among them, the network can be the network trained in the above training stage. 0.5, 1, and 2 respectively represent the scaling multiples.
[0222] It should be noted that scaling the image to be recognized to different sizes and performing image recognition based on different sizes is an optional step. In this way, it helps to improve the accuracy of image recognition.
[0223] S403: The image recognition device uses the network to perform forward inference to obtain score maps (S1, S2) and feature maps (F1, F2) at different scales.
[0224] For example, in combination with Figure 2b , the score maps S1 and S2 can be considered as the first score map obtained by inputting the image to be recognized in S401 into the network, and the first score map obtained by inputting the sample image in S401 into the network. The feature maps F1 and F2 can be considered as the fourth feature map obtained by inputting the image to be recognized in S401 into the network, and the fourth feature map obtained by inputting the sample image in S401 into the network.
[0225] Regarding the working principle of the network in the image recognition stage, reference can be made to the training process of the network in the above text, which will not be elaborated here. It should be noted that compared with the working principle of the network in the training process, the network in the image recognition stage does not include an adversarial network, a siamese network, and may also not include a network that uses a loss function for feedback regulation.
[0226] S404: The image recognition device uses image retrieval technology (specifically, reference can be made to the prior art), and based on the score maps (S1, S2) at different scales, performs the following steps to determine the number of matching feature point pairs in F1 and F2: For the feature point f1 in F1, its corresponding score s1 in S1 > T, where T is the score threshold, and a feature point is not considered if the score is lower than this threshold. Search for the most similar feature f2 to f1 in F2. For example, take the two features with the closest Euclidean distance as the most similar features, where f2 corresponds to the score s2 in S2 > T. f1 and f2 are a pair of matching feature points.
[0227] S405: If the number of matching feature point pairs in F1 and F2 is greater than or equal to the preset threshold, the image recognition device uses the sample image used to obtain F2 as the recognition result of the image to be recognized. Otherwise, update the sample image and re - execute S401 - S405.
[0228] This embodiment provides a specific application example of an image recognition method, and the actual implementation is not limited to this.
[0229] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of the method. To implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0230] The embodiments of the present application can divide the functional modules of the image recognition device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above - integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0231] As shown Figure 16 in the figure Figure 16 shows a schematic structural diagram of the image recognition device 160 provided in an embodiment of the present application. The image recognition device 160 is used to execute the above-mentioned image recognition method. For example, it executes Figure 14 the image recognition method shown in the figure. Exemplarily, the image recognition device 160 may include a first acquisition unit 1601, a feature extraction unit 1602, a second acquisition unit 1603, and a recognition unit 1604.
[0232] The first acquisition unit 1601 is used to acquire an image to be recognized. The feature extraction unit 1602 is used to extract features from the image to be recognized using a first neural network to obtain a first feature map; and, use a second neural network to extract features from the first feature map to obtain a second feature map, and perform a dot product of the second feature map and the first feature map to obtain a third feature map; wherein, the third feature map represents the feature map obtained after transforming the features of the image to be recognized to the main direction. The second acquisition unit 1603 is used to obtain a first score map of the image to be recognized based on the third feature map. The recognition unit 1604 is used to recognize the image to be recognized based on the third feature map and the first score map.
[0233] As an example, the first neural network may be the first neural network 411 described above, and the second neural network may be the second neural network 412 described above. Combining Figure 14 this, the first acquisition unit 1601 may execute S301, the feature extraction unit 1602 may execute S302 and S303, the second acquisition unit 1603 may execute S304, and the recognition unit 1604 may execute S305.
[0234] Optionally, the feature extraction unit 1602 is specifically used for: using the second neural network to perform at least one layer of convolution operation on the first feature map to obtain a second feature map.
[0235] Optionally, the feature extraction unit 1602 is specifically used for: using the second neural network to perform at least one layer of convolution operation on the first feature map; performing at least one layer of pooling operation and / or fully connected operation on the first feature map after the convolution operation to obtain a second feature map.
[0236] Optionally, the size of the third feature map is M1*N1*P1, the size of the first score map is M1*N1, P1 is the size of the feature direction dimension, M1*N1 is the size perpendicular to the feature direction dimension, and M1, N1, and P1 are all positive integers.
[0237] Optionally, the size of the third feature map is M2*N2*P2, and the size of the first score map is M1*N1. P2 is the size of the feature direction dimension, and M1, N1, P1, M2, N2, and P2 are all positive integers. The recognition unit 1604 is specifically configured to: perform feature extraction on the third feature map to obtain a fourth feature map; wherein, the size of the fourth feature map is M1*N1*P1; P1 is the size of the feature direction dimension, and P1 is a positive integer; based on the fourth feature map and the first score map, perform recognition on the image to be recognized; wherein, the size of the first score map is M1*N1.
[0238] Optionally, M1*N1 < M2*N2.
[0239] Optionally, the second acquisition unit 1603 is specifically configured to: use a 1-channel convolution kernel to perform a convolution operation on the third feature map to obtain X fifth feature maps; wherein, the size of the feature direction of the fifth feature map is smaller than the size of the feature direction of the third feature map; X is an integer greater than 2; perform weighted summation on the elements of the X fifth feature maps to obtain a sixth feature map; perform feature extraction on the sixth feature map to obtain the first score map.
[0240] Optionally, the second acquisition unit 1603 is specifically configured to: perform feature extraction on the third feature map to obtain a seventh feature map; wherein, the dimension size perpendicular to the feature direction of the third feature map is greater than the dimension size perpendicular to the feature direction of the seventh feature map; X is an integer greater than 2; use a 1-channel convolution kernel to perform a convolution operation on the seventh feature map to obtain X fifth feature maps; perform weighted summation on the elements of the X fifth feature maps to obtain a sixth feature map; perform feature extraction on the sixth feature map to obtain the first score map.
[0241] Optionally, the size of the image to be recognized is greater than the size of the first score map.
[0242] For the specific descriptions of the above optional manners, reference may be made to the foregoing method embodiments, which will not be elaborated herein. In addition, the explanations and descriptions of the beneficial effects of any of the above-provided image recognition devices 160 may refer to the corresponding method embodiments above, and will not be elaborated.
[0243] As an example, in combination with Figure 1 , the functions implemented by the first acquisition unit 1601, the feature extraction unit 1602, the second acquisition unit 1603, and the recognition unit 1604 in the image recognition device 160 can be executed by Figure 1 the processor 101 in Figure 1 the program code in the memory 102 in
[0244] The embodiment of the present application further provides a chip system, such as Figure 17As shown, the chip system includes at least one processor 111 and at least one interface circuit 112. As an example, when the chip system 110 includes one processor and one interface circuit, the one processor can be Figure 11 the processor 111 shown by the solid line box in Figure 11 (or the processor 111 shown by the dashed line box), and the one interface circuit can be Figure 11 the interface circuit 112 shown by the solid line box in Figure 11 (or the interface circuit 112 shown by the dashed line box). When the chip system 110 includes two processors and two interface circuits, the two processors include Figure 11 the processor 111 shown by the solid line box and the processor 111 shown by the dashed line box in Figure 11 , and the two interface circuits include Figure 11 the interface circuit 112 shown by the solid line box and the interface circuit 112 shown by the dashed line box in Figure 11 . There is no limitation on this.
[0245] The processor 111 and the interface circuit 112 can be interconnected by lines. For example, the interface circuit 112 can be used to receive signals (such as receiving signals from a vehicle speed sensor or an edge service unit). For another example, the interface circuit 112 can be used to send signals to other devices (such as the processor 111). Exemplarily, the interface circuit 112 can read instructions stored in a memory and send the instructions to the processor 111. When the instructions are executed by the processor 111, the image recognition device can execute each step in the above embodiments. Of course, the chip system can also include other discrete devices, and there is no specific limitation in the embodiments of the present application.
[0246] Another embodiment of the present application also provides a computer-readable storage medium, in which instructions are stored. When the instructions run on an image recognition device, the image recognition device executes each step that the image recognition device executes in the method flow shown in the above method embodiment.
[0247] In some embodiments, the disclosed method can be implemented as computer program instructions encoded in a computer-readable storage medium in a machine-readable format or encoded in other non-transitory media or articles.
[0248] Figure 18 Schematically shows a conceptual partial view of a computer program product provided by an embodiment of the present application. The computer program product includes a computer program for executing a computer process on a computing device.
[0249] In one embodiment, the computer program product is provided using a signal-bearing medium 120. The signal-bearing medium 120 can include one or more program instructions, which when run by one or more processors can provide the functions or partial functions described above for Figure 14 description. Therefore, for example, referring toFigure 14 One or more features of S401 - S405 can be borne by one or more instructions associated with the signal - bearing medium 120. Additionally, Figure 18 The program instructions in also describe example instructions.
[0250] In some examples, the signal - bearing medium 120 can include a computer - readable medium 121, such as but not limited to, a hard - disk drive, a compact disc (CD), a digital video disc (DVD), a digital tape, a memory, a read - only memory (ROM), or a random access memory (RAM), etc.
[0251] In some embodiments, the signal - bearing medium 120 can include a computer - recordable medium 122, such as but not limited to, a memory, a read / write (R / W) CD, an R / W DVD, etc.
[0252] In some embodiments, the signal - bearing medium 120 can include a communication medium 123, such as but not limited to, a digital and / or analog communication medium (e.g., an optical fiber cable, a waveguide, a wired communication link, a wireless communication link, etc.).
[0253] The signal - bearing medium 120 can be conveyed by a wireless form of the communication medium 123 (e.g., a wireless communication medium compliant with the IEEE 802.11 standard or other transmission protocols). One or more program instructions can be, for example, computer - executable instructions or logic implementation instructions.
[0254] In some examples, such as for Figure 14 The image - recognition device described can be configured to provide various operations, functions, or actions in response to one or more program instructions via the computer - readable medium 121, the computer - recordable medium 122, and / or the communication medium 123.
[0255] It should be understood that the arrangements described herein are for illustrative purposes only. Thus, those skilled in the art will understand that other arrangements and other elements (e.g., machines, interfaces, functions, sequences, and groups of functions, etc.) can be used instead, and some elements can be omitted altogether depending on the desired results. Additionally, many of the elements described can be implemented as discrete or distributed components, or as functional entities combined with other components in any suitable combination and location.
[0256] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer execution instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more media integrated therein. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0257] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An image recognition method, applied to a high homography transformation scenario, characterized in that including: The high homography transformation scenario is a scenario with large feature differences before and after transformation; obtain the image to be recognized; use a first neural network to extract features from the image to be recognized to obtain a first feature map; use a second neural network to extract features from the first feature map to obtain a second feature map, and perform element-wise multiplication on the second feature map and the first feature map to obtain a third feature map; wherein, the third feature map represents the feature map obtained by transforming the features of the image to be recognized to the main direction; performing a first operation on the third feature map to obtain a first score map of the image to be recognized, the first operation including: convolution operation and / or feature extraction; recognize the image to be recognized based on the third feature map and the first score map.
2. The method according to claim 1, characterized in that, The step of using a second neural network to extract features from the first feature map to obtain a second feature map includes: using the second neural network to perform at least one layer of convolution operation on the first feature map to obtain the second feature map.
3. The method according to claim 1, wherein The step of using a second neural network to extract features from the first feature map to obtain a second feature map includes: using the second neural network to perform at least one layer of convolution operation on the first feature map; performing at least one layer of pooling operation and / or fully connected operation on the first feature map after the convolution operation to obtain the second feature map.
4. The method according to any one of claims 1 to 3, wherein the size of the third feature map is M1*N1*P1, the size of the first score map is M1*N1, P1 is the size of the feature direction dimension, and M1*N1 is the size perpendicular to the feature direction dimension, and M1, N1, and P1 are all positive integers.
5. The method according to any one of claims 1 to 3, characterized in that, the size of the third feature map is M2*N2*P2, the size of the first score map is M1*N1, P2 is the size of the feature direction dimension, and M1, N1, P1, M2, N2, and P2 are all positive integers; The step of recognizing the image to be recognized based on the third feature map and the first score map includes: extracting features from the third feature map to obtain a fourth feature map; wherein, the size of the fourth feature map is M1*N1*P1; P1 is the size of the feature direction dimension and P1 is a positive integer; recognize the image to be recognized based on the fourth feature map and the first score map; wherein, the size of the first score map is M1*N1.
6. The method according to claim 5, characterized in that, M1*N1 < M2*N2.
7. The method according to any one of claims 1 to 3, characterized in that The step of performing a first operation on the third feature map to obtain a first score map of the image to be recognized includes: using a 1-channel convolution kernel to perform a convolution operation on the third feature map to obtain X fifth feature maps; wherein, the size of the feature direction of the fifth feature map is smaller than the size of the feature direction of the third feature map; X is an integer greater than 2; performing weighted summation of the elements of the X fifth feature maps to obtain a sixth feature map; extracting features from the sixth feature map to obtain the first score map.
8. The method according to any one of claims 1 to 3, characterized in that The step of performing a first operation on the third feature map to obtain a first score map of the image to be recognized includes: Feature extraction is performed on the third feature map to obtain a seventh feature map; wherein, the dimension size of the third feature map perpendicular to the feature direction is greater than the dimension size of the seventh feature map perpendicular to the feature direction; X is an integer greater than 2; Using a 1-channel convolution kernel, a convolution operation is performed on the seventh feature map to obtain X fifth feature maps; The elements of the X fifth feature maps are weighted and summed to obtain a sixth feature map; Feature extraction is performed on the sixth feature map to obtain the first score map.
9. The method according to any one of claims 1 to 3, characterized in that The size of the image to be recognized is greater than the size of the first score map.
10. An image recognition device is applied to a high homography transformation scenario, and is characterized in that Includes: The high homography transformation scenario is a scenario with large feature differences before and after transformation; A first acquisition unit for acquiring an image to be recognized; A feature extraction unit for performing feature extraction on the image to be recognized using a first neural network to obtain a first feature map; and, performing feature extraction on the first feature map using a second neural network to obtain a second feature map, and performing a dot product of the second feature map and the first feature map to obtain a third feature map; wherein, the third feature map represents the feature map obtained by transforming the features of the image to be recognized to the main direction; A second acquisition unit for performing a first operation on the third feature map to obtain a first score map of the image to be recognized, the first operation including: a convolution operation and / or feature extraction; An identification unit for identifying the image to be recognized based on the third feature map and the first score map.
11. The apparatus according to claim 10, wherein The feature extraction unit is specifically configured to: perform at least one layer of convolution operation on the first feature map using the second neural network to obtain the second feature map.
12. The device according to claim 10, wherein The feature extraction unit is specifically configured to: Perform at least one layer of convolution operation on the first feature map using the second neural network; Perform at least one layer of pooling operation and / or fully connected operation on the first feature map after the convolution operation to obtain the second feature map.
13. The apparatus according to any one of claims 10 to 12, wherein The size of the third feature map is M1*N1*P1, the size of the first score map is M1*N1, P1 is the size of the feature direction dimension, M1*N1 is the size perpendicular to the feature direction dimension, and M1, N1, and P1 are all positive integers.
14. The device according to any one of claims 10 to 12, characterized in that The size of the third feature map is M2*N2*P2, the size of the first score map is M1*N1, P2 is the size of the feature direction dimension, and M1, N1, P1, M2, N2, and P2 are all positive integers; the identification unit is specifically configured to: Perform feature extraction on the third feature map to obtain a fourth feature map; wherein, the size of the fourth feature map is M1*N1*P1; P1 is the size of the feature direction dimension, and P1 is a positive integer; Based on the fourth feature map and the first score map, identify the image to be recognized; wherein, the size of the first score map is M1*N1.
15. The device according to claim 14, characterized in that, M1*N1 < M2*N2.
16. The device according to any one of claims 10 to 12, characterized in that, The second acquisition unit is specifically configured to: Perform a convolution operation on the third feature map using a 1-channel convolution kernel to obtain X fifth feature maps; wherein, the size of the feature direction of the fifth feature map is smaller than the size of the feature direction of the third feature map; X is an integer greater than 2; Weightedly sum the elements of the X fifth feature maps to obtain a sixth feature map; Perform feature extraction on the sixth feature map to obtain the first score map.
17. The device according to any one of claims 10 to 12, characterized in that, The second acquisition unit is specifically configured to: Perform feature extraction on the third feature map to obtain a seventh feature map; wherein, the dimension size perpendicular to the feature direction of the third feature map is greater than the dimension size perpendicular to the feature direction of the seventh feature map; X is an integer greater than 2; Perform a convolution operation on the seventh feature map using a 1-channel convolution kernel to obtain X fifth feature maps; Weightedly sum the elements of the X fifth feature maps to obtain a sixth feature map; Perform feature extraction on the sixth feature map to obtain the first score map.
18. The device according to any one of claims 10 to 12, characterized in that The size of the image to be recognized is larger than the size of the first score map.
19. An image recognition device, characterized in that, Comprising: A memory and a processor, the memory is used to store a computer program, and the processor is used to call the computer program to execute the method according to any one of claims 1-9.
20. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program runs on a computer, the computer is caused to execute the method according to any one of claims 1-9.
Citation Information
Patent Citations
Spatial target ISAR image processing method for template identification
CN105678311A
Visual identification method, device and system and storage medium
CN109978077A