Image recognition method, device, computer equipment, storage medium and product
By performing key feature extraction and correlation information processing on images, the problem of difficulty in feature extraction in complex image recognition is solved, and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202111419426.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-11-26
AI Technical Summary
In the process of image recognition, especially map element recognition in complex scenarios, the prior art has a high recognition error rate due to insufficient feature extraction.
By obtaining the image to be identified, key image features are extracted, key feature information of the key object is determined, and information joint processing is performed based on the associated information, and image recognition is finally performed.
Improve the accuracy of image recognition and avoid feature extraction difficulties and recognition errors caused by excessive complexity of image content.
Smart Images

Figure CN114332599B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to an image recognition method, apparatus, computer equipment, storage medium and product. Background Art
[0002] In the process of identifying the image content, it is necessary to classify and process the image according to the feature information corresponding to the image to determine the image content contained in the image. When the classification categories are numerous and complex, the image content recognition may be wrong due to inaccurate feature extraction. For example, in the process of identifying map elements, due to the large number of map elements, such as map elements containing multiple types of signboards, and each type of signboard contains different content in different scenarios, there will be a large number of recognition type errors due to the large number of signboards and the complex content contained. Summary of the invention
[0003] Embodiments of the present application provide an image recognition method, apparatus, computer device, storage medium, and product, which can improve the accuracy of image recognition.
[0004] An image recognition method provided in an embodiment of the present application includes:
[0005] Obtain an image to be recognized;
[0006] Extract key image features from the image to be identified to obtain key feature information of key objects contained in the image to be identified;
[0007] Determining the associated information corresponding to the key object according to the key feature information of the key object;
[0008] Performing information joint processing on the associated information of the key object to obtain associated feature information of the image to be identified;
[0009] Image recognition is performed on the image to be recognized based on the associated feature information to obtain image information of the image to be recognized.
[0010] Accordingly, an embodiment of the present application further provides an image recognition device, comprising:
[0011] An image acquisition unit, used for acquiring an image to be recognized;
[0012] A feature extraction unit, used to extract key image features from the image to be identified, and obtain key feature information of key objects contained in the image to be identified;
[0013] An information determination unit, configured to determine the associated information corresponding to the key object according to the key feature information of the key object;
[0014] An information combination unit, used for performing information combination processing on the associated information of the key object to obtain associated feature information of the image to be identified;
[0015] The image recognition unit is used to perform image recognition on the image to be recognized based on the associated feature information to obtain content information of the image to be recognized.
[0016] In one embodiment, the key feature information includes at least two sub-key feature information, and the feature extraction unit includes:
[0017] A feature extraction subunit, used to extract key image features from the image to be identified, and obtain at least two sub-key feature information of the key object;
[0018] The feature fusion subunit is used to perform feature fusion processing on at least two key feature information corresponding to the key object to obtain the key feature information of the key object.
[0019] In one embodiment, the at least two sub-key feature information include object feature information, location feature information and category feature information, and the feature extraction subunit includes:
[0020] A feature extraction module is used to extract key image features from the image to be identified, and obtain object feature information and position feature information corresponding to each key object in the image to be identified;
[0021] The feature information determination module is used to determine the category feature information of the key object based on the object feature information of each key object.
[0022] In one embodiment, the feature information determination module includes:
[0023] An information acquisition submodule is used to obtain initial category feature information;
[0024] A category determination submodule, used to determine the target object category corresponding to the key object based on the object feature information of each key object;
[0025] The selection submodule is used to select the initial category feature information according to the target object category to obtain the category feature information.
[0026] In one embodiment, the information determination unit includes:
[0027] A generating subunit, configured to generate an object network diagram according to the key feature information of the key object, wherein the object network diagram includes object nodes corresponding to the key feature information;
[0028] The association information determination subunit is used to determine the association information of the key object according to the object node corresponding to the key feature information in the object network diagram.
[0029] In one embodiment, the image recognition device further includes:
[0030] A first acquisition unit, used to acquire an initial image to be recognized, wherein the initial image to be recognized contains recognition elements;
[0031] A detection unit, used to detect identification elements on the initial image to be identified, and determine the position information of the identification elements in the initial image to be identified;
[0032] The second acquisition unit is used to acquire the image to be recognized from the initial image to be recognized according to the position information of the recognition element.
[0033] In one embodiment, the detection unit comprises:
[0034] An image feature extraction subunit, configured to extract image features from the initial image to be recognized, and obtain at least one feature point contained in the initial recognition image;
[0035] The position information determination subunit is used to determine the position information of the recognition element in the initial image to be recognized based on the feature points and a preset detection frame.
[0036] Correspondingly, an embodiment of the present application also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute any image recognition method provided in the embodiment of the present application.
[0037] Correspondingly, an embodiment of the present application also provides a computer-readable storage medium, which is used to store a computer program, and the computer program is loaded by a processor to execute any image recognition method provided in the embodiment of the present application.
[0038] The embodiments of the present application acquire an image to be identified; extract key image features of the image to be identified to obtain key feature information of key objects contained in the image to be identified; determine associated information corresponding to the key objects based on the key feature information of the key objects; perform information joint processing on the associated information of the key objects to obtain associated feature information of the image to be identified; and perform image recognition on the image to be identified based on the associated feature information to obtain image information of the image to be identified.
[0039] This solution extracts key image features from the image to be identified to obtain key feature information corresponding to key objects, and performs image recognition on the image to be identified based on the key feature information corresponding to the key objects, so as to perform image recognition based on the key objects contained in the image to be identified, thereby avoiding image recognition errors caused by difficulty in feature extraction due to the overly complex content contained in the image to be identified when performing image recognition based on the complete image, thereby improving the accuracy of image recognition on the image to be identified. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0041] Figure 1 is a scene graph of the image recognition method provided in an embodiment of the present application;
[0042] Figure 2 is a flow chart of an image recognition method provided by an embodiment of the present application;
[0043] Figure 3 is another flow chart of the image recognition method provided by an embodiment of the present application;
[0044] Figure 4 It is a schematic diagram of the identification element categories provided in the embodiment of the present application;
[0045] Figure 5 is a schematic diagram of key object categories provided in an embodiment of the present application;
[0046] Figure 6 is a schematic diagram of a candidate detection frame provided in an embodiment of the present application;
[0047] Figure 7 is a schematic diagram of an image recognition device provided in an embodiment of the present application;
[0048] Figure 8 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0050] The present application provides an image recognition method, apparatus, computer device and computer readable storage medium. The image recognition apparatus can be integrated in a computer device, which can be a server or a terminal.
[0051] The terminal may include a mobile phone, a wearable smart device, a tablet computer, a laptop computer, a personal computer (PC), an intelligent voice interaction device, and a vehicle-mounted computer.
[0052] Among them, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0053] For example, Figure 1 As shown, a computer device obtains an image to be identified, performs key image feature extraction on the image to be identified, obtains key feature information of at least one key object contained in the image to be identified, determines the associated information of each key object according to the associated relationship between the key feature information corresponding to each key object, performs information joint processing on the associated information of each key object, obtains the associated feature information of the image to be identified, and performs image recognition on the image to be identified based on the associated feature information to obtain image information of the image to be identified. This solution extracts key image features from the image to be identified to obtain key feature information corresponding to the key object, and performs image recognition on the image to be identified based on the key feature information corresponding to the key object, so as to perform image recognition based on the key object contained in the image to be identified, thereby avoiding image recognition errors caused by difficulty in feature extraction due to overly complex content contained in the image to be identified when performing image recognition based on the complete image, thereby improving the accuracy of image recognition of the image to be identified.
[0054] It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments.
[0055] This embodiment will be described from the perspective of an image recognition device. The image recognition device may be integrated into a computer device, which may be a server or a terminal.
[0056] An image recognition method provided in an embodiment of the present application is as follows: Figure 2 As shown, the specific process of the image recognition method can be as follows:
[0057] 101. Obtain an image to be identified.
[0058] For example, the image to be identified may be an image for image recognition, and the image to be identified may be obtained from the cloud, a database, or a blockchain.
[0059] The image to be identified may be an image cropped from an initial image to be identified. By cropping the image to be identified from the initial image to be identified, the problem of large amount of calculation caused by large image size when extracting features from the image can be reduced, and the key features of the key object can be more easily learned when extracting key feature information. That is, in one embodiment, before the step of "obtaining the image to be identified", the image recognition method provided in the embodiment of the present application may further include:
[0060] Acquire an initial image to be identified, where the initial image to be identified contains identification elements;
[0061] Performing identification element detection on the initial image to be identified, and determining the position information of the identification element in the initial image to be identified;
[0062] The image to be recognized is obtained from the initial image to be recognized according to the position information of the recognition element.
[0063] Among them, the initial image to be identified may be an image containing identification elements, and the identification elements may include the content to be identified by image recognition. For example, when the application scenario is road sign recognition, the identification element may be a road sign in the initial image to be identified, and the image to be identified is the image area where the identification element is located in the initial image to be identified.
[0064] For example, the initial image to be identified can be obtained from the cloud, database or blockchain. The image to be identified can also be obtained from different channels according to the application scenario. For example, when the application scenario is to perform image recognition on road signs, the initial image to be identified can be obtained by collecting on-board cameras and other devices.
[0065] The initial image to be identified is detected for identification elements through a neural network. For example, the image area to be identified that is similar to the template image is identified from the image to be identified through template matching. Template matching refers to finding the part of the initial image to be identified that is similar to the template image by comparing the template image with the initial image to be identified. The template image can be flexibly set according to the characteristics of the identification elements in the actual application scenario. The part similar to the template image is the image area where the identification elements are located. The identification elements are detected from the image to be identified and the position information of the location of the identification elements is determined. The image to be identified is captured from the initial image to be identified based on the position information.
[0066] Determining the identification elements from the initial image to be identified may also be determined based on feature points and candidate detection frames. Specifically, the step of “detecting image elements on the initial image to be identified to determine the position information of the image elements in the initial image to be identified” may include:
[0067] Extracting image features from the initial image to be recognized to obtain at least one feature point contained in the initial recognition image;
[0068] Based on the feature points and the preset detection frame, the position information of the recognition element in the initial image to be recognized is determined.
[0069] The preset detection frame may be used to determine the image region where the image to be identified is located from the initial image to be identified, and the preset detection frame may be a square, rectangle, circle, triangle or other shapes.
[0070] For example, the image features of the initial image to be identified can be extracted through a neural network model. The network structure of the neural network model may include a convolution layer, a normalization layer (Batch Normalization, BN), and an activation layer. Specifically, the convolution layer in the neural network model can be used to extract image features such as edges and textures of the initial image to be identified to obtain the image features of the initial image to be identified. The normalization layer in the neural network model normalizes the image features extracted by the convolution layer according to a normal distribution to filter out noise features in the image features. The activation layer performs nonlinear mapping on the normalized image features to obtain feature points in the initial image to be identified.
[0071] With each feature point as the center point and the preset detection frame, an image area is cropped from the initial image to be identified as the image to be identified. Optionally, multiple preset detection frames can be set. For example, the aspect ratios of the preset detection frames are {1:1, 2:1, 1:2}, respectively. With each feature point as the midpoint, a preset detection frame of one feature point (three ratios), a preset detection frame of two feature points (three ratios), and a preset detection frame of three feature points (three ratios) are included as candidate detection frames. It can be understood that the ratio of the preset detection frame can be flexibly set as needed. For example, the aspect ratio of the preset detection frame can be set to 4:1, or it can be a circle or other shapes, and the number of preset detection frames is adjusted according to the scene recognition accuracy, which is not limited here.
[0072] The image area obtained by the nine candidate detection frames is matched with each recognition element in the preset recognition element set, for example, the shape of the object contained in the candidate detection frame is matched with the shape of the recognition element, etc., to obtain a matching score for each candidate detection frame. For example, edge features are extracted from the image area in the candidate detection frame, the shape of the object contained in the candidate detection frame is determined, and the shape is matched with the shape of each recognition element in the preset recognition element set to obtain a matching score. The matching score can represent the degree of similarity between the object contained in the candidate detection frame and the recognition element in the preset recognition element set.
[0073] The image area where the candidate detection box with the highest matching score is located is captured as the image to be recognized.
[0074] 102. Extract key image features from the image to be identified to obtain key feature information of key objects contained in the image to be identified.
[0075] The key feature information may include feature information of the key object, and the key object may be mapped into the feature space through the key feature information.
[0076] For example, the image features such as edges and textures of the image to be identified may be specifically extracted to obtain a feature map of the image to be identified, key object detection may be performed based on the feature map corresponding to the image to be identified, and the key object contained in the image to be identified may be determined from the feature map. When the image to be identified contains multiple key objects, multiple key objects may be determined from the image to be identified. Based on the image area of the key object in the image to be identified, the feature information of the corresponding position in the feature map may be determined as the key feature information of the key object.
[0077] Optionally, the key object detection from the image to be identified can also refer to the method of detecting identification elements of the initial image to be identified, so as to determine the key object from the image to be identified. The specific implementation process is described in detail above and will not be repeated here.
[0078] Optionally, after determining the image region where the key object is located in the image to be identified, object features of the image region are extracted to obtain key feature information that can characterize the key object.
[0079] Optionally, other features such as position features and category features of the key object may also be extracted to determine key feature information corresponding to the key object based on multiple feature information, that is, the step of "extracting key image features from the image to be identified to obtain key feature information of the key object contained in the image to be identified" may specifically include:
[0080] Extract key image features from the image to be identified to obtain at least two sub-key feature information of the key object;
[0081] Feature fusion processing is performed on at least two key feature information corresponding to the key object to obtain the key feature information of the key object.
[0082] The sub-key feature information may include information indicating key object features. For example, the sub-key feature information may be object feature information, location feature information, or category feature information.
[0083] The object feature information can represent the semantic information of the key object. According to the object feature information, the information indicated by the key object can be determined. For example, according to the object feature information corresponding to the key object in the road sign, it can be determined that the key object indicates a left turn at the next intersection.
[0084] The position feature information may represent the position of the key object in the image to be identified, as well as information such as the size of the key object. The size may include the size of the candidate detection box determined when detecting the key object.
[0085] The category characteristic information may characterize the category to which the key object belongs. For example, multiple categories may be preset according to the information conveyed by different key objects, and the category characteristic information may indicate that the key object belongs to a corresponding category among the preset multiple categories.
[0086] For example, in one embodiment, key image features such as object features and position features of the image to be identified can be extracted to obtain object feature information and position feature information corresponding to the key object, and the object feature information and position feature information can be feature fused to obtain key feature information of the key object.
[0087] In one embodiment, key image features such as object features and category features of the image to be identified can be extracted to obtain object feature information and category feature information corresponding to the key object, and the object feature information and category feature information can be feature fused to obtain key feature information of the key object.
[0088] In one embodiment, key image features such as object features, position features, and category features of the image to be identified may be extracted to obtain object feature information, position feature information, and category feature information corresponding to the key object, and feature fusion processing is performed on the object feature information, the position feature information, and the category feature information to obtain key feature information of the key object, that is, the step of "extracting key image features from the image to be identified to obtain at least two key feature information corresponding to each key object in the image to be identified" may specifically include:
[0089] Extract key image features from the image to be identified, and obtain object feature information and position feature information corresponding to each key object in the image to be identified;
[0090] Based on the object characteristic information of each key object, category characteristic information of the key object is determined.
[0091] For example, key image features are extracted from the image to be identified to obtain object feature information and position feature information of each key object in the object to be identified. Since the object feature information represents the semantic information of the key object, the object category corresponding to the key object can be determined based on the object feature information, and the corresponding category feature information can be determined based on the object category to which the key object belongs.
[0092] The category feature information can be obtained based on the initial category feature information and the object category of the key object, so that the category feature information corresponding to different key objects is unified in form, thereby improving the processing speed of the category feature information and thus improving the image recognition speed. That is, in one embodiment, the step of "determining the category feature information of the key object based on the object feature information of each key object" can specifically include:
[0093] Initial category feature information;
[0094] Based on the object feature information of each key object, determine the target object category corresponding to the key object;
[0095] The initial category feature information is selected and processed according to the target object category to obtain category feature information.
[0096] The initial category feature information can be obtained based on at least one object category included in the object category set. The object category set includes multiple object categories. The initial category feature information is generated based on the number of object categories included in the object category set. The dimension of the initial category feature information is the same as the number of categories. The multiple object categories in the object category set are sorted according to a preset rule, and each object category corresponds to a dimension of data in the initial category feature information based on the sorting order. For example, the object category set includes 5 object categories, which are ABCDE after sorting. Then, the corresponding initial category feature information is (0,0,0,0,0).
[0097] For example, since the object feature information represents the semantic information of the key object, the target object category corresponding to the key object can be determined based on the object feature information. The value of the dimension corresponding to the target object category in the initial category feature information is 1 to obtain the category feature information of the key object. For example, if the target object category of the key object is E, the category feature information is (0, 0, 0, 0, 1). Optionally, the category feature information can be dimensionalized and other operations can be performed based on information such as the dimensions of other sub-key feature information, so that different sub-key feature information are consistent in size to facilitate calculation.
[0098] The position of the key object in the image to be identified can be determined through the position feature information, and the category of the key object can be determined based on the category feature information. The role of the key object in the image to be identified can be better determined based on the category of the key object and its position in the image to be identified, so as to better determine the image information of the image to be identified.
[0099] When the image to be identified contains multiple key objects, the position feature information of the key objects can also be obtained to determine the positional relationship between different key objects in the image to be identified based on the position feature information. The image to be identified can be identified more accurately based on the positional relationship between key objects. The category feature information of the key objects can also be obtained to better determine the relationship between key objects based on the category feature information, thereby improving the accuracy of image recognition.
[0100] The embodiment of the present application converts the image recognition task of the image to be recognized into the recognition of smaller elements in the image to be recognized, namely the key objects, by extracting the key feature information of the key objects contained in the image to be recognized. This can avoid image recognition errors or failure to recognize due to inaccurate feature extraction caused by the difficulty of feature extraction when the content of the image to be recognized is too complex and the feature extraction is based on the entire image. Moreover, the expression form of the recognition elements of the same category is not fixed. For example, the text content, arrow direction, color, and symbols contained therein will change, which will result in the need to obtain samples of different expressions for training, which makes the training difficult and has poor application effects.
[0101] Converting image recognition into recognition of key objects can not only reduce the difficulty of feature extraction, but also when the recognition elements of the same category change, the category of the key object can be determined through the key feature information of multiple key objects in the image to be identified, and the association relationship between the key objects can be predicted. The category of the image to be identified can still be identified. Compared with recognition based on the entire image, the image recognition method provided in the embodiment of the present application improves the accuracy of image recognition and has flexibility.
[0102] 103. Determine the associated information corresponding to the key object based on the key feature information of the key object.
[0103] The association information may include the association relationship between the key object and other key objects. When there is only one key object in the image to be identified, the association information corresponding to the key object is 0 or empty.
[0104] For example, when there is only one key object, other unrelated key feature information is determined based on the key feature information, and the associated information corresponding to the key object is determined to be 0 or empty; when there are at least two key objects, the associated information between each two key objects is determined based on the key feature information of each key object. When the associated information is 0 or empty, it indicates that there is no associated relationship between the two key objects, otherwise, there is an associated relationship.
[0105] Optionally, the associated information of the key objects may be determined by using a Graph Neural Network (GNN), that is, the step of “determining the associated information of the key objects according to the key feature information of the key objects” may specifically include:
[0106] Generate an object network diagram according to the key feature information of the key object, wherein the object network diagram includes object nodes corresponding to the key feature information;
[0107] According to the object nodes corresponding to the key feature information in the object network diagram, the associated information of the key objects is determined.
[0108] The object network diagram may be a neural network diagram generated based on key feature information, and the object network diagram includes nodes generated based on key feature information, and each node corresponds to a key object.
[0109] For example, the key feature information of key objects can be mapped to nodes on a network graph to obtain an object graph network, and the association relationship between key objects can be predicted based on the object nodes in the object graph network through a graph neural network to obtain the association information between key objects.
[0110] 104. Perform information joint processing on the associated information of the key object to obtain associated feature information of the image to be identified.
[0111] The associated feature information may include key feature information of key objects in the image to be identified, and associated information between key objects.
[0112] For example, if there is only one key object, after information joint processing is performed on the associated information of the key object, the associated characteristic information obtained is the associated information of the key object.
[0113] When there are n (n>1) key objects, there are n (n-1) pairs of association relationships. The n (n-1) pairs of relationships in the association information of the key objects are spliced to obtain the association feature information of the image to be identified.
[0114] In addition to the concatenation method, key information can also be processed jointly by feature fusion, for example, by performing bitwise sum or bitwise multiplication calculations on the related information to process the related information jointly.
[0115] 105. Perform image recognition on the image to be identified based on the associated feature information to obtain image information of the image to be identified.
[0116] The image information may be the image category of the image to be identified, or may be the content information conveyed by the image to be identified.
[0117] For example, the image to be recognized may be recognized based on associated feature information through a classification network, and the image category of the image to be recognized may be determined to determine the image information of the image to be recognized.
[0118] As can be seen from the above, the embodiments of the present application obtain an image to be identified; extract key image features of the image to be identified to obtain key feature information of key objects contained in the image to be identified; determine the associated information corresponding to the key objects based on the key feature information of the key objects; perform information joint processing on the associated information of the key objects to obtain associated feature information of the image to be identified; and perform image recognition on the image to be identified based on the associated feature information to obtain image information of the image to be identified.
[0119] This solution extracts key image features from the image to be identified to obtain key feature information corresponding to key objects, and performs image recognition on the image to be identified based on the key feature information corresponding to the key objects, so as to perform image recognition based on the key objects contained in the image to be identified, thereby avoiding image recognition errors caused by difficulty in feature extraction due to the overly complex content contained in the image to be identified when performing image recognition based on the complete image, thereby improving the accuracy of image recognition on the image to be identified.
[0120] Based on the above embodiments, further detailed description will be given below with examples.
[0121] This embodiment will be described from the perspective of an image recognition device and an image recognition method applied to a road sign recognition scenario. The image recognition device may be integrated into a computer device, which may be a terminal or other device.
[0122] An image recognition method provided in an embodiment of the present application is as follows: Figure 3 As shown, the specific process of the image recognition method can be as follows:
[0123] 201. The terminal obtains an initial image to be recognized.
[0124] For example, it can be an initial image to be identified that is collected by the terminal through a vehicle-mounted camera or other device. The initial image to be identified may contain road signs to be identified, such as speed limit signs, traffic restriction signs, etc. The road signs in the initial image to be identified are the identification elements.
[0125] During the map road data collection process, problems such as poor image quality and uneven coverage of map elements may occur. The training samples for the deep learning classification network are not of high quality and the number of samples is insufficient, resulting in a large number of false detections and recognition type errors.
[0126] Optionally, the classification categories of the identification elements may include Figure 4 There are multiple categories shown in the figure: road name-sign, sign vehicle information, direction name, and road name. It can be seen from the figure that each category contains complex content. For example, category a: road name-sign contains information such as arrows, symbols, and text. The modality of the content information contained in each road sign category is not fixed. It is very difficult to extract features for classification through a simple convolutional neural network, resulting in inaccurate classification results.
[0127] The image recognition method provided in the embodiment of the present application takes these complex content information as key objects, extracts key image features, and performs subsequent information joint processing to predict categories, thereby improving the recognition ability of the image to be recognized, thereby effectively improving the accuracy of road sign recognition.
[0128] Among them, the categories of key objects can include Figure 5 The categories shown in Figure 5 It includes multiple categories such as going straight, turning left, turning right, U-turn, and crosswalk. Among them, the category with serial number 20 can include text information in addition to other arrow forms.
[0129] 202. The server obtains an image to be recognized from the initial image to be recognized according to the position information of the recognition element contained in the initial image to be recognized.
[0130] For example, the server may extract image features of the initial image to be identified through a neural network model. Specifically, the image features such as edges and textures of the initial image to be identified may be extracted through the convolution layer in the neural network model to obtain the image features of the initial image to be identified. The image features extracted by the convolution layer are normalized according to the normal distribution by the normalization layer (BatchNormalization, BN) in the neural network model to filter out the noise features in the image features, so that the training convergence of the neural network model in the training process is faster. The normalized image features are nonlinearly mapped through the activation layer to enhance the generalization ability of the neural network model and obtain the feature points in the initial image to be identified.
[0131] The server uses each feature point as the center point and the preset detection frame to crop an image area from the initial image to be identified as the image to be identified. Optionally, multiple preset detection frames can be set. For example, the aspect ratios of the preset detection frames are {1:1, 2:1, 1:2} respectively. Figure 6 As shown, with one feature point as the midpoint, different line types are preset detection boxes containing different numbers of feature points, for example, a preset detection box containing one feature point (three ratios), a preset detection box with two feature points (three ratios), and a preset detection box with three feature points (three ratios) are used as candidate detection boxes.
[0132] The server matches the image area obtained from the nine candidate detection frames with each recognition element in the preset recognition element set, for example, matching the shape of the object contained in the candidate detection frame with the shape of the recognition element, etc., to obtain a matching score for each candidate detection frame.
[0133] The server intercepts the position information of the image area where the candidate detection frame with the highest matching score is located, determines it as the position information of the recognition element, and obtains the image to be recognized from the initial image to be recognized based on the position information.
[0134] Among them, the convolution layer includes a convolutional neural network. The convolutional neural network is a type of feedforward neural network (Feedforward Neural Networks) that includes convolution calculations and has a deep structure. It is one of the representative algorithms of deep learning. The convolutional neural network can be a Resnet convolutional neural network.
[0135] When the initial image to be recognized contains multiple recognition elements, multiple images to be recognized can be obtained from the initial image to be recognized.
[0136] 203. The server extracts key image features from the image to be identified, and obtains at least two sub-key feature information of a key object.
[0137] For example, the server may extract image features such as edge texture of the image to be identified to obtain a feature map of the image to be identified, perform key object detection based on the feature map corresponding to the image to be identified, and determine the key objects contained in the image to be identified from the feature map to be identified. When the image to be identified contains multiple key objects, the server may determine multiple key objects from the image to be identified.
[0138] According to the image area of the key object in the image to be identified, the key image features of the image area are extracted to obtain the object feature information and position feature information of each key object in the object to be identified. Since the object feature information represents the semantic information of the key object, the server can perform classification processing according to the object feature information through the classification network to determine the target object category corresponding to the key object, and set the value of the dimension corresponding to the target object category in the initial category feature information to 1 to obtain the category feature information of the key object. For example, if the target object category of the key object is E, the category feature information is (0, 0, 0, 0, 1).
[0139] The position of the key object in the image to be identified can be determined through the position feature information, and the category of the key object can be determined based on the category feature information. The role of the key object in the image to be identified can be better determined based on the category of the key object and its position in the image to be identified, so as to better determine the image information of the image to be identified.
[0140] When the image to be identified contains multiple key objects, the position feature information of the key objects can also be obtained to determine the positional relationship between different key objects in the image to be identified based on the position feature information. The image to be identified can be identified more accurately based on the positional relationship between key objects. The category feature information of the key objects can also be obtained to better determine the relationship between key objects based on the category feature information, thereby improving the accuracy of image recognition.
[0141] 204. The server performs feature fusion processing on at least two corresponding sub-key feature information for each key object to obtain key feature information of the key object.
[0142] For example, the server may perform feature fusion processing on the object feature information, the location feature information and the category feature information. For example, the object feature information, the location feature information and the category feature information may be bitwise ANDed to perform feature fusion processing on at least two sub-key feature information to obtain key feature information of the key object.
[0143] 205. The server determines an association relationship between at least one key object according to the key feature information of each key object.
[0144] For example, the server may map the key feature information of the key objects to the nodes on the network graph to obtain the object graph network, and predict the association relationship between the key objects based on the object nodes in the object graph network through the graph neural network to obtain the association relationship between the key objects. When the association relationship R between every two object nodes is 0 When it is 0, it means there is no relationship between the two key objects, otherwise, there is an association relationship.
[0145] 206. The server performs information joint processing on the association relationship between key objects to obtain the association feature information of the image to be identified.
[0146] For example, if there is only one key object, after the server performs information joint processing on the association relationship of the key object, the obtained association feature information is the association information of the key object.
[0147] When there are n (n>1) key objects, there are n (n-1) pairs of association relationships. The server concatenates the n (n-1) pairs of relationships in the association information of the key objects to obtain the association feature information of the image to be identified. For example, there is an association relationship R between every two object nodes. 0 , R 0 ∈512, concatenate n(n-1) pairs of association relations to obtain the association feature information R, R∈512×n(n-1).
[0148] In addition to the splicing method, the server can also process the information relationship of key information in a feature fusion method, for example, performing bitwise sum calculation or bitwise multiplication calculation on the association relationship to perform information joint processing on the association relationship.
[0149] 207. The server performs image recognition on the image to be recognized based on the associated feature information, obtains image information of the image to be recognized, and sends the image information to the terminal.
[0150] For example, the server may transform the associated feature information into 1×n dimensions through dimensionality conversion, where n represents the total number of road sign categories. The road sign categories are shown in the figure, and the category corresponding to the maximum value of the n numbers is selected as the final predicted image category to determine the image information of the image to be identified.
[0151] The server sends the image information obtained by image recognition to the terminal.
[0152] 208. The terminal performs display based on the image information.
[0153] For example, the terminal may specifically receive image information sent by the server and display the image information in an initial image to be identified acquired by the terminal, so that the user can obtain the type of the road sign in the initial image to be identified and obtain the type of the road sign displayed in the display device based on the image information.
[0154] Optionally, the server may determine the types of road signs included in the road based on the image information, so as to perform road planning or autonomous driving and other transportation fields and applications in the transportation field according to the road information.
[0155] The image recognition method provided in the embodiment of the present application can be applied to the field of artificial intelligence. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.
[0156] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0157] Among them, Computer Vision (CV) is a science that studies how to make machines "see". To put it more specifically, it refers to machine vision that uses cameras and computers to replace human eyes to identify, track and measure targets, and further performs graphic processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.
[0158] As can be seen from the above, the embodiment of the present application obtains an initial image to be identified through a terminal; the server obtains the image to be identified from the initial image to be identified based on the position information of the identification elements contained in the initial image to be identified; the server extracts key image features of the image to be identified to obtain at least two sub-key feature information of the key object; the server performs feature fusion processing on the corresponding at least two sub-key feature information for each key object to obtain the key feature information of the key object; the server determines the association relationship between at least one key object based on the key feature information of each key object; the server performs information joint processing on the association relationship between the key objects to obtain the associated feature information of the image to be identified; the server performs image recognition on the image to be identified based on the associated feature information to obtain image information of the image to be identified, and sends the image information to the terminal; the terminal displays based on the image information.
[0159] This solution extracts key image features from the image to be identified to obtain key feature information corresponding to key objects, and performs image recognition on the image to be identified based on the key feature information corresponding to the key objects, so as to perform image recognition based on the key objects contained in the image to be identified, thereby avoiding image recognition errors caused by difficulty in feature extraction due to the overly complex content contained in the image to be identified when performing image recognition based on the complete image, thereby improving the accuracy of image recognition on the image to be identified.
[0160] In order to facilitate better implementation of the image recognition method provided in the embodiment of the present application, an image recognition device is also provided in one embodiment. The meanings of the terms are the same as those in the above-mentioned image recognition method, and the specific implementation details can refer to the description in the method embodiment.
[0161] The image recognition device can be integrated into a computer device, such as Figure 7 As shown, the image recognition device may include: an image acquisition unit 301, a feature extraction unit 302, an information determination unit 303, an information combination unit 304 and an image recognition unit 305, which are specifically as follows:
[0162] (1) Image acquisition unit 301: used to acquire the image to be recognized.
[0163] (2) Feature extraction unit 302: used to extract key image features from the image to be identified, and obtain key feature information of key objects contained in the image to be identified.
[0164] In one embodiment, the key feature information includes at least two sub-key feature information, and the feature extraction unit 302 includes a feature extraction sub-unit and a feature fusion sub-unit, specifically:
[0165] Feature extraction subunit: used to extract key image features from the image to be identified, and obtain at least two sub-key feature information of the key object;
[0166] Feature fusion subunit: used for performing feature fusion processing on at least two key feature information corresponding to the key object to obtain the key feature information of the key object.
[0167] In one embodiment, the at least two sub-key feature information include object feature information, location feature information and category feature information, and the feature extraction subunit includes a feature extraction module and a feature information determination module, specifically:
[0168] Feature extraction module: used to extract key image features from the image to be identified, and obtain object feature information and position feature information corresponding to each key object in the image to be identified;
[0169] Feature information determination module: used to determine the category feature information of the key object based on the object feature information of each key object.
[0170] In one embodiment, the feature information determination module includes an information acquisition submodule, a category determination submodule and a selection submodule, specifically:
[0171] Information acquisition submodule: used to obtain initial category feature information;
[0172] Category determination submodule: used to determine the target object category corresponding to the key object based on the object feature information of each key object;
[0173] Selection submodule: used to select and process the initial category feature information according to the target object category to obtain category feature information.
[0174] (3) Information determination unit 303: used to determine the associated information corresponding to the key object according to the key feature information of the key object.
[0175] In one embodiment, the information determination unit 303 includes a generation subunit and an associated information determination subunit, specifically:
[0176] A generating subunit, used to generate an object network diagram according to the key feature information of the key object, wherein the object network diagram includes object nodes corresponding to the key feature information;
[0177] The association information determination subunit is used to determine the association information of the key object according to the object node corresponding to the key feature information in the object network diagram.
[0178] (4) Information combination unit 304: used to perform information combination processing on the associated information of the key object to obtain the associated feature information of the image to be identified.
[0179] (5) Image recognition unit 305: used to perform image recognition on the image to be recognized based on the associated feature information to obtain content information of the image to be recognized.
[0180] In one embodiment, the image recognition device further includes a first acquisition unit, a detection unit, and a second acquisition unit, specifically:
[0181] A first acquisition unit is used to acquire an initial image to be recognized, where the initial image to be recognized contains recognition elements;
[0182] Detection unit: used to detect the identification elements of the initial image to be identified, and determine the position information of the identification elements in the initial image to be identified;
[0183] The second acquisition unit is used to acquire the image to be recognized from the initial image to be recognized according to the position information of the recognition element.
[0184] In one embodiment, the detection unit includes an image feature extraction subunit and a position information determination subunit, specifically:
[0185] Image feature extraction subunit: used to extract image features from the initial image to be recognized, and obtain at least one feature point contained in the initial recognition image;
[0186] Position information determination subunit: used to determine the position information of the recognition element in the initial image to be recognized based on the feature points and the preset detection frame.
[0187] As can be seen from the above, the image recognition device of the embodiment of the present application acquires the image to be recognized through the image acquisition unit 301; the feature extraction unit 302 extracts key image features of the image to be recognized, and obtains key feature information of the key object contained in the image to be recognized; the information determination unit 303 determines the associated information corresponding to the key object according to the key feature information of the key object; the information combination unit 304 performs information combination processing on the associated information of the key object, and obtains the associated feature information of the image to be recognized; finally, the image recognition unit 305 performs image recognition on the image to be recognized based on the associated feature information, and obtains the image information of the image to be recognized.
[0188] This solution extracts key image features from the image to be identified to obtain key feature information corresponding to key objects, and performs image recognition on the image to be identified based on the key feature information corresponding to the key objects, so as to perform image recognition based on the key objects contained in the image to be identified, thereby avoiding image recognition errors caused by difficulty in feature extraction due to the overly complex content contained in the image to be identified when performing image recognition based on the complete image, thereby improving the accuracy of image recognition on the image to be identified.
[0189] The present application also provides a computer device, which may be a terminal or a server. Figure 8 As shown, it shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically:
[0190] The computer device may include components such as a processor 1001 with one or more processing cores, a memory 1002 with one or more computer-readable storage media, a power supply 1003, and an input unit 1004. Those skilled in the art will appreciate that Figure 8 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:
[0191] The processor 1001 is the control center of the computer device. It uses various interfaces and lines to connect various parts of the entire computer device. By running or executing software programs and / or modules stored in the memory 1002 and calling data stored in the memory 1002, it executes various functions of the computer device and processes data, thereby monitoring the computer device as a whole. Optionally, the processor 1001 may include one or more processing cores; preferably, the processor 1001 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and computer programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1001.
[0192] The memory 1002 can be used to store software programs and modules. The processor 1001 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002. The memory 1002 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, a computer program required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 1002 may also include a memory controller to provide the processor 1001 with access to the memory 1002.
[0193] The computer device also includes a power supply 1003 for supplying power to various components. Preferably, the power supply 1003 can be logically connected to the processor 1001 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 1003 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.
[0194] The computer device may further include an input unit 1004, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0195] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail herein. Specifically, in this embodiment, the processor 1001 in the computer device will load the executable files corresponding to the processes of one or more computer programs into the memory 1002 according to the following instructions, and the processor 1001 will run the computer programs stored in the memory 1002, thereby realizing various functions, as follows:
[0196] Obtain an image to be recognized;
[0197] Extract key image features from the image to be identified to obtain key feature information of key objects contained in the image to be identified;
[0198] Determine the associated information corresponding to the key object according to the key feature information of the key object;
[0199] Perform information joint processing on the associated information of key objects to obtain the associated feature information of the image to be identified;
[0200] Image recognition is performed on the image to be identified based on the associated feature information to obtain image information of the image to be identified.
[0201] The specific implementation of the above operations can be found in the previous embodiments and will not be described in detail here.
[0202] As can be seen from the above, the embodiments of the present application obtain an image to be identified; extract key image features of the image to be identified to obtain key feature information of key objects contained in the image to be identified; determine the associated information corresponding to the key objects based on the key feature information of the key objects; perform information joint processing on the associated information of the key objects to obtain associated feature information of the image to be identified; and perform image recognition on the image to be identified based on the associated feature information to obtain image information of the image to be identified.
[0203] This solution extracts key image features from the image to be identified, obtains key feature information corresponding to the key object, and performs image recognition on the image to be identified based on the key feature information corresponding to the key object, so as to perform image recognition based on the key object contained in the image to be identified, thereby avoiding image recognition errors caused by difficulty in feature extraction due to the overly complex content contained in the image to be identified when performing image recognition based on the complete image, thereby improving the accuracy of image recognition of the image to be identified. According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementations in the above-mentioned embodiments.
[0204] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by a computer program, or by controlling related hardware through a computer program. The computer program may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0205] To this end, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program can be loaded by a processor to execute any image recognition method provided in the embodiment of the present application.
[0206] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0207] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0208] Since the computer program stored in the computer-readable storage medium can execute any image recognition method provided in the embodiments of the present application, the beneficial effects that can be achieved by any image recognition method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0209] The above is a detailed introduction to an image recognition method, device, computer equipment and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An image recognition method, It is characterized in that include: Obtain an image to be recognized; Extract key image features from the image to be identified to obtain key feature information of key objects contained in the image to be identified; Determine, according to the key feature information of the key object, the associated information corresponding to the key object, wherein the associated information includes an associated relationship between the key object and other key objects; Performing information joint processing on the association information of the key objects to obtain association feature information of the image to be identified, wherein the association feature information includes key feature information of the key objects in the image to be identified and association information between the key objects; Image recognition is performed on the image to be recognized based on the associated feature information to obtain image information of the image to be recognized.
2. The method according to claim 1, It is characterized in that The key feature information includes at least two sub-key feature information, and the key feature extraction of the image to be identified to obtain the key feature information of the key object contained in the image to be identified includes: Extract key image features from the image to be identified to obtain at least two sub-key feature information of the key object; Feature fusion processing is performed on at least two key feature information corresponding to the key object to obtain the key feature information of the key object.
3. The method according to claim 2, It is characterized in that The at least two sub-key feature information include object feature information, location feature information and category feature information, and the key image feature extraction is performed on the image to be identified to obtain at least two key feature information corresponding to each key object in the image to be identified, including: Extract key image features from the image to be identified to obtain object feature information and position feature information corresponding to each key object in the image to be identified; Based on the object feature information of each of the key objects, the category feature information of the key objects is determined.
4. The method according to claim 3, It is characterized in that The determining, based on the object feature information of each of the key objects, the category feature information of the key objects comprises: Obtain initial category feature information; Based on the object feature information of each of the key objects, determining the target object category corresponding to the key object; The initial category feature information is selected and processed according to the target object category to obtain the category feature information.
5. The method according to claim 1, It is characterized in that The determining, according to the key feature information of the key object, the associated information corresponding to the key object includes: Generate an object network diagram according to the key feature information of the key object, wherein the object network diagram includes object nodes corresponding to the key feature information; Determine the association information of the key object according to the object node corresponding to the key feature information in the object network diagram.
6. The method according to any one of claims 1 to 5, It is characterized in that Before acquiring the image to be identified, the method further includes: Acquire an initial image to be identified, wherein the initial image to be identified includes identification elements; Performing identification element detection on the initial image to be identified, and determining position information of the identification element in the initial image to be identified; The image to be recognized is obtained from the initial image to be recognized according to the position information of the recognition element.
7. The method according to claim 6, It is characterized in that The performing image element detection on the initial image to be identified and determining the position information of the image element in the initial image to be identified includes: Performing image feature extraction on the initial image to be identified to obtain at least one feature point contained in the initial image to be identified; Based on the feature points and the preset detection frame, the position information of the recognition element in the initial image to be recognized is determined.
8. An image recognition device, It is characterized in that include: An image acquisition unit, used for acquiring an image to be recognized; A feature extraction unit, used to extract key image features from the image to be identified, and obtain key feature information of key objects contained in the image to be identified; An information determination unit, configured to determine, based on key feature information of the key object, associated information corresponding to the key object, wherein the associated information includes an associated relationship between the key object and other key objects; An information combination unit, used for performing information combination processing on the associated information of the key object to obtain associated feature information of the image to be identified; An image recognition unit is used to perform image recognition on the image to be recognized based on the associated feature information, wherein the associated feature information includes key feature information of the key object in the image to be recognized and associated information between the key objects, so as to obtain content information of the image to be recognized.
9. A computer device, It is characterized in that It comprises a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the image recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium is used to store a computer program, and the computer program is loaded by a processor to execute the image recognition method according to any one of claims 1 to 7.
11. A computer product, It is characterized in that The computer product comprises a computer program, and when the computer program is executed by a processor, the image recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Electric meter reading identification method based on YOLOV3 network
CN111461121A
Image recognition method, device and equipment and readable storage medium
CN111553419A
Image recognition method, device and equipment and computer storage medium
CN113052159A