Image recognition method and device, and computer-readable storage medium
By performing multi-resolution downsampling and feature fusion on images and combining them with multi-scale target detection, the problem of limited image recognition results in existing technologies is solved, and accurate recognition of target objects is achieved to meet the needs of real-world application scenarios.
Patent Information
- Application Number
- CN202210331824.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-03-30
AI Technical Summary
Existing image recognition technology cannot meet the application requirements of real-world scenarios, has limited recognition results, and cannot accurately identify the position, area percentage, and number of target objects in the image.
By performing multi-resolution downsampling on the image to be identified, feature maps of different resolutions are obtained. After feature fusion, multi-scale target detection is performed, target candidate frames are screened out, and the target objects are identified to obtain category information, location information, area ratio and quantity information.
The richness of image recognition results has been improved, and it can accurately identify the position, area proportion and number of target objects in the image to meet the needs of real-world application scenarios.
Smart Images

Figure CN114693918B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to an image recognition method and device, and a computer-readable storage medium. Background Art
[0002] Image recognition technology refers to the technology that processes, analyzes and understands images to identify various target objects. Among them, in the input image, the category of the trademark in the image is determined. It can be used to interact with users in product activities and provide users with information, promotional information and other application scenarios, which can help enhance the brand influence of the product and increase product sales.
[0003] Currently, target objects in images are usually identified based on image matching algorithms, but the recognition results are limited and cannot meet the needs of real-world scenarios. Summary of the Invention
[0004] The embodiments of the present invention are intended to provide an image recognition method and apparatus, and a computer-readable storage medium, which can improve the richness of image recognition results to meet the application needs of real-world scenarios.
[0005] The technical solution of the present invention is achieved as follows:
[0006] An embodiment of the present invention provides an image recognition method, including:
[0007] Performing multi-resolution downsampling on the image to be identified, thereby obtaining first feature maps corresponding to at least two resolutions;
[0008] Performing feature fusion on the first feature maps corresponding to the at least two resolutions, thereby obtaining a second feature map;
[0009] Performing multi-scale object detection on the second feature map to obtain a preset number of candidate boxes;
[0010] Screening the preset number of candidate frames to obtain a target candidate frame; the target candidate frame represents an area in the second feature map that contains a target object;
[0011] The target object corresponding to the target candidate frame is identified to obtain a recognition result.
[0012] In the above technical solution, the recognition result includes category information, location information, area ratio and quantity information.
[0013] In the above technical solution, the multi-resolution downsampling of the image to be identified, thereby obtaining first feature maps corresponding to at least two resolutions, includes:
[0014] Determining a downsampling ratio coefficient based on the side length of the image to be identified; the downsampling ratio coefficient is a common divisor of the side lengths of the image to be identified;
[0015] Downsampling the image to be recognized at least two resolutions according to the downsampling ratio coefficient to obtain the first feature maps corresponding to the at least two resolutions.
[0016] In the above technical solution, performing multi-scale target detection on the second feature map to obtain a preset number of candidate frames includes:
[0017] Detecting the sub-features of each decomposition layer in the second feature map through sliding windows until a preset number of sliding windows have detected the target sub-features, thereby achieving multi-scale target detection;
[0018] According to the target sub-feature mapping to the area of the image to be identified, a corresponding candidate frame is generated, thereby obtaining the preset number of candidate frames.
[0019] In the above technical solution, the identification of the area corresponding to the target candidate frame to obtain the identification result includes:
[0020] Identify the target object corresponding to the target candidate frame to obtain position information of the target candidate frame;
[0021] The area ratio information and the target position information in the recognition result are obtained according to the position information and the size information of the image to be recognized.
[0022] In the above technical solution, the identification of the area corresponding to the target candidate frame to obtain the identification result includes:
[0023] Identify the target object corresponding to the target candidate frame to obtain a predicted classification label of the target object;
[0024] The predicted classification label is matched with the preset classification label to obtain the target quantity information and the category information in the recognition result.
[0025] In the above technical solution, the method further includes:
[0026] When the quantity information is for one target object, item activity information corresponding to the one target object is displayed.
[0027] In the above technical solution, the method further includes:
[0028] When the quantity information is at least two target objects, receiving an activity information acquisition request, the activity information acquisition request carrying any one of the at least two target objects;
[0029] In response to the activity information acquisition request, corresponding item activity information is displayed according to the any one target object.
[0030] In the above technical solution, the method further includes:
[0031] receiving a product tracking request, wherein the product tracking request carries the identification result;
[0032] In response to the product tracking request, corresponding tracking information is obtained based on the identification result, and prompt information is displayed according to the tracking information and the area ratio information in the identification result.
[0033] An embodiment of the present invention provides an image recognition device, comprising a downsampling unit, a fusion unit, a detection unit, a screening unit, and a recognition unit; wherein,
[0034] The downsampling unit is configured to perform multi-resolution downsampling on the image to be identified, thereby obtaining first feature maps corresponding to at least two resolutions;
[0035] The fusion unit is configured to perform feature fusion on the first feature maps corresponding to at least two resolutions, thereby obtaining a second feature map;
[0036] The detection unit is configured to perform multi-scale target detection on the second feature map, thereby obtaining a preset number of candidate boxes;
[0037] The screening unit is configured to screen the preset number of candidate frames to obtain a target candidate frame; the target candidate frame represents an area in the second feature map containing a target object;
[0038] The recognition unit is used to recognize the target object corresponding to the target candidate frame, thereby obtaining a recognition result.
[0039] In the above technical solution, the downsampling unit is further used to downsample the image to be identified according to a preset resolution, so as to obtain the first feature maps corresponding to at least two resolutions.
[0040] In the above technical solution, the downsampling unit is also used to determine a downsampling ratio coefficient based on the side length of the image to be identified; the downsampling ratio coefficient is a common divisor of the side length of the image to be identified; and the image to be identified is downsampled to at least two resolutions according to the downsampling ratio coefficient to obtain the first feature maps corresponding to at least two resolutions.
[0041] In the above technical solution, the detection unit is also used to detect the sub-features of each decomposition layer in the second feature map through a sliding window until a preset number of sliding windows detect the target sub-features, thereby realizing multi-scale target detection; according to the target sub-features mapped to the area of the image to be identified, a corresponding candidate box is generated, thereby obtaining the preset number of candidate boxes.
[0042] In the above technical solution, the recognition unit is also used to identify the target object corresponding to the target candidate frame to obtain the position information of the target candidate frame; based on the position information and the size information of the image to be identified, the area ratio information and the position information in the recognition result are obtained.
[0043] In the above technical solution, the recognition unit is further used to recognize the target object corresponding to the target candidate frame to obtain a predicted classification label of the target object;
[0044] The predicted classification label is matched with the preset classification label to obtain the quantity information and the category information in the recognition result.
[0045] In the above technical solution, the device further includes a display unit, which is used to display item activity information corresponding to the target object when the quantity information is a target object.
[0046] In the above technical solution, the device further includes a receiving unit, wherein:
[0047] The receiving unit is configured to, when the quantity information indicates at least two target objects, receive an activity information acquisition request, wherein the activity information acquisition request carries any one of the at least two target objects;
[0048] The display unit is configured to display the corresponding item activity information according to the any one target object in response to the activity information acquisition request.
[0049] In the above technical solution, the receiving unit is further configured to receive a product tracking request, wherein the product tracking request carries the identification result;
[0050] The display unit is further configured to obtain corresponding tracking information based on the identification result in response to the product tracking request, and display prompt information based on the tracking information and the area ratio information in the identification result.
[0051] An embodiment of the present invention provides an image recognition device, comprising:
[0052] A memory for storing executable data instructions;
[0053] The processor is configured to implement an image recognition method as described in an embodiment of the present invention when executing the executable instructions stored in the memory.
[0054] An embodiment of the present invention provides a computer-readable storage medium storing executable instructions for causing a processor to execute the instructions to implement an image recognition method according to an embodiment of the present invention.
[0055] An embodiment of the present invention provides an image recognition method and device, and a computer-readable storage medium. The method includes performing multi-resolution downsampling on an image to be recognized, then performing feature fusion on the multi-resolution feature map obtained by downsampling, and then extracting features of different resolutions in the feature map after feature fusion through a multi-scale method, so as to determine the area that may contain the target object and generate a candidate box to select the above area. Finally, the above area that may contain the target object is screened to determine the area with the highest probability of containing the target object and the target candidate box. The recognition result can be obtained by identifying the target candidate box and the area within the target candidate box.
[0056] In the embodiment of the present invention, through multi-resolution downsampling and multi-scale methods, target objects of different sizes in the image to be identified can be identified, thereby improving the richness of the recognition results and meeting the needs of real application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 The process of an image recognition method provided by an embodiment of the present invention is Figure 1 ;
[0058] Figure 2 A schematic diagram of the structure of a preset detection network provided by an embodiment of the present invention;
[0059] Figure 3 The process of an image recognition method provided by an embodiment of the present invention is Figure 2 ;
[0060] Figure 4 A flowchart of an image recognition method provided by an embodiment of the present invention;
[0061] Figure 5 A schematic diagram of the structure of an image recognition device provided by an embodiment of the present invention Figure 1 ;
[0062] Figure 6 A schematic diagram of the structure of an image recognition device provided by an embodiment of the present invention Figure 2 . DETAILED DESCRIPTION
[0063] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0064] Figure 1 This is a process of an image recognition method provided by an embodiment of the present invention Figure 1 ,like Figure 1 As shown, an embodiment of the present invention provides an image recognition method, including:
[0065] S101 : Perform multi-resolution downsampling on an image to be recognized, thereby obtaining first feature maps corresponding to at least two resolutions.
[0066] The embodiment of the present invention is applicable to a scenario in which an image to be recognized is sampled for subsequent recognition.
[0067] In an embodiment of the present invention, an image recognition device performs multi-resolution downsampling on an image to be recognized through a preset detection network, thereby obtaining a first feature map corresponding to at least two resolutions; wherein the first feature map represents feature information in the image to be recognized, and the feature information refers to the inherent features of the image to be recognized that can be distinguished from other categories of images, such as brightness, edges, texture, and color.
[0068] In an embodiment of the present invention, after the image recognition device obtains the image to be recognized, the image to be recognized is input into a preset detection network, so that the preset detection network performs multi-resolution downsampling on the image to be recognized; multi-resolution downsampling refers to downsampling the image to be recognized according to different resolutions to obtain images to be recognized with different resolutions, and digitizing the images to be recognized with different resolutions to obtain feature information in the images to be recognized with different resolutions, thereby obtaining first feature maps corresponding to each of the different resolutions.
[0069] In an embodiment of the present invention, when performing multi-resolution downsampling on an image to be identified, the image to be identified is adjusted to different resolutions, and feature information in the image to be identified at different resolutions is extracted; wherein, feature information of the same resolution (size) is combined to form a feature vector, that is, a first feature map corresponding to the above resolution is obtained; the first feature maps corresponding to different resolutions are arranged in a preset order, which can be arranged from high resolution to low resolution, or from low resolution to high resolution. For example, Figure 2 is a schematic diagram of a structure of a preset detection network provided by an embodiment of the present invention, such as Figure 2As shown, the preset detection network includes Backbone, Neck, and Head. Backbone is the backbone network, used to extract information from the image to be identified; Head is the detection head, primarily responsible for predicting the type and location of the target. Neck, located between Backbone and Head, adds network layers for collecting feature maps from different stages to better utilize the information extracted by Backbone. In actual use, Backbone implements multi-resolution downsampling of the image to be identified, thereby obtaining first feature maps corresponding to at least two resolutions. When the resolution of the image to be identified is 256×256 and downsampling is required five times, first feature maps with resolutions of 256×256, 128×128, 64×64, 32×32, 16×16, and 8×8 can be obtained, respectively. These first feature maps are arranged in order from high to low resolution. Backbone is also used to output the collected first feature maps corresponding to at least two resolutions to Neck for collection.
[0070] It can be understood that by performing multi-resolution downsampling on the image to be identified, a first feature map corresponding to at least two resolutions collected at different resolutions can be obtained, and due to the different resolutions, the features corresponding to the feature information in the size of the first feature map corresponding to the above at least two resolutions are also different, which can improve the subsequent feature extraction to identify the recognition effect and recognition accuracy of the target object.
[0071] In some embodiments of the present invention, S101 may further include S1011 and S1012, which are as follows:
[0072] S1011 : Determine a downsampling ratio coefficient based on the side length of the image to be recognized.
[0073] In some embodiments of the present invention, a downsampling scale factor is determined based on the side length of the image to be recognized, wherein the downsampling scale factor is a common divisor of the side lengths of the image to be recognized.
[0074] In some embodiments of the present invention, for example, an image with a resolution of M×N (image to be identified) is downsampled by s times (downsampling ratio coefficient) to obtain an image with a resolution of (M / s)×(N / s) (first feature map), where s is a common divisor of M and N; when the resolution of the image to be identified is 256×256 and the image to be identified needs to be downsampled 5 times, the downsampling ratio coefficient can be determined to be 2.
[0075] S1012: Downsample the image to be recognized to at least two resolutions according to the downsampling ratio coefficient to obtain first feature maps corresponding to the at least two resolutions.
[0076] In some embodiments of the present invention, the image to be identified is downsampled to at least two resolutions according to a downsampling ratio coefficient; since each downsampling will obtain a layer of first feature map, downsampling the image to be identified to at least two resolutions can obtain first feature maps corresponding to at least two resolutions.
[0077] In some embodiments of the present invention, when downsampling the image to be identified, each downsampling means that the image to be identified is reduced once according to the downsampling ratio coefficient. As the size of the image to be identified is reduced, the resolution of the image to be identified will also be reduced. According to the feature information in the images to be identified with different resolutions, the corresponding first feature map can be obtained, and the resolution of each layer of the first feature map is different and in a decreasing state. For example, after downsampling an image with a resolution of M×N (image to be identified) by s times (downsampling ratio coefficient) and obtaining an image with a resolution of (M / s)×(N / s), it means that the first feature map of the first layer is obtained, and the obtained image with a resolution of (M / s)×(N / s) is downsampled by s times, thereby obtaining (M / s)×(N / s). 2 )×(M / s 2 ) resolution image, and the first feature map of the second layer is obtained.
[0078] It can be understood that since the resolution of the first feature map of each layer is different, the size of the feature information included in the first feature map of each layer will also be different. In this way, when the first feature maps corresponding to the above-mentioned at least two resolutions are detected, the comprehensiveness of the obtained feature information can be improved and the recognition effect can be improved.
[0079] S102: Perform feature fusion on the first feature maps corresponding to at least two resolutions to obtain a second feature map.
[0080] The embodiment of the present invention is applicable to scenarios where features are enhanced to facilitate subsequent recognition processing.
[0081] In an embodiment of the present invention, feature fusion is performed on the first feature maps corresponding to at least two resolutions, that is, the feature information in the first feature maps corresponding to at least two resolutions are added bit by bit, thereby completing the feature fusion and obtaining a second feature map, wherein the second feature map includes all feature information extracted from the first feature maps corresponding to at least two resolutions, and all feature information is distributed hierarchically according to the corresponding resolutions.
[0082] In this embodiment of the present invention, the first feature map corresponding to at least two resolutions will contain more positional information and detail information. However, the first feature map with a higher resolution undergoes less convolution and therefore contains less semantic information. The first feature map with a lower resolution, on the other hand, contains more semantic information, but due to its lower resolution, its perception of detail is poor. Positional information, detail information, and semantic information all constitute feature information.
[0083] In the embodiment of the present invention, Figure 2 As shown, the Neck part represents the process of adding the first feature maps corresponding to at least two resolutions bit by bit to obtain the second feature map, wherein the arrows represent the order of adding the first feature maps corresponding to at least two resolutions bit by bit. When performing feature fusion, the deep information (semantic information) of the first feature map with low resolution is first extracted, and then the shallow information (position information, detail information) of the first feature map with high resolution is extracted. Finally, the deep information and the shallow information are added element by element to obtain multiple layers of sub-features, and each layer is called a decomposition layer. The sub-feature resolutions within the decomposition layers in the second feature map are the same, and the sub-feature resolutions between layers decrease. For example, the Neck part performs feature fusion on the first feature maps with resolutions of 64×64, 32×32, 16×16 and 8×8 to obtain the second feature map, so that the sub-feature resolutions between layers in the second feature map obtained will decrease from 64×64 to 8×8.
[0084] It is understandable that this can improve the calculation amount of Neck and reduce it. When the second feature map is subsequently processed to obtain the target object, more feature information can be obtained based on the sub-features of multiple resolutions, thereby improving the accuracy of identifying the target object.
[0085] S103: Perform multi-scale object detection on the second feature map to obtain a preset number of candidate boxes.
[0086] The embodiment of the present invention is applicable to a scenario where a region of interest is selected from a second feature map.
[0087] In an embodiment of the present invention, multi-scale target detection is performed on the second feature map, that is, target detection is performed on the second feature map through a multi-scale method, so as to obtain feature maps of different resolutions. Feature maps of different resolutions represent target objects of different sizes. After detecting the above-mentioned target object, a candidate box is generated to select the target object, thereby obtaining a preset number of candidate boxes.
[0088] In this embodiment of the present invention, the second feature map is composed of at least two decomposition layers, each of which contains sub-features of different resolutions. The sub-features of different resolutions represent different feature information of the image to be recognized. The resolution of the sub-features in each of the at least two decomposition layers is the same as the resolution of the corresponding first feature map.
[0089] In an embodiment of the present invention, a multi-scale method, also known as a multi-resolution method, is used to obtain sub-features of different resolutions in the second feature map, and to determine the target object and its location through the sub-features of different resolutions. Approximate features of the target object can be obtained through low-resolution sub-features, and detailed features of the target object can be obtained through high-resolution sub-features. For example, a decomposition layer with a resolution of 8×8 in the second feature map corresponds to a first feature map with a resolution of 8×8. During target detection, since the size of the area corresponding to each sub-feature with a resolution of 8×8 is relatively large after the sub-features of the decomposition layer with a resolution of 8×8 in the second feature map are mapped to the image to be identified, a target object with a larger area can be obtained by detecting the sub-features with a resolution of 8×8. A decomposition layer with a resolution of 64×64 in the second feature map corresponds to a first feature map with a resolution of 64×64. Since, after the sub-features of the decomposition layer with a resolution of 64×64 are mapped to the image to be identified, the size of the area corresponding to each sub-feature with a resolution of 64×64 will be smaller than the area corresponding to the sub-feature with a resolution of 8×8; therefore, by detecting the sub-features with a resolution of 64×64, target objects with a smaller area can be obtained.
[0090] In an embodiment of the present invention, candidate frames are generated to search for locations in the image to be identified that may contain target objects. These locations are also called regions of interest (ROIs). In actual use, a sliding window (sliding window) is used to scan the sub-features in each decomposition layer in the second feature map. The size and number of the sliding window are preset, and the size and number of the sliding window are set according to each decomposition layer in the second feature map. During the scanning process, the sliding window will detect the sub-features of each decomposition layer in the second feature map, thereby obtaining the target sub-features in each decomposition layer and capturing the semantic information of the area corresponding to the above target sub-features. The target sub-features are then mapped to the image to be identified and the corresponding area is selected to generate a candidate frame.
[0091] In the embodiment of the present invention, Figure 2As shown, for example, if the second feature map is obtained by feature fusion of the first feature maps with resolutions of 64×64, 32×32, 16×16 and 8×8, and it is preset that there are 6 candidate boxes at each position of the decomposition layer with a resolution of 8×8 in the second feature map, and 2 candidate boxes at each position of the other decomposition layers. When the sliding window captures the target sub-features at the decomposition layer with a resolution of 8×8 in the second feature map, the captured sub-features will be scored during the capture process, and 6 candidate boxes will be generated at each position based on the score. As the target sub-features of the decomposition layers with resolutions of 64×64, 32×32 and 16×16 in the second feature map are captured by the sliding window and 2 candidate boxes are generated at each position, the second feature map will obtain 8×8×6+16×16×2+32×32×2+64×64×2=11136 candidate boxes. Among them, Figure 2 The 8×8×6, 16×16×2, 32×32×2, and 64×64×2 resolutions between the Neck and Head represent the number of candidate boxes generated at each position in the second feature map at resolutions of 8×8, 16×16, 32×32, and 64×64, respectively. The scoring criteria for generating candidate boxes is automatically generated by the preset detection network during training. After the second feature map completes multi-scale object detection, the preset number of candidate boxes is fed into the Boxes function.
[0092] It can be understood that by obtaining sub-features of different resolutions in the second feature map through multi-scale target detection, it is possible to effectively simulate the size changes caused by the distance from the target object in real life and improve the recognition accuracy.
[0093] In some embodiments of the present invention, the sub-features of each decomposition layer in the second feature map are detected by sliding windows until a preset number of sliding windows detect the target sub-features, thereby achieving multi-scale target detection.
[0094] In some embodiments of the present invention, by presetting a sliding window for each decomposition layer in the second feature map, the sub-features of each decomposition layer in the second feature map are detected until each sliding window detects the corresponding target sub-features in each decomposition layer, thereby completing multi-scale target detection in the second feature map.
[0095] In some embodiments of the present invention, corresponding candidate frames are generated based on the target sub-features mapped to the area of the image to be identified, thereby obtaining a preset number of candidate frames.
[0096] In some embodiments of the present invention, the area where the target sub-feature is mapped to the image to be identified is an area that may contain the target object, and a candidate box with the same size as the sliding window is generated based on the above area, thereby obtaining a preset number of candidate boxes.
[0097] It can be understood that by obtaining sub-features of different resolutions in the second feature map through multi-scale target detection, more feature information can be obtained, and the position of the target object can be determined based on the feature information.
[0098] S104: Screen a preset number of candidate frames to obtain a target candidate frame. The target candidate frame represents the area containing the target object in the second feature map.
[0099] In the embodiment of the present invention, it is applicable to a scenario where a preset number of candidate frames are screened to determine a candidate frame with the highest probability of containing a target object as the target candidate frame.
[0100] In an embodiment of the present invention, a preset number of candidate frames are screened to obtain a target candidate frame; the target candidate frame is a candidate frame with the highest probability of containing the target object, and therefore, the target candidate frame represents the area containing the target object in the second feature map.
[0101] In an embodiment of the present invention, when a preset number of candidate boxes are screened, since there is a preset number of candidate boxes at each position of each decomposition layer in the second feature map, the preset number of candidate boxes may overlap with each other. According to the overlapping of the candidate boxes, the positioning accuracy of the candidate boxes for the target object position can be obtained by intersection over union (IOU).
[0102] In the embodiment of the present invention, Figure 2 As shown, a preset number of candidate frames are screened by NMS (Non-Maximum Suppression). The screening process is: first, the preset number of candidate frames are sorted according to the score of each candidate frame, wherein the score of each candidate frame is obtained when the candidate frame is generated in S103. Secondly, the candidate frames with a score greater than the preset score are screened out, and the IOU between the candidate frames with a score greater than the preset score is calculated. Finally, the candidate frames with an IOU higher than the threshold are removed. Through the above screening process, the preset number of candidate frames are iteratively selected until the target candidate frame is obtained. Among them, the number of target candidate frames is at least one.
[0103] It can be understood that by screening the candidate frames, the target candidate frames are determined, and each target candidate frame represents a target object. In the process of screening the candidate frames, at least one candidate frame can be obtained, which means that the embodiment of the present invention can identify multiple target objects from the image to be identified, thereby expanding the scope of application scenarios.
[0104] S105: Identify the target object corresponding to the target candidate frame to obtain a recognition result.
[0105] The embodiment of the present invention is applicable to the scenario where the recognition result is output through the target candidate box.
[0106] In an embodiment of the present invention, the target object selected by the target candidate frame and the corresponding target sub-features in the second feature map are identified, and the target object and the corresponding recognition result are obtained according to the target sub-features.
[0107] In an embodiment of the present invention, feature information of the target object selected by the target candidate box is extracted. This feature information is the target sub-feature corresponding to the target candidate box in the second feature map. In actual use, a classifier can be used to identify the target sub-features, obtain a predicted classification label corresponding to each target sub-feature, and then match the predicted classification label with a preset classification label. If the predicted classification label matches the preset classification label, the content selected by the target candidate box is the target object. In this case, the predicted classification label represents the category information of the target object. For example, if the target object to be identified is a trademark, and the preset classification labels are "wine," "beverage," "condiment," etc., and the classifier identifies the target sub-feature as "wine," then because the predicted classification label matches the preset classification label, the category information of the target object will be: wine. If three predicted classification labels are obtained, and two of the three predicted classification labels successfully match the preset classification label, then the number of target objects is two, and the category information of these two target objects can be the same or different.
[0108] In this embodiment of the present invention, the coordinates of the four corners of the target candidate frame are identified to obtain the location information of the target object, that is, the location information is obtained. At the same time, the side length of the target candidate frame can be obtained based on the coordinates of the four corners of the target candidate frame, and thus the area of the target candidate frame can be obtained. Combined with the area of the image to be recognized, the quotient of the area of the target candidate frame and the area of the image to be recognized is taken to obtain the area ratio of the target object, that is, the area ratio in the recognition result.
[0109] Understandably, this can enrich image recognition results, solving the problem of being unable to determine the location, area percentage, and number of target objects when identifying target objects using image matching technology, thereby meeting the needs of real-world application scenarios.
[0110] In some embodiments of the present invention, the recognition result includes category information, location information, area ratio, and quantity information.
[0111] In some embodiments of the present invention, the number of target objects and their category information are obtained by identifying the target candidate frame, and the area percentage and location information of the target objects are determined based on the size and location information of the target candidate frame. This results in the following recognition results: category information, location information, area percentage, and number information.
[0112] In some embodiments of the present invention, the target candidate box can be recognized by the same Convolutional Neural Networks (CNN) model and the recognition results can be output simultaneously. Alternatively, the target candidate box can be recognized by multiple cascaded CNN models, and different sub-recognition results in the recognition result can be output separately.
[0113] In some embodiments of the present invention, the target object corresponding to the target candidate frame is identified to obtain position information of the target candidate frame.
[0114] In some embodiments of the present invention, the target object corresponding to the target candidate frame is identified, and a coordinate system is constructed according to the image to be identified, thereby obtaining coordinate information of the target candidate frame.
[0115] In some embodiments of the present invention, the area ratio information and the position information in the recognition result are obtained based on the position information and the size information of the image to be recognized.
[0116] In some embodiments of the present invention, position information can be obtained based on the coordinate information of the target candidate box. After calculating the area of the target candidate box based on the coordinate information, the area of the image to be recognized is calculated based on the size information of the image to be recognized, thereby obtaining area ratio information in the recognition result.
[0117] In some embodiments of the present invention, the target object corresponding to the target candidate frame is identified to obtain a predicted classification label of the target object.
[0118] In some embodiments of the present invention, the target object corresponding to the target candidate frame is identified, and a coordinate system is constructed according to the image to be identified, thereby obtaining coordinate information of the target candidate frame.
[0119] In some embodiments of the present invention, the predicted classification labels are matched with the preset classification labels to obtain the quantity information and category information of the target objects in the recognition results.
[0120] In some embodiments of the present invention, the predicted classification label is matched with the preset classification label. If the matching result is that the predicted classification label successfully matches the preset classification label, that is, the predicted classification label belongs to the preset classification label, then the target object is in the target candidate frame. If the target object is a trademark, and the predicted classification label is the category information corresponding to the trademark; if the matching result is that the predicted classification label fails to match the preset classification label, that is, the predicted classification label does not belong to the preset classification label, then the target object is not in the target candidate frame. Based on the matching results of the output predicted classification label and the preset classification label, trademark quantity information can be obtained.
[0121] In some embodiments of the present invention, the image recognition method further includes: when the quantity information is a target object, displaying item activity information corresponding to the target.
[0122] In some embodiments of the present invention, when the obtained number is one target object, that is, when there is only one target object in the image to be identified, item activity information related to the category information is displayed based on the category information corresponding to the one target object.
[0123] In some embodiments of the present invention, item activity information can be displayed in the form of a pop-up window, and the display terminal can be a mobile phone or computer, etc. The item activity information is stored in the client database. When the above item activity information is displayed, it can be called from the client database.
[0124] In some embodiments of the present invention, the image recognition method further includes: when the quantity information is at least two target objects, receiving an activity information acquisition request, where the activity information acquisition request carries any one of the at least two target objects.
[0125] In some embodiments of the present invention, when the obtained quantity information is at least two target objects, that is, there are at least two target objects in the image to be identified, the client receives an activity information acquisition request, wherein the activity information acquisition request is initiated through interaction between the user and the client, that is, the activity information acquisition request is initiated by the user selecting any one of the at least two target objects; therefore, the activity information acquisition request carries any one of the at least two target objects.
[0126] In some embodiments of the present invention, in response to an activity information acquisition request, corresponding item activity information is displayed according to any target object.
[0127] In some embodiments of the present invention, item activity information related to any target object is displayed based on the location information and / or category information of any target object, where the item activity information is a response result in response to the activity information acquisition request.
[0128] In some embodiments of the present invention, an activity information acquisition request is used to request item activity information related to any one of at least two target objects. For example, when at least two target objects are present in the image to be recognized, the client displays the item activity information related to the target object selected by the user in the activity information acquisition request in a pop-up window.
[0129] It is understandable that, through the activity information acquisition request, the item activity information required by the user can be displayed, providing the user with more interactive operations and meeting more of the user's needs.
[0130] In some embodiments of the present invention, the image recognition method further includes: receiving a product tracking request, the product tracking request carrying a recognition result;
[0131] In some embodiments of the present invention, an item tracking request is received, and the item tracking request carries area ratio information of a target corresponding to the item.
[0132] In some embodiments of the present invention, a camera captures the current image, identifies the target object in the current image, and captures the target object associated with the item corresponding to the item tracking request. Because the target object's area percentage in the current image changes with the distance between the camera and the item, the location of the item can be tracked based on the target object's area percentage in the current image. During the tracking process, different prompts may appear as the target object's area percentage changes. For example, information about the target object's category and the distance between the target object and the user may be displayed.
[0133] In some embodiments of the invention, in response to a product tracking request, corresponding tracking information is obtained based on the identification result, and prompt information is displayed according to the tracking information and the area ratio information in the identification result.
[0134] In some embodiments of the invention, corresponding tracking information is obtained based on the items related to the identification results, and prompt information is displayed based on the tracking information and area ratio information to inform the user of the current distance from the item.
[0135] It is understandable that the object can be tracked through the recognition results, and the distance between the user and the object can be known based on the area ratio information of the target object. This is suitable for more application scenarios and expands the scope of application.
[0136] Figure 3 This is a process of an image recognition method provided by an embodiment of the present invention Figure 2 ,like Figure 3 As shown, an embodiment of the present invention provides an image recognition method. The above image recognition method is described by taking a trademark as an example. The above method includes:
[0137] S201 , performing detection preprocessing on the image to be recognized to obtain a processed image to be recognized.
[0138] In some embodiments of the present invention, Figure 4 is a flowchart of an image recognition method provided by an embodiment of the present invention, such as Figure 4 As shown in FIG, after obtaining the image to be recognized, it is necessary to perform detection preprocessing on the image to be recognized. Detection preprocessing is used to adjust the size, format, etc. of the image to be recognized.
[0139] It is understandable that by processing the image to be identified to meet the preset requirements through detection preprocessing, it is more convenient to subsequently identify the image to be identified and the recognition effect is guaranteed.
[0140] S202 , down-sampling the processed image to be recognized, sampling the processed image to be recognized from 256×256 to 8×8, and obtaining a multi-layer feature map (first feature maps corresponding to at least two resolutions).
[0141] In some embodiments of the present invention, the feature map is downsampled by the anchor-base detection network (preset detection network). Figure 4 As shown in FIG, the processed image to be identified is input into the detection network for trademark recognition. When the detection network recognizes the trademark, the recognition effect is ensured by the data set (LogoDet) in the detection network.
[0142] In some embodiments of the present invention, the data set obtains image samples to be identified through offline shooting or online posting of single pictures. After the trademarks in the image samples to be identified are marked, they are classified according to the trademarks in the images to be identified. Finally, the image samples to be identified are input into the detection network and the detection network is trained. When the detection network completes the training, a data set will be generated in the detection network.
[0143] It is understandable that the data set in the above detection network is trained by collecting image samples containing trademarks to be identified, which can ensure the pertinence of the data set for trademark identification and the accuracy of the detection network in identifying trademarks.
[0144] S203: Add the features of the multiple layers of feature maps bit by bit to obtain a second feature map.
[0145] In some embodiments of the present invention, the detection network performs feature fusion by bit-by-bit addition of features of multiple layers of feature maps.
[0146] It can be understood that through feature fusion, the bottom-level position information and high-level semantic information in the multi-layer feature map can be effectively combined.
[0147] S204 : Detect target objects of different sizes and set anchors (candidate boxes) using a multi-scale method and sub-features of different resolutions in the second feature map.
[0148] In some embodiments of the present invention, the detection network uses sub-features of different resolutions to be responsible for target objects of different sizes. For example, by extracting sub-features with a resolution of 8×8 in the second feature map, larger target objects are detected; by extracting sub-features with a resolution of 64×64 in the second feature map, smaller target objects are detected.
[0149] It is understandable that this can effectively simulate the size change caused by the distance between the camera and the trademark when the client obtains the image to be recognized in real life.
[0150] S205: Obtain the target anchor and recognition result through NMS.
[0151] In some embodiments of the present invention, Figure 4 As shown in the figure, by screening the anchors through NMS, the target anchor (detection box) can be obtained. In actual use, the required recognition results, such as the category (category information) and area ratio (area ratio information) of the trademark, can be obtained through the target anchor.
[0152] It can be understood that this not only can identify the category of the trademark, but also can obtain the area information of the trademark as needed, which can expand the use scenarios of product identification in activities.
[0153] S206: Output the target solution based on the recognition result.
[0154] In some embodiments of the present invention, Figure 4 As shown, after obtaining the recognition result, the corresponding target solution can be generated according to the recognition result and output.
[0155] It is understandable that customizing and generating target activities based on recognition results can achieve the goals of enhancing the brand influence of products, increasing user stay time, and promoting product sales.
[0156] Figure 5 This is a schematic diagram of the structure of an image recognition device provided by an embodiment of the present invention. Figure 1 ,like Figure 5 As shown, an embodiment of the present invention provides an image recognition device 5, comprising a downsampling unit 51, a fusion unit 52, a detection unit 53, a screening unit 54 and a recognition unit 55; wherein,
[0157] The downsampling unit 51 is configured to perform multi-resolution downsampling on the image to be identified, thereby obtaining first feature maps corresponding to at least two resolutions;
[0158] The fusion unit 52 is configured to perform feature fusion on the first feature maps corresponding to at least two resolutions, thereby obtaining a second feature map;
[0159] The detection unit 53 is configured to perform multi-scale object detection on the second feature map, thereby obtaining a preset number of candidate boxes;
[0160] The screening unit 54 is configured to screen the preset number of candidate frames to obtain a target candidate frame; the target candidate frame represents an area in the second feature map that contains a target object;
[0161] The recognition unit 55 is configured to recognize the target object corresponding to the target candidate frame, thereby obtaining a recognition result.
[0162] In some embodiments of the present invention, the downsampling unit 51 is further configured to downsample the image to be recognized according to a preset resolution, thereby obtaining the at least two layers of first feature maps.
[0163] In some embodiments of the present invention, the downsampling unit 51 is further used to determine a downsampling ratio coefficient based on the side length of the image to be identified; the downsampling ratio coefficient is a common divisor of the side length of the image to be identified; and the image to be identified is downsampled to at least two resolutions according to the downsampling ratio coefficient to obtain the first feature maps corresponding to at least two resolutions.
[0164] In some embodiments of the present invention, the detection unit 53 is further used to detect the sub-features of each decomposition layer in the second feature map through a sliding window until a preset number of sliding windows detect the target sub-features, thereby realizing multi-scale target detection; and generate corresponding candidate boxes based on the target sub-features mapped to the area of the image to be identified, thereby obtaining the preset number of candidate boxes.
[0165] In some embodiments of the present invention, the recognition unit 55 is further used to identify the target object corresponding to the target candidate frame to obtain the position information of the target candidate frame; based on the position information and the size information of the image to be identified, the area ratio information and the position information in the recognition result are obtained.
[0166] In some embodiments of the present invention, the recognition unit 55 is also used to identify the target object corresponding to the target candidate frame to obtain a predicted classification label of the target object; match the predicted classification label with a preset classification label to obtain the quantity information and the category information in the recognition result.
[0167] In some embodiments of the present invention, the device further includes a display unit 56 for displaying item activity information corresponding to one target object when the quantity information is one target object.
[0168] In some embodiments of the present invention, the apparatus further comprises a receiving unit 57, wherein:
[0169] The receiving unit 57 is configured to receive an activity information acquisition request when the quantity information indicates at least two target objects, wherein the activity information acquisition request carries any one of the at least two target objects;
[0170] The display unit 56 is configured to display the corresponding item activity information according to the any one target object in response to the activity information acquisition request.
[0171] In some embodiments of the present invention, the receiving unit 57 is further configured to receive a product tracking request, wherein the product tracking request carries the identification result;
[0172] The display unit 56 is further configured to respond to the product tracking request, obtain corresponding tracking information based on the identification result, and display prompt information according to the tracking information and the area ratio information in the identification result.
[0173] Figure 6 This is a schematic diagram of the structure of an image recognition device provided by an embodiment of the present invention. Figure 2 As shown in Figure 6, an embodiment of the present invention provides an image recognition device 6, including: a processor 61, a memory 62 and a communication bus 64. The memory 62 communicates with the processor 61 via the communication bus 64. The memory 62 stores one or more programs executable by the processor 61. When the one or more programs are executed, the processor 61 executes the image recognition method according to the embodiment of the present invention. Specifically, the image recognition device 6 also includes a communication component 63 for data transmission, wherein the processor 61 is provided with at least one.
[0174] In the embodiment of the present invention, the components in the image recognition device 6 are coupled together via a bus 64. It is understood that the bus 64 is used to realize the connection and communication between these components. In addition to the data bus, the bus 64 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 6 The various buses are labeled as passing through bus 64.
[0175] An embodiment of the present invention provides a computer-readable storage medium storing executable instructions for causing a processor to execute the instructions to implement an image recognition method according to an embodiment of the present invention.
[0176] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0177] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0178] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0180] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.
Claims
1. An image recognition method, characterized in that: The method is applied to a client, and includes: Performing multi-resolution downsampling on the image to be identified, thereby obtaining first feature maps corresponding to at least two resolutions; Performing feature fusion on the first feature maps corresponding to the at least two resolutions to obtain a second feature map; the second feature map includes at least two decomposition layers; detecting sub-features of each decomposition layer in the second feature map using sliding windows corresponding to the at least two decomposition layers, until a preset number of sliding windows detect target sub-features, thereby achieving multi-scale target detection; the size and number of the sliding windows are determined based on the corresponding decomposition layers; Generate a corresponding candidate frame having the same size as the sliding window according to the target sub-feature mapped to the area of the image to be recognized, thereby obtaining a preset number of candidate frames; Screening the preset number of candidate frames to obtain a target candidate frame; the target candidate frame represents an area in the second feature map containing a target object; the target object includes a trademark; Identify the target object corresponding to the target candidate frame to obtain a recognition result; the recognition result includes the number information of the target object, the category information of the target object, and the location information of the target object; When the number of pieces of information is at least two, the method further includes: In response to a selection operation on any one of the at least two target objects, receiving an activity information acquisition request carrying the any one target object; In response to the activity information acquisition request, the item activity information of the arbitrary target object is displayed according to the location information and category information of the arbitrary target object.
2. The method according to claim 1, characterized in that The recognition result includes area ratio.
3. The method according to claim 1, characterized in that The multi-resolution downsampling of the image to be identified, thereby obtaining first feature maps corresponding to at least two resolutions, includes: Determining a downsampling ratio coefficient based on the side length of the image to be identified; the downsampling ratio coefficient is a common divisor of the side lengths of the image to be identified; Downsampling the image to be recognized into at least two resolutions according to the downsampling ratio coefficient to obtain the first feature maps corresponding to the at least two resolutions respectively.
4. The method according to claim 2, characterized in that The identifying the area corresponding to the target candidate frame to obtain the identification result includes: Identify the target object corresponding to the target candidate frame to obtain position information of the target candidate frame; The area ratio information and the position information in the recognition result are obtained according to the position information and the size information of the image to be recognized.
5. The method according to claim 2, characterized in that The identifying the area corresponding to the target candidate frame to obtain the identification result includes: Identify the target object corresponding to the target candidate frame to obtain a predicted classification label of the target object; The predicted classification label is matched with the preset classification label to obtain the quantity information and the category information in the recognition result.
6. The method according to claim 2, 4 or 5, characterized in that The method further comprises: When the quantity information is for one target object, item activity information corresponding to the one target object is displayed.
7. The method according to claim 2, 4 or 5, characterized in that The method further comprises: receiving a product tracking request, wherein the product tracking request carries the identification result; In response to the product tracking request, corresponding tracking information is obtained based on the identification result, and prompt information is displayed according to the tracking information and the area ratio information in the identification result.
8. An image recognition device, characterized in that: It includes a downsampling unit, a fusion unit, a detection unit, a screening unit, a recognition unit, a receiving unit and a display unit; wherein, The downsampling unit is configured to perform multi-resolution downsampling on the image to be identified, thereby obtaining first feature maps corresponding to at least two resolutions; The fusion unit is configured to perform feature fusion on the first feature maps corresponding to at least two resolutions, thereby obtaining a second feature map; the second feature map includes at least two decomposition layers; The detection unit is configured to detect sub-features of each decomposition layer in the second feature map using sliding windows corresponding to the at least two decomposition layers, respectively, until a preset number of sliding windows detect target sub-features, thereby achieving multi-scale target detection; the size and number of the sliding windows are determined based on the corresponding decomposition layers; and based on the target sub-features being mapped to the area of the image to be identified, generating corresponding candidate boxes of the same size as the sliding windows, thereby obtaining a preset number of candidate boxes; The screening unit is configured to screen the preset number of candidate frames to obtain a target candidate frame; the target candidate frame represents an area in the second feature map containing a target object; the target object includes a trademark; The recognition unit is used to recognize the target object corresponding to the target candidate frame, thereby obtaining a recognition result; the recognition result includes the number information of the target objects, the category information of the target objects, and the location information of the target objects; In a case where the number information is at least two, the receiving unit is configured to receive, in response to a selection operation on any one of the at least two target objects, an activity information acquisition request carrying the any one target object; The display unit is configured to display the item activity information of the arbitrary target object according to the location information and category information of the arbitrary target object in response to the activity information acquisition request.
9. An image recognition device, characterized in that: include: a memory for storing executable instructions; A processor, configured to implement the method according to any one of claims 1 to 7 when executing the executable instructions stored in the memory.
10. A computer-readable storage medium, characterized in that Executable instructions are stored, which are used to cause a processor to execute and implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and apparatus for detecting target
CN110084257A
Equipment state detection method and device based on image recognition, equipment and medium
CN111666958A