An image recognition method, device, electronic equipment and storage medium

By performing object detection, local feature extraction, and reconstruction on complex images, the problem of inter-class interference is solved, and the accuracy and stability of image recognition are improved.

CN114332809BActive Publication Date: 2026-05-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-12-01
Publication Date
2026-05-22

Smart Images

  • Figure CN114332809B_ABST
    Figure CN114332809B_ABST
Patent Text Reader

Abstract

The application discloses an image recognition method and device, electronic equipment and storage medium. The method can perform object detection on a to-be-processed image to obtain an object detection image, input the object detection image into a local feature extraction network for feature extraction to obtain a plurality of local feature information, input the plurality of local feature information into a local feature reorganization network for feature reorganization to obtain reorganized feature information, and input the reorganized feature information into an image recognition network for type recognition to obtain target type information corresponding to the object detection image. The method can extract local feature information in the object detection image, reorganize the local feature information, improve the model's recognition ability of the local feature information, reduce the inter-class interference of the object detection image, and improve the accuracy and stability of the object detection image recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular to an image recognition method, apparatus, electronic device and storage medium. Background Technology

[0002] Image recognition refers to the technology of using computers to process, analyze, and understand images in order to identify targets and objects of various patterns. It is a practical application of deep learning algorithms. In existing technologies, when recognizing complex images, the entire complex image is often used as annotation information to train a convolutional neural network to extract and recognize high-level semantic features. However, when complex images contain multiple similar types, this method leads to significant inter-class interference, resulting in decreased accuracy and stability of image recognition, and causing problems such as false detections and incorrect classification. Summary of the Invention

[0003] This application provides an image recognition method, apparatus, electronic device, and storage medium, which can reduce inter-class interference in object detection images and improve the accuracy and stability of object detection image recognition.

[0004] On one hand, this application provides an image recognition method, the method comprising:

[0005] Perform object detection on the image to be processed to obtain an object detection image, wherein the object detection image is an object image of at least two objects located in the same connected region in the image to be processed;

[0006] The object detection image is input into a local feature extraction network for feature extraction to obtain multiple local feature information, wherein the local feature information is the feature information corresponding to each of the at least two objects;

[0007] The multiple local feature information is input into a local feature recombination network for feature recombination to obtain recombined feature information;

[0008] The recombined feature information is input into an image recognition network for type recognition to obtain the target type information corresponding to the object detection image.

[0009] On the other hand, an image recognition device is provided, the device comprising:

[0010] An object detection module is used to perform object detection on the image to be processed and obtain an object detection image, wherein the object detection image is an object image of at least two objects located in the same connected region in the image to be processed.

[0011] The feature extraction module is used to input the object detection image into a local feature extraction network for feature extraction to obtain multiple local feature information, wherein the local feature information is the feature information corresponding to each of the at least two objects;

[0012] The feature recombination module is used to input the multiple local feature information into the local feature recombination network for feature recombination to obtain recombined feature information;

[0013] The type recognition module is used to input the recombined feature information into an image recognition network for type recognition, so as to obtain the target type information corresponding to the object detection image.

[0014] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement an image recognition method as described above.

[0015] On the other hand, a computer-readable storage medium is provided, the storage medium including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement an image recognition method as described above.

[0016] On the other hand, a computer program product is provided, including a computer program that, when executed by a processor, implements the image recognition method described above.

[0017] This application provides an image recognition method, apparatus, electronic device, and storage medium. The method performs object detection on an image to be processed, obtaining an object detection image. The object detection image is then input into a local feature extraction network for feature extraction, yielding multiple local feature information. This local feature information is then input into a local feature reconstruction network for feature reconstruction, obtaining reconstructed feature information. Finally, the reconstructed feature information is input into an image recognition network for type recognition, obtaining the target type information corresponding to the object detection image. This method can extract local feature information from object detection images and reconstruct this local feature information, thereby improving the model's ability to identify local feature information, reducing inter-class interference in object detection images, and improving the accuracy and stability of object detection image recognition. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram illustrating an application scenario of an image recognition method provided in an embodiment of this application;

[0020] Figure 2 A flowchart of an image recognition method provided in an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of the target object detection network in an image recognition method provided in an embodiment of this application;

[0022] Figure 4 This is a schematic diagram of the candidate bounding box of the target object detection network in an image recognition method provided in an embodiment of this application;

[0023] Figure 5 This is a flowchart illustrating the determination of local feature information in an image recognition method provided in an embodiment of this application;

[0024] Figure 6 This is a flowchart illustrating the acquisition of reconstructed feature information in an image recognition method provided in an embodiment of this application;

[0025] Figure 7 A flowchart illustrating feature fusion based on target distance in an image recognition method provided in this application embodiment;

[0026] Figure 8 A flowchart illustrating a method for model training in an image recognition method provided in this application embodiment;

[0027] Figure 9 A flowchart illustrating the calculation of target loss information in an image recognition method provided in this application embodiment;

[0028] Figure 10 This is a schematic diagram illustrating an image recognition method provided in this application for use in a traffic sign recognition scenario.

[0029] Figure 11 This is a schematic diagram of the structure of an image recognition device provided in an embodiment of this application;

[0030] Figure 12 This is a schematic diagram of the hardware structure of a device for implementing the method provided in the embodiments of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0032] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. Furthermore, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein.

[0033] Please see Figure 1 This diagram illustrates an application scenario of an image recognition method provided in this application. The application scenario includes a client 110 and a server 120. The client 110 acquires an image to be processed. The server 120 receives the image to be processed from the client 110. The server 120 obtains an object detection image from the image to be processed and extracts features from the object detection image to obtain local feature information corresponding to each object. The server 120 reconstructs the local feature information to obtain reconstructed feature information and performs type recognition on the reconstructed feature information to obtain target type information corresponding to the object detection image. The server 120 sends the target type information to the client 110.

[0034] In this embodiment, client 110 includes physical devices such as smartphones, desktop computers, tablets, laptops, digital assistants, and smart wearable devices, and may also include software running on the physical device, such as applications. The operating system running on the physical device in this embodiment may include, but is not limited to, Android, iOS, Linux, Unix, and Windows. Client 110 includes a UI (User Interface) layer, through which it provides the display of the image to be processed and the target type information. Additionally, it sends the data required for image recognition to server 120 based on API (Application Programming Interface).

[0035] In the embodiments of the application, server 120 may include a standalone server, a distributed server, or a server cluster consisting of multiple servers. Server 120 may include a network communication unit, a processor, and a memory, etc. Specifically, server 120 can be used to extract features from object detection images to obtain local feature information corresponding to each object, and to reconstruct the local feature information to obtain reconstructed feature information. Then, the reconstructed feature information is used for type recognition to obtain target type information corresponding to the object detection image.

[0036] Please see Figure 2 It demonstrates an image recognition method applicable to the server side, which includes:

[0037] S210. Perform object detection on the image to be processed to obtain an object detection image. The object detection image is the object image of at least two objects located in the same connected region in the image to be processed.

[0038] In some embodiments, a connected region is a region in the image to be processed that has a closed boundary, meaning the boundary of the region is continuous. A connected region may include at least two objects that completely belong to the connected region. For example, traffic signs are typically square, circular, or triangular in shape, and their boundaries are all closed boundaries. Since a traffic sign may include at least two objects, such as patterns or text, and these patterns or text are completely located within the boundaries of the traffic sign, the traffic sign can be considered a connected region in the image to be processed. Object detection is performed on the image to be processed to obtain an object detection image, which is an object image of at least two objects located in the same connected region in the image to be processed. These at least two objects can be objects of the same category or objects of different categories. The objects can be elements in the image to be processed, such as traffic signs in a traffic road image. For example, if the at least two objects are both images, they can be at least two different types of image information. If the at least two objects include both images and text, they can also be at least one type of image information and at least one type of text information. Object detection images are diverse, with significant inter-class interference, and contain rich information, such as traffic signs and shop signs. These images often consist of multiple patterns, and some may also include text. For example, the directional arrows on traffic signs can include various types such as straight ahead, left turn, right turn, left turn straight ahead, right turn straight ahead, U-turn, left turn U-turn, diagonally up right, diagonally up left, diagonally down right, diagonally down left, left, right, down, etc. In addition, there are many irregular partial images, such as bus images indicating bus lanes, bicycle images indicating non-motorized vehicle lanes, car images indicating motorized vehicle lanes, step images indicating pedestrian overpasses and underpasses, and images indicating speed cameras. These partial images can also be combined according to practical applications, leading to significant recognition difficulties. Shop signs are similar; they can include different fonts and patterns, and depending on the type of shop, there may be irregular fonts or patterns embedded in the text, further increasing the recognition challenge.

[0039] In some embodiments, object detection is performed on the image to be processed to obtain an object detection image, including:

[0040] The image to be processed is input into the target object detection network for object detection, and the object detection image is obtained.

[0041] In some embodiments, see Figure 3 ,like Figure 3The diagram shows the structure of a target object detection network, which includes convolutional layers, normalization layers, and activation layers. Convolutional layers extract basic features such as edges and textures. Normalization layers normalize the features extracted by the convolutional layers according to a normal distribution, filtering out noisy features. Activation layers perform non-linear mapping on the features extracted by the convolutional layers. The image to be processed is input into the convolutional layers for feature extraction, obtaining initial feature information. This initial feature information is then input into the normalization layers for normalization, and finally input into the activation layers for non-linear mapping to obtain the target feature information.

[0042] Please see Figure 4 ,like Figure 4 The diagram illustrates candidate bounding boxes used by an object detection network for object detection. Each feature point in the target feature information can be used as the center point to select three candidate bounding boxes with aspect ratios of 1:1, 2:1, and 1:2. Each candidate bounding box includes three additional bounding boxes with scales of 1, 2, and 3 feature points, respectively. Based on the target feature information and the corresponding candidate bounding boxes, the object detection image can be determined.

[0043] The object detection image is determined from the image to be processed, thereby distinguishing between regions that are difficult to identify and regions that are easy to identify. This allows for feature recombination and re-identification of regions that are difficult to identify, thus improving the accuracy of object detection image recognition.

[0044] S220. Input the object detection image into a local feature extraction network to extract features and obtain multiple local feature information, wherein the local feature information is the feature information corresponding to each of at least two objects;

[0045] In some embodiments, when an object detection image is input into a local feature extraction network for feature extraction, local feature information corresponding to each object in the object detection image can be obtained, while information irrelevant to object recognition is removed. If the object detection image includes at least two objects, at least two pieces of local feature information can be obtained. This local feature information can be semantically rich feature information, and based on this local feature information, each object in the object detection image can be identified.

[0046] In some embodiments, see Figure 5 The local feature extraction network includes an image feature extraction network and a local semantic recognition network. The object detection image is input into the local feature extraction network for feature extraction, yielding multiple local feature information, including:

[0047] S510. Input the object detection image into the image feature extraction network to extract features and obtain the detection image feature information;

[0048] S520. Input the detected image feature information into the local semantic recognition network for semantic recognition to obtain the local feature information corresponding to each object;

[0049] Multiple local feature information is input into a local feature reconstruction network for feature reconstruction, resulting in reconstructed feature information including:

[0050] S530. Input the detected image feature information and multiple local feature information into the local feature reconstruction network for feature reconstruction to obtain reconstructed feature information.

[0051] In some embodiments, before inputting the object detection image into the image feature extraction network, the object detection image can be processed by scaling it to match a preset size. For example, if the preset size is set to 300x300 pixels, the object detection image can be scaled to 300x300 pixels before being used as input to the image feature extraction network.

[0052] In some embodiments, the local feature extraction network includes an image feature extraction network and a local semantic recognition network. The image feature extraction network is used to extract features from objects in the object detection image and determine the positions of different objects in the object detection image based on preset bounding boxes and extracted feature information. Inputting the object detection image into the image feature extraction network for feature extraction yields detection image feature information. This detection image feature information represents the overall characteristic information corresponding to the object detection image. Different regions in the detection image feature information have different feature values, and objects in the object detection image can be determined based on these different feature values. The image feature extraction network can be a ResNet convolutional neural network.

[0053] In some embodiments, the detected image feature information is input into a local semantic recognition network for semantic recognition. Based on the different feature values ​​in the detected image feature information, local regions corresponding to each object are determined. Feature information corresponding to each local region is obtained, and mean pooling is performed on the feature information corresponding to each local region to unify the feature information of each local region to the same scale, thereby obtaining the local feature information corresponding to each object. For example, a traffic sign with an arrow pattern and a highway exit name is input into an image feature extraction network for feature extraction to obtain the detected image feature information. This detected image feature information is then input into a local semantic recognition network for semantic recognition, thus obtaining the local feature information corresponding to the arrow pattern and the local feature information corresponding to the highway exit name.

[0054] In some embodiments, after acquiring the detection image feature information and local feature information, the detection image feature information and local feature information are input into the local feature reconstruction network. Based on the image feature information, the detection position information corresponding to the local feature information is calibrated. Then, based on the calibrated detection position information, the local feature information is reconstructed to obtain the reconstructed feature information.

[0055] Local feature information of each object in the object detection image is extracted, thereby extracting regions with rich semantic information in the image to be processed. This allows the local feature reconstruction network to pay more attention to the features of regions with rich semantic information, remove useless information in the object detection image, and improve the effectiveness of local feature extraction.

[0056] S230. Input multiple local feature information into a local feature recombination network to perform feature recombination and obtain recombined feature information;

[0057] In some embodiments, multiple local feature information is input into a local feature reconstruction network, and the local feature information is spliced ​​together based on the location information corresponding to each local feature information to obtain reconstructed feature information.

[0058] In some embodiments, see Figure 6 The local feature reconstruction network includes a location feature extraction network and a feature fusion network. It inputs the detected image feature information and multiple local feature information into the local feature reconstruction network for feature reconstruction, resulting in reconstructed feature information including:

[0059] S610. Input the detected image feature information into the location feature extraction network to extract location features and obtain the location calibration information corresponding to each object. The location calibration information is the calibration information of the detection position of each object in the object detection image.

[0060] S620. Input multiple local feature information and the position calibration information corresponding to each local image into the feature fusion network to perform feature fusion and obtain reconstructed feature information.

[0061] In some embodiments, the detected image feature information is input into a location feature extraction network for location feature extraction to obtain the offset of the detection position of each object in the detected image feature information, thereby obtaining location calibration information. For example, if the detection position information of object A in the detected image feature information is coordinates (x, y), when the detected image feature information is input into the location feature extraction network for location feature extraction, the location calibration information of object A is obtained as (x1, y1), then the actual position information of object A is (x+x1, y+y1).

[0062] Position calibration information can be represented as a position potential field, which is a two-dimensional direction vector that characterizes the positional trend of the corresponding object in the object detection image. In other words, it represents the position that local feature information should be in the corresponding image feature information of the object detection image. For example, if an object should be located in the upper left corner of the object detection image, then the direction of the position potential field tends to point in the upper left direction.

[0063] Based on the position calibration information, we can determine how far the actual position has shifted compared to the detection position. Therefore, based on the position calibration information, we can adjust the detection position information corresponding to each local feature. For example, if the detection position information of object B is shifted two pixels to the left compared to the actual position information, we can determine the position calibration information of object B and adjust it so that the detection position information of object B is shifted two pixels to the right, thereby matching the actual position information.

[0064] In some embodiments, multiple local feature information and position calibration information corresponding to each local image are input into a feature fusion network. Based on the position calibration information, the detection position information corresponding to each local feature information is calibrated. Then, based on the calibrated detection position information, the multiple local feature information are fused to obtain reconstructed feature information.

[0065] Feature fusion of multiple local features rich in semantic information can integrate the complex spatial information within an object detection image, thereby improving the effectiveness of feature fusion. Furthermore, by first identifying local features and then reconstructing them, the model's ability to recognize local features can be improved, thus enhancing the accuracy and stability of object detection image recognition in subsequent steps.

[0066] In some embodiments, see Figure 7 Multiple local feature information and the position calibration information corresponding to each local image are input into a feature fusion network for feature fusion to obtain reconstructed feature information, including:

[0067] S710. Input multiple local feature information and the position calibration information corresponding to each local image into the feature fusion network, and determine the target distance between each local feature information and the preset starting fusion position based on the detection position information corresponding to each local feature information and the position calibration information corresponding to each local feature information.

[0068] S720. Based on the target distance, feature fusion is performed on multiple local feature information to obtain recombined feature information.

[0069] In some embodiments, the calibrated position information is obtained by adding the coordinates corresponding to the detection position information and the coordinates corresponding to the position calibration information. When performing feature fusion, a preset starting fusion position can be determined first. The target distance from each local feature information to the preset starting fusion position can then be calculated using the calibrated position information obtained by adding the detection position information and the position calibration information, as shown in the following formula:

[0070]

[0071] Where d represents the target distance, (x, y) represents the detection location information, and (x1, y1) represents the location calibration information. A smaller target distance indicates that the local feature information is closer to the preset starting fusion position, while a larger target distance indicates that the local feature information is farther from the preset starting fusion position. For example, if the preset starting fusion position is set to the top-left vertex of the object detection image, then the target distance from each local feature information to the top-left vertex is calculated. The smaller the target distance, the closer the local feature information is to the top-left vertex.

[0072] In some embodiments, a recombination weight for each local feature can be determined based on the target distance. A fusion sequence is then determined based on the recombination weight, and the local feature information is fused. The recombination weight represents the probability that each local feature is selected during feature fusion. The smaller the target distance, the larger the recombination weight, and the greater the probability that the local feature is selected. Therefore, the local feature appears earlier in the fusion sequence. Thus, when fusing local feature information, local feature information with a large recombination weight is placed at the beginning of the fusion sequence and is selected for fusion first. Local feature information with a small recombination weight is placed later in the fusion sequence and is selected for fusion later. That is, the local feature information is fused from the preset starting fusion position to the diagonal vertex opposite the preset starting fusion position. The recombination weight is calculated as shown in the following formula:

[0073]

[0074] Where, p i Let represent the recombination weight, n represent the number of local features, and i represent the i-th local feature. After calculating the recombination weight corresponding to each local feature, the fusion sequence corresponding to the local feature can be determined sequentially according to the magnitude of the recombination weight.

[0075] In some embodiments, the local feature information of fusion sequence 1 can be determined first. After removing the local feature information of fusion sequence 1, the local feature information of fusion sequence 2 can be determined based on the other local feature information besides the local feature information of fusion sequence 1, and so on. After determining the local feature information corresponding to each fusion sequence, the local feature information is deleted, and the remaining local feature information is used to determine the local feature information corresponding to the next fusion sequence. This process continues until the fusion sequence for which all local feature information has been determined is obtained.

[0076] After calibrating the position of local feature information, feature fusion can make the position of the local feature information match the original position better, thereby improving the accuracy of feature fusion.

[0077] S240. Input the recombined feature information into the image recognition network for type recognition to obtain the target type information corresponding to the object detection image.

[0078] In some embodiments, identifying the type information corresponding to the reconstructed feature information in an image recognition network can yield the target type information corresponding to the object detection image. The reconstructed feature information is the feature information of a semantically rich region in the object detection image, such as patterns and text. The target type information is the overall recognition result of the object detection image; for example, if the object detection image is a traffic sign, then the target type information is the type of that traffic sign.

[0079] In some embodiments, when training the model, please refer to Figure 8 The method also includes:

[0080] S810. Obtain the sample detection image, the annotation type information corresponding to the sample detection image, and the annotation position information corresponding to each sample object in the sample detection image. The sample detection image is the object image of at least two sample objects located in the same connected region in the sample image.

[0081] S820. Input the sample detection image into the first network to be trained for feature extraction to obtain multiple sample local feature information and training detection location information corresponding to each sample local feature information. The sample local feature information is the feature information corresponding to each of at least two sample objects.

[0082] S830. Input the local feature information of multiple samples into the second network to be trained for feature recombination to obtain training recombination feature information;

[0083] S840. Input the training recombination feature information into the third network to be trained for type recognition, and obtain the training type information corresponding to the sample detection image;

[0084] S850. Determine the target loss information based on training type information, annotation type information, annotation location information, and training detection location information;

[0085] S860. Based on the target loss information, the first network to be trained, the second network to be trained, and the third network to be trained are trained to obtain a local feature extraction network, a local feature reconstruction network, and an image recognition network.

[0086] In some embodiments, a sample detection image is determined from a sample image. The sample detection image is an object image of at least two sample objects located in the same connected region within the sample image. The sample detection image is image information of a known type, and the type information corresponding to the sample detection image is used as the annotation type information. The original position corresponding to each sample object in the sample detection image is used as the annotation position information.

[0087] The first network to be trained includes a network for extracting features from images to be trained and a network for semantic recognition to be trained. The sample detection image is input into the network for extracting features to be trained to obtain the feature information of the training image. Then, the feature information of the training image is input into the network for semantic recognition to be trained to perform local semantic recognition, which can obtain the local feature information of each sample object and the training detection location information corresponding to each local feature information of the sample.

[0088] Training image feature information and sample local feature information are input into a second network to be trained for feature recombination, resulting in training recombined feature information. The second network to be trained includes a training location feature extraction network and a training feature fusion network. Training image feature information is input into the training location feature extraction network for location feature extraction, resulting in training location calibration information. Training location calibration information and training detection location information corresponding to sample local feature information are input into the training feature fusion network. Based on the training location calibration information and training detection location information, the sample distance between the sample local feature information and the preset starting fusion position is determined. Based on this sample distance, feature fusion is performed on the local feature information to obtain training recombined feature information.

[0089] By inputting the recombined training feature information into the third network to be trained for type recognition, the training type information corresponding to the sample detection image can be obtained.

[0090] Based on training type information, annotation type information, annotation location information, and training detection location information, target loss information can be determined. Based on this target loss information, models are trained on the first, second, and third networks to be trained, thereby obtaining a local feature extraction network, a local feature reconstruction network, and an image recognition network.

[0091] Based on training type information, annotation type information, annotation location information, and training detection location information, the target loss information is determined and the model is trained. This allows for simultaneous training of local feature extraction and classification results, thereby improving the accuracy of the local feature extraction network, the local feature reconstruction network, and the image recognition network.

[0092] In some embodiments, see Figure 9 Based on training type information, annotation type information, annotation location information, and training / detection location information, the target loss information is determined to include:

[0093] S910. Determine the classification loss information based on the training type information and the annotation type information;

[0094] S920. Determine the location loss information based on the labeled location information and the training and detection location information;

[0095] S930. Determine the target loss information based on the classification loss information and the location loss information.

[0096] In some embodiments, classification loss information can be determined based on training type information and labeled type information. The classification loss information is the difference between the training type information and the labeled type information; therefore, it measures the accuracy of the first, second, and third training networks in type recognition. The classification loss information can be cross-entropy.

[0097] Based on the labeled location information and the training detection location information, the location loss information can be determined. The location loss information is the difference between the detected location and the actual location of each sample object in the sample detection image. Therefore, the location loss information can measure the accuracy of the local feature information of the sample. The location loss information can be regression loss information, such as smooth L1, which is the L1 norm loss function after smoothing.

[0098] By fusing classification loss information and location loss information, target loss information can be obtained. Then, based on the target loss information, the first, second, and third training networks are trained.

[0099] In some embodiments, the formula for calculating the target loss information is as follows:

[0100]

[0101] Where L represents the target loss information, For location loss information, To classify loss information, t i Indicates the location information of the annotation, t' iThis represents the training location information, M represents the number of sample detection image types, and y represents the training location information. ic p is an indicator variable; it indicates 1 when the training type information and the annotation type information are the same, and 0 when they are different. ic This represents the probability that sample object i belongs to category c.

[0102] The formula for calculating smooth L1 is as follows:

[0103] smooth L1 (x)=0.5x 2 if |x| < 1

[0104] smooth L1 (x) = |x| - 0.5 otherwise

[0105] In the image recognition method proposed in the embodiments of this application, x is (t i -t' i ).

[0106] By using location information and classification loss information, local feature extraction and classification recognition can be corrected, thereby improving the accuracy of model training.

[0107] In some embodiments, see Figure 10 ,like Figure 10 The diagram illustrates an image recognition method applied to a traffic sign recognition scenario. The client can be a vehicle-mounted terminal, the image to be processed is a road image captured by the vehicle-mounted terminal, and the object detection image is a traffic sign. The vehicle-mounted terminal captures the road image and sends it to the server. The server determines the image information of the traffic sign from the road image. The traffic sign may include arrows, text, roads, and non-motorized vehicle symbols, etc.

[0108] The local feature extraction network includes an image feature extraction network and a local semantic recognition network. By inputting the traffic sign into the image feature extraction network for feature extraction, the image feature information corresponding to the traffic sign can be obtained. Then, the image feature information is input into the local semantic recognition network for local semantic recognition, which can obtain local feature information corresponding to arrow patterns, text, road patterns, and non-motor vehicle patterns, etc.

[0109] By inputting the image feature information corresponding to the traffic sign and the local feature information corresponding to each object into a local feature reconstruction network for feature reconstruction, reconstructed feature information can be obtained. The local feature reconstruction network includes a location feature extraction network and a feature fusion network. Inputting the image feature information into the location feature extraction network for location feature extraction yields location calibration information for arrow patterns, text, road patterns, and non-motorized vehicle patterns, etc. For example, if a non-motorized vehicle pattern is located in the lower right corner of the traffic sign, the potential field direction corresponding to the location calibration information tends to point towards the lower right.

[0110] The detection location information corresponding to the position calibration information and local feature information is input into the feature fusion network. Based on the position calibration information and detection location information, the target distance between the local feature information and the preset starting fusion position is determined. Based on this target distance, feature fusion is performed on the local feature information to obtain reconstructed feature information. The reconstructed feature information is input into the image recognition network for type recognition to obtain the target type information corresponding to the traffic sign. The server sends the target type information to the vehicle terminal, which displays the target type information and provides a prompt to the user.

[0111] This application provides an image recognition method, which includes: performing object detection on an image to be processed to obtain an object detection image; inputting the object detection image into a local feature extraction network for feature extraction to obtain multiple local feature information; inputting the multiple local feature information into a local feature reconstruction network for feature reconstruction to obtain reconstructed feature information; and inputting the reconstructed feature information into an image recognition network for type recognition to obtain target type information corresponding to the object detection image. This method can extract local feature information from the object detection image and reconstruct the local feature information, thereby improving the model's ability to identify local feature information, reducing inter-class interference in the object detection image, and improving the accuracy and stability of object detection image recognition.

[0112] This application also provides an image recognition device; please refer to [link to relevant documentation]. Figure 11 ,like Figure 11 As shown, the device includes:

[0113] Object detection module 1110 is used to perform object detection on the image to be processed and obtain an object detection image. The object detection image is an object image of at least two objects located in the same connected region in the image to be processed.

[0114] The feature extraction module 1120 is used to input the object detection image into the local feature extraction network for feature extraction, and obtain multiple local feature information, wherein the local feature information is the feature information corresponding to each of at least two objects;

[0115] The feature recombination module 1130 is used to input multiple local feature information into the local feature recombination network for feature recombination to obtain recombined feature information;

[0116] The type recognition module 1140 is used to input the recombined feature information into the image recognition network for type recognition, and obtain the target type information corresponding to the object detection image.

[0117] In some embodiments, the local feature extraction network includes an image feature extraction network and a local semantic recognition network, and the feature extraction module 1120 includes:

[0118] The image feature extraction unit is used to input the object detection image into the image feature extraction network for feature extraction to obtain the detection image feature information;

[0119] The semantic recognition unit is used to input the detected image feature information into the local semantic recognition network for semantic recognition, and obtain the local feature information corresponding to each object;

[0120] Feature recombination module 1130 includes:

[0121] The feature reconstruction subunit is used to input the detected image feature information and multiple local feature information into the local feature reconstruction network for feature reconstruction to obtain reconstructed feature information.

[0122] In some embodiments, the local feature reconstruction network includes a location feature extraction network and a feature fusion network, and the feature reconstruction subunit includes:

[0123] The location feature extraction unit is used to input the detection image feature information into the location feature extraction network for location feature extraction, and obtain the location calibration information corresponding to each object. The location calibration information is the calibration information of the detection position of each local image in the object detection image.

[0124] The feature fusion unit inputs multiple local feature information and the position calibration information corresponding to each object into the feature fusion network to perform feature fusion and obtain reconstructed feature information.

[0125] In some embodiments, the feature fusion unit includes:

[0126] The target distance determination unit is used to input multiple local feature information and the position calibration information corresponding to each object into the feature fusion network, and determine the target distance between each local feature information and the preset starting fusion position based on the detection position information and the position calibration information corresponding to each local feature information.

[0127] The local feature fusion unit is used to fuse multiple local feature information based on the target distance to obtain recombined feature information.

[0128] In some embodiments, the object detection module 1110 includes:

[0129] The object detection subunit is used to input the image to be processed into the target object detection network for object detection, and obtain the object detection image.

[0130] In some embodiments, the device further includes:

[0131] The sample information acquisition module is used to acquire the sample detection image, the annotation type information corresponding to the sample detection image, and the annotation position information corresponding to each sample object in the sample detection image. The sample detection image is the object image of at least two sample objects located in the same connected region in the sample image.

[0132] The sample feature extraction module is used to input the sample detection image into the first network to be trained for feature extraction, and to obtain multiple sample local feature information and training detection location information corresponding to each sample local feature information. The sample local feature information is the feature information corresponding to each of at least two sample objects.

[0133] The sample feature recombination module is used to input local feature information of multiple samples into the second network to be trained for feature recombination to obtain training recombination feature information.

[0134] The training type recognition module is used to input the training recombination feature information into the third network to be trained for type recognition, and obtain the training type information corresponding to the sample detection image;

[0135] The target loss calculation module is used to determine the target loss information based on training type information, annotation type information, annotation location information, and training detection location information;

[0136] The model training module is used to train the first, second, and third networks to be trained based on the target loss information, resulting in a local feature extraction network, a local feature reconstruction network, and an image recognition network.

[0137] In some embodiments, the target loss calculation module includes:

[0138] The classification loss calculation unit is used to determine the classification loss information based on the training type information and the label type information;

[0139] The location loss calculation unit is used to determine the location loss information based on the labeled location information and the training and detection location information;

[0140] The target loss determination unit is used to determine the target loss information based on the classification loss information and the location loss information.

[0141] The apparatus provided in the above embodiments can execute the method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the above embodiments can be found in an image recognition method provided in any embodiment of this application.

[0142] This embodiment also provides a computer-readable storage medium storing computer-executable instructions, which are loaded by a processor and executed by the image recognition method described above in this embodiment.

[0143] This embodiment also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the image recognition described above.

[0144] This embodiment also provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program adapted to be loaded by the processor and executed by the image recognition method described above in this embodiment.

[0145] The device can be a computer terminal, a mobile terminal, or a server, and it can also participate in constituting the apparatus or system provided in the embodiments of this application. For example... Figure 12 As shown, server 12 may include one or more processors 1202 (shown as 1202a, 1202b, ..., 1202n in the figure) 1202 (processor 1202 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 1204 for storing data, and a transmission device 1206 for communication functions. In addition, it may also include: input / output interfaces (I / O interfaces) and network interfaces. Those skilled in the art will understand that... Figure 12 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 12 may also include components that are more... Figure 12 The more or fewer components shown, or having the same Figure 12 The different configurations shown.

[0146] It should be noted that the aforementioned one or more processors 1202 and / or other data processing circuitry are generally referred to herein as "data processing circuitry". This data processing circuitry may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the server 12.

[0147] The memory 1204 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method in the embodiments of this application. The processor 1202 executes various functional applications and data processing by running the software programs and modules stored in the memory 1204, thereby realizing the above-described method for generating temporal behavior capture boxes based on self-attention networks. The memory 1204 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1204 may further include memory remotely located relative to the processor 1202, and these remote memories can be connected to the server 12 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0148] The transmission device 1206 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 12. In one example, the transmission device 1206 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet.

[0149] This specification provides method operation steps as shown in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operation steps may be included. The steps and order listed in the embodiments are merely one possible execution order among many steps and do not represent the only execution order. In actual system or interrupt product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0150] The structures shown in this embodiment are only partial structures related to the solution of this application and do not constitute a limitation on the device to which the solution of this application is applied. Specific devices may include more or fewer components than shown, or combinations of certain components, or arrangements of different components. It should be understood that the methods, apparatuses, etc., disclosed in this embodiment can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection between devices or unit modules through some interfaces.

[0151] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0152] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0153] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An image recognition method, characterized in that, The method includes: Perform object detection on the image to be processed to obtain an object detection image, wherein the object detection image is an object image of at least two objects located in the same connected region in the image to be processed; The object detection image is input into a local feature extraction network for feature extraction to obtain multiple local feature information, wherein the local feature information is the feature information corresponding to each of the at least two objects; The detection image feature information corresponding to the object detection image is input into the position feature extraction network for position feature extraction to obtain the position calibration information corresponding to each object image. The position calibration information is the calibration information of the detection position of each object image in the object detection image. The multiple local feature information and the position calibration information corresponding to each object are input into the feature fusion network. Based on the detection position information corresponding to each local feature information and the position calibration information corresponding to each local feature information, the target distance between each local feature information and the preset starting fusion position is determined. Based on the target distance, feature fusion is performed on the multiple local feature information to obtain reconstructed feature information; The recombined feature information is input into an image recognition network for type recognition to obtain the target type information corresponding to the object detection image.

2. The image recognition method according to claim 1, characterized in that, The local feature extraction network includes an image feature extraction network and a local semantic recognition network. The object detection image is input into the local feature extraction network for feature extraction to obtain multiple local feature information, including: The object detection image is input into the image feature extraction network for feature extraction to obtain the detection image feature information; The detected image feature information is input into the local semantic recognition network for semantic recognition to obtain the local feature information corresponding to each object.

3. The image recognition method according to claim 1, characterized in that, The process of performing object detection on the image to be processed to obtain an object detection image includes: The image to be processed is input into a target object detection network for object detection to obtain the object detection image.

4. The image recognition method according to claim 1, characterized in that, The method further includes: The sample detection image, the annotation type information corresponding to the sample detection image, and the annotation position information corresponding to each sample object in the sample detection image are obtained. The sample detection image is an object image of at least two sample objects located in the same connected region in the sample image. The sample detection image is input into the first network to be trained for feature extraction to obtain multiple sample local feature information and training detection location information corresponding to each sample local feature information. The sample local feature information is the feature information corresponding to each of the at least two sample objects. The local feature information of the multiple samples is input into the second network to be trained for feature recombination to obtain training recombination feature information; The training recombination feature information is input into the third network to be trained for type recognition, thereby obtaining the training type information corresponding to the sample detection image; Based on the training type information, the annotation type information, the annotation location information, and the training detection location information, the target loss information is determined; Based on the target loss information, the first network to be trained, the second network to be trained, and the third network to be trained are trained to obtain the local feature extraction network, the local feature reconstruction network, and the image recognition network. The local feature extraction network includes the location feature extraction network and the feature fusion network.

5. The image recognition method according to claim 4, characterized in that, The determination of target loss information based on the training type information, the annotation type information, the annotation location information, and the training detection location information includes: Based on the training type information and the annotation type information, the classification loss information is determined; Based on the labeled location information and the training detection location information, determine the location loss information; The target loss information is determined based on the classification loss information and the location loss information.

6. An image recognition device, characterized in that, The device includes: An object detection module is used to perform object detection on the image to be processed and obtain an object detection image, wherein the object detection image is an object image of at least two objects located in the same connected region in the image to be processed. The feature extraction module is used to input the object detection image into a local feature extraction network for feature extraction to obtain multiple local feature information, wherein the local feature information is the feature information corresponding to each of the at least two objects; The feature reconstruction module is used to input the detection image feature information corresponding to the object detection image into the position feature extraction network for position feature extraction, thereby obtaining position calibration information corresponding to each object image, wherein the position calibration information is the calibration information of the detection position of each object image in the object detection image; input the multiple local feature information and the position calibration information corresponding to each object into the feature fusion network, and determine the target distance between each local feature information and the preset starting fusion position based on the detection position information and the position calibration information corresponding to each local feature information; and perform feature fusion on the multiple local feature information based on the target distance to obtain reconstructed feature information. The type recognition module is used to input the recombined feature information into an image recognition network for type recognition, so as to obtain the target type information corresponding to the object detection image.

7. The apparatus according to claim 6, characterized in that, The local feature extraction network includes an image feature extraction network and a local semantic recognition network, and the device includes: An image feature extraction unit is used to input the object detection image into an image feature extraction network for feature extraction to obtain the detection image feature information; The semantic recognition unit is used to input the detected image feature information into the local semantic recognition network for semantic recognition, and obtain the local feature information corresponding to each object.

8. The apparatus according to claim 6, characterized in that, The object detection module includes: The object detection subunit is used to input the image to be processed into the target object detection network for object detection, and obtain the object detection image.

9. The apparatus according to claim 6, characterized in that, The device further includes: The sample information acquisition module is used to acquire the sample detection image, the annotation type information corresponding to the sample detection image, and the annotation position information corresponding to each sample object in the sample detection image. The sample detection image is the object image of at least two sample objects located in the same connected region in the sample image. The sample feature extraction module is used to input the sample detection image into the first network to be trained for feature extraction, and to obtain multiple sample local feature information and training detection location information corresponding to each sample local feature information. The sample local feature information is the feature information corresponding to each of at least two sample objects. The sample feature recombination module is used to input local feature information of multiple samples into the second network to be trained for feature recombination to obtain training recombination feature information. The training type recognition module is used to input the training recombination feature information into the third network to be trained for type recognition, and obtain the training type information corresponding to the sample detection image; The target loss calculation module is used to determine the target loss information based on training type information, annotation type information, annotation location information, and training detection location information; The model training module is used to train the first network to be trained, the second network to be trained, and the third network to be trained based on the target loss information, so as to obtain a local feature extraction network, a local feature reconstruction network, and an image recognition network. The local feature extraction network includes the location feature extraction network and the feature fusion network.

10. The apparatus according to claim 9, characterized in that, The target loss calculation module includes: The classification loss calculation unit is used to determine the classification loss information based on the training type information and the label type information; The location loss calculation unit is used to determine the location loss information based on the labeled location information and the training and detection location information; The target loss determination unit is used to determine the target loss information based on the classification loss information and the location loss information.

11. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement an image recognition method as described in any one of claims 1-5.

12. A computer-readable storage medium, characterized in that, The storage medium includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement an image recognition method as described in any one of claims 1-5.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the image recognition method according to any one of claims 1-5.