An image recognition method, device, equipment, and storage medium

A two-branch deep learning model with Gaussian loss function improves image recognition by accurately separating connected target objects, enhancing precision and efficiency in identifying individual objects.

CN115222939BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210627140.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-07-15
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

When existing deep learning models recognize that multiple target objects are adhered in the image, they cannot effectively segment out individual target objects, resulting in low recognition accuracy and low efficiency.

Method used

The recognition model with a two-branch structure is adopted. The first branch is used to accurately identify the target object, and the second branch is used to segment the adhesion area. The model is trained in combination with the Gaussian loss function to improve the accuracy and efficiency of target object recognition.

Benefits of technology

Automatic segmentation of the adhered target objects in the image is realized, which improves the recognition accuracy and efficiency, and avoids the need for manual segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115222939B_ABST
    Figure CN115222939B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image recognition method, apparatus, device, and storage medium, relating to the technical field of image processing, and particularly to the technical field of computer vision. The specific implementation solution is as follows: obtaining an image to be recognized; recognizing a target object in the image to be recognized according to a first recognition model to obtain a first recognition result; cropping the target object image according to the first recognition result; and segmenting an adhesion region in the target object image according to a second recognition model to obtain a second recognition result. The image recognition method, apparatus, device, and storage medium provided by the present disclosure can improve the efficiency and accuracy of image recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and particularly to an image recognition method, apparatus, device, and storage medium in the field of computer vision technologies. Background Art

[0002] With the rapid development of artificial intelligence technologies, image recognition is often required in various industries, such as recognizing target objects in images and the position information of the target objects in the images. Existing technologies generally use manual methods or deep learning models for image recognition. Summary of the Invention

[0003] The present disclosure provides an image recognition method, apparatus, device, and storage medium for improving the accuracy of image recognition.

[0004] According to one aspect of the present disclosure, there is provided an image recognition method, which includes: obtaining an image to be recognized; recognizing a target object in the image to be recognized according to a first recognition model to obtain a first recognition result; cropping to obtain a target object image according to the first recognition result; and segmenting an adhesion region in the target object image according to a second recognition model to obtain a second recognition result.

[0005] According to another aspect of the present disclosure, there is provided an image recognition apparatus, which includes: an obtaining module for obtaining an image to be recognized; a recognition module for recognizing a target object in the image to be recognized according to a first recognition model to obtain a first recognition result; a cropping module for cropping to obtain a target object image according to the first recognition result; and a segmentation module for segmenting an adhesion region in the target object image according to a second recognition model to obtain a second recognition result.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, enable the at least one processor to execute the method of the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method of the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which, when executed by a processor, implements the method of the present disclosure.

[0009] An image recognition method, device, equipment and storage medium provided by the present disclosure can improve the efficiency of image recognition and the accuracy of image recognition results.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 is a schematic diagram of an application scenario of image recognition in the field of intelligent transportation;

[0013] Figure 2 is a schematic flowchart of an image recognition method according to the first embodiment of the present disclosure;

[0014] Figure 3 is a schematic flowchart of an image recognition method according to the second embodiment of the present disclosure;

[0015] Figure 4 is a schematic flowchart of an image recognition method according to the third embodiment of the present disclosure;

[0016] Figure 5 is a schematic flowchart of an image recognition method according to the seventh embodiment of the present disclosure;

[0017] Figure 6 is a schematic diagram of an application scenario of an image recognition method according to the seventh embodiment of the present disclosure Figure 1 ;

[0018] Figure 7 is a schematic diagram of an application scenario of an image recognition method according to the seventh embodiment of the present disclosure Figure 2 ;

[0019] Figure 8 is a schematic structural diagram of an image recognition device according to the eighth embodiment of the present disclosure;

[0020] Figure 9 is a block diagram of an electronic device for implementing the image recognition method of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0022] With the rapid development of artificial intelligence technology, it is often necessary to identify target objects in images in various industries. For example, in the field of intelligent transportation, it is often necessary to identify various traffic identifiers in the road images collected; in the field of intelligent agriculture, it is often necessary to identify various crops in the farm images collected; in the field of intelligent healthcare, it is often necessary to identify various diseased parts in the medical images collected, etc. The prior art generally uses deep learning models for image recognition. However, if there are multiple target objects in the image and the multiple target objects are adhered, the general deep learning model can only identify the adhered area of the target objects and cannot segment the adhered area of the target objects to obtain the recognition results of individual target objects. Figure 1 is a schematic diagram of an application scenario of image recognition in the field of intelligent transportation, as Figure 1 shown. Taking the field of intelligent transportation as an example, the zebra crossings at general intersections may be adhered. As Figure 1 the zebra crossing a and zebra crossing b in Figure 1 If the zebra crossings in

[0023] Figure 2 is a schematic flowchart of an image recognition method according to the first embodiment of the present disclosure, as Figure 2 shown. The method mainly includes:

[0024] Step S101, obtaining an image to be recognized.

[0025] In this embodiment, first, an image to be recognized needs to be obtained. According to different application scenarios, the image to be recognized can be an image in any scenario. For example, in the field of intelligent transportation, the image to be recognized can be an aerial view of a road; in the field of intelligent agriculture, the image to be recognized can be a farm surveillance image; in the field of intelligent healthcare, the image to be recognized can be a medical imaging picture, etc. The image to be recognized can also be other types of images, and the present disclosure does not limit the image to be recognized.

[0026] In an implementable embodiment, the image to be recognized may be a high-precision point cloud map generated based on point cloud data. The image to be recognized may be directly collected by an image acquisition device such as a camera, a scanner, etc., or the monitoring video related to the application scenario may be collected first, and then the monitoring video is converted frame by frame into the image to be recognized. The present disclosure does not limit the acquisition method of the image to be recognized.

[0027] Step S102: Recognize the target object in the image to be recognized according to the first recognition model to obtain a first recognition result.

[0028] In this embodiment, after obtaining the image to be recognized, the image to be recognized needs to be input into the first recognition model to recognize the target object in the image to be recognized, so as to obtain a first recognition result. Specifically, the first recognition result may include the name of the target object, the position information of the target object, etc.

[0029] In an implementable embodiment, a training set may be obtained first. The training set includes sample images that have been labeled with target objects, and then the training set is input into a deep learning model for training to obtain a first recognition model. Specifically, the deep learning model may be a fully convolutional network (FCN), a Unet model, a Deeplab model, etc. The Unet model and the Deeplab model are both semantic segmentation models proposed based on the FCN.

[0030] Step S103: Crop the target object image according to the first recognition result.

[0031] In this embodiment, after obtaining the first recognition result, the target object image needs to be cropped from the image to be recognized processed by the first recognition model according to the position information of the target object in the first recognition result.

[0032] In an implementable embodiment, the image to be recognized processed by the first recognition model may be first subjected to image morphological operations, such as removing recognition noise, etc., and then the target object image is cropped from the image to be recognized processed by the first recognition model. Specifically, methods such as neighborhood averaging method, median filtering, and low-pass filtering may be used to remove recognition noise.

[0033] In an implementable embodiment, since the first recognition model performs coarser-grained recognition on the image to be recognized, the first recognition result may not be accurate. For example, the range corresponding to the position information of the target object in the first recognition result may not fully contain the target object. Therefore, the range corresponding to the position information of the target object in the first recognition result can be expanded, for example, expanded by 50 pixel ranges, and then the expanded range can be cropped out as the target object image. This can ensure that if the position information of the target object in the first recognition result is inaccurate, the cropped target object image can still fully contain the target object. Among them, the expanded pixel range can be preset according to the actual application scenario, and the present disclosure does not limit it.

[0034] Step S104, segment the adhesion area in the target object image according to the second recognition model to obtain a second recognition result.

[0035] In this embodiment, after cropping the target object image, the target object image needs to be input into the second recognition model to segment the adhesion area in the target object image to obtain a second recognition result. Specifically, there may be multiple target objects in the image to be recognized and the multiple target objects may be adhered to each other. Therefore, there may also be a situation where the target objects in the target object image obtained by the first recognition model are adhered to each other, and the second recognition model needs to segment the area where the target objects in the target object image are adhered to each other.

[0036] In an implementable embodiment, a deep learning model can be trained using a sample image with the adhesion area already segmented, so as to obtain the second recognition model. Among them, the deep learning model can be an FCN, Unet model, Deeplab model, etc. The Unet model and Deeplab model are both semantic segmentation models proposed based on the FCN.

[0037] In the first embodiment of the present disclosure, first, according to the first recognition model, the target object in the image to be recognized is recognized and cropped to obtain the target object image, and then, according to the second recognition model, the adhesion area in the target object image is segmented to obtain the second recognition result. In this way, the second recognition result obtained only contains individual target objects, improving the accuracy of image recognition. In addition, the second recognition model automatically segments the adhesion area, improving the efficiency of image recognition.

[0038] Figure 3 is a schematic flowchart of an image recognition method according to the second embodiment of the present disclosure. As Figure 3 shown, the second recognition model is obtained in the following manner:

[0039] Step S201: Obtain a first training sample set and a second training sample set. The first training sample set includes first sample images with the target object already annotated, and the second training sample set includes second sample images with the adhesion regions already segmented.

[0040] In this embodiment, the second recognition model has a dual-branch structure. The first branch is used to accurately recognize the target object in the target object image, and the second branch is used to segment the adhesion regions in the target object image. Therefore, it is first necessary to obtain the first training sample set and the second training sample set for training the second recognition model. Among them, the first training sample set includes first sample images with the target object already annotated, and the first training sample set is used to train the first branch; the second training sample set includes second sample images with the adhesion regions already segmented, and the second training sample set is used to train the second branch.

[0041] In an implementable manner, various annotation tools, such as Labelimg or Labelme, etc., can be used to annotate the target object in the first sample images and segment the adhesion regions in the second sample images, so as to obtain the first training sample set and the second training sample set.

[0042] Step S202: Train the deep learning model according to the Gaussian loss function and the first training sample set to obtain the first branch of the second recognition model.

[0043] In this embodiment, the first training sample set can be directly input into the deep learning model, and the deep learning model is trained in combination with the Gaussian loss function. The trained deep learning model is used as the first branch of the second recognition model. Specifically, the Gaussian loss function is a loss function that strengthens the weights of the outer edges of the target object. During training, the parameters of the deep learning model can be adjusted according to the Gaussian loss function until the training result reaches the preset conditions. The use of the Gaussian loss function can further improve the recognition accuracy of the first branch for the target object.

[0044] In an implementable manner, the deep learning model can be an FCN, Unet model, Deeplab model, etc. The Unet model and the Deeplab model are both semantic segmentation models proposed based on the FCN.

[0045] Step S203: Add a segmentation head after the backbone network of the first branch to obtain the initial second branch of the second recognition model. The segmentation head is the part of the deep learning model used for segmentation prediction.

[0046] In this embodiment, after the first branch is trained, a segmentation head needs to be added after the backbone network of the first branch to obtain the initial second branch of the second recognition model. Specifically, the first branch is trained by a deep learning model, and a deep learning model generally includes a backbone network and a segmentation head. The backbone network is the part of the deep learning model for feature extraction, and the segmentation head is the part of the deep learning model for segmentation prediction using the features extracted by the backbone network. Adding a segmentation head after the backbone network of the first branch is equivalent to having two segmentation heads in the deep learning model. These two segmentation heads are connected in parallel and share the backbone network of the deep learning model. The branch composed of the newly added segmentation head and the backbone network of the deep learning model is the initial second branch of the second recognition model.

[0047] In an implementable manner, a segmentation head may include a convolutional layer and a batch normalization layer (BN, Batch Normalization), etc. The composition of the segmentation head corresponding to different deep learning models may be different.

[0048] Step S204: Train the initial second branch according to the second training sample set to obtain the second branch of the second recognition model.

[0049] In this embodiment, after obtaining the initial second branch of the second recognition model, the second training sample set needs to be input into the initial second branch for training, so as to obtain the second branch of the second recognition model.

[0050] In an implementable manner, before training the initial second branch, the already trained first branch needs to be frozen, that is, the first branch does not participate in the training of the initial second branch. This can ensure that when the initial second branch is trained using the second training sample set, the parameters of the first branch will not change, so as not to affect the recognition effect of the first branch on the target object.

[0051] In an implementable manner, when training the initial second branch, the loss function corresponding to the deep learning model can be directly used to adjust the parameters of the initial second branch, or the Gaussian loss function can be used to adjust the parameters corresponding to the initial second branch, so as to improve the accuracy of the second branch in segmenting the adhesion area in the target object image.

[0052] In the second embodiment of the present disclosure, according to the Gaussian loss function and the first training sample set, the deep learning model is trained to obtain the first branch of the second recognition model; then the first branch is transformed to obtain the initial second branch of the second recognition model; finally, the initial second branch is trained according to the second training sample set to obtain the second branch of the second recognition model. The second recognition model obtained in this way has two branches. The first branch is used to accurately recognize the target object in the target object image, and the second branch is used to segment the adhesion area in the target object image, which can avoid the situation of target object adhesion in the image recognition result. Moreover, training the second recognition model in combination with the Gaussian loss function can further improve the accuracy of the recognition result.

[0053] Figure 4 is a schematic flowchart of an image recognition method according to the third embodiment of the present disclosure, as Figure 4 shown. The method obtains the Gaussian loss function in the following manner:

[0054] Step S301, calculate the first Gaussian matrix corresponding to each target object in the first sample image.

[0055] In this embodiment, first, it is necessary to traverse each target object in the first sample image and calculate the first Gaussian matrix corresponding to each target object. Specifically, the first Gaussian matrix is a matrix composed of the weight values corresponding to each pixel point in the first sample image, and the weight value can be calculated from the shortest pixel distance from each pixel point to the outer edge of the target object and the circumscribed rectangle of the target object.

[0056] Step S302, stack multiple first Gaussian matrices to obtain the second Gaussian matrix.

[0057] In this embodiment, after calculating the first Gaussian matrix corresponding to each target object in the first sample image, it is also necessary to stack the obtained multiple first Gaussian matrices to obtain the second Gaussian matrix. Specifically, the weight values corresponding to each pixel point in different first Gaussian matrices can be directly added to obtain the second Gaussian matrix.

[0058] Step S303, perform a translation process on the second Gaussian matrix to obtain the Gaussian coefficient matrix.

[0059] Step S304, adjust the loss function corresponding to the deep learning model according to the Gaussian coefficient matrix to obtain the Gaussian loss function.

[0060] In this embodiment, it is also necessary to perform a translation process on the second Gaussian matrix to obtain the Gaussian coefficient matrix, and then adjust the loss function corresponding to the deep learning model according to the Gaussian coefficient matrix, that is, multiply the Gaussian coefficient matrix by the loss function corresponding to the deep learning model to obtain the Gaussian loss function.

[0061] In an implementable embodiment, since the value range of the weight value is from 0 to 1, if there is only one target object in the first sample image, then the weight value corresponding to the pixel points at infinity in the first sample image is 0, and thus the weight value corresponding to the pixel points at infinity in the second Gaussian matrix is 0. If the second Gaussian matrix is directly multiplied by the loss function corresponding to the deep learning model, then the loss corresponding to the pixel points at infinity in the Gaussian loss function is 0. However, during the training process, it is necessary to ensure that the loss corresponding to the pixel points at infinity is equal to the loss in the deep learning model. This requires the weight value corresponding to the pixel points at infinity in the second Gaussian matrix to be 1. Therefore, it is necessary to perform a translation process on the second Gaussian matrix, that is, add 1 to the weight value corresponding to each pixel point in the second Gaussian matrix, and translate the value range of the weight value of the second Gaussian matrix from between 0 and 1 to between 1 and 2 to obtain the Gaussian coefficient matrix. Then multiply the Gaussian coefficient matrix by the loss function corresponding to the deep learning model to obtain the Gaussian loss function.

[0062] In the third embodiment of the present disclosure, calculate the first Gaussian matrix corresponding to each target object in the first sample image, and stack multiple first Gaussian matrices to obtain the second Gaussian matrix; then perform a translation process on the second Gaussian matrix to obtain the Gaussian coefficient matrix; finally, adjust the loss function corresponding to the deep learning model according to the Gaussian coefficient matrix to obtain the Gaussian loss function. The Gaussian loss function strengthens the weight of the outer edge of the target object. By training the second recognition model using the Gaussian loss function, it can be ensured that the second recognition model has accurate recognition of the target object in the target object image.

[0063] In the fourth embodiment of the present disclosure, step S301 mainly includes:

[0064] Calculate the shape parameter of the target object according to the shortest side of the circumscribed rectangle of the target object; calculate the first Gaussian matrix corresponding to the target object according to the shortest pixel distance from each pixel point in the first sample image to the outer edge of the target object and the shape parameter.

[0065] In this embodiment, first, it is necessary to calculate the shape parameter of the target object according to the shortest side of the circumscribed rectangle of the target object, and then calculate the first Gaussian matrix corresponding to the target object according to the shortest pixel distance from each pixel point in the first sample image to the outer edge of the target object and the shape parameter.

[0066] In an implementable embodiment, the shape parameter of the target object is calculated using the following formula: Among them, var is the shape parameter, w is the width of the circumscribed rectangle, h is the height of the circumscribed rectangle, min(w, h) represents the minimum value between w and h, n is a fixed constant, and α is an adjustment coefficient. Specifically, the value of the fixed constant n is any natural number greater than or equal to 1, and the values of the fixed constant n and the adjustment coefficient α can be set according to the actual situation. For example, if the value of the fixed constant n is set to 4 and the value of the adjustment coefficient α is set to 0.54, the formula for calculating the shape parameter is: var = 0.135 × min(w, h).

[0067] In an implementable manner, the pointPolygonTest function in OpenCV can be called to obtain the shortest pixel distance from each pixel point in the first sample image to the outer edge of the target object. OpenCV is an open-source computer vision library.

[0068] In an implementable manner, after obtaining the shape parameter and the shortest pixel distance, the following formula can be used to calculate the first Gaussian matrix corresponding to the target object: Among them, map is the first Gaussian matrix, dist is the shortest pixel distance, and var is the shape parameter.

[0069] In an implementable manner, the following formula can also be used to calculate the first Gaussian matrix corresponding to the target object: Among them, map is the first Gaussian matrix, and dist is the shortest pixel distance.

[0070] In an implementable manner, the following formula can also be used to calculate the first Gaussian matrix corresponding to the target object: Among them, map is the first Gaussian matrix, dist is the shortest pixel distance, N is a constant greater than 1, r is the distance threshold, and the values of N and r can be set according to the actual situation.

[0071] In the fourth and fifth embodiments of the present disclosure, first, the shape parameter of the target object is calculated according to the shortest side of the circumscribed rectangle of the target object, and then, according to the shortest pixel distance from each pixel point in the first sample image to the outer edge of the target object and the shape parameter, the first Gaussian matrix corresponding to the target object is calculated. The first Gaussian matrix assigns corresponding weight values to each pixel point, strengthens the weights of the pixel points at the outer edge of the target object, and thus can improve the recognition accuracy of the second recognition model trained using the Gaussian loss function for the target object.

[0072] In the sixth embodiment of the present disclosure, step S102 mainly includes:

[0073] According to the first recognition model, feature extraction is performed on the image to be recognized to obtain the feature map of the image to be recognized; according to the feature map, the target object in the image to be recognized is recognized to obtain the first recognition result; the first recognition result includes the position information of the target object in the image to be recognized; correspondingly, step S103 mainly includes: according to the position information, the target object image is cropped.

[0074] In this embodiment, when recognizing the image to be recognized according to the first recognition model, first, the features of the image to be recognized need to be extracted to obtain the feature map of the image to be recognized, and then according to the feature map, the target object in the image to be recognized is recognized to obtain the first recognition result. The first recognition result includes the position information of the target object in the image to be recognized, and the target object image can be cropped according to the position information.

[0075] In an implementable manner, the first recognition model can divide the feature map into regional blocks with certain semantic meanings, identify the semantic categories of each regional block, and finally obtain a segmentation image with per-pixel semantic annotation, thereby obtaining the first recognition result.

[0076] In an implementable manner, the range corresponding to the position information can be directly cropped to obtain the target object image, or the range corresponding to the position information can be expanded first, and then the expanded range can be cropped as the target object image.

[0077] In the sixth embodiment of the present disclosure, according to the first recognition model, the feature map of the image to be recognized is extracted, and then according to the feature map, the target object in the image to be recognized is recognized to obtain the first recognition result, which is equivalent to performing a coarser-grained recognition on the image to be recognized, so that the subsequent second recognition model can directly recognize and segment according to the target object image.

[0078] Figure 5 is a schematic flowchart of an image recognition method according to the seventh embodiment of the present disclosure, as Figure 5 shown, step S104 mainly includes:

[0079] Step S401, recognizing the target object in the target object image according to the first branch to obtain the target image.

[0080] In this embodiment, first, the target object in the target object image needs to be recognized according to the first branch of the second recognition model to obtain the target image. Specifically, in the training process of the first branch, the Gaussian loss function is combined. The Gaussian loss function strengthens the weights of the pixel points at the outer edge of the target object. Therefore, the first branch can perform high-precision recognition on the target object image. Compared with the first recognition model, the first branch can recognize the target object in a finer-grained manner, thereby obtaining an accurate target image.

[0081] Step S402: According to the second branch, identify the adhesion boundary line in the target object image, where the adhesion boundary line is the boundary line for dividing the adhesion area.

[0082] In this embodiment, it is also necessary to identify the adhesion boundary line in the target object image according to the second branch of the second recognition model, where the adhesion boundary line is the boundary line for dividing the adhesion area. Specifically, use the second sample image that has already segmented the adhesion area, that is, the second sample image with the adhesion boundary line marked, to train the second branch. Therefore, the second branch can identify the adhesion boundary line for dividing the adhesion area of the target object image.

[0083] Figure 6 It is a schematic diagram of an application scenario of an image recognition method according to the seventh embodiment of the present disclosure Figure 1 , such as Figure 6 shown, in the field of intelligent transportation, the zebra crossings at general intersections may be adhered. For example, Figure 6 the zebra crossing a and zebra crossing b in Figure 6 . If Figure 6 is used as the target object image and input into the second recognition model, then the second branch of the second recognition model will identify the adhesion boundary line in the target object image. For example, identify the adhesion boundary line c in

[0084] The adhesion boundary line c can divide the adhesion area composed of the zebra crossing a and zebra crossing b into separate zebra crossing a and zebra crossing b.

[0085] In this embodiment, after obtaining the accurate target image and the adhesion boundary line, combine the target image with the adhesion boundary line to generate the second recognition result.

[0086] Figure 7 It is a schematic diagram of an application scenario of an image recognition method according to the seventh embodiment of the present disclosure Figure 2 , such as Figure 7 shown. If the first branch identifies the target image d composed of the zebra crossing a and zebra crossing b, and the second branch identifies the adhesion boundary line c, then after combining the target image d with the adhesion boundary line c, separate zebra crossing a and separate zebra crossing b can be obtained, that is, the second recognition result.

[0087] In the seventh embodiment of the present disclosure, first, according to the first branch of the second recognition model, an accurate target image of the target object is recognized; then, according to the second branch of the second recognition model, the adhesion boundary line in the target object image is recognized; finally, the target image and the adhesion boundary line are combined to generate a second recognition result. In this way, the accuracy of the recognition result of the image to be recognized can be improved; in addition, there is no need for manual segmentation of the adhesion area, which can improve the recognition efficiency of the image to be recognized.

[0088] Figure 8 FIG. is a schematic structural diagram of an image recognition device according to the eighth embodiment of the present disclosure, as Figure 8 shown, the device mainly includes:

[0089] An acquisition module 80, configured to acquire an image to be recognized; a recognition module 81, configured to recognize a target object in the image to be recognized according to a first recognition model to obtain a first recognition result; a cropping module 82, configured to crop a target object image according to the first recognition result; a segmentation module 83, configured to segment an adhesion area in the target object image according to a second recognition model to obtain a second recognition result.

[0090] In an implementable manner, the device further includes:

[0091] A training module 84, configured to train the second recognition model. Specifically, the training module 84 mainly includes: an acquisition sub-module 840, configured to acquire a first training sample set and a second training sample set, the first training sample set includes first sample images with the target object already labeled, and the second training sample set includes second sample images with the adhesion area already segmented; a first training sub-module 841, configured to train a deep learning model according to a Gaussian loss function and the first training sample set to obtain the first branch of the second recognition model; an adding sub-module 842, configured to add a segmentation head after the backbone network of the first branch to obtain an initial second branch of the second recognition model, and the segmentation head is a part of the deep learning model for segmentation prediction; a second training sub-module 843, configured to train the initial second branch according to the second training sample set to obtain the second branch of the second recognition model.

[0092] In an implementable manner, the device further includes:

[0093] The calculation module 85 is used to calculate the Gaussian loss function. Specifically, the calculation module 85 mainly includes: a calculation sub-module 850 for calculating a first Gaussian matrix corresponding to each target object in the first sample image; a superposition sub-module 851 for superposing a plurality of first Gaussian matrices to obtain a second Gaussian matrix; a translation sub-module 852 for performing a translation process on the second Gaussian matrix to obtain a Gaussian coefficient matrix; and an adjustment sub-module 853 for adjusting the loss function corresponding to the deep learning model according to the Gaussian coefficient matrix to obtain the Gaussian loss function.

[0094] In an implementable manner, the calculation sub-module 850 mainly includes:

[0095] A first calculation unit 8500 for calculating the shape parameter of the target object according to the shortest side of the circumscribed rectangle of the target object; a second calculation unit 8501 for calculating the first Gaussian matrix corresponding to the target object according to the shortest pixel distance from each pixel point in the first sample image to the outer edge of the target object and the shape parameter.

[0096] In an implementable manner, the first calculation unit 8500 calculates the shape parameter of the target object using the following formula: where var is the shape parameter, w is the width of the circumscribed rectangle, h is the height of the circumscribed rectangle, min(w, h) represents the minimum value between w and h, n is a fixed constant, and α is an adjustment coefficient; the second calculation unit 8501 calculates the first Gaussian matrix corresponding to the target object using the following formula: where map is the first Gaussian matrix, dist is the shortest pixel distance, and var is the shape parameter.

[0097] In an implementable manner, the recognition module 81 mainly includes:

[0098] A feature extraction sub-module 810 for extracting features of the image to be recognized according to the first recognition model to obtain a feature map of the image to be recognized; a first recognition sub-module 811 for recognizing the target object in the image to be recognized according to the feature map to obtain a first recognition result; the first recognition result includes the position information of the target object in the image to be recognized; correspondingly, the cropping module 82 is further used to crop the target object image according to the position information.

[0099] In an implementable manner, the segmentation module 83 mainly includes:

[0100] The second recognition sub-module 830 is configured to recognize a target object in the target object image according to the first branch to obtain a target image; the third recognition sub-module 831 is configured to recognize an adhesion boundary line in the target object image according to the second branch, where the adhesion boundary line is the boundary line for dividing the adhesion region; the generation sub-module 832 is configured to generate a second recognition result according to the target image and the adhesion boundary line.

[0101] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0102] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0103] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0104] As Figure 9 shown, the device 900 includes a computing unit 901, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0105] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0106] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as an image recognition method. For example, in some embodiments, an image recognition method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the image recognition method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute an image recognition method in any other suitable manner (e.g., by means of firmware).

[0107] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0108] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed as an independent software package partially on the machine and partially on a remote machine, or executed entirely on a remote machine or server.

[0109] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0111] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0112] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0113] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0114] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An image recognition method, comprising: Obtaining an image to be recognized; Recognizing a target object in the image to be recognized according to a first recognition model to obtain a first recognition result; Cropping to obtain a target object image according to the first recognition result; Segmenting an adhesion region in the target object image according to a second recognition model to obtain a second recognition result; Wherein, the second recognition model is obtained by the following method: Obtaining a first training sample set and a second training sample set, the first training sample set includes first sample images with labeled target objects, and the second training sample set includes second sample images with segmented adhesion regions; Training a deep learning model according to a Gaussian loss function and the first training sample set to obtain a first branch of the second recognition model; Adding a segmentation head after the backbone network of the first branch to obtain an initial second branch of the second recognition model, and the segmentation head is a part of the deep learning model for segmentation prediction; Training the initial second branch according to the second training sample set to obtain a second branch of the second recognition model; Wherein, the Gaussian loss function is obtained by the following method: Calculating a first Gaussian matrix corresponding to each target object in the first sample image; Superposing a plurality of the first Gaussian matrices to obtain a second Gaussian matrix; Performing a translation process on the second Gaussian matrix to obtain a Gaussian coefficient matrix; Adjusting a loss function corresponding to the deep learning model according to the Gaussian coefficient matrix to obtain the Gaussian loss function.

2. The method according to claim 1, wherein The calculating the first Gaussian matrix corresponding to each target object in the first sample image includes: Calculating a shape parameter of the target object according to the shortest side of the circumscribed rectangle of the target object; Calculating the first Gaussian matrix corresponding to the target object according to the shortest pixel distance from each pixel point in the first sample image to the outer edge of the target object and the shape parameter.

3. The method according to claim 2, wherein Using the following formula to calculate the shape parameter of the target object: , where var is the shape parameter, w is the width of the circumscribed rectangle, and h is the height of the circumscribed rectangle, represents the minimum value between w and h, n is a fixed constant, is an adjustment coefficient; And, using the following formula to calculate the first Gaussian matrix corresponding to the target object: , where map is the first Gaussian matrix, is the shortest pixel distance, is the shape parameter.

4. The method according to any one of claims 1 to 3, wherein The recognizing a target object in the image to be recognized according to a first recognition model to obtain a first recognition result includes: Performing feature extraction on the image to be recognized according to the first recognition model to obtain a feature map of the image to be recognized; Recognizing a target object in the image to be recognized according to the feature map to obtain the first recognition result; The first recognition result includes position information of the target object in the image to be recognized; the cropping to obtain a target object image according to the first recognition result includes: cropping to obtain the target object image according to the position information.

5. The method according to claim 1, wherein The segmenting an adhesion region in the target object image according to a second recognition model to obtain a second recognition result includes: Recognizing a target object in the target object image according to the first branch to obtain a target image; Recognizing an adhesion boundary line in the target object image according to the second branch, and the adhesion boundary line is a boundary line for segmenting the adhesion region; Generate the second recognition result according to the target image and the adhesion junction line.

6. An image recognition device, comprising: An acquisition module, configured to acquire an image to be recognized; A recognition module, configured to recognize a target object in the image to be recognized according to a first recognition model, and obtain a first recognition result; A cropping module, configured to crop a target object image according to the first recognition result; A segmentation module, configured to segment an adhesion region in the target object image according to a second recognition model, and obtain a second recognition result; Wherein, the second recognition model is obtained by the following method: Acquire a first training sample set and a second training sample set, where the first training sample set includes first sample images with labeled target objects, and the second training sample set includes second sample images with segmented adhesion regions; Train a deep learning model according to a Gaussian loss function and the first training sample set to obtain a first branch of the second recognition model; Add a segmentation head after the backbone network of the first branch to obtain an initial second branch of the second recognition model, where the segmentation head is a part for segmentation prediction in the deep learning model; Train the initial second branch according to the second training sample set to obtain a second branch of the second recognition model; Wherein, the Gaussian loss function is obtained by the following method: Calculate a first Gaussian matrix corresponding to each target object in the first sample image; Superimpose multiple first Gaussian matrices to obtain a second Gaussian matrix; Perform a translation process on the second Gaussian matrix to obtain a Gaussian coefficient matrix; Adjust the loss function corresponding to the deep learning model according to the Gaussian coefficient matrix to obtain the Gaussian loss function.

7. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-5.

8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.

9. A computer program product, comprising a computer program, where the computer program implements the method according to any one of claims 1-5 when executed by a processor.

Citation Information

Patent Citations

  • Oral cavity image multi-tissue full-automatic segmentation method and system

    CN113223010A

  • Method, device and equipment for identifying lane line, storage medium and unmanned vehicle

    CN113392793A