Image recognition method, device, equipment and storage medium

By converting images to HSV color space, grayscale images, and Gamma-transformed images to generate target images, the problem of image recognition accuracy under illumination is solved, and the accuracy of recognition results is improved.

CN116092171BActive Publication Date: 2026-07-24CHINA UNITED NETWORK COMM GRP CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA UNITED NETWORK COMM GRP CO LTD
Filing Date
2023-01-12
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

During image recognition, lighting conditions can cause reflective spots in the image, reducing the accuracy of the recognition results.

Method used

By converting images to HSV color space, grayscale images, and Gamma-transformed images, and combining the color values ​​of these images to generate target images, the reference amount for object recognition is increased, and the influence of reflective points is reduced.

Benefits of technology

It improves the accuracy of image recognition results and reduces the impact of reflective points on the recognition process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116092171B_ABST
    Figure CN116092171B_ABST
Patent Text Reader

Abstract

The application provides a kind of image recognition method, device, equipment and storage medium, it is related to information technology field, for improving the accuracy of image recognition result.The method comprises: obtaining a first image, the first image includes the image of target object.Processing first image, generate first image set, first image set includes: first image, second image and third image, first image is the image of first image after HSV color space conversion, second image is the gray image of first image, third image is the image of first image after Gamma transformation.Target image is generated, the color value of each pixel point in target image is determined based on first image, first image, second image and third image.Processing target image, determine target object in target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular to an image recognition method, apparatus, device, and storage medium. Background Technology

[0002] With the development of image recognition technology, devices using this technology (such as servers and terminals) have gradually entered people's lives and are widely used in fields such as medicine, entertainment, and education. For example, servers can recognize images of human eyes.

[0003] Currently, when a server identifies human eye images, it needs to acquire the captured image of the human eye. The server then classifies the pixels in the image to identify different objects within it. However, in this technical solution, the presence of reflective points in the captured image due to lighting conditions during the shooting process can affect the server's recognition of the eye image, reducing the accuracy of the image recognition results. Summary of the Invention

[0004] This application provides an image recognition method, apparatus, device, and storage medium to improve the accuracy of image recognition results.

[0005] To achieve the above objectives, this application adopts the following technical solution:

[0006] In a first aspect, this application provides an image recognition method, comprising: an image recognition device (hereinafter referred to as "recognition device") acquiring a first image, the first image including an image of a target object; the recognition device processing the first image to generate a first image set, the first image set including: a first type of image, a second type of image, and a third type of image, wherein the first type of image is an image of the first image after HSV color space conversion, the second type of image is a grayscale image of the first image, and the third type of image is an image of the first image after gamma transformation; the recognition device generating a target image, wherein the color value of each pixel in the target image is determined based on the first image, the first type of image, the second type of image, and the third type of image; and the recognition device processing the target image to determine the target object in the target image.

[0007] Optionally, the image recognition method further includes: the recognition device segmenting the first image to generate a second image set, the second image set including at least one second image. For each of the at least one second image, the recognition device determines a weight value of the second image based on the color value of the pixels in the second image, the weight value of the second image indicating the degree of association between the second image and the image of the target object. The recognition device stitches together the at least one second image to generate a third image, the third image being labeled with the weight value of each second image. The above-mentioned method of "the recognition device generating a target image" includes: the recognition device generating a target image based on the first image, a first type of image, a second type of image, a third type of image, and the third image, the target image also being labeled with the weight value of each second image.

[0008] Optionally, the image recognition method further includes: the recognition device processes the first type of image, the second type of image, and the third type of image to generate a fourth image, wherein the fourth image is an edge image of the target object. The method of "the recognition device generating a target image" includes: the recognition device generating a target image based on the first image, the first type of image, the second type of image, the third type of image, the third image, and the fourth image, wherein the target image includes the fourth image.

[0009] Optionally, the image recognition method further includes: the recognition device marking the target object in the first image to generate a fifth image.

[0010] Optionally, the target object may include at least one of the following: eyelid, iris, sclera, pupil, and canthus.

[0011] Secondly, this application provides an image recognition device, which includes an acquisition module and a processing module.

[0012] The module includes an acquisition module for acquiring a first image, which includes an image of the target object. The processing module processes the first image to generate a first image set, which includes three types of images: a first type of image, a second type of image, and a third type of image. The first type of image is the first image after HSV color space conversion; the second type of image is a grayscale version of the first image; and the third type of image is the first image after gamma transformation. The processing module also generates a target image, where the color value of each pixel is determined based on the first image, the first type of image, the second type of image, and the third type of image. Finally, the processing module processes the target image to determine the target object within it.

[0013] Optionally, the processing module is further configured to segment the first image to generate a second image set, the second image set including at least one second image. The processing module is further configured to, for each of the at least one second image, determine a weight value for the second image based on the color values ​​of pixels in the second image, the weight value of the second image indicating the degree of association between the second image and the image of the target object. The processing module is further configured to stitch the at least one second image to generate a third image, the third image being labeled with the weight value of each second image. Specifically, the processing module is configured to generate a target image based on the first image, the first type of image, the second type of image, the third type of image, and the third image, the target image also being labeled with the weight value of each second image.

[0014] Optionally, the processing module is further configured to process the first type of image, the second type of image, and the third type of image to generate a fourth image, which is an edge image of the target object. Specifically, the processing module is configured to generate a target image, including the fourth image, based on the first image, the first type of image, the second type of image, the third type of image, the third image, and the fourth image.

[0015] Optionally, the processing module is also used to mark the target object in the first image to generate a fifth image.

[0016] Optionally, the target object may include at least one of the following: eyelid, iris, sclera, pupil, and canthus.

[0017] Thirdly, this application provides an image recognition device, the device comprising: a processor and a memory coupled together, the memory for storing one or more programs, the one or more programs including computer-executable instructions, wherein when the image recognition device is running, the processor executes the computer-executable instructions stored in the memory to implement the image recognition method as described in the first aspect and any possible implementation thereof.

[0018] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the image recognition method described in the first aspect and any possible implementation thereof.

[0019] Fifthly, this application provides a computer program product applied to a server, the computer program product including computer instructions, which, when executed on the server, enable the server to implement the image recognition method as described in the first aspect and any possible implementation thereof.

[0020] The technical problems that the image recognition device, equipment, computer storage medium or computer program product can solve and the technical effects it can achieve can be found in the technical problems and effects solved in the first aspect above, and will not be repeated here.

[0021] The technical solution provided in this application offers at least the following beneficial effects: The recognition device can acquire a first image, which includes an image of the target object. Then, the recognition device processes the first image to generate a first image set, which includes: a first type of image, a second type of image, and a third type of image. The first type of image is the first image converted to HSV color space, the second type of image is a grayscale image of the first image, and the third type of image is the first image converted to Gamma. Next, the recognition device generates a target image, where the color value of each pixel is determined based on the first image, the first type of image, the second type of image, and the third type of image. Then, the recognition device processes the target image to determine the target object within it. In other words, the recognition device can add the color values ​​of the pixels converted to HSV color space, the pixel values ​​of the grayscale image, and the color values ​​of the pixels converted to Gamma to the color values ​​of each pixel in the first image. This increases the reference quantity for the recognized object, reduces the influence of reflective points on the recognition process, and improves the accuracy of the image recognition results. Attached Figure Description

[0022] Figure 1 A schematic diagram of a communication system provided in an embodiment of this application;

[0023] Figure 2 A schematic diagram of a server system architecture provided in this application embodiment;

[0024] Figure 3 A flowchart illustrating an image recognition method provided in an embodiment of this application;

[0025] Figure 4 A flowchart illustrating another image recognition method provided in an embodiment of this application;

[0026] Figure 5 A schematic diagram of an image example provided in this application embodiment;

[0027] Figure 6 A flowchart illustrating another image recognition method provided in an embodiment of this application;

[0028] Figure 7 A flowchart illustrating another image recognition method provided in an embodiment of this application;

[0029] Figure 8A flowchart illustrating another image recognition method provided in an embodiment of this application;

[0030] Figure 9 This is a schematic diagram illustrating an example of image recognition provided in an embodiment of this application;

[0031] Figure 10 A schematic diagram of the structure of an image recognition device provided in an embodiment of this application;

[0032] Figure 11 A schematic diagram of the structure of an image recognition device provided in an embodiment of this application;

[0033] Figure 12 A conceptual partial view of a computer program product provided for an embodiment of this application. Detailed Implementation

[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] In this article, the character " / " generally indicates that the objects before and after it are in an "or" relationship. For example, A / B can be understood as A or B.

[0036] The terms “first” and “second” in the specification and claims of this application are used to distinguish different objects, rather than to describe a specific order of objects.

[0037] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include other steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or device.

[0038] Furthermore, in the embodiments of this application, the words "exemplary" or "for example" are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design that is described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design options. Specifically, the use of the words "exemplary" or "for example" is intended to present concepts in a concrete manner.

[0039] Before providing a detailed description of the image recognition method provided in the embodiments of this application, the implementation environment and application scenarios of the embodiments of this application will be introduced first.

[0040] First, the application scenarios of the embodiments of this application will be introduced.

[0041] The image recognition method of this application embodiment is applied to image recognition scenarios. In related technologies, when a server recognizes a human eye image, the server needs to acquire a captured human eye image. Then, the server classifies the pixels in the human eye image to identify different objects in the human eye image.

[0042] For example, the server acquires images of human eyes through a camera and uses a Mobile Network (MobileNet) model in a neural network to identify objects such as eyelids, irises, and sclera in the human eye image.

[0043] In summary, the current technical solution suffers from the presence of reflective points in the captured human eye image due to lighting conditions during the shooting process. This affects the server's recognition of the human eye image and reduces the accuracy of the recognition results.

[0044] To address the aforementioned problems, this application provides an image recognition method. The recognition device acquires an image including a first image, where the first image includes a target object. Then, the recognition device converts the first image to a Hue Saturation Value (HSV) color space to generate a first-type image; furthermore, the recognition device converts the first image to a grayscale image to generate a second-type image; and finally, the recognition device performs a Gamma transformation on the first image to generate a third-type image. Next, the recognition device adds the color values ​​of pixels at the same positions in the first image, the first-type image, the second-type image, and the third-type image to generate a target image. Then, the recognition device determines the target object in the target image based on the color value of each pixel in the target image. In other words, the recognition device can add the color values ​​of pixels converted from the HSV color space, the color values ​​of pixels in the grayscale image, and the color values ​​of pixels after Gamma transformation to the color values ​​of pixels in the first image. This increases the reference amount for the identified object, reduces the influence of reflective points on the recognition process, and improves the accuracy of the image recognition results.

[0045] The implementation environment of the embodiments of this application is described below.

[0046] like Figure 1The diagram illustrates a communication system provided in an embodiment of this application. The communication system includes: an identification device (such as server 101) and at least one terminal (such as terminal 102 or terminal 103). Server 101 can communicate with terminal 102 (or terminal 103) via wired / wireless communication.

[0047] Terminal 102 can acquire images by taking a picture. Terminal 102 can also send the captured images to server 101. Server 101 can receive images from terminal 102. Furthermore, server 101 can identify objects in the images and send the identification results to terminal 102. Similarly, terminal 103 can acquire images by taking a picture. Terminal 103 can also send the captured images to server 101. Server 101 can receive images from terminal 103. Furthermore, server 101 can identify objects in the images and send the identification results to terminal 103.

[0048] The terminal (such as terminal 102, terminal 103) can be a device with transmitting and receiving functions and a camera function. The terminal can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; it can also be deployed on water (such as on ships); and it can also be deployed in the air (e.g., on airplanes, balloons, and satellites). Terminals include handheld devices, vehicle-mounted devices, wearable devices, or computing devices with wireless communication and camera functions. For example, the terminal can be a mobile phone, a tablet computer, or a computer with wireless transmitting and receiving and camera functions. The terminal device can also be a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in autonomous driving, a wireless terminal in telemedicine, a wireless terminal in a smart grid, a wireless terminal in a smart city, a wireless terminal in a smart home, etc.

[0049] A server (such as server 101) can be a physical server or a cloud server.

[0050] It should be noted that, as Figure 2 As shown, the system architecture of a server (such as server 101) may include: a light cancellation module, an edge enhancement module, and a position relationship module.

[0051] The illumination removal module corrects glare points in the image. The edge enhancement module filters edge information belonging to objects within the image. The positional relationship module determines image elements related to the objects.

[0052] After introducing the application scenarios and implementation environments of the embodiments of this application, the image recognition method provided by the embodiments of this application will be described in detail below in conjunction with the above implementation environment.

[0053] The methods in the following embodiments can all be implemented in the above-described application scenarios and implementation environments. The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0054] Figure 3 This is a schematic flowchart illustrating an image recognition method provided in an embodiment of this application. Figure 3 As shown, the method may include: S301-S304.

[0055] S301, The server obtains the first image.

[0056] The first image includes an image of the target object.

[0057] In one possible design, the target object may include at least one of the following: eyelid, iris, sclera, pupil, and canthus.

[0058] In one possible implementation, the server can receive the first image from the terminal.

[0059] In another possible implementation, the server stores a preset database. The server can receive a sixth image from the terminal, identify the sixth image using the preset database, and retrieve the first image from the sixth image. The sixth image includes the first image.

[0060] It should be noted that the embodiments of this application do not limit the preset database. For example, the preset database can be the Dlib open-source database. Another example is the Open Computer Vision (OpenCV) open-source computer vision library. Yet another example is the Image Network (ImageNet) database.

[0061] For example, the server receives a face image captured by a mobile phone (i.e., the sixth image) and uses the Dlib open-source database (i.e., the preset database) to recognize the face image, identifying 68 feature points in the face image. The server selects feature points 37-42 for the right eye and 43-48 for the left eye from the 68 feature points, and uses these 12 feature points to crop the face image, obtaining a 416×416 eye image (i.e., the first image).

[0062] S302. The server processes the first image to generate the first image set.

[0063] The first image set may include: a first type of image. The first type of image is the first image converted to the HSV color space.

[0064] In one possible implementation, the server can convert the red-green-blue (RGB) color space of the first image to the HSV color space to generate the first type of image.

[0065] It should be noted that the size of the first type of image is 416×416×3.

[0066] It should be noted that the process of the server converting the RGB color space of the first image to the HSV color space can be found in the introduction of the conventional technology for converting the RGB color space to the HSV color space, and will not be elaborated here.

[0067] In one possible design, the server can convert the red-green-blue (RGB) color space of the first image to the HSV color space, and extract the H channel image from the converted image to generate the first type of image.

[0068] It should be noted that the size of the first type of image is 416×416×1.

[0069] It should be noted that the process of extracting the H channel from the converted image by the server can be found in the introduction of conventional techniques for extracting the H channel, and will not be elaborated here.

[0070] In this embodiment of the application, the first image set may further include a second type of image. The second type of image is a grayscale image of the first image.

[0071] In one possible implementation, the server can convert the RGB color space of the first image to the grayscale color space to generate the second type of image.

[0072] It should be noted that the process of the server converting the RGB color space of the first image to the Gray color space can be found in the introduction of the conventional technology for converting the RGB color space to the Gray color space, and will not be elaborated here.

[0073] In one possible design, the server can convert the RGB color space of the first image to the grayscale color space, and then perform histogram equalization on the converted image to generate the second type of image.

[0074] It should be noted that the process of the server performing histogram equalization on the converted image can be found in the introduction of conventional image histogram equalization techniques, and will not be elaborated here.

[0075] In this embodiment of the application, the first image set may further include a third type of image. The third type of image is the first image after Gamma transformation.

[0076] In one possible implementation, the server can convert the RGB color space of the first image to a grayscale color space, and perform a Gamma transformation on the color value of each pixel in the converted image to generate a third type of image.

[0077] In one possible design, the color value of a pixel in a third type of image can be represented by Formula 1.

[0078] G = C·I γ Formula 1.

[0079] Where G indicates the color value of a pixel in the third type of image, C is a constant (with a value of 1), I indicates the color value of a pixel in the image after Gray color space conversion, and γ is the gamma factor (with a value of 1). ).

[0080] S303, The server generates the target image.

[0081] The color value of each pixel in the target image is determined based on the first image, the first type of image, the second type of image, and the third type of image.

[0082] For example, the first image is composed of pixel A, pixel B, and pixel C. If the color value of pixel A in the first image is 101, the color value of pixel B is 97, and the color value of pixel C is 31; the color value of pixel A in the first type of image is 52, the color value of pixel B is 27, and the color value of pixel C is 146; the color value of pixel A in the second type of image is 303, the color value of pixel B is 46, and the color value of pixel C is 116; and the color value of pixel A in the third type of image is 152, the color value of pixel B is 307, and the color value of pixel C is 186, then the server generates a target image in which the color value of pixel A is 608, the color value of pixel B is 377, and the color value of pixel C is 479.

[0083] In one possible implementation, the server can perform convolution processing on the first type of images to generate a first vector. Furthermore, the server can also perform convolution processing on the second type of images to generate a second vector. Then, the server multiplies the first vector and the second vector to generate a seventh image. The seventh image has a size of 416×416.

[0084] It should be noted that during the convolution process of the first type of image, the server can multiply the 416×416×1 image along its x-axis by a 1×1 convolution, resulting in a 1×1×416 vector and generating a first vector of size 1×416×1. Similarly, during the convolution process of the second type of image, the server can multiply the 416×416×1 image along its y-axis by a 1×1 convolution, resulting in a 1×1×416 vector and generating a second vector of size 416×1×1.

[0085] In one possible design, the seventh image can be represented by Formula 2.

[0086]

[0087] Where 'a' indicates the seventh image in matrix representation. Used to indicate the second vector, Used to indicate the first vector.

[0088] In one possible implementation, the server can perform element-wise product of the seventh image in matrix form and the third image in matrix form to generate the eighth image.

[0089] In one possible design, the eighth image can be represented by Formula 3.

[0090] G′=a i ⊙G i Formula 3.

[0091] Where G′ is used to indicate the eighth image in matrix representation, a i G is used to indicate the color value of the i-th pixel in a. i Used to indicate the color value of the i-th pixel in G. Where i ∈ {1, 2, ..., 416} 2}

[0092] In one possible implementation, the server can further perform convolution processing on the first, second, and third type images to generate a ninth image. Then, the server can perform a residual concatenation between the eighth image (in matrix representation) and the ninth image (in matrix representation) to generate a fourth type image. This fourth type image is the image after correcting the reflective points in the first image.

[0093] It should be noted that during the convolution process of the first, second, and third type of images, the server can multiply the first, second, and third type of images by a 1×1 convolution to reduce their dimensionality and generate a ninth image with a size of 416×416×1.

[0094] In one possible design, the fourth type of image can be represented by Formula 4.

[0095] L i =G′ i +A′ i Formula 4.

[0096] Among them, L i G is used to indicate the color value of the i-th pixel in the fourth type of image in matrix representation. i A' is used to indicate the color value of the i-th pixel in G'. i Used to indicate the color value of the i-th pixel in the ninth image in matrix representation.

[0097] In one possible implementation, the server can perform sum fusion on the first image in matrix representation and the fourth image in matrix representation to generate the target image.

[0098] In one possible design, after the server performs Sum fusion processing on the first image and the fourth image in matrix representation, the server can perform convolution processing on the Sum fusion-processed image to generate the target image.

[0099] It should be noted that during the convolution process of the image processed by Sum fusion, the server can multiply the image processed by Sum fusion by a 1×1 convolution, resulting in a size of 1×1×4, to generate a target image with a size of 416×416×1.

[0100] In another possible implementation, the server stores a pre-defined neural network model and includes an upsampling layer. The server can use the pre-defined neural network model to identify the first image, generate a fifth type of image, and input the fifth type of image into the upsampling layer for upsampling processing to generate a sixth type of image. Then, the server can perform Sum fusion processing on the sixth type of image (in matrix representation) and the fourth type of image (in matrix representation) to generate the target image.

[0101] It should be noted that the embodiments of this application do not limit the preset neural network model. For example, the preset neural network model can be the MobileNet model with the average pooling layer, fully connected layer, and Softmax layer removed. Another example is the semantic segmentation network (SegNet) model. Yet another example is the convolutional network for biomedical image segmentation (U-Net) model.

[0102] S304. The server processes the target image to determine the target object in the target image.

[0103] In one possible implementation, the server stores a first preset function and preset color values. Based on the color value of each pixel in the target image and the preset color values, the server can classify the pixels in the target image using the first preset function to determine the image of the target object. In other words, the server can identify pixels belonging to the target object from the target image, thereby achieving the recognition of the target object in the target image.

[0104] It should be noted that the first preset function is not limited in the embodiments of this application. For example, the first preset function can be the Softmax function. Another example is that the first preset function can be the hardmax function, or even the sigmoid function.

[0105] For example, the target image consists of pixel A, pixel B, and pixel C. Pixel A has a color value of 123, pixel B has a color value of 55, and pixel C has a color value of 303. If the server stores the Softmax function and a preset color threshold of 100, the server uses the Softmax function to determine that pixel A has a probability of 0.7 of being a pixel in the pupil image, 0.11 of being a pixel in the iris image, and 0.09 of being a pixel in the sclera image; pixel B has a probability of 0.5 of being a pixel in the pupil image, 0.37 of being a pixel in the iris image, and 0.13 of being a pixel in the sclera image; and pixel C has a probability of 0.22 of being a pixel in the pupil image, 0.41 of being a pixel in the iris image, and 0.37 of being a pixel in the sclera image. The server then determines that pixels A and B are both pixels in the pupil image, and pixel C is a pixel in the iris image.

[0106] The technical solution provided by the above embodiments brings at least the following beneficial effects: The server can acquire a first image, which includes an image of the target object. Then, the server processes the first image to generate a first image set, which includes: a first type of image, a second type of image, and a third type of image. The first type of image is the first image after HSV color space conversion, the second type of image is a grayscale image of the first image, and the third type of image is the first image after Gamma transformation. Then, the server generates a target image, where the color value of each pixel is determined based on the first image, the first type of image, the second type of image, and the third type of image. Then, the server processes the target image to determine the target object in the target image. That is, the recognition device can add the color values ​​of the pixels converted from HSV color space, the pixel values ​​of the grayscale image, and the color values ​​of the pixels after Gamma transformation to the color values ​​of each pixel in the first image. This increases the reference amount for the identified object, reduces the influence of reflective points on the recognition process, and improves the accuracy of the image recognition results.

[0107] In some embodiments, such as Figure 4 As shown, after S304, the image recognition method also includes S401.

[0108] S401. The server marks the target object in the first image and generates the fifth image.

[0109] For example, such as Figure 5 As shown, the first image 501 includes object 502, object 503, and object 504. If the server identifies object 503, the server marks object 503 in the first image 501 and generates the fifth image 505.

[0110] Understandably, by tagging the recognition results in the image, the server can present the recognition results in the image, thus improving the practicality and operability of image recognition.

[0111] In some embodiments, such as Figure 6 As shown, after S301, the image recognition method also includes: S601-S603.

[0112] S601. The server segments the first image and generates a second image set.

[0113] The second image set includes at least one second image.

[0114] It should be noted that the process of the server segmenting the first image can be found in the introduction of image segmentation in conventional techniques, and will not be elaborated here.

[0115] It should be noted that the embodiments of this application do not limit the order in which S601 and S302 are executed. For example, the server may execute S601 first and then S302. Or, the server may execute S302 first and then S601. Or, the server may execute S601 and S302 simultaneously.

[0116] For each of at least one second image, the server executes S602.

[0117] S602. The server determines the weight value of the second image based on the color value of the pixels in the second image.

[0118] The weight value of the second image is used to indicate the degree of association between the second image and the image of the target object.

[0119] It should be noted that the larger the weight value of the second image, the greater the correlation between the second image and the target object image.

[0120] In one possible implementation, the weight values ​​of the second image may include the weight value of each pixel in the second image. The pixel weight value indicates the degree of association between the pixel and the image of the target object. The server can determine the weight value of each pixel based on its color value in the second image.

[0121] In one possible design, the weight value of each pixel in the second image can be represented by Formula 5.

[0122]

[0123] Among them, a k The weight value of the k-th pixel in the second image is used to indicate the weight value of the second image, h is used to indicate the height of the second image, w is used to indicate the width of the second image, and F is used to indicate the weight value of the k-th pixel in the second image. k F is used to indicate the color value of the k-th pixel in the second image. j Used to indicate the color value of the j-th pixel in the second image. Where j∈{1,2,...,k,...,h·w}.

[0124] In another possible implementation, the server can use a pre-set neural network model to recognize the second image, generate the tenth image, and determine the weight value of each pixel based on the color value of each pixel in the tenth image. Here, one second image corresponds to one tenth image.

[0125] In one possible design, the weight value of each pixel in the tenth image can be represented by Formula 6.

[0126]

[0127] Among them, a k’ The weight value of the k-th pixel in the tenth image is used to indicate the weight value of the tenth image. h′ indicates the height of the tenth image, w′ indicates the width of the tenth image, and F... k ′ is used to indicate the color value of the k-th pixel in the tenth image, F j ' is used to indicate the color value of the j-th pixel in the tenth image. Where j∈{1,2,...,k,...,h'·w'}.

[0128] In one possible implementation, after the server determines the weight value of each pixel in the second image, the server can arrange the weight values ​​of each pixel according to the position of each pixel's color value in the matrix representation of the second image, thus generating the weight values ​​of the second image. In other words, the weight values ​​of the second image are in matrix representation.

[0129] S603. The server stitches together at least one second image to generate a third image.

[0130] In one possible implementation, the server can stitch together at least one second image according to the position of each second image in the first image to generate a third image.

[0131] In this embodiment of the application, the third image is labeled with the weight value of each second image.

[0132] For example, the first image in matrix representation is At least one second image includes: image A, image B, image C, and image D. Wherein, image A in matrix representation is 0, image B in matrix representation is 5, image C in matrix representation is 13, and image D in matrix representation is 9. If the weight value of image A is 1, the weight value of image B is 0, the weight value of image C is 2, and the weight value of image D is 11, then the third image generated by the server in matrix representation is...

[0133] In some embodiments, the server can generate the weight values ​​of the third image based on the weight values ​​of each of the second images marked in the third image. That is, the weight values ​​of the third image are expressed in matrix form.

[0134] For example, if the third image in matrix form is The server then determines the weight value of the third image.

[0135] In this embodiment of the application, S303 includes: S604.

[0136] S604. The server generates a target image based on the first image, the first type of image, the second type of image, the third type of image, and the third image.

[0137] In one possible implementation, the server stores a second preset function. The server can use the second preset function to process the weight values ​​of the third image and the fifth type of image to generate the seventh type of image.

[0138] It should be noted that the second preset function is not limited in the embodiments of this application. For example, the second preset function can be a linear rectification function (ReLU). Another example is that the second preset function can be a softplus function. Yet another example is that the second preset function can be a Maxout function.

[0139] In one possible design, the seventh type of image can be represented by Formula 7.

[0140] O m =ReLU(b m ⊙R m R m )+R m Formula 7.

[0141] Among them, O m b is used to indicate the color value of the m-th pixel in the seventh type of image. m R is used to indicate the weight value of the m-th pixel in the weight values ​​of the third image. m Indicates the color value of the m-th pixel in the fifth type of image. Where m∈{1, 2, ..., 416} 2}

[0142] In one possible implementation, the server can upsample the seventh type of image to generate an eighth type of image. Then, the server can perform Sum fusion processing on the sixth type of image (in matrix representation), the fourth type of image (in matrix representation), and the eighth type of image (in matrix representation) to generate the target image. The target image is also labeled with the weight values ​​of each second image.

[0143] Understandably, by labeling the weight value of each second image in the target image, the server can indicate the degree of association between each second image in the target image and the target object image, and thus select the second image with a high degree of association for priority recognition, thereby improving the efficiency of image recognition.

[0144] In some embodiments, such as Figure 7 As shown, after S302, the image recognition method also includes S701.

[0145] S701 The server processes the first type of image, the second type of image, and the third type of image to generate the fourth image.

[0146] The fourth image is the edge image of the target object.

[0147] In one possible implementation, the server stores multiple preset thresholds. The server can determine a first threshold from the multiple preset thresholds based on a first type of image, and perform binarization processing on the first type of image based on the first threshold to generate a ninth type of image. Similarly, the server can determine a second threshold from the multiple preset thresholds based on a second type of image, and perform binarization processing on the second type of image based on the second threshold to generate a ninth type of image corresponding to the second type of image. The server can also determine a third threshold from the multiple preset thresholds based on a third type of image, and perform binarization processing on the third type of image based on the third threshold to generate a ninth type of image corresponding to the third type of image.

[0148] In one possible implementation, the server stores a third preset function. The server can process the ninth type of image corresponding to the first type of image according to the third preset function to generate the tenth type of image corresponding to the first type of image.

[0149] It should be noted that the third preset function is not limited in the embodiments of this application. For example, the third preset function can be the Roberts operator. Another example is the Prewitt operator. Yet another example is the Sobel operator.

[0150] In one possible design, the tenth image corresponding to the first image can be represented by Formula 8.

[0151]

[0152] Where p is used to indicate the tenth type image corresponding to the first type image, p x p is used to indicate the image of the ninth category corresponding to the first category image in the x-axis direction. y The image used to indicate the ninth type of image corresponding to the first type of image in the y-axis direction.

[0153] Similarly, in conjunction with the above embodiments, the server can generate the tenth type image corresponding to the second type image and the tenth type image corresponding to the third type image.

[0154] In one possible implementation, the server stores a preset matrix. The server can process the tenth-category image corresponding to the first-category image, the tenth-category image corresponding to the second-category image, and the tenth-category image corresponding to the third-category image based on the preset matrix to generate a fourth image.

[0155] It should be noted that during the process of the server processing the tenth image corresponding to the first image, the tenth image corresponding to the second image, and the tenth image corresponding to the third image, the server can combine the tenth images corresponding to the first image, the tenth images corresponding to the second image, and the tenth images corresponding to the third image into a matrix representation, and multiply them by a preset matrix to generate a fourth image in a matrix representation.

[0156] For example, the preset matrix can be a 416×416 matrix A. Wherein, matrix A is... The values ​​in each row of matrix A from row 1 to row 104 are all 0, the values ​​in each row of matrix A from row 105 to row 312 are all 1, and the values ​​in each row of matrix A from row 313 to row 416 are all 0.

[0157] In this embodiment of the application, S303 includes: S702.

[0158] S702. The server generates a target image based on the first image, the first type of image, the second type of image, the third type of image, the third image, and the fourth image.

[0159] In one possible implementation, the server can perform Sumfusion processing on the sixth type image, the fourth type image, the eighth type image, and the fourth type image in matrix representation to generate a target image. The target image includes the fourth type image.

[0160] Understandably, the server determines the ninth image category for each of the first, second, and third image categories, and then determines the tenth image category for each of the same categories. Next, the server aggregates these tenth image categories in matrix form and multiplies them by a preset matrix to generate a fourth image in matrix form. Finally, the server generates the target image based on the first image, the first image category, the second image category, the third image category, the third image category, and the fourth image. In other words, the server can determine the edge information in the first image and remove edge information that does not belong to the target object, retaining only the edge information belonging to the target image. This increases the edge information of the object in the image, improving the completeness of the image of the identified object.

[0161] In some embodiments, after the server acquires the first image, the server can process the first image according to a preset processing method to generate at least one first image. Then, the server can identify each of the at least one first image to determine the target object in each first image.

[0162] It should be noted that the embodiments of this application do not limit the preset processing method. For example, the preset processing method can be random flipping. Another example is adding Gaussian noise. Yet another example is a combination of random flipping and adding Gaussian noise.

[0163] In one possible implementation, the server can divide at least one first image into a first image set, a second image set, and a third image set according to a preset ratio. Then, the server uses the first image set to train its image recognition capability. Next, the server uses the second image set to test its image recognition capability and uses the third image set to verify the test results.

[0164] It should be noted that the number of images in the second image set is the same as the number of images in the third image set. The preset ratio can be 6:2:2.

[0165] Understandably, the server expands the information content of the first image through a preset processing method and recognizes the expanded first image, thereby enhancing the server's generalization ability.

[0166] In some embodiments, after the terminal sends the first image to the server, the terminal can receive user-inputted operation instructions through a preset annotation tool, annotate the target objects in the first image, generate a mask image, and send the mask image to the server. The server can then receive the mask image and identify the target objects in the first image based on the mask image.

[0167] It should be noted that the default annotation tool can be the deep learning annotation tool in MATLAB, and the Mask image format can be Portable Network Graphics (PNG) format.

[0168] In one possible implementation, the target object may further include a background area. The terminal can receive user input commands via a preset annotation tool to mark the background area in the first image.

[0169] Understandably, the terminal generates a mask image by marking the target objects in the first image and sends it to the server. The server then receives the mask image and identifies the target objects in the first image based on it. This improves the server's efficiency in recognizing objects in images. Furthermore, since objects in some images are in the same position, the server can identify unlabeled images based on labeled images. This enhances the server's generalization ability.

[0170] The following flowchart illustrates the image recognition method provided in this application, using specific examples. Figure 8 As shown, the image recognition method includes S801-S813.

[0171] S801, The server obtains an image of the human eye (i.e., the server executes S301).

[0172] S802, The server generates a tone channel image, a histogram equalized image, and a gamma transform image (i.e., the server executes S302).

[0173] S803, the server performs image binarization, edge detection, and edge merging operations (i.e., the server performs S701).

[0174] S804. The server removes redundant edges (i.e., the server processes the tenth image corresponding to the first image, the tenth image corresponding to the second image, and the tenth image corresponding to the third image according to the preset matrix to generate the fourth image).

[0175] S805. The server performs convolution and multiplication operations to generate a cross-attention matrix (i.e., the server can perform convolution processing on the first type of image to generate a first vector. Furthermore, the server can also perform convolution processing on the second type of image to generate a second vector. Then, the server multiplies the first vector and the second vector to generate the seventh image).

[0176] It should be noted that the embodiments of this application do not limit the order in which S803 and S805 are executed. For example, the server may execute S803 first and then S805. Alternatively, the server may execute S805 first and then S803. Or, the server may execute S803 and S805 simultaneously.

[0177] S806. The server multiplies the cross-attention matrix with the gamma-transformed image (i.e., the server can perform element-wise product of the seventh image in matrix representation and the third image in matrix representation to generate the eighth image), and adds the original features through residual connection (i.e., the server can also perform convolution processing on the first, second, and third images to generate the ninth image. After that, the server can perform residual connection between the eighth image in matrix representation and the ninth image in matrix representation to generate the fourth image).

[0178] S807, The server divides the human eye image into nine slices (i.e., the server executes S601).

[0179] It should be noted that the embodiments of this application do not limit the order in which S802 and S807 are executed. For example, the server may execute S802 first and then S807. Alternatively, the server may execute S807 first and then S802. Yet another example is that the server may execute S802 and S807 simultaneously.

[0180] S808. The server extracts features through the backbone network (i.e., the server can identify the second image through a preset neural network model, generate the tenth image, and determine the weight value of each pixel based on the color value of each pixel in the tenth image).

[0181] S809, The server calculates the attention map based on the slice feature map (i.e., the server executes S603).

[0182] S810, The server multiplies the attention map with the feature map of the human eye image (that is, the server can process the weight value of the third image and the fifth type image through the second preset function to generate the seventh type image).

[0183] S811. The server performs an upsampling operation (i.e., the server can upsample the seventh type of image to generate the eighth type of image, and upsample the fifth type of image to generate the sixth type of image).

[0184] S812, The server performs feature fusion operation (that is, the server can perform Sum fusion processing on the sixth type image, the fourth type image, the eighth type image, and the fourth type image in matrix representation to generate the target image).

[0185] S813, The server generates the segmentation result (i.e., the server executes S401).

[0186] The image recognition method will be introduced below with specific examples.

[0187] For example, such as Figure 9As shown, the server can acquire a human eye image 901. Then, the server processes the human eye image 901 to generate a tone channel image 902, a histogram equalized image 903, a gamma transform image 904, and slices 905, 906, 907, 908, 909, 910, 911, 912, and 913. Next, the server multiplies the tone channel image 902 by a 1×1 convolution to generate a first vector, and multiplies the histogram equalized image 903 by a 1×1 convolution to generate a second vector. Finally, combining this with Formula 2 above, the server multiplies the first vector and the second vector to generate a cross-attention matrix 914. Then, combining with Formula 3 above, the server multiplies the cross-attention matrix 914 with the gamma-transformed image 904, and multiplies the tone channel image 902, histogram equalized image 903, and gamma-transformed image 904 by a 1×1 convolution. Combining with Formula 4 above, the server adds the result of multiplying the tone channel image 902, histogram equalized image 903, and gamma-transformed image 904 by a 1×1 convolution to the result of multiplying the cross-attention matrix 914 with the gamma-transformed image 904, thus obtaining the output result of the illumination cancellation module.

[0188] Furthermore, the server uses the MobileNet model to identify slices 905, 906, 907, 908, 909, 910, 911, 912, 913, and human eye image 901. Combined with Formula Six above, it generates feature maps 915 for slice 905, 916 for slice 906, 917 for slice 907, 918 for slice 908, 919 for slice 909, 920 for slice 910, 921 for slice 911, 922 for slice 912, 923 for slice 913, and 924 for human eye image 901. Next, the server concatenates feature maps 915, 916, 917, 918, 919, 920, 921, 922, and 923 according to the positions of slices 905, 906, 907, 908, 909, 910, 911, 912, and 913 in the human eye image 901, generating attention map 925. Then, combining this with formula seven above, the server multiplies feature map 924 with attention map 925, and upsamples both the result of this multiplication and feature map 924 itself to obtain the output of the positional relationship module.

[0189] Furthermore, the server performs binarization processing on the tone channel image 902, histogram equalization image 903, and gamma transform image 904, along with edge detection, edge merging, and edge culling based on the Prewitt operator, to obtain the output of the edge enhancement module, which is edge map 926. Next, the server performs Sum fusion processing on the output of the illumination cancellation module, the output of the positional relationship module, and edge map 926, and multiplies the Sum fusion result by a 1×1 convolution. Then, the server uses the Softmax function to process the Sum fusion result multiplied by the 1×1 convolution, generating segmentation map 927.

[0190] The foregoing primarily describes the solutions provided in the embodiments of this application from the perspective of computer devices. It is understood that, in order to achieve the aforementioned functions, the computer device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the image recognition method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0191] This application also provides an image recognition device. The image recognition device can be a computer device, a CPU within the aforementioned computer device, a processing module for image recognition within the aforementioned computer device, or a client application for image recognition within the aforementioned computer device.

[0192] This application embodiment can divide the image recognition device into functional modules or functional units according to the above method examples. For example, each function can be divided into a separate functional module or functional unit, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or in software functional modules or functional units. The module or unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0193] like Figure 10 The diagram shown is a structural schematic of an image recognition device provided in an embodiment of this application. The image recognition device is used to perform... Figure 3 , Figure 4 , Figure 6 or Figure 7The image recognition method shown. The image recognition device 1000 may include: an acquisition module 1001 and a processing module 1002.

[0194] The acquisition module 1001 is used to acquire a first image, which includes an image of the target object. The processing module 1002 is used to process the first image to generate a first image set, which includes: a first type of image, a second type of image, and a third type of image. The first type of image is the first image after HSV color space conversion, the second type of image is a grayscale image of the first image, and the third type of image is the first image after gamma transformation. The processing module 1002 is also used to generate a target image, where the color value of each pixel in the target image is determined based on the first image, the first type of image, the second type of image, and the third type of image. The processing module 1002 is also used to process the target image to determine the target object in the target image.

[0195] Optionally, the processing module 1002 is further configured to segment the first image to generate a second image set, the second image set including at least one second image. The processing module 1002 is further configured to, for each of the at least one second image, determine a weight value for the second image based on the color values ​​of pixels in the second image, the weight value of the second image indicating the degree of association between the second image and the image of the target object. The processing module 1002 is further configured to stitch the at least one second image to generate a third image, the third image being labeled with the weight value of each second image. Specifically, the processing module 1002 is configured to generate a target image based on the first image, the first type of image, the second type of image, the third type of image, and the third image, the target image also being labeled with the weight value of each second image.

[0196] Optionally, the processing module 1002 is further configured to process the first type of image, the second type of image, and the third type of image to generate a fourth image, wherein the fourth image is an edge image of the target object image. Specifically, the processing module 1002 is configured to generate a target image based on the first image, the first type of image, the second type of image, the third type of image, the third image, and the fourth image, wherein the target image includes the fourth image.

[0197] Optionally, the processing module 1002 is also used to mark the target object in the first image to generate a fifth image.

[0198] Optionally, the target object may include at least one of the following: eyelid, iris, sclera, pupil, and canthus.

[0199] Figure 11This is a schematic diagram of the hardware structure of an image recognition device according to an exemplary embodiment. The image recognition device may include a processor 1102, which is used to execute application code to implement the image recognition method of this application.

[0200] The processor 1102 may be a central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application.

[0201] like Figure 11 As shown, the image recognition device may further include a memory 1103. The memory 1103 stores the application code that executes the scheme of this application, and its execution is controlled by the processor 1102.

[0202] Memory 1103 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory 1103 may exist independently and be connected to processor 1102 via bus 1104. Memory 1103 may also be integrated with processor 1102.

[0203] like Figure 11 As shown, the image recognition device may further include a communication interface 1101, wherein the communication interface 1101, processor 1102, and memory 1103 may be coupled to each other, for example, through a bus 1104. The communication interface 1101 is used for information exchange with other devices, for example, supporting information exchange between the image recognition device and other devices.

[0204] It should be pointed out that, Figure 11The device structure shown does not constitute a limitation on the image recognition device, except... Figure 11 In addition to the components shown, the image recognition device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0205] In actual implementation, the functions implemented by the processing module 1002 can be derived by... Figure 11 The processor 1102 shown calls the program code in memory 1103 to implement this.

[0206] This application also provides a computer-readable storage medium storing instructions that, when executed by a processor of a computer device, enable the computer to perform the image recognition method provided in the embodiments described above. For example, the computer-readable storage medium may be a memory 1103 including instructions, which may be executed by a processor 1102 of a computer device to complete the method. Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium, such as a ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0207] Figure 12 A conceptual partial view of a computer program product provided in an embodiment of this application is shown as an example. The computer program product includes a computer program for executing computer processes on a computing device.

[0208] In one embodiment, a computer program product is provided using a signal bearer medium 1200. The signal bearer medium 1200 may include one or more program instructions that, when executed by one or more processors, can provide the above-mentioned... Figure 3 , Figure 4 , Figure 6 or Figure 7 The described function or part of the function. Therefore, for example, refer to... Figure 3 In the embodiment shown, one or more features of S301 to S304 can be fulfilled by one or more instructions associated with the signal carrying medium 1200. Furthermore, Figure 12 The program instructions in the document also describe example instructions.

[0209] In some examples, the signal carrying medium 1200 may include a computer-readable medium 1201, such as, but not limited to, a hard disk drive, a compact disc (CD), a digital video disc (DVD), a digital magnetic tape, a memory, a read-only memory (ROM), or a random access memory (RAM), etc.

[0210] In some implementations, the signal carrying medium 1200 may include a computer recordable medium 1202, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, and so on.

[0211] In some implementations, the signal carrying medium 1200 may include a communication medium 1203, such as, but not limited to, digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.).

[0212] The signal-bearing medium 1200 can be transmitted by a wireless communication medium 1203. One or more program instructions may be, for example, computer-executable instructions or logical implementation instructions.

[0213] In some examples, such as targeting Figure 10 The image recognition device described can be configured to provide various operations, functions, or actions in response to one or more program instructions in a computer-readable medium 1201, a computer-recordable medium 1202, and / or a communication medium 1203.

[0214] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0215] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0216] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be in one place or distributed in multiple different locations. Some or all of the constituent units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0217] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0218] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, essentially, or the part that contributes to the prior art, or a complete or partial classification of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0219] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image recognition method, characterized in that, The method includes: Acquire a first image, wherein the first image includes an image of the target object; The first image is processed to generate a first image set, which includes: a first type of image, a second type of image, and a third type of image. The first type of image is the first image after HSV color space conversion, the second type of image is the grayscale image of the first image, and the third type of image is the first image after gamma transformation. The first type of image is convolved to generate a first vector; the second type of image is convolved to generate a second vector; the first vector and the second vector are multiplied to generate a seventh image; the seventh image in matrix form is element-wise multiplied with the third image in matrix form to generate an eighth image; the first, second, and third types of images are convolved to generate a ninth image; the eighth image in matrix form and the ninth image in matrix form are residually concatenated to generate a fourth type of image; the fourth type of image is the image after correcting the reflective points in the first image; A target image is generated, which is obtained by summing and fusing the first image in matrix form and the fourth type of image in matrix form. The target image is processed to determine the target object in the target image.

2. The method according to claim 1, characterized in that, After determining the target object in the target image, the method further includes: The target object is marked in the first image to generate the fifth image.

3. The method according to claim 1, characterized in that, The target object includes at least one of the following: eyelid, iris, sclera, pupil, and canthus.

4. An image recognition device, characterized in that, The device includes: The acquisition module is used to acquire a first image, wherein the first image includes an image of the target object; The processing module is used to process the first image to generate a first image set, which includes: a first type of image, a second type of image, and a third type of image. The first type of image is the first image after HSV color space conversion, the second type of image is the grayscale image of the first image, and the third type of image is the first image after gamma transformation. The processing module is further configured to perform convolution processing on the first type of image to generate a first vector; perform convolution processing on the second type of image to generate a second vector; multiply the first vector and the second vector to generate a seventh image; perform element-wise product of the seventh image in matrix form and the third image in matrix form to generate an eighth image; perform convolution processing on the first type of image, the second type of image, and the third image to generate a ninth image; and perform residual concatenation between the eighth image in matrix form and the ninth image in matrix form to generate a fourth type of image; the fourth type of image is the image after correcting the reflective points in the first image; The processing module is also used to generate a target image, which is obtained by summing and fusing the first image in matrix form and the fourth type of image in matrix form; The processing module is further configured to process the target image to determine the target object in the target image.

5. The apparatus according to claim 4, characterized in that, The processing module is further configured to mark the target object in the first image to generate a fifth image.

6. The apparatus according to claim 4, characterized in that, The target object includes at least one of the following: eyelid, iris, sclera, pupil, and canthus.

7. An image recognition device, characterized in that, include: Processor and memory; The processor and the memory are coupled; The memory is used to store one or more programs, the one or more programs including computer execution instructions. When the image recognition device is running, the processor executes the computer execution instructions stored in the memory to cause the image recognition device to perform the image recognition method as described in any one of claims 1-3.

8. A computer-readable storage medium storing instructions, characterized in that, When the computer executes the instructions, the computer performs the image recognition method as described in any one of claims 1-3.