Image recognition method and device, electronic equipment and storage medium

By generating a mask on the upper layer of the labeled image after image recognition and generating adjustable identification box controls on the mask, the memory overhead problem caused by adjustment of multiple object recognition results during image recognition is solved, and efficient memory management is achieved.

CN120047741APending Publication Date: 2025-05-27BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510127498.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

During the image recognition process, the recognition results of multiple objects in the image need to be manually marked and adjusted, resulting in a large memory overhead, especially when the number of objects is large.

Method used

By generating a mask on the upper layer of the annotation image after image recognition by the AI ​​algorithm, users can select the target identification box on the mask through human-computer interaction events and generate the corresponding identification box adjustment control. Users can adjust these controls independently to reduce unnecessary memory usage.

Benefits of technology

This implements the generation of adjustable controls for only the recognition results of a specific object without the need to generate adjustable controls for all the recognition results of an object, thus greatly reducing memory overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047741A_ABST
    Figure CN120047741A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining a labeled image obtained after image recognition through an AI algorithm, and generating a mask on the upper layer of the labeled image; wherein a plurality of identification frames are marked on the marked image, and each identification frame is used for marking the initial position of an object corresponding to the identification frame in the marked image; a mask is displayed based on the man-machine interaction event, after the man-machine interaction event for a target identification frame is monitored on the mask, an identification frame adjustment control is generated on the mask, and the target identification frame is at least one of a plurality of identification frames; updating coordinate information of the identification box adjustment control based on a man-machine interaction event of a user on the identification box adjustment control; and adjusting the target identification frame to a position matched with the updated identification frame adjustment control to serve as a target position of an object corresponding to the target identification frame in the annotated image. And the memory overhead can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition, and particularly to an image recognition method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] Currently, during the image recognition process, for the recognition results corresponding to several objects in the image, separate annotations (such as marking with a rectangular box) are required, and for inaccurate annotation results, the user needs to manually adjust them. Summary of the Invention

[0003] In view of the above technical problems, the present application provides an image recognition method, apparatus, electronic device, and computer-readable storage medium, and the technical solutions are as follows:

[0004] According to a first aspect of the present application, there is provided an image recognition method; the method includes:

[0005] Obtain an annotated image obtained by performing image recognition through an AI algorithm, and generate a mask layer on the upper layer of the annotated image; wherein, a plurality of identification frames are marked on the annotated image, and each identification frame is used to mark the initial position of the object corresponding to the identification frame in the annotated image;

[0006] Display the mask layer based on a human-computer interaction event, and generate an identification frame adjustment control on the mask layer after a human-computer interaction event for a target identification frame is detected on the mask layer, wherein the target identification frame is at least one of the plurality of identification frames;

[0007] Update the coordinate information of the identification frame adjustment control based on a human-computer interaction event of the user with respect to the identification frame adjustment control;

[0008] Adjust the target identification frame to a position matching the updated identification frame adjustment control, so as to serve as the target position of the object corresponding to the target identification frame in the annotated image.

[0009] According to a second aspect of the present application, there is provided an image recognition apparatus; the apparatus includes:

[0010] A generation unit, configured to obtain an annotated image obtained by performing image recognition through an AI algorithm, and generate a mask layer on the upper layer of the annotated image; wherein, a plurality of identification frames are marked on the annotated image, and each identification frame is used to mark the initial position of the object corresponding to the identification frame in the annotated image;

[0011] The generating unit is further configured to display the mask based on a human-computer interaction event, and generate a control for adjusting the identification box on the mask after a human-computer interaction event for a target identification box is detected on the mask, where the target identification box is at least one of the several identification boxes;

[0012] An updating unit, configured to update the coordinate information of the control for adjusting the identification box based on a human-computer interaction event of a user with the control for adjusting the identification box;

[0013] An adjusting unit, configured to adjust the target identification box to a position matching the updated control for adjusting the identification box, so as to serve as a target position of an object corresponding to the target identification box in the labeled image.

[0014] According to a third aspect of the present application, there is provided an electronic device, including:

[0015] A processor;

[0016] A memory for storing instructions executable by the processor;

[0017] Wherein, the processor is configured to implement the method described in the first aspect.

[0018] According to a fourth aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the method described in the first aspect are implemented.

[0019] The technical solution provided by the present application obtains a labeled image obtained after image recognition by an AI algorithm, generates a mask on the upper layer of the labeled image. A user can select, through a human-computer interaction event, an identification box that needs to be manually adjusted on the labeled image on the mask, and only generate a control for adjusting the identification box corresponding to the target identification box selected by the user through the human-computer interaction event on the mask. Since the labeled image and the mask on its upper layer are in different layers, the user can independently adjust the updated control for adjusting the identification box through a human-computer interaction event on the mask, and adjust the target identification box to a position matching the updated control for adjusting the identification box, so as to realize generating adjustable controls only for the recognition results of specific objects, without generating adjustable controls for the recognition results of all objects on the labeled image, thereby greatly reducing the memory overhead.

[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments described in the present application. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings.

[0022] Figure 1 is a schematic diagram of a specific scenario of image annotation in related technologies;

[0023] Figure 2 is a schematic flowchart of an image recognition method according to an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of an application scenario of a marking frame according to an embodiment of the present application;

[0025] Figure 4 is a schematic diagram of an application scenario of a mask according to an embodiment of the present application;

[0026] Figure 5 is a schematic diagram of an image identification adjustment scenario according to an embodiment of the present application;

[0027] Figure 6 is a schematic structural diagram of an image recognition device according to an embodiment of the present application;

[0028] Figure 7 is a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0029] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will describe the technical solutions in the embodiments of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the present application.

[0030] Image recognition is a subfield of computer vision, whose goal is to enable a computer to automatically recognize objects, scenes, actions, or people in an image. Image recognition can be applied to various scenarios such as face recognition, object classification, license plate recognition, and medical image analysis. Image annotation is the process of adding descriptive information to objects, regions, or points in an image, usually done manually and sometimes with the help of automated tools. Annotated datasets are crucial for training and evaluating computer vision models because they provide the "answers" required for the models to learn and optimize. High-quality annotated data can significantly improve the accuracy and robustness of the models.

[0031] Currently, during the image recognition process, for the recognition results corresponding to several objects in the image, separate annotations are required (for example, annotation with a rectangular box). For inaccurate annotation results, the user needs to manually adjust them. Therefore, the annotation results need to be presented to the user with manually adjustable controls for the user to manually adjust.

[0032] Taking the image recognition process under the Windows Presentation Foundation (WPF) framework as an example, WPF is a user interface framework for creating Windows client applications. As part of the.NET Framework, it is designed to build Graphical User Interfaces (GUIs) and is also supported in.NET Core and.NET 5 and later versions. WPF enables developers to create feature-rich and visually appealing desktop applications.

[0033] For WPF programs, or for the vast majority of client software, functions such as image recognition and annotation are extremely resource-consuming. Since it is not known in advance how many objects are in the image before recognition, nor which object's annotation result may be inaccurate, it is necessary to generate corresponding manually adjustable controls (such as adjustable size, position, and color) for all objects recognized in the image. However, when the number of objects to be recognized is large, the number of manually adjustable controls to be generated is also large, resulting in a large memory overhead. The following is an illustration in a specific scenario. Please refer to Figure 1 , taking this specific scenario as the cell detection scenario as an example, for image recognition and annotation in the cell detection scenario, the objects to be recognized and annotated are cells (such as single cells or cell spheres). Cell detection is often complex and involves a large number of cells. When there are limitations on the memory size, with the incremental accumulation of detection data, it will cause the system to be unable to meet the memory required for program operation, affecting the stability of the program, and at the same time, it will also cause data loss and unforeseen rollbacks.

[0034] To address the above problems, this application provides an image recognition method that can greatly reduce the memory overhead.

[0035] As Figure 2 shown, this method includes the following steps:

[0036] S201. Obtain the annotated image obtained after image recognition by the AI algorithm, and generate a mask layer on the upper layer of the annotated image.

[0037] Among them, a plurality of identification frames are marked on the annotation image, and each of the identification frames is used to mark the initial position of the object corresponding to the identification frame in the annotation image.

[0038] S202. Display the mask based on the human-computer interaction event, and after a human-computer interaction event for the target identification frame is detected on the mask, generate an identification frame adjustment control on the mask.

[0039] Among them, the target identification frame is at least one of the plurality of identification frames.

[0040] S203. Update the coordinate information of the identification frame adjustment control based on the human-computer interaction event of the user on the identification frame adjustment control.

[0041] S204. Adjust the target identification frame to the position matching the updated identification frame adjustment control, so as to serve as the target position of the object corresponding to the target identification frame in the annotation image.

[0042] The technical solution provided by the embodiment of the present application obtains an annotation image obtained after image recognition by an AI algorithm, generates a mask on the upper layer of the annotation image. The user can select, through a human-computer interaction event, the identification frame that needs to be manually adjusted on the annotation image on the mask, and only generate an identification frame adjustment control corresponding to the target identification frame selected by the user through the human-computer interaction event on the mask. Since the annotation image and the mask on its upper layer are in different layers, the user can independently adjust the updated generated identification frame adjustment control through the human-computer interaction event on the mask, and adjust the target identification frame to the position matching the updated identification frame adjustment control, so as to realize generating adjustable controls only for the recognition results of specific objects, without generating adjustable controls for the recognition results of all objects on the annotation image, thereby greatly reducing the memory overhead.

[0043] As an example, a plurality of objects included in the original image can be recognized through an Artificial Intelligence (AI) algorithm to obtain the recognition results corresponding to the plurality of objects respectively. As another example, according to the recognition results corresponding to the plurality of objects respectively, identification frames for marking the initial positions of the objects can be drawn on the original image respectively to obtain an annotation image including a plurality of identification frames. As another example, the above AI algorithm can provide the ability to recognize objects (size, position, and / or type) in the image.

[0044] The above-mentioned identification frame can have multiple specific implementations. As an example, the identification frame can represent the size information and position information of an object. As another example, several identification frames can be displayed in different colors respectively, so as to represent different types of objects through different colors. As another example, the position information, size information, and / or color information of the identification frames corresponding to each object can be recorded in the initial data source, and the initial data source can be maintained in the form of a table.

[0045] As an example, the above-mentioned several identification frames can be pixels drawn on an image based on the recognition results of several objects in the annotated image. As pixels, the memory overhead of the identification frame is the same as that of other parts of the annotated image, and the memory overhead of the identification frame itself is much smaller than the memory occupancy of adjustable controls.

[0046] Considering the issue of accuracy, when performing image recognition through an AI algorithm, generally the original image is used for recognition, which can improve the recognition accuracy. However, the size of the original image is often relatively large (often exceeding the size of the container that holds the image). For example, when using an original image with a size of 6000*4000 for recognition, the size of 6000*4000 often exceeds the size of the container that holds the image, and even exceeds the size of the display of the client device. If the mask is also generated using the size of the original image and then scaled according to the size of the container that holds the mask, it is obvious that resources are wasted.

[0047] To address this problem, there are multiple ways to generate a mask on top of the annotated image. As an example, the size data of the container that holds the image in the client program can be obtained first. After obtaining the recognition result of the original image through an AI algorithm, based on the recognition result, the original image can be scaled according to the size data of the container that holds the image, and corresponding identification frames can be drawn in the scaled image according to the recognition result to obtain the above-mentioned annotated image. At the same time, based on the size data of the container that holds the image, the above-mentioned mask is generated on top of the annotated image, thus avoiding waste of resources caused by the overly large sizes of the annotated image and the mask on top of the annotated image.

[0048] It can be understood that the annotated image and the mask are accommodated in different containers and are not on the same layer. As an example, the container that holds the annotated image can be placed at the bottom layer in the client program, while the container that holds the mask can be placed at the top layer.

[0049] To ensure the smooth implementation of user operations and interactions, as an example, the human-computer interaction event can be bound to the container that holds the mask, that is, the mask can be displayed based on the human-computer interaction event, and the human-computer interaction event can be monitored and corresponding responses can be made on the mask.

[0050] The above-mentioned human-computer interaction events can have various specific implementations. As an example, the above-mentioned mask can be displayed based on a mouse event. The human-computer interaction event for the target identification box can also include a mouse event. The target identification box can be determined based on the position information of the mouse event triggered by the user on the mask and the position information of several identification boxes on the annotated image. As another example, the position information of the mouse event triggered by the user on the mask can be compared with the position information of several identification boxes on the annotated image respectively, and the identification box corresponding to the position information that matches the position information of the mouse event is determined as the above-mentioned target identification box. As another example, the above-mentioned mask can also be displayed based on a keyboard event. The human-computer interaction event for the target identification box can also include a keyboard event. It should be noted that the above introduction to the specific implementation of the human-computer interaction event is only an exemplary display. In actual applications, there may be other specific implementations, which are not limited herein.

[0051] The above-mentioned mouse event can have various specific implementations. As an example, the above-mentioned mouse event can be a click event. As another example, the above-mentioned mouse event can also be a selection event. The specific implementation of the mouse event is not limited.

[0052] The following is an exemplary introduction to the click event:

[0053] As an example, the target identification box selected by the user can be determined among several identification boxes by listening for click events on the mask.

[0054] There are various ways to determine the target identification box based on the click event. As an example, the coordinate points on the annotated image indicated by the click event triggered by the user on the mask can be compared with the coordinate information of several identification boxes on the annotated image respectively, and the identification box corresponding to the coordinate information that matches the coordinate points of the click event is determined as the above-mentioned target identification box.

[0055] Considering that it is difficult to listen for click events, especially in a scenario where the object in the annotated image is a cell, the identification boxes used to identify cells often occupy fewer pixels, and may even be smaller than the pixels occupied by the mouse cursor. Therefore, the difficulty of identifying the identification box targeted by the click event is relatively high. Because there are no fully specified requirements for clicking, one can click inside the visible identification box on the user side, click on the edge of a certain identification box, or click on an area where there is no identification box. Due to the variability and uncertainty of the click event, precise calculations are required to determine which identification box the click event targets.

[0056] Regarding the above problems, as an example, starting from the coordinate point on the labeled image indicated by the click event, rays can be emitted in four directions (up, down, left, and right) from this coordinate point until these rays intersect with the identification boxes on the labeled image. Select the identification box that is closest to the coordinate point of the click event from the identification boxes intersected by the four rays as the target identification box.

[0057] There are various ways to detect whether a ray intersects with an identification box on the labeled image. As an example, for each direction (up, down, left, and right) of the coordinate point of the click event, it can be detected whether any identification box intersects with the ray. Taking the identification box as a rectangular box as an example, the conditions for the ray to intersect with the identification box can include: the ray passes through any side of the rectangular box but does not pass through its two corner points.

[0058] There are various ways to determine the identification box that is closest to the coordinate point of the click event. As an example, for each rectangular box intersected by the above rays, the shortest distance from the coordinate point of the click event to this rectangular box can be calculated. This shortest distance can be the distance from the coordinate point of the click event to a certain point on the boundary of the rectangular box or the distance to the center of the rectangular box. Calculate the shortest distances between all the rectangular boxes intersected by the rays in the above four directions and the coordinate point of the click event, and determine the rectangular box with the minimum shortest distance as the identification box that is closest to the coordinate point of the click event, that is, the target identification box. As another example, the formula for calculating the shortest distance from the coordinate point of the click event to the rectangular box is as follows:

[0059]

[0060] where d is the shortest distance from the coordinate point of the click event to the rectangular box, (x 0 , y 0 ) are the coordinates of the coordinate point of the click event. (x 1 , y 1 ) are the coordinates of the lower left corner point of the rectangular box, (x 2 , y 2 ) are the coordinates of the upper right corner point of the rectangular box, or, (x 1 , y 1 ) can also be the coordinates of the upper left corner point of the rectangular box, (x 2 , y 2 ) can also be the coordinates of the lower right corner point of the rectangular box.

[0061] The following gives an exemplary introduction to a specific application of determining whether a ray intersects with a rectangular box:

[0062] Considering that if a ray intersects with the side of a rectangular box, it may be necessary to additionally determine whether the ray intersects with the corner points of the rectangle or whether the ray completely passes through the rectangle. To address this issue, as an example, assume there is a ray R starting from point P(xp, yp) and extending along the direction vector V(vx, vy). Any rectangular box is defined by its lower-left corner point A(xa, ya) and upper-right corner point B(xb, yb). The four sides of the rectangular box can be represented as line segments. For example, the upper side of the rectangular box can be represented as the line segment from point (xa, yb) to point B, and the left side can be represented as the line segment from point A to point (xa, yb), and so on. For any line segment S defined as starting from point Q1(xq1, yq1) to point Q2(xq2, yq2), the intersection of ray R and line segment S can be represented by parametric equations. Let t be the ratio of the distance from a point on the ray to point P to the length of the direction vector V, and u be the ratio of the position of a point on the line segment to the starting point Q1 to the length of the line segment direction vector Q2 - Q1. The condition for ray R and line segment S to intersect is that there exists a pair of t and u that satisfy the following system of equations:

[0063] x p +tv x =x q1 +u(x q2 -x q1 ) (2)

[0064] y p +tv y =y q1 +u(y q2 -y q1 ) (3)

[0065] If the values of t and u satisfy the above system of equations (2) and (3), it indicates that ray R and line segment S intersect. For each side of the rectangular box, it can be calculated through the above steps whether there is an intersection point with the ray. If there is, the ray intersects with the rectangle. If the ray does not intersect with any of the four sides of the rectangular box, but the starting point of the ray is inside the rectangular box, it means the ray completely passes through the rectangle.

[0066] The following is an exemplary introduction to the lasso event:

[0067] As an example, the lasso event can be listened for on the mask to determine the target identification box selected by the user among several identification boxes.

[0068] There are various ways to determine the target identification box based on the lasso event. As an example, the coordinate range on the labeled image indicated by the lasso event triggered by the user on the mask can be compared with the coordinate information of several identification boxes on the labeled image respectively, and the identification box corresponding to the coordinate information within this coordinate range can be determined as the above-mentioned target identification box.

[0069] It can be understood that the number of target identification boxes that can be selected by a click event is usually one, while the number of target identification boxes that can be selected by a lasso selection event can be one or more.

[0070] There can be various specific implementations for the position of the identification box adjustment control generated on the mask. As an example, the position of the generated identification box adjustment control can be adapted to the position of the above-mentioned target identification box. As another example, the position of the generated identification box adjustment control being adapted to the position of the target identification box can only mean that the center coordinates are the same, or it can mean that the center coordinates, corner coordinates, border coordinates, etc. are all the same (i.e., the identification box adjustment control coincides with the target identification box), and no specific limitation is made in this regard. The position of the identification box adjustment control generated on the mask corresponds to the target identification box, making it more convenient for the user to adjust the identification box adjustment control.

[0071] As an example, another specific implementation of the position of the identification box adjustment control generated on the mask can include: the position of the generated identification box adjustment control can be at a preset position on the mask. As another example, the preset position can be the center position of the mask or other positions of the mask, and no specific limitation is made in this regard. Generating the identification box adjustment control directly at a pre-set fixed position (such as the center of the mask) eliminates the need to compare the position of the target identification box during generation, reducing the computational amount. (The identification box adjustment control is movable, and the user can manually move the identification box adjustment control to the position of the target object corresponding to the target identification box).

[0072] It can be understood that the number of identification box adjustment controls generated on the mask is the same as the number of target identification boxes. For example, if the number of target identification boxes that can be selected by a click event is usually one, then the number of identification box adjustment controls that can be generated on the mask is also one. While the number of target identification boxes that can be selected by a lasso selection event can be one or more, and correspondingly, the number of generated identification box adjustment controls is also one or more. As another example, when the number of generated identification box adjustment controls includes multiple ones, the user can separately adjust each identification box adjustment control through a human-computer interaction event.

[0073] After generating the identification box adjustment control on the mask, the user can adjust the identification box adjustment control through a human-computer interaction event. There can be various specific implementations for this human-computer interaction event. As an example, this human-computer interaction event can be a mouse event or a keyboard event, and no specific limitation is made in this regard.

[0074] As an example, the adjustment of the identification box adjustment control by the user through a human-computer interaction event can be to adjust the size of the identification box adjustment control, or to move the identification box adjustment control, that is, to adjust the center coordinates of the identification box adjustment control, or to adjust the color of the identification box adjustment control. There is no specific limitation on this.

[0075] After the user adjusts the identification box adjustment control through a human-computer interaction event, the coordinate information of the identification box adjustment control can be updated. There are various specific implementations of this coordinate information. As an example, the coordinate information can include the center coordinate information of the identification box adjustment control. As another example, the coordinate information can include the border coordinate information of the identification box adjustment control. As another example, the coordinate information can include the corner coordinate information of the identification box adjustment control. As another example, the coordinate information can include one or more of the center coordinate information, border coordinate information, and corner coordinate information of the identification box adjustment control. There is no specific limitation on this. And updating the coordinate information of the identification box adjustment control can mean updating one or more of the center coordinate information, border coordinate information, and corner coordinate information of the identification box adjustment control.

[0076] If the adjustment of the identification box adjustment control by the user through a human-computer interaction event includes adjusting the color of the identification box adjustment control, then after the color of the identification box adjustment control is adjusted, the color information of the identification box adjustment control can be updated, and the color of the target identification box can be adjusted to the color matching the updated identification box adjustment control. For example, in the cell detection scenario, assume that the above AI algorithm defaults to setting the identification box of the cell sphere to green and the identification box of the single cell to blue during the image recognition process. For the actual recognition result, if the identification box of a certain cell sphere is misrecognized as blue by mistake, the user can adjust the identification box adjustment control corresponding to the identification box of the cell sphere to green through a human-computer interaction event, and then adjust the identification box of the cell sphere to green based on the identification box adjustment control with the updated color.

[0077] The target identification box can be adjusted to the position matching the updated identification box adjustment control in various ways. As an example, the pixels at the corresponding position of the target identification box on the annotation image can be restored to the original pixels of the annotation image, and based on the position matching the updated identification box adjustment control on the annotation image, drawing can be performed, and the drawn pixels can be used as the adjusted target identification box, and the adjusted target identification box can be used as the target position of the object corresponding to the target identification box.

[0078] There are various ways to draw by adjusting the position where the control matches on the labeled image based on the updated identification box. As an example, the position information, size information, and / or color information corresponding to the updated identification box adjustment control can be used to replace the position information, size information, and / or color information recorded in the initial data source for the target identification box corresponding to the identification box adjustment control. Then, based on the replaced position information, size information, and / or color information, redrawing is performed on the labeled image. The pixels drawn are used as the adjusted target identification box, and the adjusted target identification box is used as the target position of the object corresponding to the target identification box.

[0079] As an example, an interaction ability attribute can also be added to the identification box corresponding to each object. This interaction ability attribute can be recorded in the above-mentioned initial data source together with the position information, size information, and / or color information of the identification box. The value of the interaction ability attribute can be used to indicate whether the position information, size information, and / or color information recorded in the initial data source for the identification box can be adjusted or replaced.

[0080] There can be various specific implementations of the above interaction ability attribute. As an example, this interaction ability attribute can be named Visable (i.e., a Bool type that controls whether the information of the identification box can be adjusted), and its value can be "True" (adjustable) or "False" (non-adjustable). As another example, the value of this interaction ability attribute can also be "1" (adjustable) or "0" (non-adjustable). For example, when a human-computer interaction event for the target identification box is detected on the mask, the interaction ability attribute of the target identification box can be set to "True", and after adjusting the target identification box to the position where it matches the updated identification box adjustment control, the interaction ability attribute of the adjusted target identification box can be set to "False".

[0081] As an example, when a close instruction for the identification box adjustment control is detected, the identification box adjustment control can be deleted. It can be understood that deleting the identification box adjustment control means deleting the identification box adjustment control generated on the mask from the memory it occupies. At the same time, the deleted identification box adjustment control is no longer rendered and displayed on the mask (it can be understood as becoming invalid). Deleting the unnecessary identification box adjustment controls can reduce memory occupancy.

[0082] The close instruction for the identification box adjustment control can be triggered in various ways. As an example, the close instruction for the identification box adjustment control can be triggered based on a human-computer interaction event before the identification box adjustment control is updated. For example, after the identification box adjustment control is generated, before or during the adjustment of the identification box adjustment control, for some reasons (such as no longer needing to adjust, or the user selects the target identification box incorrectly), the user can trigger the close instruction through a mouse event to delete the identification box adjustment control. At this time, since the update of the identification box adjustment control is not completed, the target identification box corresponding to the identification box adjustment control will not be adjusted.

[0083] As an example, another triggering method for the close instruction for the identification box adjustment control can include: the close instruction for the identification box adjustment control is automatically triggered after the target identification box is adjusted to a position matching the updated identification box adjustment control. For example, after the user updates the identification box adjustment control and the target identification box corresponding to the identification box adjustment control is also adjusted to a position matching the updated identification box adjustment control, if the user no longer needs the identification box adjustment control at this time, the above close instruction can be automatically triggered to delete the updated identification box adjustment control. Of course, in addition to the automatic triggering method, the user can also manually trigger the above close instruction through a human-computer interaction event, and specific details are not limited here.

[0084] As an example, the mask described in any of the above embodiments can be a transparent mask, and the transparent mask can increase the visibility of the underlying annotation image, facilitating the user to observe the objects and identification boxes in the annotation image.

[0085] Next, in combination with Figure 3 、 Figure 4 and Figure 5 an exemplary introduction is given to a specific application scenario of the embodiments of the present application:

[0086] As Figure 3 shown, taking the cell detection scenario as an example, the annotation image described in any of the above embodiments can be a cell annotation image, and the identification box in the cell annotation image can be used to identify the initial position of an object (such as a cell). The identification box can be a rectangular box. As Figure 4 shown, the generated mask is on the upper layer of the annotation image, that is, the annotation image can be at the bottom layer, the mask can be at the top layer, and the two are not in the same layer. As Figure 5As shown, if the user observes a bounding box that needs to be adjusted in the annotated image (for example, there is a deviation between the size and position of the bounding box and the actual object position), the user can, through a human-computer interaction event such as a click event on the mask, select the bounding box, and then only the bounding box adjustment control corresponding to the selected bounding box can be generated on the mask, without generating the controls corresponding to other bounding boxes. The user can adjust and update the bounding box adjustment control on the mask, and the adjusted and updated bounding box adjustment control can be used as the adjustment basis for its corresponding bounding box.

[0087] Corresponding to the above method embodiments, an embodiment of the present application also provides an image recognition device. Refer to Figure 6 As shown, the device may include:

[0088] A generating unit 601, configured to obtain an annotated image obtained by performing image recognition through an AI algorithm, and generate a mask on the upper layer of the annotated image; wherein, a plurality of bounding boxes are marked on the annotated image, and each bounding box is used to mark the initial position of the object corresponding to the bounding box in the annotated image;

[0089] The generating unit 601 is further configured to display the mask based on a human-computer interaction event, and generate a bounding box adjustment control on the mask after a human-computer interaction event for a target bounding box is monitored on the mask, where the target bounding box is at least one of the plurality of bounding boxes;

[0090] An updating unit 602, configured to update the coordinate information of the bounding box adjustment control based on a human-computer interaction event of the user on the bounding box adjustment control;

[0091] An adjusting unit 603, configured to adjust the target bounding box to a position matching the updated bounding box adjustment control, so as to serve as the target position of the object corresponding to the target bounding box in the annotated image.

[0092] As an example, the human-computer interaction event for the target bounding box includes a mouse event; the mouse event includes a click event or a selection event to select the target bounding box through the mouse event.

[0093] As an example, the human-computer interaction event of the user on the bounding box adjustment control includes a mouse event or a keyboard event.

[0094] As an example, the position of the generated bounding box adjustment control is adapted to the position of the target bounding box.

[0095] As an example, the position of the generated bounding box adjustment control is at a preset position on the mask.

[0096] As an example, the several identification frames are pixels drawn on the labeled image based on the recognition results of several objects in the labeled image.

[0097] As an example, the updating unit 602 is specifically configured to update the center coordinate information of the identification frame adjustment control, and / or update the border coordinate information of the identification frame adjustment control.

[0098] As an example, the updating unit 602 is further configured to update the color information of the identification frame adjustment control based on a human-computer interaction event of the user with respect to the identification frame adjustment control; the adjustment unit 603 is further configured to adjust the color of the target identification frame to the color matching the updated identification frame adjustment control.

[0099] As an example, the apparatus further includes a deletion unit, and the deletion unit is configured to delete the identification frame adjustment control when a closing instruction for the identification frame adjustment control is detected.

[0100] As an example, the closing instruction is triggered based on a human-computer interaction event before the identification frame adjustment control is updated, or the closing instruction is automatically triggered after the target identification frame is adjusted to a position matching the updated identification frame adjustment control.

[0101] As an example, the image is a cell labeled image, and the identification frame is used to identify the initial position of cells in the cell labeled image.

[0102] As an example, the adjustment unit 603 is specifically configured to restore the pixels at the corresponding position of the target identification frame on the labeled image to the original pixels of the labeled image, and perform drawing based on the position on the labeled image that matches the updated identification frame adjustment control, use the drawn pixels as the adjusted target identification frame, and use the adjusted target identification frame as the target position.

[0103] The present application further provides an electronic device, as Figure 7 shown, the electronic device includes:

[0104] A processor 701;

[0105] A memory 702 for storing instructions executable by the processor;

[0106] Wherein, the processor 701 is configured to implement the image recognition method described in any one of the above embodiments.

[0107] The present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the image recognition method described in any one of the above embodiments is implemented.

[0108] The above are only specific embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. An image recognition method, characterized in that: The method comprises: Acquire an annotated image obtained after image recognition by an AI algorithm, and generate a mask on an upper layer of the annotated image; wherein the annotated image is marked with a plurality of identification frames, each of which is used to mark an initial position of an object corresponding to the identification frame in the annotated image; The mask is displayed based on a human-computer interaction event, and after a human-computer interaction event for a target identification frame is monitored on the mask, an identification frame adjustment control is generated on the mask, wherein the target identification frame is at least one of the plurality of identification frames; Based on a human-computer interaction event of the user on the identification frame adjustment control, updating coordinate information of the identification frame adjustment control; The target identification frame is adjusted to a position matched with the updated identification frame adjustment control to serve as a destination position of an object corresponding to the target identification frame in the annotated image.

2. The method according to claim 1, characterized in that The human-computer interaction event for the target identification frame includes a mouse event; the mouse event includes a click event or a circle selection event, so that the target identification frame is selected through the mouse event.

3. The method according to claim 1, characterized in that The human-computer interaction event of the user adjusting the control of the identification frame includes a mouse event or a keyboard event.

4. The method according to claim 1, characterized in that: The position of the generated identification frame adjustment control is adapted to the position of the target identification frame.

5. The method according to claim 1, characterized in that The position of the generated identification frame adjustment control is at a preset position on the mask.

6. The method according to claim 1, characterized in that The plurality of identification boxes are pixels drawn on the annotated image based on recognition results of a plurality of objects in the annotated image.

7. The method according to claim 1, characterized in that The updating of the coordinate information of the identification frame adjustment control includes: Update the center coordinate information of the identification frame adjustment control, and / or update the frame coordinate information of the identification frame adjustment control.

8. The method according to claim 1, characterized in that The method further comprises: Based on a human-computer interaction event of a user on the identification frame adjustment control, updating color information of the identification frame adjustment control; The color of the target identification frame is adjusted to the color matched by the updated identification frame adjustment control.

9. The method according to claim 1, characterized in that: The method further comprises: When a closing instruction for the identification frame adjustment control is detected, the identification frame adjustment control is deleted.

10. The method according to claim 9, characterized in that The closing instruction is triggered based on a human-computer interaction event before the identification frame adjustment control is updated, or The closing instruction is automatically triggered after the target identification frame is adjusted to a position matching the updated identification frame adjustment control.

11. The method according to claim 1, characterized in that: The annotated image is a cell annotated image, and the identification frame is used to identify the initial position of the cell in the cell annotated image.

12. The method according to claim 1, characterized in that The step of adjusting the target identification frame to a position matching the updated identification frame adjustment control includes: The pixels at the corresponding position of the target identification frame on the annotated image are restored to the original pixels of the annotated image, and the matching position on the annotated image is drawn based on the updated identification frame adjustment control, and the drawn pixels are used as the adjusted target identification frame, and the adjusted target identification frame is used as the destination position.

13. An image recognition device, characterized in that: The device comprises: A generating unit, used to obtain an annotated image obtained after image recognition by an AI algorithm, and generate a mask on an upper layer of the annotated image; wherein the annotated image is marked with a plurality of identification frames, each of which is used to mark an initial position of an object corresponding to the identification frame in the annotated image; The generating unit is further configured to display the mask based on a human-computer interaction event, and generate an identification frame adjustment control on the mask after a human-computer interaction event for a target identification frame is monitored on the mask, wherein the target identification frame is at least one of the plurality of identification frames; An updating unit, configured to update coordinate information of the identification frame adjustment control based on a human-computer interaction event of the user on the identification frame adjustment control; The adjusting unit is used to adjust the target identification frame to a position matched with the updated identification frame adjustment control, so as to serve as a destination position of the object corresponding to the target identification frame in the annotated image.

14. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 12.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.