Training method, image processing method and related equipment
By using the elliptical intersection-union ratio in object recognition model training, the problem of noise introduced by background information in rectangular areas is solved, the recognition accuracy of the model is improved, and more accurate feature learning is achieved.
Patent Information
- Application Number
- CN202511259532.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-04
AI Technical Summary
During the training process of the neural network-based object recognition model, the four corners of the rectangular area may contain background information, resulting in additional noise and reducing the recognition accuracy.
The parameters of the initial object recognition model are adjusted using the ellipse intersection-union ratio (first intersection-union ratio). By obtaining the recognition boundary and annotation boundary information of the sample object, the boundaries of the first elliptical area and the second elliptical area are determined, their intersection-union ratio is calculated, and the model parameters are adjusted.
This reduces background information interference, improves the recognition accuracy of the trained object recognition model, and ensures that the model can learn the true characteristics of the sample objects.
Smart Images

Figure CN120807961A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to a training method, an image processing method, and related equipment. BACKGROUND
[0002] Currently, in the training process of an object recognition model based on a neural network, an intersection over union (IoU) between a labeled rectangular region of a sample object and a recognized rectangular region of the sample object recognized by the model is calculated, and the model parameters are adjusted based on the IoU to complete the training.
[0003] However, the four corners of the recognized rectangular region may contain background information. These background information introduces additional noise in the model training process, reducing the recognition accuracy of the trained object recognition model. SUMMARY
[0004] The embodiments of the present application provide a training method, an image processing method, and related equipment, which can reduce the additional noise introduced in the model training process and improve the recognition accuracy of the trained object recognition model.
[0005] In a first aspect, a training method is provided. The method includes: obtaining labeled boundary information of a sample object of a sample image and recognized boundary information of the sample object, the labeled boundary information representing a boundary of a labeled rectangular region in which the sample object is located in the sample image, and the recognized boundary information representing a boundary of a recognized rectangular region of the sample object in the sample image, the shape of the sample object being related to an ellipse; determining first boundary information representing a boundary of a first elliptical region based on the recognized boundary information, the first elliptical region being a region with the same center and the same size as a maximum inscribed ellipse of the recognized rectangular region, and an included angle between a major axis of the first elliptical region and a major axis of the maximum inscribed ellipse of the recognized rectangular region being a target angle; determining second boundary information representing a boundary of a second elliptical region based on the labeled boundary information, the second elliptical region being a region with the same center and the same shape as a maximum inscribed ellipse of the labeled rectangular region, and an included angle between a major axis of the second elliptical region and a major axis of the maximum inscribed ellipse of the labeled rectangular region being the target angle; determining a first intersection over union of the first elliptical region and the second elliptical region based on the first boundary information and the second boundary information; and adjusting parameters of an initial object recognition model based on the first intersection over union to obtain an object recognition model.
[0006] In one possible implementation, the target angle is obtained by: acquiring an image of the identified rectangular area; performing pixel distribution analysis on the image of the identified rectangular area to obtain a main direction of the pixel distribution; and calculating the angle between the main direction and a reference direction to obtain the target angle.
[0007] In a possible implementation, the target angle is 0 degrees.
[0008] In one possible implementation, determining a first intersection-and-union ratio (IOR) of the first elliptical area and the second elliptical area based on the first boundary information and the second boundary information includes: determining a minimum circumscribed rectangular area containing the marked rectangular area and the identified rectangular area based on the annotated boundary information and the identified boundary information; obtaining a set of sampling points containing a target number of sampling points in the minimum circumscribed rectangular area; determining a first number of sampling points in the set that belong to the first elliptical area and the second elliptical area based on the first boundary information and the second boundary information, and determining a second number of sampling points in the set that belong to the union area of the first elliptical area and the second elliptical area based on the first boundary information and the second boundary information; and calculating a ratio of the first number to the second number to obtain the first IOR.
[0009] In one possible implementation, obtaining the annotated boundary information of a sample object of a sample image and the identification boundary information of the sample object includes: obtaining candidate annotated boundary information of multiple candidate objects of the sample image, the candidate annotated boundary information representing the boundary of a candidate annotated rectangular area where the candidate object is located in the sample image; using an initial object recognition model to process the sample image to obtain candidate identification boundary information of the multiple candidate objects, the candidate identification boundary information representing the boundary of the candidate identification rectangular area where the candidate object is located in the sample image; determining the annotated boundary information of the sample object from the multiple candidate annotated boundary information, and determining the identification boundary information of the sample object from the multiple candidate identification boundary information.
[0010] In a possible implementation, determining the labeled boundary information of the sample object from the plurality of candidate labeled boundary information, and determining the identification boundary information of the sample object from the plurality of candidate identification boundary information, includes: determining, for each candidate identification boundary information and candidate labeled boundary information of the candidate object, a second intersection-over-union ratio between the candidate labeled rectangular area and the candidate identification rectangular area based on the candidate identification boundary information and the candidate labeled boundary information; The second intersection-over-union ratios corresponding to the plurality of candidate objects are filtered to obtain a target intersection-over-union ratio; candidate recognition boundary information corresponding to the target intersection-over-union ratio in candidate recognition boundary information of the plurality of candidate objects is taken as the recognition boundary information, and candidate annotation boundary information corresponding to the target intersection-over-union ratio in candidate annotation boundary information of the candidate object is taken as the annotation boundary information.
[0011] In a possible implementation, the filtering of the second intersection-over-union ratios corresponding to the plurality of candidate objects to obtain a target intersection-over-union ratio comprises: sorting the plurality of second intersection-over-union ratios in descending order to obtain sorted second intersection-over-union ratios, and taking the first K second intersection-over-union ratios in the sorted second intersection-over-union ratios as the target intersection-over-union ratio, where K is a positive integer greater than 1; or determining a second intersection-over-union ratio greater than or equal to an intersection-over-union ratio threshold from the plurality of second intersection-over-union ratios to obtain the target intersection-over-union ratio.
[0012] In a second aspect, an image processing method is provided. The method comprises: obtaining a to-be-processed image; and identifying the to-be-processed image by using an object recognition model to determine a target object in the to-be-processed image, the shape of the target object being related to an ellipse, the object recognition model being trained by the method of any one of the first aspect.
[0013] In a third aspect, a training apparatus is provided. The apparatus comprises: an obtaining unit configured to obtain annotation boundary information of a sample object of a sample image and recognition boundary information of the sample object, the annotation boundary information representing a boundary of an annotation rectangular region in which the sample object is located in the sample image, and the recognition boundary information representing a boundary of a recognition rectangular region of the sample object in the sample image, the shape of the sample object being related to an ellipse; a processing unit configured to determine, based on the recognition boundary information, first boundary information representing a boundary of a first elliptical region, the first elliptical region being a region with the same center and the same size as a maximum inscribed ellipse of the recognition rectangular region, and an included angle between a major axis of the first elliptical region and a major axis of the maximum inscribed ellipse of the recognition rectangular region being a target angle; determine, based on the annotation boundary information, second boundary information representing a boundary of a second elliptical region, the second elliptical region being a region with the same center and the same shape as a maximum inscribed ellipse of the annotation rectangular region, and an included angle between a major axis of the second elliptical region and a major axis of the maximum inscribed ellipse of the annotation rectangular region being the target angle; and determine a first intersection-over-union ratio of the first elliptical region and the second elliptical region based on the first boundary information and the second boundary information; and an adjusting unit configured to adjust parameters of the initial object recognition model based on the first intersection-over-union ratio to obtain an object recognition model.
[0014] In a fourth aspect, an electronic device is provided, comprising a memory for storing a computer program and a processor for invoking and running the computer program from the memory, so that the electronic device executes the method of any one of the first aspect or the second aspect.
[0015] In a fifth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed to execute the method of any one of the first aspect or the second aspect.
[0016] In a sixth aspect, a computer program product is provided, which comprises computer program instructions, and the computer program instructions are executed to execute the method of any one of the first aspect or the second aspect.
[0017] In a seventh aspect, a chip is provided, comprising a processor and a data interface, and the processor reads instructions stored on a memory through the data interface to implement the method of any one of the first aspect or the second aspect.
[0018] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: the first boundary information representing the boundary of the first elliptical region is determined by using the obtained recognition boundary information, and the second boundary information representing the boundary of the second elliptical region is determined by using the obtained labeling boundary information, the first intersection-over-union ratio of the first elliptical region and the second elliptical region is determined based on the first boundary information and the second boundary information, and the parameters of the initial object recognition model are adjusted based on the first intersection-over-union ratio to obtain the object recognition model. In this way, since the contour of the sample object is more fitted to the ellipse, the model training is performed by using the ellipse intersection-over-union ratio (i.e., the first intersection-over-union ratio) in the model training process, the background information interference is reduced, the model can learn the real features of the sample object in the training process, and the recognition accuracy of the object recognition model obtained by training is improved. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is a flowchart of a training method provided by the embodiments of the present application; Figure 2 is a schematic diagram of a minimum circumscribed rectangular region in a training method provided by the embodiments of the present application; Figure 3 is a schematic diagram of a scene for calculating an ellipse intersection-over-union ratio in a training method provided by the embodiments of the present application; Figure 4 is a schematic diagram of a rectangular intersection ratio and an ellipse intersection-over-union ratio in a training method provided by the embodiments of the present application; Figure 5is a flowchart of an image processing method provided by an embodiment of the present application; Figure 6 is a structural diagram of an image processing device provided by an embodiment of the present application; Figure 7 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present application will be clearly and completely described in combination with the drawings in the embodiments of the present application.
[0021] In the following description, specific details are set forth in connection with the particular system structures, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, persons skilled in the art understand that the embodiments of the present application can be achieved in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary details.
[0022] It should be understood that, when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0023] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0024] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for differentiation of description, and cannot be understood as indicating or implying relative importance.
[0025] In the specification of the present application, the reference "one embodiment" or "some embodiments" means that the specific features, structures, or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in further some embodiments", etc. appearing in different places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized.
[0026] Currently, in a training process of a neural network-based object recognition model, an intersection over union between a labeled rectangular region of a sample object and a recognized rectangular region of the sample object is calculated, and based on the intersection over union, model parameters of the object recognition model are adjusted to complete the training. However, when the shape of the sample object is approximately elliptical, the four corner points of the rectangular region can contain background information. The background information introduces additional noise in the model training process, and reduces the recognition accuracy of the trained object recognition model.
[0027] In an example, taking a screen defect detection scenario as an example, the screen defects are, for example, bright and dark spots, bright and dark groups, etc. A deep learning-based natural image object detection algorithm (for example, you only look once (Yolo), mask region-based convolutional neural network (MaskR-CNN), etc.) is applied to the screen defect detection scenario, an object recognition model for recognizing screen defects can be trained, and the object recognition model is used to detect defects in the screen. In the training process, a rectangular intersection over union is usually used to measure the fitting degree between the recognized rectangular region and the labeled rectangular region of the sample object, and based on this, the model parameters of the initial object recognition model are adjusted to obtain the object recognition model. However, when the rectangular region is used to frame the sample object, the background information is often contained in the four corners of the rectangular region, which introduces additional noise in the training, is not conducive to the convergence of the model, and reduces the recognition accuracy of the trained object recognition model.
[0028] Based on this, the embodiments of the present application provide an image processing method. Since the shapes of the screen defects are more consistent with ellipses, in the model training process, the model parameters of the initial object recognition model are adjusted based on an elliptical intersection over union to complete the training to obtain the object recognition model. In this way, the interference of the background information is reduced, the real features of the sample object can be learned in the training process, and the recognition accuracy of the trained object recognition model is improved.
[0029] The image processing method provided by the present application will be explained and described in detail below.
[0030] Figure 1 is a flowchart of a training method provided by the embodiments of the present application. The training method 100 can be applied to an electronic device, such as a terminal device or a server. The training method 100 is not limited to the specific order of Figure 1 It should be understood that in other embodiments, the order of some steps of the training method 100 can be exchanged with each other according to actual needs, or some steps can be omitted or deleted. For example, Figure 1As shown, the training method 100 includes S101-S105, and each step is explained in detail as follows.
[0031] S101, obtaining labeled boundary information of a sample object of a sample image and recognized boundary information of the sample object, the labeled boundary information representing a boundary of a labeled rectangular region where the sample object is located in the sample image, the recognized boundary information representing a boundary of a recognized rectangular region of the sample object in the sample image, and the shape of the sample object being related to an ellipse.
[0032] The number of the sample images is multiple. One sample image includes at least one sample object.
[0033] In the embodiments of the present application, the labeled boundary information of the sample object of the sample image can be obtained by the electronic device in response to an input operation of the terminal device; or the labeled boundary information of the sample object of the sample image can be obtained by other devices in response to an input operation of a user and then sent to the electronic device. The recognized boundary information of the sample object can be obtained by the terminal device by processing the sample image using an initial object recognition model; or the recognized boundary information of the sample object can be obtained by other devices by processing the sample image using the initial object recognition model and then sent to the electronic device.
[0034] It should be understood that the shape of the sample object being related to an ellipse can be understood as the sample object being an ellipse or an approximately ellipse. The similarity between the shape of the sample object and the shape of the ellipse is greater than or equal to a similarity threshold. The sample object can be an object in the sample image that needs to be focused on in the training process.
[0035] In a possible implementation, the labeled boundary information includes position information of the labeled rectangular region and size information of the labeled rectangular region. The position information of the labeled rectangular region includes coordinates of a center point of the labeled rectangular region; wherein the coordinates of the center point of the labeled rectangular region can be determined by establishing a coordinate system with the top-left vertex of the sample image as the origin (0, 0), the horizontal direction to the right as the x-axis, and the vertical direction upward as the y-axis. The size information of the labeled rectangular region includes the width and the height of the labeled rectangular region.
[0036] Exemplarily, the labeled boundary information can be represented as , wherein represents the coordinates of the center point of the labeled rectangular region; 1 represents the width of the labeled rectangular region, 1 represents the height of the labeled rectangular region.
[0037] In a possible implementation, the identifying boundary information includes identifying position information of the rectangular region and identifying size information of the rectangular region. The identifying position information of the rectangular region includes coordinates of a center point of the rectangular region; wherein the coordinates of the center point of the rectangular region can be determined by establishing a coordinate system with a top-left vertex of the sample image as an origin (0, 0), a horizontal direction to the right as an x-axis, and a vertical direction upward as a y-axis. The identifying size information of the rectangular region includes a width and a height of the rectangular region. The rectangular region is also referred to as a predicted rectangular region.
[0038] Exemplarily, the rectangular region can be represented as . , wherein (x0, y0) represents the coordinates of the center point of the rectangular region; 2 represents the width of the rectangular region, 2 represents the height of the rectangular region.
[0039] Exemplarily, taking a screen defect detection scenario as an example, the sample image can be a sample screen defect image, and the sample object can be a bright-dark point, a bright-dark spot or the like defect. The labeled rectangular region can be, for example, a rectangular region in which a bright-dark point in the sample screen defect image is located, which is obtained by a human labeling manner. The identified rectangular region can be a rectangular region in which the bright-dark point is located, which is obtained by processing the sample screen defect image by using an initial detection model.
[0040] In a possible implementation, obtaining the labeled boundary information of the sample object of the sample image and the identified boundary information of the sample object includes: obtaining candidate labeled boundary information of a plurality of candidate objects of the sample image, the candidate labeled boundary information representing boundaries of candidate labeled rectangular regions in which the candidate objects are located in the sample image; processing the sample image by using an initial object identification model to obtain candidate identified boundary information of the plurality of candidate objects, the candidate identified boundary information representing boundaries of candidate identified rectangular regions in which the candidate objects are located in the sample image; determining the labeled boundary information of the sample object from the plurality of candidate labeled boundary information, and determining the identified boundary information of the sample object from the plurality of candidate identified boundary information.
[0041] It should be understood that a sample image can include a plurality of candidate objects, and each candidate object corresponds to a candidate labeled boundary information and a candidate identified boundary information. The candidate object can be an object that needs to be focused on and is selected from the sample image by a human.
[0042] The candidate annotation boundary information of the multiple candidate objects in the sample image is obtained by marking the rectangular region where each candidate object is located in the sample image, and the candidate annotation boundary information of each candidate object is obtained. In this marking process, the electronic device can record the coordinates of the four vertices and the center point of the candidate annotation rectangular region to obtain the candidate annotation boundary information. The candidate annotation boundary information includes the position information and size information of the candidate annotation rectangular region. The electronic device can use the initial object recognition model to process the sample image to obtain the candidate recognition boundary information of the candidate object. In this way, the candidate annotation rectangular region and the candidate recognition boundary information of each candidate object can be obtained for the same sample image.
[0043] The recognition boundary information of the sample object can be part or all of the multiple candidate recognition boundary information. The annotation boundary information of the sample object can be the candidate annotation boundary information corresponding to the recognition boundary information of the sample object in the multiple candidate annotation boundary information.
[0044] In a possible implementation, when the number of the multiple candidate recognition boundary information is less than or equal to the first preset number, the multiple candidate recognition boundary information are all used as the recognition boundary information. When the number of the multiple candidate recognition boundary information is greater than the first preset number, a second preset number of candidate recognition boundary information are randomly selected from the multiple candidate recognition boundary information as the recognition boundary information, and the second preset number is less than or equal to the first preset number.
[0045] Based on the above scheme, the problem that the model training process is complex and time-consuming due to too many recognition boundary information and annotation boundary information is avoided, the model training efficiency is improved, and the complexity of model training is reduced.
[0046] In a possible implementation, the multiple candidate recognition boundary information can be screened based on the second intersection-over-union ratio between the candidate annotation rectangular region and the candidate recognition rectangular region of the multiple candidate objects to obtain the recognition boundary information.
[0047] For example, when the number of the multiple candidate recognition boundary information is greater than or equal to the first preset number, the multiple candidate recognition boundary information are screened based on the multiple second intersection-over-union ratios to obtain a second preset number of recognition boundary information. The second preset number is less than the first preset number. The second preset number can be represented by K, and K is a positive integer greater than 1.
[0048] Based on the above scheme, the candidate recognition boundary information is filtered to obtain the recognition boundary information, so as to obtain the labeled boundary information corresponding to the recognition boundary information. In this way, the interference information existing in the multiple candidate recognition boundary information can be eliminated, the reliability and accuracy of the obtained recognition boundary information and labeled boundary information are improved, and then the model can be provided with an accurate learning direction, and the recognition accuracy of the trained object recognition model is improved.
[0049] In a possible implementation, for the candidate recognition boundary information and the candidate labeled boundary information of each candidate object, a second intersection-over-union ratio between the candidate labeled rectangular region and the candidate recognition rectangular region can be determined based on the candidate recognition boundary information and the candidate labeled boundary information.
[0050] It should be understood that the second intersection-over-union ratio is a ratio between an intersection area and a union area of the candidate labeled rectangular region and the candidate recognition rectangular region. The second intersection-over-union ratio represents the degree of overlap between the candidate labeled rectangular region and the candidate recognition rectangular region, and the greater the second intersection-over-union ratio, the more accurate the candidate recognition rectangular region detected by the initial object recognition model.
[0051] Exemplarily, for the candidate recognition boundary information and the candidate labeled boundary information of each candidate object, an intersection area between the candidate labeled rectangular region and the candidate recognition rectangular region can be calculated based on the candidate recognition boundary information and the candidate labeled boundary information, and a sum of an area of the candidate labeled rectangular region and an area of the candidate recognition rectangular region can be calculated based on the candidate recognition boundary information and the candidate labeled boundary information to obtain an area sum value. The area sum value is subtracted by the intersection area to obtain a union area between the candidate labeled rectangular region and the candidate recognition rectangular region. Subsequently, a ratio between the intersection area and the union area is calculated to obtain the second intersection-over-union ratio.
[0052] In a possible implementation, the multiple candidate recognition boundary information is filtered based on the second intersection-over-union ratios between the candidate labeled rectangular regions and the candidate recognition rectangular regions of the multiple candidate objects to obtain the recognition boundary information, including: the second intersection-over-union ratios corresponding to the multiple candidate objects are filtered to obtain a target intersection-over-union ratio; candidate recognition boundary information corresponding to the target intersection-over-union ratio in the candidate recognition boundary information of the multiple candidate objects is taken as the recognition boundary information, and candidate labeled boundary information corresponding to the target intersection-over-union ratio in the candidate labeled boundary information of the candidate object is taken as the labeled boundary information.
[0053] The screening of the second intersection-over-union ratios corresponding to the plurality of candidate objects to obtain the target intersection-over-union ratio can specifically be: removing second intersection-over-union ratios with lower values in the second intersection-over-union ratios to obtain the target intersection-over-union ratio. The second intersection-over-union ratio with a lower value indicates that the overlapping part between the candidate annotation rectangular region and the candidate recognition rectangular region is less, which can be caused by the misrecognition of the initial object recognition model, positioning deviation and other error factors. Removing the second intersection-over-union ratio with a lower value can effectively avoid the interference of the error factors on the subsequent model training, and improve the accuracy of the trained object recognition model.
[0054] Exemplarily, an index of the target intersection-over-union ratio can be obtained, candidate recognition boundary information corresponding to the index in the candidate recognition boundary information of the plurality of candidate objects is taken as the recognition boundary information, and candidate annotation boundary information corresponding to the index in the candidate annotation boundary information of the candidate object is taken as the annotation boundary information.
[0055] In a possible implementation, the screening of the second intersection-over-union ratios corresponding to the plurality of candidate objects to obtain the target intersection-over-union ratio can be implemented in the following manner: Manner one: The plurality of second intersection-over-union ratios are sorted in descending order to obtain a plurality of sorted second intersection-over-union ratios, and the first K second intersection-over-union ratios in the plurality of sorted second intersection-over-union ratios are taken as the target intersection-over-union ratio.
[0056] Exemplarily, the number of the plurality of second intersection-over-union ratios is 20, and K is 10. The 20 second intersection-over-union ratios can be sorted in descending order, and the first 10 second intersection-over-union ratios in the 20 sorted second intersection-over-union ratios are determined, which are taken as the target intersection-over-union ratio.
[0057] It should be noted that K is exemplified by 10 in the above example, which is only an example and does not limit the number of the plurality of second intersection-over-union ratios and K. In actual application, the value of K can be flexibly set according to actual application, and the specific value of K is not limited in the embodiment of the present application.
[0058] Manner two: The second intersection-over-union ratios greater than or equal to an intersection-over-union ratio threshold value are determined from the plurality of second intersection-over-union ratios to obtain the target intersection-over-union ratio.
[0059] The intersection-over-union ratio threshold value is set in advance.
[0060] Based on the above scheme, the lower second intersection-over-union ratios in the plurality of second intersection-over-union ratios can be removed to obtain the target intersection-over-union ratio, which ensures that the obtained target intersection-over-union ratio can truly and accurately reflect the fitting degree between the real feature region of the sample object and the recognition feature region of the sample object, avoids the interference of accidental factors on the model training, and further improves the training accuracy.
[0061] In a possible implementation, the first boundary information characterizing the boundary of the first elliptical region is determined based on the target angle and the identified boundary information.
[0062] It should be understood that when the target angle is 0 or the target angle is not considered, the first elliptical region is the maximum inscribed ellipse of the identified rectangular region. When the target angle is not 0, the first elliptical region is obtained by rotating the maximum inscribed ellipse of the identified rectangular region around the center of the identified rectangular region based on the target angle.
[0063] In a possible implementation, the first boundary information characterizing the boundary of the first elliptical region is determined based on the target angle and the identified boundary information.
[0064] In a possible implementation, the target angle can be preset or automatically determined by the electronic device.
[0065] For example, the target angle is 0 degree.
[0066] Based on the above scheme, when the target angle is 0 degree, the first elliptical region is the region of the maximum inscribed ellipse of the identified rectangular region. In this way, the first elliptical region no longer has the background information carried by the top points of the four corners of the identified rectangular region, avoiding the interference of the background information on subsequent model training, and improving the recognition accuracy of the trained object recognition model.
[0067] In a possible implementation, the target angle is obtained by: obtaining an image of the identified rectangular region; performing pixel distribution analysis on the image of the identified rectangular region to obtain a main direction of the pixel distribution; and calculating an included angle between the main direction and a reference direction to obtain the target angle.
[0068] For example, principal component analysis can be used to perform pixel distribution analysis on the image of the identified rectangular region to obtain the main direction of the pixel distribution. A two-dimensional rectangular coordinate system is established with the top left corner of the sample image as the origin (0, 0) and the horizontal direction to the right as the x axis and the vertical direction upward as the y axis. The included angle between the main direction of the pixel distribution and the positive direction of the x axis is calculated and taken as the target angle.
[0069] Based on the above scheme, the target angle can be determined based on the angle between the main direction of the pixel distribution in the identified rectangular area and the reference direction. The first boundary information can then be determined based on the target angle and the identified boundary information. In this way, the first elliptical area indicated by the first boundary information can conform to the actual extension direction of the sample object, preserving the pixel information of the core area of the sample object while eliminating background information interference from the four vertices of the identified rectangular area, thereby improving the accuracy of the determined area where the sample object is located.
[0070] For example, the image for identifying the rectangular region may be a binary image, which may be referred to as a segmentation mask for identifying the rectangular region. When the initial object recognition model also performs a segmentation task, the segmentation mask may be output by the initial object recognition model.
[0071] In one possible implementation, in Mapped to the minor axis b of the ellipse, Mapped to the major axis of the ellipse, the center point ( ) as the center of the ellipse, the maximum inscribed ellipse of the identified rectangular area can be obtained. The equation of the maximum inscribed ellipse of the identified rectangular area can be expressed by formula (1): Formula (1) Where a is the length of the major semi-axis of the largest inscribed ellipse in the identified rectangular area, and b is the length of the minor semi-axis of the largest inscribed ellipse in the identified rectangular area. The first ellipse represented by the first elliptical area can be expressed using the following formula (2).
[0072] Formula (2) Assuming that the target angle is not 0 degrees, the process of determining formula (2) is explained in detail below.
[0073] Assume that the maximum inscribed ellipse is rotated counterclockwise around the center of the identification rectangular area , let the coordinates of a point on the largest inscribed ellipse before rotation be expressed as ( x , y ), the coordinates of the point after rotation are expressed as ( x ′, y ′). Set , , , .
[0074] Will x 'and y ′ is expanded according to the cosine trigonometric function and sine trigonometric function identities to obtain the rotated coordinates ( x ′, y) and the relationship between the rotation matrix x , y ) and the rotation matrix Equation (3) Equation (4) Equation (3) and (4) are written in matrix form, and the following equation (5) about the rotation matrix Equation (5) If there is a matrix , such that , where is the identity matrix and and are non-zero, it can be proved that is an invertible matrix. It is derived that the rotation matrix can satisfy equations (6) to (8). Wherein: Equation (6) Equation (7) Equation (8) It can be proved that is a rotation matrix, and the transpose of in equation (5) is equation (9) Equation (9) The equation (9) is brought into the original ellipse equation (1), and the rotated ellipse equation can be obtained, which can be expressed as the equation shown in the above equation (2).
[0075] S103, based on the labeled boundary information, determine the second boundary information representing the boundary of the second elliptical region, the second elliptical region is a region with the same center and shape as the maximum inscribed ellipse of the labeled rectangular region, and the angle between the major axis of the second elliptical region and the major axis of the maximum inscribed ellipse of the labeled rectangular region is the target angle.
[0076] It should be understood that when the target angle is 0 or the target angle is not considered, the second elliptical region is the maximum inscribed ellipse of the labeled rectangular region. When the target angle is not 0, the second elliptical region is obtained by rotating the maximum inscribed ellipse of the labeled rectangular region around the center of the labeled rectangular region based on the target angle.
[0077] In a possible implementation, the second boundary information representing the boundary of the second elliptical region is determined based on the target angle and the labeled boundary information, the second elliptical region being obtained by rotating the largest inscribed ellipse of the labeled rectangular region based on the target angle.
[0078] In a possible implementation, the target angle can be preset or automatically determined by the electronic device.
[0079] For example, the target angle is 0 degree.
[0080] Based on the above scheme, when the target angle is 0 degree, the second elliptical region is the region of the largest inscribed ellipse of the labeled rectangular region. In this way, the second elliptical region no longer has the background information carried by the four corners of the labeled rectangular region, avoiding the interference of the background information on subsequent model training, and improving the recognition accuracy of the trained object recognition model.
[0081] In a possible implementation, the target angle is obtained by: obtaining an image of the recognition rectangular region; performing pixel distribution analysis on the image of the recognition rectangular region to obtain a main direction of the pixel distribution; and calculating an included angle between the main direction and a reference direction to obtain the target angle.
[0082] Based on the above scheme, the target angle can be determined based on the included angle between the main direction of the pixel distribution of the recognition rectangular region and the reference direction, so as to determine the second boundary information based on the target angle and the labeled boundary information. In this way, the second elliptical region indicated by the second boundary information can fit the actual extension direction of the recognized sample object, retaining the pixel information of the core region of the sample object and eliminating the interference of the background information carried by the four corners of the labeled rectangular region, thereby improving the accuracy of the labeled region of the determined sample object.
[0083] For example, the second ellipse represented by the second elliptical region can be represented by formula (10): Formula (10) The derivation process of formula (10) can refer to the derivation process of formula (2) of the first ellipse, which will not be described here again.
[0084] S104, determine the first intersection-over-union of the first elliptical region and the second elliptical region based on the first boundary information and the second boundary information.
[0085] In a possible implementation, from the aspect of image drawing, the first ellipse region can be drawn based on the first boundary information, and the second ellipse region can be drawn in the same coordinate plane based on the second boundary information. Then, the pixel area of the intersection region between the first ellipse region and the second ellipse region is calculated, and the pixel area of the first ellipse region and the pixel area of the second ellipse region are calculated. The first intersection union ratio can be calculated by using the following formula (11): Formula (11) It should be understood that an image is composed of discrete pixel points. The pixel area of the first ellipse region refers to the total number of pixel points located in the first ellipse region. The pixel area of the second ellipse region refers to the total number of pixel points located in the second ellipse region. The pixel area of the intersection region between the first ellipse region and the second ellipse region refers to the total number of pixel points located in both the first ellipse region and the second ellipse region.
[0086] Based on the above scheme, the first boundary information can be converted into a visual graph, and the second boundary information can be converted into a visual graph, so as to directly obtain the pixel area of the intersection region, the pixel area of the first ellipse region, and the pixel area of the second ellipse region by using image processing technology, and calculate the first intersection union ratio, thereby improving the accuracy of the calculated first intersection union ratio and further ensuring the reliability of the first intersection union ratio.
[0087] In a possible implementation, based on the first boundary information and the second boundary information, the first intersection union ratio of the first ellipse region and the second ellipse region can be determined by the following manner: based on the labeled boundary information and the recognized boundary information, a minimum circumscribed rectangle region containing the labeled rectangular region and the recognized rectangular region is determined; a set containing a target number of sampling points in the minimum circumscribed rectangle region is obtained; based on the first boundary information and the second boundary information, a first number of sampling points belonging to the first ellipse region and belonging to the second ellipse region in the set is determined, and based on the first boundary information and the second boundary information, a second number of sampling points belonging to the union region of the first ellipse region and the second ellipse region in the set is determined; a ratio of the first number to the second number is calculated to obtain the first intersection union ratio.
[0088] Exemplarily, Figure 2 The region of the minimum circumscribed rectangle containing the labeled rectangular region and the recognized rectangular region is shown. As Figure 2 shown, a two-dimensional rectangular coordinate system can be established with the center of the minimum circumscribed rectangle as the origin, the horizontal right direction as the positive direction of the x axis, and the vertical upward direction as the positive direction of the y axis. A set containing a target number of sampling points in the minimum circumscribed rectangle region is obtained from the region of the minimum circumscribed rectangle, and the target number can be represented by t. The set can be represented by Point{ ( ), ( )…( )}. t≤10000.
[0089] The first ellipse represented by the first ellipse area can be expressed by the above formula (2). The second ellipse represented by the second ellipse area can be expressed by the above formula (10). Point{( ), ( )…( )} are substituted into formula (2) and formula (10) respectively. When the coordinates of the sampling points satisfy ≤1, and ≤1, it means that the sampling point belongs to the first elliptical area and the second elliptical area. ≤1, and does not satisfy When ≤1, it indicates that ( ) does not belong to the first elliptical area and does not belong to the second elliptical area, that is, the sampling point belongs to the background area. Taking the target number of 1000 sampling points as an example, according to this method, the first number of sampling points belonging to the first and second elliptical areas among the 10,000 sampling points is, for example, 4000, and the number of sampling points belonging to the background area is 5000. Then, the second number of sampling points belonging to the union of the first and second elliptical areas is 10,000 - 5000 = 5000. The first intersection-over-union ratio is first number / second number = 4000 / 5000 = 4 / 5.
[0090] For example, the target number is 10,000, and these 10,000 points may be all points in the minimum circumscribed rectangular area. Figure 3 The first elliptical area, the second elliptical area, the intersection area, and the minimum bounding rectangle area are shown. When the first intersection-over-union ratio is 4 / 5, this means that a large number of the 10,000 points belong to the background area. Figure 3 also clearly shows that background points account for a large proportion and are numerous.
[0091] Based on the above solution, the first intersection-over-union ratio is calculated by using sampling point distribution statistics, which simplifies the process of determining the first intersection-over-union ratio and improves the efficiency of determining the first intersection-over-union ratio.
[0092] It should be noted that for the same sample object, the ellipse intersection-and-union ratio (first intersection-and-union ratio) and the rectangle intersection-and-union ratio can be different. Figure 4 Explanation is given by way of example.
[0093] like Figure 4As shown in (a) in FIG. 6, the rectangular intersection-over-union ratio = (area of intersection region) / (area of identified rectangular region + area of labeled rectangular region - intersection region) = 0.6954. As shown in (b) in FIG. 6, the calculated elliptical intersection-over-union ratio is 0.6473. Obviously, the rectangular intersection-over-union ratio and the elliptical intersection-over-union ratio are different. Figure 4
[0094] The above is an example of the rectangular intersection-over-union ratio being greater than the elliptical intersection-over-union ratio, which is only an example and does not constitute a limitation on the size relationship between the rectangular intersection-over-union ratio and the elliptical intersection-over-union ratio. In actual application, the elliptical intersection-over-union ratio can also be greater than the rectangular intersection-over-union ratio.
[0095] In S105, the parameters of the initial object recognition model are adjusted based on the first intersection-over-union ratio, and an object recognition model is obtained.
[0096] It should be understood that the first intersection-over-union ratio and the loss value are negatively correlated, that is, the higher the first intersection-over-union ratio, the lower the loss value, and the lower the first intersection-over-union ratio, the higher the loss value.
[0097] In the embodiments of the present application, a plurality of first intersection-over-union ratios corresponding to the sample images can be obtained in the above manner, and the first intersection-over-union ratio is used to adjust the model parameters of the initial object recognition model to obtain an object recognition model.
[0098] For example, the number of first intersection-over-union ratios is K.
[0099] In the embodiments of the present application, the image to be processed can be input into the object recognition model, and the object recognition model can process the image to be processed to locate the position of the target object existing in the object recognition model.
[0100] For example, taking the screen defect detection scenario as an example, the image to be processed can be a screen image to be recognized, and the object recognition model can be a screen defect recognition model. After processing the screen image to be recognized using the screen defect recognition model, the region where the defect exists can be recognized. The target object is a defect, such as a bright-dark spot or a bright-dark spot.
[0101] In the training process, the sample screen defect image can be input into the initial object recognition model, and the initial object recognition model can output the recognition boundary information of the defect in one training. Then, based on the recognition boundary information of the defect and the labeled boundary information of the defect, the elliptical intersection-over-union ratio can be determined to quantify the detection result, and the model parameters in the initial object recognition model can be adjusted. After multiple iterations of training, the optimal model is obtained, and the model parameter file of the optimal one-time training is saved. The optimal model file is output as the final model. When the initial object recognition model also performs a segmentation task, the segmentation mask of the defect can be output.
[0102] It should be understood that the training method provided in this paper can not only be applied to the screen defect detection scene, but also to the detection scene of the shape of the detection object related to the ellipse, such as the screw defect detection scene. The specific application scene is not limited in the embodiments of the application.
[0103] The embodiments of the application provide a training method, which determines first boundary information representing the boundary of the first elliptical region by using the obtained identification boundary information, and determines second boundary information representing the boundary of the second elliptical region by using the obtained labeling boundary information, determines the first intersection-over-union of the first elliptical region and the second elliptical region based on the first boundary information and the second boundary information, adjusts the parameters of the initial object recognition model based on the first intersection-over-union, and obtains the object recognition model. In this way, since the outline of the sample object is more consistent with the ellipse, the model training is performed by using the ellipse intersection-over-union (i.e. the first intersection-over-union) in the model training process, which reduces the interference of background information, so that the model can learn the real features of the sample object in the training process, and the recognition accuracy of the object recognition model obtained by training is improved.
[0104] The above detailed explanation is the training process of the object recognition model, and the application process of the object recognition model is explained in detail below.
[0105] It should be understood that the electronic device used to train the object recognition model and the electronic device used to apply the object recognition model can be the same electronic device, or different electronic devices.
[0106] For example, the electronic device used to train the object recognition model is a server, and the electronic device used to apply the object recognition model is a terminal device. It can also be that the electronic device used to train the object recognition model is a server, and the electronic device used to apply the object recognition model is a terminal device. Of course, the electronic device used to train the object recognition model and the electronic device used to apply the object recognition model can both be servers, or the electronic device used to train the object recognition model and the electronic device used to apply the object recognition model can both be terminal devices.
[0107] Figure 5 FIG. 1 is a flowchart of an image processing method provided by the embodiments of the application. The image processing method can be applied to an electronic device, and the image processing method includes S501 and S502.
[0108] S501, obtaining a to-be-processed image.
[0109] In the embodiments of the application, the to-be-processed image can be obtained in response to the input operation of a user; or the to-be-processed image is downloaded from a network by the electronic device; or the to-be-processed image can be sent to the electronic device by another device.
[0110] S502, the object recognition model is used to recognize the to-be-processed image to determine a target object in the to-be-processed image, and a shape of the target object is related to an ellipse.
[0111] Exemplarily, taking a screen defect detection scenario as an example, the to-be-processed image can be a screen image to be recognized, and the object recognition model can be a screen defect recognition model. After the screen defect recognition model is used to process the screen image to be recognized, a region in which a defect exists is recognized from the screen image to be recognized.
[0112] The training process of the object recognition model can refer to the training process of the object recognition model in the corresponding embodiment. Details are not described herein again. Figure 1 The training process of the object recognition model in the corresponding embodiment. Details are not described herein again.
[0113] The image processing method provided in the embodiment of the application is used to determine a target object in a to-be-processed image. Since the contour of the sample object is more close to an ellipse, the ellipse intersection-over-union ratio (i.e., the first intersection-over-union ratio) is used for model training in the model training process, the interference of background information is reduced, the model can learn the real features of the sample object in the training process, and the recognition accuracy of the object recognition model is improved.
[0114] The training method and the image processing method of the embodiment of the application are described in detail above in combination with Figure 1 and Figure 5 The device embodiment of the application will be described in detail below in combination with Figure 6 and Figure 7 It should be understood that the image processing device in the embodiment of the application can execute various methods of the foregoing embodiments of the application, that is, the specific working processes of various products below, and the corresponding processes in the foregoing method embodiments can be referred to.
[0115] Figure 6 is a schematic structural diagram of an image processing device provided in the embodiment of the application. The image processing device 600 can include an acquisition unit 610, a processing unit 620, and an adjustment unit 630.
[0116] In some embodiments, the image processing device 600 can be used to implement each step of the method shown in Figure 1
[0117] The acquisition unit 610 is configured to acquire labeled boundary information of a sample object of a sample image and recognition boundary information of the sample object, the labeled boundary information represents a boundary of a labeled rectangular region in which the sample object is located in the sample image, the recognition boundary information represents a boundary of a recognition rectangular region of the sample object in the sample image, and a shape of the sample object is related to an ellipse.
[0118] The processing unit 620 is configured to determine first boundary information representing a boundary of a first elliptical region based on the recognition boundary information, the first elliptical region being a region with the same center and the same size as a maximum inscribed ellipse of the recognition rectangular region, and an included angle between a major axis of the first elliptical region and a major axis of the maximum inscribed ellipse of the recognition rectangular region being a target angle; determine second boundary information representing a boundary of a second elliptical region based on the annotation boundary information, the second elliptical region being a region with the same center and the same shape as a maximum inscribed ellipse of the annotation rectangular region, and an included angle between a major axis of the second elliptical region and a major axis of the maximum inscribed ellipse of the annotation rectangular region being the target angle; and determine a first intersection-over-union of the first elliptical region and the second elliptical region based on the first boundary information and the second boundary information.
[0119] The adjusting unit 630 is configured to adjust parameters of an initial object recognition model based on the first intersection-over-union to obtain the object recognition model.
[0120] Optionally, the obtaining unit 610 is further configured to obtain the image of the recognition rectangular region.
[0121] The processing unit 620 is further configured to perform pixel distribution analysis on the image of the recognition rectangular region to obtain a main direction of pixel distribution, and calculate an included angle between the main direction and a reference direction to obtain the target angle.
[0122] Optionally, the target angle is 0 degrees.
[0123] Optionally, the processing unit 620 is further configured to determine a minimum circumscribed rectangular region containing the annotation rectangular region and the recognition rectangular region based on the annotation boundary information and the recognition boundary information.
[0124] The obtaining unit 610 is further configured to obtain a set containing a target number of sampling points in the minimum circumscribed rectangular region.
[0125] The processing unit 620 is further configured to determine a first number of sampling points in the set that belong to the first elliptical region and belong to the second elliptical region based on the first boundary information and the second boundary information, and determine a second number of sampling points in the set that belong to a union region of the first elliptical region and the second elliptical region based on the first boundary information and the second boundary information; and calculate a ratio of the first number and the second number to obtain the first intersection-over-union.
[0126] Optionally, the obtaining unit 610 is further configured to obtain candidate annotation boundary information of a plurality of candidate objects in the sample image, the candidate annotation boundary information representing boundaries of candidate annotation rectangular regions in which the candidate objects are located in the sample image.
[0127] The processing unit 620 is further configured to process the sample image by using the initial object recognition model to obtain candidate recognition boundary information of the plurality of candidate objects.
[0128] determining the annotation boundary information of the sample object from the plurality of candidate annotation boundary information, and determining the recognition boundary information of the sample object from the plurality of candidate recognition boundary information.
[0129] Optionally, the processing unit 620 is further configured to, for the candidate recognition boundary information and the candidate annotation boundary information of each candidate object, determine a second intersection-over-union ratio between the candidate annotation rectangular region and the candidate recognition rectangular region based on the candidate recognition boundary information and the candidate annotation boundary information; filter the second intersection-over-union ratios corresponding to the plurality of candidate objects to obtain a target intersection-over-union ratio; take the candidate recognition boundary information corresponding to the target intersection-over-union ratio from the candidate recognition boundary information of the plurality of candidate objects as the recognition boundary information, and take the candidate annotation boundary information corresponding to the target intersection-over-union ratio from the candidate annotation boundary information of the candidate object as the annotation boundary information.
[0130] Optionally, the processing unit 620 is further configured to sort the plurality of second intersection-over-union ratios in descending order to obtain a plurality of sorted second intersection-over-union ratios, and take the first K second intersection-over-union ratios in the plurality of sorted second intersection-over-union ratios as the target intersection-over-union ratio, where K is a positive integer greater than 1; or determine the second intersection-over-union ratio greater than or equal to an intersection-over-union ratio threshold from the plurality of second intersection-over-union ratios to obtain the target intersection-over-union ratio.
[0131] In some other embodiments, the image processing apparatus 600 can not include the adjusting unit 630. The image processing apparatus 600 can be configured to implement each step of the method shown in FIG. 7. Figure 5
[0132] The obtaining unit 610 is configured to obtain a to-be-processed image.
[0133] The processing unit 620 is configured to recognize the to-be-processed image by using an object recognition model to determine a target object in the to-be-processed image, where the shape of the target object is related to an ellipse.
[0134] It should be noted that the image processing apparatus 600 described above is embodied in the form of functional units. The term “unit” herein can be implemented in the form of software and / or hardware, and no specific limitation is made thereto.
[0135] For example, the “unit” can be a software program, a hardware circuit, or a combination of the two, which implements the above functions. The hardware circuit can include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combination logic circuit, and / or other suitable components that support the described functions.
[0136] Therefore, the units of each example described in the embodiments of the present application can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0137] Figure 7 is a structural schematic diagram of an electronic device 700 provided by an embodiment of the present application. As shown in the figure, the electronic device 700 of this embodiment includes at least one processor 701 (only one processor is shown in the figure), a memory 702, and a computer program 703 stored in the memory 702 and executable on the at least one processor 701, and the processor 701 implements the steps of any method embodiment described above when executing the computer program 703. Figure 7 Figure 7 The processor 701 can be a central processing unit (CPU), and the processor 701 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0138] Those skilled in the art can understand that Figure 7 The electronic device 700 is only an example and does not constitute a limitation on the electronic device 700, and can include more or fewer components than shown, or combine certain components, or different components, for example, can also include input / output devices, network access devices, etc.
[0139] The processor 701 can be a central processing unit (CPU), and the processor 701 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0140] The memory 702 may, in some embodiments, be an internal storage unit of the electronic device 700, such as a hard disk or a memory of the electronic device 700. The memory 702 may, in other embodiments, also be an external storage device of the electronic device 700, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, or the like, equipped on the electronic device 700. Further, the memory 702 may also include both an internal storage unit and an external storage device of the electronic device 700. The memory 702 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as program codes of computer programs, and the like. The memory 702 may also be used to temporarily store data that has been output or is to be output.
[0141] It should be noted that the information interaction, execution process, and the like between the above apparatuses / units are based on the same concept as the method embodiments of the present application, and specific functions and technical effects brought by the same can be referred to the method embodiments part for description, which will not be repeated here.
[0142] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the purpose of mutual distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0143] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement any of the above method embodiments.
[0144] The embodiments of the present application provide a computer program product, which can implement any of the above method embodiments when running.
[0145] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the computer program for instructing the related hardware to complete all or part of the processes in the above-mentioned embodiments can be stored in a computer readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium at least includes any entity or device capable of carrying the computer program code to the photographing device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some possible implementation manners, the computer readable medium can not be an electrical carrier signal and a telecommunication signal.
[0146] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.
[0147] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0148] In the embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other ways. For example, the above-described system embodiments are only schematic, for example, the division of modules or units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0149] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may also be distributed to multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.
[0150] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A training method, characterized in that: include: Acquire annotated boundary information of a sample object in a sample image and recognized boundary information of the sample object, wherein the annotated boundary information represents a boundary of a marked rectangular region where the sample object is located in the sample image, and the recognized boundary information represents a boundary of the recognized rectangular region of the sample object in the sample image, wherein a shape of the sample object is related to an ellipse; Determining, based on the identified boundary information, first boundary information representing a boundary of a first elliptical area, where the first elliptical area is an area having the same center and the same size as a maximum inscribed ellipse of the identified rectangular area, and an angle between a major axis of the first elliptical area and a major axis of the maximum inscribed ellipse of the identified rectangular area is a target angle; Determining, based on the annotated boundary information, second boundary information representing a boundary of a second elliptical area, where the second elliptical area is an area having the same center and the same shape as a maximum inscribed ellipse of the annotated rectangular area, and an angle between a major axis of the second elliptical area and a major axis of the maximum inscribed ellipse of the annotated rectangular area is the target angle; determining a first intersection-over-union ratio (IOR) of the first elliptical area and the second elliptical area based on the first boundary information and the second boundary information; Based on the first intersection-over-union ratio, parameters of the initial object recognition model are adjusted to obtain an object recognition model.
2. The method according to claim 1, characterized in that The target angle is obtained by: Acquire an image of the identified rectangular area; Performing pixel distribution analysis on the image of the identified rectangular area to obtain a main direction of pixel distribution; The angle between the main direction and the reference direction is calculated to obtain the target angle.
3. The method according to claim 1, characterized in that The target angle is 0 degrees.
4. The method according to any one of claims 1 to 3, characterized in that The determining, based on the first boundary information and the second boundary information, a first intersection-over-union ratio between the first elliptical area and the second elliptical area includes: Determining a minimum circumscribed rectangular area including the marked rectangular area and the identified rectangular area based on the marked boundary information and the identified boundary information; Obtain a set of sampling points containing the target number in the minimum bounding rectangular area; Determining, based on the first boundary information and the second boundary information, a first number of sampling points in the set that belong to the first elliptical area and the second elliptical area, and determining, based on the first boundary information and the second boundary information, a second number of sampling points in the set that belong to an area that is a union of the first elliptical area and the second elliptical area; The ratio of the first number to the second number is calculated to obtain the first intersection-over-union ratio.
5. The method according to any one of claims 1 to 3, characterized in that: The obtaining of the labeled boundary information of the sample object of the sample image and the identified boundary information of the sample object includes: Acquire candidate annotation boundary information of a plurality of candidate objects of the sample image, wherein the candidate annotation boundary information represents a boundary of a candidate annotation rectangular region where the candidate objects are located in the sample image; Using an initial object recognition model, the sample image is processed to obtain candidate recognition boundary information of the multiple candidate objects, where the candidate recognition boundary information represents boundaries of candidate recognition rectangular regions where the candidate objects are located in the sample image; The labeled boundary information of the sample object is determined from the plurality of candidate labeled boundary information, and the recognized boundary information of the sample object is determined from the plurality of candidate recognized boundary information.
6. The method according to claim 5, characterized in that The determining the labeling boundary information of the sample object from the plurality of candidate labeling boundary information, and determining the recognition boundary information of the sample object from the plurality of candidate recognition boundary information, includes: For each candidate object, the candidate recognition boundary information and the candidate annotation boundary information are used to determine a second intersection-over-union ratio between the candidate annotation rectangular area and the candidate recognition rectangular area based on the candidate recognition boundary information and the candidate annotation boundary information. Screening the second intersection-over-union ratios corresponding to the multiple candidate objects to obtain a target intersection-over-union ratio; The candidate recognition boundary information of the candidate objects corresponding to the target IoU ratio is used as the recognition boundary information, and the candidate annotation boundary information of the candidate objects corresponding to the target IoU ratio is used as the annotation boundary information.
7. The method according to claim 6, characterized in that The step of screening the second IoU corresponding to the plurality of candidate objects to obtain a target IoU includes: Sorting the plurality of second intersection-over-union ratios in descending order to obtain a plurality of sorted second intersection-over-union ratios, and using the top K second intersection-over-union ratios among the plurality of sorted second intersection-over-union ratios as target intersection-over-union ratios, where K is a positive integer greater than 1; or A second IoU that is greater than or equal to an IoU threshold is determined from the plurality of second IoUs to obtain the target IoU.
8. An image processing method, characterized in that: The method comprises: Get the image to be processed; An object recognition model is used to identify the image to be processed to determine a target object in the image to be processed, wherein the shape of the target object is related to an ellipse, and the object recognition model is trained using the method described in any one of claims 1 to 7.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the electronic device executes the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 8 is executed.
Citation Information
Patent Citations
Method and system for detecting target in image based on Libra-RCNN and elliptical features
CN114445482A
Target detection method and device, equipment and storage medium
CN116778452A
Image target identification method based on elliptical anchor frame
CN117036673A
Cell nucleus instance segmentation model and method based on attention and ellipse regularization
CN120013912A
Learning apparatus, learning method, and learning program
JP2025067344A