Image processing method, corresponding device and storage medium
By using angled annotation lines instead of traditional horizontal annotation boxes, the training images are annotated, which solves the problem of low linear defect detection accuracy in the prior art, and achieves higher annotation accuracy and model training accuracy.
Patent Information
- Application Number
- CN202510421273.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing rectangular frame defect detection methods are low when detecting linear defects such as slip lines, and the pixel proportion of real defects is very low.
Angled annotation lines are used instead of the traditional horizontal annotation boxes to mark the training images, and preset machine learning models are trained using the annotation lines to improve the training accuracy of the model and the defect detection accuracy after deployment.
The labeling accuracy is improved, so that the pixel proportion of the true defects marked is high, thereby improving the training accuracy and defect detection accuracy of the model.
Smart Images

Figure CN119941720A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semiconductor technology, and in particular to an image processing method, and corresponding equipment and storage medium. Background Art
[0002] When surface inspection devices are used for semiconductor testing, image processing methods based on ML (Machine Learning) models are usually used to determine the defects on the sample surface (such as wafers). Specifically, defects are marked with rectangular boxes, a training data set is established, and the ML model is trained. The image to be tested is then inspected by the ML model, and the ML model will use a rectangular box to frame the defect. However, when the defect is a linear defect such as a slip line, the pixel ratio of the real defect in the area framed by the rectangular box is very low, resulting in low detection accuracy. Summary of the invention
[0003] The technical solution of the present application is to provide an image processing method, and corresponding equipment and storage medium, in order to solve the problem of low detection accuracy in the existing rectangular frame defect detection method.
[0004] The first aspect of the present application provides an image processing method, including: obtaining a training data set based on an image to be trained, and obtaining the training data set based on the image to be trained includes: using a marking line to mark a first linear target of a first sample in the image to be trained, wherein the marking line is set with a first angle; using the training data set to train a preset machine learning model to obtain a trained machine learning model; and detecting a second linear target of a second sample in the image to be tested by using the trained machine learning model.
[0005] The present application also provides an image processing device, which includes a processor and is configured to: obtain a training data set based on the image to be trained, wherein obtaining the training data set based on the image to be trained includes: marking a first linear target of a first sample in the image to be trained using a marking line, wherein the marking line is set with a first angle; training a preset machine learning model using the training data set to obtain a trained machine learning model; and detecting a second linear target of a second sample in the image to be tested using the trained machine learning model.
[0006] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image processing method described above are implemented.
[0007] The technical solution of the present invention has the following beneficial effects.
[0008] In the invention provided by the technical solution of the present invention, the present application replaces the traditional horizontal annotation box with an angled annotation line to annotate the training image, so that the pixel ratio of the annotated real defects is high, the annotation accuracy is improved, and the training accuracy of the model and the defect detection accuracy after deployment are further improved. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0009] Figure 1 It is a schematic diagram of a first embodiment of the image processing method of the present invention; Figure 2 A schematic diagram of an embodiment of a sample image in the present invention; Figure 3 is a schematic diagram of a second embodiment of the image processing method of the present invention; Figure 4 A schematic diagram of an embodiment of a priori marking of targets in the present invention; Figure 5 is a schematic diagram of another embodiment of the sample image in the present invention; Figure 6 is a schematic diagram of a third embodiment of the image processing method of the present invention; Figure 7 It is a structural schematic diagram of the control system in the present invention. DETAILED DESCRIPTION
[0010] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0011] In the description of the present invention, it should be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance. In the present invention, "each" includes one and more than two quantities.
[0012] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0013] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not constitute a conflict with each other, and should all be considered to be within the scope of this specification.
[0014] The present invention relates to a preset machine learning model, including a model training stage and a model application stage. Except that model optimization needs to be performed in the model training stage, the main processing flows involved in the two stages are basically the same. In the following embodiments, the two stages may involve: a general description of related feature descriptions such as between an image to be trained and an image to be tested, between a first sample and a second sample, between a first linear target and a second linear target, and between a first angle and a third angle. Therefore, when performing a general description, the aforementioned corresponding related feature descriptions are represented by sample images, samples, linear targets, angles, etc., which are applicable to the application scenarios of the aforementioned corresponding two related feature descriptions.
[0015] For the sake of convenience, the basic process of the image processing method in the present invention is described below. Figure 1 , provides a first embodiment of the image processing method in the present invention, which is specifically as follows: 110. Acquire a training data set according to the image to be trained, wherein acquiring the training data set according to the image to be trained includes: annotating a first linear target of a first sample in the image to be trained by using an annotation line, wherein the annotation line is set at a first angle; The present invention aims to specifically detect linear targets in samples, wherein linear targets refer to defects with linear shapes (exemplarily quantified as: length-width ratio ≥ preset ratio threshold), such as slip lines (length 100-1000μm, width 0.1-1μm, length-width ratio ≥1000), electron migration lines (length 10-100μm, width 10-50nm, length-width ratio ≥1000), laser annealing scan lines (length is the full sample size, width is 1-5nm, length-width ratio>1000), etc.; exemplarily, the preset ratio threshold is set to 1000. Linear targets may include at least one of slip lines, electron migration lines, laser annealing scan lines, etc. In addition, scratches and cracks may also be present.
[0016] Based on this, Figure 2 As shown, it can be determined that when the linear target 210 is formed in the sample image 220 of the sample, it has the following planar characteristics: it has a very large aspect ratio (extremely small width), so that the linear target has an extremely small characteristic area; the generated position has an angle randomness, and it is difficult to strictly maintain an angle parallel or perpendicular to the sample image; each different linear target has a distinguishing significant feature (linear feature) between them, or relative to the non-linear target; different linear targets may be densely distributed or overlapped. Using angled annotation lines to train a preset machine learning model is conducive to overcoming the problems of small characteristic area, random angles, and clustered appearance of linear targets, while using the distinguishing significant features of linear targets to speed up detection efficiency.
[0017] Therefore, when the first linear target is marked by marking lines, the area range of the first linear target can be accurately covered; in the traditional rectangular marking box / anchor box marking, the first linear target only occupies the diagonal position of the rectangular marking box / anchor box, so that the pixel occupancy ratio is low; by replacing the rectangular marking box / anchor box with marking lines, problems such as low model learning efficiency, low positioning accuracy, low classification performance, high risk of overfitting, unstable training, and low data utilization in the training process can be avoided, thereby greatly improving the deep learning effect of the linear target in the first sample.
[0018] At the same time, due to the extremely large aspect ratio of the first linear target, the width of the first linear target may not be involved in the annotation positioning. Therefore, when characterizing the position of the first linear target on the image to be trained, we can mainly focus on two parameters: length and angle. Compared with the annotation box / anchor box that needs to annotate 4 vertex coordinates (x1, y1; x2, y2; x3, y3; x4; y4), the annotation line only needs two coordinate point coordinates and one angle value (x1, y1; x2, y2; A) to annotate the position of the first linear target, thereby reducing the amount and dimension of annotation data.
[0019] 120. Use the training data set to train a preset machine learning model to obtain a trained machine learning model, and use the trained machine learning model to detect a second linear target of a second sample in the image to be tested.
[0020] In this embodiment, at least an optical imaging system is provided to collect a sample image by utilizing the optical characteristics of the sample's linear target, and the sample's linear target and the sample's background in the sample image can be distinguished. The optical characteristics include the difference in refractive index between the sample's linear target and the background, and the edge of the linear target having enhanced light scattering.
[0021] Specifically, the preset machine learning model learns the distinguishing features of the first linear target in the first sample relative to the background through the training image of the first sample, including at least the brightness, contour, and grayscale of the difference. Based on the distinguishing features learned in the training image, the first linear target is detected, and the detection error is checked through the annotation line, so as to adjust the preset machine learning model until the error between the detection result and the annotation line is reduced to a preset minimum threshold.
[0022] In this embodiment, each marked line is also marked with at least a defect category / non-defect category, and the preset machine learning model also outputs the detected defect category / non-defect category at the same time; at the same time, the defect category / non-defect category is used to verify the detection error of the preset machine learning model.
[0023] Specifically, in the aforementioned optical imaging system, interference phenomena and phase difference images are used to characterize the difference in optical properties between the linear target and the background of the sample, and based on this, the image features of the linear target are highlighted. Exemplarily, differential interference contrast (DIC) and dark field illumination can be used to capture multi-frame phase difference images of the sample in coordination with interference phenomena such as Wollaston prisms and phase modulation such as piezoelectric ceramic micro-shift prisms, and reconstruct phase gradient images to highlight the edge contrast of the linear target relative to the background. Based on the phase gradient map, the continuous phase is restored by integrating the gradient field and converted into a spatial domain image to form a sample image in a two-dimensional plane represented by grayscale.
[0024] In this embodiment, after the preset machine learning model is trained, in addition to the step of performing model optimization using the marked lines, the second linear target of the second sample in the image to be tested is detected according to the same step process.
[0025] The following is an explanation of the annotation line annotation process in the first embodiment of the image processing method of the present invention. Figure 3 , as shown below: 310. Using a labeling line, label only one of the first state feature or the second state feature of the first sample in the to-be-trained image; 320. Obtain a training data set according to the annotation of the first state feature or the second state feature; wherein a difference between a first angle of the annotation line in the image to be trained and a second angle of the first state feature or the second state feature in the image to be trained is less than a preset difference threshold; or In this embodiment, the first linear target presents at least one of the associated first state features and second state features in the image to be trained; indicating that the first linear target can present the first state feature alone, the second state feature alone, or the first state feature and the second state feature at the same time; the first state feature and / or the second state feature are used for the subsequent preset machine learning model to detect the first linear target.
[0026] In one embodiment, the pixel value of the first state feature is greater than a first threshold, the pixel value of the second state feature is less than a second threshold, the first threshold is greater than or equal to the second threshold, exemplarily, the pixel value of the first state feature is greater than 200 grayscale values of the first threshold, the pixel value of the second state feature is less than 50 grayscale values of the second threshold, the first state feature and the second state feature present different degrees of brightness, the first feature state presents a bright line, and the second feature state presents a dark line; at least part of the first state feature and the second state feature are presented simultaneously, wherein the first state feature or the second state feature is used as a common feature, and only the common feature can be annotated. It is further limited that the first linear target can present the first state feature or the second state feature alone, or can present the first state feature and the second state feature simultaneously.
[0027] Exemplarily, the interference phenomenon and phase difference image are used to characterize the difference in optical properties between the linear target and the background of the sample, based on which the image features of the linear target are highlighted. If the interference fringes are offset based on a phase difference that is a multiple of 2π, the linear target forms bright fringes in the sample image; if the interference fringes are offset based on a phase difference that is an odd multiple of π, the linear target forms dark fringes in the sample image, which may be accompanied by bright fringes.
[0028] Interference fringes migrate according to different phase difference multiples, forming strip defects with different significant features on the surface of the sample image, such as significant features including bright lines (bright fringes) and dark lines (dark fringes); the preset machine learning model can detect the linear target of the sample in the sample image based on bright lines or dark lines. Among them, with regard to the formation of dark lines, the dark lines come from the shadows of concave defects or convex defects, so there must be bright lines near convex defects, and there are no bright lines near concave defects; therefore, dark lines are preferred to characterize linear targets, and there is no ambiguity in characterization. That is, the first state feature can be a dark line, and the second state feature can be a bright line.
[0029] In one embodiment, if the annotation line is generated by manual annotation, the difference between the first angle and the second angle is almost 0; if the annotation line is generated by an automated annotation tool or assisted in generating the annotation line, an error within a preset difference threshold is allowed between the first angle and the second angle.
[0030] For example, an automated annotation tool such as a pre-annotation model based on deep learning can be used to apply multi-angle annotation lines for annotation to adapt to the first linear target with arbitrary angles. Multi-task learning can also be introduced to simultaneously detect different categories of the first linear target, such as slip lines, electron migration lines, laser annealing scan lines, etc.
[0031] The first linear target mentioned above includes two main parameters: length and angle. After the initial annotation line is generated by using the multi-angle annotation line of the automatic annotation tool, the angle and length of the annotation line can be dynamically adjusted according to the special morphological characteristics of the first linear target (no need to pay attention to the morphology in the width direction, only the morphology in the length direction and the morphology of the inclination angle).
[0032] According to the physical characteristics of the sample and the processing technology, the formation position of the first linear target may have certain regularity; for example, the slip line (first linear target) of the wafer (sample) usually has the following distribution characteristics: 1) Due to the stress concentration on the edge of the wafer, the slip line is often formed first in the edge area; 2) Due to the symmetry of the wafer, the extension line of the slip line usually points to the center of the circle or near the center of the circle. Therefore, in the pre-annotation model based on deep learning, a priori annotation template for multi-angle training tasks can be further learned.
[0033] For example, Figure 4 As shown, a standard prior annotation template M1 is provided for the distribution characteristics of the slip line on the wafer; during the annotation process / detection process, any center line (the line segment between the center C1 and the edge) in the prior annotation template M1 where the first state feature / second state feature is located is determined, and the second angle of the first state feature / second state feature is determined based on the known angle of the center line, so as to obtain the angle of the annotation line / first linear target. Among them, the center C1 and edge of the prior annotation template M1 are mapped to the center and edge of the wafer respectively.
[0034] For further information, please refer to Figure 4 , the extension line of the sliding line may also point to the vicinity of the circle center C1. Therefore, when the deep learning pre-annotation model learns the prior annotation template, the center of the extension line can be fitted to each annotated line to obtain the prior annotation target M2. Each center line passes through the circle center C1 and moves to the extension line center C2. Based on this, the angle of each center line (the line segment between the extension line center C2 and the edge) is re-determined.
[0035] 330. Acquire a labeling frame according to the labeling line, wherein the labeling frame at least covers the labeling line; In this embodiment, the first linear target can also be annotated using an annotation box generated based on the annotation line. Compared with the model training based on the annotation line, the model training using the annotation box is more conducive to accelerating the model convergence speed and optimizing the generalization ability and overfitting of the model. At the same time, compared with the traditional annotation box, the area of the annotation line can be strictly located, and the effective area of the first linear target in the annotation box can be increased.
[0036] In this embodiment, along the preset direction relative to the first angle may include: only along the first vertical direction and / or the second vertical direction relative to the first angle, or along the first vertical direction and / or the second vertical direction relative to the first angle, and the first extension direction and / or the second extension direction. According to different preset directions relative to the first angle, the relative position of the annotation line and the annotation box is determined, that is, the relative position of the first linear target and the annotation box is represented; when the model is optimized later, the relative position of the annotation box can be further used to calculate the detection error of the preset machine learning model. The preset annotation range can use any size unit suitable for measuring the image size.
[0037] In one embodiment, the preset annotation range includes a preset number of pixels; the annotation line is extended by the preset number of pixels along the extension direction and the vertical direction of the first angle, and a annotation frame is generated based on the extended pixels. The generated annotation frame follows the angular randomness of the annotation line and strictly annotates the area where the annotation line is located.
[0038] For example, Figure 5 As shown, the image to be trained includes annotation lines A, annotation lines B and annotation lines C. For example, the annotation line A is expanded by a first number of m pixels in the horizontal direction and by a second number of n pixels in the vertical direction to obtain a rectangular annotation frame L1 of size m+n. The annotation lines A and C have a large angle difference relative to the parallel or vertical direction of the image to be trained. For example, the annotation line A generates the annotation frame L1 by the annotation line expansion method, minimizing the area of the annotation frame L1, that is, maximizing the effective area of the labeled linear object; if the conventional anchor frame L2 is used for annotation, the area of the anchor frame L2 will be increased, resulting in an increase in the invalid area and a reduction in the positioning accuracy of the linear object.
[0039] In addition, the preset direction can also have an acute angle with the extension direction of the annotation line, so that the distance between the side length of the generated annotation box and the annotation line is reduced, and the invalid area is reduced accordingly, thereby improving the positioning accuracy of the linear target.
[0040] In one embodiment, the preset annotation range is set based on the preset generation deviation of the image to be trained, wherein the preset generation deviation refers to the deviation of the generation position of the first linear target in the image to be trained relative to the actual position during the generation of the image to be trained. Based on the principle of interference phenomenon and phase difference image, in the process of generating the sample image in the spatial domain by using integral function transformation, the generation position of the first linear target in the sample image will be offset relative to the actual position; based on this annotation range used to set the annotation box, the actual position of the first linear target can be found in the annotation box.
[0041] In another embodiment, the preset annotation range is set based on at least one of the width of the first linear target and the preset annotation error. The first linear target has a first width, and the annotation line has a second width. The second width may be smaller than the first width, so that the annotation line cannot completely cover the first linear target; even if the second width is equal to the first width, the annotation line may not be able to exactly cover the first linear target during the annotation process; the annotation range is set based on these two situations so that the annotation box can completely cover the entire first linear target.
[0042] 340. Annotate a first linear target of a first sample in the to-be-trained image by using the annotation frame to obtain the training data set; 350. Use the training data set to train a preset machine learning model to obtain a trained machine learning model, and use the trained machine learning model to detect a second linear target of a second sample in the image to be tested.
[0043] In this embodiment, the first linear target is labeled based on a priori labeling template or labeling box to update the training data set and improve labeling accuracy and labeling efficiency.
[0044] The following is an explanation of the training process of the preset machine in the first embodiment of the image processing method of the present invention. Please refer to Figure 6 , as shown below: 610. Extract a feature map of the image to be trained in the training data set; In this embodiment, the image to be trained can be a grayscale image (grayscale value is 0-255), including a first linear target represented by a dark line or a bright line; when the first linear target is represented by a dark line, the first linear target presents a lower grayscale value (leaning toward grayscale value 0) relative to the background of the first sample; when the first linear target is represented by a bright line, the first linear target presents a higher grayscale value (leaning toward grayscale value 255) relative to the background of the first sample; based on this, a feature map is extracted.
[0045] Prior to this, in the input training data set, the orientation of the annotation line can be expressed in the format of two-point coordinates + the first angle (x1, y1, x2, y2, A). Then, the training image is enhanced, including rotation, scaling, translation, flipping, etc., to improve the generalization ability of the machine learning model. And the pixel value of the training image is normalized from [0, 255] to [0, 1] to meet the input requirements of the machine learning model.
[0046] Specifically, after the training image is input into the preset machine learning model, multi-level features can be extracted through the backbone network (such as CSPDarknet, (Cross Stage Partial Darknet)). The features are fused using a feature fusion network (such as FPN (Feature Pyramid Networks) or PANet (Path Aggregation Network)) to generate a multi-scale feature map.
[0047] 620. Generate prediction information of the first linear target according to the feature map, wherein the prediction information at least includes a prediction area for the first linear target and a third angle of the prediction area; In this embodiment, the prediction network of the preset machine learning model includes a line prediction module and an angle prediction module; based on the input multi-scale feature map, the line prediction module gradually predicts the positioning line of the first linear target according to the scale from large to small; wherein, the larger scale prediction result is transmitted back to the smaller scale line prediction module, and fused with the smaller scale feature map to assist in locating the approximate position of the first linear target, and quickly perform more precise positioning line prediction on the smaller scale feature map according to the positioning position. The positioning line output by the line prediction module is predicted by the angle prediction module at a third angle of the positioning line in the image to be trained. The prediction network may further include a category prediction module, which is deployed with a Softmax function or a Sigmoid function, and is used to predict the linear target category / non-linear target category of the positioning line.
[0048] 630. Adjust the weight parameters of the preset machine learning model based at least on the prediction area and the third angle, and the annotation line and the first angle.
[0049] In this embodiment, the positioning error between the predicted area and the annotated line, and the angular error between the third angle and the first angle are determined; based on the positioning error and the angular error, the loss of the predicted area relative to the annotated line is calculated, and the weight parameters of the preset machine learning model are adjusted based on the loss. It is also possible to calculate the category error of the linear target category / non-linear target category predicted by the positioning line and the category of the annotated line, determine the classification loss, and then perform back propagation in combination with the aforementioned angular loss and positioning line loss to update the weight parameters of the preset machine learning model.
[0050] Specifically, the weights W1, W2 and W3 of the classification loss Loss1, the annotation line loss Loss2 and the angle loss Loss3 are preset, and the three losses are fused based on the weights to obtain the total loss value: Loss = Loss1*W1+Loss2*W2+Loss3*W4; then use automatic differentiation (Autograd) to calculate the gradient of the loss function to the model parameters, and the gradient calculation includes all trainable parameters of the backbone network, feature fusion network, and prediction network; perform gradient clipping (such as setting the maximum gradient norm) to avoid gradient explosion; finally use optimizers such as SGD and Adm to update the model parameters according to the gradient, where the update formula is: model parameters = original model parameters - learning rate * gradient.
[0051] In one embodiment, the predicted area includes a predicted line, a first offset value of the predicted line relative to the marked line is determined, and a positioning error is determined according to the first offset value. The first offset value includes a distance offset vector.
[0052] In another embodiment, the prediction area includes a prediction box, wherein the prediction box is a linear box whose width is less than a preset width threshold; determining the coverage length of the prediction box on the annotation line, and determining the positioning error based on the coverage length; or determining a second offset value of a center line along the third angle in the prediction box relative to the annotation line, and determining the positioning error based on the second offset value.
[0053] In this embodiment, the prediction box is a linear box with a width less than a preset width threshold, so that the linear box is in a long strip shape, reducing the positioning range of the first linear target; the preset width threshold is set small enough, such as 1 / 1000 of the length, or 2 / 3 / 4 times the number of pixels of the first linear target, or a fixed 10 / 20 / 50 pixels; the positioning error of the prediction box covering the annotation line in the width direction is ignored (the width is small enough), and only the positioning error in the length direction is calculated, which reduces the amount of calculation and can effectively represent the overall error with the annotation line.
[0054] In this embodiment, the center line of the prediction box is set along the center line of the third angle as the positioning position of the detected first linear target; based on the second offset value of the center line relative to the marked line, which is equivalent to the first offset value of the prediction line relative to the marked line, and also including the distance offset vector, the positioning error between the two in distance is determined.
[0055] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned function allocation can be completed by different functional units and modules based on needs, that is, the internal structure of the application device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0056] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0057] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0058] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0059] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected based on actual needs to achieve the purpose of the solution of this embodiment.
[0060] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0061] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0062] The present application also provides a control system. Figure 7 , which shows a schematic diagram of the structure of a control system provided by an embodiment of the present application. Figure 7 As shown, the control system 700 includes: a processor 70, a memory 71, a bus 72 and a communication interface 73, and the processor 70, the communication interface 73 and the memory 71 are connected via the bus 72; the memory 71 stores computer program instructions that can be executed by the processor 70, and the processor 70 executes the image processing method provided in any of the aforementioned embodiments of the present application when executing the computer program instructions.
[0063] The memory 71 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the device network element and at least one other network element is realized through at least one communication interface 73 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0064] The bus 72 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 71 is used to store programs, and the processor 70 executes the programs after receiving execution instructions. The image processing method disclosed in any implementation of the above-mentioned embodiment of the present application may be applied to the processor 70, or implemented by the processor 70.
[0065] The processor 70 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 70. The above processor 70 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 71, and the processor 70 reads the information in the memory 71 and completes the steps of the above method in combination with its hardware.
[0066] The control system provided in the embodiment of the present application and the image processing method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented therein.
[0067] An embodiment of the present application also provides a computer-readable storage medium corresponding to the image processing method provided in the aforementioned embodiment, on which computer program instructions are stored. When the computer program instructions are executed by a processor, they will implement the image processing method provided in any of the aforementioned embodiments.
[0068] It should be noted that examples of the computer-readable storage medium may include, but are not limited to, optical disks, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0069] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the image processing method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0070] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0071] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.
Claims
1. An image processing method, characterized in that: The method comprises: Acquiring a training data set according to the image to be trained, wherein acquiring the training data set according to the image to be trained comprises: marking a first linear target of a first sample in the image to be trained by using a marking line, wherein the marking line is set with a first angle; Using the training data set to train a preset machine learning model to obtain a trained machine learning model; The second linear target of the second sample in the image to be tested is detected by the trained machine learning model.
2. The image processing method according to claim 1, characterized in that: Obtaining a training data set based on the image to be trained also includes: Acquire a labeling frame according to the labeling line, wherein the labeling frame at least covers the labeling line; The first linear object of the first sample in the image to be trained is annotated by the annotation frame to obtain the training data set.
3. The image processing method according to claim 1, characterized in that: Acquiring a labeling frame according to the labeling line also includes: The preset annotation range of the annotation line is extended toward both sides of the annotation line along a preset direction to obtain an annotation frame, wherein the preset direction is perpendicular to the extension direction of the annotation line or has an acute angle with the extension direction of the annotation line.
4. The image processing method according to claim 2, characterized in that: The method further comprises: The preset annotation range is set based on a preset generation deviation of the image to be trained, wherein the preset generation deviation refers to a deviation of a generated position of the first linear target in the image to be trained relative to an actual position during the generation process of the image to be trained; or The preset annotation range is set based on at least one of a width of the first linear object and a preset annotation error.
5. The image processing method according to claim 1, characterized in that: The first linear target has at least one of a first state feature and a second state feature in the to-be-trained image, a pixel value of the first state feature is greater than a first threshold, a pixel value of the second state feature is less than a second threshold, and the first threshold is greater than or equal to the second threshold; Using the annotation lines to annotate the first linear target of the first sample in the image to be trained to obtain a training data set, including: Using a labeling line to label only one of the first state feature or the second state feature of the first sample in the to-be-trained image; Obtaining a training data set according to the labeling of the first state feature or the second state feature; The difference between a first angle of the annotation line in the image to be trained and a second angle of the first state feature or the second state feature in the image to be trained is less than a preset difference threshold.
6. The image processing method according to claim 5, characterized in that: The first linear target includes a common feature, which is a first state feature or a second state feature. Using a labeling line to label only one of the first state feature or the second state feature of the first sample in the image to be trained includes: only labeling the common feature.
7. The image processing method according to claim 1, characterized in that: The first linear target is a scratch, a crack or a slip line.
8. The image processing method according to any one of claims 1 to 7, characterized in that: Using the training data set to train a preset machine learning model includes: Extracting a feature map of an image to be trained from the training data set; Generate prediction information of the first linear target according to the feature map, wherein the prediction information at least includes a prediction area for the first linear target and a third angle of the prediction area; Adjust the weight parameters of the preset machine learning model based at least on the prediction area and the third angle, and the annotation line and the first angle.
9. The image processing method according to claim 8, characterized in that: Adjusting a weight parameter of the preset machine learning model based at least on the prediction area and the third angle, and the annotation line and the first angle, comprises: Determining a positioning error between the predicted area and the marked line, and an angular error between the third angle and the first angle; According to the positioning error and the angle error, the loss of the predicted area relative to the marked line is calculated, and the weight parameters of the preset machine learning model are adjusted based on the loss.
10. The image processing method according to claim 9, characterized in that: The prediction area includes a prediction line; Determining the positioning error between the predicted area and the marked line includes: A first offset value of the predicted line relative to the marked line is determined, and a positioning error is determined according to the first offset value.
11. The image processing method according to claim 9, characterized in that: The prediction area includes a prediction box, wherein the prediction box is a linear box with a width less than a preset width threshold; Determining the positioning error between the predicted area and the marked line includes: Determine the coverage length of the prediction box on the annotation line, and determine the positioning error according to the coverage length; or, A second offset value of a center line along the third angle in the prediction frame relative to the annotation line is determined, and a positioning error is determined according to the second offset value.
12. An image processing device, characterized in that: The device comprises a processor configured to: Acquire an image to be trained, and annotate a first linear target of a first sample in the image to be trained using an annotation line to obtain a training data set, wherein the annotation line is set with a first angle; Using the training data set to train a preset machine learning model to obtain a trained machine learning model; The second linear target of the second sample in the image to be tested is detected by the trained machine learning model.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image processing method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Defect detecting device and method for non-elevation reflective surface workpieces
CN109557101A
Image labeling method and device, storage medium and computer equipment
CN109902672A
Target detection model training method and device, electronic equipment and medium
CN116245193A
Rolling metal surface defect automatic labeling method based on multi-task self-adaptive model
CN119444759A