Mask image editing method and apparatus, electronic device, and storage medium
By using mask image editing methods in industrial automation, the problem of missing or redundant areas of mask image in workpiece detection is solved, and the detection accuracy and sorting accuracy are improved.
Patent Information
- Application Number
- PCT/CN2023/135640
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-05
AI Technical Summary
In industrial automation, when robots are used to sort workpieces, due to noise, clutter and occlusion, the workpiece detection results are often inaccurate, resulting in missing or redundant areas in the mask image, affecting the sorting accuracy.
A mask image editing method is proposed. By acquiring the workpiece image and inputting the trained workpiece detection model, after outputting the mask image, it is input to the trained mask image editing model associated with the type for editing. The edited mask image can supplement the missing area or remove the redundant area.
By introducing AI capabilities, the quality of masked images is improved, the accuracy of workpiece detection and the accuracy of industrial sorting are enhanced, and the editing workload is reduced.
Smart Images

Figure CN2023135640_05062025_PF_FP_ABST
Abstract
Description
Mask image editing method, device, electronic device and storage medium Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence (AI) technology, and in particular to a method, device, electronic device, and storage medium for editing a mask image. Background Art
[0002] With the increasing degree of industrial automation, intelligent industrial robots are gradually replacing manual labor in sorting tasks in various industrial scenarios, improving work efficiency while effectively reducing production costs. To achieve intelligent sorting, robots need to use machine vision technology to automatically identify and locate the mixed materials to be sorted.
[0003] With significant breakthroughs in feature extraction using deep learning technology, object detection based on deep learning can automatically extract more comprehensive image features. However, due to the complexity of industrial scenes (such as the presence of noise, clutter, and occlusion), workpiece detection results are often inaccurate, with areas often being missed or redundant.
[0004] Summary of the Invention
[0005] The embodiments of the present invention provide a method, device, electronic device and storage medium for editing a mask image.
[0006] A mask image editing method, comprising:
[0007] acquiring an artifact image including the artifact;
[0008] Inputting the workpiece image into a trained workpiece detection model, wherein the workpiece detection model is adapted to output a mask image of the workpiece and a type of the workpiece based on the workpiece image;
[0009] inputting the mask image into a trained mask image editing model associated with the type, the mask image editing model being suitable for editing mask images of artifacts of the type;
[0010] An edited mask image is received from the mask image editing model.
[0011] Therefore, introducing AI capabilities into the editing process of mask images improves the quality of mask images.
[0012] In one embodiment, inputting the mask image into a trained mask image editing model associated with the type includes:
[0013] The mask image and a primary prompt are input into the mask image editing model, wherein the primary prompt includes a reference image of the type of workpiece; wherein the mask image editing model is adapted to edit the mask image with reference to the reference image.
[0014] It can be seen that the mask image can be edited based on a single prompt, which improves the accuracy of the edited mask image.
[0015] In one embodiment, before inputting the mask image into the trained mask image editing model associated with the type, the method includes a training process of the mask image editing model associated with the type, the training process comprising:
[0016] Acquire a training sample, wherein the training sample includes a mask test image of the type and a primary prompt, wherein the primary prompt includes a reference image of the type of workpiece;
[0017] Determine the neural network model;
[0018] Inputting the training sample into the neural network model so that the neural network model edits the mask image to obtain an edited mask image;
[0019] Determining a loss function value of the neural network model based on a difference between the reference image and the edited mask image;
[0020] Configuring model parameters of the neural network model so that the loss function value is lower than a preset threshold;
[0021] The configured neural network model is determined as the trained mask image editing model associated with the type.
[0022] Therefore, the embodiments of the present invention propose a training process for a mask image editing model, which can train a mask image editing model associated with the type of workpiece.
[0023] In one embodiment, the mask image editing model is a transformer model comprising a plurality of blocks; before inputting the mask image into the trained mask image editing model associated with the type, the method comprises:
[0024] Lightweight processing is performed on each block in the trained mask image editing model associated with the type.
[0025] Therefore, the lightweight mask image editing model is easy to deploy and reduces resource pressure on the deployment side, making it particularly suitable for edge devices.
[0026] In one embodiment, performing lightweight processing on each block in the trained mask image editing model associated with the type includes:
[0027] Based on the identity transformation, convert the first batch normalization unit, the first adder connected to the first batch normalization unit, the 1x1 convolution kernel connected in parallel with the first batch normalization unit, and the short connection from the model input to the first adder in each block of the trained mask image editing model into a new first batch normalization unit;
[0028] Based on the identity transformation method, converting the pooling unit in each block of the trained mask image editing model, the second adder connected to the pooling unit, the 1x1 convolution kernel connected in parallel with the pooling unit, and the short connection from the first adder to the second adder into a new pooling unit;
[0029] Based on the identity transformation method, converting the drop unit, the third adder connected to the drop unit, the 1x1 convolution kernel connected in parallel with the drop unit, and the short connection from the input of the drop unit to the third adder in each block of the trained mask image editing model into a new drop unit;
[0030] Converting, based on an identity transformation, a second batch normalization unit, a fourth adder connected to the second batch normalization unit, a 1x1 convolution kernel connected in parallel to the second batch normalization unit, and a short connection from an input of the second batch normalization unit to the fourth adder in each block of the trained mask image editing model into a new second batch normalization unit;
[0031] Based on the identity transformation, convert the convolution unit, the fifth adder, the 1x1 convolution kernel connected in parallel with the convolution unit, and the short connection from the input of the convolution unit to the fifth adder in each block of the trained mask image editing model into a new convolution unit;
[0032] The new first batch normalization unit, the new pooling unit, the new dropout unit, the new second batch normalization unit, and the new convolution unit are sequentially connected to form each block of a lightweight, trained mask image editing model.
[0033] Therefore, by performing identity transformation on the components in the block and replacing them with lightweight components, the detection speed can be improved and the storage requirements can be reduced, which is particularly suitable for edge devices.
[0034] In one embodiment, it includes:
[0035] The lightweight, trained mask image editing model and the trained workpiece detection model are centrally deployed in an industrial edge device, which is suitable for picking up and / or placing the workpiece.
[0036] Therefore, based on the integrated collaboration of the mask image editing model and the workpiece detection model, workpieces can be picked up and / or placed quickly and accurately on industrial edge devices.
[0037] In one embodiment, inputting the mask image into the mask image editing model includes:
[0038] When it is determined that the shape parameters of the mask image do not fall within a predetermined reasonable range, the mask image is input into the mask image editing model; wherein the shape parameters include at least one of the following:
[0039] Aspect ratio; profile; coaxiality; cylindricity; flatness.
[0040] Therefore, whether to edit the mask image is determined by the shape parameters of the mask image, thereby reducing the editing workload.
[0041] In one embodiment, the editing mask image includes at least one of the following:
[0042] Filling in the missing areas in the mask image;
[0043] Remove redundant areas in the mask image.
[0044] It can be seen that the embodiments of the present invention can not only supplement the missing areas but also remove the redundant areas, thereby fully improving the quality of the mask image.
[0045] A mask image editing device, comprising:
[0046] an acquisition module, for acquiring an artifact image containing the artifact;
[0047] A first input module is configured to input the workpiece image into a trained workpiece detection model, wherein the workpiece detection model is adapted to output a mask image of the workpiece and a type of the workpiece based on the workpiece image;
[0048] A second input module, configured to input the mask image into a trained mask image editing model associated with the type, wherein the mask image editing model is suitable for editing the mask image of the workpiece of the type;
[0049] The receiving module is configured to receive the edited mask image from the mask image editing model.
[0050] Therefore, introducing AI capabilities into the editing process of mask images improves the quality of mask images.
[0051] An electronic device, comprising:
[0052] processor;
[0053] a memory for storing executable instructions of the processor;
[0054] The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement any one of the above mask image editing methods.
[0055] A computer-readable storage medium stores computer instructions, wherein the computer instructions, when executed by a processor, implement the mask image editing method as described in any one of the above items.
[0056] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the method for editing a mask image as described in any one of the above items is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, so that those skilled in the art will understand the above and other features and advantages of the present invention more clearly. In the accompanying drawings:
[0058] FIG1 is a schematic diagram of a mask image with a missing portion of the area.
[0059] FIG2 is a schematic diagram of a mask image with redundant regions.
[0060] FIG3 is an exemplary flowchart of a method for editing a mask image according to an embodiment of the present invention.
[0061] FIG4 is a schematic diagram of a processing procedure of a trained mask image editing model according to an embodiment of the present invention.
[0062] FIG5 is an exemplary structural diagram of a block according to an embodiment of the present invention.
[0063] FIG6 is an exemplary schematic diagram of an identity transformation according to an embodiment of the present invention.
[0064] FIG. 7 is an exemplary structural diagram of a lightweight block according to an embodiment of the present invention.
[0065] FIG. 8 is an exemplary schematic diagram of a pick-and-place workpiece according to an embodiment of the present invention.
[0066] FIG. 9 is an exemplary structural diagram of a mask image editing apparatus according to an embodiment of the present invention.
[0067] FIG. 10 is an exemplary structural diagram of an electronic device according to an embodiment of the present invention.
[0068] The accompanying drawings are numerals as follows: DETAILED DESCRIPTION
[0069] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail with reference to the following examples.
[0070] For the sake of brevity and intuitiveness in description, the solution of the present invention is explained below by describing several representative implementations. A large number of details in the implementations are only used to help understand the solution of the present invention. However, it is obvious that the technical solution of the present invention may not be limited to these details when implemented. In order to avoid unnecessarily obscuring the solution of the present invention, some implementations are not described in detail, but only a framework is given. Hereinafter, "including" means "including but not limited to", and "according to..." means "at least according to..., but not limited to only according to...". Due to the language habits of Chinese, when the number of a component is not specifically specified below, it means that the component can be one or more, or can be understood as at least one.
[0071] In the industrial field, picking and placing workpieces for self-assembly, printed circuit boards (PCBs) for automatic testing, and electronic waste sorting, etc., usually rely on AI-based workpiece detection. For flexible gripping solutions, they are mainly based on deep learning models running on industrial edge (IE) devices. Masking is a common operation in deep learning, which is equivalent to covering the original tensor with a mask to shield or select some specific elements, and is often used to construct tensor filters. When noise, clutter, and occlusion occur, the masked image in the detection result output by the workpiece detection model often misses some areas in the workpiece or unexpectedly has additional areas.
[0072] Figure 1 is a schematic diagram of a mask image with a missing region. As can be seen, due to occlusion between workpieces, the mask image of workpiece 10 is incomplete, lacking missing region 11. Figure 2 is a schematic diagram of a mask image with an excess region. As can be seen, due to occlusion between workpieces, the mask image of workpiece 13 has an excess region 13.
[0073] FIG3 is an exemplary flow chart of a method for editing a mask image according to an embodiment of the present invention. As shown in FIG3 , the method includes:
[0074] Step 101: Acquire a workpiece image containing a workpiece.
[0075] A workpiece refers to the object being machined during machining. It can be a single part or a combination of several parts fixed together. Workpieces in industrial scenes can be photographed to obtain a workpiece image containing the workpiece. Due to the complexity of industrial scenes, the workpiece in the workpiece image may be partially obscured by other workpieces (or the working environment). Furthermore, during the capture of workpiece images, environmental factors may introduce noise (e.g., excessively bright or dark light) or clutter, which can adversely affect the quality of the workpiece image.
[0076] Step 102: Input the workpiece image into a trained workpiece detection model. The workpiece detection model is adapted to output a mask image of the workpiece and the type of the workpiece based on the workpiece image.
[0077] Based on an input artifact image, an artifact detection model can output an artifact mask image and artifact type. For example, the artifact detection model can be implemented as a Mask R-CNN network model, a Unet+ network model, a UNet++ network model, or a Yolo-based network model. For example, Mask R-CNN uses the instance segmentation algorithm to output an artifact mask image and artifact type.
[0078] Step 103: Input the mask image into a trained mask image editing model associated with the type, where the mask image editing model is suitable for editing the mask image of the artifact of the type.
[0079] Here, the mask image is input into the trained mask image editing model associated with the type of the workpiece identified in step 102 , so that the mask image editing model edits the mask image of the type of workpiece.
[0080] In one embodiment, editing the mask image includes at least one of the following: supplementing a missing area in the mask image; and removing a redundant area in the mask image.
[0081] It can be seen that the embodiments of the present invention can not only supplement the missing areas but also remove the redundant areas, thereby fully improving the quality of the mask image.
[0082] Step 104: Receive the edited mask image from the mask image editing model.
[0083] In one embodiment, inputting a mask image into a trained mask image editing model can output an edited mask image. In other words, the trained mask image editing model can output the edited mask image with a zero-shot prompt.
[0084] Optionally, given the high demand for artifact detection accuracy in industrial scenarios, mask image editing can be implemented using a one-shot prompt. In one embodiment, inputting the mask image into a trained mask image editing model associated with a type includes: inputting the mask image and a one-shot prompt into the mask image editing model, where the one-shot prompt includes a reference image of the artifact type; wherein the mask image editing model is adapted to edit the mask image with reference to the reference image. Thus, editing the mask image based on the one-shot prompt improves the accuracy of the edited mask image.
[0085] In one embodiment, before inputting the mask image into the trained mask image editing model associated with the type, the method includes a training process of the mask image editing model associated with the type, the training process comprising:
[0086] Step (1): Obtain training samples, which include a masked test image of the type and a prompt, and the prompt includes a reference image of the type of artifact. For example, the reference image only contains the artifact of the type and does not contain other artifacts or scenes.
[0087] Step (2): Determine the neural network model. For example, a Transformer model or a Meta-Transformer model can be used.
[0088] Step (3): Input the training sample into the neural network model so that the neural network model edits the mask image to obtain the edited mask image.
[0089] Step (4): Based on the difference between the reference image and the edited mask image, determine the loss function value of the neural network model.
[0090] Step (5): Configure the model parameters of the neural network model so that the loss function value is lower than a preset threshold. For example, perform back propagation based on the loss function value to modify the weights and bias values in the neural network model from back to front.
[0091] Step (6): Determine the configured neural network model as a trained mask image editing model associated with the type.
[0092] Therefore, the embodiments of the present invention propose a training process for a mask image editing model, which can train a mask image editing model associated with the type of workpiece.
[0093] In one embodiment, the mask image editing model is a transformer model comprising multiple blocks. The method includes performing lightweight processing on each block in the trained mask image editing model before inputting the mask image into the trained mask image editing model associated with the type. Therefore, the lightweight mask image editing model facilitates deployment and reduces resource pressure on the deployment side, making it particularly suitable for edge devices.
[0094] In one embodiment, performing lightweight processing on each block in the trained mask image editing model includes:
[0095] (1) Based on the identity transformation method, the first batch normalization unit, the first adder connected to the first batch normalization unit, the 1x1 convolution kernel connected in parallel with the first batch normalization unit, and the short connection from the model input to the first adder in each block of the trained mask image editing model are converted into a new first batch normalization unit;
[0096] (2) Based on the identity transformation method, the pooling unit in each block of the trained mask image editing model, the second adder connected to the pooling unit, the 1x1 convolution kernel connected in parallel with the pooling unit, and the short connection from the first adder to the second adder are converted into a new pooling unit;
[0097] (3) Based on the identity transformation method, the discarding unit, the third adder connected to the discarding unit, the 1x1 convolution kernel connected in parallel with the discarding unit, and the short connection from the input of the discarding unit to the third adder in each block of the trained mask image editing model are converted into a new discarding unit;
[0098] (4) Based on the identity transformation method, the second batch normalization unit, the fourth adder connected to the second batch normalization unit, the 1x1 convolution kernel connected in parallel with the second batch normalization unit, and the short connection from the input of the second batch normalization unit to the fourth adder in each block of the trained mask image editing model are converted into a new second batch normalization unit;
[0099] (5) Based on the identity transformation method, the convolution unit, the fifth adder, the 1x1 convolution kernel connected in parallel with the convolution unit, and the short connection from the input of the convolution unit to the fifth adder in each block of the trained mask image editing model are converted into a new convolution unit;
[0100] (6) The new first batch normalization unit, the new pooling unit, the new dropout unit, the new second batch normalization unit, and the new convolution unit are sequentially connected to form each block of the lightweight, trained mask image editing model.
[0101] Therefore, by performing identity transformation on the components in the block and replacing them with lightweight components, the detection speed can be improved and the storage requirements can be reduced, which is particularly suitable for edge devices.
[0102] In one embodiment, the method further includes centrally deploying a lightweight, trained mask image editing model and a trained workpiece detection model on an industrial edge device suitable for picking and / or placing workpieces. Thus, based on the integrated collaboration of the mask image editing model and the workpiece detection model, workpieces can be quickly and accurately picked and / or placed on the industrial edge device.
[0103] For example, a lightweight, trained mask image editing model and a trained artifact detection model can be centrally deployed on various types of industrial edge devices, so that these industrial edge devices can pick and place artifacts (such as printed circuit boards (PCBs) or electronic waste), etc.
[0104] In one embodiment, inputting the mask image into the mask image editing model includes: inputting the mask image into the mask image editing model when it is determined that shape parameters of the mask image do not fall within a predetermined reasonable range; the shape parameters include at least one of the following: aspect ratio; contour; coaxiality; cylindricity; flatness, etc. Preferably, when it is determined that the shape parameters of the mask image fall within the predetermined reasonable range, the mask image is not input into the mask image editing model, and the process is exited. Therefore, determining whether the mask image needs to be edited based on the shape parameters of the mask image can reduce editing workload.
[0105] Example 1: Assuming the shape parameter is aspect ratio, and the predetermined reasonable range is [1 / 5 to 1 / 3]. If the aspect ratio of the mask image is found to be 1 / 2, it can be seen that this aspect ratio does not fall within the reasonable range. Therefore, the mask image quality is determined to be unsatisfactory. The mask image is input into the mask image editing model to perform editing on the mask image.
[0106] Example 2: Assuming the shape parameter is aspect ratio, and the predetermined reasonable range is [1 / 5 to 1 / 3]. If the aspect ratio of the mask image is found to be 1 / 4, it can be seen that this aspect ratio is within the reasonable range. Therefore, the mask image quality is determined to be qualified and the mask image is not input into the mask image editing model.
[0107] The following describes the embodiments of the present invention in detail by taking the mask image editing model specifically adopting the transformer model as an example.
[0108] Figure 4 is a schematic diagram illustrating the processing of a trained mask image editing model according to an embodiment of the present invention, wherein the mask image editing model employs a transformer model. The transformer model's backbone network 71 includes n blocks, namely blocks 80 through 8n, each having the same structure, where n is a positive integer.
[0109] Mask image 70 is input into backbone network 71 in the mask image editing model. After mask image 70 passes through n blocks, an edited, vectorized mask image 72 is obtained. Vectorized mask image 72 and primary hint 74 are input into mask image generator 73 in the mask image editing model. Mask image generator 73 refers to primary hint 74 and outputs a further edited mask image.
[0110] The structure of a block is described in detail below by taking block 80 as an example. Fig. 5 is an exemplary structural diagram of a block according to an embodiment of the present invention.
[0111] Block 80 includes: a first batch normalization (BN) unit 41, a first adder 42, a pooling unit 43, a first 1x1 convolution kernel 44, a second adder 45, a drop unit 46, a third adder 47, a second BN unit 48, a fourth adder 49, a convolution unit 50 (for example, a 3x3 convolution kernel), a second 1x1 convolution kernel 51, a fifth adder 52, a third 1x1 convolution kernel 53, a fourth 1x1 convolution kernel 54 and a fifth 1x1 convolution kernel 55.
[0112] The model input has a short connection to the input of the first adder 42. The output of the first adder 42 has a short connection to the input of the second adder 45. The output of the second adder 45 has a short connection to the input of the third adder 47. The output of the third adder 47 has a short connection to the input of the fourth adder 49. The output of the fourth adder 49 has a short connection to the input of the fifth adder 52. The first 1x1 convolution kernel 44 is connected in parallel with the pooling unit 43, the second 1x1 convolution kernel 51 is connected in parallel with the convolution unit 50, the third 1x1 convolution kernel 53 is connected in parallel with the first BN unit 41, the fourth 1x1 convolution kernel 54 is connected in parallel with the drop unit 46, and the fifth 1x1 convolution kernel 55 is connected in parallel with the second BN unit 48.
[0113] Based on the identity transformation, lightweighting is performed on the block 80. Fig. 7 is an exemplary structural diagram of a lightweight block according to an embodiment of the present invention. After lightweighting is performed on the block 80, a lightweight block 90 is obtained.
[0114] Specifically:
[0115] (1) Based on the identity transformation method, the first BN unit 41 in block 80, the first adder 42 connected to the first BN unit 41, the third 1x1 convolution kernel 53 connected in parallel with the first BN unit 41, and the short connection from the model input to the first adder 42 are converted into the first BN unit 60 in block 90.
[0116] (2) Based on the identity transformation method, the pooling unit 43 in block 80, the second adder 45 connected to the pooling unit 43, the first 1x1 convolution kernel 44 in parallel with the pooling unit 43, and the short connection from the first adder 43 to the second adder 43 are converted into the pooling unit 61 in block 90.
[0117] (3) Based on the identity transformation method, the discarding unit 46 in block 80, the third adder 47 connected to the discarding unit 46, the fourth 1x1 convolution kernel connected in parallel with the discarding unit 46, and the short connection from the input of the discarding unit 46 to the third adder 47 are converted into the discarding unit 62 in block 90.
[0118] (4) Based on the identity transformation method, the second BN unit 48 in block 80, the fourth adder 49 connected to the second BN unit 48, the fifth 1x1 convolution kernel 55 connected in parallel with the second BN unit 48, and the short connection from the input of the second BN unit 48 to the fourth adder 49 are converted into the second BN unit 63 in block 90.
[0119] (5) Based on the identity transformation method, the convolution unit 50, the fifth adder 52, the second 1x1 convolution kernel 51 in parallel with the convolution unit 50, and the short connection from the input of the convolution unit 50 to the fifth adder 52 in block 80 are converted into a new convolution unit 64.
[0120] FIG6 is an exemplary schematic diagram of an identity transformation according to an embodiment of the present invention. In FIG6 , the pooling unit 43 in block 80, the second adder 45 connected to the pooling unit 43, the first 1×1 convolution kernel 44 connected in parallel with the pooling unit 43, and the short connection from the first adder 43 to the second adder 43 are converted into a new pooling unit 61. Among them: the short connection matrix (shortcut), the pooling matrix (pool), the matrix of the first 1×1 convolution kernel (conv(1×1)), and the new pooling matrix (pool”) are shown in FIG6 .
[0121] As shown in FIG7 , the first BN unit 60 , the pooling unit 61 , the discarding unit 62 , the second BN unit 63 and the convolution unit 64 are sequentially connected to form a lightweight block 90 .
[0122] The above describes the lightweighting process for a single block, using block 80 as an example. Similarly, similar operations can be performed on the remaining blocks 81 through 8n in the backbone network to generate their own lightweight blocks. These lightweight blocks are then connected in sequence to form the backbone network of the lightweight mask image editing model.
[0123] FIG. 8 is an exemplary schematic diagram of a pick-and-place workpiece according to an embodiment of the present invention.
[0124] First, steps 201 and 202 are executed to generate a primary prompt library for storing primary prompts corresponding to each type of workpiece.
[0125] Step 101: Capture a reference image of each type of workpiece, where the reference image only contains the workpiece of that type without any occlusion.
[0126] Step 102: Save the reference image of each type of workpiece in a one-time prompt library.
[0127] Then, steps 301 to 305 are performed at the industrial edge device.
[0128] Step 301: Capture a workpiece image containing a workpiece. Due to the complexity of industrial scenes, a workpiece may be partially blocked by other workpieces (or working environments), so the workpiece in the workpiece image may be partially blocked.
[0129] Step 302: The workpiece image is input into a trained workpiece detection model in the industrial edge device. The workpiece detection model outputs a workpiece mask image and workpiece type based on the workpiece image. Here, because the workpiece in the workpiece image is partially occluded, the output may include a workpiece mask 30 with excess areas or a workpiece mask 31 with missing areas.
[0130] Step 303: Detecting the quality of the mask image, specifically including: comparing the shape parameters of the mask image with a predetermined reasonable range.
[0131] Step 304: Determine whether the quality is acceptable. If the shape parameters of the mask image fall within a reasonable range, the mask image is deemed acceptable (corresponding to the "Y branch") and the process proceeds to step 305. Otherwise, the mask image is deemed unacceptable (corresponding to the "N branch"), and the mask image is input into the trained mask image editing model 203, exiting the process. The mask image editing model 203 is suitable for editing mask images for this type of workpiece.
[0132] Step 305 : Utilize the mask image to perform pick and place of the workpiece to achieve accurate pick and place 34 .
[0133] The trained mask image editing model 203 is also deployed in the industrial edge device. After receiving the mask image, the trained mask image editing model 203 retrieves a reference image of the workpiece type from the primary prompt library. Based on the reference image, the mask image editing model 203 outputs an edited mask image 204. As can be seen, the edited mask image 204 may include: a workpiece mask 32 with excess areas removed and a workpiece mask 33 with missing areas added. The edited mask image 204 can then be used to perform workpiece picking and placement, achieving the same precise picking and placement 34 as a qualified mask image.
[0134] FIG9 is an exemplary structural diagram of a mask image editing apparatus according to an embodiment of the present invention. As shown in FIG9 , the mask image editing apparatus 700 includes:
[0135] An acquisition module 701 is used to acquire an artifact image containing an artifact; a first input module 702 is used to input the artifact image into a trained artifact detection model, which is adapted to output a mask image of the artifact and the type of the artifact based on the artifact image; a second input module 703 is used to input the mask image into a trained mask image editing model associated with the type, which is adapted to edit the mask image of the artifact type; and a receiving module 704 is used to receive the edited mask image from the mask image editing model.
[0136] In one embodiment, the second input module 703 is used to input the mask image and a prompt into the mask image editing model, wherein the prompt includes a reference image of the workpiece type; wherein the mask image editing model is suitable for editing the mask image with reference to the reference image.
[0137] In one embodiment, the device 700 includes: a training module 705, which is used to perform a training process of the mask image editing model associated with the type before the second input module 703 inputs the mask image into the trained mask image editing model associated with the type, the training process including: obtaining training samples, the training samples including a mask test image of the type and a prompt, the prompt including a reference image of the workpiece of the type; determining a neural network model; inputting the training samples into the neural network model so that the mask image is edited by the neural network model to obtain an edited mask image; determining the loss function value of the neural network model based on the difference between the reference image and the edited mask image; configuring the model parameters of the neural network model so that the loss function value is lower than a preset threshold; and determining the configured neural network model as the trained mask image editing model associated with the type.
[0138] In one embodiment, the mask image editing model is a transformer model comprising multiple blocks; the device 700 includes: a lightweight module 706, which is used to perform lightweight processing on each block in the trained mask image editing model associated with the type before the second input module 703 inputs the mask image into the trained mask image editing model associated with the type.
[0139] In one embodiment, the lightweight module 706 is used to convert the first BN unit in each block of the trained mask image editing model, the first adder connected to the first BN unit, the 1x1 convolution kernel connected in parallel with the first BN unit, and the short connection from the model input to the first adder into a new first BN unit based on an identity transformation method; based on an identity transformation method, convert the pooling unit in each block of the trained mask image editing model, the second adder connected to the pooling unit, the 1x1 convolution kernel connected in parallel with the pooling unit, and the short connection from the first adder to the second adder into a new pooling unit; based on an identity transformation method, convert the discarding unit in each block of the trained mask image editing model, the third adder connected to the discarding unit, the 1x1 convolution kernel connected in parallel with the discarding unit, and the short connection from the input of the discarding unit to the third adder into a new discarding unit; based on an identity transformation method, convert the second BN unit in each block of the trained mask image editing model, The fourth adder connected to the second BN unit, the 1x1 convolution kernel connected in parallel with the second BN unit, and the short connection from the input of the second BN unit to the fourth adder are converted into a new second BN unit; based on the identity transformation method, the convolution unit, the fifth adder, the 1x1 convolution kernel connected in parallel with the convolution unit, and the short connection from the input of the convolution unit to the fifth adder in each block of the trained mask image editing model are converted into a new convolution unit; the new first BN unit, the new pooling unit, the new discarding unit, the new second BN unit and the new convolution unit are connected in sequence to form each block of the lightweight, trained mask image editing model.
[0140] In one embodiment, a lightweight, trained mask image editing model and a trained workpiece detection model are centrally deployed in an industrial edge device suitable for picking and / or placing workpieces.
[0141] In one embodiment, the second input module 703 is used to input the mask image into the mask image editing model when it is determined that the shape parameters of the mask image do not fall within a predetermined reasonable range; wherein the shape parameters include at least one of the following: aspect ratio; contour; coaxiality; cylindricity; and flatness.
[0142] In one embodiment, editing the mask image includes at least one of the following: supplementing a missing area in the mask image; and removing a redundant area in the mask image.
[0143] The embodiments of the present invention propose solutions that run on industrial edge devices when there is noise, clutter, and partial occlusion in the factory, which can promote the upgrade of traditional factories to digital factories and assist in intelligent manufacturing. Industrial edge devices and neural processing units (for example, NPU modules equipped with Myriad MX vision processing unit chips with AI functions) are entering the automation industry through full integration. The embodiments of the present invention can work on a variety of industrial edge devices, so that these industrial edge devices can handle industrial scenes with noise, clutter, and partial occlusion without the need for resource-intensive programming.
[0144] The embodiment of the present invention also proposes an electronic device with a processor-memory architecture. Figure 10 is a structural diagram of an electronic device according to an embodiment of the present invention. As shown in Figure 10, the electronic device 800 includes a processor 801, a memory 802, and a computer program stored on the memory 802 and executable on the processor 801. When the computer program is executed by the processor 801, any of the above mask image editing methods is implemented. Among them, the memory 802 can be specifically implemented as a variety of storage media such as an electrically erasable programmable read-only memory (EEPROM), a flash memory (Flash memory), and a programmable read-only memory (PROM). The processor 801 can be implemented to include one or more central processing units or one or more field programmable gate arrays, wherein the field programmable gate array integrates one or more central processing unit cores. Specifically, the central processing unit or the central processing unit core can be implemented as a CPU, an MCU, or a DSP, etc.
[0145] It should be noted that not all steps and modules in the above processes and structure diagrams are required, and certain steps or modules can be omitted based on actual needs. The execution order of the steps is not fixed and can be adjusted as needed. The division of the modules is merely for the convenience of describing the functional division adopted. In actual implementation, a module can be implemented by multiple modules, and the functions of multiple modules can be implemented by the same module. These modules can be located in the same device or in different devices.
[0146] The hardware modules in each embodiment can be implemented mechanically or electronically. For example, a hardware module may include a specially designed permanent circuit or logic device (such as a dedicated processor, such as an FPGA or ASIC) for performing a specific operation. The hardware module may also include a programmable logic device or circuit (such as a general-purpose processor or other programmable processor) temporarily configured by software to perform a specific operation. As for whether to implement the hardware module mechanically, or using a dedicated permanent circuit, or using a temporarily configured circuit (such as configured by software), it can be decided based on cost and time considerations.
[0147] The above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for editing a masked image, characterized in that, comprising: obtaining (101) a workpiece image including a workpiece; inputting (102) the workpiece image into a trained workpiece detection model, the workpiece detection model being adapted to output a masked image of the workpiece and the type of the workpiece based on the workpiece image; inputting (103) the masked image into a trained masked image editing model associated with the type, the masked image editing model being adapted to edit the masked image of the workpiece of the type; receiving (104) the edited masked image from the masked image editing model.
2. The method according to claim 1, characterized in that, the inputting (103) the masked image into a trained masked image editing model associated with the type comprises: inputting the masked image and a first prompt into the masked image editing model, the first prompt including a reference image of the workpiece of the type; wherein the masked image editing model is adapted to edit the masked image with reference to the reference image.
3. The method according to claim 1, characterized in that, before inputting (103) the masked image into a trained masked image editing model associated with the type, the method includes a training process of the masked image editing model associated with the type, the training process comprising: obtaining training samples, the training samples including masked test images of the type and a first prompt, the first prompt including a reference image of the workpiece of the type; determining a neural network model; inputting the training samples into the neural network model to edit the masked image by the neural network model to obtain an edited masked image; determining a loss function value of the neural network model based on a difference between the reference image and the edited masked image; configuring model parameters of the neural network model to make the loss function value lower than a preset threshold; determining the configured neural network model as the trained masked image editing model associated with the type.
4. The method according to claim 1, characterized in that, the masked image editing model is a transformer model including multiple blocks; before inputting (103) the masked image into a trained masked image editing model associated with the type, the method includes: performing a lightweight processing on each block in the trained masked image editing model associated with the type.
5. The method according to claim 4, characterized in that, the performing a lightweight processing on each block in the trained masked image editing model associated with the type includes: based on an identity transformation method, converting a first batch normalization unit, a first adder connected to the first batch normalization unit, a 1x1 convolution kernel connected in parallel with the first batch normalization unit, and a short connection from the model input to the first adder in each block of the trained masked image editing model into a new first batch normalization unit; Based on the identity transformation method, convert the pooling unit in each block of the trained mask image editing model, the second adder connected to the pooling unit, the 1x1 convolutional kernel connected in parallel with the pooling unit, and the short connection from the first adder to the second adder into a new pooling unit; Based on the identity transformation method, convert the dropout unit in each block of the trained mask image editing model, the third adder connected to the dropout unit, the 1x1 convolutional kernel connected in parallel with the dropout unit, and the short connection from the input of the dropout unit to the third adder into a new dropout unit; Based on the identity transformation method, convert the second batch normalization unit in each block of the trained mask image editing model, the fourth adder connected to the second batch normalization unit, the 1x1 convolutional kernel connected in parallel with the second batch normalization unit, and the short connection from the input of the second batch normalization unit to the fourth adder into a new second batch normalization unit; Based on the identity transformation method, convert the convolutional unit in each block of the trained mask image editing model, the fifth adder, the 1x1 convolutional kernel connected in parallel with the convolutional unit, and the short connection from the input of the convolutional unit to the fifth adder into a new convolutional unit; Connect the new first batch normalization unit, the new pooling unit, the new dropout unit, the new second batch normalization unit, and the new convolutional unit in sequence to form each block of the lightweight and trained mask image editing model.
6. The method according to claim 5, wherein, it includes: Deploy the lightweight and trained mask image editing model and the trained workpiece detection model in an industrial edge device, and the industrial edge device is suitable for picking up and / or placing the workpiece.
7. The method according to any one of claims 1-6, wherein, the inputting (103) the mask image into the mask image editing model includes: When it is determined that the shape parameter of the mask image does not belong to a predetermined reasonable range, input the mask image into the mask image editing model; wherein the shape parameter includes at least one of the following: Aspect ratio; contour; coaxiality; cylindricity; flatness.
8. The method according to any one of claims 1-6, wherein, the editing the mask image includes at least one of the following: Supplementing the missing area in the mask image; Removing the redundant area in the mask image.
9. An editing device for a mask image, wherein, it includes: An acquisition module (701) for acquiring a workpiece image containing a workpiece; A first input module (702) for inputting the workpiece image into a trained workpiece detection model, and the workpiece detection model is suitable for outputting a mask image of the workpiece and the type of the workpiece based on the workpiece image; A second input module (703) for inputting the mask image into a trained mask image editing model associated with the type, and the mask image editing model is suitable for editing the mask image of the workpiece of the type; A receiving module (704) for receiving the edited masked image from the masked image editing model.
10. An electronic device, characterized in that, comprising: a processor (801); a memory (802) for storing executable instructions of the processor (801); the processor (801) for reading the executable instructions from the memory (802) and executing the executable instructions to implement the method for editing a masked image according to any one of claims 1-8.
11. A computer-readable storage medium having computer instructions stored thereon, characterized in that, when the computer instructions are executed by a processor, the method for editing a masked image according to any one of claims 1-8 is implemented.
12. A computer program product, characterized in that, comprising a computer program which, when executed by a processor, implements the method for editing a masked image according to any one of claims 1-8.
Citation Information
Patent Citations
Method for semantic segmentation of image with occlusion area
CN110992367A
Image segmentation result optimization method and device, intelligent terminal and storage medium
CN111192252A
High-precision image semantic segmentation method for industrial part measurement
CN111488882A
Image processing method and device
CN112967199A
Metal part rapid segmentation method based on deep learning
CN114494272A