Target area identification method, model training method, target area positioning method, medium, equipment and device

By combining CNN and Transformer networks to extract features and perform edge information enhancement processing, the complexity and recognition accuracy issues of image recognition technology when switching product models are resolved, efficient target area recognition and positioning are achieved, and target images of different shapes and appearances are adapted, thereby improving production efficiency.

CN120599280APending Publication Date: 2025-09-05JOMOO KITCHEN & BATHROOM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510590483.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing image recognition technology requires replacing templates when switching between different product models, which increases the changeover time and product recognition complexity. It also has high requirements for image edge contrast and is easily affected by noise and image transformation, resulting in reduced recognition accuracy.

Method used

The CNN network is used to extract local features and the Transformer network is used to extract global features. Feature fusion and edge information are enhanced, and the weighted boundary loss and Euclidean error center loss functions are combined to optimize model training to achieve accurate recognition and positioning of the target area.

Benefits of technology

The accuracy and adaptability of target area recognition are improved, and it can quickly adapt to target images of different shapes and appearances, reducing changeover costs and improving production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599280A_ABST
    Figure CN120599280A_ABST
Patent Text Reader

Abstract

A target area identification, model training and positioning method, medium, device and apparatus wherein the target area identification method comprises: extracting local features from an original image containing a target area by using a CNN network, and extracting global features from the original image by using a Transform network; fusing the local features with the global features to obtain fused features; carrying out edge information enhancement processing on the fusion features; decoding the fusion features after enhancement processing to obtain a prediction map of the target area; the target area identification precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This article relates to image recognition technology, particularly a method, medium, equipment and device for target area recognition, model training and positioning. Background Art

[0002] Image recognition technology is widely used in the field of automation. For example, on production lines, image recognition technology can help robotic arms locate and sort target objects.

[0003] Edge extraction and template matching are commonly used image recognition methods for industrial robotic arms. They identify the target object by comparing an object template with an image of the real object. Because different products correspond to different templates, template changes are required when switching between product models, increasing changeover time. Adding new products also requires creating corresponding templates, further complicating product recognition. Summary of the Invention

[0004] The embodiments of the present application provide a method, medium, equipment and device for target area identification, model training and positioning, which can improve the accuracy of target area identification.

[0005] The target area recognition method provided in the embodiment of the present application includes:

[0006] Using a CNN network to extract local features from an original image containing a target area, and using a Transformer network to extract global features from the original image;

[0007] Fusing the local features with the global features to obtain fused features;

[0008] Performing edge information enhancement processing on the fusion features;

[0009] The enhanced fusion features are decoded to obtain a prediction map of the target area.

[0010] The training method of the target region recognition model in an image provided by the embodiment of the present application includes:

[0011] Identifying target area training samples using a target area recognition model to obtain a prediction map of the target area training samples;

[0012] Determining a loss value based on the target area training sample and the prediction graph of the target area training sample, and a preset loss function;

[0013] Updating parameters of the target area recognition model based on the loss value;

[0014] The target area recognition model includes:

[0015] The CNN network module is configured to extract local features from the original image containing the target area;

[0016] A Transformer network module configured to extract global features from the original image;

[0017] a feature fusion module configured to fuse the local features with the global features to obtain fused features;

[0018] The edge information enhancement module is configured to enhance the edge information of the fusion feature:

[0019] The decoding module is configured to decode the fused features after the enhanced processing to obtain a prediction map of the target area.

[0020] The visual positioning method provided in the embodiment of the present application includes:

[0021] Obtaining a prediction map of the target area from an original image containing the target area based on the target area recognition method described in the aforementioned embodiment;

[0022] A grasping center and an offset angle of visual positioning are determined based on the predicted image of the target area.

[0023] The non-volatile computer-readable storage medium provided in the embodiments of the present application stores one or more program instructions, and the one or more program instructions can be executed by one or more processors to implement the target area recognition method as described in the aforementioned embodiments, or the training method of the target area recognition model in the image as described in the aforementioned embodiments, or the visual positioning method as described in the aforementioned embodiments.

[0024] The control device provided in the embodiment of the present application includes:

[0025] a memory configured to store computer program instructions executable on the processor;

[0026] A processor is configured to execute the computer program instructions to implement the target area recognition method as described in the aforementioned embodiment, or the training method of the target area recognition model in the image as described in the aforementioned embodiment, or the visual positioning method as described in the aforementioned embodiment.

[0027] The control device provided in the embodiment of the present application includes:

[0028] robotic arm;

[0029] A visual sensor provided on the robotic arm;

[0030] A control device as described in the above embodiment that transmits information with the visual sensor.

[0031] The embodiment of the present application integrates the advantages of the CNN network and the Transformer network, which can comprehensively capture the local and global features of the target area, which is conducive to improving the accuracy of target area recognition; by strengthening the processing of edge information, the contrast of the edge of the target area is improved, and the target area and the background can be distinguished more accurately, further improving the accuracy of target area recognition; it has stronger versatility and adaptability, and can be applied to target images of different shapes, sizes and appearances, can achieve rapid changeover and adapt to the needs of new products, improve production efficiency and reduce costs.

[0032] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. Other advantages of the present application can be realized and obtained by the solutions described in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings are used to provide an understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.

[0034] Figure 1 A flow chart of a method for identifying a target area in an image provided by an embodiment of the present application;

[0035] Figure 2 A schematic diagram of a network structure for performing image recognition on a toilet provided in an embodiment of the present application;

[0036] Figure 3 A schematic diagram of enhancing edge information of fused features provided in an embodiment of the present application;

[0037] Figure 4 Schematic diagram of weighted features obtained by using Sigmoid function and ReLU function as activation functions respectively provided in the embodiment of the present application;

[0038] Figure 5 This is a gradient thermal distribution diagram of the fusion feature after enhanced processing provided in an embodiment of the present application;

[0039] Figure 6 A flowchart of a method for training a target region recognition model in an image provided in an embodiment of the present application;

[0040] Figure 7 A flow chart of a visual positioning method provided in an embodiment of the present application;

[0041] Figure 8 A schematic diagram of determining an offset angle when positioning a toilet provided in an embodiment of the present application;

[0042] Figure 9 A module diagram of a control device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] This application describes multiple embodiments, but this description is exemplary rather than restrictive, and it is obvious to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described in this application. Although many possible feature combinations are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with any other feature or element in any other embodiment, or may replace any other feature or element in any other embodiment.

[0044] The present application includes and contemplates combinations of features and elements known to those of ordinary skill in the art. The embodiments, features, and elements disclosed in this application may also be combined with any conventional features or elements to form a unique inventive solution. Any features or elements of any embodiment may also be combined with features or elements from other inventive solutions to form another unique inventive solution. Therefore, it should be understood that any feature shown and / or discussed in this application may be implemented individually or in any appropriate combination. Therefore, except for the limitations made according to the appended claims and their equivalents, the embodiments are not subject to other limitations. In addition, various modifications and changes may be made within the scope of protection of the appended claims.

[0045] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be understood by those skilled in the art, other orders of steps are also possible. Therefore, the specific order of the steps set forth in the specification should not be interpreted as a limitation to the claims. In addition, the claims for the method and / or process should not be limited to performing their steps in the order written, and those skilled in the art can readily understand that these orders can be changed and still remain within the spirit and scope of the embodiments of the present application.

[0046] This study found that the edge extraction and template matching methods still have the following deficiencies:

[0047] It has high requirements for the contrast of the edge of the image. When the edge contrast is small or there is component noise, it is easy to cause inaccurate edge extraction, which in turn affects the image recognition accuracy. It is sensitive to image transformations. When the image undergoes changes such as rotation, scaling or occlusion, it is easy to cause template matching failure.

[0048] The present invention provides a method for identifying a target area in an image. Figure 1 As shown, the identification method includes:

[0049] Step S101 uses a CNN network to extract local features from an original image containing a target area, and uses a Transformer network to extract global features from the original image;

[0050] Among them, the local features of an image refer to the features of a specific area or local position in the image, which are usually related to the details and local structure of the image. Local features can be used to describe image edges, corners, textures and other details; the global features of an image refer to the features that describe the entire image, which are usually related to the overall structure and content of the image. Global features can be used to describe the overall layout, color distribution, overall texture and other information of the image;

[0051] The core operation of the CNN network is the convolution operation. When each convolution kernel slides on the input original image, it only calculates with the pixels in the local area of ​​the original image, so that the CNN network can capture the local subtle changes and details in the original image. Therefore, using the CNN network to obtain the local features of the original image has an advantage;

[0052] The Transformer network is able to capture the dependencies between all positions in the input sequence through the self-attention mechanism. This means that no matter how far apart two elements are in the sequence, the Transformer can directly calculate the relationship between them, thereby better understanding the global context. Therefore, using the Transformer network to obtain the global features of the original image has the advantage;

[0053] Step S102: fusing the local features with the global features to obtain fused features;

[0054] Step S103: performing edge information enhancement processing on the fusion features;

[0055] Step S104 decodes the enhanced fusion features to obtain a prediction map of the target area.

[0056] The embodiment of the present application integrates the advantages of the CNN network and the Transformer network, which can more comprehensively capture the local and global features of the target area, which is conducive to improving the accuracy of target area recognition; by strengthening the processing of edge information, the contrast of the edge of the target area is improved, and the target area and background can be more accurately distinguished, further improving the accuracy of target area recognition. Compared with the traditional template matching method, the feature-based image recognition method of the embodiment of the present application has stronger versatility and adaptability, and can be applied to target images of different shapes, sizes and appearances, while achieving rapid changeover and adapting to the needs of new products, improving production efficiency and reducing costs.

[0057] In an exemplary embodiment, extracting local features from an original image containing a target area using a CNN network and extracting global features from the original image using a Transformer network include:

[0058] A CNN network is used to extract multi-layer local features from the original image, and a Transformer network is used to extract multi-layer global features from the original image; the number of layers of the multi-layer local features and the number of layers of the multi-layer global features may be the same or different.

[0059] For example, Resnet18 is selected as the feature extractor of CNN network, and 5 layers of features are extracted and recorded as {f c1 ,f c2 ,f c3 ,f c4 ,f c5}, the 5-layer feature can cover from low-level edge information to high-level semantic information, which can meet the needs of different visual tasks; T2T-ViT-14 is selected as the feature extractor of the Transformer network, and 4 layers of features are extracted, which are respectively denoted as {f t1 ,f t2 ,f t3 ,f t4}.

[0060] Table 1 shows the mean absolute error (MAE) of target image recognition using different feature extractors for CNN and Transformer networks. x 、MAE y and MAE A They represent the mean absolute error in the x and y directions and the mean absolute error in the rotation angle, respectively. As can be seen from Table 1, the CNN network uses Resnet18 as the feature extractor, and the Transformer network uses T2T-ViT-14 as the feature extractor, which can reduce the recognition error of the target image.

[0061] CNN network Transformer Network <![CDATA[MAE x ]]> <![CDATA[MAE y ]]> <![CDATA[MAE A ]]> VGG16 T2T-ViT-14 3.07 6.45 3.34° Densenet121 T2T-ViT-14 6.58 9.3 7.03° Resnet18 Swin B 3.14 6.33 3.46° Resnet18 - 4.66 8.65 4.13° - T2T-ViT-14 4.78 8.55 4.73° Resnet18 T2T-ViT-14 2.78 5.75 2.43°

[0062] Table 1

[0063] In an exemplary embodiment, fusing the local features with the global features includes:

[0064] All or part of the local features are selected from the multi-layer local features, and all or part of the global features are selected from the multi-layer global features; after adjusting the selected local features and the selected global features to the same dimension, they are fused in each dimension to obtain multi-layer fused features.

[0065] For example, CNN network is used to obtain 5 layers of local features {f c1 ,f c2 ,f c3 ,f c4 ,f c5}, using the Transformer network to obtain 4 layers of global features {f t1 ,f t2 ,f t3 ,f t4}; Considering the local feature f c1 and the global feature f t1 The semantics is low, the understanding of the image content is relatively superficial, and it may contain more noise. Therefore, from the 5-layer local feature {f c1 ,f c2 ,f c3 ,f c4 ,f c5}Select the 2nd to 5th layer features {f c2 ,f c3 ,f c4 ,f c5}, from the 4-layer global feature {f t1 ,f t2 ,f t3 ,f t4}Select the 2nd to 4th layer features {f t2 ,f t3 ,f t4}, as the feature involved in the fusion. In order to adjust the selected local features and global features to the same dimension, f t4 The feature is downsampled twice to get f t4 ', that is, the final participants in the fusion are {f c2 ,f c3 ,f c4 ,f c5} and {f t2 ,f t3 ,f t4 ,f t4 '}.

[0066] For example, a mutual consensus module (MCM) and a consensus complementary module (CCM) can be introduced in the feature fusion stage. CCM generates complementary features by calculating the complementarity between different features, and fuses the generated complementary features with the original features to form a richer feature representation. The fusion method can be a simple weighted sum or a more complex nonlinear fusion. MCM fuses the high-level features obtained from the CNN network and the Transformer network by learning or dynamically adjusting the weights to obtain global consensus information. The global consensus information can also be used as the input of the consensus complementary module (CCM) to guide more fine-grained feature fusion.

[0067] Figure 2 A network structure diagram for toilet image recognition is shown. The diagram shows an example of using a CNN network for local feature extraction, a Transformer network for global feature extraction, and the introduction of MCM and CCM modules to fuse local and global features.

[0068] In an exemplary embodiment, when the fused feature is a multi-layer fused feature, performing edge information enhancement processing on the fused feature includes:

[0069] For each layer of fusion features except the last layer of fusion features, perform the following operations:

[0070] The next layer of fusion features of the fusion features of the layer is sequentially upsampled and activated, and the processing results are used as the weights of the fusion features of the layer; or, the next layer of fusion features of the fusion features of the layer is sequentially upsampled, activated, and color-inverted, and the processing results are used as the weights of the fusion features of the layer; exemplarily, the activation function corresponding to the activation is a ReLU function;

[0071] The enhancement process is performed on the fusion features of the layer based on the weights.

[0072] In an exemplary embodiment, the step of performing the enhancement processing on the layer fusion feature based on the weight includes:

[0073] Using the weights to weight the fusion features of the layer;

[0074] The fusion features of this layer are superimposed on the weighted fusion features of this layer at the pixel level, and the superposition result is used as the enhanced fusion feature corresponding to the fusion features of this layer.

[0075] Figure 3 A schematic diagram of enhancing edge information of fusion features is given. In the figure, f fusion,i Represents the fusion features of the i-th layer.

[0076] First, the deep feature f fusion,i+1 Perform double upsampling to ensure that it is consistent with the shallow feature f fusion,i The deep and shallow features of an image typically have different dimensions, primarily referring to the spatial resolution (height and width) and number of channels (depth) of the features. Shallow features typically have larger spatial dimensions (height and width), close to the original image size, while deep features typically have smaller dimensions. Upsampling refers to increasing the resolution of features; a common upsampling method is bilinear interpolation. This operation increases image resolution, making the image clearer and more detailed; it also scales low-resolution deep features to the same size as shallow features for subsequent processing.

[0077] Next, the ReLU function is used to activate the deep features output after upsampling. The ReLU function can better highlight the feature weight area, making the black and white boundaries more distinct. It can also darken the background area, highlighting only the boundary segmentation area, thereby reducing background interference. The selection of the ReLU function facilitates the subsequent edge feature extraction and segmentation process.

[0078] Then, the output of the ReLU function is color-inverted. By color-inverting, the black in the feature map can be inverted to white, and the white to black. The color-inverted feature map is used as the weight matrix weight i+1 , weight i+1 It can be calculated according to formula (1):

[0079] weight i+1 =Reverse(ReLU(Upsample2(f fusion,i+1 ))) (1)

[0080] Upsample2 represents double upsampling based on the bilinear interpolation method, ReLU() represents the ReLU function, and Reverse() represents a function for reversing color.

[0081] The Reverse() function is shown in formula (2):

[0082] weight i+1,ij =Reverse(x ij )=max(X)-min(X)-x ij(2)

[0083] weight i+1,ij Represents the elements in the weight matrix.

[0084] The weight matrix weight i+1 With shallow features f fusion,i Perform element-wise multiplication to obtain a weighted feature map. Figure 4 The weighted feature maps obtained by using the Sigmoid function and the ReLU function as activation functions are given respectively. Among them, A is the original image of the toilet, B is the weighted feature map using the Sigmoid function as the activation function, and C is the weighted feature map using the ReLU function as the activation function. Figure 4 It can be seen that in the weighted feature map using the ReLU function as the activation function, the black and white boundaries are more obvious, and the white is completely concentrated at the boundaries, which can effectively improve the accuracy of segmentation at the boundaries.

[0085] Finally, f fusion,i The result of pixel-level superposition with the weighted feature map is used as the fusion feature after enhanced processing. The pixel-level superposition can refer to formula (3):

[0086]

[0087] Figure 5 The gradient heat map of the enhanced fusion features is shown. It can be observed that the heat values ​​of the enhanced fusion features (the closer to red, the higher the heat value) are concentrated in the center, while the edges are blue, and the heat values ​​disperse from the center of the segmented area to the edge. This shows that the enhanced fusion features can better guide the model to focus on the details of the segmentation edge, helping to improve the model's prediction accuracy for the edge.

[0088] The solution described in the embodiments of the present application can effectively enhance the features of the segmentation boundary, better capture the boundary information of the target object, and improve the accuracy and robustness of the segmentation.

[0089] In an exemplary embodiment, step S104 decodes the enhanced fused features to obtain a prediction map of the target area, including: decoding the enhanced fused features using a Transformer-based consistent progressive decoder (DCPG).

[0090] by Figure 2 As an example of the network structure diagram shown in the figure, assuming that the local feature {f c2 ,f c3 ,f c4 ,f c5} and global features {f t2 ,ft3 ,f t4 ,f t4 '}After fusion, the fusion feature {f fusion2 ,f fusion3 ,f fusion4 ,f fusion5}, for the fusion feature {f fusion2 ,f fusion3 ,f fusion4 ,f fusion5} is enhanced to obtain the feature {f enhance2 ,f enhance3 ,f enhance4}; Replace {f enhance2 ,f enhance3 ,f enhance4} and f fusion5 Connect and decode with DCPGg to obtain a one-dimensional output sequence; combine the one-dimensional output sequence with the feature f fusion5 The first two sequences are connected and used as input to the next DCPG module. Similarly, DCPG5, DCPG4, DCPG3, and DCPG2 are decoded sequentially, ultimately obtaining the one-dimensional output sequence of DCPG2, which is then reshaped into a two-dimensional prediction graph. This decoding structure can enhance the consistency of features within an image and achieve image saliency prediction.

[0091] The present application also provides a method for training a target region recognition model in an image. Figure 6 As shown, the training method includes:

[0092] Step S601 uses the target region recognition model to identify the target region training sample to obtain a prediction map of the target region training sample;

[0093] Step S602 determines a loss value based on the target area training sample, the prediction graph of the target area training sample, and a preset loss function;

[0094] Step S603 updates the parameters of the target area recognition model based on the loss value;

[0095] The target area recognition model includes:

[0096] The CNN network module is configured to extract local features from the original image containing the target area;

[0097] A Transformer network module configured to extract global features from the original image;

[0098] a feature fusion module configured to fuse the local features with the global features to obtain fused features;

[0099] The edge information enhancement module is configured to enhance the edge information of the fusion feature:

[0100] The decoding module is configured to decode the fused features after the enhanced processing to obtain a prediction map of the target area.

[0101] In an exemplary embodiment, the preset loss function includes: a weighted margin loss function and a Euclidean error center loss function.

[0102] In the robotic arm visual positioning task, in order to achieve accurate grasping center and rotation angle calculation, it is necessary to accurately segment the image edges during image recognition. To this end, the weighted boundary loss (wBoundary Loss) function can be used to perform the loss between the predicted image and the real image. Different from other boundary loss functions, the weighted boundary loss function assigns a weight α to the corresponding pixel based on the importance of each pixel, which can better reflect the importance difference between pixels. For example, α is regarded as an indicator to measure the importance of pixels, where more important pixels correspond to larger α values, and less important pixels correspond to smaller α values. The method for determining the α value can refer to formula (4):

[0103]

[0104] In formula (4), A ij represents the area around the point with coordinates (i, j), Represents the true value of the point with coordinates (m, n) in this area, It represents the true value of the point with coordinates (i, j). Represents the boundary weight of the (i, j)th point. For all pixels, if The larger it is, the greater the difference between the pixel at (i, j) and the surrounding environment.

[0105] For example, the weighted boundary loss function calculated based on the k-th layer features is recorded as Loss_B wMSE,k , as shown in formula (5):

[0106]

[0107] In formula (5), H and W represent the height and width of the boundary image respectively, Gt_bound ij Represents the true value of point (i, j) in the boundary graph, Pred_bound ij represents the predicted value of point (i, j) in the boundary graph, represents the boundary weight of the (i, j)th point, and γ is a hyperparameter used to balance the influence of boundary loss in the overall loss function. The weighted boundary loss function can better guide the model to learn image edge information, thereby improving the accuracy and stability of the robot arm visual positioning task.

[0108] In the robot visual positioning task, the introduction of the Euclidean error center loss (MSECenter) function is also crucial to improving positioning accuracy. The accuracy of the grasping center directly affects the accuracy and stability of the robot when performing the grasping task. Therefore, it is necessary to use the grasping center as an output to directly participate in the calculation of the loss function. The Euclidean error center loss function uses the Euclidean distance loss to calculate the error between the predicted image and the true image, and comprehensively calculates the error in the x and y directions. For example, the Euclidean error center loss function calculated based on the k-th layer feature is denoted as Loss_C MSE,k , as shown in formula (6):

[0109]

[0110] In formula (6), x gt and y gt are the true values ​​of the horizontal and vertical coordinates of the center point, respectively, pred and y pred are the predicted values ​​of the horizontal and vertical coordinates of the center point respectively.

[0111] In an exemplary embodiment, the preset loss function Loss can be determined by a weighted boundary loss function and a Euclidean error center loss function, as shown in formula (7):

[0112]

[0113] Loss wBCE,k Usually represents the weighted binary cross entropy loss function (Weighted Binary Cross Entropy, WBCE); Loss wIoU,k Typically, it represents a weighted intersection over union (wIoU) loss function, which is used to optimize bounding box regression in object detection tasks. k = 2, 3, 4, 5 represent layers 2-5, namely the uniform progressive decoder DCPG2, DCPG3, DCPG4, and DCPG5, respectively, while k = 6 represents the connected layers g of layers 2, 3, 4, and 5, namely DCPGg. λ1 and λ2 represent the weight coefficients of the two loss functions, respectively. For example, λ1 is set to 10 and λ2 is set to 1 / 90000. By introducing these weight coefficients, the importance of different bounding boxes can be adjusted, thereby better handling samples of different sizes or qualities.

[0114] The embodiments of this application introduce a weighted boundary loss function that quantifies the deviation between the segmentation boundary and the true value, and uses the Euclidean error center loss to measure the deviation between the center point and the true value. The design of these loss functions enables the network to better optimize the target, further improving the training effect and performance of the network. By effectively quantifying the deviation between the segmentation boundary and the grasp center, the network can more accurately learn the characteristics of the target object, thereby improving overall performance and robustness.

[0115] The present application also provides a visual positioning method, such as Figure 7 As shown, the positioning method includes:

[0116] Step S701 obtains a prediction map of the target area from an original image containing the target area based on the method for identifying the target area in an image described in any of the above embodiments;

[0117] Step S702 determines the grasping center and offset angle of visual positioning based on the prediction map of the target area.

[0118] In an exemplary embodiment, determining a grasping center and an offset angle of visual positioning based on a predicted image of the target area includes:

[0119] determining a minimum circumscribed circle of the target area based on the predicted map of the target area;

[0120] The center of the minimum circumscribed circle is used as the grasping center;

[0121] Determine a straight line passing through the center of the minimum circumscribed circle and having a chord with the maximum length intercepted by the outer contour of the target area as a first straight line;

[0122] The target area is segmented using a second straight line passing through the center of the minimum circumscribed circle and perpendicular to the horizontal direction of the predicted image of the target area. The position of the end with smaller curvature in the target area is determined based on the areas of the two segments. For example, the position of the smaller figure in the two segments is used as the position of the smaller curvature in the target area.

[0123] Based on the first straight line and the position of the end with smaller curvature in the target area, an angle at which the end with smaller curvature in the target area deviates from its preset position is determined as the offset angle.

[0124] The visual positioning method described in the embodiment of the present application can be applied to toilet positioning; the target area is the concave area of ​​the toilet body. Figure 8For example, the white area in the figure shows the recessed area of ​​the main part of the toilet. The angles between the first straight line and the x-axis (the horizontal straight line in the figure) are θ1 and θ2. Since the end with the smaller curvature in the recessed area of ​​the toilet is on the left, θ1 is retained and θ2 is discarded. θ1 is the offset angle.

[0125] The visual positioning method described in the embodiment of the present application can accurately determine the grasping position and posture of the target object, providing key positioning information for the grasping operation of the robotic arm.

[0126] The visual positioning method described in the examples of this application was used to verify the accuracy of different product models in real-world scenarios. The results showed a mean absolute error of 2.78 pixels in the x-direction, 5.75 pixels in the y-direction, and 2.43 degrees in rotation angle. Compared to other methods, the mean absolute error of the grasp center and angle decreased by 2.27 pixels, 2.62 pixels, and 2.43 degrees, respectively, representing 45%, 31.3%, and 70.3% reductions, meeting the requirements of real-world scenarios.

[0127] An embodiment of the present application also provides a non-volatile computer-readable storage medium, which stores one or more program instructions, and the one or more program instructions can be executed by one or more processors to implement the method for identifying the target area in the image as described in any of the previous embodiments.

[0128] An embodiment of the present application also provides a non-volatile computer-readable storage medium, which stores one or more program instructions, and the one or more program instructions can be executed by one or more processors to implement the training method of the target area recognition model in the image as described in any of the previous embodiments.

[0129] An embodiment of the present application also provides a non-volatile computer-readable storage medium, which stores one or more program instructions, and the one or more program instructions can be executed by one or more processors to implement the visual positioning method as described in any of the previous embodiments.

[0130] The present application also provides a control device, such as Figure 9 As shown, the control device includes:

[0131] Memory 901, configured to store computer program instructions executable on the processor;

[0132] The processor 902 is configured to execute the computer program instructions to implement the method for identifying a target area in an image as described in any of the previous embodiments.

[0133] The present application also provides a control device, the control device comprising:

[0134] a memory configured to store computer program instructions executable on the processor;

[0135] A processor is configured to execute the computer program instructions to implement the training method for the target area recognition model in the image as described in any of the previous embodiments.

[0136] The present application also provides a control device, the control device comprising:

[0137] a memory configured to store computer program instructions executable on the processor;

[0138] A processor is configured to execute the computer program instructions to implement the visual positioning method as described in any of the previous embodiments.

[0139] The present application also provides a control device, which includes:

[0140] robotic arm;

[0141] A visual sensor provided on the robotic arm;

[0142] The control device as described in the previous embodiment transmits information with the visual sensor.

[0143] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term "computer storage medium" includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

Claims

1. A method for identifying a target area in an image, the method comprising: Using a CNN network to extract local features from an original image containing a target area, and using a Transformer network to extract global features from the original image; Fusing the local features with the global features to obtain fused features; Performing edge information enhancement processing on the fusion features; The enhanced fusion features are decoded to obtain a prediction map of the target area.

2. The method according to claim 1, characterized in that The method of extracting local features from an original image containing a target area using a CNN network and extracting global features from the original image using a Transformer network includes: Extracting multiple layers of local features from the original image using a CNN network, and extracting multiple layers of global features from the original image using a Transformer network; the number of layers of the multiple layers of local features is the same as or different from the number of layers of the multiple layers of global features; Fusing the local features with the global features, including: Selecting all or part of the local features from the multiple layers of local features, and selecting all or part of the global features from the multiple layers of global features; After adjusting the selected local features and the selected global features to the same dimension, they are fused in each dimension to obtain multi-layer fusion features.

3. The method according to claim 1, characterized in that In the case where the fused feature is a multi-layer fused feature, the step of enhancing edge information of the fused feature includes: For each layer of fusion features except the last layer of fusion features, perform the following operations: The next layer of fusion features of this layer of fusion features is upsampled and activated in sequence, and the processing results are used as the weights of the fusion features of this layer; The enhancement process is performed on the fusion features of the layer based on the weights.

4. The method according to claim 1, wherein In the case where the fused feature includes multiple layers of features, the step of enhancing edge information of the fused feature includes: For each layer of fusion features except the last layer of fusion features, perform the following operations: The next layer of fusion features of this layer of fusion features is sequentially upsampled, activated, and color inverted, and the processing results are used as the weights of the fusion features of this layer; The enhancement process is performed on the fusion features of the layer based on the weights.

5. The method according to claim 3 or 4, characterized in that The performing the strengthening process on the layer fusion feature based on the weight includes: Using the weights to weight the fusion features of the layer; The fusion features of this layer are superimposed on the weighted fusion features of this layer at the pixel level, and the superposition result is used as the enhanced fusion feature corresponding to the fusion features of this layer.

6. The method according to claim 3 or 4, characterized in that The activation function corresponding to the activation is the ReLU function.

7. A method for training a model for identifying target regions in an image, the method comprising: Identifying target area training samples using a target area recognition model to obtain a prediction map of the target area training samples; Determining a loss value based on the target area training sample and the prediction graph of the target area training sample, and a preset loss function; Updating parameters of the target area recognition model based on the loss value; The target area recognition model includes: The CNN network module is configured to extract local features from the original image containing the target area; A Transformer network module configured to extract global features from the original image; a feature fusion module configured to fuse the local features with the global features to obtain fused features; The edge information enhancement module is configured to enhance the edge information of the fusion feature: The decoding module is configured to decode the fused features after the enhanced processing to obtain a prediction map of the target area.

8. The training method according to claim 7, characterized in that: The preset loss functions include: weighted margin loss function and Euclidean error center loss function.

9. A visual positioning method, comprising: Obtaining a prediction map of the target area from an original image containing the target area based on the method according to any one of claims 1 to 6; A grasping center and an offset angle of visual positioning are determined based on the predicted image of the target area.

10. The visual positioning method according to claim 9, characterized in that: Determining a grasping center and an offset angle of visual positioning based on the predicted image of the target area includes: determining a minimum circumscribed circle of the target area based on the predicted map of the target area; The center of the minimum circumscribed circle is used as the grasping center; Determine a straight line passing through the center of the minimum circumscribed circle and having a chord with the maximum length intercepted by the outer contour of the target area as a first straight line; Segmenting the target area using a second straight line passing through the center of the minimum circumscribed circle and perpendicular to the horizontal direction of the predicted image of the target area, and determining the position of the end with the smaller curvature in the target area based on the areas of the two segments; Based on the first straight line and the position of the end with smaller curvature in the target area, an angle at which the end with smaller curvature in the target area deviates from its preset position is determined as the offset angle.

11. The visual positioning method according to claim 10, characterized in that: The visual positioning method is applied to toilet positioning; The target area is the recessed area of ​​the toilet bowl.

12. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores one or more program instructions, and the one or more program instructions can be executed by one or more processors to implement the recognition method according to any one of claims 1 to 6, or the training method according to any one of claims 7 to 8, or the visual positioning method according to any one of claims 9 to 11.

13. A control device, characterized in that: The control device includes: a memory configured to store computer program instructions executable on the processor; A processor is configured to execute the computer program instructions to implement the recognition method according to any one of claims 1 to 6, or the training method according to any one of claims 7 to 8, or the visual positioning method according to any one of claims 9 to 11.

14. A control device, characterized in that: The control device comprises: robotic arm; A visual sensor provided on the robotic arm; The control device according to claim 13, which transmits information to the visual sensor.

Citation Information

Patent Citations

  • Target object grabbing method and equipment based on instance segmentation

    CN115797332A

  • Polyp cutting method and device

    CN119251244A

  • Medical image segmentation method fusing SAM global modeling and U-Net local optimization

    CN119832012A