Image segmentation method and device, electronic equipment and computer program product

By using image segmentation models to be processed in the industrial field, the problem that traditional methods are difficult to accurately distinguish micro information codes from backgrounds is solved, and higher image segmentation accuracy and information code recognition accuracy are achieved.

CN120219404APending Publication Date: 2025-06-27SHENZHEN QIANHAI EVOC ASIA-PACIFIC ELECTRONIC EQUIP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510299699.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the industrial field, traditional image segmentation methods are difficult to accurately distinguish micro information codes from backgrounds, resulting in inaccurate information code segmentation, affecting the accuracy of identification.

Method used

An image segmentation method is adopted to segment information codes on the image to be processed using the information code segmentation model. The model includes an image feature extraction module, an image feature processing module and an image feature fusion module. By extracting graph structure features, multi-scale features and performing feature fusion, the accurate distinction between information code and background is achieved.

Benefits of technology

The image feature fusion module integrates multi-scale features, which can more accurately distinguish the information code and background, improve the accuracy of image segmentation, and thus improve the accuracy of information code recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219404A_ABST
    Figure CN120219404A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of image segmentation, and provides an image segmentation method and device, electronic equipment and a computer program product, and the method comprises the steps: carrying out the information code segmentation of a to-be-processed image through an information code segmentation model, and obtaining an information code segmentation result, the information code segmentation model comprises an image feature extraction module, an image feature processing module and an image feature fusion module; the image feature extraction module is used for extracting image structure features in the to-be-processed image; the image feature processing module is used for extracting multi-scale features in the image structure features; and the image feature fusion module is used for carrying out feature fusion on the multi-scale features to obtain an information code segmentation result. The information code segmentation accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image segmentation, and particularly relates to an image segmentation method, apparatus, electronic device, and computer program product. Background Art

[0002] An information code refers to a defined code obtained by encoding information or data in a specific form, including two-dimensional codes, barcodes, etc. Information codes (such as two-dimensional codes) are widely used in the industrial field and can be used for production traceability and quality management, production line data collection, product anti-counterfeiting and after-sales support, etc.

[0003] Currently, when identifying information codes in the industrial field, the information codes are mainly segmented from the image to be recognized through an image segmentation method. However, due to the harsh environment in the industrial field and the small volume of electronic components, the information codes are usually extremely tiny, and it is difficult for traditional image segmentation methods to accurately distinguish the information codes from the background, resulting in inaccurate segmentation of the information codes and thus affecting the accuracy of information code recognition. Summary of the Invention

[0004] Embodiments of this application provide an image segmentation method, apparatus, and electronic device, which can improve the accuracy of information code segmentation.

[0005] In a first aspect, embodiments of this application provide an image segmentation method, including:

[0006] Obtain an image to be processed, where the image to be processed includes one or more information codes;

[0007] Use an information code segmentation model to perform information code segmentation on the image to be processed to obtain an information code segmentation result, where the information code segmentation model includes an image feature extraction module, an image feature processing module, and an image feature fusion module;

[0008] The image feature extraction module is used to extract the graph structure features in the image to be processed;

[0009] The image feature processing module is used to extract multi-scale features from the graph structure features;

[0010] The image feature fusion module is used to perform feature fusion on the multi-scale features to obtain the information code segmentation result.

[0011] In a second aspect, embodiments of this application provide an image segmentation apparatus, including:

[0012] An image acquisition module, configured to obtain an image to be processed, where the image to be processed includes one or more information codes;

[0013] An information code segmentation module is configured to perform information code segmentation on the image to be processed by using an information code segmentation model, so as to obtain an information code segmentation result, where the information code segmentation model includes an image feature extraction module, an image feature processing module, and an image feature fusion module; the image feature extraction module is configured to extract graph structure features in the image to be processed; the image feature processing module is configured to extract multi-scale features from the graph structure features; and the image feature fusion module is configured to perform feature fusion on the multi-scale features to obtain the information code segmentation result.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the image segmentation method described in the first aspect above are implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the steps of the image segmentation method described in the first aspect above are implemented.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, which when running on an electronic device causes the electronic device to execute the image segmentation method described in any item of the first aspect above.

[0017] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:

[0018] In the present application, first, the graph structure features in the image to be processed are extracted by the image feature extraction module in the information code segmentation model, and the multi-scale features in the graph structure features are extracted by the image feature processing module in the information code segmentation model. Finally, feature fusion is performed on the multi-scale features according to the image feature fusion module in the information code segmentation model to obtain the information code segmentation result. Since the above graph structure features are used to characterize the features in the image to be processed, and the above multi-scale features can further characterize the features at different scales, it means that the features of the image to be processed can be extracted through the above graph structure features, and at the same time, the features can be further distinguished at different scales through the above multi-scale features during feature processing, so as to accurately distinguish the background and information code in the image to be processed during information code segmentation and improve the accuracy of image segmentation. Therefore, by performing feature fusion on the multi-scale features through the above image feature fusion module, the features at different scales can be fused more accurately to obtain a more accurate information code segmentation result. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 is a schematic flowchart of an image segmentation method provided by an embodiment of the present application;

[0021] Figure 2 is a schematic structural diagram of an information code segmentation model provided by an embodiment of the present application;

[0022] Figure 3 is a schematic structural diagram of an image feature processing module provided by an embodiment of the present application;

[0023] Figure 4 is a schematic structural diagram of a graph feature fusion module provided by an embodiment of the present application;

[0024] Figure 5 is a schematic structural diagram of an image segmentation device provided by an embodiment of the present application;

[0025] Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0026] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0027] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0028] It should also be understood that the term " / and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0029] As used in the specification of this application and the appended claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, to mean "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0030] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0031] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0032] With the continuous development of the industrial Internet and intelligent manufacturing, information code technology plays a more important role in the industrial field. For example, taking the information code as a two-dimensional code, by attaching a two-dimensional code to industrial products, information such as the production process, raw material source, and quality inspection report of the product can be understood through scanning the product two-dimensional code, thereby enabling the traceability of the entire product life cycle; by attaching a two-dimensional code to industrial equipment, information such as the real-time status, operation guide, and maintenance record of the relevant equipment can be obtained through scanning the equipment two-dimensional code, thereby realizing the refined management and optimization of the production process.

[0033] Currently, when a reader scans an image containing an information code (such as a QR code), the information code (such as a QR code) in the image is mainly segmented and recognized through image segmentation techniques. Image segmentation techniques mainly include traditional image segmentation methods and neural network-based image segmentation methods. Among them, traditional image segmentation methods usually highlight the target by suppressing the background or designing a fixed sliding window; neural network-based image segmentation methods mainly use a neural network (such as a CNN) to learn the information code features and background features in the image and segment the information code from them. For example, in the backbone feature extraction network of the neural network, the global information of the information code can be retained by increasing the receptive field through dilated convolution, or the information code can be segmented by means of a visual attention mechanism.

[0034] However, the production environment in the industrial field may be relatively harsh, and the volume of electronic components used in various industrial devices may be small, which results in the information code in the image being usually small when scanned by the reader. Traditional image segmentation methods rely heavily on the setting of hyperparameters, are difficult to apply to different industrial scenarios, and lack robustness; while neural network-based image segmentation methods are usually also difficult to accurately distinguish small information codes and the background when facing tiny information codes. For example, increasing the receptive field is likely to cause undersampling of the features of small QR codes, resulting in the loss of information code features; the method based on the visual attention mechanism is prone to confusing information code features and background features when learning the features of tiny information codes due to its local limitation.

[0035] Therefore, the above image segmentation methods are difficult to effectively segment the information code (such as a tiny QR code in the industrial field) and the background in the image, resulting in inaccurate information code segmentation.

[0036] To improve the accuracy of image segmentation, the present application provides an image segmentation method. In this method, an information code segmentation model is used to segment the information code in the image to be processed, and an information code segmentation result is obtained. Among them, the above information code segmentation model includes an image feature extraction module, an image feature processing module, and an image feature fusion module; the above image feature extraction module is used to extract the graph structure features in the image to be processed, and the above graph structure features are used to represent background pixels and information code pixels with similar features; the above image feature processing module is used to extract multi-scale features in the graph structure features, and the above multi-scale features are used to represent background pixels and information code pixels with similar features at different scales; the above image feature fusion module is used to perform feature fusion on the multi-scale features to obtain an information code segmentation result.

[0037] Figure 1 The flowchart of an image segmentation method provided by an embodiment of the present application is shown and described in detail as follows:

[0038] S11. Obtain the image to be processed, where the image to be processed includes one or more information codes.

[0039] Among them, the image to be processed can be an image collected or received by a code reader, and the code reader can include one of a mobile phone, a tablet, a barcode scanner, etc. The information code refers to a defined code formed by a specific coding technology, and the information code can include one or more of the following: barcodes, two-dimensional codes, information tags, etc.

[0040] For example, a user can use a code reader to scan a nameplate with a two-dimensional code on an industrial device to obtain the image to be processed including the two-dimensional code.

[0041] S12. Use the information code segmentation model to segment the information codes in the image to be processed to obtain an information code segmentation result. Among them, the information code segmentation model includes an image feature extraction module, an image feature processing module, and an image feature fusion module; the image feature extraction module is used to extract the graph structure features in the image to be processed; the image feature processing module is used to extract multi-scale features from the graph structure features; the image feature fusion module is used to perform feature fusion on the multi-scale features to obtain the information code segmentation result.

[0042] It should be understood that the information code segmentation model can be composed of the image feature extraction module, the image feature processing module, and the image feature fusion module connected in series. The information code segmentation refers to segmenting one or more information codes included in the image to be processed. The graph structure features are used to represent background pixels and information code pixels with similar features; the multi-scale features are used to represent background pixels and information code pixels with similar features at different scales.

[0043] Specifically, after obtaining the image to be processed, the image to be processed is input into the image feature extraction module, and the features in the image to be processed are extracted through the image feature extraction module. Then, the features corresponding to the information code pixels and the background pixels are classified by means of feature classification to obtain the graph structure features. Since the features of the background pixels or the information code pixels are often different, the background pixels and the information code pixels can be distinguished during feature extraction by distinguishing the background pixels or the information pixels with similar features through feature classification methods; then the graph structure features are input into the image feature processing module, and the image feature processing module is used to extract the features in the graph structure features from different scales to obtain the multi-scale features, so as to further distinguish the features corresponding to the information code pixels and the background pixels in the graph structure features from different scales; finally, the multi-scale features are input into the image feature fusion module, and the image feature fusion module fuses the multi-scale features into an image with a unified channel and outputs it to obtain the information code segmentation result.

[0044] For example, refer toFigure 2 As shown, it is a schematic structural diagram of an information code segmentation model. Among them, the above-mentioned image feature extraction module can be a Graph Representation Embedding (GREmbed) module, the above-mentioned image feature processing module can be a Graph Representation Evolution Block (GREvol) module, and the above-mentioned image feature fusion module can be a Graph Feature Multi-scale Fusion Block (GraphMFu) module. These three modules are connected in series from the GREmbed module to the GraphMFu module. The image to be processed outputs the information code segmentation result after passing through the above three modules. By connecting in series the above-mentioned Graph Representation Embedding (GREmbed) module, Graph Representation Evolution (GREvol) module, and Graph Feature Multi-scale Fusion (GraphMFu) module of the series structure, the depth and complexity of the entire information code segmentation model can be significantly increased, so as to better learn the useful information in the image to be processed and improve the accuracy of information code segmentation.

[0045] In this application, first, the graph structure features in the image to be processed are extracted through the image feature extraction module in the information code segmentation model, and the multi-scale features in the graph structure features are extracted through the image feature processing module in the information code segmentation model. Finally, the multi-scale features are feature fused according to the image feature fusion module in the information code segmentation model to obtain the information code segmentation result. Since the above-mentioned graph structure features are used to characterize the features in the image to be processed, and the above-mentioned multi-scale features can further characterize the features at different scales, it means that the features of the image to be processed can be extracted through the above-mentioned graph structure features, and at the same time, the features can be further distinguished at different scales through the above-mentioned multi-scale features during feature processing, so as to accurately distinguish the background and information code in the image to be processed during information code segmentation and improve the accuracy of image segmentation. Therefore, by feature fusing the multi-scale features through the above-mentioned image feature fusion module, the features at different scales can be more accurately fused to obtain a more accurate information code segmentation result.

[0046] In some embodiments, the above-mentioned information code segmentation of the above-mentioned image to be processed using the information code segmentation model to obtain the information code segmentation result includes:

[0047] Extracting the image features of the above-mentioned image to be processed using the above-mentioned image feature extraction module to obtain the above-mentioned graph structure features;

[0048] Updating the above-mentioned graph structure features from different scales using the above-mentioned image feature processing module to obtain the above-mentioned multi-scale features;

[0049] The above multi-scale features are subjected to multi-channel feature fusion by the above image feature fusion module to obtain the above information code segmentation result.

[0050] Among them, the above image feature extraction module may include a feature extraction network and a classification operator; the above image feature processing network may include multiple layers of feature update networks; the above image feature fusion module may include multiple layers of feature fusion networks. It should be noted that in order to improve the accuracy of graph structure feature processing and fusion, the number of layers of the above feature update network and the number of layers of the above feature fusion network need to be the same.

[0051] Specifically, the above feature extraction network can be used to extract the image features of the above image to be processed, and the image features are classified according to the above classification operator to obtain the graph structure features representing the background pixels and information code pixels with similar features in the above image to be processed; then the above graph structure features are continuously updated by multiple layers of feature update networks, and the graph structure features output by each layer of feature update network are aggregated to obtain the above multi-scale features; finally, the graph structure features output by each layer are unified and fused in multiple channels through multiple layers of feature fusion networks to obtain the above information code segmentation result.

[0052] In some embodiments, the above feature extraction network may be one or more of a convolutional neural network (CNN), a residual neural network (ResNet), etc.; the above classification operator may be one of the following: a classification operator based on a decision tree, a classification operator based on a support vector machine, and a classification operator based on a similarity distance, etc. Among them, the classification operator based on a decision tree classifies by constructing a decision tree corresponding to the above graph structure features, the classification operator based on a support vector machine classifies by mapping the graph structure features to a high-dimensional space, and the classification operator based on a similarity distance classifies by calculating the similarity distance of each graph structure feature; the above multiple layers of feature update networks may include multi-scale CNNs, etc.; the above multiple layers of feature fusion networks may include multiple layers of convolutional layers, etc.

[0053] In the embodiments of the present application, the background pixels and information code pixels are initially distinguished by the feature extraction network and the classification operator in the image feature extraction module, then the background pixels and information code pixels are further distinguished from different scales and depths by the multiple layers of feature update networks in the image feature processing module, and finally the feature fusion is performed by the multiple layers of feature fusion networks in the image feature fusion module, so that the background pixels and information code pixels can be continuously learned and deeply distinguished, thereby improving the accuracy of information code segmentation.

[0054] In some embodiments, the above-mentioned image feature extraction module includes a graph feature extraction module and a graph feature conversion module. The above-mentioned graph structure features include graph node features and an adjacency matrix corresponding to the above-mentioned graph node features. Among them, the above-mentioned graph feature extraction module can be a residual neural network (ResNet), and the above-mentioned graph feature conversion module can be a classification operator based on similarity distance. The above-mentioned graph node features are the features represented by nodes in the image to be processed, and the adjacency matrix includes graph node features with similar features.

[0055] Correspondingly, extracting the image features of the above-mentioned image to be processed by using the above-mentioned image feature extraction module to obtain the above-mentioned graph structure features includes:

[0056] Extracting the image features of the above-mentioned image to be processed by using the above-mentioned graph feature extraction module to obtain the above-mentioned graph node features;

[0057] For each of the above-mentioned graph node features, using the above-mentioned graph feature conversion module to calculate the distance between the above-mentioned graph node feature and other above-mentioned graph node features, and determining the above-mentioned adjacency matrix corresponding to the above-mentioned graph node feature according to the above-mentioned distance.

[0058] Among them, the extracted above-mentioned image features at least include texture features, and the above-mentioned graph node features are used to represent the above-mentioned background pixels or the above-mentioned information code pixels. The adjacency matrix includes graph node features corresponding to background pixels or information code pixels with similar features.

[0059] Specifically, the above-mentioned image to be processed can be input into a residual neural network (ResNet), and the texture features corresponding to the image to be processed are output through the above-mentioned residual neural network. Then, each output texture feature is used as a node to obtain graph node features X, X ∈ R D×H / 4×W / 4 , where R represents the set of real numbers, D represents the channel dimension, H represents the height of the image to be processed, and W represents the width of the image to be processed. Then, for each graph node feature, the distance between the above-mentioned graph node feature and other above-mentioned graph node features is calculated through the graph feature conversion module, and a preset number of graph node features are selected according to the above-mentioned distance to form the above-mentioned adjacency matrix. Of course, since the information code is usually a defined code with unified color or format. For example, a two-dimensional code is a black-and-white pattern distributed regularly on a plane (in two-dimensional directions). Therefore, in order to better distinguish the information code (such as a two-dimensional code) and the background, the above-mentioned image features can also include one or more of color features, shape features, etc., which are not limited here.

[0060] It should be noted that when selecting nodes, it can be selected according to the actual situation. For example, the image features of each row are selected as nodes, or the image to be processed is divided into regions, and the image features of each region are used as nodes.

[0061] In the embodiments of the present application, by using graph node features to represent background pixels and information code pixels, it is more convenient to process the background and information code in the image to be processed; at the same time, the graph feature transformation module distinguishes similar graph node features through the distances between graph node modules. Since background pixels and information code pixels usually have different features, the information code and the background can be better distinguished.

[0062] In some embodiments, the above distance can be one of the Euclidean distance, Manhattan distance, etc. Assuming the above distance is the Euclidean distance, the above uses the above graph feature transformation module to calculate the distances between the above graph node features and other above graph node features, and determines the above adjacency matrix corresponding to the above graph node features according to the above distances, including:

[0063] Use the above graph feature transformation module to calculate the Euclidean distances between the above graph node features and other above graph node features;

[0064] Select K graph node features from other above graph node features according to the above Euclidean distances, and form the above adjacency matrix according to the selected above K graph node features, where K is a positive integer greater than 1.

[0065] Specifically, for any graph node feature X i , the Euclidean distances between this graph node feature and other graph node features can be calculated, and then according to the above Euclidean distances, find the K graph node features closest to this graph node feature, and form the adjacency matrix A with this graph node feature X i and its K closest surrounding graph node features, A ∈ R HW / 16×HW / 16 , where R represents the set of real numbers, H represents the height of the image to be processed, and W represents the width of the image to be processed.

[0066] It should be noted that if the value of K is too large, it may lead to inaccurate distinction between the background and the information code. If the value of K is too small, it may lead to classification errors or overfitting. Therefore, methods such as cross-validation can be used to determine the value of K. Exemplarily, the value of the above K can be 9.

[0067] In the embodiments of the present application, by using the Euclidean distance to select the K graph node features closest to each graph node feature, the information code and the background can be more accurately distinguished.

[0068] In some embodiments, the above image feature processing module includes a plurality of graph filtering modules and a plurality of graph updating modules, and the above graph filtering modules and the above graph updating modules are alternately connected in series; wherein, the above graph filtering module is used to update the above graph node features from different scales, and the above graph updating module may include a classification operator for updating the above adjacency matrix corresponding to the above graph node features.

[0069] Updating the above graph structure features from different scales by using the above image feature processing module to obtain the above multi-scale features, including:

[0070] Taking the above graph node features and the adjacency matrix corresponding to the above graph node features as the input of the above graph filtering module, obtaining the graph node feature matrix output by the above graph filtering module, and updating the above graph node feature matrix by using the above graph updating module to obtain the target graph structure features;

[0071] Taking the above target graph structure features output by the previous above graph updating module as the input of the next above graph filtering module, obtaining the new graph node feature matrix output by the next above graph filtering module, updating the above new graph node feature matrix by using the next above graph updating module to obtain the new target graph structure features, and returning to the step of taking the above target graph structure features output by the previous above graph updating module as the input of the next above graph filtering module and subsequent steps until the above graph filtering module and the above graph updating module in the alternating series connection both output results;

[0072] Aggregating the graph node feature matrices output by each of the above graph filtering modules to obtain the above multi-scale features.

[0073] Specifically, the mathematical model of the above graph filtering matrix can be shown as follows: Y = AXW, where Y is the graph node feature matrix output by the above graph filtering matrix, X is the input graph node feature, A is the adjacency matrix corresponding to the input graph node feature, and W is a matrix with learnable parameters; the above graph updating module can include a classification operator. After receiving the graph node feature matrix Y output by the graph filtering matrix, the target graph structure features of the graph node feature matrix are further extracted through the classification operator. The classification operator is similar to that in the above embodiment and will not be elaborated here. Since the above graph filtering matrix and the graph updating module are connected in an alternating series, the above target graph structure features output by the previous graph updating module will be used as the input of the next graph filtering module, thereby continuously updating the above graph node features and the target graph structure features until the above graph filtering module and the above graph updating module in the alternating series connection both output results, and aggregating the graph node feature matrices output by each of the above graph filtering modules to obtain the above multi-scale features.

[0074] It should be noted that when the above graph filtering module and the graph updating module are connected in an alternating series, the number of graph filtering modules and the number of graph updating modules can be set according to the actual situation. However, in order to ensure that the graph node feature matrix is output at the last layer, the number of the above graph filtering modules needs to be more than the number of the above graph updating modules. Refer to Figure 3As shown in the figure, assume that the above-mentioned image feature processing module is composed of four graph filtering modules and three graph update modules connected in series alternately. After the graph filtering module in the first layer receives the graph node feature X and the adjacency matrix A corresponding to the graph node feature, it updates to obtain a graph node feature matrix and inputs it to the subsequent graph update module until the last graph filtering module outputs the final graph node feature matrix. The graph feature matrices output by each layer are summarized to obtain multi-scale features, namely Z1, Z2, Z3, and Z4.

[0075] In the embodiments of the present application, through the graph filtering module and the graph update module connected in series alternately, the above-mentioned target graph structure features can be continuously updated, so as to further distinguish the background and the information code from different scales and improve the accuracy of information code segmentation.

[0076] In some embodiments, in order to more accurately extract features at different scales, and corresponding to the above-mentioned image feature extraction module, the above-mentioned graph update module can select the same graph feature conversion module as in the above-mentioned image feature extraction module. The above-mentioned graph update module includes one above-mentioned graph feature conversion module and one pooling module. Among them, the above-mentioned pooling module includes a pooling layer for obtaining deeper features through pooling operations. The above-mentioned graph update module is composed of a graph feature conversion module and a pooling module connected in parallel.

[0077] The above-mentioned target graph structure features include a new graph node feature matrix and a new adjacency matrix. The above-mentioned use of the above-mentioned graph update module to update the above-mentioned graph node feature matrix to obtain the target graph structure features includes:

[0078] Determine the background pixels and information code pixels with similar features in the above-mentioned graph node feature matrix according to the above-mentioned graph feature conversion module to obtain the above-mentioned new adjacency matrix;

[0079] Perform pooling processing on the above-mentioned graph node feature matrix by a preset multiple according to the above-mentioned pooling module to obtain the above-mentioned new graph node feature matrix, so as to obtain the above-mentioned target graph structure features.

[0080] Specifically, for the graph node feature matrix Y, the graph feature conversion module can also find the K graph node features with the closest distance to each graph node feature in it according to the Euclidean distance to form the above-mentioned new adjacency matrix A'; at the same time, the above-mentioned pooling module (i.e., the pooling layer) performs average pooling processing on the graph node feature matrix Y by a preset multiple (assumed to be 2) to obtain the above-mentioned new graph node feature matrix Y'. The processing method of the above-mentioned graph conversion module can refer to the above-mentioned embodiments and will not be elaborated here.

[0081] It should be noted that with reference to Figure 3As shown, for the new graph node feature matrix Y' and the new adjacency matrix A' output by the first graph update module, after being input into the next-layer graph filtering module, the next-layer graph filtering module will update the above new graph node feature matrix in the same way, that is, substituting the new graph node feature matrix Y' and the new adjacency matrix A' into the corresponding mathematical model of the graph filtering module to obtain the graph node feature matrix output by this layer. For example, the graph node feature matrix output by the second graph filtering module can be Y" = A'Y'W, where Y" is the graph node feature matrix output by this layer of the graph filtering module, and W is the parameter matrix of this layer of the graph filtering module. The subsequent update process is similar and will not be elaborated here.

[0082] In the embodiments of the present application, through the above pooling module and the same graph feature conversion module, deeper target graph structure features can be further deeply extracted, that is, the features corresponding to background pixels and information code pixels can be more deeply mined, so as to better distinguish the background and the information code. At the same time, through the above pooling process, the amount of calculation can also be reduced, and the robustness of the information code segmentation model can be improved.

[0083] In some embodiments, the above-mentioned multi-channel feature fusion of the above-mentioned multi-scale features by the above-mentioned image feature fusion module to obtain the above-mentioned information code segmentation result includes:

[0084] Using the above-mentioned image feature fusion module to unify the number of channels of the above-mentioned multi-scale features, and performing channel stacking on the multi-scale features with the unified number of channels to obtain a multi-scale feature map;

[0085] Performing upsampling of a preset multiple on the above-mentioned multi-scale feature map to obtain the above-mentioned information code segmentation result.

[0086] Among them, the above-mentioned image feature fusion module may include multiple convolutional layers, multiple upsampling layers, and channel stacking; the convolutional layer is used to unify the number of channels of the multi-scale features, and the upsampling layer is used to adjust the resolution, where the multiple of upsampling corresponds to the multiple of the pooling module. The above-mentioned information code segmentation result can be an image with the same resolution size as the original image, which is composed of pixel points with gray values of 0 or 255.

[0087] Specifically, assuming that the image feature processing module is composed of four graph filtering modules and three graph update modules connected in series alternately, and the multiple of the pooling process of each graph update module is 2, then the above-mentioned scale features include 4 graph node feature matrices: Z1, Z2, Z3, and Z4, where, C1, C2, C3, and C4 are the corresponding number of channels. Refer to Figure 4As shown in the figure, it is a schematic structural diagram of the graph feature fusion module. First, each graph node feature matrix will pass through a 1×1 convolutional layer to unify to the same number of channels. Then, Z2, Z3, and Z4 are respectively upsampled by 2 times, 4 times, and 8 times to unify the resolution, obtaining four updated graph node feature matrices Z’1, Z’2, Z’3, and Z’4. Among them, Z’1 ∈ R C×HW / 16×HW / 16 , Z’2 ∈ R C×HW / 16×HW / 16 , Z’3 ∈ R C×HW / 16×HW / 16 , Z’4 ∈ R C×HW / 16×HW / 16 , C is the unified number of channels. Then, Z’1, Z’2, Z’3, and Z’4 are stacked in the channel dimension to obtain a multi-scale feature map Z. Among them, Z ∈ R 4C ×HW / 16×HW / 16 , 4C is the number of channels of the multi-scale feature map Z. Finally, the multi-scale feature map Z passes through a 1×1 convolutional layer to reduce the number of channels to 1, obtaining a single-channel image (i.e., an image composed of pixel points with gray values of 0 or 255), and is upsampled by 4 times to the size of the image to be processed, obtaining the above information code segmentation result.

[0088] In the embodiments of the present application, through the convolutional layer, upsampling, and channel stacking in the image feature fusion module, the feature matrices of each graph node can be feature-fused and unifiedly reduced to a single channel, so as to fuse the features corresponding to the background pixels and information code pixels of different scales learned by the previous network, and obtain a more accurate information code segmentation result.

[0089] It should also be noted that in the actual application of the information code segmentation model in the embodiments of the present application, since there is no visual attention mechanism, etc., it is not necessary to perform corresponding parameter training for different scenarios, so that it can effectively adapt to different scenarios, improve the applicability of the information code segmentation model, and also improve the robustness of the application of the information code segmentation model. In addition, the model structures and complexities of each module (including the image feature extraction module, the image feature processing module, and the image feature fusion module) in the above information code segmentation model are relatively low, so the overall scale of the information code segmentation model is small, and it can be applied to devices with small memory, which further improves the applicability of the information code segmentation model.

[0090] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0091] Corresponding to the image segmentation method described in the above embodiments, Figure 5 The figure shows a schematic structural diagram of the image segmentation device provided by the embodiments of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0092] Reference Figure 5 The device may be an image segmentation device 51, and the image segmentation device 51 may include an image acquisition module 511 and an information code segmentation module 512.

[0093] Reference Figure 5 The image segmentation device 51 includes:

[0094] The above-mentioned image acquisition module 511 is used to acquire an image to be processed, and the image to be processed contains one or more information codes;

[0095] The above-mentioned information code segmentation module 512 is used to perform information code segmentation on the image to be processed by using an information code segmentation model to obtain an information code segmentation result. Among them, the information code segmentation model includes an image feature extraction module, an image feature processing module, and an image feature fusion module; the image feature extraction module is used to extract the graph structure features in the image to be processed; the image feature processing module is used to extract multi-scale features from the graph structure features; the image feature fusion module is used to perform feature fusion on the multi-scale features to obtain the information code segmentation result.

[0096] In some embodiments, when the above-mentioned information code segmentation module 512 performs information code segmentation on the image to be processed by using an information code segmentation model to obtain an information code segmentation result, it includes:

[0097] Using the above-mentioned image feature extraction module to extract the image features of the image to be processed to obtain the graph structure features;

[0098] Using the above-mentioned image feature processing module to update the graph structure features from different scales to obtain the multi-scale features;

[0099] Using the above-mentioned image feature fusion module to perform multi-channel feature fusion on the multi-scale features to obtain the information code segmentation result.

[0100] In some embodiments, the above-mentioned image feature extraction module includes a graph feature extraction module and a graph feature conversion module. The graph structure features include graph node features and an adjacency matrix corresponding to the graph node features; among them, the graph feature extraction module may be a residual neural network (ResNet), and the graph feature conversion module may be a classification operator based on a similarity distance. The graph node features are features represented by nodes in the image to be processed, and the adjacency matrix includes graph node features with similar features.

[0101] Correspondingly, when the above-mentioned information code segmentation module 512 uses the above-mentioned image feature extraction module to extract the image features of the image to be processed to obtain the graph structure features, it includes:

[0102] Extract the image features of the to-be-processed image using the above-mentioned graph feature extraction module to obtain the above-mentioned graph node features;

[0103] For each of the above-mentioned graph node features, use the above-mentioned graph feature transformation module to calculate the distance between the above-mentioned graph node feature and other above-mentioned graph node features, and determine the above-mentioned adjacency matrix corresponding to the above-mentioned graph node feature according to the above-mentioned distance, where the above-mentioned adjacency matrix includes the graph node features corresponding to background pixels or information code pixels with similar features.

[0104] In some embodiments, the above-mentioned distance can be one of Euclidean distance, Manhattan distance, etc. Assuming the above-mentioned distance is Euclidean distance, when the above-mentioned information code segmentation module 512 uses the above-mentioned graph feature transformation module to calculate the distance between the above-mentioned graph node feature and other above-mentioned graph node features and determines the above-mentioned adjacency matrix corresponding to the above-mentioned graph node feature according to the above-mentioned distance, it includes:

[0105] Use the above-mentioned graph feature transformation module to calculate the Euclidean distance between the above-mentioned graph node feature and other above-mentioned graph node features;

[0106] Select K graph node features from other above-mentioned graph node features according to the above-mentioned Euclidean distance, and form the above-mentioned adjacency matrix according to the selected above-mentioned K graph node features, where K is a positive integer greater than 1.

[0107] In some embodiments, the above-mentioned image feature processing module includes a plurality of graph filtering modules and a plurality of graph updating modules, and the above-mentioned graph filtering modules and the above-mentioned graph updating modules are alternately connected in series; wherein, the above-mentioned graph filtering module is used to update the above-mentioned graph node features from different scales, and the above-mentioned graph updating module may include a classification operator for updating the above-mentioned adjacency matrix corresponding to the above-mentioned graph node features.

[0108] When the above-mentioned information code segmentation module 512 uses the above-mentioned image feature processing module to update the above-mentioned graph structure features from different scales to obtain the above-mentioned multi-scale features, it includes:

[0109] Use the above-mentioned graph node features and the above-mentioned adjacency matrix corresponding to the above-mentioned graph node features as the input of the above-mentioned graph filtering module to obtain the graph node feature matrix output by the above-mentioned graph filtering module, and use the above-mentioned graph updating module to update the above-mentioned graph node feature matrix to obtain the target graph structure feature;

[0110] Take the above-mentioned target graph structure feature output by the previous graph update module as the input of the next graph filtering module, obtain a new graph node feature matrix output by the next graph filtering module, use the next graph update module to update the above-mentioned new graph node feature matrix, obtain a new target graph structure feature, return the step of taking the above-mentioned target graph structure feature output by the previous graph update module as the input of the next graph filtering module and subsequent steps until both the above-mentioned graph filtering module and the above-mentioned graph update module in the alternating series output results;

[0111] Summarize the graph node feature matrices output by each of the above graph filtering modules to obtain the above multi-scale features.

[0112] In some embodiments, in order to more accurately extract features of different scales, and corresponding to the above image feature extraction module, the above graph update module may select the same graph feature conversion module as in the above image feature extraction module. The above graph update module includes a graph feature conversion module and a pooling module. Among them, the above pooling module includes a pooling layer for obtaining deeper features through pooling operations. The above graph update module is composed of a graph feature conversion module and a pooling module in parallel.

[0113] The above target graph structure feature includes a new graph node feature matrix and a new adjacency matrix. When the information code segmentation module 512 uses the above graph update module to update the above graph node feature matrix to obtain the target graph structure feature, it includes:

[0114] Determine the background pixels and information code pixels with similar features in the above graph node feature matrix according to the above graph feature conversion module to obtain the above new adjacency matrix;

[0115] Perform pooling processing on the above graph node feature matrix by a preset multiple according to the above pooling module to obtain the above new graph node feature matrix, thereby obtaining the above target graph structure feature.

[0116] In some embodiments, when the information code segmentation module 512 uses the above image feature fusion module to perform multi-channel feature fusion on the above multi-scale features to obtain the above information code segmentation result, it includes:

[0117] Use the above image feature fusion module to unify the number of channels of the above multi-scale features and perform channel stacking on the multi-scale features with unified channel numbers to obtain a multi-scale feature map;

[0118] Perform upsampling on the above multi-scale feature map by a preset multiple to obtain the above information code segmentation result.

[0119] It should be noted that for the content such as information interaction and execution process between the devices / units, since it is based on the same concept as the method embodiments of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not elaborated herein.

[0120] Figure 6 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 6 shown, the electronic device 6 of this embodiment includes: at least one processor 60 ( Figure 6 only one is shown in the figure), a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the steps in any of the method embodiments are implemented.

[0121] The electronic device 6 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 6 this is only an example of the electronic device 6 and does not constitute a limitation on the electronic device 6. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may further include an input and sending device, a network access device, a bus, etc.

[0122] The so-called processor 60 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0123] In some embodiments, the memory 61 may be an internal storage unit of the electronic device 6, such as a hard disk or memory of the electronic device 6. The memory 61 may also be an external storage device of the electronic device 6, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 6. Further, the memory 61 may also include both the internal storage unit and the external storage device of the electronic device 6. The memory 61 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as program codes of the computer program. The memory 61 may also be used to temporarily store data that has been sent or will be sent.

[0124] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the functional units and modules is used as an example for illustration. In practical applications, the functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0125] An embodiment of the present application further provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps in any of the foregoing method embodiments are implemented.

[0126] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.

[0127] An embodiment of the present application provides a computer program product, and when the computer program product runs on an electronic device, the electronic device is caused to implement the steps in the foregoing method embodiments when executed.

[0128] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the method of the above embodiments in this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0129] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0130] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0131] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0132] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0133] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. An image segmentation method, characterized in that: include: Acquire an image to be processed, wherein the image to be processed includes one or more information codes; Performing information code segmentation on the image to be processed by using an information code segmentation model to obtain an information code segmentation result, wherein the information code segmentation model includes an image feature extraction module, an image feature processing module and an image feature fusion module; The image feature extraction module is used to extract the image structure features in the image to be processed; The image feature processing module is used to extract multi-scale features from the graph structure features; The image feature fusion module is used to perform feature fusion on the multi-scale features to obtain the information code segmentation result.

2. The image segmentation method according to claim 1, characterized in that: The step of performing information code segmentation on the image to be processed by using the information code segmentation model to obtain the information code segmentation result includes: Utilizing the image feature extraction module to extract image features of the image to be processed, and obtaining the graph structure features; Using the image feature processing module to update the graph structure features from different scales to obtain the multi-scale features; The image feature fusion module is used to perform multi-channel feature fusion on the multi-scale features to obtain the information code segmentation result.

3. The image segmentation method according to claim 2, characterized in that: The image feature extraction module includes a graph feature extraction module and a graph feature conversion module, the graph structure feature includes a graph node feature and an adjacency matrix corresponding to the graph node feature; the step of extracting the image feature of the image to be processed by the image feature extraction module to obtain the graph structure feature includes: The graph feature extraction module is used to extract the image features of the image to be processed to obtain the graph node features. For each of the graph node features, the graph feature conversion module is used to calculate the distance between the graph node feature and other graph node features, and the adjacency matrix corresponding to the graph node feature is determined based on the distance.

4. The image segmentation method according to claim 3, characterized in that: The using the graph feature conversion module to calculate the distance between the graph node feature and other graph node features, and determining the adjacency matrix corresponding to the graph node feature according to the distance, includes: Utilizing the graph feature conversion module to calculate the Euclidean distance between the graph node feature and other graph node features; K graph node features are selected from the other graph node features according to the Euclidean distance, and the adjacency matrix is ​​formed according to the selected K graph node features, wherein K is a positive integer greater than 1.

5. The image segmentation method according to claim 3, characterized in that: The image feature processing module includes a plurality of image filtering modules and a plurality of image updating modules, wherein the image filtering modules and the image updating modules are alternately connected in series; and the image feature processing module is used to update the image structure features from different scales to obtain the multi-scale features, including: Using the graph node features and the adjacency matrix corresponding to the graph node features as inputs of the graph filtering module to obtain a graph node feature matrix output by the graph filtering module, and using the graph updating module to update the graph node feature matrix to obtain target graph structure features; Using the target graph structure features output by the previous graph update module as the input of the next graph filtering module to obtain a new graph node feature matrix output by the next graph filtering module, using the next graph update module to update the new graph node feature matrix to obtain a new target graph structure feature, returning to the step of using the target graph structure features output by the previous graph update module as the input of the next graph filtering module and subsequent steps, until the graph filtering modules and graph update modules connected alternately in series all output results; The graph node feature matrix output by each of the graph filtering modules is aggregated to obtain the multi-scale features.

6. The image segmentation method according to claim 5, characterized in that: The graph updating module includes a graph feature conversion module and a pooling module, the target graph structural features include a new graph node feature matrix and a new adjacency matrix, and the graph node feature matrix is ​​updated by the graph updating module to obtain the target graph structural features, including: Determine the background pixels and information code pixels with similar features in the graph node feature matrix according to the graph feature conversion module to obtain the new adjacency matrix; The graph node feature matrix is ​​pooled with a preset multiple according to the pooling module to obtain the new graph node feature matrix, thereby obtaining the target graph structure feature.

7. The image segmentation method according to any one of claims 2 to 6, characterized in that: The using the image feature fusion module to perform multi-channel feature fusion on the multi-scale features to obtain the information code segmentation result includes: The image feature fusion module is used to unify the number of channels of the multi-scale features, and the multi-scale features with the unified number of channels are stacked to obtain a multi-scale feature map; The multi-scale feature map is upsampled by a preset multiple to obtain the information code segmentation result.

8. An image segmentation device, characterized in that: include: An image acquisition module, used for acquiring an image to be processed, wherein the image to be processed includes one or more information codes; An information code segmentation module is used to perform information code segmentation on the image to be processed using an information code segmentation model to obtain an information code segmentation result, wherein the information code segmentation model includes an image feature extraction module, an image feature processing module and an image feature fusion module; the image feature extraction module is used to extract graph structure features in the image to be processed; the image feature processing module is used to extract multi-scale features in the graph structure features; the image feature fusion module is used to perform feature fusion on the multi-scale features to obtain the information code segmentation result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 7 when being executed.