Code Recognition Method, Device, Equipment and Readable Storage Medium

By performing connectivity area analysis, tone division and fusion of Verilog code images, and combining with a pre-trained code recognition model, automatic code recognition is realized, solving the problems of low efficiency and error-prone manual input.

CN119863485BActive Publication Date: 2025-06-17BEIJING INSTITUTE OF OPEN SOURCE CHIP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510350623.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-17
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

During Verilog code development, the devices on the development side usually do not support direct code copying, resulting in inefficient and error-prone manual input methods.

Method used

By obtaining the image containing the code to be identified, performing connection area analysis and tone division, fusing the foreground area to segment the image, and finally inputting the sub-image to the pre-trained code recognition model to automatically identify the code.

Benefits of technology

It realizes accurate and efficient automatic identification of code, avoiding the problems of low efficiency and high error rate caused by manual input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863485B_ABST
    Figure CN119863485B_ABST
Patent Text Reader

Abstract

The present application provides a code recognition method, apparatus, device and readable storage medium, relating to the field of computer technology. The method includes: obtaining an image including the code to be recognized, and performing connected region analysis on the image to obtain a first foreground region; the first foreground region includes the connected region where the code to be recognized is located; obtaining the hue of the pixel points in the image, and dividing the image according to the hue to obtain a second foreground region; the second foreground region includes the region composed of the pixel points of the code to be recognized; fusing the first foreground region and the second foreground region to obtain a fused foreground region, and segmenting the image based on the fused foreground region to obtain at least one sub-image; inputting the at least one sub-image into a pre-trained code recognition model to obtain the code to be recognized in the image. The code method of the present application can quickly and accurately extract the code to be recognized from the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a code recognition method, apparatus, device, and readable storage medium. Background Art

[0002] Hardware Description Language (Verilog Hardware Description Language) is a language that describes the hardware structure and behavior of a digital system in text form.

[0003] During the development process of Verilog code, it is usually necessary to summarize and communicate Verilog code between different devices. However, the devices on the development side usually do not support the direct copying of Verilog code to local devices. In the related art, the Verilog code to be summarized is usually input manually in the devices on the development side. However, this method has problems of high error rate and low efficiency. Summary of the Invention

[0004] Embodiments of this application provide a code recognition method, apparatus, device, and readable storage medium to solve the problems of low efficiency and high error rate in obtaining code in the prior art.

[0005] In a first aspect, embodiments of this application provide a code recognition method, including:

[0006] Obtain an image including the code to be recognized, and perform connected region analysis on the image to obtain a first foreground region; the first foreground region includes the connected region where the code to be recognized is located;

[0007] Obtain the hue of the pixel points in the image, and divide the image according to the hue to obtain a second foreground region; the second foreground region includes the region composed of the pixel points of the code to be recognized;

[0008] Fuse the first foreground region and the second foreground region to obtain a fused foreground region, and segment the image based on the fused foreground region to obtain at least one sub-image;

[0009] Input at least one of the sub-images into a pre-trained code recognition model to obtain the code to be recognized in the image.

[0010] In a second aspect, embodiments of this application provide a code recognition apparatus, including:

[0011] A first acquisition module, configured to obtain an image including the code to be recognized, and perform connected region analysis on the image to obtain a first foreground region; the first foreground region includes the connected region where the code to be recognized is located;

[0012] A second acquisition module, configured to acquire the hue of the pixel points in the image, and divide the image according to the hue to obtain a second foreground region; the second foreground region includes a region composed of the pixel points of the code to be recognized;

[0013] A third acquisition module, configured to fuse the first foreground region and the second foreground region to obtain a fused foreground region, and segment the image based on the fused foreground region to obtain at least one sub-image;

[0014] A fourth acquisition module, configured to input at least one of the sub-images into a pre-trained code recognition model to obtain the code to be recognized in the image.

[0015] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor;

[0016] A memory for storing executable instructions of the processor;

[0017] Wherein, the processor is configured to execute the instructions to implement the method of the first aspect.

[0018] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the method of the first aspect.

[0019] In this embodiment, a connected region analysis is performed on the image including the code to be recognized to obtain a first foreground region; the first foreground region obtained thereby is a region of the code to be recognized determined from the shape level. The hue of the pixel points in the image is acquired, and the image is regionally divided according to the hue to obtain a second foreground region; the second foreground region obtained thereby is a region of the code to be recognized determined from the color level. The color of the code and the background color are usually different. Therefore, the first foreground region obtained based on the connected region analysis and the second foreground region obtained based on the hue of the pixel points can accurately reflect the region where the code to be recognized is located, and the fused foreground region obtained by fusing the first foreground region and the second foreground region further improves the accuracy of reflecting the position where the code to be recognized is located. Further, the image is segmented based on the fused foreground region, and then the obtained sub-images are input into a pre-trained code recognition model, and the code to be recognized in the image can be accurately and efficiently obtained. It solves the problems of low input efficiency and easy error caused by manually inputting codes in the related art.

[0020] The above description is only an overview of the technical solution of the present application. In order to better understand the technical means of the present application, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0022] Figure 1 It is a schematic diagram of an application scenario of a code recognition method provided by an embodiment of the present application;

[0023] Figure 2 It is a flowchart of the steps of a code recognition method provided by an embodiment of the present application;

[0024] Figure 3 It is a flowchart of the steps of a code recognition method provided by an embodiment of the present application;

[0025] Figure 4 It is a flowchart of the steps of a code recognition method provided by an embodiment of the present application;

[0026] Figure 5 It is a flowchart of the steps of a code recognition method provided by an embodiment of the present application;

[0027] Figure 6 It is a flowchart of the steps of a code recognition method provided by an embodiment of the present application;

[0028] Figure 7 It is a flowchart of the steps of a code recognition method provided by an embodiment of the present application;

[0029] Figure 8 It is a block diagram of a code recognition device provided by an embodiment of the present invention;

[0030] Figure 9 It is a block diagram of an electronic device provided by an embodiment of the present invention;

[0031] Figure 10 It is a block diagram of another electronic device according to another embodiment of the present invention. Detailed Embodiments

[0032] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0033] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The term "multiple" in the embodiments of the present application means two or more, and other quantifiers are similar.

[0034] Figure 1 It is a schematic diagram of an application scenario of a code recognition method provided by an embodiment of the present application. Referring to Figure 1 it, the application scenario at least includes a server platform 10 and a local computing device 20.

[0035] Among them, the server platform 10 can be a server on the development side. During the development of some codes (such as Verilog codes), a large amount of retrievals strongly related to the codes, summarization of the codes, and text communication are required. The server platform 10 on the development side may not support the direct copying of development information to the local computing device 20. Therefore, in the related art, the codes developed in the server platform 10 need to be input into the local computing device 20 by manual input. In application scenarios involving a large number of code processing, such as the scenario of summarizing code signals in a document, manual input will not only lead to repetitive and inefficient work, but also may cause spelling mistakes and missing input problems.

[0036] To solve the problems of related technologies, the local computing device 20 in this embodiment obtains an image containing a code to be recognized provided by the service platform 10, and obtains the code to be recognized in the image according to the following method: performing connected component analysis on the image including the code to be recognized to obtain a first foreground region; the first foreground region is the connected component where the code to be recognized is located; performing region division on the image according to the gray values of the pixel points in the image to obtain a second foreground region; the second foreground region is the region composed of the pixel points of the code to be recognized; fusing the first foreground region and the second foreground region to obtain a fused foreground region, and segmenting the image based on the fused foreground region to obtain at least one sub-image; inputting the at least one sub-image into a pre-trained code recognition model to obtain the code to be recognized in the image. Based on this embodiment, the code to be recognized in the image can be automatically recognized, avoiding the problems of low efficiency and high error rate caused by manually inputting the code in related technologies.

[0037] It should be noted that Figure 1 The shown is just an application scenario of the method in this embodiment. The pictures containing the code to be recognized processed in this embodiment can also come from devices such as cloud servers, mobile terminals, tablet computers, storage devices, etc., and are not limited to the server platform. The device for recognizing the code to be recognized in the image in this embodiment can be an electronic device such as a server, a mobile terminal, a tablet computer, etc., which is not limited here.

[0038] Figure 2 The shown is a flowchart of the steps of a code recognition method provided in this embodiment. Referring to Figure 2 The method may include the following steps:

[0039] Step 101, obtain an image including the code to be recognized, and perform connected component analysis on the image to obtain a first foreground region.

[0040] Among them, the first foreground region is the connected component where the code to be recognized is located.

[0041] Exemplarily, the image in this step is the original image including the code to be recognized. The image is preprocessed to obtain a preprocessed image. Morphological processing is performed on the preprocessed image to obtain a morphologically processed image, and then connected component analysis is performed on the morphologically processed image to obtain a first foreground region. Exemplarily, the preprocessing method may include grayscale processing and binarization processing; morphological processing can be performed on the preprocessed image through closing operation.

[0042] Exemplarily, according to the gray values of each pixel in the image, through the connected region analysis method, the mutually connected pixels are identified as a connected domain, thereby obtaining the first foreground region where the code to be recognized is located. The first foreground region has a corresponding identifier. Specifically, through the connected region analysis, the centroid of the connected domain is obtained, and the centroid is determined as the seed point of the first foreground region, and an identifier is set for the seed point. For example, the identifier can be set to 1.

[0043] Furthermore, the seed points of the background region can be set according to empirical data. For example, since the code to be recognized is usually located in the middle region of the image, the four corner vertices and the midpoints of the four sides of the image can be determined as the seed points of the background region, and an identifier is set for the background region. For example, the identifier can be set to 0.

[0044] Even further, in addition to the first foreground region obtained through the connected region analysis and the determined background region, the image also includes other regions, and the other regions are uncertain regions where it is uncertain whether the code to be recognized is included. An identifier is set for the other regions. For example, the identifier can be set to a negative number (for example, -1).

[0045] Thus, based on the connected region analysis method, the first foreground region, the background region, and the other regions are obtained. Each region has a corresponding identifier, and these identifiers and the pixels corresponding to the identifiers can form an identifier image with the same size as the image, and this identifier image is a shape marker identifier image for reflecting the shapes of the regions in the image.

[0046] Step 102, obtain the hue of the pixels in the image, and perform region division on the image according to the hue to obtain the second foreground region.

[0047] Among them, the second foreground region is the region composed of the pixels of the code to be recognized.

[0048] Exemplarily, perform a color space conversion on the original image to obtain the hue of the pixels in the image; among them, the original image is the original Red Green Blue (RGB) image. Specifically, perform a color space conversion on the original image to obtain an image based on the Hue-Saturation-Value (HSV) color model. Among the three channels of the HSV image, the H (Hue) channel represents the hue, and the hue H can be used to distinguish color types, and different colors can be effectively perceived based on the hue H.

[0049] According to different parts or functions of the code to be recognized, corresponding different colors can be set in the compiler, and different colors correspond to different color marker identifications. Further, after obtaining an image in the HSV color model, different preset hue thresholds are set for the H channel to distinguish different color regions in the code to be recognized according to the hue thresholds. For example, the hues of the background, comments, and the main body of the code in the code to be recognized are different, and corresponding preset hue thresholds are set for different hues respectively; then foreground seed points are set for different second foreground regions in the image. Further, the identifications of the seed points of different second foreground regions are different from each other and different from the identifications of the seed points of the first foreground region.

[0050] For example, the seed point identification used to identify the first foreground region in the shape marker is 1, and the positive integers used to identify regions such as comments and the main body of the code in the second foreground region need to be different from the positive integers in the shape marker. For example, they can be set to 2, 3 to N.

[0051] Further, the background seed points can be set by referring to the method in step 101, and other regions except the second foreground region and the background seed points are determined, and then negative numbers (such as -1) are set as the identifications for the pixel points in other regions. Based on the second foreground region, the background region, and other regions, a color marker identification image of the same size as the input image can be generated.

[0052] Step 103, fuse the first foreground region and the second foreground region to obtain a fused foreground region, and segment the image based on the fused foreground region to obtain at least one sub-image.

[0053] Among them, fusing the first foreground region and the second foreground region includes: if a pixel point belongs to at least one of the first foreground region and the second foreground region, it is determined as a pixel point in the fused foreground region, and an identification is set for it.

[0054] Specifically, if a pixel point belongs to both the first foreground region and the second foreground region, a new identification is set for it, where the new identification is different from the identifications of the first foreground region and the second foreground region. If a pixel point belongs to only one of the foreground regions, the identification of the foreground region to which the pixel point belongs is determined as the identification of the pixel point.

[0055] Further, if a pixel point is a pixel point in other regions, the identification of other regions is determined as the identification of the pixel point; if a pixel point is a pixel point in the background region, the identification of the background region is determined as the identification of the pixel point.

[0056] Exemplarily, the identifications of the fused foreground region, other regions, and the background region form a fused marker identification. The fused marker identification is input into the watershed segmentation model to obtain at least one sub-image.

[0057] Step 104: Input at least one sub-image into a pre-trained code recognition model to obtain the code to be recognized in the image.

[0058] Exemplarily, through the pre-trained code recognition model, visual feature extraction and semantic feature extraction are performed on the sub-image, and then the visual features and semantic features are fused. According to the fused features obtained by the fusion, the code to be recognized in the image is predicted.

[0059] Exemplarily, the code to be recognized in the sub-image can be extracted by Optical Character Recognition (OCR) technology. Among them, OCR technology can convert the text materials in the image file into electronic text.

[0060] In the related art, a general model that can be used to directly recognize characters in an image can be formed through OCR technology. Then, the image containing the text materials is input into the model, and the recognition content is output. In the related art, there is a lack of a text recognition model dedicated to code extraction (such as the hardware description language Verilog code), and the existing general recognition model is directly used to extract the code. There are at least the following problems: The format of the code is different from that of other texts. For example, the code usually has indents, comments, etc.; the general model has poor adaptability to the indents, comments, etc. in the code, resulting in low accuracy in obtaining the code using the general model; the general model lacks the understanding and learning of the semantics of the code, resulting in misrecognition of the symbolic elements in the code and neglect of the hierarchical structure. The general model is not applicable to code recognition. Therefore, in the related art, the code is usually input into the local computing device by manual input, which will cause problems of low efficiency and high error rate.

[0061] In this embodiment, connected component analysis is performed on an image including a code to be recognized to obtain a first foreground region; the first foreground region obtained thereby is the region of the code to be recognized determined from the shape level. The hue of the pixel points in the image is obtained, and the image is regionally divided according to the hue to obtain a second foreground region; the second foreground region obtained thereby is the region of the code to be recognized determined from the color level. The color of the code and the background color are usually different. Therefore, the first foreground region obtained based on connected component analysis and the second foreground region obtained based on the hue of the pixel points can accurately reflect the region where the code to be recognized is located, and the fused foreground region obtained by fusing the first foreground region and the second foreground region further improves the accuracy of reflecting the location of the code to be recognized. Further, the image is segmented based on the fused foreground region, and then the obtained sub-images are input into a pre-trained code recognition model, and the code to be recognized in the image can be accurately and efficiently obtained. This solves the problems of low input efficiency and easy error caused by manually inputting codes in the related art.

[0062] Figure 3 The flowchart of steps of another code recognition method provided by this embodiment is shown. Referring to FIG. 3, the method may include the following steps:

[0063] Step 201, obtain an image including a code to be recognized.

[0064] Exemplarily, the code to be recognized may be a Verilog code or other codes recorded in image format.

[0065] Step 202, perform grayscale processing on the image to obtain a grayscale image.

[0066] An image containing a code (such as a Verilog code) usually has the characteristics of a single background color, less noise, and an obvious color contrast between the foreground characters and the background.

[0067] For example, if the image is an RGB image, grayscale processing can be performed on the RGB image to further highlight the contrast between the foreground characters and the background in the image, thereby improving the accuracy of subsequent segmentation and recognition. In addition, grayscale processing can reduce the computational complexity of subsequent steps and improve the processing efficiency.

[0068] Further, grayscale processing is the process of converting a color image into a grayscale image. A color image usually contains color information of three channels: red (R), green (G), and blue (B). Among them, the grayscale value range of each channel is generally 0-255, while a grayscale image has only one channel, and the grayscale value is also between 0-255, representing different grayscale levels from black (0) to white (255).

[0069] For example, a weighted average method can be used to grayscale the image including the code. Specifically, corresponding weights are set for the R, G, and B channels, and the image is grayscaled according to the set weights. The weights can be set according to the sensitivity of the human eye to color. For example, the weights of the R, G, and B channels can be set to 0.299, 0.578, and 0.114, respectively.

[0070] Step 203, performing image binarization processing on the grayscale image to obtain a binary image;

[0071] Specifically, according to analysis, images including code (e.g., Verilog code) usually have a single background color, and there is a large difference between the background color and the foreground text color. Based on different user settings, code characters with different functions may appear in different colors. For example, the color of the comment part in the code can be set to red, blue, etc. Through binarization processing, the foreground characters and background of the image can be more clearly distinguished.

[0072] Furthermore, in order to ensure the integrity of the text after binarization, the adaptive threshold method can be used to simplify the grayscale image. Binarization is an image processing technology used to convert the grayscale values ​​of pixels in an image into only two values, which can be 0 and 255, where grayscale values ​​0 and 255 represent black and white, respectively.

[0073] Specifically, the grayscale image is divided into multiple sub-regions, and for each sub-region: according to the grayscale value of the pixel in the sub-region and the grayscale value threshold set for the sub-region, the grayscale value in the sub-region is binarized. The grayscale value threshold is equivalent to the binarization threshold for binarization. For example, if the grayscale value of the pixel is greater than the grayscale value threshold, it is directly set to 255, otherwise it is set to 0.

[0074] Furthermore, the codes are usually distributed in straight lines in the image, so the image is divided into rectangular sub-regions, and then the average grayscale value of all pixels in the sub-region is calculated, and the average grayscale value threshold corresponding to the sub-region is determined by the average grayscale value.

[0075] Step 204: Perform connected region analysis on the binary image to obtain a first foreground region.

[0076] The first foreground area is a connected area where the code to be identified is located.

[0077] The connected region analysis is used to mark interconnected pixels as a connected domain to determine the foreground character region. The foreground character region obtained based on the connected region analysis is the first foreground region in this step.

[0078] Further, after the connected region analysis is completed, the centroid of each connected region will be output. The centroid is the seed point of the first foreground region, and this seed point is the shape maker of the first foreground region.

[0079] Step 205: Obtain the hue of the pixel points in the image, and divide the image according to the hue to obtain the second foreground region.

[0080] Among them, the second foreground region is the region composed of the pixel points of the code to be recognized.

[0081] Exemplarily, obtaining the hue of the pixel points in the image in step 205 may include: performing a color space conversion on the image to obtain the hue of the pixel points in the image.

[0082] Step 206: Fuse the first foreground region and the second foreground region to obtain a fused foreground region;

[0083] The first foreground region and the second foreground region respectively have a shape marker and a color marker. By fusing the shape marker and the color marker, the fusion of the first foreground region and the second foreground region is realized to obtain the foreground seed point corresponding to the fused foreground region.

[0084] Exemplarily, the pixel points in the fused foreground region have a first identifier; the pixel points in the first foreground region have a corresponding first preset identifier, and the pixel points in the second foreground region have a corresponding second preset identifier. Among them, the following method can be used to obtain the corresponding second preset identifier of the pixel points in the second foreground region: If the hue reaches the preset hue threshold, then according to the corresponding relationship between the hue threshold and the identifier, obtain the identifier corresponding to the preset hue threshold; determine the identifier corresponding to the preset hue threshold as the second preset identifier of the pixel point.

[0085] During the programming process, different colors can be set for codes with different functions. For example, the code body and code comments are respectively set to white and blue. Among them, different colors respectively have corresponding preset hue thresholds. If the hue of the pixel point reaches the preset hue threshold, it means that the pixel point belongs to the color corresponding to the preset hue threshold. Among them, different colors correspond to different identifiers. Correspondingly, the preset hue thresholds corresponding to different colors also have different identifiers. According to the method of this embodiment, corresponding second preset identifiers can be respectively set for the pixel points corresponding to the codes in the second foreground region.

[0086] The first identifier of the pixel points in the fused foreground region can be obtained according to sub-steps A1 to A2:

[0087] Sub-step A1: If the pixel point in the image belongs to the first foreground region and the second foreground region, obtain the maximum value of the first preset identifier and the second preset identifier, and set the sum result of the maximum value and the preset value as the first identifier of the pixel point.

[0088] For a pixel point in the image, if the pixel point includes the first preset identifier and the second preset identifier, it is determined that the pixel point belongs to the first foreground region and the second foreground region.

[0089] For example, the first preset identifier of the pixel point in the first foreground region is 1, and the second preset identifiers of the pixel points in the second foreground region include 2, 3 to N. In one embodiment, if the identifiers of the pixel point are 1 and 3, then the sum result of the maximum value N of the first preset identifier (1) and the second preset identifiers (2, 3 to N) and the preset value (for example, 1), which is N + 1, is determined as the first identifier of the pixel point.

[0090] Sub-step A2: If the pixel point in the image belongs to the first foreground region or the second foreground region, determine the preset identifier of the foreground region to which the pixel point belongs as the first identifier of the pixel point.

[0091] For example, if the pixel point has the first preset identifier 1 and the identifier -1 of other regions, it is determined that the pixel point belongs to the first foreground region, then retain its identifier in the first foreground region and set the identifier 1 as its first identifier. For example, if the pixel point has the second preset identifier N and the identifier -1 of other regions, it is determined that the pixel point belongs to the first foreground region, then retain its identifier in the first foreground region and set the identifier N as its first identifier.

[0092] Step 207: Perform edge detection on the image to obtain an edge detection result;

[0093] The font of the code characters in the image is usually relatively thin. Therefore, before performing edge detection on the preprocessed image, perform morphological processing on it first. For example, morphological processing can be performed on it through closing operation. Through the closing operation, small holes and broken parts in the characters can be filled, preventing small characters from being overly divided, making adjacent characters easier to be recognized as a whole, and thus improving the accuracy of the edge detection result.

[0094] For example, step 207 may include sub-steps B1 to B3:

[0095] Sub-step B1: Perform grayscale processing on the image to obtain a grayscale image.

[0096] The method of this sub-step can refer to the description of the aforementioned step 202 and will not be elaborated here.

[0097] Sub-step B2: Perform image binarization processing on the grayscale image to obtain a binarized image.

[0098] The method of this sub-step can refer to the description of the aforementioned step 203 and will not be elaborated here.

[0099] Sub-step B3: Perform edge detection on the binary image to obtain the edge detection result.

[0100] Exemplarily, according to the gray values of each pixel point in the binary image, perform edge detection on the binary image to obtain the edge detection result.

[0101] Step 208: Obtain the second identifier of the pixel points in the background area of the image and the third identifier of the pixel points in other areas;

[0102] Other areas refer to the areas in the image other than the fused foreground area and the background area.

[0103] Among them, the pixel points in the background area have the second identifier, and the pixel points in other areas have the third identifier;

[0104] This step also includes obtaining background seed points. Specifically, generate background seed points based on the speculated background position. Exemplarily, the four corner vertices and the midpoints of the four sides of the image can be selected as background seed points.

[0105] Furthermore, the foreground seed points, background seed points, and other pixel points except the foreground seed points and background seed points can be identified with different numerical values. Exemplarily, the foreground seed points and background seed points are respectively identified as different positive integers. For example, they can be respectively identified as 1 and 0; the remaining pixel points can be identified as a negative number. For example, they can be identified as -1.

[0106] The numerical identifications of the foreground seed points, background seed points, and other pixel points can form a marker identification image with the same size as the input image.

[0107] Exemplarily, when obtaining the first foreground area by processing the image based on the connected component analysis method, set the first preset identifier (for example, identified as 1) for the first foreground area, and also set identifiers for the background area and other areas. The identifier set for the background area is 0, and the identifier set for other areas is a negative number (for example, -1); when obtaining the second foreground area based on the hue and setting the second preset identifier (for example, 2, 3...), also set identifiers for the background area and other areas except the second foreground area and the background area. For example, they are respectively set as 0 and a negative number (for example, -1).

[0108] The second identifier of the pixel points in the background area of the image and the third identifier of the pixel points in other areas can be obtained according to the following method:

[0109] After fusing the shape marker and the color marker, if a pixel point in the image is marked as 0 in both markers, it indicates that the pixel point is a pixel point in the background area, and 0 is determined as the second identifier for other areas.

[0110] After fusing the shape marker and the color marker, if a pixel point in the image is marked as a negative number in both markers, it indicates that the pixel point is a pixel point in other areas, and the negative number (for example, -1) in the foregoing embodiment is determined as the third identifier for other areas.

[0111] Step 209: Input the first identifier, the second identifier, the third identifier, and the edge detection result into the watershed segmentation model to obtain at least one sub-image.

[0112] Specifically, when performing image segmentation through the watershed segmentation model, seed points need to be set and the image gradient needs to be calculated. Further, the watershed algorithm is a method for image segmentation based on the topological structure of the image. The Canny operator is an algorithm for edge detection. The Canny operator is specifically used to find edges based on the characteristic that the gray value changes violently at the edges in the image. In this embodiment, the Canny operator can be used to perform edge detection on the preprocessed image for the code area to generate a gradient image, and use it as the boundary input for the watershed algorithm.

[0113] Exemplarily, step 209 may include sub-steps C1 to C2:

[0114] Sub-step C1: Input the first identifier, the second identifier, the third identifier, and the edge detection result into the watershed segmentation model to obtain at least one sub-region.

[0115] Further, according to the first identifier, the second identifier, and the third identifier, a seed point marking marker with the same size as the image can be obtained. This marker is the fused marker. Input the fused marker and the edge detection result into the watershed segmentation model to obtain at least one sub-region.

[0116] Sub-step C2: If the area of the sub-region is less than or equal to the preset region area threshold, merge the sub-region and the sub-region adjacent to it, and determine the image part corresponding to the merged region as the sub-image.

[0117] The number of pixel points in the sub-region can be obtained, and the number of pixel points is used to represent the area of the sub-region.

[0118] Before performing the merging, it is necessary to confirm the adjacency relationship of the segmentation results. Specifically, in the segmentation results of the watershed algorithm, the same region is marked with the same region identifier, and the identifiers between different regions are different. All pixels in the segmentation results can be traversed. If the region identifier of a pixel is different from that of its surrounding pixels, it can be determined that the region where the pixel is located is adjacent to the regions where the surrounding pixels are located.

[0119] Furthermore, sub-step C2 may include the following sub-steps:

[0120] Sub-step C21, obtaining the difference degree value between the gray values of the pixels in the sub-region and the adjacent sub-region.

[0121] Obtain the color Euclidean distance between the gray values of the pixels in the sub-region and the gray values of the pixels in the adjacent sub-region, and determine the color Euclidean distance as the difference degree value. Specifically, obtain the RGB values of the pixels, and use the R values, B values, and G values of the pixels in the two sub-regions as the input parameters of the Euclidean distance function to obtain the color Euclidean distance between the gray values of the pixels in the sub-region and the gray values of the pixels in the adjacent sub-region.

[0122] Sub-step C22, if the area of the sub-region is less than the preset region area threshold, or the difference degree value is less than or equal to the preset difference degree threshold, then merge the sub-region and the sub-region adjacent to the sub-region.

[0123] Through region merging, over-segmentation can be reduced, and the situation of splitting the same character or the same line of characters into different sub-regions can be avoided, thus avoiding the problem of low code prediction accuracy caused by code prediction based on sub-regions.

[0124] Further, merging the sub-region and the sub-region adjacent to the sub-region in sub-step C2, and determining the obtained merged region as a sub-image may include the following sub-steps:

[0125] Sub-step C23, obtaining the region contour of the merged region.

[0126] Exemplarily, the corner points of the merged region can be obtained, and the region contour of the merged region can be judged based on the corner points.

[0127] Sub-step C24, if the region contour is other than a rectangle, then obtain the circumscribed rectangle of the merged region, and determine the image part corresponding to the circumscribed rectangle as the sub-image.

[0128] Codes are usually distributed in straight lines in the image. Therefore, the ideal result of segmentation is to split the codes into long and straight rectangles by rows. Therefore, after region merging, obtain the external contour of each merged region and determine it as the region contour of the merged region.

[0129] If the external contour is not rectangular, the minimum circumscribed matrix is fitted according to the corner points of the merged region, and the four vertices of the rectangle are output. The rectangle determined by the four vertices is applied to the original image, and the image part corresponding to the rectangle region is the final sub-image obtained.

[0130] Further, in sub-step C2, merging the sub-region and the sub-regions adjacent to the sub-region, and determining the image part corresponding to the obtained merged region as the sub-image may include the following sub-steps:

[0131] Sub-step C25, merging the sub-region and the sub-regions adjacent to the sub-region, and performing image enhancement processing on the merged region to obtain an enhanced image.

[0132] Exemplarily, methods such as histogram equalization and sharpening can be adopted to perform image enhancement processing on the merged region. Through histogram equalization processing, the contrast between the bright and dark regions of the image can be enhanced, making the background and characters easier to distinguish; through sharpening processing, the high-frequency regions such as edges and details in the image can be enhanced, making the character edges clearer. Exemplarily, a high-pass filter sharpening method can be used to perform sharpening processing on the image.

[0133] Sub-step C26, performing size normalization processing on the enhanced image to obtain the sub-image.

[0134] Exemplarily, a bilinear interpolation algorithm can be adopted to perform size normalization processing on the enhanced image.

[0135] Step 210, extracting the image features and semantic features of the sub-image;

[0136] Exemplarily, a text recognition model (Convolutional Recurrent Neural Network, CRNN) can be adopted to perform feature extraction on the input image to obtain visual features. Specifically, referring to Figure 4 , the visual features of the image are extracted through the convolutional layer (Convolutional Layer). Then, the label distribution of the feature vector is predicted through the recurrent and transcription layers (Recurrent LayerTranscription Layer) and the output results are integrated. Further, the visual features are extracted through the convolutional layer, and the integrated results output by the recurrent layer and the transcription layer are used for semantic feature extraction. Further, attention feature fusion is performed on the output structure of the convolutional layer and the semantic features obtained by semantic feature extraction, and the fusion result is input to the output layer for prediction to obtain the predicted code.

[0137] Specifically, CRNN is a deep learning model for image sequence recognition, mainly composed of two parts: a Convolutional Neural Network (CNN) and a Recurrent Neural Network (RNN). It combines the feature extraction ability of CNN with the sequence modeling ability of RNN to process image data with sequence features, such as text images, speech spectrograms, etc.

[0138] In this embodiment, the extraction of semantic features is achieved based on the preliminary prediction results of the visual part. Through the input representation and encoder layer of the pre-trained model (Bidirectional Encoder Representations from Transformers, BERT), visual features are extracted. Among them, the BERT model is a pre-trained language model based on the Transformer architecture. The pre-trained BERT model is based on large-scale data training and has strong robustness and generalization ability. Based on the pre-trained BERT model, downstream tasks can converge quickly based on the initial parameters during the fine-tuning stage.

[0139] Step 211: Perform feature fusion on the image features and semantic features to obtain fused features.

[0140] Feature fusion is used to establish cross-modal feature associations between images and semantics. The convolutional features of the visual model and the features extracted from the semantic sequence after fusion are used for the final prediction.

[0141] Exemplarily, an attention mechanism is used for feature fusion. Specifically, the visual features are input as the query feature Q, and the semantic features are input as the key feature K and value feature V. Through softmax(QK T / )V+H V feature fusion is performed.

[0142] Where Q = W Q H V , K = W K H L , V = W V H L , W Q , W K and W V are learnable weight matrices respectively, H V and H L are visual features and semantic features respectively; is the vector dimension.

[0143] Feature fusion using the attention mechanism enables the model to focus on the parts most relevant to the current task when processing a large amount of information (such as multiple elements in sequence data). In the field of deep learning, the attention mechanism is mainly used to process sequential or structured data such as sentences in natural language processing and pixel regions in images.

[0144] Step 212, perform code prediction based on the fused features to obtain the code to be recognized in the image

[0145] Exemplarily, a fully connected layer is used to predict the code to be recognized. The fully connected layer is the same as the output layer of BERT. Specifically, the prediction is made through softmax(hW + b), where W is the weight matrix and b is the bias vector.

[0146] Furthermore, semantic fine-tuning is performed on the output layer and the semantic feature extraction model.

[0147] Specifically, first manually generate Verilog signals, and ensure that the content of the Verilog signals fully complies with the code syntax rules. When naming the signals, use common abbreviations of the naming rules, such as set, clr (indicating clear), clk (indicating clock), rst / rstn (indicating reset or reset - negative), addr (indicating address), req (indicating request), ack (indicating acknowledge), etc., and use underscores for separation.

[0148] When training the model, randomly select 15% of the generated Verilog signals as training data, mark some internal characters, and apply masking to these characters, or replace these characters with incorrect characters for semantic training of the model. Among them, the purpose of semantic training is to predict the original characters that are masked or replaced. On the basis of the aforementioned semantic training, train with a small amount of labeled image data.

[0149] In this embodiment, methods of image processing and deep learning are adopted to realize the automatic recognition of code signals based on images, avoiding the repetitive labor and human errors caused by manual signal integration, and improving work efficiency. This method combines the visual features of images and the semantic features of signal sequences in model design, has better feature representation for Verilog code, has a higher recognition accuracy compared with existing general OCR recognition models, and has a stronger perception ability for symbols, abbreviations, formats, etc. in signals. This method will segment different signal regions to achieve batch recognition of multiple signals, without the need to input multiple times according to single - signal images, and has higher work efficiency.

[0150] The following refers to Figure 5, for further exemplary illustration of the method of the present application, refer to Figure 5 , the method may include the following steps:

[0151] Step S1, preprocess the image of the Verilog code to obtain a preprocessed image. Among them, the preprocessing process includes grayscale processing and binary simplification processing.

[0152] The method of this step has been described in the foregoing steps 202 to 203, and will not be elaborated here.

[0153] Further, step S1 may include the following sub-steps:

[0154] Sub-step S11, perform grayscale processing on the image of the Verilog code to obtain a grayscale image.

[0155] The method of this step has been described in the foregoing step 202, and will not be elaborated here.

[0156] Sub-step S12, perform binary simplification on the grayscale image.

[0157] The method of this step has been described in the foregoing step 203, and will not be elaborated here.

[0158] Step S2, perform morphological processing, connected component analysis, and edge detection on the preprocessed image to obtain a shape marker identification.

[0159] Exemplarily, perform morphological processing on the preprocessed image through closing operation.

[0160] Sub-step S21, perform edge detection processing on the preprocessed image to obtain the edge detection processing result.

[0161] The method of this step can refer to the description of the foregoing step 208, and will not be elaborated here.

[0162] Sub-step S22, perform connected component analysis processing on the morphological processing result, and numerically identify the foreground region, background region, and other regions to obtain a shape marker identification.

[0163] The method of this step can refer to the description of the foregoing steps 204 to 207, and will not be elaborated here.

[0164] Step S3, perform color space conversion on the original image to obtain a color marker identification.

[0165] The method of this step can refer to the description of the foregoing step 205, and will not be elaborated here.

[0166] Step S4, fuse the shape marker identification and the color marker identification to obtain the final marker.

[0167] For the method in this step, reference can be made to the description of the aforementioned step 206, which will not be elaborated here.

[0168] Exemplarily, the fusion method may include the following sub-steps:

[0169] Sub-step S41, if the pixel points are both identified as foreground seed points at the corresponding points of the shape marker and the color marker, then they are identified as a new positive integer (such as N + 1).

[0170] Wherein, N is the maximum value in the shape marker and the color marker.

[0171] Sub-step S42, if the pixel point is only identified as foreground in one marker and is identified as negative in the other marker, then retain the foreground identification.

[0172] For example, if the pixel point is marked as foreground in the shape marker, the mark in the shape marker is 1, and is marked as negative in the color marker, then the identification of the pixel point is retained as the mark 1 in the shape marker.

[0173] Sub-step S43, if the pixel points are both identified as negative in the two markers, then retain the negative identification.

[0174] Sub-step S44, if the pixel points are both identified as 0 in the two markers, then retain this identification.

[0175] Being identified as 0 in the marker indicates that the pixel point is a background seed point. Since the background seed points in the two markers are set the same, there is no conflict and it is still retained.

[0176] According to the methods shown in sub-step S41 to sub-step S44, a new seed point identification marker with the same size as the input image can be generated, that is, the final marker.

[0177] Step S5, use the edge detection result and the final marker as inputs for watershed segmentation, and output the preliminary region segmentation result.

[0178] The segmentation result in this step is the sub-region obtained by segmentation in the aforementioned embodiment.

[0179] Step S6, perform post-processing on the segmented image, specifically including region merging and shape regularization.

[0180] When merging regions, it is necessary to confirm the adjacency relationship of pixel points. After confirming the adjacency relationship, the adjacency relationship can be stored in the form of a dictionary, such as {a: b1, b2, …}, indicating that sub-regions b1 and b2 are adjacent regions of sub-region a.

[0181] Furthermore, it is possible to determine whether to merge with neighboring regions from two perspectives: the area of the segmented region and the average color of the code. The specific steps are as follows:

[0182] Sub-step S61: Traverse the pixel points, calculate the area of each segmented region with a label, and set an area threshold. Regions smaller than the threshold will be merged into adjacent regions.

[0183] Exemplarily, the area threshold can be set according to user requirements. For instance, it can be set to 50.

[0184] Sub-step S62: Based on the original RGB image, calculate the Euclidean distance of colors between adjacent regions. If the distance is less than the set color threshold, the two regions will be merged.

[0185] According to the above two merging strategies, traverse all labeled regions and update the adjacency relationship dictionary. Furthermore, the merging process can be repeated to avoid over-segmented regions still existing after a single merge.

[0186] After region merging, the external contour of each region will be obtained, and on this basis, the minimum bounding rectangle will be fitted and the four vertices of the rectangle will be output. Project the rectangle represented by the four vertices onto the original image, and the region corresponding to the rectangle in the original image is the final result of the segmented region.

[0187] Step S7: Perform image enhancement and normalization on the segmented image.

[0188] Through the image segmentation, segmented region enhancement, and size normalization in the foregoing steps, the signal region can be highlighted.

[0189] Specifically, the segmented sub-images will be sequentially subjected to image enhancement and size normalization to meet the requirements of the recognition model and improve the recognition accuracy.

[0190] Step S8: Adopt a method combining visual features and semantic features for code character recognition.

[0191] The method in this step has been described in the foregoing steps 210 to 211 and will not be elaborated here.

[0192] Specifically, combine Figure 6, preprocess the original image to obtain a preprocessed image, perform image segmentation and segmentation region enhancement on the preprocessed image to obtain sub-images with enhanced images, normalize the size of the sub-images to obtain processed sub-images. Extract visual features from the processed sub-images, perform text prediction and semantic feature extraction, fuse the visual features and semantic features, input the fusion result into the output layer, and perform code prediction through the output layer to obtain the recognized code.

[0193] Step S8 may include the following sub-steps:

[0194] Sub-step S81, extract visual features from the image with size normalization through a visual feature model and output the recognition result.

[0195] Specifically, perform text prediction based on the extracted visual features to obtain the recognition result.

[0196] Sub-step S82, input the recognition result into the semantic model and extract semantic features.

[0197] Among them, the semantic model can be a BERT model.

[0198] Sub-step S83, fuse the semantic features and the initially extracted visual features through an attention fusion mechanism.

[0199] Sub-step S84, input the fused features into the output layer to generate a prediction result with semantic correction.

[0200] In this embodiment, first, an image containing Verilog signal definitions or declarations, such as a signal screenshot of the code interface, needs to be input into the system. The system will automatically preprocess, segment, and recognize the input image. After the processing is completed, the recognition result of the signal will be directly output.

[0201] Among them, the input image may contain multiple signal instances. This system supports multiple signals in the image, and different signals in the recognition result will be listed separately by row.

[0202] Refer to Figure 7 , obtain the preprocessed image, perform morphological processing and connected component analysis on the preprocessed image in sequence, and obtain the shape marker; perform edge detection on the preprocessed image to obtain the edge detection result. Perform color space conversion on the original image to obtain the color marker. Fuse the shape marker and the color marker, and use the watershed segmentation method to segment the image according to the edge detection result and the fused marker, and perform post-processing on the segmented image, and then perform image enhancement and normalization processing on the image after post-processing, thereby obtaining the processed image. Combine Figure 6, perform model recognition processing on the processed image to identify the code to be recognized in the image. In this embodiment, for the signal extraction scenario of code (such as Verilog code), a text recognition method based on images is proposed, and efficient and highly accurate code signal recognition is achieved through steps such as image preprocessing, signal region positioning, visual feature extraction, and semantic feature extraction.

[0203] Reference Figure 8 , which shows a code recognition device provided by an embodiment of the present application. The device 30 includes:

[0204] The first acquisition module 301 is used to acquire an image including the code to be recognized, and perform connected region analysis on the image to obtain a first foreground region; the first foreground region includes the connected region where the code to be recognized is located;

[0205] The second acquisition module 302 is used to acquire the hue of the pixel points in the image, and divide the image according to the hue to obtain a second foreground region; the second foreground region includes the region composed of the pixel points of the code to be recognized;

[0206] The third acquisition module 303 is used to fuse the first foreground region and the second foreground region to obtain a fused foreground region, and segment the image based on the fused foreground region to obtain at least one sub-image;

[0207] The fourth acquisition module 304 is used to input at least one sub-image into a pre-trained code recognition model to obtain the code to be recognized in the image.

[0208] Exemplarily, the pixel points in the fused foreground region have a first identifier; the third acquisition module 303 may include:

[0209] The first acquisition sub-module is used to perform edge detection on the image to obtain an edge detection result;

[0210] The second acquisition sub-module is used to acquire a second identifier of the pixel points in the background region of the image and a third identifier of the pixel points in other regions; other regions are the regions in the image except the fused foreground region and the background region;

[0211] The third acquisition sub-module is used to input the first identifier, the second identifier, the third identifier, and the edge detection result into a watershed segmentation model to obtain at least one sub-image.

[0212] Optionally, the pixel points in the first foreground region have a corresponding first preset identifier, and the pixel points in the second foreground region have a corresponding second preset identifier; the device 30 further includes:

[0213] A fifth acquisition module, configured to, if a pixel point in an image belongs to a first foreground region and a second foreground region, acquire the maximum value between a first preset identifier and a second preset identifier, and set the sum result of the maximum value and a preset value as the first identifier of the pixel point;

[0214] A sixth acquisition module, configured to, if a pixel point in an image belongs to the first foreground region or the second foreground region, determine the preset identifier of the foreground region to which the pixel point belongs as the first identifier of the pixel point.

[0215] Optionally, the second acquisition module 302 includes:

[0216] A fourth acquisition sub-module, configured to perform color space conversion on the image to obtain the hue of the pixel points in the image;

[0217] A fifth acquisition sub-module, configured to, if the hue reaches a preset hue threshold, acquire the identifier corresponding to the preset hue threshold according to the correspondence between the hue threshold and the identifier;

[0218] A first determination sub-module, configured to determine the identifier corresponding to the preset hue threshold as the second preset identifier of the pixel point.

[0219] Optionally, the third acquisition sub-module includes:

[0220] A first acquisition unit, configured to input the first identifier, the second identifier, the third identifier, and the edge detection result into a watershed segmentation model to obtain at least one sub-region;

[0221] A first merging unit, configured to, if the area of the sub-region is less than or equal to a preset area threshold of the region, merge the sub-region and the sub-region adjacent to the sub-region, and determine the obtained merged region as a sub-image.

[0222] Optionally, the first merging unit includes:

[0223] A first acquisition sub-unit, configured to acquire the difference degree value between the gray values of the pixel points in the sub-region and the adjacent sub-region;

[0224] A first merging sub-unit, configured to, if the area of the sub-region is less than the preset area threshold of the region, or the difference degree value is less than or equal to a preset difference degree threshold, merge the sub-region and the sub-region adjacent to the sub-region.

[0225] Optionally, the first merging unit includes: includes:

[0226] A second acquisition sub-unit, configured to acquire the region contour of the merged region;

[0227] A second merging sub-unit, configured to, if the region contour is other than a rectangle, acquire the circumscribed rectangle of the merged region, and determine the region corresponding to the circumscribed rectangle as the sub-image.

[0228] Optionally, the first merging unit includes:

[0229] A first processing subunit, configured to merge a sub-region and a sub-region adjacent to the sub-region, perform image enhancement processing on the merged region, and obtain an enhanced image;

[0230] A second processing subunit, configured to perform size normalization processing on the enhanced image to obtain a sub-image.

[0231] Optionally, the first obtaining sub-module includes:

[0232] A second obtaining unit, configured to perform grayscale processing on the image to obtain a grayscale image;

[0233] A second obtaining unit, configured to perform image binarization processing on the grayscale image to obtain a binary image;

[0234] A third obtaining unit, configured to perform edge detection on the binary image to obtain an edge detection result.

[0235] Optionally, the fourth obtaining module 304 includes:

[0236] An extraction sub-module, configured to extract image features and semantic features of the sub-image;

[0237] A feature fusion sub-module, configured to perform feature fusion on the image features and semantic features to obtain a fusion feature;

[0238] A sixth obtaining sub-module, configured to perform code prediction based on the fusion feature to obtain a code to be recognized in the image.

[0239] Optionally, the first obtaining module 301 obtains a first foreground region, including:

[0240] A seventh obtaining sub-module, configured to perform grayscale processing on the image to obtain a grayscale image;

[0241] An eighth obtaining sub-module, configured to perform image binarization processing on the grayscale image to obtain a binary image;

[0242] A ninth obtaining sub-module, configured to perform connected region analysis on the binary image to obtain a first foreground region.

[0243] In this embodiment, connected component analysis is performed on an image including a code to be recognized to obtain a first foreground region. The first foreground region obtained thereby is the region of the code to be recognized determined from the shape level. The hue of the pixel points in the image is obtained, and the image is regionally divided according to the hue to obtain a second foreground region. The second foreground region obtained thereby is the region of the code to be recognized determined from the color level. The color of the code is usually different from the background color. Therefore, the first foreground region obtained based on connected component analysis and the second foreground region obtained based on the hue of the pixel points can accurately reflect the region where the code to be recognized is located, and the fused foreground region obtained by fusing the first foreground region and the second foreground region further improves the accuracy of reflecting the location where the code to be recognized is located. Further, the image is segmented based on the fused foreground region, and then the obtained sub-images are input into a pre-trained code recognition model, and the code to be recognized in the image can be accurately and efficiently obtained. This solves the problems of low input efficiency and easy error caused by manually inputting the code in the related art.

[0244] Figure 9 FIG. 4 is a block diagram of an electronic device 400 shown according to an exemplary embodiment. For example, the electronic device 400 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0245] Referring to Figure 9 FIG. 4, the electronic device 400 may include one or more of the following components: a processing component 402, a memory 404, a power component 406, a multimedia component 408, an audio component 410, an input / output (I / O) interface 412, a sensor component 414, and a communication component 416.

[0246] The processing component 402 generally controls the overall operation of the electronic device 400, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 402 may include one or more processors 420 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 402 may include one or more modules to facilitate the interaction between the processing component 402 and other components. For example, the processing component 402 may include a multimedia module to facilitate the interaction between the multimedia component 408 and the processing component 402.

[0247] The memory 404 is used to store various types of data to support the operation of the electronic device 400. Examples of such data include instructions for any application or method operating on the electronic device 400, contact data, phone book data, messages, images, multimedia, and the like. The memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0248] The power supply component 406 provides power to various components of the electronic device 400. The power supply component 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 400.

[0249] The multimedia component 408 includes a screen that provides an output interface between the electronic device 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 408 includes a front camera and / or a rear camera. When the electronic device 400 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0250] The audio component 410 is used to output and / or input audio signals. For example, the audio component 410 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 400 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 404 or transmitted via the communication component 416. In some embodiments, the audio component 410 further includes a speaker for outputting audio signals.

[0251] The I / O interface 412 provides an interface between the processing component 402 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a start button, and a lock button.

[0252] The sensor assembly 414 includes one or more sensors for providing status assessments of various aspects for the electronic device 400. For example, the sensor assembly 414 can detect the on / off state of the electronic device 400, the relative positioning of components, such as components for the display and keypad of the electronic device 400. The sensor assembly 414 can also detect a change in the position of the electronic device 400 or a component of the electronic device 400, the presence or absence of user contact with the electronic device 400, the orientation or acceleration / deceleration of the electronic device 400, and the temperature change of the electronic device 400. The sensor assembly 414 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 414 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 414 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0253] The communication component 416 is used to facilitate communication between the electronic device 400 and other devices in a wired or wireless manner. The electronic device 400 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 416 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0254] In an exemplary embodiment, the electronic device 400 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for implementing a code recognition method provided by the embodiments of the present application.

[0255] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions. The above instructions can be executed by the processor 420 of the electronic device 400 to complete the above method. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0256] Figure 10is a block diagram of an electronic device 500 shown according to an exemplary embodiment. For example, the electronic device 500 may be provided as a server. Referring to Figure 10 , the electronic device 500 includes a processing component 522, which further includes one or more processors, and memory resources represented by a memory 532 for storing instructions executable by the processing component 522, such as application programs. The application programs stored in the memory 532 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 522 is configured to execute instructions to perform a code recognition method provided by an embodiment of the present application.

[0257] The electronic device 500 may further include a power component 526 configured to perform power management of the electronic device 500, a wired or wireless network interface 550 configured to connect the electronic device 500 to a network, and an input / output (I / O) interface 558. The electronic device 500 may operate based on an operating system stored in the memory 532, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM or the like.

[0258] An embodiment of the present application also provides a computer program product, including a computer program, which implements a code recognition method when executed by a processor.

[0259] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0260] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

[0261] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0262] The above has introduced in detail a code recognition method, device, electronic device and computer-readable storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A code recognition method, characterized in that: include: Acquire an image including a code to be identified, and perform connected region analysis on the image to obtain a first foreground region; The first foreground area includes a connected area where the code to be identified is located; Acquire the hue of the pixel points in the image, and divide the image according to the hue to obtain a second foreground area; the second foreground area includes an area composed of the pixel points of the code to be identified; Fusing the first foreground area and the second foreground area to obtain a fused foreground area, and performing edge detection on the image to obtain an edge detection result; obtaining a second identifier of pixel points in a background area of ​​the image, and a third identifier of pixel points in other areas; the other areas are areas in the image other than the fused foreground area and the background area; Inputting the first identifier, the second identifier, the third identifier and the edge detection result into a watershed segmentation model to obtain at least one sub-image; the pixel points in the fused foreground area have the first identifier; At least one of the sub-images is input into a pre-trained code recognition model to obtain a code to be recognized in the image.

2. The method according to claim 1, characterized in that The pixels in the first foreground area have a corresponding first preset identifier, and the pixels in the second foreground area have a corresponding second preset identifier; the method further includes: If a pixel point in the image belongs to the first foreground area and the second foreground area, obtaining a maximum value between the first preset identifier and the second preset identifier, and setting a sum of the maximum value and the preset value as the first identifier of the pixel point; If a pixel point in the image belongs to the first foreground area or the second foreground area, a preset identifier of the foreground area to which the pixel point belongs is determined as a first identifier of the pixel point.

3. The method according to claim 1, characterized in that Inputting the first identifier, the second identifier, the third identifier and the edge detection result into a watershed segmentation model to obtain at least one sub-image includes: Inputting the first identifier, the second identifier, the third identifier and the edge detection result into a watershed segmentation model to obtain at least one sub-region; If the area of ​​the sub-region is less than or equal to a preset area threshold, the sub-region and a sub-region adjacent to the sub-region are merged, and the obtained merged region is determined as the sub-image.

4. The method according to claim 3, characterized in that If the area of ​​the sub-region is less than or equal to a preset area threshold, merging the sub-region with a sub-region adjacent to the sub-region includes: Obtaining a difference value between the grayscale values ​​of pixels in the sub-region and the adjacent sub-region; If the area of ​​the sub-region is less than or equal to the preset area threshold, or the difference value is less than or equal to the preset difference threshold, the sub-region and the sub-region adjacent to the sub-region are merged.

5. The method according to claim 3, characterized in that: The merging of the sub-region and a sub-region adjacent to the sub-region, and determining the obtained merged region as the sub-image, includes: Obtaining a region outline of the merged region; If the region outline is an outline other than a rectangle, the bounding rectangle of the merged region is obtained, and the region corresponding to the bounding rectangle is determined as the sub-image.

6. The method according to claim 1, characterized in that Inputting at least one of the sub-images into a pre-trained code recognition model to obtain a code to be recognized in the image, comprising: Extracting image features and semantic features of the sub-image; Performing feature fusion on the image feature and the semantic feature to obtain a fused feature; Code prediction is performed based on the fusion features to obtain the code to be recognized in the image.

7. A code recognition device, characterized in that: The device comprises: A first acquisition module is used to acquire an image including a code to be identified, and perform a connected region analysis on the image to obtain a first foreground region; the first foreground region includes a connected region where the code to be identified is located; A second acquisition module is used to acquire the hue of the pixel points in the image, and divide the image according to the hue to obtain a second foreground area; the second foreground area includes an area composed of the pixel points of the code to be identified; A third acquisition module, configured to fuse the first foreground area and the second foreground area to obtain a fused foreground area, and segment the image based on the fused foreground area to obtain at least one sub-image; A fourth acquisition module, used for inputting at least one of the sub-images into a pre-trained code recognition model to obtain a code to be recognized in the image; The pixel points in the fused foreground area have a first identifier; and the third acquisition module includes: The first acquisition submodule is used to perform edge detection on the image to obtain edge detection results; The second acquisition submodule is used to acquire the second identifier of the pixel points in the background area of ​​the image and the third identifier of the pixel points in other areas; the other areas are the areas in the image except the fused foreground area and the background area; The third acquisition submodule is used to input the first identifier, the second identifier, the third identifier and the edge detection result into a watershed segmentation model to obtain at least one sub-image.

8. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • RGBD-based portrait segmentation method and device

    CN113139983A

  • Character recognition method and device, equipment and storage medium

    CN113792741A