Method for identifying character information in electronic component symbol diagram

By using DBNet++ network, upscoring extraction algorithm, character segmentation algorithm and CNN network recognition algorithm in the symbol diagram and circuit schematic diagram of electronic components, combined with the PIN Name lexicon and spelling error correction language model, the problem of low recognition accuracy in the existing technology is solved, and fast and accurate character information recognition is achieved.

CN120071379AActive Publication Date: 2025-05-30WUHAN UNIV OF TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510530344.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately identify character information in the symbol diagram of electronic components and circuit schematic diagrams of picture type, resulting in low recognition accuracy and a large number of missed and missed detections.

Method used

Character information recognition method in symbol diagram of electronic components is adopted, including text detection module, text recognition module and PIN Name text semantic repair module. Text detection is performed using DBNet++ network, text recognition is performed by combining the upper marking extraction algorithm, character segmentation algorithm and single-character classification CNN network recognition algorithm, and semantic repair is performed through the PIN Name lexicon and spelling error correction language model.

Benefits of technology

It realizes the rapid and accurate extraction of character information in the symbol diagram of electronic components and circuit schematic diagrams, improves the recognition accuracy and reduces missed and missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071379A_ABST
    Figure CN120071379A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electronic design automation, and discloses a method for identifying character information in a symbol diagram of an electronic component. Comprising an electronic component symbol graph character detection module for detecting an image area with character information, an electronic component symbol graph character recognition module for recognizing each character to obtain the character information in the character, and a construction PIN Name word library, and the electronic component symbol graph PIN Name text semantic restoration module is used for performing spelling correction on the PIN Name information by using a spelling error correction language model and outputting a final recognition result. According to the method for identifying the character information in the electronic component symbol diagram, various character information in the electronic component symbol diagram of the picture type and the schematic circuit diagram can be rapidly and accurately extracted, and an electronic component symbol diagram model is accurately obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic design automation, and particularly relates to a method for identifying character information in an electronic component symbol diagram. Background Art

[0002] In the schematic design tool and component selection tool of PCB EDA, the component symbol diagram parameter library is essential. Electrical design engineers hope to quickly obtain the component symbol diagram model from the component data sheet (containing a single component symbol diagram) or the existing circuit schematic diagram (containing multiple component symbol diagrams), and the main content is the number (PIN No) and name (PIN Name) of each pin of the component.

[0003] Currently, some semi-automatic symbol diagram library building tools have emerged in the market, but it is still very difficult to identify component symbols in the form of pictures. The reasons are as follows: there are a large number of combined symbols such as overlines, underlines, and slashes in the component symbol diagrams in the electronic component manuals and circuit schematic diagrams, and there are also some rare special symbols. Although they are printed fonts, the recognition accuracy is not high when using general OCR tools for recognition, and there will be a large number of missed detections and misdetections. Summary of the Invention

[0004] The purpose of the present invention is to address the above-mentioned deficiencies in the technology, and provide a method for identifying character information in an electronic component symbol diagram, which can quickly and accurately extract various character information in the electronic component symbol diagram and circuit schematic diagram in the form of pictures, and accurately obtain the electronic component symbol diagram model.

[0005] To achieve the above purpose, the method for identifying character information in an electronic component symbol diagram according to the present invention includes an electronic component symbol diagram text detection module, an electronic component symbol diagram text recognition module, and an electronic component symbol diagram PIN Name text semantic repair module; The electronic component symbol diagram text detection module, for the electronic component symbol diagram and circuit schematic diagram in the form of pictures, uses the segmentation-based scene text detection network DBNet++ to detect the image area where text information exists; The electronic component symbol diagram text recognition module is composed of an overline extraction algorithm, a character segmentation algorithm, and a single-character classification CNN network recognition algorithm. First, use the overline extraction algorithm to extract the overlines in the character area, then use the character segmentation algorithm to segment the characters in the image area obtained by the electronic component symbol diagram text detection module, and finally use the classification CNN network recognition algorithm to recognize each character to obtain the text information therein; The PIN Name semantic logic repair module for the electronic component symbol diagram includes a PIN Name thesaurus construction and a spelling correction language model. After constructing the PIN Name thesaurus, the spelling correction language model is used to correct the spelling of the PIN Name information obtained by the electronic component symbol diagram text recognition module, and the final recognition result is output.

[0006] Preferably, when the electronic component symbol diagram text detection module is used, several electronic component symbol diagrams of the picture type are selected from various electronic component manuals of the chip manufacturer, and several circuit schematic diagrams provided by the chip manufacturer are also selected. All the text information in the electronic component symbol diagrams and the circuit schematic diagrams is labeled and sent to the pre-trained DBNet++ network for training to obtain a proprietary DBNet++ network, so as to accurately detect the text information in the electronic component symbol diagrams and the circuit schematic diagrams.

[0007] Preferably, for the overline extraction algorithm, horizontal projection and vertical projection histograms are used to extract the overline. First, the horizontal projection histogram is used to determine whether there is an overline in the string. If there is an overline, the overline is cut, and then the cut part is vertically projected to accurately locate the character corresponding to the overline.

[0008] Preferably, the character segmentation algorithm is the connected component analysis method.

[0009] Preferably, the single-character classification CNN network recognition algorithm is a three-layer CNN network. 52 English upper and lower case characters, 10 digital characters, and several special characters are collected and sorted out. A total of several single-character images in the common font forms in the electronic component manual and the circuit schematic diagram are used to form a CNN training set for single-character images, and the single-character classification CNN network recognition algorithm in the electronic component symbol diagram is trained.

[0010] Preferably, when constructing the PIN Name thesaurus, the electronic component symbol information of various device types of the chip manufacturer is collected and sorted out, all the PIN Names appearing therein are statistically analyzed, the PIN Name thesaurus is constructed, and the word frequency of each PIN Name is counted. For the PIN Name obtained by the electronic component symbol diagram text recognition module, according to whether it appears in the PIN Name thesaurus and the corresponding word frequency, the confidence level of the PIN Name is set.

[0011] Preferably, the spelling correction language model collects and collates the electronic component symbol information of various device types of chip manufacturers, constructs sentence forms, with each electronic component corresponding to a piece of text. This text includes the name, type, manufacturer, and number of pins of the electronic component, and all PIN Names are arranged in the text according to the pin numbers. Then, a data set is made based on these text collections, and the migrated pre-training data set and labeled data are sent into the Soft Masked BERT network for training to obtain a spelling correction language model for semantic logic repair of the electronic component PIN Names obtained by the electronic component symbol graph text recognition module.

[0012] Preferably, the overline extraction algorithm includes the following steps: Step S101: Read the string image within the coordinate range of the text information detected by the DBNet++ network; Step S102: Perform binarization processing; Step S103: Draw the horizontal projection histogram of the string image and count the number of pixel points in the horizontal direction; Step S104: Determine whether there is an obvious pixel-free area in the horizontal projection histogram. If there is no such area, there is no overline in the string image; if there is such an area, there is an overline in the string image; Step S105: Use the pixel-free area to segment the string image and select the narrow part as the overline part image; Step S106: Perform vertical histogram statistics on the overline part image to obtain the coordinate range of each overline segment.

[0013] Preferably, the character segmentation algorithm and the single-character classification CNN network recognition algorithm include the following steps: Step S201: Read the string image obtained after overline extraction; Step S202: Perform connected component analysis on the image to obtain N connected components, and divide the string image into N single-character images according to the vertical midlines between adjacent connected components; Step S203: Take the next single-character image from left to right in sequence; Step 204: If the aspect ratio of the single-character image is greater than 1.5, there are connected characters, and go to Step S205; otherwise, go to Step S206; Step S205: Use the connected component method to segment the single-character image into single-character images without connected characters, and then use the single-character classification CNN network recognition algorithm to recognize the single-character images, and go to Step S207; Step S206: Use the single-character classification CNN network recognition algorithm to recognize the single-character image; Step S207: Append the recognized characters to the end of the previously recognized result string. If there are still single-character images that have not been recognized, go to step S203; otherwise, go to step S208; Step S208: Match the recognized result string with the result obtained by the overline extraction algorithm to get the final recognition result.

[0014] Preferably, the training of the spelling correction language model includes the following steps: Step S301: Collect and organize the symbol information of several electronic components, including device name, manufacturer name, device type, number of pins, and PIN Name of all pins; Step S302: Construct the symbol information of the electronic component into an English sentence, including device name, manufacturer, device type, number of pins, and list all PIN Names of the electronic component in order as the Origin_text; Step S303: Construct a batch of Random_text according to the experience of PIN Name recognition errors and random methods; Step S304: Make labeled data; Step S305: Feed the Origin_text, Random_text, and labeled data into the Soft Masked BERT network for training. During the fine-tuning process, keep the learning rate at 2×10^-5, the number of hidden units in the bidirectional GRU is 256, and the batch size used by the model is 320 to obtain a spelling correction language model for semantic repair of the PIN Name obtained by the electronic component symbol graph text recognition module.

[0015] Compared with the prior art, the present invention has the following advantages: 1. It can quickly and accurately extract various character information in the electronic component symbol graph and circuit schematic diagram of the picture type, and accurately obtain the electronic component symbol graph model; 2. According to the electronic component symbol graph scenario, train the self-owned DBNet++ network of the electronic component symbol graph, which can realize accurate detection of the text in the electronic component symbol graph; 3. According to the characteristics that the characters in the electronic component symbol graph are printed characters and only include English letters, numbers, and some special characters, use the CNN network, which has a simple and effective structure and high recognition accuracy; 4. According to the characteristics that there are a large number of characters with overlines in the electronic component symbol graph, use horizontal and vertical projection histograms to detect the overlines, which can accurately locate the overlines; 5. Construct a common PIN Name library for electronic components, and set the confidence of the PIN Name obtained by the electronic component symbol graph text recognition module accordingly; 6. Train a Soft Masked BERT language model dedicated to the PIN Name of electronic components, enabling it to perform overall semantic logic repair on the previously recognized PIN Name of electronic components according to the characteristics of the electronic component types and perform effective error correction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a module diagram of the method for identifying character information in the electronic component symbol diagram of the present invention; Figure 2 It is a flowchart of the upper line extraction algorithm in the text recognition module of the electronic component symbol diagram of the present invention; Figure 3 It is a flowchart of the character segmentation algorithm and the single-character classification CNN network recognition algorithm in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions of the present invention will be described clearly and completely in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] As Figure 1 shown, a method for identifying character information in an electronic component symbol diagram includes an electronic component symbol diagram text detection module, an electronic component symbol diagram text recognition module, and an electronic component symbol diagram PIN Name text semantic repair module; The electronic component symbol diagram text detection module uses the segmentation-based scene text detection network DBNet++ for the electronic component symbol diagram and circuit schematic diagram of the picture type to detect the image area with text information; The electronic component symbol diagram text recognition module consists of an upper line extraction algorithm, a character segmentation algorithm, and a single-character classification CNN network recognition algorithm. First, use the upper line extraction algorithm to extract the upper line in the character area, and then use the character segmentation algorithm to segment the characters in the image area obtained by the electronic component symbol diagram text detection module. Finally, use the classification CNN network recognition algorithm to recognize each character and obtain the text information therein; The electronic component symbol diagram PIN Name semantic logic repair module includes PIN Name thesaurus construction and a spelling correction language model. After constructing the PIN Name thesaurus, use the spelling correction language model to perform spelling correction on the PIN Name information obtained by the electronic component symbol diagram text recognition module and output the final recognition result.

[0019] In this embodiment, when the electronic component symbol diagram text detection module is used, 1000 electronic component symbol diagrams in the form of pictures are selected from various electronic component manuals of common chip manufacturers around the world. At the same time, 100 circuit schematic diagrams provided by the chip manufacturers are selected, and all the text information in the electronic component symbol diagrams and the circuit schematic diagrams is labeled and sent to the pre-trained DBNet++ network for training to obtain a proprietary DBNet++ network, so as to accurately detect the text information in the electronic component symbol diagrams and the circuit schematic diagrams.

[0020] In this embodiment, for the overline extraction algorithm, horizontal projection and vertical projection histograms are used to extract the overline. First, the horizontal projection histogram is used to determine whether there is an overline in the string. If there is an overline, the overline is cut, and then the vertical projection is performed on the cut part to accurately locate the character corresponding to the overline.

[0021] Specifically, as Figure 2 shown, the overline extraction algorithm includes the following steps: Step S101: Read the string image within the coordinate range of the text information detected by the DBNet++ network; Step S102: Perform binarization processing; Step S103: Draw the horizontal projection histogram of the string image and count the number of pixel points in the horizontal direction; Step S104: Determine whether there is an obvious pixel-free area in the horizontal projection histogram. If there is no such area, there is no overline in the string image. If there is such an area, there is an overline in the string image; Step S105: Use the pixel-free area to segment the string image, and select the narrow part as the overline part image; Step S106: Perform vertical histogram statistics on the overline part image to obtain the coordinate range of each overline segment.

[0022] In this embodiment, the character segmentation algorithm is the connected component analysis method, and the single-character classification CNN network recognition algorithm is a three-layer CNN network. 52 English upper and lower case characters, 10 numeric characters, and 20 special characters are collected and sorted. Each single-character image includes two types, itself and rotated 180°. For each single character, 1000 pictures of common printed fonts in various electronic component manuals are included, totaling 164,000 as the training set, forming the CNN training set of single-character images. Using a three-layer CNN network, a total of 70 rounds of training are performed to train the single-character classification CNN network recognition algorithm in the electronic component symbol diagram.

[0023] Specifically, in this embodiment, as Figure 3As shown in the figure, the character segmentation algorithm and the single-character classification CNN network recognition algorithm include the following steps: Step S201: Read the string image obtained after overline extraction; Step S202: Perform connected component analysis on the image to obtain N connected components. According to the vertical midlines between adjacent connected components, divide the string image into N single-character images; Step S203: Take the next single-character image in sequence from left to right; Step 204: If the aspect ratio of the single-character image is greater than 1.5, there are adhesive characters, and go to step S205; otherwise, go to step S206; Step S205: Use the connected component method to segment the single-character image into single-character images without adhesive characters, and then use the single-character classification CNN network recognition algorithm to recognize the single-character images, and go to step S207; Step S206: Use the single-character classification CNN network recognition algorithm to recognize the single-character image; Step S207: Append the recognized character to the end of the previously recognized result string. If there are still unrecognized single-character images, go to step S203; otherwise, go to step S208; Step S208: Match the recognized result string with the result obtained by the overline extraction algorithm to obtain the final recognition result.

[0024] In this embodiment, when constructing the PIN Name library, 200,000 pieces of electronic component symbol information of various device types of chip manufacturers are collected and sorted out. All the PIN Names that appear are statistically analyzed to construct the PIN Name library, and the word frequency of each PIN Name is counted. For the PIN Names obtained by the electronic component symbol text recognition module, according to whether they appear in the PIN Name library and the corresponding word frequency, the confidence level of the PIN Name is set.

[0025] The spelling correction language model collects and sorts out 200,000 pieces of electronic component symbol information of various device types of chip manufacturers, constructs sentence forms, and each electronic component corresponds to a paragraph of text. This text contains the name, type, manufacturer, and number of pins of the electronic component, and all the PIN Names are arranged in the text according to the pin numbers. Then, a data set is made based on these text collections, and the pre-trained data set and the labeled data are sent into the Soft Masked BERT network for training to obtain the spelling correction language model for semantic logic repair of the electronic component PIN Names obtained by the electronic component symbol text recognition module.

[0026] Specifically, the training of the spelling correction language model includes the following steps: Step S301: Collect and organize the symbol information of 200,000 electronic components, including device name, manufacturer name, device type, number of pins, and PIN Name of all pins. Step S302: Construct the symbol information of the electronic component into an English sentence, including device name, manufacturer, device type, number of pins, and list all PIN Names of the electronic component in order as the Origin_text. For example, for the ADC device ADS1240 of TI company, the corresponding text sentence is: "The ADS1240 is an ADC of TI company. It is composed of 24 pins. The name of the pins from 1 to 24 in an ascending order are DVDD, DGND, XIN, XOUT, RESET, DSYNC, PDWN, DGND, VREF+, VREF–, AIN0 / D0, AIN1 / D1, AIN2 / D2, AIN3 / D3, AINCOM, AGND, AVDD, POL, CS, DIN, DOUT, SCLK, DRDY, BUFEN." Step S303: Construct a batch of Random_text according to the error recognition experience of PIN Name and the random method. For example, the Random_text corresponding to the example sentence in Step S302 is as follows: "ADS1240 is an ADC of TI company. It is composed of 24 pins. The name of the pins from 1 to 24 in an ascending order are DVDD, DGND, XIN, X*, RESET, DSYNC, PDWN, DGND, VREF+, VREF*, AIN0 / D0, AIN1 / D1, AIN* / D2, AIN3 / D3, AINCOM, AGND, AVDD, POL, CS, DIN, DOUT, SCLK, DRDY, BUFEN.," Step S304: Generate labeled data. The labeled data corresponding to Origin_text in step S302 and Random_text in step S303 is "0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 00 0 1 0 0 0 0 0 1 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0", where 0 indicates that the corresponding PIN Name is correct and 1 indicates that the corresponding PIN Name is incorrect; Step S305: Feed Origin_text, Random_text, and the labeled data into the Soft Masked BERT network for training. During fine-tuning, keep the learning rate at 2×10^-5, set the number of hidden units in the bidirectional GRU to 256, and use a batch size of 320 for the model to obtain a spelling correction language model for semantic repair of the PIN Names obtained by the electronic component symbol graph text recognition module.

[0027] The method for identifying character information in the electronic component symbol graph of the present invention can quickly and accurately extract various character information in the electronic component symbol graph and circuit schematic diagram of the picture type, and accurately obtain the electronic component symbol graph model; according to the electronic component symbol graph scenario, train the self-owned DBNet++ network of the electronic component symbol graph to achieve accurate detection of the text in the electronic component symbol graph; according to the characteristics that the characters in the electronic component symbol graph are printed characters and only include English letters, numbers, and some special characters, adopt a CNN network, which is simple and effective with high recognition accuracy; according to the characteristics that there are a large number of characters with overlines in the electronic component symbol graph, use horizontal and vertical projection histograms to detect the overlines and can accurately locate the overlines; construct a common electronic component PIN Name library, and set the confidence level of the PIN Names obtained by the electronic component symbol graph text recognition module accordingly; train a dedicated SoftMasked BERT language model for the electronic component PIN Name, so that it can perform overall semantic logic repair on the previously recognized PIN Names of the electronic components according to the characteristics of the electronic component type and perform effective error correction.

[0028] Meanwhile, it should be noted that the description of the above technical solutions is exemplary. This specification can be embodied in different forms and should not be construed as limited to the technical solutions set forth herein. On the contrary, providing these descriptions will make the disclosure of the present invention thorough and complete, and will fully convey the scope disclosed in this specification to those skilled in the art. In addition, the technical solutions of the present invention are only defined by the scope of the claims. The features of the various embodiments of the present invention can be combined or spliced partially or fully with each other and can be implemented in various different ways as can be fully understood by those skilled in the art. The embodiments of the present invention can be implemented independently of each other or can be implemented together in a mutually dependent relationship.

[0029] For those of ordinary skill in the art to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can also be made, which should all be regarded as falling within the protection scope of the present invention.

Claims

1. A method for identifying character information in an electronic component symbol diagram, characterized in that: It includes an electronic component symbol image text detection module, an electronic component symbol image text recognition module, and an electronic component symbol image PIN Name text semantic repair module; The electronic component symbol diagram text detection module uses the segmentation-based scene text detection network DBNet++ to detect image areas with text information for electronic component symbol diagrams and circuit schematic diagrams of picture type; The electronic component symbol image text recognition module is composed of an overline extraction algorithm, a character segmentation algorithm, and a single character classification CNN network recognition algorithm. The overline extraction algorithm is first used to extract the overlines in the character area, and then the character segmentation algorithm is used to segment the characters in the image area obtained by the electronic component symbol image text detection module. Finally, the classification CNN network recognition algorithm is used to recognize each character to obtain the text information therein; The electronic component symbol PIN Name semantic logic repair module includes PIN Name word library construction and spelling error correction language model. After the PIN Name word library is constructed, the spelling error correction language model is used to perform spelling correction on the PIN Name information obtained by the electronic component symbol text recognition module, and the final recognition result is output.

2. The method for recognizing character information in an electronic component symbol diagram as claimed in claim 1, characterized in that: When the electronic component symbol diagram text detection module is used, several electronic component symbol diagrams of picture type are selected from various electronic component manuals of chip manufacturers, and several circuit schematic diagrams provided by the chip manufacturers are selected at the same time. All text information in the electronic component symbol diagrams and circuit schematic diagrams is annotated and sent to a pre-trained DBNet++ network for training to obtain a proprietary DBNet++ network, thereby realizing accurate detection of text information in electronic component symbol diagrams and circuit schematic diagrams.

3. The method for character information recognition in an electronic component symbol diagram as claimed in claim 1, characterized in that: The overline extraction algorithm uses horizontal projection and vertical projection histograms to extract overlines. First, the horizontal projection histogram is used to determine whether there is an overline in the character string. If there is an overline, the overline is cut. Then, a vertical projection is performed on the cut part to accurately locate the character corresponding to the overline.

4. The method for recognizing character information in an electronic component symbol diagram according to claim 1, characterized in that: The character segmentation algorithm is a connected domain analysis method.

5. The method for recognizing character information in an electronic component symbol diagram according to claim 1, characterized in that: The single-character classification CNN network recognition algorithm is a three-layer CNN network, which collects and organizes 52 uppercase and lowercase English characters, 10 numeric characters and several special characters, and a total of several single-character images in common fonts in electronic component manuals and circuit schematics to form a CNN training set of single-character images, and trains a single-character classification CNN network recognition algorithm in electronic component symbol diagrams.

6. The method for recognizing character information in an electronic component symbol diagram according to claim 1, characterized in that: When building the PINName thesaurus, collect and organize the electronic component symbol information of various device types from chip manufacturers, statistically analyze all the PIN Names that appear in them, build the PIN Name thesaurus, and count the word frequency of each PIN Name. For the PIN Name obtained by the electronic component symbol image text recognition module, set the confidence of the PIN Name based on whether it appears in the PIN Name thesaurus and the corresponding word frequency.

7. The method for recognizing character information in an electronic component symbol diagram according to claim 1, characterized in that: The spelling correction language model collects and organizes the electronic component symbol information of various device types of chip manufacturers, constructs a sentence form, and each electronic component corresponds to a text, which contains the name, type, manufacturer and number of pins of the electronic component, and arranges all PIN Names according to the pin number in the text. Then, a data set is created based on these text sets, and the pre-trained data set and annotated data are migrated and sent to the Soft Masked BERT network for training. The spelling correction language model is used to perform semantic logic repair on the electronic component PIN Name obtained by the electronic component symbol graph text recognition module.

8. The method for recognizing character information in an electronic component symbol diagram as claimed in claim 3, characterized in that: The overline extraction algorithm includes the following steps: Step S101: reading a character string image within the coordinate range of the text information detected by the DBNet++ network; Step S102: perform binarization processing; Step S103: draw a horizontal projection histogram of the character string image and count the number of pixels in the horizontal direction; Step S104: determining whether there is an obvious pixel-free area in the horizontal projection histogram; if there is no such area, there is no overline in the string image; if there is such an area, there is an overline in the string image; Step S105: segmenting the character string image using the pixel-free area, and selecting the narrow part as the overline part image; Step S106: Perform vertical histogram statistics on the overlined portion of the image to obtain the coordinate range of each overline segment.

9. The method for recognizing character information in an electronic component symbol diagram according to claim 1, characterized in that: The character segmentation algorithm and the single character classification CNN network recognition algorithm include the following steps: Step S201: reading the character string image obtained after overline extraction; Step S202: Performing a connected domain analysis on the image to obtain N connected domains, and dividing the character string image into N blocks of single-character images according to vertical midlines between adjacent connected domains; Step S203: taking a single character image from left to right in sequence; Step 204: If the aspect ratio of the single character image is greater than 1.5, there are contiguous characters, and the process proceeds to step S205; otherwise, the process proceeds to step S206; Step S205: using the connected domain method to segment the single character image into single character images without adhesion characters, and then using the single character classification CNN network recognition algorithm to recognize the single character images respectively, and then proceeding to step S207; Step S206: using a single character classification CNN network recognition algorithm to identify the single character image; Step S207: insert the recognized character into the end of the previous recognition result string. If there is still an unrecognized single character image, proceed to step S203; otherwise, proceed to step S208; Step S208: Match the recognition result character string with the result obtained by the overline extraction algorithm to obtain the final recognition result.

10. The method for recognizing character information in an electronic component symbol diagram according to claim 1, characterized in that: The spelling correction language model training includes the following steps: Step S301: collect and organize symbol information of several electronic components, including device name, manufacturer name, device type, number of pins and PIN Name of all pins; Step S302: construct the symbol information of the electronic component into an English sentence, including the device name, manufacturer, device type, number of pins, and list all the PIN Names of the electronic component in order as Origin_text; Step S303: construct a batch of Random_text according to the PIN Name recognition error experience and random method; Step S304: creating annotation data; Step S305: Origin_text, Random_text and annotated data are sent to the Soft Masked BERT network training. During the fine-tuning process, the learning rate is kept at 2×10^-5, the number of hidden units in the bidirectional GRU is 256, and the batch size used by the model is 320. The spelling correction language model is used for semantic repair of the PINName obtained by the electronic component symbol image text recognition module.

Citation Information

Patent Citations

  • Underline removal apparatus

    CN101859379A

  • A machine vision detection algorithm for logic circuit diagram information extraction

    CN109508676A

  • Chip surface character recognition method based on deep learning

    CN112115948A

  • Visual multi-mode character detection recognition and error correction method in network public opinion analysis

    CN116229482A

  • Picture table recognition method and picture table recognition device

    CN116935421A