OCR (Optical Character Recognition) device based on artificial intelligence recognition and recognition method

Through the OCR device and method based on artificial intelligence, the problems of inefficient cable verification and insufficient accuracy are solved, and the automated efficient identification and rapid alarm functions are realized.

CN120356222APending Publication Date: 2025-07-22STATE GRID JIANGSU ELECTRIC POWER CO ZHENJIANG POWER SUPPLY CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510423892.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the inefficiency and insufficient accuracy of cable verification will increase the probability of failure in testing or starting.

Method used

Using an OCR device based on artificial intelligence, including a visual camera, telescopic rod body, base bracket and support feet, text positioning, correction and recognition through image enhancement algorithms and depth relationship inference graph networks, and model character relationships with graph convolution neural networks to automatically identify line lines and tabular texts.

Benefits of technology

Improve cable recognition efficiency and verification accuracy, reduce the need for manual verification, and achieve rapid identification and output inconsistent text alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356222A_ABST
    Figure CN120356222A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and particularly discloses an OCR (Optical Character Recognition) device based on artificial intelligence recognition and a recognition method.The OCR device comprises a visual camera, a telescopic rod body, a first base support, a second base support, a first supporting leg, a second supporting leg and a triangular supporting block; the first base support is in sliding connection with the telescopic rod body and arranged on the outer surface wall of the telescopic rod body in a sleeving mode, the second base support is fixedly connected with the telescopic rod body, and one end of each first supporting leg is rotationally connected with the first base support and located on the outer surface wall of the first base support. Each triangular supporting block is rotationally connected with the corresponding first supporting leg, and one end of each second supporting leg is rotationally connected with the second base support. The cable characters and the table characters are recognized through artificial intelligence, manual visual inspection is not needed, the cable recognition efficiency is effectively improved, and the checking accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to an OCR device and an identification method based on artificial intelligence identification. Background Art

[0002] Currently, during the wiring process of cabinets and electrical cabinets, a variety of cables are required. After installation, multiple inspections and checks need to be carried out. Only after ensuring accuracy can the machine be tested or started. To ensure the accuracy of the cable verification process, marks composed of different combinations of letters and numbers are printed on the corresponding cables and interfaces, facilitating the docking of the corresponding cables and corresponding interfaces, improving the docking accuracy, and also improving the convenience of subsequent verification.

[0003] However, in the prior art, when verifying and docking cables, manual verification is used, which results in low verification efficiency and insufficient verification accuracy, increasing the probability of machine test failure or startup failure. Summary of the Invention

[0004] The purpose of the present invention is to provide an OCR device and an identification method based on artificial intelligence identification, aiming to solve the technical problems of low verification efficiency and insufficient verification accuracy in the prior art when using manual verification.

[0005] To achieve the above purpose, an OCR device based on artificial intelligence identification of the present invention includes a vision camera, a telescopic rod body, a first base bracket, a second base bracket, a first support foot, a second support foot, and a triangular support block. The vision camera is fixedly connected to the telescopic rod body and is located at the upper end of the telescopic rod body. The first base bracket is slidably connected to the telescopic rod body and is sleeved on the outer wall of the telescopic rod body. The second base bracket is fixedly connected to the telescopic rod body and is located on the outer wall of the telescopic rod body. The number of the first support feet is multiple, and one end of each first support foot is respectively rotatably connected to the first base bracket and is respectively located on the outer wall of the first base bracket. The number of the triangular support blocks is multiple, and each triangular support block is respectively rotatably connected to the corresponding first support foot and is respectively located at one end of the corresponding first support foot. The number of the second support feet is multiple, and one end of each second support foot is respectively rotatably connected to the second base bracket and is respectively located on the outer wall of the second base bracket, and the other end of each second support foot is respectively rotatably connected to the corresponding triangular support block.

[0006] Among them, the telescopic rod body includes a bottom rod, an intermediate rod and a top rod. The first base bracket and the second base bracket are respectively sleeved on the outer wall of the bottom rod. The intermediate rod is slidably connected to the bottom rod and is embedded inside the bottom rod. The top rod is slidably connected to the intermediate rod and is embedded inside the intermediate rod. The vision camera is arranged at the upper end of the top rod.

[0007] The present invention also provides an OCR recognition method based on artificial intelligence recognition, which is applied to the OCR device based on artificial intelligence recognition as described above, and includes the following steps:

[0008] S1: Perform text positioning on the text number at the wire arrangement through the vision camera;

[0009] S2: Perform text correction on the recognized text number of the wire arrangement;

[0010] S3: Perform text recognition on the text number of the wire arrangement;

[0011] S4: Mine the positional relationship of the text on the wire arrangement and match the corresponding text and number;

[0012] S5: Perform text positioning on the text number at the table through the vision camera;

[0013] S6: Perform text correction on the recognized text number of the table;

[0014] S7: Perform text recognition on the text number of the table;

[0015] S8: Output the structured relationship of the text number in the table;

[0016] S9: Compare the text and number matched by the wire arrangement with the text number in the table, and highlight and output the inconsistent text numbers.

[0017] Among them, in the step S1, it includes the following steps:

[0018] S101: Calculate through an image enhancement algorithm, including contrast processing and color domain processing, to make the text clearer and reduce adhesion;

[0019] S102: Adapt the image to the input of the text detection framework. For text detection, use the deep relationship reasoning graph network DRRG, model each character as a small rectangle, and then use the graph convolutional neural network to model the relationship between each character;

[0020] S103: Finally, output the position coordinates of the four corners of the minimum bounding box of the text line.

[0021] Among them, in the step S101, it includes the following steps:

[0022] Calculate the maximum value data_max and minimum value data_min of the pixels of the image Image. The exponential stretching constant C = pow((data_max - data_min), erp), and erp usually takes a value between 1.5 and 2;

[0023] Map each channel of the image to a new domain ImageNew, where ImageNew(i, j) = pow((erpImage(i, j) - data_min), erp) / C * 255;

[0024] The color domain transformation includes image white balance and color transformation. ImageWB is the image after white balance processing of ImageNew, and ImageColorScale is the image after color transformation. ImageColorScale(k) = ImageWB(k) * a + b, where a and b are constants, and k is the RGB three-component index value. When k = 0, perform blue component transformation on ImageWB, a takes 0.485, and b takes 0.229; when k = 1, perform green component transformation on ImageWB, a takes 0.456, and b takes 0.224; when k = 2, perform red component transformation on ImageWB, a takes 0.406 and b takes 0.225;

[0025] In the step S102, the following steps are included:

[0026] Normalize the graph to the 0 - 1 interval, resize it to 960 * 960 pixels, and rearrange it by columns to generate a Tensor;

[0027] Input the Tensor into the DRRG model and output the Box of a single-line character.

[0028] Among them, in the step S2, the following steps are included:

[0029] S201: Sort the Boxes of the input text area in ascending order of the x coordinate of the upper left corner point;

[0030] S202: Perform elliptical region detection on the image and calculate the positions of the screw holes;

[0031] S203: Draw a dividing line Line from the center point of the screw hole area, take out BoxLeft of the part falling on the left side of the dividing line, sort BoxLeft in ascending order of the y coordinate of the upper left corner point to obtain BoxVal, and record the remaining part as BoxDiff;

[0032] S204: Divide the BoxDiff area again, perform morphological processing, and take out a certain area on the left side of Line and record it as BoxTagOrg;

[0033] S205: Refine BoxVal to obtain the dilated element, perform dilation operation on BoxVal in the horizontal direction to obtain BoxValExpand;

[0034] S206: Use Otsu's method to perform threshold segmentation on the input image to obtain the Light region RegionLight;

[0035] S207: Take out the region RegionLight1 near BoxVal on the left side of the separator line Line of RegionLight and on the right side of BoxVal, and take out the region RegionLight2 near Line on the left side of the separator line Line of RegionLight and on the right side of BoxVal;

[0036] S208: Use opening operation to remove the adhesion of RegionLight1 to obtain RegionLightErosion, and sequentially take out the region RegionGuanOrg(i) with the largest overlapping area with BoxValExpand(i), where BoxValExpand(i) represents the i-th expanded region of the cable text;

[0037] S209: Use RegionGuanOrg(i) as the seed and use the connected component growth method to obtain the part within the RegionLight2 region, which is RegionGuan(i);

[0038] S210: Use Otsu's method to perform threshold segmentation on the left part of the image Line, take out the Dark part, which is RegionDark, and perform morphological processing on RegionDark to separate it into individual holes to obtain RegionHole;

[0039] S211: Refine RegionGuan to obtain the dilated element2, and perform dilation in the horizontal direction to obtain RegionGuanExpand;

[0040] S212: Sequentially take out the part with the largest overlapping area between RegionGuanExpand(i) and the RegionHole region, denoted as RegionHole(i), and perform horizontal expansion on RegionHole(i), denoted as RegionHoleExpand(i);

[0041] S213: Sequentially take out the part with the largest overlapping area with RegionHoleExpand(i) from BoxTagOrg, denoted as BoxTag(i), and the region corresponding to BoxVal(i) is BoxTag(i), and output them in pairs.

[0042] Among them, in the step S3, the following steps are included:

[0043] S301: In the step of performing text correction, the BoxTag and BoxVal are transformed into a regular rectangle through affine transformation, and the transformation matrix H is calculated;

[0044] S302: Take out the BoxTag and BoxVal parts in the ImageColorScale, and use the inverse matrix H_inv of the transformation matrix to correct the image into a regular rectangle ImageTagRect and ImageValRect;

[0045] S303: The corrected image is normalized, scaled to 192*48, and sorted by column to obtain Tsesor_Cal;

[0046] S304: Input Tsesor_Cal into the trained MobileNet model, and output the text direction correction result. The output value is 0, no correction is required; 90, rotate 90 degrees clockwise; -90, rotate 90 degrees counterclockwise; 180, flip vertically;

[0047] S305: Rotate the ImageTagRect and ImageVal by the center according to their respective output angles to obtain ImageTag and ImageVal.

[0048] Among them, in the step S4, the following steps are included:

[0049] S401: Normalize the ImageTag and ImageVal, scale them to 320*48, and rearrange the image by column to obtain TensorTag and TensorVal;

[0050] S402: Input TensorTag and TensorVal into the trained SVTR model, and output the index and score;

[0051] S403: Use the index to obtain characters from the font library dictionary, discard the characters with too low scores, and output them as a string.

[0052] Among them, in the step S6, the following steps are included:

[0053] S601: Perform charthreshold segmentation on the table image, and take out the text and wireframe parts RegionDark;

[0054] S602: Dilate the Box of the text to obtain BoxDila;

[0055] S603: Calculate the corner points of RegionDark, remove the overlapping part with BoxDila, and obtain haarCell;

[0056] S604: Input haarCell into the SLANet model to obtain the structured relationship of the table and express it in html format;

[0057] S605: Analyze the table structured relationship in html format, use the input prior knowledge to obtain the table text and the table number column, analyze the region BoxVal of the table text string and the region BoxTag of the table number column, and output BoxVal(i) and BoxTag(i) in pairs according to the corresponding order.

[0058] Among them, in the step S9, the following steps are included:

[0059] S901: The recognition results of the cable layout text and the cable layout number are LineResultVal and LineResultTag, and the recognition results of the table text and the table number are TableResultVal and TableResultTag;

[0060] S902: Taking TableResultVal(i) and TableResultTag(i) in turn based on the table, search in the cable layout detection result LineResultTag. If LineResultTag(j)≠TableResultTag(i), then search for LineResultTag(j + 1) until all of LineResultTag is searched. If LineResultTag(j) = TableResultTag(i) and TableResultVal(i)≠LineResultVal(j), mark it as a value inequality defect. If LineResultTag(j) = TableResultTag(i) and TableResultVal(i) = LineResultVal(j), mark it as normal and remove it from the candidates to prevent repeated searches;

[0061] S903: If no same item is found in LineResultTag in TableResultTag, mark it as missing.

[0062] The beneficial effects of an OCR device and an identification method based on artificial intelligence recognition of the present invention are as follows: By using artificial intelligence to recognize the text on the cable harness and the text in the table, visual inspection by humans is not required, effectively improving the recognition efficiency of the cable and the verification accuracy. Through automatic recognition by artificial intelligence, the text is quickly read, recognized, and corrected, and the text on the cable harness and the text in the table are quickly compared, and the text in inconsistent groups is quickly output to achieve the purpose of rapid alarm. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0064] Figure 1 is a schematic structural diagram of an OCR device based on artificial intelligence recognition of the present invention.

[0065] 1 - Vision camera, 2 - Bottom rod, 3 - Intermediate rod, 4 - Top rod, 5 - First base bracket, 6 - Second base bracket, 7 - First support foot, 8 - Second support foot, 9 - Triangular support block.

[0066] Figure 2 is a flowchart of an OCR recognition method based on artificial intelligence recognition of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] Please refer to Figure 1, the present invention provides an OCR device based on artificial intelligence recognition, including a vision camera 1, a telescopic rod body, a first base bracket 5, a second base bracket 6, a first support foot 7, a second support foot 8, and a triangular support block 9. The vision camera 1 is fixedly connected to the telescopic rod body and is located at the upper end of the telescopic rod body. The first base bracket 5 is slidably connected to the telescopic rod body and is sleeved on the outer wall of the telescopic rod body. The second base bracket 6 is fixedly connected to the telescopic rod body and is located on the outer wall of the telescopic rod body. The number of the first support feet 7 is multiple, and one end of each first support foot 7 is respectively rotatably connected to the first base bracket 5 and is respectively located on the outer wall of the first base bracket 5. The number of the triangular support blocks 9 is multiple, and each triangular support block 9 is respectively rotatably connected to the corresponding first support foot 7 and is respectively located at one end of the corresponding first support foot 7. The number of the second support feet 8 is multiple, and one end of each second support foot 8 is respectively rotatably connected to the second base bracket 6 and is respectively located on the outer wall of the second base bracket 6, and the other end of each second support foot 8 is respectively rotatably connected to the corresponding triangular support block 9.

[0068] Further, the telescopic rod body includes a bottom rod 2, an intermediate rod 3, and a top rod 4. The first base bracket 5 and the second base bracket 6 are respectively sleeved on the outer wall of the bottom rod 2. The intermediate rod 3 is slidably connected to the bottom rod 2 and is embedded in the interior of the bottom rod 2. The top rod 4 is slidably connected to the intermediate rod 3 and is embedded in the interior of the intermediate rod 3. The vision camera 1 is arranged at the upper end of the top rod 4.

[0069] The present invention also provides an OCR recognition method based on artificial intelligence recognition, which is applied to the OCR device based on artificial intelligence recognition as described above, as Figure 2 shown, and includes the following steps:

[0070] S1: Perform text positioning on the text numbers at the cable arrangement through the vision camera 1;

[0071] S2: Perform text correction on the recognized text numbers of the cable arrangement;

[0072] S3: Perform text recognition on the text numbers of the cable arrangement;

[0073] S4: Explore the positional relationship of the text at the cable arrangement and match the corresponding text and numbers;

[0074] S5: Perform text positioning on the text numbers at the table through the vision camera 1;

[0075] S6: Perform text correction on the recognized text numbers of the table;

[0076] S7: Perform optical character recognition on the text numbers in the table;

[0077] S8: Output the structured relationship of the text numbers in the table;

[0078] S9: Compare the text and numbers matched by the cable with those in the table, and highlight and output the inconsistent text numbers.

[0079] Further, in the step S1, the following steps are included:

[0080] S101: Calculate through an image enhancement algorithm, including contrast processing and color domain processing,

[0081] to make the text clearer and reduce adhesion;

[0082] S102: Adapt the image to the input of the text detection framework. For text detection, use the Deep Relationship Reasoning Graph Network (DRRG). Model each character as a small rectangle, and then use a graph convolutional neural network to model the relationship between each character;

[0083] S103: Output the coordinates of the minimum bounding box of the text line.

[0084] After the above processing, the system will find the specific position of each line of text in the picture and draw a minimum box (Box) that can just completely enclose this line of text. The final output result is the

[0085] coordinate values of the four corners of this box, in the format of:

[0086] [Upper left corner (x1, y1), upper right corner (x2, y2), lower right corner (x3, y3), lower left corner (x4, y4)].

[0087] Further, in the step S101, the following steps are included:

[0088] Calculate the maximum value data_max and minimum value data_min of the image Image pixels. The exponential stretching constant C = pow((data_max - data_min), erp), and erp usually takes 1.5 - 2;

[0089] Map each channel of the image to a new domain ImageNew, ImageNew(i, j) = pow((erpImage(i, j) - data_min), erp) / C * 255;

[0090] Color domain transformation includes image white balance and color transformation. ImageWB is the image after white balance processing of ImageNew, and ImageColorScale is the image after color transformation. ImageColorScale(k) = ImageWB(k) * a + b, where a and b are constants, and k is the RGB three-component index value. When k = 0, the blue component of ImageWB is transformed, a takes 0.485, and b takes 0.229; when k = 1, the green component of ImageWB is transformed, a takes 0.456, and b takes 0.224; when k = 2, the red component of ImageWB is transformed, a takes 0.406 and b takes 0.225;

[0091] In the step S102, the following steps are included:

[0092] Normalize the graph to the 0-1 interval, resize it to 960 * 960 pixels, and rearrange it by column to generate a Tensor;

[0093] Input the Tensor into the DRRG model and output the Box of a single-line character

[0094] . Further, in the step S2, the following steps are included:

[0095] S201: Sort the Boxes of the input text area in ascending order of the x coordinate of the upper left corner point;

[0096] S202: Perform elliptical region detection on the image and calculate the positions of the screw holes;

[0097] S203: Draw a dividing line Line from the center point of the screw hole area, take out the BoxLeft on the left side of the dividing line, sort BoxLeft in ascending order of the y coordinate of the upper left corner point to obtain BoxVal, and record the remaining part as BoxDiff;

[0098] S204: Re-segment the BoxDiff area, perform morphological processing, and take out a certain area on the left side of Line and record it as BoxTagOrg;

[0099] S205: Refine BoxVal to obtain a dilation element, perform a dilation operation on BoxVal in the horizontal direction to obtain BoxValExpand;

[0100] S206: Use the Otsu method to perform threshold segmentation on the input image to obtain the Light region RegionLight;

[0101] Otsu's method is based on the grayscale histogram of an image, assuming that the pixels in the image can be divided into two categories: foreground (target) and background. By traversing all possible thresholds and calculating the between-class variance of the foreground and background pixel classes for each threshold, the threshold that maximizes the between-class variance is the optimal threshold. The larger the between-class variance, the more obvious the difference between the foreground and the background, and the better the segmentation effect. Otsu's method divides the image pixels into foreground and background classes by determining an optimal threshold, maximizing the variance between these two classes.

[0102] S207: Take out the region RegionLight1 near BoxVal on the left side of the RegionLight dividing line Line and on the right side of BoxVal, and take out the region RegionLight2 near Line on the left side of the RegionLight dividing line Line and on the right side of BoxVal;

[0103] S208: Use opening operation to remove the adhesion of RegionLight1 to obtain RegionLightErosion, and sequentially take out the region RegionGuanOrg(i) with the largest overlapping area with BoxValExpand(i), where BoxValExpand(i) represents the i-th cable text expansion region;

[0104] S209: Use RegionGuanOrg(i) as a seed and use the connected component growth method to obtain the part within the RegionLight2 region, which is RegionGuan(i);

[0105] S210: Use Otsu's method to perform threshold segmentation on the left part of the image Line, take out the Dark part, which is RegionDark, and perform morphological processing on RegionDark to separate it into individual holes to obtain RegionHole;

[0106] S211: Thin RegionGuan to obtain the dilation element2, and perform dilation in the horizontal direction to obtain RegionGuanExpand;

[0107] S212: Sequentially take out the part with the largest overlapping area between the RegionGuanExpand(i) domain and the RegionHole region, denoted as RegionHole(i), and perform horizontal expansion on RegionHole(i), denoted as RegionHoleExpand(i);

[0108] S213: Sequentially take out the part with the largest overlapping area with RegionHoleExpand(i) from BoxTagOrg, denoted as BoxTag(i). The area corresponding to BoxVal(i) is BoxTag(i), and output them in pairs.

[0109] Further, in the step S3, the following steps are included:

[0110] S301: In the step of performing text correction, transform BoxTag and BoxVal into a regular rectangle through affine transformation, and calculate the transformation matrix H;

[0111] S302: Take out the BoxTag and BoxVal parts in ImageColorScale, and use the inverse matrix H_inv of the transformation matrix to correct the image into regular rectangles ImageTagRect and ImageValRect;

[0112] S303: The corrected image is normalized, scaled to 192*48, and sorted by column to obtain Tsesor_Cal;

[0113] S304: Input Tsesor_Cal into the trained MobileNet model, and output the text direction correction result. The output value is 0, no correction is required; 90, rotate 90 degrees clockwise; -90, rotate 90 degrees counterclockwise; 180, flip vertically;

[0114] S305: Rotate ImageTagRect and ImageVal around the center according to their respective output angles to obtain ImageTag and ImageVal.

[0115] Further, in the step S4, the following steps are included:

[0116] S401: Normalize ImageTag and ImageVal, scale them to 320*48, and rearrange the image by column to obtain TensorTag and TensorVal;

[0117] S402: Input TensorTag and TensorVal into the trained SVTR model, and output the index and score;

[0118] S403: Use the index to obtain the characters from the font library dictionary, discard the characters with too low scores, and output them as a string.

[0119] Further, in the step S6, the following steps are included:

[0120] S601: Perform charthreshold segmentation on the table image, and extract the text and wireframe part RegionDark;

[0121] S602: Dilate the Box of the text to obtain BoxDila;

[0122] S603: Calculate the corner points of RegionDark, remove the overlapping part with BoxDila, and obtain haarCell;

[0123] S604: Input haarCell into the SLANet model to obtain the structured relationship of the table, and express it in html form;

[0124] S605: Analyze the table structured relationship in html format, use the input prior knowledge to obtain the table text and table number columns, analyze the regions BoxVal of the table text string and the regions BoxTag of the table number column, and output BoxVal(i) and BoxTag(i) in pairs according to the corresponding order.

[0125] Further, in the step S9, the following steps are included:

[0126] S901: The recognition results of the wire arrangement text and wire arrangement number are LineResultVal and LineResultTag, and the recognition results of the table text and table number are TableResultVal and TableResultTag;

[0127] S902: Based on the table, sequentially extract TableResultVal(i) and TableResultTag(i), search in the wire arrangement detection result LineResultTag. If LineResultTag(j)

[0128] ≠TableResultTag(i), then search for LineResultTag(j + 1) until all of LineResultTag is searched; if LineResultTag(j) = TableResultTag(i), and TableResultVal(i) ≠ LineResultVal(j), mark it as a value inequality defect. If LineResultTag(j)

[0129] =TableResultTag(i), and TableResultVal(i) = LineResultVal(j), mark it as normal and remove it from the candidates to prevent repeated search;

[0130] S903: The item not found in the LineResultTag in the TableResultTag is marked as missing.

[0131] In this embodiment, the visual camera 1 is supported and fixed by the ejector rod 4. By stretching the bottom rod 2, the middle rod 3 and the ejector rod 4, the height of the visual camera 1 can be adjusted telescopically. Meanwhile, when folding and storing, pulling the first base bracket 5 upward or downward can cause the first support leg 7 and the second support leg 8 to complete folding. During the use and support process, the second support leg 8 can be adjusted to be horizontal with the ground and then placed on the ground for use.

[0132] The present invention uses artificial intelligence to identify cable layout text and table text, eliminating the need for manual visual inspection, effectively improving the recognition efficiency of cables, and enhancing the verification accuracy. Through automatic recognition by artificial intelligence, text is quickly read, recognized, and corrected, and the cable layout text and table text are quickly compared, and the text of inconsistent groups is quickly output to achieve the purpose of rapid alarm.

[0133] The above-disclosed is only a preferred embodiment of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. An OCR device based on artificial intelligence recognition, characterized in that it includes a vision camera, a telescopic rod body, a first base bracket, a second base bracket, a first support foot, a second support foot and a triangular support block. The vision camera is fixedly connected to the telescopic rod body and is located at the upper end of the telescopic rod body. The first base bracket is slidably connected to the telescopic rod body and is sleeved on the outer wall of the telescopic rod body. The second base bracket is fixedly connected to the telescopic rod body and is located on the outer wall of the telescopic rod body. The number of the first support feet is multiple, and one end of each first support foot is respectively rotatably connected to the first base bracket and is respectively located on the outer wall of the first base bracket. The number of the triangular support blocks is multiple, and each triangular support block is respectively rotatably connected to the corresponding first support foot and is respectively located at one end of the corresponding first support foot. The number of the second support feet is multiple, and one end of each second support foot is respectively rotatably connected to the second base bracket and is respectively located on the outer wall of the second base bracket, and the other end of each second support foot is respectively rotatably connected to the corresponding triangular support block.

2. The OCR device based on artificial intelligence recognition according to claim 1, characterized in that the telescopic rod body includes a bottom rod, an intermediate rod and a top rod. The first base bracket and the second base bracket are respectively sleeved on the outer wall of the bottom rod. The intermediate rod is slidably connected to the bottom rod and is embedded in the interior of the bottom rod. The top rod is slidably connected to the intermediate rod and is embedded in the interior of the intermediate rod. The vision camera is arranged at the upper end of the top rod.

3. An OCR recognition method based on artificial intelligence recognition, applied to the OCR device based on artificial intelligence recognition as described in claim 2, characterized in that, It includes the following steps: S1: Perform text positioning on the text numbers at the wire arrangement through the vision camera; S2: Perform text correction on the recognized text numbers at the wire arrangement; S3: Perform text recognition on the text numbers at the wire arrangement; S4: Explore the positional relationship of the wire arrangement text and match the corresponding text and numbers; S5: Perform text positioning on the text numbers at the table through the vision camera; S6: Perform text correction on the recognized text numbers at the table; S7: Perform text recognition on the text numbers at the table; S8: Output the structured relationship of the text numbers at the table; S9: Compare the text and numbers matched by the wire arrangement with the text numbers in the table, and highlight and output the inconsistent text numbers.

4. The OCR recognition method based on artificial intelligence recognition according to claim 3, characterized in that in the step S1, it includes the following steps: S101: Calculate through an image enhancement algorithm, including contrast processing and color domain processing, to make the text clearer and reduce adhesion; S102: Adapt the image to the input of the text detection framework. For text detection, use the deep relationship reasoning graph network DRRG, model each character as a small rectangle, and then use the graph convolutional neural network to model the relationship between each character; S103: Finally, output the position coordinates of the four corners of the minimum bounding box of the text line.

5. The OCR recognition method based on artificial intelligence recognition according to claim 4, wherein In the step S101, the following steps are included: Calculate the maximum value data_max and minimum value data_min of the pixels of the image Image. The exponential stretching constant C = pow((data_max - data_min), erp), and the value of erp is 1.5 - 2; Map each channel of the image to a new domain ImageNew, ImageNew(i,j) = pow((erpImage(i,j) - data_min), erp) / C * 255; The color domain transformation includes image white balance and color transformation. ImageWB is the image after white balance processing of ImageNew, ImageColorScale is the image after color transformation, ImageColorScale(k) = ImageWB(k) * a + b, where a and b are constants, and k is the RGB three-component index value. When k = 0, perform blue component transformation on ImageWB, a takes 0.485, and b takes 0.229; when k = 1, perform green component transformation on ImageWB, a takes 0.456, and b takes 0.224; when k = 2, perform red component transformation on ImageWB, a takes 0.406 and b takes 0.225; In the step S102, the following steps are included: Normalize the figure to the 0 - 1 interval, resize the pixels to 960 * 960 pixels, and rearrange them by column to generate a tensor Tensor; Input Tensor into the DRRG model and output the Box of a single-line character.

6. The OCR recognition method based on artificial intelligence recognition according to claim 3, wherein In the step S2, the following steps are included: S201: Sort the Boxes of the input text area in ascending order according to the x coordinate of the upper left corner point; S202: Perform elliptical area detection on the image and calculate the positions of the screw holes; S203: Draw a dividing line Line from the center point of the screw hole area, take out the BoxLeft on the left side of the dividing line, sort BoxLeft in ascending order according to the y coordinate of the upper left corner point to obtain BoxVal, and record the remaining part as BoxDiff; S204: Divide the BoxDiff area again, perform morphological processing, and take out a certain area on the left side of Line as BoxTagOrg; S205: Refine BoxVal to obtain an expansion element, perform horizontal expansion operation on BoxVal to obtain BoxValExpand; S206: Use the Otsu method to perform threshold segmentation on the input image to obtain the Light area RegionLight; S207: Take out the area RegionLight1 near BoxVal on the left side of the RegionLight dividing line Line and on the right side of BoxVal, and take out the area RegionLight2 near Line on the left side of the RegionLight dividing line Line and on the right side of BoxVal; S208: Use opening operation to remove the adhesion of RegionLight1 to obtain RegionLightErosion, and sequentially take out the region RegionGuanOrg(i) with the largest overlapping area with BoxValExpand(i). BoxValExpand(i) represents the i-th expanded area of the cable text; S209: Use RegionGuanOrg(i) as a seed and use the connected domain growth method to obtain the part within the RegionLight2 area, which is RegionGuan(i); S210: Use the Otsu method to perform threshold segmentation on the left part of the image Line, take out the Dark part, denoted as RegionDark, and perform morphological processing on RegionDark to separate it into individual holes to obtain RegionHole; S211: Thin RegionGuan to obtain the dilation element2, and perform dilation in the horizontal direction to obtain RegionGuanExpand; S212: Sequentially take out the part with the largest overlapping area between the RegionGuanExpand(i) domain and the RegionHole area, denoted as RegionHole(i), and expand RegionHole(i) in the horizontal direction, denoted as RegionHoleExpand(i); S213: Sequentially take out the part with the largest overlapping area with RegionHoleExpand(i) from BoxTagOrg, denoted as BoxTag(i). The area corresponding to BoxVal(i) is BoxTag(i), and output them in pairs.

7. A method for OCR recognition based on artificial intelligence recognition according to claim 3, characterized in that In the step S3, the following steps are included: S301: In the step of performing text correction, transform BoxTag and BoxVal into a regular rectangle through affine transformation, and calculate the transformation matrix H; S302: Take out the BoxTag and BoxVal parts in ImageColorScale, and use the inverse matrix H_inv of the transformation matrix to correct the image into a regular rectangle ImageTagRect and ImageValRect; S303: The corrected image is normalized and scaled to 192*48, and sorted by column to obtain Tsesor_Cal; S304: Input Tsesor_Cal into the trained MobileNet model, output the text direction correction result, and the output value is 0, no correction is required; 90, rotate 90 degrees clockwise; -90 rotate 90 degrees counterclockwise; 180, flip vertically; S305: Rotate ImageTagRect and ImageValRect around the center according to their respective output angles to obtain ImageTag and ImageVal.

8. An OCR recognition method based on artificial intelligence recognition according to claim 3, characterized in that In the step S4, the following steps are included: S401: Normalize ImageTag and ImageVal, scale them to 320 * 48, and rearrange the image by columns to obtain TensorTag and TensorVal; S402: Input TensorTag and TensorVal into the trained SVTR model to output indexes and scores; S403: Use the indexes to obtain characters from the font library dictionary, discard the characters with too low scores, and output them as a string.

9. An OCR recognition method based on artificial intelligence recognition according to claim 3, characterized in that In the step S6, the following steps are included: S601: Perform charthreshold segmentation on the table image, and extract the text and wireframe part RegionDark; S602: Dilate the Box of the text to obtain BoxDila; S603: Calculate the corner points of RegionDark, remove the overlapping part with BoxDila, and obtain haarCell; S604: Input haarCell into the SLANet model to obtain the structured relationship of the table, and express it in html form; S605: Analyze the structured relationship of the html-formatted table, use the input prior knowledge to obtain the table text and the table number column, analyze the region BoxVal of the table text string and the region BoxTag of the table number column, and output BoxVal(i) and BoxTag(i) in the corresponding order in pairs.

10. An OCR recognition method based on artificial intelligence recognition according to claim 3, characterized in that In the step S9, the following steps are included: S901: The recognition results of the wire arrangement text and the wire arrangement number are LineResultVal and LineResultTag, and the recognition results of the table text and the table number are TableResultVal and TableResultTag; S902: Taking TableResultVal(i) and TableResultTag(i) sequentially with the table as the reference, search in the cable detection result LineResultTag. If LineResultTag(j) ≠ TableResultTag(i), then search for LineResultTag(j + 1) until all of LineResultTag has been searched; if LineResultTag(j) = TableResultTag(i) and TableResultVal(i) ≠ LineResultVal(j), mark it as a value inequality defect. If LineResultTag(j) = TableResultTag(i) and TableResultVal(i) = LineResultVal(j), mark it as normal and remove it from the candidates to prevent repeated searches; S903: If no same item is found in LineResultTag in TableResultTag, mark it as missing.