Intelligent identification method and system for complex images
By constructing a segmentation scheme scoring function to obtain the optimal segmentation information, the text recognition of circuit diagrams or CAD drawings is segmented, enhanced, and recombined, solving the problem of complex image text recognition, improving recognition accuracy and completeness, and supporting equipment maintenance and fault diagnosis in the coal industry.
Patent Information
- Application Number
- CN202510608516.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Traditional text recognition methods struggle to accurately capture text information in complex and ever-changing circuit diagrams or CAD drawings, resulting in poor recognition performance.
The process involves acquiring an image to be recognized, which may be a circuit diagram or a computer-aided design (CAD) drawing; obtaining target block information; constructing a block scoring function based on the text density and/or background interference level of the image blocks obtained after block processing of the image to be recognized according to the block information; solving for the target block information; performing block processing; image enhancement processing; and text recognition; and reconstructing the recognized text.
It enables precise segmentation of circuit diagrams or CAD drawings, reduces the difficulty of text recognition, and improves the accuracy and completeness of text recognition. It supports coal industry technicians in quickly and accurately interpreting images, thereby improving the efficiency of equipment maintenance and troubleshooting.
Smart Images

Figure CN120126154B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a complex image intelligent recognition method and system. BACKGROUND
[0002] Some industries such as the coal industry gather a large number of precision mechanical equipment, the smooth operation of these equipment cannot be separated from the strong support of the electrical control system, therefore, the coal industry has generated a large amount of equipment circuit diagram and computer aided design (Computer Aided Design, CAD) diagram, covering coal mining machine, scraper conveyor and other key equipment, these diagrams play a decisive role in coal production, specifically, once the equipment fails, the circuit diagram is like a "navigation diagram" for technicians, which can help technicians quickly lock the problem source and avoid blind troubleshooting, thereby greatly saving time and labor cost. In addition, by accurately identifying the text in the circuit diagram or CAD diagram, the technician can also deeply analyze the fault signal, accurately judge the fault type, and accordingly take targeted maintenance measures to ensure that the equipment can quickly resume normal operation and ensure the continuity and stability of production.
[0003] As can be seen, the text recognition of the circuit diagram or CAD diagram has a decisive position in the coal industry, which is not only an indispensable basis for equipment maintenance and fault troubleshooting, but also a key to equipment modification and upgrading, safety production and energy efficiency improvement, and training and skill improvement, so the coal industry must attach great importance to the text recognition of the circuit diagram or CAD diagram and continuously improve the recognition ability and level to adapt to the increasingly complex and variable equipment environment.
[0004] However, in the complex and variable circuit diagram or CAD diagram, the text information is often presented in various scales, directions and layout forms, and the traditional text recognition method is difficult to accurately capture the text information when facing such diagrams, and the text recognition effect is poor. SUMMARY
[0005] The present application aims to at least solve one of the technical problems in the related art to some extent.
[0006] To this end, the first purpose of the present application is to propose a complex image intelligent recognition method.
[0007] The second purpose of the present application is to propose a complex image intelligent recognition system.
[0008] To achieve the above purpose, the first aspect embodiment of the present application proposes a complex image intelligent recognition method, comprising:
[0009] obtaining an image to be recognized, the image to be recognized being a circuit diagram or a computer aided design (CAD) diagram;
[0010] taking the block information used for indicating the block scheme as a variable, constructing a block scheme score function based on the text density and / or the background interference degree in the image block obtained after block processing of the image to be recognized according to the block information, and solving the block scheme score function to obtain target block information;
[0011] performing block processing on the image to be recognized according to the target block information to obtain a plurality of image blocks; wherein the target block information is used for indicating a block scheme;
[0012] performing image enhancement processing on the image block to obtain an enhanced image block, and performing text recognition based on the enhanced image block to obtain recognized text;
[0013] reorganizing the recognized text corresponding to at least one enhanced image block based on the position of the recognized text in the image to be recognized to obtain target text corresponding to the image to be recognized.
[0014] To achieve the above object, a second aspect of the embodiment of the present application provides a complex image intelligent recognition system, comprising:
[0015] a first acquisition module configured to acquire an image to be recognized, wherein the image to be recognized is a circuit diagram or a computer-aided design (CAD) diagram;
[0016] a second acquisition module configured to take block information used for indicating a block scheme as a variable, construct a block scheme score function based on the text density and / or the background interference degree in the image block obtained after block processing of the image to be recognized according to the block information, and solve the block scheme score function to obtain target block information;
[0017] a block processing module configured to perform block processing on the image to be recognized according to the target block information to obtain a plurality of image blocks; wherein the target block information is used for indicating a block scheme;
[0018] an image processing module configured to perform image enhancement processing on the image block to obtain an enhanced image block, and perform text recognition based on the enhanced image block to obtain recognized text;
[0019] a reorganization module configured to reorganize the recognized text corresponding to at least one enhanced image block based on the position of the recognized text in the image to be recognized to obtain target text corresponding to the image to be recognized.
[0020] The technical scheme provided by the present application at least brings the following beneficial effects:
[0021] The application obtains a to-be-recognized image, the to-be-recognized image is a circuit diagram or a computer-aided design (CAD) diagram; obtains target block information, and performs block processing on the to-be-recognized image according to the target block information to obtain a plurality of image blocks; wherein the target block information is used to indicate a block scheme; for any image block, image enhancement processing is performed on the image block to obtain an enhanced image block, and text recognition is performed based on the enhanced image block to obtain recognized text; based on the position of the recognized text in the to-be-recognized image, the recognized text corresponding to at least one enhanced image block is reorganized to obtain target text corresponding to the to-be-recognized image. By constructing and solving a block scheme score function, optimal block information is obtained, accurate block of the circuit diagram or the CAD diagram is realized, which helps to reduce the difficulty of text recognition of complex images; based on the enhanced image block, the recognized text is reorganized to obtain the target text, which helps to improve the accuracy and integrity of text recognition. In addition, in the coal industry, by using advanced image processing technology and deep learning algorithm, accurate and efficient recognition of image text information is realized, which can provide strong support for technical personnel in the coal industry, so that they can more quickly and accurately interpret the circuit diagram or the CAD diagram, thereby improving the efficiency of equipment maintenance and fault troubleshooting, and escorting the safe production and efficient operation of the coal industry.
[0022] Additional aspects and advantages of the application will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following and / or can be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0023] The above and / or additional aspects and advantages of the application will become apparent and be readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:
[0024] Figure 1 A flowchart of a complex image intelligent recognition method provided by an embodiment of the application;
[0025] Figure 2 A flowchart of a complex image intelligent recognition method provided by another embodiment of the application;
[0026] Figure 3 A structural diagram of a complex image intelligent recognition system provided by an embodiment of the application. DETAILED DESCRIPTION
[0027] Embodiments of the application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the accompanying drawings are exemplary and are intended to explain the application, and cannot be understood as limiting the application.
[0028] A complex image intelligent recognition method and system are described below with reference to the accompanying drawings.
[0029] Figure 1 A flowchart of a complex image intelligent recognition method provided by an embodiment of the present application is shown in Figure 1 The complex image intelligent recognition method includes the following steps:
[0030] In step 101, an image to be recognized is obtained. The image to be recognized is a circuit diagram or a computer-aided design (CAD) diagram.
[0031] In the embodiment of the present application, the image to be recognized can be a circuit diagram or a CAD diagram including text, or a simple image including text, such as a document screenshot or a photographed image including text.
[0032] It should be noted that, in addition to the circuit diagram and the CAD diagram, the image to be recognized can also be an image including a large number of lines, symbols, and text.
[0033] In step 102, a block information indicating a block scheme is taken as a variable, a block scheme score function is constructed based on a text density and / or a background interference degree in an image block obtained by performing block processing on the image to be recognized according to the block information, and the block scheme score function is solved to obtain target block information.
[0034] In the embodiment of the present application, the block information is used to indicate the number of image blocks or the size of image blocks. For example, the block information can include the number of block rows and the number of block columns, or information such as the number of blocks or the size of blocks. The text density can be used to indicate the text distribution, and the background interference degree can be used to indicate the image complexity of a non-text region in the image block.
[0035] In an optional embodiment, the block scheme score function can be constructed based on the text density in the image block obtained by performing block processing on the image to be recognized according to the block information.
[0036] In an optional embodiment, the block scheme score function can be constructed based on the background interference degree in the image block obtained by performing block processing on the image to be recognized according to the block information.
[0037] For example, the text density mean value can be calculated based on the text density in at least one image block obtained by performing block processing on the image to be recognized according to the block information, to obtain the block scheme score function; or the background interference degree mean value can be calculated based on the background interference degree in at least one image block obtained by performing block processing on the image to be recognized according to the block information, to obtain the block scheme score function.
[0038] In an optional embodiment, the block scheme score function can also be constructed based on the text density and the background interference degree in the image block obtained after the to-be-recognized image is processed according to the block information.
[0039] The block scheme score function is a target function to be solved, and the block information has a corresponding constraint condition. When the block scheme score function is solved, the target block information can be solved in combination with the constraint condition.
[0040] In step 103, the to-be-recognized image is processed according to the target block information to obtain a plurality of image blocks; wherein the target block information is used to indicate a block scheme.
[0041] As an example, the to-be-recognized image can be uniformly processed according to the target block information to obtain a plurality of image blocks with the same size.
[0042] In order to integrate the text corresponding to each image block and ensure the integrity of the text information, there can be an overlapping area between adjacent image blocks. The length of the overlapping area is denoted as , and the width is denoted as . It should be noted that the overlap only occurs between adjacent image blocks, and each image block overlaps with the image blocks above, below, left and right.
[0043] In step 104, for any image block, the image block is processed for image enhancement to obtain an enhanced image block, and text recognition is performed based on the enhanced image block to obtain recognized text.
[0044] In the circuit diagram or CAD diagram, a large number of lines, symbols and text descriptions are usually included. The image enhancement of the image block can be targeted to optimize the local details, such as improving the clarity and contrast of lines, symbols and text, removing noise, avoiding blurring of details caused by overall processing, and thus helping to improve the text recognition effect.
[0045] In the embodiments of the present application, the text recognition model can be used to perform text recognition on the enhanced image block to obtain the recognized text.
[0046] In step 105, the recognized text corresponding to at least one enhanced image block is reorganized based on the position of the recognized text in the to-be-recognized image to obtain the target text corresponding to the to-be-recognized image.
[0047] In the embodiments of the present application, the recognized text corresponding to at least one enhanced image block can be spliced and integrated to obtain the target text corresponding to the to-be-recognized image.
[0048] In this embodiment, a to-be-recognized image is obtained, and the to-be-recognized image is a circuit diagram or a computer-aided design (CAD) diagram; a block information indicating a block scheme is taken as a variable, a block scheme score function is constructed based on text density and / or background interference degree in an image block obtained after the to-be-recognized image is processed according to the block information, and the block scheme score function is solved to obtain target block information, and the to-be-recognized image is processed according to the target block information to obtain a plurality of image blocks; the target block information is used to indicate the block scheme; for any image block, image enhancement processing is performed on the image block to obtain an enhanced image block, and text recognition is performed based on the enhanced image block to obtain recognized text; and the recognized text in the to-be-recognized image is used to reorganize the recognized text corresponding to at least one enhanced image block to obtain target text corresponding to the to-be-recognized image. By constructing and solving the block scheme score function, optimal block information is obtained, accurate block of the circuit diagram or the CAD diagram is realized, and the text recognition difficulty of a complex image is reduced. The text is recognized based on the enhanced image block, and the target text is obtained by reorganization, which helps to improve the text recognition accuracy and completeness.
[0049] The embodiment provides another complex image intelligent recognition method, Figure 2 A flowchart of the complex image intelligent recognition method provided in the embodiment of the application is shown in FIG. 1.
[0050] As shown in FIG. 1, the complex image intelligent recognition method can include the following steps: Figure 2
[0051] In step 201, a to-be-recognized image is obtained.
[0052] In step 202, for any image block obtained after the to-be-recognized image is processed according to block information, a difference between text density and background interference degree in the image block is obtained; a difference mean value is calculated based on the difference corresponding to at least one image block obtained after the to-be-recognized image is processed according to the block information, to obtain a block scheme score function, and the block scheme score function is solved to obtain target block information.
[0053] In the embodiment of the application, the block information can include a block row number and a block column number .
[0054] The target block information is determined according to the complexity and the text distribution of the to-be-recognized image, which can ensure that the background interference in each image block is reduced as much as possible, and more text information is contained.
[0055] As an example, the block scheme score function can be as shown in the following formula:
[0056]
[0057] in, This indicates the score for the block partitioning scheme. Indicates the first Line number Text density in the image blocks of the column, Indicates the first Line number The degree of background interference in the image blocks of the column, α Indicates the weighting coefficient. The range of values for is [1, ...]. ], The range of values for is [1, ...]. It should be noted that, in the formula, and For variables.
[0058] In some embodiments, the number of rows and columns in a block are subject to value constraints.
[0059] In some embodiments, the maximum value of the partitioning scheme score function is obtained by combining the value constraints corresponding to the number of partitioned rows and columns, thereby obtaining the target partitioning information, wherein the target partitioning information includes the target number of partitioned rows and the target number of partitioned columns. The value constraints can refer to a range of values.
[0060] In this embodiment, the maximum value of the segmentation scheme score function is calculated, which is to maximize the segmentation scheme score function (i.e., maximize the text density). Minimize background interference Using ) as the objective, the target block information is solved.
[0061] In some embodiments, the target segmentation information can be obtained by aiming to achieve a set score threshold for the segmentation scheme score.
[0062] When iteratively solving the block segmentation scheme scoring function, the mean or median text density can be obtained. Taking the median text density as an example, if the median text density is less than the first density threshold (image block too large or text too little), then the density is increased. and Value; if the median text density is greater than the second density threshold (image patch too small or text too dense), then decrease it. and Value, until the optimal value is obtained. and The iteration stops. The first density threshold is less than the second density threshold.
[0063] It should be noted that the block information has an initial value, that is, in the above formula. and has an initial value, which can be calculated as an example by the following formula and :
[0064]
[0065]
[0066] wherein, denotes the initial value of , denotes the initial value of , denotes the length of the image to be recognized, denotes the width of the image to be recognized, AvgTextWidth denotes the average width of the text area, and AvgTextHeight denotes the average height of the text area. is an empirical coefficient for ensuring that the image block can accommodate multiple text areas.
[0067] In some embodiments, for any image block obtained after the image to be recognized is processed according to the block information, the text area in the image block is obtained, and the text density is determined based on the ratio of the text area to the area of the image block; and / or, for any image block obtained after the image to be recognized is processed according to the block information, the non-text area in the image block is obtained, and the number of edge pixels obtained through edge detection in the non-text area is obtained, and the background interference degree is determined based on the ratio of the number of edge pixels to the area of the non-text area.
[0068] As an example, the text density in the image block can be calculated by the following formula:
[0069]
[0070] wherein, TextArea( cell ij ) denotes the area of the text area in the image block at the th row and the th column, denotes the length of the image block at the th row and the th column, denotes the width of the image block at the th row and the th column.
[0071] As an example, the background interference degree in the image block can be calculated by the following formula:
[0072]
[0073] wherein, EdgePixels(cell ij ) represents the number of edge pixels extracted by edge detection on the image block in the row and the column, TextEdgePixels( cell ij ) represents the number of edge pixels extracted in the text region of the image block in the row and the column, in the above formula, the numerator represents the number of edge pixels obtained by edge detection in the non-text region of the image block in the row and the column, and the denominator represents the area of the non-text region in the image block in the row and the column.
[0074] Step 203, according to the target block information, the to-be-recognized image is block processed to obtain a plurality of image blocks.
[0075] Among them, after the to-be-recognized image is block processed according to the target block information, the adjacent image blocks have an overlapping area.
[0076] Suppose that the target block information obtained by solving the block scheme score function includes the number of rows and the number of columns , then the size of the image block considering the overlapping area is:
[0077]
[0078]
[0079] Step 204, for any image block, the image block is subjected to image enhancement processing to obtain an enhanced image block, and text recognition is performed based on the enhanced image block to obtain recognized text.
[0080] In some embodiments, the image block is subjected to filtering processing by a low-pass filter to obtain a blurred image block; based on the image block and the blurred image block, a detail image block is obtained; the detail image block is subjected to image enhancement to obtain an enhanced detail image block; and based on the image block and the enhanced detail image block, an enhanced image block is obtained.
[0081] Among them, the specific implementation manner of the image enhancement processing of the image block is:
[0082] (1) applying a low-pass filter to the image block to generate a blurred image block as an anti-sharpening mask reference; specifically, , wherein represents the original image block, represents the low-pass filter, represents the blurred image block.
[0083] (2) Subtract the blurred image block from the original image block to obtain a detail image block (a high-frequency image block), specifically, wherein, represents the detail image block.
[0084] (3) Multiply the detail image block by an enhancement factor to adjust the intensity of enhancement, specifically, wherein, represents the enhancement factor, is the enhanced detail image block.
[0085] (4) Add the enhanced detail image block to the original image block to obtain a final enhanced image block, specifically, wherein represents the final enhanced image block, which retains the low-frequency components of the original image block and enhances the high-frequency components (edges and details).
[0086] In some embodiments, based on a text detection model, text detection is performed on the enhanced image block to obtain a text region; based on a text recognition model, text recognition is performed on the text region to obtain recognized text.
[0087] wherein, the text region can refer to a region where text in the enhanced image block is located; detecting the text region in the enhanced image block first, and then performing text recognition on the text region, can accurately locate the text position and reduce background interference, thereby improving the reliability, accuracy and efficiency of text recognition.
[0088] Before text detection is performed on the enhanced image block based on the text detection model, feature extraction can be performed on the enhanced image block based on a convolutional neural network, and then the text detection model is used to perform text detection based on the extracted features.
[0089] wherein, the formula for feature extraction by the convolutional neural network is T = f(W w *t+b), wherein t is the input enhanced image block, W w is a convolution kernel, b is a bias term, and T is an output feature map.
[0090] In some embodiments, the text detection model is a Feature Pyramid Networks (FPN) model based on convolution, and / or the text recognition model is an Attention-based model.
[0091] The FPN captures text regions of different sizes by constructing feature maps of different scales, and a prediction layer is arranged on each scale of feature map to output the detection result of the text region. The prediction layer can include a convolution layer and a classification layer (or a regression layer) to generate the coordinates of the text region bounding box.
[0092] The core of the Attention-based model is the calculation formula of the Attention mechanism. The Attention mechanism determines the weight of each feature vector by calculating the similarity (such as cosine similarity) between the current state of the decoder and the encoder feature vector, and then performs weighted summation on the feature vectors according to the weights to obtain the context vector of the current time of the decoder. The decoder generates the prediction of the next character according to the context vector and the current state.
[0093] It should be noted that during the training of the text detection model and the text recognition model, the recognition result of the model can be evaluated by using evaluation indexes such as accuracy and recall:
[0094] Accuracy = (TP + TN) / (TP + FP + FN + TN)
[0095] Recall = TP / (TP + FN)
[0096] True Positives (TP) : The number of correctly predicted positive samples by the model;
[0097] True Negatives (TN) : The number of correctly predicted negative samples by the model;
[0098] False Positives (FP) : The number of incorrectly predicted positive samples by the model;
[0099] False Negatives (FN) : The number of incorrectly predicted negative samples by the model.
[0100] In step 205, the identified text is reorganized based on the position of the identified text in the to-be-recognized image to obtain the target text corresponding to the to-be-recognized image.
[0101] In some embodiments, for any identified text, a first position of the identified text in the corresponding enhanced image block is obtained, and a second position of the enhanced image block in the to-be-recognized image is obtained; the second position is corrected according to the size of the overlapping region to obtain a corrected second position; and the identified texts are reorganized based on the first position and the corrected second position to obtain the target text.
[0102] The reorganization of the identified text can realize integration of the identified text, so that the target text is complete and smooth.
[0103] The second position of the image block in the i-th row and the j-th column is (i, j). The second position of the image block in the i-th row and the j-th column is (i, j). The second position of the image block in the i-th row and the j-th column is (i, j). , , where i represents a row number (1≤i≤n), j represents a column number (1≤j≤m), and the second position of the image block in the i-th row and the j-th column is (i, j). i , , j , j , The second position can be corrected according to the size of the overlapping area by the following formula:
[0104]
[0105] It is assumed that the overlapping areas are uniformly distributed between adjacent cells, and only the overlaps with the left and right cells are considered. The corrected absolute starting horizontal coordinate is represented by xstart. The left and right boundary conditions are used to ensure that the first column and the last column do not exceed the boundary due to overlapping.
[0106]
[0107] It is assumed that the overlapping areas are uniformly distributed between adjacent cells, and only the overlaps with the top and bottom cells are considered. The corrected absolute starting vertical coordinate is represented by ystart. The top and bottom boundary conditions are used to ensure that the first row and the last row do not exceed the boundary due to overlapping.
[0108] The second position of the image block in the i-th row and the j-th column is (i, j). i The second position of the image block in the i-th row and the j-th column is (i, j). j The second position of the image block in the i-th row and the j-th column is (i, j). The second position of the image block in the i-th row and the j-th column is (i, j).
[0109] It should be noted that after obtaining the target text, the target text can be stored. Specifically, each identified text can be stored in association with the position of the corresponding image block (the second position and / or the corrected second position). For example, the identified texts can be stored in association with the row number, the column number, and the corrected second position of the corresponding image block. Alternatively, the identified texts, the first position of the identified texts in the corresponding image block, and the position of the identified texts in the corresponding image block (the second position and / or the corrected second position) can be stored in association.
[0110] Storing the identified texts in association with the corresponding position information can help to locate the actual position of the identified texts in the image to be recognized.
[0111] In some embodiments, the method further comprises: using a large language model, performing a verification and adjustment process on the integrated text to obtain a processed target text; wherein the verification and adjustment process includes at least one of syntax verification processing, error character / error word replacement processing, and text adjustment processing combined with the context semantic information of the integrated text.
[0112] Wherein, using a large language model to check the syntax of the target text can correct possible syntax errors; matching the recognized target text with a predefined dictionary can replace incorrectly recognized characters or words; and according to the context semantic information, further verifying and adjusting the target text can reduce errors and omissions, ensuring the accuracy and reliability of the recognition result.
[0113] It should be noted that after obtaining the target text, it can be manually proofread and analyzed in depth to ensure the accuracy of the text.
[0114] In this embodiment, a to-be-recognized image is obtained; for any image block obtained by performing a blocking process on the to-be-recognized image according to the blocking information, the difference between the text density and the background interference degree in the image block is obtained; based on the difference corresponding to at least one image block obtained by performing a blocking process on the to-be-recognized image according to the blocking information, the difference mean value is calculated to obtain a blocking scheme score function; the blocking scheme score function is solved to obtain target blocking information; the to-be-recognized image is blocked according to the target blocking information to obtain a plurality of image blocks; for any image block, the image block is subjected to image enhancement processing to obtain an enhanced image block, and text recognition is performed based on the enhanced image block to obtain recognized text; based on the position of the recognized text in the to-be-recognized image, the recognized text corresponding to at least one enhanced image block is reorganized to obtain target text corresponding to the to-be-recognized image. The difference between the text density and the background interference degree in the image block is used to construct the blocking scheme score function, and then the target blocking information is solved, which can ensure that the background interference in each image block is reduced as much as possible while containing more text information, thereby realizing fast and accurate text recognition for to-be-recognized images of multiple scales, complex backgrounds, and difficult text extraction.
[0115] To implement the above-mentioned embodiments, an embodiment of the present application further provides a complex image intelligent recognition system.
[0116] Figure 3 A structural schematic diagram of a complex image intelligent recognition system provided by an embodiment of the present application.
[0117] As shown in Figure 3 the complex image intelligent recognition system 300 includes:
[0118] The first obtaining module 310 is configured to obtain a to-be-recognized image, where the to-be-recognized image is a circuit diagram or a computer-aided design (CAD) diagram.
[0119] The second obtaining module 320 is configured to, taking block information indicating a block scheme as a variable, construct a block scheme score function based on a text density and / or a background interference degree in an image block obtained by performing block processing on the to-be-recognized image according to the block information, and solve the block scheme score function to obtain target block information.
[0120] The block processing module 330 is configured to perform block processing on the to-be-recognized image according to the target block information to obtain a plurality of image blocks, where the target block information is used to indicate a block scheme.
[0121] The image processing module 340 is configured to, for any image block, perform image enhancement processing on the image block to obtain an enhanced image block, and perform text recognition based on the enhanced image block to obtain recognized text.
[0122] The reorganization module 350 is configured to reorganize the recognized text corresponding to at least one enhanced image block based on a position of the recognized text in the to-be-recognized image to obtain target text corresponding to the to-be-recognized image.
[0123] Optionally, the second obtaining module 320 is specifically configured to: for any image block obtained by performing block processing on the to-be-recognized image according to the block information, obtain a difference between a text density and a background interference degree in the image block; and based on the difference corresponding to at least one image block obtained by performing block processing on the to-be-recognized image according to the block information, calculate a mean value of the differences to obtain the block scheme score function.
[0124] Optionally, the block information includes a block row number and a block column number, and the block row number and the block column number have value constraint conditions respectively, and the second obtaining module 320 is specifically configured to: in combination with the value constraint conditions, solve a maximum value of the block scheme score function to obtain the target block information, where the target block information includes a target block row number and a target block column number.
[0125] Optionally, the system further includes a third obtaining module configured to:
[0126] for any image block obtained by performing block processing on the to-be-recognized image according to the block information, obtain a text region in the image block, and determine the text density based on a ratio of the text region to an area of the image block; and / or, for any image block obtained by performing block processing on the to-be-recognized image according to the block information, obtain a non-text region in the image block, and obtain a number of edge pixels in the non-text region obtained by edge detection, and determine the background interference degree based on a ratio of the number of edge pixels to an area of the non-text region.
[0127] Optionally, the image processing module 340 is specifically configured to: filter the image block through a low-pass filter to obtain a blurred image block; obtain a detail image block based on the image block and the blurred image block; perform image enhancement on the detail image block to obtain an enhanced detail image block; and obtain an enhanced image block based on the image block and the enhanced detail image block.
[0128] Optionally, the image processing module 340 is specifically configured to: perform text detection on the enhanced image block based on a text detection model to obtain a text region; and perform text recognition on the text region based on a text recognition model to obtain recognized text.
[0129] Optionally, the text detection model is a convolution-based feature pyramid network model, and / or the text recognition model is a model based on an attention mechanism.
[0130] Optionally, the adjacent image blocks have an overlapping region, and the reorganization module 350 is specifically configured to: obtain a first position of the recognized text in the corresponding enhanced image block, and obtain a second position of the enhanced image block in the image to be recognized; correct the second position according to the size of the overlapping region to obtain a corrected second position; and reorganize each recognized text based on the first position and the corrected second position to obtain target text.
[0131] Optionally, the system further includes a verification and adjustment processing module configured to:
[0132] adopting a large language model to perform verification and adjustment processing on the target text to obtain processed target text; wherein the verification and adjustment processing includes at least one of syntax verification processing, error character / error word replacement processing, and text adjustment processing combined with context semantic information of the text.
[0133] It should be noted that the foregoing explanation and description of the embodiment of the complex image intelligent recognition method also apply to the embodiment of the complex image intelligent recognition system, which will not be described here.
[0134] In the embodiments of the present application, a to-be-recognized image is obtained, the to-be-recognized image is a circuit diagram or a computer-aided design (CAD) diagram; target block information is obtained, and the to-be-recognized image is subjected to block processing according to the target block information, to obtain a plurality of image blocks; wherein the target block information is used to indicate a block scheme; for any image block, the image block is subjected to image enhancement processing, to obtain an enhanced image block, and text recognition is performed based on the enhanced image block, to obtain recognized text; and the recognized text in the to-be-recognized image is used as a basis to recombine the recognized text corresponding to at least one enhanced image block, to obtain target text corresponding to the to-be-recognized image. By constructing and solving a block scheme score function, optimal block information is obtained, precise block of the circuit diagram or the CAD diagram is realized, and this helps to reduce the difficulty of text recognition of complex images; the text is recognized based on the enhanced image block, and then the target text is obtained by recombination, which helps to improve the accuracy and completeness of text recognition.
Claims
1. A method for intelligent recognition of complex images, characterized in that, Includes the following steps: Acquire an image to be identified, wherein the image to be identified is a circuit diagram or a computer-aided design (CAD) drawing; Using block information indicating the block segmentation scheme as variables, a block segmentation scheme scoring function is constructed based on the text density and / or background interference level in the image blocks obtained after the image to be identified is segmented according to the block information. The block segmentation scheme scoring function is then solved to obtain the target block information. The block information includes the number of block rows and the number of block columns. The number of block rows and the number of block columns are respectively subject to value constraints. The maximum value of the block segmentation scheme scoring function is then solved by combining the value constraints. The scoring function for the block-based scheme is: This indicates the score for the block partitioning scheme. Indicates the first Line 1 Text density in the image blocks of a column, Indicates the first Line 1 The degree of background interference in the image blocks of the column, α Indicates the weighting coefficient. The range of values for is [1, ...]. ], The range of values for is [1, ...]. ], and As a variable, and and It has an initial value; Based on the target segmentation information, the image to be identified is segmented to obtain multiple image blocks; wherein, the target segmentation information is used to indicate the segmentation scheme; For any of the aforementioned image blocks, image enhancement processing is performed on the image blocks to obtain enhanced image blocks, and text recognition is performed based on the enhanced image blocks to obtain recognized text; Based on the position of the identified text in the image to be identified, the identified text corresponding to at least one enhanced image block is reconstructed to obtain the target text corresponding to the image to be identified.
2. The method according to claim 1, characterized in that, The step of constructing a segmentation scheme scoring function based on the text density and / or background interference level in the image blocks obtained after segmenting the image to be recognized according to the segmentation information includes: For any image block obtained after dividing the image to be identified according to the block information, the difference between the text density and the degree of background interference in the image block is obtained; Based on the differences corresponding to at least one image block obtained after dividing the image to be identified into blocks according to the block information, the mean difference is calculated to obtain the block division scheme score function.
3. The method according to claim 2, characterized in that, The process of solving the block segmentation scheme score function to obtain the target block information includes: The target block information is obtained based on the maximum value of the scoring function of the block segmentation scheme, wherein the target block information includes the number of target block rows and the number of target block columns.
4. The method according to claim 1, characterized in that, The step of performing image enhancement processing on the image block to obtain an enhanced image block includes: The image block is filtered by a low-pass filter to obtain a blurred image block; Based on the image patch and the blurred image patch, obtain the detail image patch; Image enhancement is performed on the detailed image block to obtain the enhanced detailed image block; The enhanced image block is obtained based on the image block and the enhanced detail image block.
5. The method according to claim 1, characterized in that, The text recognition based on the enhanced image patch to obtain the recognized text includes: Based on the text detection model, text detection is performed on the enhanced image block to obtain the text region; Based on the text recognition model, text recognition is performed on the text region to obtain the recognized text.
6. The method according to claim 5, characterized in that, The text detection model is a convolutional feature pyramid network model, and / or the text recognition model is an attention-based model.
7. The method according to claim 1, characterized in that, Adjacent image patches have overlapping areas. The step of reconstructing the recognized text corresponding to at least one enhanced image patch based on the position of the recognized text in the image to be recognized, to obtain the target text corresponding to the image to be recognized, includes: For any of the identified texts, obtain the first position of the identified text in the corresponding enhanced image block, and obtain the second position of the enhanced image block in the image to be identified; The second position is corrected according to the size of the overlapping area to obtain the corrected second position; Based on the first position and the corrected second position, the identified texts are spliced and recombined to obtain the target text.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: A large language model is used to perform verification and adjustment processing on the target text to obtain the processed target text; The verification and adjustment process includes at least one of the following: syntax verification processing, error character / error word replacement processing, and text adjustment processing based on the contextual semantic information of the integrated text.
9. A complex image intelligent recognition system, characterized in that, include: The first acquisition module is used to acquire an image to be identified, wherein the image to be identified is a circuit diagram or a computer-aided design (CAD) drawing. The second acquisition module is used to construct a segmentation scheme scoring function based on the text density and / or background interference level in the image blocks obtained after segmenting the image to be identified according to the segmentation information as a variable, and to solve the segmentation scheme scoring function to obtain target segmentation information. The segmentation information includes the number of segmentation rows and the number of segmentation columns. The number of segmentation rows and the number of segmentation columns are respectively subject to value constraints. The maximum value of the segmentation scheme scoring function is solved by combining the value constraints. The scoring function for the block-based scheme is: This indicates the score for the block partitioning scheme. Indicates the first Line 1 Text density in the image blocks of a column, Indicates the first Line 1 The degree of background interference in the image blocks of the column, α Indicates the weighting coefficient. The range of values for is [1, ...]. ], The range of values for is [1, ...]. ], and As a variable, and and It has an initial value; The block processing module is used to divide the image to be identified into blocks according to the target block information to obtain multiple image blocks; wherein, the target block information is used to indicate the block scheme; An image processing module is configured to perform image enhancement processing on any of the image blocks to obtain an enhanced image block, and to perform text recognition based on the enhanced image block to obtain recognized text; The recombination module is used to reconstruct the recognized text corresponding to at least one enhanced image block based on the position of the recognized text in the image to be recognized, so as to obtain the target text corresponding to the image to be recognized.
Citation Information
Patent Citations
Super-large format document intelligent identification method and system based on block parallelism
CN119625766A