Complex image intelligent identification method and system
By constructing the blocking scheme score function and performing image enhancement processing in complex image intelligent recognition methods, the problem of text recognition difficulty in complex circuit diagrams or CAD diagrams is solved, and the text recognition effect with high accuracy and completeness is achieved.
Patent Information
- Application Number
- CN202510608516.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
In complex and changeable circuit diagrams or CAD diagrams, text information is presented in a variety of scales, directions and layouts. Traditional text recognition methods are difficult to accurately capture text information, resulting in poor recognition effects.
A complex image intelligent recognition method is proposed. By obtaining the image to be identified, a blocking scheme score function is constructed, blocking processing is performed according to the text density and background interference degree, image enhancement and text recognition are performed, and the target text is finally reorganized.
It realizes accurate blocking and text recognition of circuit diagrams or CAD diagrams, improves the accuracy and completeness of text recognition, and is suitable for equipment maintenance and troubleshooting of complex equipment environments such as the coal industry.
Smart Images

Figure CN120126154A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to a complex image intelligent recognition method and system. Background Art
[0002] Some industries, such as the coal industry, have gathered a large number of precision mechanical equipment. The smooth operation of these equipment is inseparable from the strong support of the electrical control system. Therefore, the coal industry has generated a huge amount of equipment circuit diagrams and Computer Aided Design (CAD) diagrams, covering a variety of key equipment such as shearers and scraper conveyors. These diagrams play a crucial role in coal production. Specifically, once a device fails, the circuit diagram is like a "navigation map" for technicians, which can help technicians quickly lock down the root cause of the problem, avoid unstructured troubleshooting, and thus greatly save time and labor costs. In addition, by accurately identifying the text in the circuit diagram or CAD diagram, technicians can also deeply analyze the fault signals, accurately judge the type of fault, and accordingly take targeted maintenance measures to ensure that the equipment can quickly return to normal operation and guarantee the continuity and stability of production.
[0003] It can be seen that the text recognition of circuit diagrams or CAD diagrams plays a crucial role in the coal industry. It is not only an indispensable foundation for equipment maintenance and fault troubleshooting, but also the key to equipment transformation and upgrading, safety production and energy efficiency improvement, as well as training and skill improvement. Therefore, the coal industry must attach great importance to the text recognition of circuit diagrams or CAD diagrams, and continuously improve the ability and level of diagram recognition to adapt to the increasingly complex and changeable equipment environment.
[0004] However, in complex and changeable circuit diagrams or CAD diagrams, text information often presents in diverse scales, directions, and layout forms. Traditional text recognition methods are difficult to accurately capture text information when facing such diagrams, and the text recognition effect is poor. Summary of the Invention
[0005] This application aims to solve at least one of the technical problems in the related art to some extent.
[0006] To this end, the first object of this application is to propose a complex image intelligent recognition method.
[0007] The second object of this application is to propose a complex image intelligent recognition system.
[0008] To achieve the above object, the first aspect embodiment of this application proposes a complex image intelligent recognition method, including: Obtain an image to be recognized, where the image to be recognized is a circuit diagram or a Computer Aided Design (CAD) diagram; Taking the chunking information used to indicate the chunking scheme as a variable, based on the text density and / or background interference degree in the image chunks obtained after chunking the image to be recognized according to the chunking information, constructing a chunking scheme scoring function, and solving the chunking scheme scoring function to obtain the target chunking information; Chunking the image to be recognized according to the target chunking information to obtain a plurality of image chunks; wherein, the target chunking information is used to indicate the chunking scheme; For any one of the image chunks, performing image enhancement processing on the image chunk to obtain an enhanced image chunk, and performing text recognition based on the enhanced image chunk to obtain recognized text; Based on the position of the recognized text in the image to be recognized, reorganizing the recognized text corresponding to at least one enhanced image chunk to obtain the target text corresponding to the image to be recognized.
[0009] To achieve the above object, an embodiment of the second aspect of the present application provides a complex image intelligent recognition system, including: A first acquisition module, configured to acquire an image to be recognized, where the image to be recognized is a circuit diagram or a computer-aided design (CAD) drawing; A second acquisition module, configured to take the chunking information used to indicate the chunking scheme as a variable, based on the text density and / or background interference degree in the image chunks obtained after chunking the image to be recognized according to the chunking information, construct a chunking scheme scoring function, and solve the chunking scheme scoring function to obtain the target chunking information; A chunking processing module, configured to chunk the image to be recognized according to the target chunking information to obtain a plurality of image chunks; wherein, the target chunking information is used to indicate the chunking scheme; An image processing module, configured to, for any one of the image chunks, perform image enhancement processing on the image chunk to obtain an enhanced image chunk, and perform text recognition based on the enhanced image chunk to obtain recognized text; A reorganization module, configured to reorganize the recognized text corresponding to at least one enhanced image chunk based on the position of the recognized text in the image to be recognized to obtain the target text corresponding to the image to be recognized.
[0010] The technical solution provided by the present application at least brings the following beneficial effects: This application obtains an image to be recognized, where the image to be recognized is a circuit diagram or a computer-aided design (CAD) drawing; obtains target block information, and performs block processing on the image to be recognized according to the target block information to obtain a plurality of image blocks; wherein, the target block information is used to indicate a block scheme; for any image block, performs image enhancement processing on the image block to obtain an enhanced image block, and performs text recognition based on the enhanced image block to obtain recognized text; based on the position of the recognized text in the image to be recognized, reorganizes the recognized text corresponding to at least one enhanced image block to obtain the target text corresponding to the image to be recognized. By constructing and solving a block scheme score function to obtain the optimal block information, accurate block division of the circuit diagram or CAD drawing is achieved, which helps to reduce the difficulty of text recognition for complex images; recognizing text based on the enhanced image block and then reorganizing to obtain the target text helps to improve the accuracy and integrity of text recognition. Additionally, in the coal industry, by adopting advanced image processing technologies and deep learning algorithms to achieve accurate and efficient recognition of text information in images, it can provide strong support for technical personnel in the coal industry, enabling them to interpret circuit diagrams or CAD drawings more quickly and accurately, thereby improving the efficiency of equipment maintenance and fault troubleshooting, and ensuring the safe production and efficient operation of the coal industry.
[0011] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and / or additional aspects and advantages of this application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where: Figure 1 is a schematic flowchart of a complex image intelligent recognition method provided by an embodiment of this application; Figure 2 is a schematic flowchart of a complex image intelligent recognition method provided by another embodiment of this application; Figure 3 is a schematic structural diagram of a complex image intelligent recognition system provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] The embodiments of this application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain this application, and should not be construed as a limitation to this application.
[0014] A complex image intelligent recognition method and system according to an embodiment of this application will be described below with reference to the accompanying drawings.
[0015] Figure 1 A flowchart of a complex image intelligent recognition method provided by an embodiment of the present application is shown as Figure 1 shown. The complex image intelligent recognition method includes the following steps: Step 101: Obtain an image to be recognized, where the image to be recognized is a circuit diagram or a computer-aided design (CAD) drawing.
[0016] In an embodiment of the present application, the image to be recognized may refer to a circuit diagram or a CAD drawing including text, or may refer to a simple image including text, such as a document screenshot, a captured image including text, etc.
[0017] It should be noted that in addition to circuit diagrams and CAD drawings, the image to be recognized may also refer to an image including a large number of lines, symbols, and text.
[0018] Step 102: Using the block information indicating the block scheme as a variable, construct a block scheme scoring function based on the text density and / or background interference degree in the image blocks obtained after performing block processing on the image to be recognized according to the block information, and solve the block scheme scoring function to obtain the target block information.
[0019] In an embodiment of the present application, the block information is used to indicate the number of image blocks or the size of the image blocks. For example, the block information may include the number of block rows and the number of block columns, or may include information such as the number of blocks or the block size; the text density may be used to indicate the text distribution, and the background interference degree may be used to indicate the image complexity of the non-text area in the image block.
[0020] In an optional embodiment, a block scheme scoring function may be constructed based on the text density in the image blocks obtained after performing block processing on the image to be recognized according to the block information.
[0021] In an optional embodiment, a block scheme scoring function may be constructed based on the background interference degree in the image blocks obtained after performing block processing on the image to be recognized according to the block information.
[0022] For example, the text density mean value may be calculated based on the text density in at least one image block obtained after performing block processing on the image to be recognized according to the block information to obtain the block scheme scoring function; or, the background interference degree mean value may be calculated based on the background interference degree in at least one image block obtained after performing block processing on the image to be recognized according to the block information to obtain the block scheme scoring function.
[0023] In an optional embodiment, a block scheme scoring function may also be constructed based on the text density and background interference degree in the image blocks obtained after performing block processing on the image to be recognized according to the block information.
[0024] Among them, the scoring function of the chunking scheme is the objective function to be solved, and the chunking information has corresponding constraint conditions. When solving the scoring function of the chunking scheme, the constraint conditions can be combined to solve the target chunking information.
[0025] Step 103: According to the target chunking information, perform chunking processing on the image to be recognized to obtain a plurality of image chunks; among them, the target chunking information is used to indicate the chunking scheme.
[0026] As an example, the image to be recognized can be evenly chunked according to the target chunking information to obtain a plurality of image chunks of the same size.
[0027] To integrate the text corresponding to each image chunk and ensure the integrity of the text information, there may be an overlapping area between adjacent image chunks. The length of the overlapping area is denoted as and the width is denoted as . It should be noted that the overlap only occurs between adjacent image chunks, and each image chunk overlaps with the image chunks above, below, to the left, and to the right of it.
[0028] Step 104: For any image chunk, perform image enhancement processing on the image chunk to obtain an enhanced image chunk, and perform text recognition based on the enhanced image chunk to obtain recognized text.
[0029] Among them, circuit diagrams or CAD drawings usually contain a large number of lines, symbols, and text descriptions. Image enhancement of the image chunk can specifically optimize local details, such as improving the clarity and contrast of lines, symbols, and text, removing noise, and avoiding blurring of details caused by overall processing, thereby helping to improve the text recognition effect.
[0030] In the embodiments of the present application, a text recognition model can be used to perform text recognition on the enhanced image chunk to obtain recognized text.
[0031] Step 105: Based on the position of the recognized text in the image to be recognized, reorganize the recognized text corresponding to at least one enhanced image chunk to obtain the target text corresponding to the image to be recognized.
[0032] In the embodiments of the present application, the recognized text corresponding to at least one enhanced image chunk can be spliced, integrated, and reorganized based on the position of the recognized text in the image to be recognized to obtain the target text corresponding to the image to be recognized.
[0033] In this embodiment, an image to be recognized is obtained. The image to be recognized is a circuit diagram or a computer-aided design (CAD) drawing. Taking the block information used to indicate the block division scheme as a variable, a block division scheme scoring function is constructed based on the text density and / or background interference degree in the image blocks obtained after dividing the image to be recognized according to the block information, and the block division scheme scoring function is solved to obtain the target block information. Then, according to the target block information, the image to be recognized is divided to obtain a plurality of image blocks. The target block information is used to indicate the block division scheme. For any image block, the image block is subjected to image enhancement processing to obtain an enhanced image block, and text recognition is performed based on the enhanced image block to obtain the recognized text. Based on the position of the recognized text in the image to be recognized, the recognized texts corresponding to at least one enhanced image block are recombined to obtain the target text corresponding to the image to be recognized. By constructing and solving the block division scheme scoring function, the optimal block information is obtained, realizing the accurate block division of the circuit diagram or CAD drawing, which helps to reduce the difficulty of text recognition for complex images. Recognizing text based on the enhanced image block and then recombining to obtain the target text helps to improve the accuracy and integrity of text recognition.
[0034] This embodiment provides another intelligent recognition method for complex images. Figure 2 It is a schematic flowchart of an intelligent recognition method for complex images provided by an embodiment of the present application.
[0035] As Figure 2 shown, the intelligent recognition method for complex images may include the following steps: Step 201, obtain an image to be recognized.
[0036] Step 202, for any image block obtained after dividing the image to be recognized according to the block information, obtain the difference between the text density and the background interference degree in the image block; calculate the difference mean based on the differences corresponding to at least one image block obtained after dividing the image to be recognized according to the block information to obtain a block division scheme scoring function, and solve the block division scheme scoring function to obtain the target block information.
[0037] In the embodiment of the present application, the block information may include the number of block rows and the number of block columns .
[0038] Determining the target block information according to the complexity of the image to be recognized and the text distribution situation can ensure that the background interference is reduced as much as possible in each image block, while containing more text information.
[0039] As an example, the block division scheme scoring function may be shown as the following formula:
[0040] Where Indicates the score of the block segmentation scheme, Indicates the row and the column of the text density in the image block, Indicates the row and the column of the background interference degree in the image block, α Indicates the weight coefficient, The value range of is [1, , The value range of is [1, . It should be noted that in the formula, and are variables.
[0041] In some embodiments, the number of block rows and the number of block columns respectively have value constraint conditions.
[0042] In some embodiments, in combination with the value constraint conditions corresponding to the number of block rows and the number of block columns, the maximum value of the block segmentation scheme score function is solved to obtain the target block information, where the target block information includes the target number of block rows and the target number of block columns. Among them, the value constraint condition may refer to a value interval.
[0043] In the embodiments of the present application, solving the maximum value of the block segmentation scheme score function is to maximize the block segmentation scheme score function (that is, to maximize the text density , and minimize the background interference degree ) as the goal, and solve the target block information.
[0044] In some embodiments, it is also possible to solve the target block information with the goal that the block segmentation scheme score reaches a set score threshold.
[0045] When iteratively solving the block segmentation scheme score function, the mean value or median value of the text density can be obtained. Taking the median value of the text density as an example, if the median value of the text density is less than the first density threshold (the image block is too large or the text is too little), then increase and values; if the median value of the text density is greater than the second density threshold (the image block is too small or the text is too dense), then decrease and values until the optimal and are obtained, and the iteration is stopped. Among them, the first density threshold is less than the second density threshold.
[0046] It should be noted that the block information has an initial value, that is, and in the above formula have initial values. As an example, and Initial value of:
[0047]
[0048] Wherein, represents the initial value of, represents the initial value of, represents the length of the image to be recognized, represents the width of the image to be recognized, AvgTextWidth represents the average width of the text area, and AvgTextHeight represents the average height of the text area. is an empirical coefficient used to ensure that an image block can accommodate multiple text areas.
[0049] In some embodiments, for any image block obtained after segmenting the image to be recognized according to the segmentation information, obtain the text area in the image block, and determine the text density based on the area ratio of the text area to the image block; and / or, for any image block obtained after segmenting the image to be recognized according to the segmentation information, obtain the non-text area in the image block, and obtain the number of edge pixels obtained by edge detection in the non-text area, and determine the background interference degree based on the ratio of the number of edge pixels to the area of the non-text area.
[0050] As an example, the text density in the image block can be calculated by the following formula:
[0051] Wherein, TextArea( cell ij ) represents the area of the text area in the image block at the th row and the th column, represents the length of the image block at the th row and the th column, represents the width of the image block at the th row and the th column.
[0052] As an example, the background interference degree in the image block can be calculated by the following formula:
[0053] Wherein, EdgePixels( cell ij ) represents the number of edge pixels obtained by edge detection for the th row and the The number of edge pixels extracted by edge detection on the image blocks in the column, TextEdgePixels( cell ij ) represents the number of edge pixels extracted from the text area of the image block in the th row and the th column. In the above formula, the numerator represents the number of edge pixels obtained by edge detection in the non-text area of the image block in the th row and the th column, and the denominator represents the area of the non-text area of the image block in the th row and the th column.
[0054] Step 203: According to the target block information, perform block processing on the image to be recognized to obtain multiple image blocks.
[0055] Among them, after performing block processing on the image to be recognized according to the target block information, there is an overlapping area between adjacent image blocks.
[0056] Assume that the target block information obtained by solving the block scheme scoring function includes the number of rows and the number of columns , then the size of the image block considering the overlapping area is:
[0057]
[0058] Step 204: For any image block, perform image enhancement processing on the image block to obtain an enhanced image block, and perform text recognition based on the enhanced image block to obtain recognized text.
[0059] In some embodiments, the image block is filtered by a low-pass filter to obtain a blurred image block; based on the image block and the blurred image block, a detail image block is obtained; the detail image block is enhanced to obtain an enhanced detail image block; based on the image block and the enhanced detail image block, an enhanced image block is obtained.
[0060] Among them, the specific implementation method of performing image enhancement processing on the image block is: (1) Apply a low-pass filter to the image block to generate a blurred image block as the unsharp masking reference; specifically, , where represents the original image block, represents the low-pass filter, and represents the blurred image block.
[0061] (2) Subtract the blurred image block from the original image block to obtain a detail image block (high-frequency image block), specifically, , where, Represents a detailed image patch.
[0062] (3) Multiply the detailed image patch by an enhancement factor to adjust the intensity of enhancement. Specifically, , where represents the enhancement factor, is the enhanced detailed image patch.
[0063] (4) Add the enhanced detailed image patch to the original image patch to obtain the final enhanced image patch. Specifically, , where represents the final enhanced image patch, which retains both the low-frequency components of the original image patch and enhances the high-frequency components (edges and details).
[0064] In some embodiments, based on a text detection model, text detection is performed on the enhanced image patch to obtain a text region; based on a text recognition model, text recognition is performed on the text region to obtain recognized text.
[0065] Among them, the text region may refer to the region where the text is located in the enhanced image patch; detecting the text region in the enhanced image patch first and then performing text recognition on the text region can accurately locate the text position, reduce background interference, and thus improve the reliability, accuracy, and efficiency of text recognition.
[0066] Before performing text detection on the enhanced image patch based on a text detection model, feature extraction can be first performed on the enhanced image patch based on a convolutional neural network, and then the text detection model can be used to perform text detection based on the extracted features.
[0067] Among them, the formula for feature extraction by the convolutional neural network is T = f(W w *t + b), where t is the input enhanced image patch, W w is the convolutional kernel, b is the bias term, and T is the output feature map.
[0068] In some embodiments, the text detection model is a Feature Pyramid Networks (FPN) model based on convolution, and / or the text recognition model is an Attention-based model.
[0069] Among them, FPN captures text regions of different sizes by constructing feature maps of different scales, and prediction layers are set on the feature maps of each scale to output the detection results of text regions. Among them, the prediction layer may include a convolutional layer and a classification layer (or regression layer) for generating the coordinates of the text region detection box.
[0070] Among them, the core of the Attention-based model is the calculation formula of the Attention mechanism. The Attention mechanism determines the weight of each feature vector by calculating the similarity (such as cosine similarity) between the current state of the decoder and the encoder feature vectors. Then, based on these weights, the feature vectors are weighted and summed to obtain the context vector at the current moment of the decoder. The decoder generates the prediction of the next character according to the context vector and the current state.
[0071] It should be noted that during the training process of the text detection model and the text recognition model, the recognition results of the model can be evaluated through evaluation metrics such as accuracy and recall rate: Accuracy = (TP + TN) / (TP + FP + FN + TN) Recall rate = TP / (TP + FN) True Positives (TP): The number of positive samples correctly predicted by the model; True Negatives (TN): The number of negative samples correctly predicted by the model; False Positives (FP): The number of positive samples wrongly predicted by the model; False Negatives (FN): The number of negative samples wrongly predicted by the model.
[0072] Step 205: Based on the position of the recognized text in the image to be recognized, reorganize the recognized text corresponding to at least one enhanced image block to obtain the target text corresponding to the image to be recognized.
[0073] In some embodiments, for any recognized text, obtain the first position of the recognized text in the corresponding enhanced image block, and obtain the second position of the enhanced image block in the image to be recognized; correct the second position according to the size of the overlapping area to obtain the corrected second position; based on the first position and the corrected second position, reorganize each recognized text to obtain the target text.
[0074] Among them, reorganizing the recognized text can achieve the integration of the recognized text, so as to obtain a complete and smooth target text.
[0075] Among them, define the relative position of the image block in the th row and the th column, that is, the second position is ( , ), where i represents the row number (1 ≤ i ≤ ), j represents the column number (1 ≤ j≤ ), the second position can be corrected according to the size of the overlapping area through the following formula:
[0076] where it is assumed here that the overlapping areas are evenly distributed between adjacent cells, and only the overlaps with the left and right cells are considered; represents the absolute starting abscissa after correction; is used to handle the boundary conditions on the left and right to ensure that the first and last columns do not exceed the boundaries due to overlap.
[0077]
[0078] where it is assumed here that the overlapping areas are evenly distributed between adjacent cells, and only the overlaps with the upper and lower cells are considered; represents the absolute starting ordinate after correction; is used to handle the boundary conditions on the upper and lower to ensure that the first and last rows do not exceed the boundaries due to overlap.
[0079] where the ( i row, j column image block corresponds to ( ) which is the corrected second position.
[0080] It should be noted that after obtaining the target text, the target text can be stored. Specifically, each recognized text can be associated and stored with the position of the image block it is in (the second position and / or the corrected second position). For example, each recognized text can be associated and stored with the row number, column number, and the corrected second position of the corresponding image block; or the recognized text, the first position of the recognized text in the corresponding image block, and the position of the image block where the recognized text is located (the second position and / or the corrected second position) can be associated and stored.
[0081] Among them, associating and storing the recognized text with the corresponding position information helps to locate the actual position of the recognized text in the image to be recognized.
[0082] In some embodiments, the method further includes: using a large language model to perform verification and adjustment processing on the integrated text to obtain the processed target text; where the verification and adjustment processing includes at least one of grammar verification processing, error character / error word replacement processing, and text adjustment processing in combination with the context semantic information of the integrated text.
[0083] Among them, using a large language model to perform grammar checking on the target text can correct possible grammar errors; matching the identified target text with a predefined dictionary can replace misidentified characters or words; further verifying and adjusting the target text according to the context semantic information can reduce errors and omissions, ensuring the accuracy and reliability of the recognition result.
[0084] It should be noted that after obtaining the target text, it can be manually proofread, and the identified information can be deeply analyzed to ensure the accuracy of the text.
[0085] In this embodiment, an image to be recognized is obtained; for any image block obtained after dividing the image to be recognized according to the block information, the difference between the text density and the background interference degree in the image block is obtained; based on the differences corresponding to at least one image block obtained after dividing the image to be recognized according to the block information, the difference mean value is calculated to obtain a block division scheme scoring function; the block division scheme scoring function is solved to obtain target block information; according to the target block information, the image to be recognized is divided to obtain a plurality of image blocks; for any image block, the image block is subjected to image enhancement processing to obtain an enhanced image block, and text recognition is performed based on the enhanced image block to obtain recognized text; based on the position of the recognized text in the image to be recognized, the recognized text corresponding to at least one enhanced image block is recombined to obtain the target text corresponding to the image to be recognized. Constructing a block division scheme scoring function based on the difference between the text density and the background interference degree in the image block, and then solving the target block information, can ensure that the background interference is reduced as much as possible in each image block, while containing more text information, so as to achieve fast and accurate text recognition for images to be recognized with multiple scales, complex backgrounds, and difficult text extraction.
[0086] To implement the above embodiment, an embodiment of the present application also proposes a complex image intelligent recognition system.
[0087] Figure 3 It is a schematic structural diagram of a complex image intelligent recognition system provided by an embodiment of the present application.
[0088] As Figure 3 shown, the complex image intelligent recognition system 300 includes: A first acquisition module 310, configured to acquire an image to be recognized, and the image to be recognized is a circuit diagram or a computer-aided design (CAD) drawing; A second acquisition module 320, configured to use the block information indicating the block division scheme as a variable, construct a block division scheme scoring function based on the text density and / or background interference degree in the image block obtained after dividing the image to be recognized according to the block information, and solve the block division scheme scoring function to obtain target block information.
[0089] The block processing module 330 is configured to perform block processing on the image to be recognized according to the target block information, so as to obtain a plurality of image blocks; wherein, the target block information is used to indicate the block scheme. The image processing module 340 is configured to perform image enhancement processing on any image block to obtain an enhanced image block, and perform text recognition based on the enhanced image block to obtain recognized text. The recombination module 350 is configured to recombine the recognized text corresponding to at least one enhanced image block based on the position of the recognized text in the image to be recognized, so as to obtain the target text corresponding to the image to be recognized.
[0090] Optionally, the second obtaining module 320 is specifically configured to: for any image block obtained by performing block processing on the image to be recognized according to the block information, obtain the difference between the text density and the background interference degree in the image block; based on the differences corresponding to at least one image block obtained by performing block processing on the image to be recognized according to the block information, calculate the difference mean value to obtain the block scheme scoring function.
[0091] Optionally, the block information includes the number of block rows and the number of block columns, and the number of block rows and the number of block columns respectively have value constraint conditions. The second obtaining module 320 is specifically configured to: in combination with the value constraint conditions, solve the maximum value of the block scheme scoring function to obtain the target block information, where the target block information includes the target number of block rows and the target number of block columns.
[0092] Optionally, the system further includes a third obtaining module, configured to: For any image block obtained by performing block processing on the image to be recognized according to the block information, obtain the text region in the image block, and determine the text density based on the area ratio of the text region to the image block; and / or, for any image block obtained by performing block processing on the image to be recognized according to the block information, obtain the non-text region in the image block, and obtain the number of edge pixels obtained by edge detection in the non-text region, and determine the background interference degree based on the ratio of the number of edge pixels to the area of the non-text region.
[0093] Optionally, the image processing module 340 is specifically configured to: filter the image block through a low-pass filter to obtain a blurred image block; obtain a detail image block based on the image block and the blurred image block; perform image enhancement on the detail image block to obtain an enhanced detail image block; obtain an enhanced image block based on the image block and the enhanced detail image block.
[0094] Optionally, the image processing module 340 is specifically configured to: perform text detection on the enhanced image block based on a text detection model to obtain a text region; perform text recognition on the text region based on a text recognition model to obtain recognized text.
[0095] Optionally, the text detection model is a convolutional-based feature pyramid network model, and / or the text recognition model is an attention mechanism-based model.
[0096] Optionally, there is an overlapping area between adjacent image patches. The recombination module 350 is specifically configured to: obtain the first position of the recognized text in the corresponding enhanced image patch, and obtain the second position of the enhanced image patch in the image to be recognized; correct the second position according to the size of the overlapping area to obtain the corrected second position; and recombine the recognized texts based on the first position and the corrected second position to obtain the target text.
[0097] Optionally, the system further includes a verification and adjustment processing module for: using a large language model to perform verification and adjustment processing on the target text to obtain the processed target text; where the verification and adjustment processing includes at least one of grammar verification processing, error character / error word replacement processing, and text adjustment processing by combining and integrating the context semantic information of the text.
[0098] It should be noted that the foregoing explanation of the embodiments of a complex image intelligent recognition method also applies to a complex image intelligent recognition system of this embodiment, and will not be elaborated here.
[0099] In the embodiments of the present application, an image to be recognized is obtained, and the image to be recognized is a circuit diagram or a computer-aided design (CAD) diagram; target block information is obtained, and according to the target block information, the image to be recognized is block-processed to obtain a plurality of image patches; where the target block information is used to indicate the block scheme; for any image patch, the image patch is subjected to image enhancement processing to obtain an enhanced image patch, and text recognition is performed based on the enhanced image patch to obtain recognized text; based on the position of the recognized text in the image to be recognized, the recognized texts corresponding to at least one enhanced image patch are recombined to obtain the target text corresponding to the image to be recognized. By constructing and solving a block scheme scoring function to obtain the optimal block information, accurate block processing of the circuit diagram or CAD diagram is realized, which helps to reduce the difficulty of text recognition of complex images; recognizing text based on the enhanced image patch and then recombining to obtain the target text helps to improve the accuracy and integrity of text recognition.
Claims
1. A complex image intelligent recognition method, characterized in that: The following steps are involved: Acquire an image to be identified, wherein the image to be identified is a circuit diagram or a computer-aided design (CAD) diagram; Taking the block information indicating the block scheme as a variable, based on the text density and / or background interference degree in the image block obtained after the image to be identified is block processed according to the block information, a block scheme scoring function is constructed, and the block scheme scoring function is solved to obtain the target block information; According to the target block information, the image to be identified is processed into blocks to obtain a plurality of image blocks; wherein the target block information is used to indicate a block scheme; For any of the image blocks, performing image enhancement processing on the image block to obtain an enhanced image block, and performing text recognition based on the enhanced image block to obtain recognized text; Based on the position of the recognition text in the image to be recognized, the recognition text corresponding to at least one enhanced image block is reorganized to obtain the target text corresponding to the image to be recognized.
2. The method according to claim 1, characterized in that The step of constructing a block scheme score function based on the text density and / or background interference degree in the image block obtained after the image to be identified is block processed according to the block information includes: For any image block obtained after the image to be identified is segmented according to the segmentation information, obtaining the difference between the text density and the background interference degree in the image block; Based on the difference corresponding to at least one image block obtained after the image to be identified is segmented according to the segmentation information, a difference mean is calculated to obtain the segmentation scheme score function.
3. The method according to claim 2, characterized in that The block information includes the number of block rows and the number of block columns, and the number of block rows and the number of block columns respectively have value constraints. The solving the block scheme score function to obtain the target block information includes: In combination with the value constraint condition, the maximum value of the score function of the partition scheme is solved to obtain the target partition information, wherein the target partition information includes the number of target partition rows and the number of target partition columns.
4. The method according to claim 1, characterized in that The method further comprises: For any image block obtained after the image to be identified is divided into blocks according to the block information, a text area in the image block is obtained, and the text density is determined based on an area ratio between the text area and the image block; and / or, For any image block obtained after the image to be identified is segmented according to the segmentation information, a non-text area in the image block is obtained, and the number of edge pixels in the non-text area obtained by edge detection is obtained, and the background interference degree is determined based on the ratio of the number of edge pixels to the area of the non-text area.
5. The method according to claim 1, characterized in that The performing image enhancement processing on the image block to obtain an enhanced image block comprises: Performing filtering processing on the image block by using a low-pass filter to obtain a blurred image block; Based on the image block and the blurred image block, obtaining a detail image block; Performing image enhancement on the detail image block to obtain an enhanced detail image block; Based on the image block and the enhanced detail image block, the enhanced image block is acquired.
6. The method according to claim 1, characterized in that The performing text recognition based on the enhanced image block to obtain recognized text includes: Based on a text detection model, performing text detection on the enhanced image block to obtain a text area; Based on the text recognition model, text recognition is performed on the text area to obtain the recognized text.
7. The method according to claim 6, characterized in that The text detection model is a convolution-based feature pyramid network model, and / or the text recognition model is a model based on an attention mechanism.
8. The method according to claim 1, characterized in that There is an overlapping area between adjacent image blocks, and the recognition text corresponding to at least one enhanced image block is reorganized based on the position of the recognition text in the image to be recognized to obtain the target text corresponding to the image to be recognized, including: For any of the recognized texts, obtaining a first position of the recognized text in a corresponding enhanced image block, and obtaining a second position of the enhanced image block in the image to be recognized; Correcting the second position according to the size of the overlapping area to obtain a corrected second position; Based on the first position and the corrected second position, each of the recognized texts is spliced and reorganized to obtain the target text.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: Using a large language model, the target text is verified and adjusted to obtain a processed target text; The verification and adjustment processing includes at least one of grammar verification processing, incorrect character / incorrect word replacement processing, and text adjustment processing combined with contextual semantic information of the integrated text.
10. A complex image intelligent recognition system, characterized in that: include: A first acquisition module is used to acquire an image to be recognized, wherein the image to be recognized is a circuit diagram or a computer-aided design (CAD) diagram; A second acquisition module is used to construct a block scheme scoring function based on the text density and / or background interference degree in the image block obtained after the image to be identified is block processed according to the block information, using the block information indicating the block scheme as a variable, and solve the block scheme scoring function to obtain the target block information; A block processing module, used for performing block processing on the image to be identified according to the target block information to obtain a plurality of image blocks; wherein the target block information is used to indicate a block scheme; An image processing module, configured to perform image enhancement processing on any of the image blocks to obtain an enhanced image block, and perform text recognition based on the enhanced image block to obtain recognized text; The reorganization module is used to reorganize the recognized text corresponding to at least one enhanced image block based on the position of the recognized text in the image to be recognized, so as to obtain the target text corresponding to the image to be recognized.
Citation Information
Patent Citations
Tibetan language recognition method and device and electronic device
CN110032938A
Image enhancement method, text detection model training method and equipment
CN118657686A
Super-large format document intelligent identification method and system based on block parallelism
CN119625766A
Method for recognizing text, device, and storage medium
US20230206667A1
Multi-language text recognition method and apparatus, computer device, and storage medium
WO2021017260A1