A method and device for mind map image recognition, analysis and reconstruction

By adaptively cutting and scaling the mind map images, combining text and line segment detection, a binary mask diagram is generated and the mind map is reconstructed, the problem of low accuracy of mind map detection is solved and the conversion efficiency is improved.

CN113449734BActive Publication Date: 2025-07-29SUZHOU ZHIXI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110757918.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-05
Publication Date
2025-07-29
Estimated Expiration
2041-07-05

AI Technical Summary

Technical Problem

The existing general text detection and recognition models have low accuracy in detecting and identifying mind map images and are large in calculations, so they cannot efficiently convert mind map styles generated by different software.

Method used

The mind map image is adaptively cut and scaled, and text and line segment detection is performed separately, a binary mask diagram is generated, and combined into the original image coordinate system, identify the text block node relationship and reconstruct the mind map.

Benefits of technology

It improves the accuracy of mind map detection and recognition, reduces the computing memory requirement, and improves the efficiency of generating or converting mind map styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113449734B_ABST
    Figure CN113449734B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for mind map image recognition and parsing and reconstruction. The method includes: adjusting the size of the mind map image and then cutting it to obtain image blocks; generating a binary mask map of the text region, and reassembling the text region mask maps generated by block division into a complete binary mask map of the text region; extracting the text image blocks corresponding to the regions of the adjusted mind map, and identifying the text information and the corresponding positional relationship in the extracted text region image blocks; recutting the image, generating a complete binary mask map of the line segments according to the detected line segment information, and finally correcting the line segment information according to the binary contour of the line segment mask map; determining the matching relationship between the text region nodes and the line segments and the connection relationship between different text region nodes according to the intersection situation of the corrected line segments and the positional and distance relationships between the text regions and the line segments, and reconstructing the mind map. The present application improves the efficiency of mind map material generation and style conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technology of text detection and recognition in images, and in particular, to a method and device for recognizing, parsing, and reconstructing mind map images. Background Art

[0002] A mind map is an effective graphic thinking tool for expressing divergent thinking. It uses the technique of combining pictures and texts to show the relationships between themes at all levels with a hierarchical diagram of mutual subordination and correlation. Mind maps are widely used because of their vividness, simplicity, and effectiveness. In recent years, a large number of mind map drawing software tools have been launched on the market, and a huge amount of mind map graphic data has been produced. The mind maps generated by different software tools have different styles, and due to the lack of representation of the original structured data relationships, the generated mind maps cannot be efficiently converted between these software tools and can only be converted in a completely manual way. This situation leads to people often generating mind maps with similar content in a very inefficient manner, reducing the usage efficiency of mind maps.

[0003] In recent years, artificial intelligence technologies such as deep learning have developed rapidly. In particular, technologies such as text detection and recognition, and line detection have made remarkable progress. These technological breakthroughs provide a basis for parsing mind maps. For ordinary documents, existing general optical character recognition (OCR) tools can already detect and recognize Chinese and English characters in graphics and texts very efficiently, such as the text detection model algorithm of the Connectionist Text Proposal Network (CPTN) and the text recognition model algorithm of the Convolutional Recurrent Neural Network (CRNN) for end-to-end variable-length text recognition. Different from ordinary documents, mind maps usually have large image sizes, inconsistent page specifications and text formats, and have the typical characteristics of large images with small characters. General text detection and recognition deep neural network models will normalize the image, and generally have certain requirements for the size and font size of the input text image. Deviating from these standards will lead to problems such as reduced detection and recognition accuracy and increased computational complexity. If the mind map image is directly scaled to a unified specification and then input into the deep neural network according to the general text detection and recognition method, the detection and recognition accuracy will be very low. This phenomenon is mainly because large-scale images will not only increase the memory and video memory requirements for model calculation, but also the texture information of the text will be severely damaged and lost after the scaling operation. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a method and device for recognizing, parsing, and reconstructing mind map images.

[0005] According to the first aspect of the present application, a method for mind map image recognition, parsing and reconstruction is provided, including:

[0006] Adjust the size of the mind map image, and perform adaptive cutting and partitioning on the adjusted image to obtain image blocks;

[0007] Detect the text in the image blocks and generate corresponding binary mask maps of the text regions. Based on the size of the adjusted image, splice the binary mask maps of the partitioned text regions into a complete binary mask map of the text;

[0008] Extract the text region image blocks at the corresponding positions in the adjusted mind map according to the text region coordinates in the binary mask map of the text, recognize the text information in the text region image blocks and extract the corresponding positional relationships;

[0009] Cut the image, detect the line segments in the image blocks, generate a complete binary mask map of the line segments according to the detected line segment information in the image, and correct the line segments again based on the binary contours in the binary mask map of the line segments;

[0010] Determine the intersection situation between the finally corrected line segments. According to the intersection situation and the positional relationship, determine the matching relationship between the finally determined line segments and the text block nodes and the parent-child connection relationship between different text block nodes. Reconstruct the mind map according to the determined parent-child relationship and position information of the text block nodes, and arrange and edit it on the editing drawing board.

[0011] As an implementation, the reconstruction of the mind map includes:

[0012] In response to an edit request for the structured text content after arrangement, edit the structured text content to obtain the initial data of the new mind map, and generate a new mind map in response to a modification request for the initial data.

[0013] As an implementation, the cutting and partitioning of the adjusted image to obtain image blocks includes:

[0014] According to the preset pixel quantity value for text detection and cutting and the overlapping pixel quantity value for text detection and cutting, use a cutting method with partial overlapping regions for the image blocks to cut the adjusted image into image blocks with consistent sizes.

[0015] As an implementation, the extraction of the text image blocks in the corresponding regions of the adjusted mind map includes:

[0016] Detect the text in the image block and generate a corresponding binary mask map of the text region. Based on the size of the adjusted image, splice the binary mask maps of the segmented text regions back into a complete binary mask map of the text.

[0017] Extract the text region image block at the corresponding position in the adjusted mind map according to the text region coordinates in the binary mask map of the text, recognize the text information in the text region image block, and extract the corresponding positional relationship.

[0018] As an implementation, the line segment includes at least one of the following:

[0019] Horizontal line segment, vertical line segment, oblique line segment;

[0020] Correspondingly, determining the relationship between the arranged text block nodes and the final line segment according to the position and distance rules includes:

[0021] Determine the spatial position distribution characteristics and statistical characteristics of the horizontal and vertical line segments, confirm the categories of the horizontal and vertical line segments and the line segment intersection situations between them; according to the intersection situations, identify the relationships between the line segments, and match the text block nodes with the corresponding line segments according to the distance rules to determine the matching relationship between the horizontal line segment and the text block nodes; among them, when determining that one vertical line segment intersects multiple horizontal line segments, the corresponding text block nodes have a matching relationship with the horizontal line segments.

[0022] According to the relationships between the determined line segments and between the text block nodes and the line segments, determine the parent-child node relationships between the text block nodes, and delete the duplicate and non-compliant line connection relationships as the relationships between the text block nodes and the final line segment.

[0023] According to the second aspect of the present application, there is provided a mind map image recognition and parsing and reconstruction device, including:

[0024] A cutting unit for adjusting the size of the mind map image and adaptively cutting and segmenting the adjusted image to obtain image blocks;

[0025] A detection unit for detecting the text or line segments in the image block and generating corresponding binary mask image blocks of the text or line segments;

[0026] A merging unit for splicing the binary mask image blocks of the text or line segments back into a complete binary mask map of the text or a binary mask map of the line segments based on the size of the adjusted image.

[0027] An identification unit, configured to extract text region image blocks at corresponding positions in the adjusted mind map according to the text region coordinates in the text binary mask map, and identify the text information in the text region image blocks; correct the line segment information according to the line segment binary mask map, and identify the matching relationship between the line segment and the text block node and the parent-child node relationship between different text block nodes according to the corrected line segment information;

[0028] An editing unit, configured to arrange and edit the identified text information on an editing drawing board according to the connection relationship and relative position information of its parent-child nodes.

[0029] As an implementation manner, the editing unit is further configured to:

[0030] In response to an editing request for the arranged structured text content, edit the structured text content to obtain the initial data of a new mind map, and generate a new mind map in response to a modification request for the initial data.

[0031] As an implementation manner, the cutting unit is further configured to:

[0032] Re-obtain the coordinate position information of the text region in the adjusted image by finding the closed contour of the spliced text binary mask map, and extract the region corresponding to the coordinate position in the adjusted image as the text region image block.

[0033] As an implementation manner, the identification unit is further configured to:

[0034] Obtain the coordinate position information of the text region in the adjusted image by finding the closed contour of the binary mask map, and extract the region corresponding to the coordinate position in the adjusted image as the text region image block.

[0035] As an implementation manner, the line segment includes at least one of the following:

[0036] Horizontal line segment, vertical line segment, oblique line segment;

[0037] Correspondingly, the editing unit is further configured to:

[0038] Determine the spatial position distribution characteristics and statistical characteristics of horizontal and vertical line segments, confirm the categories of horizontal and vertical line segments and the line segment intersection situation between them; according to the intersection situation, identify the relationship between line segments, and match the relationship between the text block node and the corresponding line segment according to the position and distance rules, and determine the matching relationship between the horizontal line segment and the text block node; wherein, when it is determined that a vertical line segment intersects with multiple horizontal line segments, the corresponding text block node has a matching relationship with the horizontal line segment;

[0039] According to the determined relationships between line segments, and between text block nodes and line segments, determine the parent-child relationships between different text block nodes, and delete duplicate and non-compliant node connection relationships as the relationships between text block nodes and the final line segments.

[0040] According to a third aspect of the present application, there is provided a storage medium storing an executable program, and when the executable program is executed by a processor, the steps of the mind map recognition and reconstruction method are implemented.

[0041] In the mind map image recognition, parsing and reconstruction method and device provided by the embodiments of the present application, before performing text detection on the mind map image, the image is first adaptively segmented, cut and scaled, and then the text detection and line segment detection are respectively performed on the segmented image blocks. Then, the detection results of the image blocks are finally merged into the coordinate system of the original image through operations such as image morphology. Finally, text recognition and the recognition of the parent-child connection relationship of text block nodes are performed, and the recognition results are displayed on the editing drawing board interface according to the relative position relationship of the original image. The user can modify and generate a new mind map with a corresponding style according to their own needs. The embodiments of the present application effectively solve the technical problem of low accuracy in text detection and recognition in common mind maps, and can not only meet the detection and recognition requirements of mind maps with different scale layout specifications, but also reduce the demand for computing video memory and improve the efficiency of users in producing mind maps or converting the style of mind maps. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic flowchart of the mind map image recognition, parsing and reconstruction method provided by the embodiments of the present application;

[0043] Figure 2 It is a schematic flowchart for parsing a straight-line connection type mind map according to the embodiments of the present application;

[0044] Figure 3 It is a schematic flowchart for text detection and recognition according to the embodiments of the present application;

[0045] Figure 4 It is a schematic flowchart for line segment detection and recognition according to the embodiments of the present application;

[0046] Figure 5 It is a schematic flowchart for matching the connection relationship of text blocks according to the embodiments of the present application;

[0047] Figure 6 It is a schematic structural diagram of the composition of the mind map image recognition, parsing and reconstruction device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The essence of the technical solution of the embodiments of the present application is elaborated in detail below with reference to examples.

[0049] Figure 1 This is a schematic flowchart of the mind map image recognition and parsing and reconstruction method provided by the embodiments of the present application. As Figure 1 shown, the mind map image recognition and parsing and reconstruction method of the embodiments of the present application includes the following processing steps:

[0050] Step 101: Adjust the size of the mind map image, and perform adaptive cutting and partitioning on the adjusted image to obtain image blocks.

[0051] In the embodiments of the present application, the mind map can be any graph generated based on relevant application software, not limited to a specific format, and the technical solutions of the embodiments of the present application can accurately recognize it.

[0052] Here, adjusting the size of the mind map image is mainly to scale the size of the mind map image according to the set scaling parameters.

[0053] Perform cutting processing on the scaled mind map image. Specifically, according to the preset text detection cutting pixel quantity value and the text detection cutting overlapping pixel quantity value, adopt a cutting method with partial overlapping regions for the image blocks, and cut the adjusted image into image blocks with consistent sizes.

[0054] Step 102: Detect the text in the image blocks and generate a corresponding binary mask map of the text region. Based on the size of the adjusted image, re - merge and splice the binary mask maps of the partitioned text regions into a complete binary mask map of the text.

[0055] In the embodiments of the present application, after the mind map image is cut to form image blocks, it is also necessary to detect the text in the image blocks, extract the text region, and merge it according to the positional relationship of the large - scale image before cutting, etc., to form a binary mask map of the text region.

[0056] Step 103: Extract the text region image blocks at the corresponding positions of the adjusted mind map according to the text region coordinates in the binary mask map of the text region, and recognize the text information in the text region image blocks and extract the corresponding positional relationships.

[0057] In the embodiments of the present application, the position coordinate information of the text region in the adjusted image is re - obtained by finding the closed contour of the merged binary mask map of the text region, and the region corresponding to the position coordinates in the adjusted image is extracted as the text region image block.

[0058] Step 104: Re - cut the image, detect the line segments in the image blocks, generate a complete binary mask map of the line segments according to the detected line segment information in the image, and correct the line segments based on the binary contours in the binary mask map of the line segments.

[0059] In the embodiments of the present application, the line segment includes at least one of the following: a horizontal line segment, a vertical line segment, and an oblique line segment.

[0060] The method for generating the line segment mask binary image is similar to the method for generating the aforementioned text mask binary image. First, line detection is performed on the image block, and then the corresponding line segment mask binary image is generated according to the detected line segment endpoints; then, the effective image mask blocks are taken and merged, and similar image morphological operations as those during detection are used to recalculate the final line segment according to the line segment binary contour.

[0061] Step 105: Determine the intersection situation between the finally corrected line segments. According to the intersection situation and the positional relationship, determine the matching relationship between the finally obtained line segments and the text block nodes and the parent-child connection relationship between different text block nodes. Reconstruct the mind map according to the determined text block node parent-child relationship and positional information, and arrange and edit it on the editing drawing board.

[0062] In the embodiments of the present application, determining the relationship between the text content after arrangement and the finally obtained line segments according to the position and distance rules includes: determining the spatial position distribution characteristics and statistical characteristics of the horizontal and vertical line segments, confirming the categories of the horizontal and vertical line segments and the intersection situation between them; according to the intersection situation, identifying the relationship between line segments, and matching the relationship between the text block nodes and the corresponding line segments according to the distance rules to determine the matching relationship between the horizontal line segments and the text block nodes; wherein, when it is determined that one vertical line segment intersects multiple horizontal line segments, the corresponding text block nodes have a matching relationship with the horizontal line segments; according to the determined relationships between line segments and between text block nodes and line segments, determine the parent-child node relationship between text block nodes, and delete the duplicate and non-compliant line segment connection relationships as the relationship between the text block nodes and the finally obtained line segments.

[0063] Among them, reconstructing the mind map includes: in response to an editing request for the structured text content after arrangement, editing the structured text content to obtain the original data of the new mind map, and in response to a modification request for the original data, generating a new mind map.

[0064] The following further elaborates on the embodiments of the present application with specific examples.

[0065] Figure 2 This is the flow chart for parsing the straight-line connection type mind map according to the embodiments of the present application. As Figure 2 shown, the parsing of the mind map according to the embodiments of the present application includes the following processing steps:

[0066] 1. Input the mind map image to be recognized and reconstructed and the preset parameters, and perform text detection and recognition

[0067] For the characteristics of large pictures and small characters in mind maps, first, an adaptive overlapping cutting method is adopted to cut the entire mind map image into M*N small image blocks of the same size, where M is the number of single-line cuts and N is the number of single-column cuts. Then, the cut image blocks are respectively input into the CPTN scene text detection deep neural network model, and the CPTN model outputs the coordinate information of single-line text blocks in the image blocks. Next, a binary mask of the text area is generated using the coordinate information of the text area box output by CPTN. The M*N binary mask image blocks are recombined into a complete binary mask image according to the size and position of the mind map image, and morphological operations such as dilation and erosion are used to correct and fine-tune the binary mask image to avoid the problem of partial area splitting caused by cutting. The coordinate position information of the text block area in the large-scale image is re-obtained by finding the closed contour of the binary mask image. Finally, after obtaining the coordinate information of the text area, the text line image blocks at the corresponding positions in the mind map image are extracted, and these text line image blocks are respectively input into the CRNN text recognition model to recognize the Chinese and English text contents of the corresponding text lines.

[0068] In the embodiment of the present application, as Figure 3 shown, the text detection and recognition in the embodiment of the present application includes the following processing:

[0069] Input the mind map image Img, and preset the image scaling factor S, the text detection cutting pixel number value C, and the text detection cutting overlapping pixel number value O, etc.

[0070] When S is less than the set threshold, if the set threshold is 0 here, the image size is fine-tuned. Specifically, according to the size of the input image and the parameter C, the length and width C_h, C_w of the cut image block are calculated, and the image size is adjusted so that the length and width of the adjusted image Img_1 can be exactly divisible by C_h and C_w respectively. When it is greater than the set threshold, the mind map image is scaled according to the specified scaling factor S.

[0071] Image adaptive cutting: Mirror padding is performed on the edges of the Img_1 image according to the parameter O to obtain the padded image Img_2. Then, the values with lengths and widths of C_h+2*O and C_w+2*O and moving length and width step sizes of C_h and C_w respectively are used to cut on the image Img_2. Finally, M*N image blocks of the same size are obtained.

[0072] Text detection: The CPTN scene text detection model is used to perform text detection on the M*N image blocks respectively, and the coordinate information of the text area is output through CPTN. A binary mask image of the corresponding text area is generated according to the coordinate information of the text detection.

[0073] Merge binary mask images: Take the valid cutting regions of each binary mask image block respectively, and merge and splice them into a complete image according to the size and position of image Img_1, and use morphological operations such as dilation and erosion to eliminate the split regions. The valid cutting region is the central region in the mask image block with a length and width of C_h and C_w respectively.

[0074] Obtain the coordinate information of the text region: By finding the closed contour of the merged binary mask image, recalculate the position coordinate information of the text block in image Img_1.

[0075] For scenarios where S is less than the set threshold, sample and statistically estimate the average font size: Sample and take the coordinate information of some detected text blocks, and statistically calculate the average value R of the text blocks; if Min_r < R < Max_r, then perform text recognition: Use the obtained text coordinate information to extract the text line image of the corresponding position region in image Img_1, and then use the RCNN text recognition model to recognize the text in the text line respectively. If R is not between Min_r and Max_r, take the empirical value to scale the image Img and repeat all the previous operations once until the new text region coordinate position information is obtained. Here, Min_r is the minimum line height of the text block, and Max_r is the maximum line height of the text block.

[0076] Output the text recognition content and related position coordinate information.

[0077] Mind maps not only have very inconsistent image sizes, but also have large variations in text font sizes and often have dense small texts. Therefore, after merging the binary mask images and re-obtaining the text region coordinate information, partial sampling will be performed on the text line coordinate data, and the line heights of the sampled text lines will be statistically calculated to calculate the average value to estimate the text size in the figure. If the sampled average line height meets the requirements, text recognition will be directly performed; otherwise, in the case of no manual interaction to specify the image scaling factor, this method will take some empirical values to scale the image, and then re-perform image adaptive cutting and text detection.

[0078] The image adaptive cutting method of the embodiment of the present application adopts a cutting method with overlapping regions. For the direct sampling non-overlapping cutting method, the edge regions of the image blocks often have low accuracy, and there will be problems such as incomplete detection and fragmentation of text regions or line segments after merging the entire image. To address this issue, the cutting method with overlapping regions in the embodiment of the present application first fills the image edges according to the pixel value size of the overlapping regions, and then cuts according to the step size of the effective image blocks. The size of the cut image block is the size after filling the pixel values of the overlapping regions around the effective image block. Therefore, the effective image block is in the central region of the cut image block and is smaller than the cut image block. After inputting the cut image block into the detection model to generate the corresponding mask image block, only the effective image block region is taken for merging. The cutting method of the embodiment of the present application effectively avoids the problem of inaccurate edge detection of image blocks and improves the detection and recognition accuracy.

[0079] During text detection, when the size of the input image after scaling is greater than the cutting set value for the effective image block cutting size of the overlapping adaptive cutting method, the size of the effective cutting side length of the image is the set value. If the image is smaller than the set cutting value, the size of the effective cutting side length of the image is the side length of the image. The actual cutting side length value is the sum of the effective side length cutting value and twice the single-sided side length value of the overlapping region. For example, assume the cutting set value is 1536, and the effective cutting side length value is equal to the cutting set value, and the single-sided side length value of the overlapping region is 32, then the size of the actual cutting side length is 1600.

[0080] 2. Line segment detection.

[0081] For line segment detection, the embodiment of the present application uses a pre-trained HAWP line segment detection model for line segment detection. This model represents a line segment using the two endpoints of the line segment, that is, the output is the two endpoints of the line segment. In view of the characteristics of large images with thin lines in mind maps, the embodiment of the present application adopts an adaptive cutting and merging method with overlapping regions similar to that used in text detection. The difference is that, to adapt to the pre-trained line segment detection model, its cutting set value is 512, and the single-sided side length value of the overlapping region is 64. Similar to text detection, first use the adaptive cutting method to cut the image into M*N image blocks with the same size; then, send the image blocks into the HAWP deep neural network model for straight line detection, and then generate the corresponding line segment mask binary map according to the detected line segment endpoints; then, take the effective image mask blocks for merging and use image morphological operations similar to those during detection. Finally, recalculate the final line segment according to the binary contour. The process of line segment detection is as Figure 4 shown, and specifically includes the following processing steps:

[0082] Input the mind map image Img, determine the coordinate information Rects of the text region, and obtain the pixel number value C for text detection cutting and the overlapping pixel number value O for text detection cutting.

[0083] Fine-tune the image size: Calculate the height C_h and width C_w of the cut image blocks according to the size of the input mind map image and parameter C, and adjust the image size so that the height and width of the adjusted image Img_1 can be exactly divisible by C_h and C_w respectively.

[0084] Adaptive image cutting: Mirror-fill the edges of the Img_1 image according to parameter O to obtain the filled image Img_2. Then, cut on the Img_2 image with the length and width of C_h + 2*O and C_w + 2*O respectively, and move the length and width step sizes of C_h and C_w respectively. Finally, obtain M*N image blocks of the same size.

[0085] Line segment detection: Use the HAWP line segment detection model to detect line segments in each of the M*N images respectively. The HAWP model outputs the two endpoints of the line segment in pairs to represent the line segment. Then, based on the results of the line segment detection, generate three binary mask images of corresponding horizontal lines, vertical lines, and diagonal lines respectively according to the slopes of the line segments.

[0086] Merge the binary mask images: Take the effective cutting regions of each binary mask image block respectively, splice them into the corresponding horizontal line, vertical line, and diagonal line whole images according to the size and position of the image Img_1, and use morphological operations such as dilation and erosion to eliminate the split regions. The effective cutting region is the central region in the mask image block with a length and width of C_h and C_w respectively.

[0087] Remove the misdetected line segments in the text region: Use the detected text coordinate position information Rects to erase the straight line binary mask in the text region and update the entire binary mask image.

[0088] Obtain line segment information: By respectively finding the closed contours of the merged horizontal line, vertical line, and diagonal line binary mask images, recalculate the line segment information in the image Img_1.

[0089] Output the line segment information.

[0090] In the embodiments of the present application, in order to accurately obtain a new complete line segment representation and morphological operations can be used. The line segments output three binary mask images of horizontal lines, vertical lines, and diagonal lines respectively according to different set slope ranges. This can better avoid the problem that the line segment contours change due to the intersection of different line segments and the straight lines cannot be recognized. In addition, for the situation where the texture of some characters is misdetected as a straight line, in order to eliminate such errors, after merging the binary mask images, the straight lines at the corresponding positions will be erased in combination with the position information of the text blocks.

[0091] 3. Match the connection relationships of text blocks.

[0092] After completing text detection and recognition and line segment detection, the embodiments of the present application further identify and match the relationships between line segments, between text blocks and line segments, and between text block and text block nodes.

[0093] First, based on the spatial position distribution characteristics and statistical characteristics of horizontal and vertical lines, confirm the categories of horizontal and vertical lines and the line segment intersection situations between them. Then, based on the intersection situations, identify and confirm the relationships between line segments. At the same time, match the relationship between text blocks and corresponding line segments according to certain distance rules, usually horizontal line segments are matched with text blocks. Next, based on the confirmed relationships between line segments and between text blocks and line segments, confirm the relationships between different text block nodes, and delete duplicate and some irregular text block node connection relationships. Finally, reconstruct the mind map on the editing drawing board according to the confirmed parent-child connection relationships between different text block nodes. The matching process of text block connection relationships is as Figure 5 shown, and specifically includes the following processing steps:

[0094] Input the detected line segments Lines and the coordinate information Rects of the detected text block rectangular frames.

[0095] Line segment sorting: Sort the line segments Lines according to the spatial position of the left endpoints of the line segments.

[0096] Statistical line segment intersection characteristics: Calculate whether each line segment intersects with other line segments respectively. If it intersects, it is 1, if it does not intersect, it is 0, and the line segment itself is 0. Suppose there are n line segments, then finally an n*n 0-1 matrix M will be generated. Count each row of the matrix M to obtain the number of intersections k of a line segment with other line segments.

[0097] Line segment cutting: Analyze the horizontal lines intersecting with vertical lines. If the line segment lengths of the horizontal lines intersecting on both sides of the vertical line both meet the conditions for being an independent single line segment, then cut the horizontal lines to generate multiple new line segments. At the same time, readjust the endpoint values of the line segments that do not meet the conditions. After completing all line segment cutting and correction, obtain a new line segment combination Lines_new.

[0098] Line segment sorting: Sort the newly obtained line segments Lines_new according to the spatial position of the left endpoints of the line segments.

[0099] Line segment to line segment matching: According to the intersection statistics of horizontal and vertical lines, if a vertical line has 2 or more horizontal lines intersecting with it, assume that the vertical line is a meaningful vertical line. Then count the number of horizontal lines on both sides of the vertical line respectively, and confirm the main node horizontal line and the secondary node horizontal line according to the number of horizontal lines on both sides, so as to complete the parent-child matching relationship between line segments.

[0100] Line segment and text block matching: According to the matching relationship between line segments, further match the corresponding text blocks according to the relationship of the sum of the minimum distances from the rectangular frames Rects of the text blocks to the horizontal lines and the corresponding meaningful vertical lines respectively. Thus, the matching relationship between the horizontal line segments and the text blocks is completed.

[0101] Text block and text block node matching: According to the matching relationship between line segments and between line segments and text blocks, obtain the hierarchical relationship between the main node text blocks and the sub-node text blocks, and then complete the matching relationship between the text blocks and the text block nodes.

[0102] Supplement missing unmatched text blocks: Between the corresponding vertical line and the sub-node horizontal line, there may be undetected horizontal lines and unmatched text blocks. Match the unmatched sub-node text blocks in this area with the corresponding main node text blocks.

[0103] Exclude duplicate and non-compliant matching relationships.

[0104] Output the parent-child connection relationships of the matching between different text blocks and text block nodes.

[0105] The spatial distribution feature of the straight-line connection type mind map referred to in the embodiments of the present application is that a vertical line often intersects multiple horizontal lines, and there is a matching relationship between the text and the horizontal lines. The statistical feature is that a meaningful vertical line intersects at least 3 horizontal lines, has 3 intersection points, and the number of intersecting horizontal lines on both sides of the vertical line has an obvious asymmetry feature, usually a one-to-many relationship, and this relationship just corresponds to the relationship between the text parent node and the sub-node. The distance rule referred to in the matching of the text block and the line segment in the embodiments of the present application is the sum of the minimum distances from the text border to the matching horizontal line and the vertical line respectively.

[0106] 4. After the reconstruction of the mind map is completed, edit the obtained structured data as the original data of the new mind map, and then generate a new style of mind map according to the user's needs.

[0107] In the embodiments of the present application, before performing text detection on the mind map image, first perform adaptive block cutting and scaling on the image, then perform text region detection on the image blocks obtained by block cutting respectively, then merge the detection results of the image blocks into the coordinate system of the original image through operations such as image morphology, and finally perform text recognition, and display the recognition results on the editing drawing board interface according to the relative position relationship of the original image. The user can modify and generate a new mind map with the corresponding style according to his own needs. The embodiments of the present application effectively solve the technical problem of low accuracy of text detection and recognition in common mind maps, can not only meet the detection and recognition requirements of mind maps with different layout specifications, but also reduce the demand for computing video memory and improve the efficiency of users in producing mind maps or converting the style of mind maps.

[0108] Figure 6 This is a schematic diagram of the composition structure of the mind map image recognition and parsing and reconstruction device provided by the embodiments of the present application. As Figure 6 shown, the mind map image recognition and parsing and reconstruction device of the embodiments of the present application includes:

[0109] A cutting unit 60, configured to adjust the size of the mind map image, and perform adaptive cutting and partitioning on the adjusted image to obtain image blocks;

[0110] A detection unit 61, configured to detect text regions or line segments in the image blocks, and generate corresponding binary mask image blocks of the text or line segments;

[0111] A merging unit 62, configured to reassemble the binary mask image blocks of the text or line segments into a complete binary mask image of the text or a binary mask image of the line segments based on the size of the adjusted image;

[0112] An identification unit 63, configured to extract text region image blocks at corresponding positions in the adjusted mind map according to the text region coordinates in the binary mask image of the text, and identify the text information in the text region image blocks; correct the line segment information according to the binary mask image of the line segments, and identify the matching relationship between the line segments and the text block nodes and the parent-child node relationship between the text block nodes according to the corrected line segment information;

[0113] An editing unit 64, configured to arrange and edit the identified text information on an editing drawing board according to the connection relationship and relative position information of its parent-child nodes.

[0114] As an implementation manner, the editing unit 64 is further configured to:

[0115] Respond to an editing request for the structured text content after arrangement, edit the structured text content to obtain the original data of a new mind map, and generate a new mind map in response to a modification request for the original data.

[0116] As an implementation manner, the cutting unit 60 is further configured to:

[0117] According to a preset number value of text detection and cutting pixels and a number value of overlapping pixels for text detection and cutting, adopt a cutting method with a partially overlapping region of the image blocks to cut the adjusted image into image blocks with consistent sizes.

[0118] As an implementation manner, the identification unit 63 is further configured to:

[0119] By finding the closed contour of the binary mask image of the spliced text, the coordinate position information of the text area in the adjusted image is obtained again, and the area corresponding to the coordinate position in the adjusted image is extracted as the text area image block.

[0120] As an implementation manner, the line segment includes at least one of the following:

[0121] Horizontal line segment, vertical line segment, oblique line segment;

[0122] Correspondingly, the editing unit 64 is further configured to:

[0123] Determine the spatial position distribution characteristics and statistical characteristics of the horizontal and vertical line segments, confirm the categories of the horizontal and vertical line segments and the line segment intersection situation between them; according to the intersection situation, identify the relationship between line segments, and match the relationship between the text block nodes and the corresponding line segments according to the position and distance rules, and determine the matching relationship between the horizontal line segment and the text block nodes; wherein, when determining that a vertical line segment intersects with multiple horizontal line segments, the corresponding text block nodes have a matching relationship with the horizontal line segments;

[0124] According to the determined relationships between line segments and between text block nodes and line segments, determine the relationships between different text block nodes, and delete duplicate and non-compliant node connection relationships as the relationships between text block nodes and the final line segments.

[0125] In an exemplary embodiment, the above-mentioned various processing units of the nested table extraction device of the embodiments of the present application can be implemented by one or more central processing units (CPUs, Central Processing Unit), graphics processing units (GPUs, Graphics Processing Unit), baseband processors (BP, Base Processor), application specific integrated circuits (ASICs, Application Specific Integrated Circuit), DSPs, programmable logic devices (PLDs, Programmable Logic Device), complex programmable logic devices (CPLDs, Complex Programmable Logic Device), field programmable gate arrays (FPGAs, Field-Programmable Gate Array), general purpose processors, controllers, microcontroller units (MCUs, Micro Controller Unit), microprocessors (Microprocessor), or other electronic components.

[0126] In the embodiments of the present disclosure, Figure 6The specific ways in which each processing unit in the shown mind map image recognition and parsing and reconstruction device performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0127] An embodiment of the present application also records a storage medium, on which an executable program is stored, and when the executable program is executed by a processor, the steps of the mind map recognition and reconstruction method of the embodiment are implemented.

[0128] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the magnitude of the serial numbers of the above processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The serial numbers of the embodiments of the present invention above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0129] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0130] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0131] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0132] In addition, each functional unit in the embodiments of the present invention may be all integrated in one processing unit, or each unit may be separately taken as one unit, or two or more units may be integrated in one unit; the above-mentioned integrated units may be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0133] As mentioned above, the above are only the embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for mind map image recognition, analysis and reconstruction, characterized in that The method includes: Adjust the size of the mind map image, and perform adaptive cutting and partitioning on the adjusted image to obtain image blocks. The step of performing cutting and partitioning on the adjusted image to obtain image blocks includes: According to a preset text detection cutting pixel quantity value and a text detection cutting overlapping pixel quantity value, adopt a cutting method with partial overlapping regions for the image blocks to cut the adjusted image into image blocks with consistent sizes; Detect the text in the image blocks and generate corresponding binary mask maps of text regions. Based on the size of the adjusted image, splice the binary mask maps of the partitioned text regions into a complete binary mask map of text; Extract the text region image blocks at the corresponding positions in the adjusted mind map according to the text region coordinates in the binary mask map of text, and identify the text information in the text region image blocks and extract the corresponding position relationships. The step of extracting the text image blocks in the corresponding regions of the adjusted mind map includes: Re-obtain the coordinate position information of the text region in the adjusted image by finding the closed contour of the spliced binary mask map of text, and extract the region corresponding to the coordinate position in the adjusted image as the text region image block; Cut the image, detect the line segments in the image blocks, generate a complete binary mask map of line segments according to the line segment information detected in the image, and correct the line segments again based on the binary contours in the binary mask map of line segments; Determine the intersection situation between the finally corrected line segments. According to the intersection situation and position relationships, determine the matching relationship between the finally determined line segments and the text block nodes and the parent-child connection relationships between different text block nodes. Reconstruct the mind map according to the determined parent-child relationships and position information of the text block nodes, and arrange it on the editing drawing board. The step of reconstructing the mind map includes: In response to an editing request for the structured text content after arrangement, edit the structured text content to obtain the initial data of a new mind map, and in response to a modification request for the initial data, generate a new mind map.

2. The method according to claim 1, characterized in that, The line segments include at least one of the following: Horizontal line segments, vertical line segments, oblique line segments; Correspondingly, the step of determining the relationship between the text block nodes after arrangement and the finally determined line segments according to the position and distance rules includes: Determine the spatial position distribution characteristics and statistical characteristics of the horizontal and vertical line segments, confirm the categories of the horizontal and vertical line segments and the intersection situation between them; according to the intersection situation, identify the relationships between the line segments, and match the relationship between the text block nodes and the corresponding line segments according to the distance rules to determine the matching relationship between the horizontal line segments and the text block nodes; when it is determined that one vertical line segment intersects multiple horizontal line segments, the corresponding text block node has a matching relationship with the horizontal line segments; According to the determined relationships between line segments and between text block nodes and line segments, determine the parent-child node relationships between text block nodes, and delete the duplicate and non-compliant line segment connection relationships as the relationship between the text block nodes and the finally determined line segments.

3. A mind map image recognition and parsing and reconstruction device, characterized in that, The device includes: A cutting unit, configured to adjust the size of the mind map image, and adaptively cut and divide the adjusted image to obtain image blocks. The cutting unit is further configured to: According to a preset text detection cutting pixel quantity value and a text detection cutting overlapping pixel quantity value, adopt a cutting method with a partially overlapping area of the image blocks to cut the adjusted image into image blocks with consistent sizes; A detection unit, configured to detect text or line segments in the image blocks and generate corresponding binary mask image blocks of the text or line segments; A merging unit, configured to reassemble the binary mask image blocks of the text or line segments into a complete binary mask image of the text or a binary mask image of the line segments based on the size of the adjusted image; An identification unit, configured to extract a text area image block at a corresponding position in the adjusted mind map according to the text area coordinates in the binary mask image of the text, and identify the text information in the text area image block; correct the line segment information according to the binary mask image of the line segments, and identify the matching relationship between the line segments and the text block nodes and the parent-child node relationship between different text block nodes. The identification unit is further configured to: Re-obtain the coordinate position information of the text area in the adjusted image by finding the closed contour of the spliced binary mask image of the text, and extract the area corresponding to the coordinate position in the adjusted image as the text area image block; An editing unit, configured to arrange and edit the identified text information on an editing drawing board according to the connection relationship and relative position information of its parent-child nodes. The editing unit is further configured to: In response to an editing request for the arranged structured text content, edit the structured text content to obtain initial data of a new mind map, and in response to a modification request for the initial data, generate a new mind map.

4. The device according to claim 3, characterized in that The line segments include at least one of the following: Horizontal line segments, vertical line segments, and oblique line segments; Correspondingly, the editing unit is further configured to: Determine the spatial position distribution characteristics and statistical characteristics of the horizontal and vertical line segments, confirm the categories of the horizontal and vertical line segments and the line segment intersection situation between them; according to the intersection situation, identify the relationship between the line segments, and match the relationship between the text block nodes and the corresponding line segments according to the position and distance rules, and determine the matching relationship between the horizontal line segments and the text block nodes; wherein, when determining that a vertical line segment intersects multiple horizontal line segments, the corresponding text block node has a matching relationship with the horizontal line segments; According to the determined relationships between the line segments and between the text block nodes and the line segments, determine the relationships between different text block nodes, and delete the duplicate and non-conforming node connection relationships as the relationships between the text block nodes and the final line segments.

Citation Information

Patent Citations

  • Mind map recognition method and device, storage medium and computer equipment

    CN108304763A

  • Workpiece metal surface character recognition method and system based on image segmentation

    CN111160352A