Multi-dimensional data processing method and system

By acquiring vector image data, parsing document structure and classifying text elements, and recognizing table row and column structures, the data structuring problem in optical character recognition technology is solved, achieving efficient and accurate data processing and storage.

CN121033876APending Publication Date: 2025-11-28LONGSHINE TECH

Patent Information

Application Number
CN202511411063.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

While existing optical character recognition technologies can recognize text when processing image data, the output is a one-dimensional text stream, which is difficult to use directly. This results in low automation of data acquisition, easy introduction of errors, and inability to effectively recover the structured information of the data.

Method used

By acquiring vector image data, parsing the document structure, extracting text and graphic elements according to grouping tags, classifying text elements into titles and data text based on preset recognition target objects, recognizing the row and column structure of tables, constructing structured data, and performing format conversion and storage.

Benefits of technology

It enables efficient and accurate processing from vector images to structured data, reduces computing resource requirements, maintains the logical structure and integrity of the data, and improves the efficiency of automated processing and recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033876A_ABST
    Figure CN121033876A_ABST
Patent Text Reader

Abstract

The invention provides a multi-dimensional data processing method and system, and belongs to the technical field of data processing, and the method comprises the steps: obtaining vector image data containing to-be-recognized data; analyzing a document structure of the vector image data, and extracting text elements and graphic elements according to grouping tags in a document; according to a preset recognition target object, the text elements are classified into title text elements and data text elements, and the recognition target object comprises at least one of a table title or a text title; according to the title text element, the data text element and the graphic element, identifying a row-column structure of the table data to construct structured data; and carrying out format conversion on the structured data and then storing the structured data. According to the method, the document structure of the vector image is analyzed, and the row-column logic structure of the table is reconstructed by utilizing the coordinate information and the spatial position relationship of the text elements and the graphic elements, so that the vector image data is efficiently and accurately processed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a multi-dimensional data processing method and system. BACKGROUND

[0002] Under the background of digital transformation, automatically extracting structured data from various images has become an important demand for enterprises and institutions to improve information processing efficiency. Especially in business systems, a large amount of information such as reports and statistical charts exists in the form of images, and the data in them needs to be extracted for subsequent analysis and processing.

[0003] In order to achieve this purpose, the technology widely used at present is optical character recognition (OCR) technology. This technology can process image files, recognize the text in them and convert them into machine-readable text format. In application, the OCR technology usually first analyzes the input image to locate the area containing the text, and then identifies the characters in these areas one by one, and finally generates a text sequence containing all the recognized characters. However, since the core goal of OCR technology is to identify characters themselves, its output result is often a one-dimensional text stream. Therefore, even if the recognized text content is completely correct, these data are difficult to use directly because they have lost their original structure. In order to restore the usability of data, subsequent manual intervention is often needed to manually process the text stream with reference to the original image, which not only greatly reduces the degree of automation and efficiency of data acquisition, but also easily introduces new errors in the processing process. SUMMARY

[0004] The present application provides a multi-dimensional data processing method, system, electronic device, storage medium and computer program product to solve the defects in the prior art and realize efficient and accurate processing of vector image data.

[0005] The present application provides a multi-dimensional data processing method, comprising the following steps: Obtaining vector image data containing data to be recognized; Parsing the document structure of the vector image data, and extracting text elements and graphic elements according to the grouping tags in the document; According to a predetermined recognition target object, the text elements are classified into title text elements and data text elements, wherein the recognition target object includes at least one of a table title or a text title; According to the title text elements, the data text elements and the graphic elements, the row and column structure of the table data is recognized to construct structured data; The structured data is stored after format conversion.

[0006] The multi-dimension data processing method according to the present application, wherein the text elements are classified into title text elements and data text elements according to a preset identification target object, comprises: When the identification target object is a table title, the text elements are traversed, the text elements with text content matching a preset table title name are classified as the title text elements, and the remaining text elements in the text elements are classified as the data text elements; When the identification target object is a text title, the text elements are traversed, the text elements with title attributes matching a preset text title name are identified as the title text elements.

[0007] The multi-dimension data processing method according to the present application, wherein a row-column structure of a table is identified according to the title text elements, the data text elements and the graphic elements, comprises: A corresponding relationship between the data text elements and the title text elements is determined according to a spatial position relationship between the title text elements and the data text elements, wherein the spatial position relationship is determined by analyzing coordinate information of the title text elements and the data text elements and / or relative positions of the title text elements, the data text elements and graphic elements; A row-column structure of a table is reconstructed according to the corresponding relationship and an arrangement rule of the data text elements.

[0008] The multi-dimension data processing method according to the present application, wherein a corresponding relationship between the data text elements and the title text elements is determined according to a spatial position relationship between the title text elements and the data text elements, comprises: When the title text elements are located in a first row of a table, a coordinate distance between the data text elements and the title text elements is calculated to determine the corresponding relationship; When the title text elements are located in a first column of a table, a containing relationship of the title text elements, the data text elements and rectangular elements in the graphic elements is analyzed to determine the corresponding relationship.

[0009] The multi-dimension data processing method according to the present application, wherein a coordinate distance between the data text elements and the title text elements is calculated to determine the corresponding relationship, comprises: A horizontal coordinate distance between each of the data text elements and each of the title text elements is calculated; Each of the data text elements is attributed to a column corresponding to a nearest title text element according to the horizontal coordinate distance; Sort according to the vertical coordinates of the data text elements belonging to the same column to determine the row sequence.

[0010] According to the multi-dimension data processing method provided by the present application, the analysis of the containing relationship between the title text elements, the data text elements and the rectangular elements in the graphic elements, and the determination of the corresponding relationship, comprises: Determine whether the coordinates of each title text element and data text element are within the boundary range defined by the rectangular element to determine the cell position to which each belongs; According to the cell position, identify the title text element in the first column as a row title; Determine the data content of each row by analyzing the data text elements in the same vertical range as the row title.

[0011] According to the multi-dimension data processing method provided by the present application, the determination of whether the coordinates of each title text element and data text element are within the boundary range defined by the rectangular element comprises: Compare the horizontal coordinates of the title text element or the data text element with the horizontal boundary of the rectangular element, wherein the horizontal boundary is defined by the horizontal coordinate of the rectangular element as the left boundary and the sum of the horizontal coordinate and the width of the rectangular element as the right boundary; Compare the vertical coordinates of the title text element or the data text element with the vertical boundary of the rectangular element, wherein the vertical boundary is defined by the vertical coordinate of the rectangular element as the upper boundary and the sum of the vertical coordinate and the height of the rectangular element as the lower boundary; When the coordinates of the title text element or the data text element are within the horizontal boundary and the vertical boundary at the same time, it is determined that the coordinates of the title text element and the data text element are within the boundary range defined by the rectangular element.

[0012] According to the multi-dimension data processing method provided by the present application, it further comprises: Detect whether there is a merged cell in the table; When a merged cell is detected, calculate the number of merged cells according to the ratio of the distance between the cells where the adjacent title text elements or data text elements are located to the standard cell size; Adjust the table structure according to the number of merged cells.

[0013] The present application further provides a multi-dimension data processing system, comprising the following modules: An acquisition module for acquiring vector image data containing data to be recognized; The processing module is configured to parse a document structure of the vector image data, and extract text elements and graphic elements according to grouping tags in the document; The processing module is further configured to classify the text elements into title text elements and data text elements according to a preset identification target object, wherein the identification target object includes at least one of a table title or a text title; The processing module is further configured to identify a row-column structure of table data according to the title text elements, the data text elements and the graphic elements, so as to construct structured data. The storage module is configured to store the structured data after format conversion.

[0014] The present application further provides an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the multi-dimensional data processing method when executing the program.

[0015] The present application further provides a non-transitory computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the multi-dimensional data processing method.

[0016] The present application further provides a computer program product, which includes a computer program, wherein the computer program is executed by a processor to implement the multi-dimensional data processing method.

[0017] In summary, the one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: By acquiring the vector image data scheme containing the data to be recognized, the structured text information embedded in the vector image can be directly utilized, thereby avoiding the large amount of computing resources required for pixel analysis of a bitmap in the traditional technology, and reducing the requirement for device performance. By analyzing the document structure of the vector image data, the text elements and graphic elements are extracted according to the grouping labels in the document, the position, content and organization relationship of the elements can be completely retained, thereby providing accurate basic information for subsequent accurate restoration of the data logical structure. By classifying the text elements into title text elements and data text elements according to the preset recognition target object, the contents of different semantic roles in the data can be distinguished in an automated manner, thereby providing a basis for understanding the data structure and establishing the row-column correspondence relationship. By recognizing the row-column structure of the table data according to the title text elements, the data text elements and the graphic elements to construct the structured data, the discrete distributed text elements are restored to the two-dimensional table structure with clear logical relationship, the accuracy of recognition is improved by comprehensively utilizing the information of multiple types of elements, and the structured data maintaining the integrity of the original data is successfully constructed. By storing the structured data after format conversion, the complete conversion process from the vector image to the database is realized, thereby realizing efficient and accurate processing of the vector image data. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0019] Figure 1 is one of the flowcharts of the multi-dimensional data processing method provided by the present application.

[0020] Figure 2 is the second flowchart of the multi-dimensional data processing method provided by the present application.

[0021] Figure 3 is the third flowchart of the multi-dimensional data processing method provided by the present application.

[0022] Figure 4 is the fourth flowchart of the multi-dimensional data processing method provided by the present application.

[0023] Figure 5 is the fifth flowchart of the multi-dimensional data processing method provided by the present application.

[0024] Figure 6 is the sixth flowchart of the multi-dimensional data processing method provided by the present application.

[0025] Figure 7 is the seventh flowchart of the multi-dimension data processing method provided by the present application.

[0026] Figure 8 is the eighth flowchart of the multi-dimension data processing method provided by the present application.

[0027] Figure 9 is the structural diagram of the multi-dimension data processing system provided by the present application.

[0028] Figure 10 is the structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0029] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. According to the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0030] It should be noted that in the description of the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover the non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element. The terms "upper", "lower" and the like indicate the orientation or positional relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the indicated system or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0031] The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a particular order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second" and the like are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / ", generally means that the front and rear associated objects are in a "or" relationship.

[0032] The application is described below Figures 1-10 The application provides a multi-dimensional data processing method, a system, an electronic device, a storage medium and a computer program product.

[0033] Referring to Figure 1 , Figure 1 is one of the flowcharts of the multi-dimensional data processing method provided by the application, and the application comprises steps 101 to 105: Step 101: acquiring vector image data containing to-be-recognized data.

[0034] The embodiment provides a configurable multi-dimensional data collection and processing method, which acquires vector image data containing to-be-recognized data as a starting step of processing.

[0035] In the prior art, a traditional image recognition method mainly processes bitmap format images, and needs to analyze pixel level data through a complex optical character recognition technology. This method not only consumes a large amount of computing resources, but also is difficult to maintain the logical relationship between data when processing images containing structured data. Therefore, the embodiment selects to acquire vector image data as a recognition object, which fundamentally changes the technical path of data recognition.

[0036] Vector image data refers to a data format that describes image content by using mathematical equations, wherein the vector image data adopts a Scalable Vector Graphics (SVG) format, and the SVG format is an open standard based on XML. Unlike a bitmap, a vector image stores attribute information of a graphic element in a text form, including text content, coordinate positions, geometric shapes and the like. To-be-recognized data refers to information that needs to be extracted from an image and converted into a structured format, and usually includes table data, text data and other business related content.

[0037] In a specific implementation, first, a URL address and authentication information of a target server are read through a configuration file. Then, an HTTP request library is used to initiate an access request to the server, and SVG format vector image data is extracted from response content. When a target page contains multiple subpages, a navigation homepage is accessed first to acquire vector image data, the access path and request parameters of a subpage are acquired by analyzing the data, and then the specific subpage is accessed according to the information and corresponding vector image data is acquired. The entire acquisition process can be automatically executed through a timing task scheduling library, and the latest image data is acquired periodically according to a preset time interval.

[0038] It should be noted that the embodiment before obtaining the vector image data containing the data to be recognized, still includes the steps of local registration authorization and server authorization authentication, to ensure the security and legality of the data acquisition process.

[0039] In the application of automated data collection, unauthorized access not only may cause data leakage, but also may cause the risk of illegal access to the target server. Therefore, a perfect authorization authentication mechanism needs to be established, which not only prevents the software from running on unauthorized devices, but also ensures the legality of access to the server.

[0040] The local registration and authorization process first generates a unique identification code based on the hard disk serial number of the deployed device. The hard disk serial number is the physical identification of the device, which has uniqueness and unchangeability, and can effectively identify a specific hardware device. After generating the identification code, the data is encrypted using the public key of the RSA asymmetric encryption technology, and the encrypted file is persistently stored locally. The RSA encryption technology uses public key encryption and private key decryption. Even if the encrypted file is copied to other devices, without the corresponding private key, the valid authorization information cannot be obtained.

[0041] In the verification phase, the program reads the encrypted file stored locally and decrypts it using the corresponding private key to restore the original registration information. After successful decryption, a double-check mechanism is executed. On the one hand, the decrypted registration code is compared with the physical identification of the current device to ensure that the software runs on a specific authorized device. On the other hand, the validity period of the authorization certificate is verified to control the software usage time. Only when the device identification matches and is within the valid period, the program is allowed to continue the subsequent operation.

[0042] After completing the local authorization, server authorization authentication is also required. First, read the URL address of the server and the authentication account and password from the configuration file. The configuration file is usually stored in an encrypted or permission-protected manner to prevent the leakage of authentication information. After reading the configuration information, use RSA public key encryption technology to encrypt the username and password. The public key used here is provided by the server, ensuring that only the server can use the corresponding private key to decrypt the authentication information.

[0043] After encryption, send a request to the server's authentication interface through the HTTP request library. The request contains encrypted authentication information. Even if it is intercepted during network transmission, attackers cannot obtain the plaintext username and password. After receiving the request, the server uses the private key to decrypt the authentication information and verifies its validity. After successful authentication, the server returns a session identifier, which the program will save for subsequent data access requests.

[0044] To maintain the persistence and security of access, the program re-authenticates as needed by setting a timeout mechanism and verifying the validity of the session. When the session times out or the service side requires re-authentication, the program automatically repeats the above authentication process to obtain a new session identifier. This mechanism not only ensures that long-running data collection tasks will not be interrupted due to session expiration, but also prevents the risk of unauthorized access using expired sessions.

[0045] Step 102: Parse the document structure of the vector image data, extract text elements and graphic elements according to the grouping tags in the document.

[0046] After obtaining the vector image data containing the data to be recognized, the embodiment needs to parse the document structure of the vector image data and extract text elements and graphic elements according to the grouping tags in the document.

[0047] Although the obtained vector image data contains complete graphic and text information, these information are stored in the form of document structure in XML format and need to be parsed to extract available data elements. Directly performing string processing on the original XML text is not only inefficient but also easy to miss the hierarchical relationship and attribute information between elements. Therefore, a structured parsing method is needed to extract various elements in the document according to their types and organizational relationships.

[0048] The document structure refers to the hierarchical organization method based on XML adopted by SVG format, in which various graphics and text contents are defined and organized through specific tags. The grouping tags in SVG are represented as <g>Elements, which are used to organize related graphical elements together to form a logical grouping. Text elements refer to the SVG's <text>The tag and its sub-elements contain the actual text content as well as the position, style, and other attribute information of the text. The graphic elements include various geometric graphic tags, such as rectangle <rect>, circular <circle>etc. In the present application, the graphic element is preferably a rectangular element.

[0049] In a specific implementation, the acquired SVG data is first parsed using an XML parsing library to construct a document object model. In this embodiment, the Beautiful Soup library is used as the parsing tool, which can convert an XML document into a traversable tree structure. During parsing, the program first locates the root element of the SVG document, and then traverses its child elements layer by layer.

[0050] When encountering a grouping tag <g>At this time, the program extracts all the sub-elements in the group as a processing unit. This group processing mode maintains the logical relationship between the elements, because the elements in the same group are usually semantically related. For each group, the program extracts the text elements and the graphic elements respectively.

[0051] When extracting the text elements, the program looks for all the <text>Tags. For each text element, not only its text content is extracted, but also its coordinate information, including x and y attribute values. If the text element contains transform attributes, the actual coordinate positions after transformation are also calculated. These coordinate information is crucial for determining the spatial relationships between text elements later.

[0052] Focus on rectangular elements when extracting graphical elements <rect>Because in table-like images, rectangles are often used to draw cell borders. For each rectangle element, the program extracts its x, y coordinates and width, height attributes, which define the position and size of the rectangle. The program records these rectangle elements according to the group they belong to, maintaining their organizational structure.

[0053] During the extraction process, the program also needs to handle nested grouping cases. When a group tag contains other group tags inside, each sub-group needs to be processed recursively to ensure that all levels of elements are correctly extracted. At the same time, the program saves the group information of each element, which helps to understand the hierarchical relationship between elements.

[0054] Step 103: According to the pre-set identification target object, the text elements are classified into title text elements and data text elements, wherein the identification target object includes at least one of table title or text title.

[0055] The text elements extracted from the vector image data contain complete text content and position information, but these elements have different roles in semantics. In images containing structured data, some text elements serve to identify the meaning of data, such as table column headers or row headers, while other text elements are actual data content. If these text elements are not classified, the organizational structure and meaning of the data cannot be understood. Therefore, a classification mechanism is needed to distinguish text elements with different functions.

[0056] The identification target object refers to the key information set by the user in advance to identify specific data. In practical applications, users usually know which data needs to be extracted, and these data are often identified by specific titles in images. Title text elements refer to text that plays an identifying role in images, used to explain the meaning or category of its associated data. Data text elements refer to text containing actual data content, which is the target information that users ultimately need to extract.

[0057] The specific implementation of step 103 will be described in detail in subsequent embodiments.

[0058] Step 104: According to the title text elements, data text elements and graphic elements, the row and column structure of the table data is identified to construct structured data.

[0059] Although the text elements have been successfully classified, these elements are still discretely distributed in the vector image, lacking clear organizational relationships. In the original vector image, the row-column structure of the table is embodied through visual presentation, while the extracted text elements only contain content and coordinate information, losing the original table structure. If this structure relationship cannot be accurately restored, it is impossible to convert the discrete text data into usable structured data. Therefore, it is necessary to reconstruct the logical structure of the table by analyzing the relationships between various elements.

[0060] The row-column structure of table data refers to the organization method of data in two-dimensional space, where rows represent data records and columns represent data attributes. Structured data refers to data organized according to a predefined pattern, with clear row-column relationships and data types, facilitating querying, analysis, and processing.

[0061] The specific implementation of step 103 will be described in detail in subsequent embodiments.

[0062] Step 105: Store the structured data after format conversion.

[0063] Although the structured data has been successfully constructed, these data still exist in the memory of the program in the form of temporary data structures. If it is not converted into a standard data format and stored persistently, the data will be lost after the program ends and cannot be used for subsequent analysis and application. At the same time, different application scenarios have different requirements for data format, and it is necessary to convert the recognition results into a format suitable for specific purposes. Therefore, a complete data conversion and storage mechanism must be established.

[0064] Format conversion refers to the conversion of structured data in memory into a specific data format to meet different application needs. Storage refers to the persistent storage of converted data to storage media, so that it can be saved for a long time and reused.

[0065] In specific implementation, first, the recognized table data and text data are converted into pandas DataFrame format. DataFrame is a two-dimensional labeled data structure with row and column indices, which is very suitable for representing table-type data. During the conversion process, the program maps the structured data in the form of a two-dimensional array to DataFrame, with the title text elements as column names and the data text elements filled into the corresponding cells.

[0066] After conversion to DataFrame format, the program further processes the data according to actual needs. This includes selective interception of rows and columns, such as retaining only specific columns or filtering out unwanted rows. When multiple recognition results need to be merged, the program merges the DataFrame data of different groups to form a unified data set.

[0067] To enhance the traceability of data and facilitate subsequent analysis, the program inserts additional metadata in the DataFrame. These metadata include the collection time, which records the timestamp of data acquisition, for building time series data. At the same time, data classification codes and names are inserted to identify the type and source of data, facilitating classification management in the database.

[0068] After completing the data format conversion, the program performs data warehousing operations. First, read the database connection information from the configuration file, including the database type, IP address, port number, username, and password. According to the configured database type, the program uses the corresponding database connection library to establish a connection. For MySQL databases, use the pymysql library, and for SQLite databases, use the sqlite3 library.

[0069] After establishing the database connection, the program creates a connection pool to improve the efficiency of database operations. The connection pool maintains a set of reusable database connections, avoiding the overhead of frequent creation and destruction of connections. The program establishes a mapping relationship between data columns and table fields based on the structure of the DataFrame and the definition of the database table.

[0070] Based on the mapping relationship, the program dynamically generates SQL insertion statements. This dynamic generation method allows the program to adapt to different table structures without the need to write fixed SQL statements for each data type. After generating the SQL statement, the program creates a database cursor and uses batch insertion to write the data in the DataFrame into the database table. Batch insertion significantly improves the warehousing efficiency compared to individual insertion.

[0071] In a possible implementation, the embodiment provides a specific implementation of classifying text elements according to a preset identification target object, including processing steps when the identification target object is a table title and processing steps when the identification target object is a text title. Referring to Figure 2 , Figure 2 is a flowchart of the multi-dimensional data processing method provided by the present application, step 103 specifically includes the following steps: Step 201: When the identification target object is a table title, traverse the text elements, and classify the text elements whose text content matches the preset table title name as title text elements, and classify the remaining text elements in the text elements as data text elements.

[0072] Step 202: When the identification target object is a text title, traverse the text elements, and identify the text elements whose title attribute matches the preset text title name as title text elements.

[0073] In practical applications, the text elements contained in vector image data can originate from different types of data structures. Table-type data often contains column headers or row headers, which are used to explain the meaning of the data. Independent text data may be identified by specific header attributes. Therefore, different classification strategies need to be adopted according to the type of the target object to accurately identify text elements with different functions.

[0074] Table headers refer to the text used to identify the meaning of columns or rows in a table, usually located in the first row or column of the table. Text content refers to the actual character information contained in the text element. Header attributes refer to special attribute markers that some text elements may have in vector image formats, used to identify the function or type of the text.

[0075] When the target object is a table header, the program performs the following processing flow. First, get the user's preset table header names, which may be one or more, representing the column names or row names of the table to be identified. For example, in a table containing sales data, the preset table headers may include "product name", "sales quantity", "sales amount", etc.

[0076] After obtaining the preset table header names, the program begins to traverse all the text elements extracted from the vector image data. For each text element, the program extracts its text content and matches it with the preset table header names. The matching process can use exact matching or fuzzy matching. Exact matching requires the text content to be exactly the same as the preset name, suitable for scenarios where the header format is fixed. Fuzzy matching allows for certain differences, such as ignoring case, removing spaces, or using similarity algorithms, suitable for scenarios where the header format may vary.

[0077] When the text content of a text element matches the preset table header name, the program marks the text element as a header text element. These header text elements will be used to determine the structure of the table in subsequent processing. After all the matching is completed, the remaining parts of the text elements that are not marked as header text elements are classified as data text elements. This exclusion method ensures that all text elements are correctly classified.

[0078] When the target object is a text header, the program uses a different processing strategy. In this case, the program does not match by text content, but by the header attribute of the text element. Some vector image formats allow text elements to be assigned specific attributes to identify the type or function of the text.

[0079] The program traverses all text elements and checks whether each element has a title attribute. If the title attribute exists, the program compares the attribute value with preset text title names. When the attribute value matches the preset name, the text element is identified as a title text element. This attribute-based identification method is more accurate because the attribute is metadata specifically used to identify the function of the text and is not affected by changes in the format of the text content.

[0080] In one possible implementation, the present embodiment provides a specific implementation of identifying the row-column structure of table data according to title text elements, data text elements, and graphical elements, including determining the correspondence between data text elements and title text elements and reconstructing the row-column structure of the table. Referring to Figure 3 , Figure 3 is a third flowchart of the multi-dimensional data processing method provided by the present application, and step 104 specifically includes the following steps: Step 301: determining the correspondence between data text elements and title text elements according to the spatial position relationship between the title text elements and the data text elements, wherein the spatial position relationship is determined by analyzing the coordinate information of the title text elements and the data text elements and / or the relative positions of the title text elements, the data text elements, and the graphical elements.

[0081] Step 302: reconstructing the row-column structure of the table according to the correspondence and the arrangement rule of the data text elements.

[0082] After the text elements are distinguished into titles and data, these elements are still discrete in the program, and their row-column relationship in the original vector image is not explicitly established. If each data text element cannot be accurately associated with its corresponding title and row, meaningful structured data cannot be formed, and thus the extracted information loses its inherent logical value. Therefore, a reliable method must be used to restore the structural relationship implied in the visual layout.

[0083] The spatial position relationship is determined by analyzing the coordinate information of the title text elements and the data text elements and / or the relative positions of the title text elements, the data text elements, and the graphical elements. The coordinate information refers to the accurate x and y coordinate values of each element extracted from the vector image data. The relative position refers to the positional relationship between elements, such as whether a text element is located within the boundary of a rectangular element. The correspondence refers to the logical link established during the identification process to associate a data text element with a title text element. The arrangement rule refers to the predictable spatial distribution pattern that the data text elements usually follow in the table layout, such as the elements in the same row having similar vertical coordinates or the elements in the same column having similar horizontal coordinates.

[0084] In a specific implementation, first, the corresponding relationship between the data text elements and the title text elements is determined according to the spatial position relationship between the title text elements and the data text elements. The purpose of this step is to find the column or row title to which each data text element belongs. Specifically, the program iterates through each data text element and analyzes its spatial position with all title text elements. This analysis may include calculating the coordinate distance between them or judging whether they are contained by the same graphical element. In this way, an explicit attribution relationship can be established for each data text element.

[0085] After determining the corresponding relationship between the data text elements and the title text elements, the row and column structure of the table is reconstructed according to the corresponding relationship and the arrangement rule of the data text elements. First, data text elements with the same corresponding relationship (i.e., corresponding to the same title text element) are grouped together, and this group of data constitutes a column or a row in the table. Then, the program analyzes the arrangement rule of these data text elements, such as determining the row order by sorting their vertical coordinates or determining the column order by sorting their horizontal coordinates. In this way, the originally discrete text elements are orderly organized to form a complete row and column structure, thereby constructing structured data.

[0086] In a possible implementation, the embodiment provides a specific implementation of determining the corresponding relationship between the data text elements and the title text elements, which adopts different strategies according to different layout modes of the table title. Referring to Figure 4 , Figure 4 is the fourth flowchart of the multi-dimensional data processing method provided by the present application, and step 301 specifically includes the following steps: Step 401: When the title text element is located in the first row of the table, calculate the coordinate distance between the data text element and the title text element to determine the corresponding relationship.

[0087] Step 402: When the title text element is located in the first column of the table, analyze the inclusion relationship between the title text element and the data text element and the rectangular element in the graphical element to determine the corresponding relationship.

[0088] In actual vector image data, the layout of the table is various, and the most common two forms are that the title is located in the first row of the table and the title is located in the first column of the table. The spatial position relationship between the data and the title has completely different characteristics in these two layout modes. If a single recognition strategy is adopted, it is likely to be accurate when processing one kind of layout, but may be wrong when processing another kind of layout. Therefore, the appropriate analysis method must be selected according to the actual layout mode of the table to ensure accurate determination of the corresponding relationship.

[0089] Coordinate distance refers to the geometric distance between two elements in a coordinate system, usually referring to horizontal or vertical distance. The containment relationship of rectangular elements refers to determining whether the coordinates of a text element lie within the boundary defined by a rectangular element.

[0090] When the header text element is located in the first row of the table, the program uses a coordinate distance-based method to determine the correspondence. In this layout, data text elements in the same column are most horizontally close to their header text elements. Therefore, the program calculates the coordinate distance between each data text element and all header text elements. This coordinate distance specifically refers to the horizontal coordinate distance, which is the absolute value of the difference between their x-axis coordinates. For each data text element, the program selects the header text element with the closest horizontal coordinate distance and assigns that data text element to that header. By traversing all data text elements, the correspondence between each data element and its column header can be established.

[0091] When the title text element is located in the first column of the table, the program determines the correspondence by analyzing the containment relationship between text elements and graphic elements (i.e., rectangular elements). In this layout, data text elements in the same row are usually located in the same or similarly vertically positioned rectangular cells as the title text element of that row. The program first analyzes the containment relationship between the title text element, data text elements, and rectangular elements in the graphic elements. Specifically, the program determines whether the coordinates of each text element fall within the boundary of a certain rectangular element, thereby determining the cell to which each text element belongs. Once the cell affiliation of the text elements is determined, the correspondence between data text elements and row titles can be established by analyzing the relative positions of the cells. For example, a data text element located in the same row cell as a row title is considered to correspond to that row title.

[0092] In one possible implementation, this embodiment provides a specific method for determining the correspondence between data text elements and title text elements when the title text element is located in the first row of a table. This method is accomplished by calculating coordinate distances and sorting the data. (Refer to...) Figure 5 , Figure 5 This is the fifth flowchart of the multi-dimensional data processing method provided by the present invention. Step 401 specifically includes the following steps: Step 501: Calculate the horizontal coordinate distance between each data text element and each title text element.

[0093] Step 502: Based on the horizontal coordinate distance, assign each data text element to the column corresponding to the nearest title text element.

[0094] Step 503: Sort the data text elements belonging to the same column according to their vertical coordinates to determine the row order.

[0095] When the table's header is located in the first row, its data content is usually visually arranged vertically below the corresponding header. This layout has a clear spatial regularity, i.e. all elements in the same column are substantially aligned in the horizontal direction. If this regularity cannot be effectively utilized, it is difficult to accurately associate a large number of data text elements with the correct column header. Therefore, a method is needed to convert this visual alignment relationship into mathematical calculation.

[0096] The horizontal coordinate distance refers to the absolute value of the coordinate difference of two text elements in the horizontal axis (x-axis), which quantifies the degree of separation of the two elements in the horizontal direction. The vertical coordinate refers to the coordinate value of the text element in the vertical axis (y-axis), which directly indicates the position of the element in the vertical direction. The row order refers to the order of data arranged from top to bottom in the table.

[0097] In specific implementation, first, the program traverses each data text element and calculates the horizontal coordinate distance between the data text element and all header text elements. This calculation process is completed by obtaining the x-coordinate values of the data text element and the header text elements, and then taking the absolute value of the difference. Through this step, the distance relationship between a data text element and all column headers in the horizontal direction can be obtained.

[0098] Next, according to the calculated horizontal coordinate distance, each data text element is attributed to the column corresponding to the nearest header text element. Specifically, for a data text element, the program compares the horizontal coordinate distance between it and all header text elements, and finds the smallest one. The header text element corresponding to the smallest distance is determined as the column header to which this data text element belongs. The program repeats this process for all data text elements, thereby assigning all data text elements to their respective columns.

[0099] Finally, after completing the column assignment, the program sorts the data text elements attributed to the same column according to their vertical coordinates to determine the row order. The program processes each formed column data group, obtains the y-coordinate values of all data text elements in the group, and sorts them in ascending order of y-coordinate values. The sorted result is the correct row order of the column data. By performing this operation on all columns, the row and column structure of the entire table can be completely reconstructed.

[0100] In one possible implementation, the present embodiment provides a specific implementation for determining the correspondence between data text elements and header text elements when the header text elements are located in the first column of the table, which is completed by analyzing the inclusion relationship between text and rectangular elements. Referring to Figure 6 , Figure 6 is a sixth flowchart of the multi-dimensional data processing method provided by the present application, and step 402 specifically comprises the following steps: Step 601: Determine whether the coordinates of each title text element and data text element are within the boundary range defined by the rectangular element, to determine the respective cell position.

[0101] Step 602: According to the cell position, identify the title text element located in the first column as a row title.

[0102] Step 603: Determine the data content of each row by analyzing the data text elements in the same vertical range as the row title.

[0103] When the title of the table is located in the first column, the data is visually spread along the horizontal direction, and the difference in horizontal coordinates of the data elements in the same row can be large, which makes it difficult to apply the method based on coordinate distance. In this layout, the row and column structure of the table is more defined by the physical boundaries of the cells, which are usually represented by rectangular elements in vector image data. Without using the structural information provided by these graphical elements, it is impossible to reliably associate data elements distributed in the same row but far apart in position.

[0104] Cell position refers to the logical position of a text element in the table grid composed of rectangular elements, which is determined by judging whether the coordinates of the text element are contained by a rectangular element. Row title refers to the title text element identified as being located in the first column of the table, which is used to define the data category of the row it is in. Vertical range refers to the vertical spatial area determined by the upper and lower boundaries of a row cell.

[0105] In specific implementation, first, the program determines whether the coordinates of each title text element and data text element are within the boundary range defined by the rectangular element, to determine the respective cell position. The program will iterate through each text element (including title and data) and each rectangular element. For any combination of text and rectangle, the program will perform a double comparison: one is to compare whether the horizontal coordinate of the text element is within the horizontal boundary of the rectangular element (i.e. greater than or equal to the x coordinate of the rectangular element and less than or equal to the sum of its x coordinate and width); the second is to compare whether the vertical coordinate of the text element is within the vertical boundary of the rectangular element (i.e. greater than or equal to the y coordinate of the rectangular element and less than or equal to the sum of its y coordinate and height). Only when both conditions are met, it is determined that the text element is located in the cell defined by the rectangular element. Through this step, each text element is assigned with its cell position information.

[0106] Next, the program identifies the header text elements in the first column as row headers based on their cell positions. The program analyzes the cell positions of all header text elements and identifies those in the leftmost column of the table. These identified header text elements are the row headers, which define the data meaning of each row in the table.

[0107] Finally, the program determines the data content of each row by analyzing the data text elements that are in the same vertical range as the row headers. For each identified row header, the program obtains the vertical range of the cell in which it is located. Then, the program iterates through all data text elements and checks whether they also fall into this vertical range. All data text elements that are in the same vertical range as a row header are considered to belong to the same row. After finding all data text elements in a row, the program sorts them in ascending order of their horizontal coordinates to determine their correct order within the row. This process is repeated for all row headers to completely reconstruct the data content and structure of the table.

[0108] In one possible implementation, the present embodiment provides a specific implementation of determining whether the coordinates of a text element are within the boundary range defined by a rectangular element by performing independent horizontal and vertical boundary comparisons of the coordinates of the text element. Referring to FIG. 7, the process of determining whether the coordinates of a text element are within the boundary range defined by a rectangular element includes the following steps: Figure 7 , Figure 7 FIG. 7 is a flowchart of a process of determining whether the coordinates of a text element are within the boundary range defined by a rectangular element according to an embodiment of the present application. Step 701: Compare the horizontal coordinate of the header text element or data text element with the horizontal boundary of the rectangular element, wherein the horizontal boundary has the horizontal coordinate of the rectangular element as the left boundary and the sum of the horizontal coordinate and the width of the rectangular element as the right boundary.

[0109] Step 702: Compare the vertical coordinate of the header text element or data text element with the vertical boundary of the rectangular element, wherein the vertical boundary has the vertical coordinate of the rectangular element as the upper boundary and the sum of the vertical coordinate and the height of the rectangular element as the lower boundary.

[0110] Step 703: Determine that the coordinates of the header text element or data text element are within the boundary range defined by the rectangular element when the coordinates of the header text element or data text element are within both the horizontal boundary and the vertical boundary.

[0111] The most core problem in the auxiliary identification of table structure by using rectangular elements is how to establish the subordination between the text elements and the rectangular elements, that is, how to accurately determine which cell a text element belongs to. The visual inclusion relationship must be converted into an accurate and calculable logical judgment process. If the judgment process is ambiguous or complex in calculation, it will directly affect the accuracy and efficiency of the entire identification method. Therefore, a simple, reliable and low-cost algorithm is needed to complete this basic operation.

[0112] The horizontal boundary is a range defined by the horizontal coordinate of the rectangular element as the left boundary and the sum of the horizontal coordinate and the width of the rectangular element as the right boundary. The vertical boundary is a range defined by the vertical coordinate of the rectangular element as the upper boundary and the sum of the vertical coordinate and the height of the rectangular element as the lower boundary. The two boundaries jointly define the closed area occupied by the rectangular element in the two-dimensional plane.

[0113] In specific implementation, the program processes the text elements and the rectangular elements that need to be judged. First, the horizontal coordinate of the title text element or the data text element is compared with the horizontal boundary of the rectangular element. The comparison process is specifically to check whether the x-coordinate value of the text element is greater than or equal to the x-coordinate value of the rectangular element, and at the same time, whether the x-coordinate value of the text element is less than or equal to the sum of the x-coordinate value and the width of the rectangular element.

[0114] Then, the vertical coordinate of the title text element or the data text element is compared with the vertical boundary of the rectangular element. Similarly, the comparison process is specifically to check whether the y-coordinate value of the text element is greater than or equal to the y-coordinate value of the rectangular element, and at the same time, whether the y-coordinate value of the text element is less than or equal to the sum of the y-coordinate value and the height of the rectangular element.

[0115] Finally, when the coordinates of the title text element or the data text element are located in the horizontal boundary and the vertical boundary at the same time, it is determined that the coordinates of the title text element and the data text element are located in the boundary range defined by the rectangular element. That is, only when the results of the above horizontal comparison and vertical comparison are both true, the program finally determines that the text element is contained by the rectangular element. If any one of the comparison results is false, it is determined that the text element is not in the boundary range of the rectangular element.

[0116] In a possible implementation, with reference to Figure 8 , Figure 8 is the eighth flowchart of the multi-dimensional data processing method provided by the application, and the method further includes the following steps: Step 801: detecting whether there is a merged cell in the table.

[0117] Step 802: When a merged cell is detected, the number of merged cells is calculated according to the ratio of the distance between the cells where the adjacent header text elements or data text elements are located and the size of a standard cell.

[0118] Step 803: Adjust the table structure according to the number of merged cells.

[0119] In many real-life tables, for the sake of visual aesthetics or information organization, there are cases where multiple row or column cells are merged into one large cell. The previous recognition step is based on a standard grid layout assumption, i.e., each cell occupies an independent row-column position. When a merged cell is encountered, this assumption is broken, which can lead to data misattribution, row-column misplacement, and a large number of blanks or missing values in the final structured data. Therefore, a special mechanism must be introduced to recognize and correctly handle this non-standard layout.

[0120] The distance between the cells where the adjacent header text elements or data text elements are located refers to the distance between the boundaries of the rectangular cells corresponding to two consecutive text elements in the same row or column. The size of a standard cell refers to the width or height of a regular cell that is not merged in the table, which can be determined by statistical analysis of the sizes of all rectangular elements, usually taking the most common width and height as the standard value. The number of merged cells is an integer calculated to represent the number of rows or columns spanned by a merged cell.

[0121] In specific implementation, first, the program detects whether there is a merged cell in the table. This step is performed after the basic row-column structure recognition. The program traverses each row and each column identified, analyzes the distance between the cells corresponding to two adjacent data text elements or header text elements, and determines whether there is a merged cell between the two cells if the distance is significantly greater than the size of a standard cell.

[0122] When a merged cell is detected, the program then calculates the number of merged cells according to the ratio of the distance between the cells where the adjacent header text elements or data text elements are located and the size of a standard cell. Specifically, for row cell merging across columns, the number of merged cells is obtained by calculating the absolute value of the difference between the horizontal coordinates of two adjacent cells and then dividing by the width of a standard cell. For column cell merging across rows, the number of merged cells is obtained by calculating the absolute value of the difference between the vertical coordinates of two adjacent cells and then dividing by the height of a standard cell. The calculation result is rounded down to obtain the exact number of merged cells.

[0123] Finally, the program adjusts the table structure according to the calculated number of merged cells. If a cell is determined to have merged multiple rows or columns, the program iteratively copies the text content of the cell when constructing the final structured data. For example, if a cell containing a text element is calculated to have merged 3 columns, the content of the text element will be filled into the corresponding 3 consecutive column cells in the final two-dimensional data structure. In this way, the originally incomplete grid structure is restored to a logically complete rectangular table, where each cell contains the correct data.

[0124] Referring to Figure 9 , Figure 9 is a structural diagram of a multi-dimensional data processing system provided by the present application, the system comprising: an acquisition module configured to acquire vector image data containing to-be-recognized data; a processing module configured to parse a document structure of the vector image data, and extract text elements and graphic elements according to grouping labels in the document; the processing module is further configured to classify the text elements into title text elements and data text elements according to a preset recognition target object, wherein the recognition target object comprises at least one of a table title or a text title; the processing module is further configured to recognize a row-column structure of table data according to the title text elements, the data text elements and the graphic elements, so as to construct structured data; a storage module configured to store the structured data after format conversion.

[0125] In a possible implementation, the processing module is further configured to: when the recognition target object is the table title, traverse the text elements, and classify text elements whose text content matches a preset table title name as title text elements, and classify the remaining text elements in the text elements as data text elements; when the recognition target object is the text title, traverse the text elements, and identify text elements whose title attribute matches a preset text title name as title text elements.

[0126] In a possible implementation, the processing module is further configured to: determine a corresponding relationship between the data text elements and the title text elements according to a spatial position relationship between the title text elements and the data text elements, wherein the spatial position relationship is determined by analyzing coordinate information of the title text elements and the data text elements and / or relative positions of the title text elements, the data text elements and the graphic elements; reconstruct the row-column structure of the table according to the corresponding relationship and an arrangement rule of the data text elements.

[0127] In a possible implementation, the processing module is further configured to: When the title text element is located in the first row of the table, calculate the coordinate distance between the data text element and the title text element to determine the corresponding relationship; When the title text element is located in the first column of the table, analyze the inclusion relationship between the title text element and the data text element and the rectangular element in the graphic element to determine the corresponding relationship.

[0128] In a possible implementation, the processing module is further configured to: Calculate the horizontal coordinate distance between each data text element and each title text element; According to the horizontal coordinate distance, attribute each data text element to the column corresponding to the title text element closest to the data text element; According to the vertical coordinates of the data text elements attributed to the same column, sort the data text elements to determine the row sequence.

[0129] In a possible implementation, the processing module is further configured to: Determine the cell position to which each title text element and data text element belongs, by judging whether the coordinates of the title text element and data text element are located in the boundary range defined by the rectangular element; According to the cell position, identify the title text element located in the first column as a row title; Determine the data content of each row by analyzing the data text elements in the same vertical range as the row title.

[0130] In a possible implementation, the processing module is further configured to: Compare the horizontal coordinates of the title text element or the data text element with the horizontal boundary of the rectangular element, wherein the horizontal boundary is defined by the horizontal coordinate of the rectangular element as the left boundary and the sum of the horizontal coordinate and the width of the rectangular element as the right boundary; Compare the vertical coordinates of the title text element or the data text element with the vertical boundary of the rectangular element, wherein the vertical boundary is defined by the vertical coordinate of the rectangular element as the upper boundary and the sum of the vertical coordinate and the height of the rectangular element as the lower boundary; When the coordinates of the title text element or the data text element are located in the horizontal boundary and the vertical boundary at the same time, determine that the coordinates of the title text element and the data text element are located in the boundary range defined by the rectangular element.

[0131] In a possible implementation, the processing module is further configured to: Detect whether there is a merged cell in the table; When the merged cell is detected, calculate the number of the merged cell according to the ratio of the distance between the cells in which the adjacent title text elements or data text elements are located to the standard cell size; Adjust the table structure according to the number of merged cells.

[0132] It should be noted that the multi-dimensional data processing system provided by the present application can execute the multi-dimensional data processing method of any of the above embodiments during specific operation, and the present embodiment will not be described here.

[0133] Figure 10 is a structural schematic diagram of an electronic device provided by the present application, as Figure 10 shown, the electronic device can include a processor 1010, a communications interface 1020, a memory 1030, and a communications bus 1040, wherein the processor 1010, the communications interface 1020, and the memory 1030 complete mutual communication through the communications bus 1040. The processor 1010 can invoke the logical instructions in the memory 1030 to execute a multi-dimensional data processing method, which includes: acquiring vector image data containing to-be-recognized data; parsing the document structure of the vector image data, and extracting text elements and graphic elements according to the grouping labels in the document; classifying the text elements into title text elements and data text elements according to a preset recognized target object, wherein the recognized target object includes at least one of a table title or a text title; identifying the row-column structure of table data according to the title text elements, the data text elements, and the graphic elements to construct structured data; and storing the structured data after format conversion.

[0134] In addition, the logical instructions in the memory 1030 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer-readable storage medium. According to such an understanding, the technical solutions of the present application or parts of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiment methods of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0135] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, when the program instructions are executed by a computer, the computer can execute the multi-dimensional data processing method provided by the above embodiments.

[0136] In still another aspect, the present application also provides a non-transitory computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the multi-dimension data processing method provided by the above embodiments.

[0137] The system embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0138] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus necessary universal hardware platforms, and of course can also be implemented by hardware. According to such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the method of each embodiment or some part of the embodiment.

[0139] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.< / rect> < / text> < / g> < / circle> < / rect> < / text> < / g>

Claims

1. A multi-dimensional data processing method, characterized in that, include: Obtain vector image data containing the data to be identified; The document structure of the vector image data is parsed, and text and graphic elements are extracted according to the grouping tags in the document; According to the preset target object for identification, the text elements are classified into title text elements and data text elements, wherein the target object for identification includes at least one of table title or text title; Based on the title text element, the data text element, and the graphic element, identify the row and column structure of the table data to construct structured data; The structured data is converted into a new format and then stored.

2. The multi-dimensional data processing method according to claim 1, characterized in that, The step of classifying the text elements into title text elements and data text elements according to the preset target object includes: When the target object is a table title, the text elements are traversed, and the text elements whose text content matches the preset table title name are classified as title text elements, and the remaining text elements are classified as data text elements. When the target object is a text title, the text elements are traversed, and the text elements whose title attributes match the preset text title name are identified as the title text elements.

3. The multi-dimensional data processing method according to claim 1, characterized in that, The step of identifying the row and column structure of the table data based on the title text element, the data text element, and the graphic element includes: Based on the spatial relationship between the title text element and the data text element, the correspondence between the data text element and the title text element is determined, wherein the spatial relationship is determined by analyzing the coordinate information of the title text element and the data text element and / or the relative positions of the title text element, the data text element and the graphic element; Based on the correspondence and the arrangement rules of the data text elements, the row and column structure of the table is reconstructed.

4. The multi-dimensional data processing method according to claim 3, characterized in that, Determining the correspondence between the data text element and the title text element based on the spatial relationship between the title text element and the data text element includes: When the title text element is located in the first row of the table, calculate the coordinate distance between the data text element and the title text element to determine the correspondence; When the title text element is located in the first column of the table, analyze the inclusion relationship between the title text element, the data text element and the rectangular element in the graphic element to determine the correspondence.

5. The multi-dimensional data processing method according to claim 4, characterized in that, The step of calculating the coordinate distance between the data text element and the title text element to determine the correspondence includes: Calculate the horizontal coordinate distance between each of the data text elements and each of the title text elements; Based on the horizontal coordinate distance, each data text element is assigned to the column corresponding to the nearest title text element; The data text elements belonging to the same column are sorted according to their vertical coordinates to determine the row order.

6. The multi-dimensional data processing method according to claim 4, characterized in that, The step of analyzing the inclusion relationship between the title text element, the data text element, and the rectangular element in the graphic element to determine the correspondence includes: Determine whether the coordinates of each title text element and data text element are within the boundary range defined by the rectangle element to determine their respective cell positions; Based on the cell position, the title text element located in the first column is identified as the row title; The data content of each row is determined by analyzing the data text elements that are within the same vertical range as the row header.

7. The multi-dimensional data processing method according to claim 6, characterized in that, The step of determining whether the coordinates of each title text element and the data text element are within the boundary range defined by the rectangle element includes: The horizontal coordinates of the title text element or the data text element are compared with the horizontal boundary of the rectangle element, wherein the horizontal boundary is defined by the horizontal coordinates of the rectangle element as the left boundary and the sum of the horizontal coordinates and the width of the rectangle element as the right boundary. The vertical coordinates of the title text element or the data text element are compared with the vertical boundary of the rectangle element, wherein the vertical boundary is defined by the vertical coordinates of the rectangle element as the upper boundary and the sum of the vertical coordinates and the height of the rectangle element as the lower boundary. When the coordinates of the title text element or the data text element are simultaneously located within the horizontal and vertical boundaries, it is determined that the coordinates of the title text element and the data text element are within the boundary range defined by the rectangle element.

8. The multi-dimensional data processing method according to claim 1, characterized in that, Also includes: Check if there are merged cells in the table; When merged cells are detected, the number of merged cells is calculated based on the ratio of the distance between adjacent cells containing the title text element or the data text element to the standard cell size. The table structure is adjusted based on the number of merged cells.

9. A multi-dimensional data processing system, characterized in that, include: The acquisition module is used to acquire vector image data containing the data to be identified; The processing module is used to parse the document structure of the vector image data and extract text and graphic elements according to the grouping tags in the document; The processing module is also used to classify the text elements into title text elements and data text elements according to a preset identification target object, wherein the identification target object includes at least one of table title or text title; The processing module is also used to identify the row and column structure of the table data based on the title text element, the data text element, and the graphic element, so as to construct structured data; The storage module is used to convert the structured data into a new format and then store it.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multi-dimensional data processing method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Portable document format (PDF) document form identification method

    CN105589841A

  • Metadata identification method and system based on fragmented document, and storage medium

    CN115168589A

  • Table structure identification method and system and electronic equipment

    CN115601773A

  • Scientific and technical literature table extraction method based on XML

    CN115935910A

  • Chart generation method and device, storage medium and electronic equipment

    CN119152072A

Cited By

  • Form extraction and reconstruction method of PDF document for model training and readable storage medium

    CN121706729A