Data type identification method and device, equipment and medium
By extracting and analyzing the first page of the target file, and utilizing layout analysis and text classification models, the problem of time-consuming OCR in existing technologies has been solved, achieving efficient and accurate file type recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE GROUP SICHUAN
- Filing Date
- 2026-01-06
- Publication Date
- 2026-04-24
AI Technical Summary
Existing file recognition technologies suffer from low efficiency when processing image files, as OCR algorithms are time-consuming and struggle to accurately identify file types.
The first page of the target file is extracted and input into a pre-trained layout analysis model for target region detection. The text in the target region is identified, and the data type is determined using a text classification model.
It improves the efficiency and accuracy of file type identification, reduces the amount of data processed, accurately locates key content areas, and reduces resource consumption.
Smart Images

Figure CN121921798A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a data type identification method, apparatus, device, and medium. Background Technology
[0002] Enterprise management data is a crucial type of data for basic telecommunications enterprises. Existing document recognition technologies involve using different methods to extract all text content from files of different formats. Due to the diversity of file formats, such as jpg / png, docx / doc, pdf, etc., different methods are needed to extract text content from different formats. Determining the file type based on the extracted text content includes: keyword matching: matching the extracted text content with specified keywords; if a match is successful, the file is determined to be a file of a certain type under that rule; and named entity recognition (NAME) matching: inputting the extracted text content into a NAME model; if one or more corresponding entities are successfully identified, the file is determined to be a file of a certain type.
[0003] However, existing solutions for image files require OCR algorithms to extract all text, a task that is inherently time-consuming due to the nature of OCR algorithms. Current techniques match the extracted text content with specified rules to obtain the recognition result. However, due to the diversity and complexity of file formats, relying solely on text content rules is insufficient for accurate file recognition. Summary of the Invention
[0004] This application provides a data type identification method, apparatus, device, and medium to specifically detect and identify areas that reflect key file data types, thereby improving the efficiency of content identification and the accuracy of subsequent data type identification.
[0005] According to one aspect of this application, a data type identification method is provided, the method comprising:
[0006] For the target file, extract the first page of the target file; wherein, the first page is the first page in the target file where the area containing content occupies a proportion of the single page area that exceeds a preset proportion;
[0007] The homepage is input into a pre-trained layout analysis model for target region detection to determine the target regions in the homepage; wherein, the target regions include text regions that summarize the entire content of the target file; the homepage input into the layout analysis model is in image format;
[0008] Text recognition is performed on the target region to obtain the target text in the target region, and the data type of the data in the target file is determined based on the target text.
[0009] According to one aspect of this application, a data type identification device is provided, the device comprising:
[0010] The homepage extraction module is used to extract the homepage of a target file; wherein, the homepage is the first page in the target file whose content area occupies a proportion exceeding a preset ratio.
[0011] The target region detection module is used to input the homepage into a pre-trained layout analysis model for target region detection, and to determine the target regions in the homepage; wherein, the target region includes a text region that summarizes the entire content of the target file; the homepage input into the layout analysis model is in image format;
[0012] The data type recognition module is used to perform text recognition on the target area to obtain the target text in the target area, and to determine the data type of the data in the target file based on the target text.
[0013] According to another aspect of this application, an electronic device is provided, the electronic device comprising:
[0014] At least one processor; and,
[0015] A memory that is communicatively connected to at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the data type identification method of any embodiment of this application.
[0017] According to another aspect of this application, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the data type identification method of any embodiment of this application.
[0018] According to another aspect of this application, a computer program product is provided, which includes a computer program that, when executed by a processor, implements the data type identification method of any embodiment of this application.
[0019] The technical solution of this application embodiment extracts the first page of the target file. The first page is the first page in the target file where the proportion of content-containing area exceeds a preset ratio. This targeted extraction of the first page, which reflects the data type of the data in the target file, avoids unnecessary resource consumption caused by processing all content in the target file, thus improving processing efficiency. The first page is input into a pre-trained layout analysis model for target area detection to determine the target areas within the first page. The target areas include text areas summarizing the entire content of the target file. The first page input into the layout analysis model is in image format. This achieves accurate detection of target areas containing content on the first page, further leveraging the advantages of the first page and accurately identifying key content areas. Text recognition is performed on the target areas to obtain the target text within them, and the data type of the data in the target file is determined based on the target text. By selectively identifying and classifying the target text in the target areas, classification accuracy is improved by recognizing the text that best reflects the data type. Furthermore, by recognizing only the target text in the target areas, the amount of data processed is effectively reduced, improving processing efficiency.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a data type identification method provided in this application embodiment;
[0023] Figure 2 A flowchart illustrating a data type identification method provided in another embodiment of this application;
[0024] Figure 3 A flowchart illustrating a data type identification method provided in another embodiment of this application;
[0025] Figure 4 This is a schematic diagram of the text classification model structure provided in another embodiment of this application;
[0026] Figure 5A flowchart illustrating a specific implementation method provided in this application embodiment;
[0027] Figure 6 This is a schematic diagram of the structure of a data type identification device provided in an embodiment of this application;
[0028] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0030] It should be noted that the terms "first," "second," "third," "fourth," "actual," "preset," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] The acquisition, storage, use, and processing of data in this application comply with relevant national laws and regulations. The acquired data is obtained with authorization and will not be disclosed without permission, used for illegal purposes, purposes detrimental to the interests of others, or for personalized analysis or product promotion. It should be noted that certain software, components, models, and other existing industry solutions may be mentioned in the embodiments of this application. These should be considered exemplary and intended only to illustrate the feasibility of implementing the technical solution of this application, but do not imply that the applicant has already used or necessarily used the relevant content of such solutions.
[0032] Figure 1This flowchart illustrates a data type identification method provided in an embodiment of this application. This embodiment is applicable to identifying data types in a target file to distinguish between different types of enterprise management data and non-enterprise management data. The method can be executed by a data type identification device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0033] S110. For the target file, extract the first page of the target file; wherein, the first page is the first page in the target file whose content area occupies a proportion exceeding a preset proportion.
[0034] The target file can be any document that requires identification and classification. Identifying the content of the target file determines whether the data belongs to enterprise management data or non-enterprise management data, and if it belongs to enterprise management data, what type of enterprise management data it is, such as meeting minutes, technological achievements, patent disclosure documents, financial statements, market analysis reports, etc. The homepage refers to the first page of the target file whose content area occupies a proportion exceeding a preset percentage. In this embodiment, the homepage is not simply the first page in a formal sense, but the first page containing substantive content. Whether it contains substantive content is determined by whether the content-containing area occupies a proportion exceeding a preset percentage. The preset percentage can be determined based on actual conditions. In this embodiment, it can be determined based on the proportion of the title area and / or the abstract area on a page under normal circumstances. That is, the preset percentage is set equal to the proportion of the title area, or the abstract area, or the combined proportion of the title and abstract on a page under normal circumstances. There may be a blank page in the target file, or a single page containing a simple identifier but no substantial content. In this case, the homepage is not the homepage that needs to be extracted in this application embodiment. Instead, the first single page in the target file that contains substantial content is taken as the homepage.
[0035] In this embodiment of the application, the target file can be detected starting from the first page, and the proportion of the area containing the content to the area of the single page can be detected. If the proportion of the currently detected single page does not exceed the preset proportion, the next single page can be detected. If the proportion of the currently detected single page exceeds the preset proportion, the currently detected single page can be taken as the first page and the first page of the target file can be extracted.
[0036] Since the target file can be in various formats, such as image, pdf, docx, wps, etc., in order to unify subsequent steps, the first page in image format can be retained, and the first pages in other formats can be uniformly converted to image format, which will facilitate unified and rapid processing based on a unified model.
[0037] S120. Input the homepage into a pre-trained layout analysis model to detect target areas and determine target areas in the homepage; wherein, the target area includes a text area that summarizes the entire content of the target file; the homepage input into the layout analysis model is in image format.
[0038] The target area is a region containing text summarizing the entire content of the target file, distinct from the entire content of the first page of the target file. It is a region containing summary text, including a title area and / or a summary area. A homepage may include both a title and a table of contents, where the title is text summarizing the entire content and the table of contents is text indexing specific sections. Therefore, the target area should be the region containing the title. Similarly, a homepage may include both a summary and a table of contents, where the summary is text summarizing the entire content. Therefore, the target area should be the region containing the summary. As described in the above embodiments, the homepage can be uniformly converted to image format; therefore, the homepage processed in this step is in image format.
[0039] In this embodiment, a pre-trained layout analysis model can be loaded. This model has the ability to detect and determine target regions contained in an image and output the location of these target regions. The extracted homepage can be input into the pre-trained layout analysis model for target region detection, determining the target regions within the homepage. This involves outputting the vertex coordinates of the target regions contained in the homepage, thus clearly identifying the areas containing the summary text of all content on the homepage.
[0040] S130. Perform text recognition on the target region to obtain the target text in the target region, and determine the data type of the data in the target file based on the target text.
[0041] Since the target region contains text summarizing the entire content, text recognition can be performed on the target region to determine the target text. Text recognition of the target region can be based on text recognition algorithms, such as SVTR algorithms or large language model-based recognition. After identifying the target text in the target region, the data type of the data in the target file is determined based on the target text. For example, this can be done by matching keywords corresponding to different data types in the target text to determine the data type corresponding to the keywords that match the target text, thereby determining the data type of the data in the target text.
[0042] In this embodiment of the application, determining the data type of the data in the target file based on the target text includes:
[0043] The target text is input into a pre-trained text classification model, and the text classification model outputs the data type of the data in the target file reflected by the target text; wherein, the data type includes multiple types of enterprise management data and non-enterprise management data.
[0044] For example, a text classification model can be pre-trained. This model has the ability to identify and determine whether the input text belongs to enterprise management data or non-enterprise management data, and if it belongs to enterprise management data, to which specific type of enterprise management data it belongs to. In practical applications, the text classification model can be loaded, the target text can be input into the text classification model, and the text classification model can output the data type of the target file reflected by the target text. By pre-training the text classification model and classifying the target text based on the text classification model, the powerful learning ability of the text classification model can be used to learn the relationship between data classification and the input target text, thereby improving the accuracy of data classification.
[0045] The technical solution of this application embodiment extracts the first page of the target file. The first page is the first page in the target file where the proportion of content-containing area exceeds a preset ratio. This targeted extraction of the first page, which reflects the data type of the data in the target file, avoids unnecessary resource consumption caused by processing all content in the target file, thus improving processing efficiency. The first page is input into a pre-trained layout analysis model for target area detection to determine the target areas within the first page. The target areas include text areas summarizing the entire content of the target file. The first page input into the layout analysis model is in image format. This achieves accurate detection of target areas containing content on the first page, further leveraging the advantages of the first page and accurately identifying key content areas. Text recognition is performed on the target areas to obtain the target text within them, and the data type of the data in the target file is determined based on the target text. By selectively identifying and classifying the target text in the target areas, classification accuracy is improved by recognizing the text that best reflects the data type. Furthermore, by recognizing only the target text in the target areas, the amount of data processed is effectively reduced, improving processing efficiency.
[0046] Figure 2 This is a flowchart illustrating a data type identification method according to another embodiment of this application. This embodiment is an optimization based on the above embodiment; solutions not described in detail in this embodiment are found in the above embodiment. Figure 2 As shown, the method in this embodiment of the application specifically includes the following steps:
[0047] S210. For the target file, extract the first page of the target file; wherein, the first page is the first page in the target file whose content area occupies a proportion exceeding a preset proportion.
[0048] S220. Input the homepage into a pre-trained layout analysis model to detect target areas and determine target areas in the homepage; wherein, the target area includes a text area that summarizes the entire content of the target file; the homepage input into the layout analysis model is in image format.
[0049] S230. Perform text recognition on the target region to obtain the target text in the target region.
[0050] S240. Determine whether the target text matches the identification rules for each type of enterprise management data. If the match is successful, proceed to S250; otherwise, proceed to S260.
[0051] For example, after obtaining the target text through text recognition of the target area, the target text can be matched with recognition rules for various types of enterprise management data. For instance, different recognition rules exist for different types of enterprise management data. For example, for type A enterprise management data, the target text needs to match keywords a1, a2, a3, and a4, thus determining that the data corresponding to the target text is type A enterprise management data. For type B enterprise management data, the target text needs to match keywords b1, b2, and b3, thus determining that the data corresponding to the target text is type B enterprise management data. If the target text fails to match any of the recognition rules for different types of enterprise management data, then the data corresponding to the target text is determined to be non-enterprise management data. Determining the data type of the data in the target file based on matching with recognition rules for different types of enterprise management data allows for faster and more convenient data type determination based on text matching, improving recognition efficiency.
[0052] S250. Determine that the data in the target text is enterprise management data, and determine the data type of the enterprise management data based on the matching results.
[0053] In this embodiment of the application, if the matching result of the target text with the identification rules of various types of enterprise management data is successful, that is, the target text can satisfy one of the identification rules of enterprise management data, then the data type of the data in the target file corresponding to the target text is determined to be enterprise management data of that type.
[0054] S260. Input the target text into a pre-trained text classification model, and output the data type of the target text reflected in the target file through the text classification model.
[0055] For example, if the target text does not match any of the recognition rules corresponding to all types of enterprise management data, the target text is input into a pre-trained text classification model. Since the text classification model has powerful learning and recognition capabilities, it can more accurately identify target text that cannot be identified based on simple rules, thereby improving recognition accuracy. For the target text that does not match in the above steps, the data type of the data in the target file is accurately determined by classifying the target text based on the text classification model.
[0056] The solution in this embodiment matches the target text with various enterprise management data identification rules. If the match is successful, the data in the target text is determined to be enterprise management data, and the data type of the enterprise management data is determined based on the matching result. Otherwise, the target text is input into a pre-trained text classification model, which outputs the data type of the target file reflected in the target text. This solution achieves efficient classification of enterprise management data through a multi-level data identification mechanism: first, the target text is matched with a preset rule base, and the rule engine is used to quickly and accurately identify structured data and determine its type; if the rule matching fails, the pre-trained text classification model is called for intelligent analysis, and unstructured text is processed through deep learning technology to ensure coverage of a wider range of data scenarios. This hybrid strategy retains the efficiency and interpretability of rule matching while improving the generalization ability of complex text with the help of the model, effectively solving the limitations of traditional methods under data diversity, and ultimately achieving comprehensive and accurate identification of enterprise management data types, providing a reliable foundation for subsequent data governance and analysis.
[0057] Figure 3 This is a flowchart illustrating a data type identification method according to another embodiment of this application. This embodiment is an optimization based on the above embodiments; solutions not described in detail in this embodiment are found in the above embodiments. Figure 3 As shown, the method in this embodiment of the application specifically includes the following steps:
[0058] S310. For the sample file, extract the first page of the sample file and set the first page to image format to obtain the sample first page image.
[0059] In this embodiment, during the pre-training of the layout analysis model, sample files of various business management data types and non-business management data types can be pre-acquired. The first page of the sample file is extracted; if it is an image, it is retained; otherwise, it is converted to an image. The process of extracting the first page from the sample file is similar to that of extracting the first page from the target file. The meaning of the first page is consistent with that in the above embodiments; it is the first page in the sample file whose content area accounts for a proportion exceeding a preset ratio within the single-page area. The first page of the image-format sample file obtained through the above scheme is used as the sample first page image, which is the training sample for training the layout analysis model.
[0060] S320. The target area in the sample homepage image is labeled to form a labeled layout analysis training sample.
[0061] For example, target areas in the sample homepage image can be specifically labeled, such as using bounding boxes to label the target areas in the sample homepage image, forming labeled layout analysis training samples. The labeled sample homepage image serves as the layout analysis training sample. In the layout analysis training sample, the target area includes keywords corresponding to the data type of the data in the sample file. That is, the text in the target area summarizing the entire content of the sample file contains keywords corresponding to the data type of the data in the sample file, reflecting the data type of the data in the sample file.
[0062] S330. The neural network model is trained based on the layout analysis training samples to obtain the layout analysis model.
[0063] In this embodiment, a neural network model is trained based on layout analysis training samples. The target region predicted by the neural network model from the sample homepage image in the layout analysis training samples is compared with the actual labeled target region. A loss function is constructed based on the deviation between the predicted target region and the actual target region, and the neural network model is optimized until the iteration stopping condition is met, thus obtaining the layout analysis model. The neural network can be a neural network constructed using a layout analysis algorithm. PicoDet is a high-speed, high-precision object detection algorithm. PicoDet achieves a small, fast, and accurate model through the following optimization strategies: a higher-performance backbone network: PicoDet uses the ESNet model as the backbone network, making the entire object detection model not only less computationally intensive, less latency, and more accurate, but also more robust; a lighter Neck and Head: In the Neck part, PicoDet proposes a CSP-PAN structure, using 1 The convolution with a value of 1 unifies the number of channels in the features with the minimum number of channels in the Backbone output, thereby reducing computational cost and ensuring that feature fusion performance is not affected. Furthermore, PicoDet downsamples again on top of CSP-PAN, adding a smaller feature scale to improve the detection performance of large objects. The trained layout analysis model has the ability to detect and identify the location of target regions on the first page image of a target document.
[0064] S340. For the target file, extract the first page of the target file; wherein, the first page is the first page in the target file whose content area occupies a proportion exceeding a preset proportion.
[0065] S350. Input the homepage into a pre-trained layout analysis model to detect target regions and determine target regions in the homepage; wherein, the target region includes a text region that summarizes the entire content of the target file; the homepage input into the layout analysis model is in image format.
[0066] S360. Perform text recognition on the target region to obtain the target text in the target region, and determine the data type of the data in the target file based on the target text.
[0067] In this embodiment, for a sample file, the first page of the sample file is extracted and set as an image to obtain a sample first page image. Target areas in the sample first page image are labeled to form labeled layout analysis training samples. A neural network model is trained based on these training samples to obtain a layout analysis model. This scheme achieves efficient and accurate document layout analysis by constructing a layout analysis training sample library and training a neural network model: First, the first page of the sample file is converted to an image format to ensure that the original layout information is preserved; second, target areas are labeled to form structured training data, providing the model with clear feature learning targets; finally, the neural network model is trained based on the labeled data, enabling it to automatically identify text, tables, charts, and other elements in the document and their layout relationships. This process significantly improves the accuracy and robustness of layout analysis, avoids the insufficient adaptability of traditional rule-based methods to complex layouts, and supports general processing of multiple document types, providing a reliable technical foundation for document digitization, information extraction, and other applications.
[0068] In this embodiment of the application, the process of constructing the text classification model includes:
[0069] For sample files of different data types, determine the summary text corresponding to all content in the sample files;
[0070] The generalized text is labeled with data types to determine text classification samples;
[0071] The neural network model is trained based on the text classification samples to construct a text classification model.
[0072] In this embodiment, combined with specific application scenarios, a text classification model can be trained to identify whether text belongs to enterprise management data or non-enterprise management data, and what type of enterprise management data it belongs to. Specifically, for sample files of different data types, a summary text corresponding to all content in the sample file can be determined, such as determining the title of the sample file, or determining the summary of the sample file, or determining both the title and summary of the sample file. The determined summary text of the sample file is labeled with data type to determine the text classification sample. The neural network model is trained based on the text classification sample. The predicted data type obtained by the neural network model in recognizing the summary text of the sample file is compared with the actual data type to determine the loss function. The neural network model is trained and optimized based on the loss function until the optimization condition is reached, resulting in a trained text classification model. The text classification model has the ability to identify the corresponding data type of the input text, including outputting a specific enterprise management data type or a non-enterprise management data type. The initial neural network of the text classification model can be constructed using the BERT algorithm. The principle diagram of the text classification model based on BERT is shown below. Figure 4 As shown, the input target text is segmented by character, features are extracted from each segmented character, and the text classification model outputs the corresponding classification and each token. Figure 4 As shown, assuming the target text is the title "Equipment Procurement Contract", it is segmented into seven independent characters. Features are extracted from each independent character to obtain corresponding features. Based on the text classification model, the features are identified to obtain the corresponding category "Contract" and the corresponding tokens are output.
[0073] In this embodiment of the application, for sample files of different data types, determining the summary text corresponding to all content in the sample file includes:
[0074] For a sample file, determine the target area marked in the sample file, or identify the target area based on the layout analysis model of the sample file;
[0075] Text recognition is performed on the target region to determine the summary text corresponding to all content in the sample file.
[0076] For example, during the training of a text classification model, the determination of text classification samples can be based on sample files. Specifically, the target regions marked in the sample files can be identified, or the target regions can be determined by recognizing the sample files using a pre-trained layout analysis model. Text recognition of the target regions determines the summary text corresponding to all content in the sample file. In other words, the summary text is determined based on the sample files used in the layout analysis model training. This allows both the layout analysis model and the text classification model to be trained on the same file samples, making both models suitable for processing the same target file and improving accuracy.
[0077] This application provides a specific implementation method, such as... Figure 5 As shown, specifically:
[0078] Step 1: Prepare relevant training data
[0079] The training data includes data used to train the layout analysis model and data used to train the text classification model.
[0080] For training the layout analysis model, the training data comes from the management data of basic telecommunications enterprises, which can be in formats such as PDF, DOCX, and WPS, or image formats. For non-image format files, the pages containing file structure information (usually the first page of the file) are selected and converted into images. For image format files, images of the first page of the file are directly selected.
[0081] For training text classification models, the training data also comes from the management data of basic telecommunications enterprises, using the title content of these documents as training data.
[0082] Step 2: Label the training data
[0083] For layout analysis model training data, data annotation tools are needed to annotate the title areas in the enterprise management data images; for text classification model training data, corresponding labels are directly added to each title data to indicate the category of enterprise management data.
[0084] Step 3: Training the model
[0085] The labeled layout analysis data was trained using the PicoDet algorithm; the title data was trained using the BERT algorithm.
[0086] PicoDet is a high-speed, high-accuracy object detection algorithm. PicoDet achieves this by employing several optimization strategies, resulting in a smaller, faster, and more accurate model: a higher-performance backbone network: PicoDet uses the ESNet model as its backbone network, making the entire object detection model not only less computationally expensive, less latency, and more accurate, but also more robust; a lighter Neck and Head: In the Neck part, PicoDet proposes a CSP-PAN structure, using 1... The convolution with a value of 1 unifies the number of channels in the features with the minimum number of channels in the Backbone output, thereby reducing computational cost and ensuring that feature fusion performance is not affected. Furthermore, PicoDet downsamples again on top of CSP-PAN, adding a smaller feature scale to improve the detection performance of large objects.
[0087] Input the labeled file image data into PicoDet to train the layout analysis model. After the model is trained, it is used to identify the title in the file image to be identified and output the coordinates of the title area.
[0088] Use BERT-trained labeled title text: such as Figure 4 As shown, the model input is the file title segmented by word; the model training involves feature extraction from the file title; and the model output includes the file title category and a token.
[0089] Step 4: Identify key areas using layout analysis models
[0090] The trained layout analysis model is used to detect the title region in the file to be identified.
[0091] Step 5: Use an optical character recognition (OCR) algorithm to identify text in key areas.
[0092] The text content in the title area was identified using a Chinese character recognition algorithm (SVTR).
[0093] The title image region detected by the layout analysis algorithm is input into the text recognition model (SVTR), and the model outputs the corresponding title text in the title image.
[0094] Step Six: Match the identified title content with the pre-defined enterprise management data identification rules. If a certain type of enterprise management data identification rule is met, the identification result is directly output. If none of the enterprise management data identification rules are met, the enterprise management data text classification model is used to classify and predict the title content, and the enterprise management data category is output.
[0095] Figure 6This is a schematic diagram of a data type identification device provided in an embodiment of this application. This device can execute the data type identification method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the method. Figure 6 As shown, the device includes:
[0096] The homepage extraction module 410 is used to extract the homepage of the target file; wherein, the homepage is the first page in the target file whose content area occupies a proportion of the single page area exceeding a preset proportion;
[0097] The target region detection module 420 is used to input the homepage into a pre-trained layout analysis model for target region detection to determine the target region in the homepage; wherein, the target region includes a text region that summarizes the entire content of the target file; the homepage input into the layout analysis model is in image format;
[0098] The data type recognition module 430 is used to perform text recognition on the target area to obtain the target text in the target area, and to determine the data type of the data in the target file based on the target text.
[0099] In this embodiment of the application, the homepage extraction module 410 extracts the homepage of the target file; wherein, the homepage is the first page in the target file whose content area occupies a proportion of the single page area exceeding a preset proportion;
[0100] The homepage is input into a pre-trained layout analysis model for target region detection to determine the target regions in the homepage; wherein, the target regions include text regions that summarize the entire content of the target file; the homepage input into the layout analysis model is in image format;
[0101] Text recognition is performed on the target region to obtain the target text in the target region, and the data type of the data in the target file is determined based on the target text.
[0102] In this embodiment of the application, the data type identification module 430 determines the data type of the data in the target file based on the target text, including:
[0103] The target text is input into a pre-trained text classification model, and the text classification model outputs the data type of the data in the target file reflected by the target text; wherein, the data type includes multiple types of enterprise management data and non-enterprise management data.
[0104] In this embodiment of the application, the data type identification module 430 determines the data type of the data in the target file based on the target text, including:
[0105] The target text is matched with various enterprise management data identification rules.
[0106] If the match is successful, the data in the target text is determined to be enterprise management data, and the data type of the enterprise management data is determined according to the matching result;
[0107] Otherwise, the target text is input into a pre-trained text classification model, and the text classification model outputs the data type of the target file reflected by the target text.
[0108] In this embodiment of the application, the apparatus further includes a training module, used for:
[0109] For the sample file, extract the first page of the sample file and set the first page to image format to obtain the sample first page image;
[0110] The target areas in the sample homepage image are labeled to form a labeled layout analysis training sample;
[0111] The neural network model is trained based on the layout analysis training samples to obtain the layout analysis model.
[0112] In this embodiment of the application, the training module is further configured to:
[0113] For sample files of different data types, determine the summary text corresponding to all content in the sample files;
[0114] The generalized text is labeled with data types to determine text classification samples;
[0115] The neural network model is trained based on the text classification samples to construct a text classification model.
[0116] In this embodiment of the application, the training module is further configured to:
[0117] For a sample file, determine the target area marked in the sample file, or identify the target area based on the layout analysis model of the sample file;
[0118] Text recognition is performed on the target region to determine the summary text corresponding to all content in the sample file.
[0119] In this embodiment of the application, the target area includes a title area or a summary area.
[0120] The data type identification device provided in this application embodiment can execute a data type identification method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the method execution.
[0121] Figure 7 A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0122] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0123] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0124] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data type recognition methods.
[0125] In some embodiments, the data type identification method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data type identification method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data type identification method by any other suitable means (e.g., by means of firmware).
[0126] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0127] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data type identification device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0128] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0130] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0131] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0132] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the data type identification method provided in any embodiment of this application.
[0133] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0134] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired information of the technical solution of this application can be achieved, and this is not limited herein.
[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A data type identification method, characterized in that, The method includes: For the target file, extract the first page of the target file; wherein, the first page is the first page in the target file where the area containing content occupies a proportion of the single page area that exceeds a preset proportion; The homepage is input into a pre-trained layout analysis model for target region detection to determine the target regions in the homepage; wherein, the target regions include text regions that summarize the entire content of the target file; the homepage input into the layout analysis model is in image format; Text recognition is performed on the target region to obtain the target text in the target region, and the data type of the data in the target file is determined based on the target text.
2. The method according to claim 1, characterized in that, Determining the data type of the data in the target file based on the target text includes: The target text is input into a pre-trained text classification model, and the text classification model outputs the data type of the data in the target file reflected by the target text; wherein, the data type includes multiple types of enterprise management data and non-enterprise management data.
3. The method according to claim 1, characterized in that, Determining the data type of the data in the target file based on the target text includes: The target text is matched with various enterprise management data identification rules. If the match is successful, the data in the target text is determined to be enterprise management data, and the data type of the enterprise management data is determined according to the matching result; Otherwise, the target text is input into a pre-trained text classification model, and the text classification model outputs the data type of the target file reflected by the target text.
4. The method according to claim 1, characterized in that, The process of constructing the layout analysis model includes: For the sample file, extract the first page of the sample file and set the first page to image format to obtain the sample first page image; The target areas in the sample homepage image are labeled to form a labeled layout analysis training sample; The neural network model is trained based on the layout analysis training samples to obtain the layout analysis model.
5. The method according to any one of claims 2-4, characterized in that, The process of constructing the text classification model includes: For sample files of different data types, determine the summary text corresponding to all content in the sample files; The generalized text is labeled with data types to determine text classification samples; The neural network model is trained based on the text classification samples to construct a text classification model.
6. The method according to claim 5, characterized in that, For sample files of different data types, determine the summary text corresponding to all content in the sample files, including: For a sample file, determine the target area marked in the sample file, or identify the target area based on the layout analysis model of the sample file; Text recognition is performed on the target region to determine the summary text corresponding to all content in the sample file.
7. The method according to claim 1, characterized in that, The target area includes the title area and / or the abstract area.
8. A data type identification device, characterized in that, The data type identification device includes: The homepage extraction module is used to extract the homepage of a target file; wherein, the homepage is the first page in the target file whose content area occupies a proportion exceeding a preset ratio. The target region detection module is used to input the homepage into a pre-trained layout analysis model for target region detection, and to determine the target regions in the homepage; wherein, the target region includes a text region that summarizes the entire content of the target file; the homepage input into the layout analysis model is in image format; The data type recognition module is used to perform text recognition on the target area to obtain the target text in the target area, and to determine the data type of the data in the target file based on the target text.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data type identification method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the data type identification method according to any one of claims 1-7.