Artificial intelligence text analysis and extraction and key point source positioning method
Through a deep learning model based on Transformer architecture, PDF documents are layout recognition and text refinement, combined with advanced language models and vector space models, the problems of missing and incomplete information in text parsing and refining are solved, and efficient and accurate information processing and retrieval are achieved.
Patent Information
- Application Number
- CN202510661807.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology has problems of missing important information and incomplete refining in text analysis and refining, and the large language model fails to combine with the original text when improving reading efficiency, resulting in poor reading results.
The deep learning model based on the Transformer architecture is used to recognize and refine PDF documents in layout and text, combine advanced language models for deep vectorization, and realize semantic retrieval through vector space model and similarity measurement algorithm.
It realizes rapid identification and conversion of PDF document content, improves information processing efficiency and accuracy, enhances the versatility and adaptability of text parsing and refining, overcomes the shortcomings of traditional methods, and effectively explores and utilizes PDF document information.
Smart Images

Figure CN120180252A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing, and particularly to an artificial intelligence text parsing, refining and key point source location method. Background Art
[0002] Currently, the technology has long been able to parse or refine text documents in formats such as pdf. Usually, natural language processing technology is used to parse and identify the more structured content in the text to achieve refinement. However, there are the following problems: 1) The refined content is too dependent on the structure of the original text, resulting in the loss of important information and incomplete refined parsing results; 2) Using large language model technology can make up for the problem of incomplete refinement, but in the scenario of improving the reading efficiency of academic literature, it fails to combine with the original text, resulting in the inability to truly improve the reading effect. Therefore, there is an urgent need for an artificial intelligence text parsing, refining and key point source location method. Summary of the Invention
[0003] The purpose of the present invention is to provide an artificial intelligence text parsing, refining and key point source location method, aiming to solve the above problems.
[0004] The present invention provides an artificial intelligence text parsing, refining and key point source location method, including: Collect PDF document data, perform manual annotation and preprocessing on the PDF document data to obtain a training set, a validation set and a test set; Based on the Transformer architecture and the hardware environment adapted for deep learning training, initialize the deep learning model and set the parameters for deep learning model training; Train the deep learning model based on the training set, and validate and test the trained deep learning model according to the validation set and the test set. If the validation and test pass, determine the deep learning model that passes the validation and test as the PDF document layout recognition model; Use the PDF document layout recognition model to recognize the PDF document to be recognized, and convert the recognition result into a text format or a chart / table format; Introduce an advanced language model based on the Transformer architecture to perform deep vectorization processing on the text converted into a text format or a chart / table format; Based on the advanced vector space model and similarity measurement algorithm, realize semantic retrieval.
[0005] Preferably, when collecting PDF document data, the source of the PDF document data is an academic paper library, an enterprise report library, an open e-book library, and government and public institution documents; The data collection strategy includes: priority of diversity and sufficient data volume.
[0006] Preferably, perform manual annotation and preprocessing on the PDF document data, including: According to a predefined tag system, manually annotate the title, text, formulas, tables, images, and captions in the PDF document data to obtain annotated data, and store the annotated data in a standardized format; Perform preprocessing on the annotated data, where the preprocessing includes duplicate removal, normalization, image resolution standardization, and sentence vector processing.
[0007] Preferably, after performing manual annotation and preprocessing on the PDF document data, perform data quality assessment on the preprocessed annotated data. The evaluation metrics used include: data diversity assessment, data integrity assessment, data accuracy assessment, data consistency assessment, and data usability assessment.
[0008] Preferably, initialize the deep learning model, including: Adjust the number of attention heads, optimize the embedding layer dimension, add a convolutional layer to fuse local features, design a multi-scale feature pyramid, initialize the model weights, and encode the spatial perception position.
[0009] Preferably, when setting the parameters for deep learning model training, the parameters include: learning rate, batch size, number of training epochs, optimizer and regularization, loss function, key parameter linkage rules, hardware resource adaptation rules, and parameter tuning verification methods; Select the Adam optimizer for the optimizer, and select the cross-entropy loss function for the loss function.
[0010] Preferably, when training the deep learning model based on the training set, calculate the loss value and update the model weights through the backpropagation algorithm to gradually reduce the loss value.
[0011] Preferably, convert the recognition result into a text format or a chart / table format, including: If the recognition result is text, use syntactic analysis and semantic understanding techniques in natural language processing to convert the recognition result into a clear and editable text format; If the recognition result is a chart or table, use an innovative algorithm based on image recognition and data mining to achieve accurate extraction and text presentation of the data.
[0012] Preferably, introduce an advanced language model based on the Transformer architecture to perform deep vectorization processing on the text converted into a text format or a chart / table format, including: The advanced language model based on the Transformer architecture includes BERT or ERNIE; Perform word segmentation on the text converted into text format or chart / table format, and input the segmented text into an advanced language model based on the Transformer architecture. Encode the text through the multi-layer Transformer encoder of the model to obtain the vector representation of each word or sub-word unit.
[0013] Preferably, based on the advanced vector space model and similarity measurement algorithm, semantic retrieval is realized, including: Construct a vector space model and store all vectorized texts in a high-dimensional vector space; When the user inputs a retrieval requirement, perform the same vectorization operation on the retrieval requirement as the text processing to obtain a retrieval vector; Use cosine similarity calculation to calculate the similarity score between the retrieval vector and all stored vectors in the vector space; Sort the words or sub-word units according to the similarity score, and return the document fragment to which the word or sub-word unit most relevant to the retrieval requirement belongs to the user.
[0014] Compared with the prior art, the beneficial effects of the present invention are that the present invention can quickly identify the content of a PDF document and convert it into text or chart format. Introduce advanced language model for in-depth vectorization processing, and combine vector space and similarity measurement algorithm to realize semantic retrieval, which can improve the information processing efficiency, enhance the accuracy and integrity, strengthen the generality and adaptability, and be widely applied in multiple fields to promote their digital and intelligent development, overcome the deficiencies of manual work and existing technologies, and effectively explore and utilize the information in PDF documents. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0016] Figure 1 It is a schematic flowchart of a method for artificial intelligence text parsing, extraction and key point source location of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0018] As Figure 1 shown, the present invention provides an artificial intelligence text parsing, refining and key point source location method, including: Collect PDF document data, perform manual annotation and preprocessing on the PDF document data to obtain a training set, a validation set and a test set, with a ratio of 7:2:1; Based on the Transformer architecture and the hardware environment adapted for deep learning training, initialize the deep learning model and set the parameters for deep learning model training; Train the deep learning model based on the training set, and verify and test the trained deep learning model according to the validation set and the test set. If the verification and test are passed, determine the deep learning model that passes the verification and test as the PDF document layout recognition model; Use the PDF document layout recognition model to recognize the PDF document to be recognized, and convert the recognition result into a text format or a chart / table format; Introduce an advanced language model based on the Transformer architecture to perform deep vectorization processing on the text converted into a text format or a chart / table format; Based on the advanced vector space model and similarity measurement algorithm, realize semantic retrieval.
[0019] The present invention can significantly improve the accuracy and efficiency of text parsing and refining, and at the same time quickly locate the key point source. By accurately recognizing the PDF document layout, the integrity of information extraction is ensured, and the tediousness and errors of manual annotation in traditional methods are avoided. The application of the deep vectorization processing technology further enhances the text semantic understanding ability, making the retrieval result more in line with the user's needs.
[0020] In some embodiments of the present application, when collecting PDF document data, the source of the PDF document data is an academic paper library, an enterprise report library, an open e-book library and government and public institution documents; the data collection strategy includes: priority for diversity and sufficient data volume.
[0021] To ensure the diversity and representativeness of the data, PDF documents can be collected from the following channels: Academic paper library: -[arXiv](https: / / arxiv.org / ): Preprint papers covering fields such as physics, mathematics, and computer science.
[0022] -[PubMed Central](https: / / www.ncbi.nlm.nih.gov / pmc / ): Open access papers in the field of biomedicine.
[0023] -[IEEE Xplore](https: / / ieeexplore.ieee.org / ): Academic papers in the fields of engineering and technology.
[0024] -[SpringerOpen](https: / / www.springeropen.com / ): Multidisciplinary open access journals.
[0025] Features: Academic papers usually contain complex typesetting, formulas, charts, and references, suitable for training models to process structured content.
[0026] Enterprise Report Library -[SEC EDGAR](https: / / www.sec.gov / edgar / searchedgar / companysearch.html): Financial reports of US publicly traded companies.
[0027] -[Bloomberg](https: / / www.bloomberg.com / ): Corporate financial reports and analysis reports.
[0028] -[McKinsey](https: / / www.mckinsey.com / ): Industry research reports.
[0029] Features: Enterprise reports usually contain tables, charts, title hierarchies, and structured text, suitable for training models to extract key information.
[0030] Open E-book Library -[Project Gutenberg](https: / / www.gutenberg.org / ): Free e-books covering fields such as literature and history.
[0031] -[Google Books](https: / / books.google.com / ): Partially open e-books.
[0032] -[Internet Archive](https: / / archive.org / ): Contains various types of e-books and documents.
[0033] Features: E-books usually contain chapters, paragraphs, illustrations, etc., suitable for training models to process long texts and diverse layouts.
[0034] Government and Public Institution Documents [United Nations Document Library](https: / / www.un.org / en / documents / ): Policy reports, resolutions, etc.
[0035] [World Bank Open Data](https: / / www.worldbank.org / en / data): Economic and Social Reports.
[0036] [Open Data Platforms of Governments of Various Countries](https: / / www.data.gov / ): Such as the open data of the US government.
[0037] Features: Government documents usually contain formal language, tables, and structured data, which are suitable for training models to process formal documents.
[0038] To ensure the diversity and quality of the data, the following strategies can be adopted: 1. Diversity First Thematic Diversity: Cover multiple fields such as technology, finance, literature, and law.
[0039] Language Diversity: Include multi-language documents such as Chinese and English.
[0040] Layout Diversity: Collect PDFs with different layouts such as single-column, multi-column, mixed text and graphics, and dense tables.
[0041] 2. Sufficient Data Volume Scale: Collect at least hundreds of thousands of PDF documents to ensure the sufficiency of model training.
[0042] Incremental Update: Regularly update the data from the sources to keep the data up-to-date.
[0043] It can be understood that by collecting data from multiple authoritative and diverse information sources, the comprehensiveness and representativeness of the training set are ensured, thereby improving the generalization ability of the deep learning model. The strategy of diversity first helps the model learn text features of different styles and structures, while sufficient data volume ensures that the model can fully learn the potential rules and patterns in the text, further improving the accuracy and efficiency of text parsing and extraction.
[0044] In some embodiments of the present application, artificial annotation and preprocessing are performed on the PDF document data, including: according to a predefined tag system, artificially annotating the title, text, formulas, tables, images, and captions in the PDF document data to obtain annotation data, and storing the annotation data in a standardized format; performing preprocessing on the annotation data, and the preprocessing includes duplicate removal processing, normalization processing, image resolution standard processing, and sentence and word vector processing.
[0045] The standardized format storage can adopt the JSON format.
[0046] It is understandable that through manual annotation and preprocessing steps, the accuracy and consistency of the training data are ensured, providing a high-quality data foundation for the training of deep learning models. The deduplication process avoids the interference of duplicate data, the normalization process enables data from different sources to be compared and learned within the same framework, the image resolution standard processing ensures the clarity and readability of image data, and the sentence vector processing converts text data into a numerical form that can be understood by the model, facilitating subsequent text parsing and extraction.
[0047] In some embodiments of the present application, after manually annotating and preprocessing the PDF document data, a data quality assessment is performed on the preprocessed annotated data. The evaluation metrics used include: data diversity evaluation, data integrity evaluation, data accuracy evaluation, data consistency evaluation, and data usability evaluation.
[0048] It is understandable that through data quality assessment, the overall quality of the annotated data can be comprehensively measured to ensure that the data meets the training requirements of the deep learning model. The data diversity evaluation ensures that the data covers various possible text types and structures, enhancing the generalization ability of the model; the data integrity evaluation guarantees the integrity and non-missingness of the data, avoiding model training biases caused by incomplete data; the data accuracy evaluation verifies the accuracy and reliability of the data, ensuring that the model learns true and effective information; the data consistency evaluation ensures the consistency of the annotation results among different annotators, avoiding model training chaos caused by annotation differences; the data usability evaluation considers the value and effect of the data in actual applications, ensuring that the data used for model training has practical application significance.
[0049] In some embodiments of the present application, the deep learning model is initialized, including: adjusting the number of attention heads, optimizing the embedding layer dimension, adding a convolutional layer to fuse local features, designing a multi-scale feature pyramid, initializing the model weights, and encoding spatially aware positions.
[0050] 1. Adjustment of the number of attention heads Objective: Enhance the model's ability to capture layout features at different scales Original configuration: Standard Transformers usually use 8 - 16 attention heads.
[0051] Transformation plan: Increase the number of heads (e.g., from 16 to 24 heads): Capture features at different granularities in the layout (such as text lines, paragraphs, table areas) through more heads in parallel.
[0052] Hierarchical attention allocation: Use more local heads in the shallow layer (focus on adjacent blocks), and increase global heads in the deep layer (analyze the overall page structure).
[0053] 2. Optimize the Embedding Layer Dimension Goal: Balance feature expression ability and computational efficiency Original configuration: The commonly used embedding dimension of ViT is 768, suitable for general image classification.
[0054] Transformation plan: Increase the dimension (1024): Enhance the encoding ability for complex layout elements (nested tables, multi-column text).
[0055] Dynamic dimension allocation: Use low dimensions in the shallow layer to capture simple features (edges, directions), and gradually increase the dimensions in the deep layer.
[0056] 3. Add Convolution Layers to Fuse Local Features Goal: Make up for the deficiency of Transformer in local feature extraction Transformation plan: Pre-convolutional network: Use a lightweight CNN (such as the first 3 layers of ResNet) to extract low-level features before chunking.
[0057] Attention-convolution hybrid module: Insert depthwise separable convolution inside the Transformer block.
[0058] 4. Design a Multi-scale Feature Pyramid Goal: Adapt to the large difference in element sizes in PDF pages Implementation plan: Hierarchical downsampling: Generate multi-scale feature maps through convolutional layers (such as 1 / 4, 1 / 8, 1 / 16 of the original resolution).
[0059] Cross-scale attention: Apply Transformer within each scale and perform feature fusion through upsampling / downsampling.
[0060] 5. Initialize Model Weights Goal: Accelerate convergence and improve the final performance Scheme selection: Pre-trained transfer: Initialize with the weights of LayoutLM or DocFormer, which have been pre-trained on document data.
[0061] Hybrid initialization: Use ImageNet pre-trained weights for the convolutional part and random initialization for the Transformer part.
[0062] 6. Encode Spatial-aware Positions Goal: Explicitly encode two-dimensional layout information Improvement plan: Relative position encoding: Calculate the relative coordinate differences (Δx, Δy) between elements as the attention bias term.
[0063] Learnable layout encoding: Map the element bounding box coordinates (xmin, ymin, xmax, ymax) to position features through an MLP.
[0064] It can be understood that through the model initialization strategy, the performance of the deep learning model can be significantly improved. Adjusting the number of attention heads can capture key information in the text more finely, and optimizing the embedding layer dimension helps the model better understand and represent text features. Adding a convolutional layer to fuse local features can enhance the model's sensitivity to local text information, and designing a multi-scale feature pyramid allows the model to capture text features at different scales, improving the model's generalization ability. Initializing the model weights and encoding space-aware positions can ensure that the model has good performance at the beginning of training and accelerate the model's convergence speed.
[0065] In some embodiments of the present application, when setting the parameters for training the deep learning model, the parameters include: learning rate, batch size, number of training epochs, optimizer and regularization, loss function, key parameter linkage rule, hardware resource adaptation rule, and parameter tuning and verification method; the optimizer selects the Adam optimizer, and the loss function selects the cross-entropy loss function.
[0066] 1. Learning Rate Initial setting: Base learning rate: 1e-3 (default value of Adam optimizer) Adjustment rule: Linear warmup: Gradually increase from 1e-6 to the target value in the first 5 epochs to avoid early gradient explosion.
[0067] Dynamic monitoring: When the mAP of the validation set does not improve for 3 consecutive epochs, the learning rate decays to 0.2 times the original.
[0068] 2. Batch Size Initial setting: Single-card batch: 16 (A100 video memory occupancy is about 22GB) Gradient accumulation: When the video memory is insufficient, set gradient_accumulation_steps = 4, equivalent batch size 64.
[0069] Adjustment rule: Gradual amplification strategy: Use a small batch (16) for stable training in the first 10 epochs, and then increase to 32 to accelerate convergence.
[0070] Batch-learning rate linkage: When the batch size doubles, the learning rate is increased by √2 times (refer to the linear scaling principle).
[0071] 3. Number of Training Epochs Initial Settings: Basic Setting: 50 epochs (based on the average convergence period of the document layout task) Dynamic Termination Conditions: Early Stopping Mechanism: Terminate training when the validation loss has not improved for 8 consecutive epochs.
[0072] Performance Threshold: Terminate early when the mAP of the validation set reaches 95% and the overfitting index (train / val loss ratio) < 1.1.
[0073] 4. Optimizer and Regularization Optimizer Configuration: optimizer = AdamW(model.parameters(), lr=1e-3, betas=(0.92, 0.99), # Smooth the momentum to adapt to sparse gradients weight_decay=0.05 # Combat overfitting in Transformer) Regularization Strategy: DropPath: Apply random depth dropout with a probability of 0.1 to Transformer blocks.
[0074] MixUp Data Augmentation: Mix samples in the image space with α = 0.2 to improve the robustness of position encoding.
[0075] 5. Loss Function Design Multi-Task Loss Combination: loss = 1.2 * cls_loss + 0.8 * bbox_loss + 0.5 * iou_loss Classification Loss: Focal Loss (γ = 2.0) to alleviate the long-tail distribution problem.
[0076] Location Regression: Joint optimization of the bounding box location using Smooth L1 Loss + GIoU Loss.
[0077] Dynamic Weight Adjustment: Re-balance the weights according to the gradient magnitudes of each task every 5 epochs (refer to GradNorm).
[0078] 6. Linkage Rules for Key Parameters Learning Rate vs Model Depth: Deep Transformers use a low learning rate (base learning rate × 0.7), while shallow CNNs use a high learning rate (× 1.2) Batch Size vs BN: When the batch size < 32, use GroupNorm instead of BatchNorm to maintain stability Number of attention heads vs Dropout: For every 4 additional attention heads, the Dropout rate of the attention matrix increases by 0.05 (with an upper limit of 0.3). 7. Hardware resource adaptation rules Single - card mode: Automatically enable Mixed Precision (AMP) and Gradient Checkpointing technology Multi - card training: Use Sharded Data Parallel to accelerate the distributed training of large models (> 1B parameters) Handling of insufficient memory: torch.cuda.empty_cache() # Clear the cache every epoch reduce_tensor_size = True # Automatically convert float32 to bfloat16 8. Parameter tuning and validation methods Hyperparameter search: Use Bayesian optimization to search within a limited range (example): search_space = {'lr': (1e - 5, 1e - 3), 'weight_decay': (0.01, 0.3), 'drop_path': (0.0, 0.3)} Ablation experiment design: Separate the influence of parameters through the orthogonal experiment method (such as fixing the learning rate and testing different combinations of the number of heads) Through this systematic parameter management strategy, the model performance can be maximized while maintaining training stability. It is recommended to use the default configuration initially to quickly verify the feasibility of the model, and then gradually enable advanced strategies for fine - tuning.
[0079] It can be understood that by setting the parameters of deep learning model training, the training process of the model can be further optimized, and the performance and stability of the model can be improved. The reasonable setting of the learning rate and batch size can balance the training speed and convergence effect of the model, and the determination of the number of training epochs ensures that the model fully learns the text features. Selecting the Adam optimizer can adaptively adjust the learning rate and improve the training efficiency of the model. The cross - entropy loss function can accurately measure the difference between the model prediction result and the true label, guiding the model to be optimized. The application of key parameter linkage rules, hardware resource adaptation rules, and parameter tuning and validation methods can ensure that the model maintains excellent performance in different scenarios, improving the practicality and generalization ability of the model.
[0080] In some embodiments of the present application, when training a deep learning model based on the training set, calculate the loss value and update the model weights through the backpropagation algorithm to gradually reduce the loss value.
[0081] It can be understood that by training a deep learning model based on a training set and calculating the loss value, the performance of the model can be evaluated in real time, and the model weights can be continuously updated through the backpropagation algorithm, thereby gradually optimizing the model. This method can ensure that the model continuously approaches the optimal solution during training, improving the accuracy and reliability of the model. At the same time, the gradual decrease of the loss value also reflects the continuous progress and optimization of the model during training, providing more accurate model support for subsequent text parsing, extraction, and key point source location.
[0082] In some embodiments of the present application, converting the recognition result into a text format or a chart / table format includes: if the recognition result is text, using syntactic analysis and semantic understanding techniques in natural language processing to convert the recognition result into a clear and editable text format; if the recognition result is a chart or a table, adopting an innovative algorithm based on image recognition and data mining to achieve accurate data extraction and text-based presentation.
[0083] It can be understood that through flexible and diverse conversion methods, the needs of different users for the presentation form of the recognition result can be met. For the recognition result in text format, the application of syntactic analysis and semantic understanding techniques not only improves the readability and editability of the text but also helps with subsequent information processing and knowledge mining. For the recognition result in chart or table format, the application of the innovative algorithm achieves accurate data extraction and text-based presentation, enabling complex data information to be presented to users in a more intuitive and understandable way.
[0084] In some embodiments of the present application, an advanced language model based on the Transformer architecture is introduced to perform deep vectorization processing on the text converted into a text format or a chart / table format, including: the advanced language model based on the Transformer architecture includes BERT or ERNIE; performing a tokenization operation on the text converted into a text format or a chart / table format, inputting the tokenized text into the advanced language model based on the Transformer architecture, and encoding the text through the multi-layer Transformer encoder of the model to obtain the vector representation of each word or sub-word unit.
[0085] Taking a news text "Today, a technology company released a brand-new smartphone with a high-performance processor and a high-definition camera" as an example, after being processed by the BERT model, each word is mapped into a 768-dimensional (BERT-base model) vector space. These vectors not only contain the semantic information of the words but also consider the position and semantic relationships of the words in the context. For example, the vector representation of "smartphone" will be affected by context words such as "technology company", "brand-new", "high-performance processor", and "high-definition camera", thus generating a vector representation with rich semantic connotations.
[0086] It can be understood that by introducing advanced language models based on the Transformer architecture, such as BERT or ERNIE, for deep vectorization processing of text, the semantic information and context relationships in the text can be captured more accurately. The tokenization operation breaks the text into smaller language units, enabling the model to understand the text content more meticulously. The multi-layer Transformer encoder gradually extracts the deep features in the text through multi-level encoding of the text, providing a more rich and accurate vector representation for subsequent text parsing and key point source location.
[0087] In some embodiments of the present application, based on the advanced vector space model and similarity measurement algorithm, semantic retrieval is implemented, including: constructing a vector space model and storing all vectorized texts in a high-dimensional vector space; when the user inputs a retrieval requirement, performing the same vectorization operation on the retrieval requirement to obtain a retrieval vector; calculating the similarity scores between the retrieval vector and all stored vectors in the vector space using cosine similarity; sorting the words or sub-word units according to the similarity scores and returning the document fragments to which the words or sub-word units most relevant to the retrieval requirement belong to the user.
[0088] For example, when the user inputs "Research report on the performance of smartphone processors" as the retrieval requirement, after the system vectorizes the retrieval text, it finds the vectors related to keywords such as "smartphone", "processor performance", and "research report" in the vector space, and returns relevant document contents such as academic literature fragments and technical blog articles that contain these keywords and have similar semantics according to the similarity scores, helping the user quickly locate the required information.
[0089] It can be understood that by constructing a vector space model and storing the vectorized representation of the text in a high-dimensional space, the effective organization and storage of text information are achieved. When a user retrieves information, the retrieval requirement is also vectorized, ensuring the consistency between the retrieval process and text processing, and improving the accuracy and efficiency of the retrieval. As an effective similarity measurement method, cosine similarity calculation can objectively reflect the similarity degree between vectors and provide a reliable basis for sorting. Sorting the words or sub-word units according to the similarity scores and returning the most relevant document fragments enables the user to quickly locate the required information, greatly enhancing the user experience.
[0090] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0091] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0092] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide for implementing the functions specified in Figure 1 one process or multiple processes and / or blocksFigure 1 Steps of functions specified in one or more boxes.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. An artificial intelligence text parsing, refining and key point source location method, characterized in that, Including: Collect PDF document data, perform manual annotation and preprocessing on the PDF document data to obtain a training set, a validation set, and a test set; Based on the Transformer architecture and the hardware environment adapted for deep learning training, initialize the deep learning model and set the parameters for training the deep learning model; Train the deep learning model based on the training set, and verify and test the trained deep learning model according to the validation set and the test set. If the verification and test pass, determine the deep learning model that passes the verification and test as the PDF document layout recognition model; Use the PDF document layout recognition model to recognize the PDF document to be recognized, and convert the recognition result into a text format or a chart / table format; Introduce an advanced language model based on the Transformer architecture to perform deep vectorization processing on the text converted into a text format or a chart / table format; Based on the advanced vector space model and the similarity measurement algorithm, implement semantic retrieval.
2. The artificial intelligence text parsing, refining and key point source location method according to claim 1, characterized in that, When collecting PDF document data, the source of the PDF document data is academic paper libraries, enterprise report libraries, publicly available e-book libraries, and government and public institution documents; The data collection strategy includes: diversity first and sufficient data volume.
3. The artificial intelligence text parsing, refining and key point source location method according to claim 1, characterized in that, Performing manual annotation and preprocessing on the PDF document data includes: According to the predefined tag system, manually annotate the title, text, formula, table, image, and caption in the PDF document data to obtain annotation data, and store the annotation data in a standardized format; Preprocess the annotation data, and the preprocessing includes duplicate removal processing, normalization processing, image resolution standard processing, and word and sentence vector processing.
4. The artificial intelligence text parsing, refining and key point source location method according to claim 1, characterized in that, After performing manual annotation and preprocessing on the PDF document data, perform data quality evaluation on the preprocessed annotation data. The evaluation indicators used include: data diversity evaluation, data integrity evaluation, data accuracy evaluation, data consistency evaluation, and data usability evaluation.
5. The artificial intelligence text parsing, refining and key point source location method according to claim 1, characterized in that, Initializing the deep learning model includes: Adjust the number of attention heads, optimize the embedding layer dimension, add a convolutional layer to fuse local features, design a multi-scale feature pyramid, initialize the model weights, and encode the spatial perception position.
6. The artificial intelligence text parsing, refining and key point source location method according to claim 1, characterized in that, When setting the parameters for training the deep learning model, the parameters include: learning rate, batch size, number of training epochs, optimizer and regularization, loss function, key parameter linkage rule, hardware resource adaptation rule, and parameter tuning verification method; The optimizer selects the Adam optimizer, and the loss function selects the cross-entropy loss function.
7. The artificial intelligence text parsing, refining and key point source location method according to claim 1, characterized in that, When training the deep learning model based on the training set, calculate the loss value and update the model weights through the backpropagation algorithm to gradually reduce the loss value.
8. The artificial intelligence text parsing, refining and key point source location method according to claim 1, characterized in that, Converting the recognition result into a text format or a chart / table format includes: If the recognition result is text, use the syntax analysis and semantic understanding techniques in natural language processing to convert the recognition result into a clear and editable text format; If the recognition result is a chart or a table, adopt an innovative algorithm based on image recognition and data mining to achieve accurate data extraction and text presentation.
9. The artificial intelligence text parsing, refining and key point source location method according to claim 1, characterized in that, Introduce advanced language models based on the Transformer architecture to perform deep vectorization on text converted into text format or chart / table format, including: The advanced language models based on the Transformer architecture include BERT or ERNIE; Perform word segmentation on the text converted into text format or chart / table format, input the segmented text into the advanced language model based on the Transformer architecture, and encode the text through the multi-layer Transformer encoder of the model to obtain the vector representation of each word or sub-word unit.
10. The artificial intelligence text parsing, refining and key point source location method according to claim 9, characterized in that,Based on the advanced vector space model and similarity measurement algorithm, implement semantic retrieval, including: Construct a vector space model and store all vectorized texts in a high-dimensional vector space; When the user inputs a retrieval requirement, perform the same vectorization operation on the retrieval requirement as the text processing to obtain a retrieval vector; Use cosine similarity calculation to calculate the similarity score between the retrieval vector and all stored vectors in the vector space; Sort the words or sub-word units according to the similarity score and return the document fragment to which the word or sub-word unit most relevant to the retrieval requirement belongs to the user.
Citation Information
Patent Citations
Structured analysis method for portable document format file and related product
CN117473980A
Document analysis method and system
CN118261140A
Intelligent data information rapid checking and question-answering system
CN119149691A
Document data processing method and device based on large model
CN119940343A
Cited By
Academic paper title grading device and method based on Bert
CN120996032A
Method for predictive analysis of psychological manipulation in digital communications
RU2869631C1