An office file processing method and system based on big data and deep learning

CN122287846BActive Publication Date: 2026-09-11GUANGXI ZHIFU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610613614.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-09-11
Estimated Expiration
2046-05-07

AI Technical Summary

Technical Problem

现有办公文件处理方式多依赖人工分类、固定模板和关键词匹配,面对多来源、大批量文件时,容易出现处理效率低、分类不准确和归档不一致的问题

Benefits of technology

首先,本发明通过对办公文本序列执行语法分析与结构标注处理,并结合预训练的BERT模型进行上下文语义建模,实现了对文件句法依赖关系、字段边界以及语法角色的精确识别,使文件主题、业务实体与关键字段的提取结果具有更高一致性与稳定性,从而提高办公文件语义理解的准确程度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122287846B_ABST
    Figure CN122287846B_ABST
Patent Text Reader

Abstract

The application discloses an office file processing method and system based on big data and deep learning, comprising the following steps: S1, collecting office file data and preprocessing to generate office text sequences; S2, performing syntax analysis and structure annotation on the office text sequences to form file structure sequences; S3, performing semantic modeling and label matching on the file structure sequences to obtain semantic label sequences; S4, extracting file topics, business entities and key fields according to the semantic label sequences to form semantic feature sequences; S5, performing node matching and path expansion based on the semantic feature sequences to generate a graph matching path set; S6, classifying and archiving according to the graph matching path set and transferring the process to form processing feedback data; and S7, updating according to the processing feedback data. The application has the advantages of high semantic understanding accuracy, high file processing automation degree and strong processing result consistency by combining syntax analysis, BERT model and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of document processing technology, and in particular to an office document processing method and system based on big data and deep learning. Background Technology

[0002] As the scale of enterprise office data continues to expand, office documents come from emails, collaboration platforms, business process nodes, and local storage locations, and these documents differ in format, structure, and expression. Current office document processing methods largely rely on manual classification, fixed templates, and keyword matching. When faced with multiple sources and large volumes of documents, these methods are prone to problems such as low processing efficiency, inaccurate classification, and inconsistent archiving.

[0003] In existing technologies, some solutions use natural language processing methods to segment office documents, extract keywords, and classify text. However, these methods usually only perform shallow text feature processing and are difficult to accurately identify syntactic dependencies, field boundaries, and grammatical roles in complex sentences, resulting in unstable extraction results for document topics, business entities, and key fields.

[0004] Some solutions introduce deep learning models such as BERT for semantic modeling, but existing methods mostly focus on text classification or entity recognition, lacking joint processing of office document structure information, syntactic analysis results and semantic tags, making it difficult to form semantic feature sequences for document classification, archiving and workflow.

[0005] In addition, existing office document circulation usually relies on fixed rules or manual configuration processes, lacking a node matching and association path expansion mechanism based on office knowledge graphs. It cannot accurately match file type nodes, business entity nodes, process nodes, and archive nodes based on the semantic features of the documents, and it is difficult to update model parameters and knowledge graphs in reverse by processing feedback data.

[0006] Therefore, how to provide an office document processing method and system based on big data and deep learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] One objective of this invention is to propose an office document processing method and system based on big data and deep learning. This invention combines syntax analysis, BERT semantic modeling, and knowledge graph association matching technology to perform structural parsing, semantic feature extraction, and graph path matching processing on multi-source office documents, realizing integrated processing of document classification and archiving and workflow. At the same time, it dynamically updates the model parameters and knowledge graph based on processing feedback data, and has the advantages of high semantic understanding accuracy, high degree of automation in document processing, and strong consistency of processing results.

[0008] An office document processing method based on big data and deep learning according to an embodiment of the present invention includes the following steps: S1. Collect and preprocess text data from multiple sources of office documents to generate a sequence of office texts; S2. Perform syntactic analysis on the office text sequence to extract the syntactic dependencies and lexical structure information of each office text, and perform structural annotation on the office text sequence based on the analysis results to identify field boundaries and syntactic roles, forming a file structure sequence; S3. Perform contextual semantic modeling on the file structure sequence using a pre-trained BERT model to form a semantic representation sequence. Perform label matching on the semantic representation sequence based on a preset semantic label set to obtain a semantic label sequence. S4. Perform field merging operation on the file structure sequence based on the semantic label sequence, and perform field semantic association calculation on the merging result in the BERT model to extract the file topic, business entity and key fields to form a semantic feature sequence. S5. Based on semantic feature sequences, perform node similarity calculation and target node matching operations in the pre-constructed office knowledge graph, and expand the associated paths of each target node to generate a graph matching path set; S6. Based on the path set matched by the graph, perform classification, archiving and workflow operations on each office document, and record the classification, archiving and workflow results to form processing feedback data; S7. Update the parameters of the BERT model and the office knowledge graph based on the processed feedback data.

[0009] Optionally, the text data of the multi-source office documents represents the file data to be processed from different storage locations and business process nodes. The preprocessing includes format unification, text extraction, noise character removal, duplicate text removal, paragraph ordering, and metadata standardization.

[0010] Optionally, S2 specifically includes: S21. Perform sentence segmentation and word segmentation operations on each office text in the office text sequence to generate an office sentence sequence and a word sequence. S22. Perform part-of-speech tagging and lexical structure recognition on the word sequence to extract noun phrases, verb phrases, time phrases, numerical phrases and business phrases to form lexical structure information; S23. Perform syntactic dependency analysis based on the office clause sequence and lexical structure information to identify subject-predicate relations, verb-object relations, attributive-head relations, adverbial-head relations and coordinate relations, and form syntactic dependency relations. S24. Perform structural annotation operations on the office text sequence based on syntactic dependencies and lexical structure information, marking the title segment, body segment, table segment, signature segment, and attachment segment to obtain the structurally annotated sequence; S25. Based on the structural annotation sequence, identify the start position, end position and paragraph to which the field belongs in each office document, and generate a field boundary sequence; S26. Perform syntactic role labeling on the content of each field within the field boundary sequence based on syntactic dependencies to determine the subject role, action role, object role, time role, numerical role, and responsibility role corresponding to the field content; S27. Sequentially associate the field boundary sequence, syntax role annotation results, syntactic dependencies, and lexical structure information to form a file structure sequence.

[0011] Optionally, S26 specifically includes: S261. Extract the field start position, field end position, field paragraph and field text content corresponding to each field content within the field boundary sequence, and generate a field content sequence; S262. Based on the start and end positions of the fields, map each field in the field content sequence to a word sequence to determine the field word set corresponding to each field. S263. Based on syntactic dependency relations, perform dependency path extraction operations on each word in the field word set to obtain the word dependency path set between each adjacent word; S264. Identify the lexical type of each lexical based on the lexical dependency path set, and map the identification results to the corresponding fields to generate field dependency features. The lexical types include head words, modifiers, governing words, and governed words. S265. Based on field dependency features, perform role candidate labeling operations on each field to obtain candidate roles for topic, action, object, time, numerical, and responsibility. S266. Based on the paragraph to which the field belongs, lexical structure information, and field dependency features, perform consistency verification operations on each role candidate and generate role verification results; S267. Based on the role verification results, perform role confirmation operations on each field to determine the subject role, action role, object role, time role, numerical role, and responsibility role corresponding to the field content, and bind them to the corresponding fields to generate grammatical role annotation results.

[0012] Optionally, S3 specifically includes: S31. Perform positional encoding on the file structure sequence to generate a structure encoding sequence; S32. Perform context semantic modeling operations on the structure encoding sequence using a pre-trained BERT model to generate a context semantic sequence; S33. Based on a preset semantic tag set, perform tag matching operation on the context semantic sequence, and combine it with the field boundary sequence to obtain the candidate tag sequence corresponding to each field; S34. Based on the grammatical role annotation results and syntactic dependencies, perform consistency filtering on each candidate tag sequence to determine the target semantic tag corresponding to each field; S35. Bind each target semantic tag to the corresponding field according to the field order in the file structure sequence to obtain the semantic tag sequence.

[0013] Optionally, the pre-training process of the BERT model specifically includes: Collect text data from historical office documents and perform operations such as format standardization, noise character removal, duplicate text removal, and paragraph order reordering to generate a pre-trained text sequence; Sentence segmentation and word segmentation are performed on the pre-trained text sequence to extract sentence boundaries and generate word sequence. Each word in the word sequence is uniquely identified and encoded to generate a word encoding sequence. Based on the sequential position of each word in the word encoding sequence in the pre-trained text sequence, position encoding operation is performed on each word, and sentence segment identification encoding operation is performed on each word according to the sentence boundary to generate a pre-trained sequence; According to the preset mask ratio, target words are randomly selected from the pre-training sequence, and the target words are replaced with preset mask markers to generate a mask sequence. The BERT model then performs prediction calculations on the original words corresponding to the target words based on the mask sequence. The prediction bias is obtained by performing a difference analysis between the prediction calculation results and the pre-trained sequence. Sentence pairs are extracted from the pre-trained text sequence based on sentence boundaries. Sentence pairs consisting of adjacent sentences are marked as positive samples, and sentence pairs consisting of non-adjacent sentences are marked as negative samples, forming a sentence pair sequence. In the BERT model, the inter-sentence relation determination is performed on the sentence pair sequence, and the relation determination result is binary-labeled according to the positive and negative sample labels to obtain the determination bias. Based on the prediction bias and decision bias, the BERT model is updated with parameters, and the random mask, prediction calculation, difference analysis, positive and negative sample construction, inter-sentence relationship determination, decision bias calculation and parameter update operations are performed iteratively until the prediction bias and decision bias reach the preset convergence condition, thus obtaining the pre-trained BERT model.

[0014] Optionally, S4 specifically includes: S41. Perform tag aggregation operation on each field in the file structure sequence according to the semantic tag sequence, extract the set of fields with consistent tags, and perform order constraint processing on the field set in combination with the positional relationship of the fields in the file structure sequence to generate a field aggregation sequence. S42. Perform a field merging operation on each field set in the field aggregation sequence, merging fields with the same semantic label and syntactic dependency relationship into field units, and generating a field merging sequence. S43. Perform contextual semantic association calculation on the field merge sequence in the BERT model, extract the semantic association strength between each field unit, and construct the field association sequence; S44. Filter multiple representative field units that represent the core content of each file, and perform topic identification operation on each representative field unit according to the field association sequence and syntax role labeling results to generate a set of file topics; S45. Based on the field association sequence and syntactic dependency relationship, perform business entity identification operation on each field unit to generate a business entity set; S46. Based on the semantic tag sequence, the syntactic role annotation results and the field association sequence, perform key field filtering operations on each field unit in the field merging sequence to obtain a set of key fields. S47. Combine and associate the file topic set, business entity set, and key field set according to the order in the file structure sequence to generate a semantic feature sequence.

[0015] Optionally, S5 specifically includes: S51. Based on semantic feature sequences, perform node matching and retrieval operations in the pre-constructed office knowledge graph according to file topic, business entity and key fields, extract the corresponding file type nodes, business entity nodes, process nodes and archive nodes, and generate a candidate node set. S52. Perform cosine similarity calculation on the semantic feature sequence and the candidate node set to obtain the node similarity corresponding to each candidate node. S53. Perform a sorting and filtering operation on the candidate node set based on node similarity to determine the target node set that matches the semantic feature sequence; S54. Based on the target node set, extract the node connection relationship, business association relationship, process association relationship and archiving association relationship corresponding to each target node in the office knowledge graph, and generate a node association sequence; S55. Perform association path expansion operation on each target node according to the node association sequence to form a candidate matching path consisting of file type nodes, business entity nodes, process nodes and archive nodes. S56. Perform path integrity verification and semantic consistency verification operations on the candidate matching paths to obtain the graph matching path set.

[0016] Optionally, S6 specifically includes: S61. Extract file type nodes, business entity nodes, process nodes and archive nodes from the graph matching path set, and generate a file processing sequence according to the path connection order. S62. Determine the file category corresponding to each office file based on the file type node in the file processing sequence, and generate file classification results; S63. Determine the archiving directory, archiving level and archiving identifier of each office document based on the archiving node in the document processing sequence, and generate the classification archiving results; S64. Based on the process nodes in the document processing sequence, determine the corresponding flow nodes, flow order and processing responsibility entities for each office document, and generate the process flow path; S65. Perform workflow status registration, processing node binding, and responsible entity binding operations on each office document according to the workflow path, and generate workflow results; S66. Link and record the document classification results, classification archiving results, and process flow results to form processing feedback data.

[0017] An office document processing system based on big data and deep learning according to an embodiment of the present invention includes: The data acquisition module is used to collect and preprocess text data from multiple sources of office documents to generate a sequence of office texts. The syntax analysis module is used to perform syntax analysis on the sequence of office texts, extract the syntactic dependencies and lexical structure information of each office text, and perform structure annotation on the sequence of office texts based on the analysis results, identify field boundaries and grammatical roles, and form a document structure sequence. The label matching module is used to perform contextual semantic modeling operations on the file structure sequence through a pre-trained BERT model to form a semantic representation sequence. Based on a preset semantic label set, a label matching operation is performed on the semantic representation sequence to obtain a semantic label sequence. The feature extraction module is used to perform field merging operations on the file structure sequence based on the semantic label sequence, and to perform field semantic association calculations on the merging results in the BERT model to extract the file topic, business entities and key fields to form a semantic feature sequence. The association matching module is used to perform node similarity calculation and target node matching operations in a pre-built office knowledge graph based on semantic feature sequences, and to expand the association paths of each target node to generate a graph matching path set; The document processing module is used to perform classification, archiving, and workflow operations on each office document based on the path set matched by the graph, and to record the classification, archiving, and workflow results to form processing feedback data. The feedback update module is used to update the parameters of the BERT model and the knowledge graph based on the processed feedback data.

[0018] The beneficial effects of this invention are: First, this invention performs syntactic analysis and structural annotation on office text sequences, and combines pre-trained BERT models for contextual semantic modeling. This enables accurate identification of document syntactic dependencies, field boundaries, and syntactic roles, resulting in higher consistency and stability in the extraction of document topics, business entities, and key fields, thereby improving the accuracy of semantic understanding of office documents.

[0019] Secondly, this invention introduces semantic feature sequences into a pre-constructed office knowledge graph to perform node matching and associated path expansion processing. By matching the path set of the graph, it realizes file classification and archiving and process flow, so that the file processing results are consistent with the actual business semantics. At the same time, it avoids the deviation problem caused by fixed rule matching, and improves the automation level of office file processing and the reliability of processing results.

[0020] Finally, this invention continuously updates the BERT model parameters and the office knowledge graph by processing feedback data, so that the semantic modeling capability and graph structure are optimized synchronously with changes in business data, enhancing the method's adaptability to multiple types of office documents, maintaining stable processing performance in large-scale data scenarios, thereby improving the overall office document processing efficiency and the system's sustainable optimization capability. Attached Figure Description

[0021] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of an office document processing method based on big data and deep learning proposed in this invention; Figure 2 This is a flowchart illustrating the knowledge graph matching and document processing of an office document processing method based on big data and deep learning proposed in this invention. Figure 3 This is a module structure diagram of an office document processing system based on big data and deep learning proposed in this invention. Detailed Implementation

[0022] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0023] refer to Figures 1-2 An office document processing method based on big data and deep learning includes the following steps: S1. Collect and preprocess text data from multiple sources of office documents to generate a sequence of office texts; S2. Perform syntactic analysis on the office text sequence to extract the syntactic dependencies and lexical structure information of each office text, and perform structural annotation on the office text sequence based on the analysis results to identify field boundaries and syntactic roles, forming a file structure sequence; S3. Perform contextual semantic modeling on the file structure sequence using a pre-trained BERT model to form a semantic representation sequence. Perform label matching on the semantic representation sequence based on a preset semantic label set to obtain a semantic label sequence. S4. Perform field merging operation on the file structure sequence based on the semantic label sequence, and perform field semantic association calculation on the merging result in the BERT model to extract the file topic, business entity and key fields to form a semantic feature sequence. S5. Based on semantic feature sequences, perform node similarity calculation and target node matching operations in the pre-constructed office knowledge graph, and expand the associated paths of each target node to generate a graph matching path set; S6. Based on the path set matched by the graph, perform classification, archiving and workflow operations on each office document, and record the classification, archiving and workflow results to form processing feedback data; S7. Update the parameters of the BERT model and the office knowledge graph based on the processed feedback data.

[0024] In this embodiment, the text data of multi-source office documents represents the file data to be processed from different storage locations and business process nodes. Preprocessing includes format unification, text extraction, noise character removal, duplicate text removal, paragraph ordering, and metadata normalization.

[0025] In this embodiment, S2 specifically includes: S21. Perform sentence segmentation and word segmentation operations on each office text in the office text sequence to generate an office sentence sequence and a word sequence. S22. Perform part-of-speech tagging and lexical structure recognition on the word sequence to extract noun phrases, verb phrases, time phrases, numerical phrases and business phrases to form lexical structure information; S23. Perform syntactic dependency analysis based on the office clause sequence and lexical structure information to identify subject-predicate relations, verb-object relations, attributive-head relations, adverbial-head relations and coordinate relations, and form syntactic dependency relations. S24. Perform structural annotation operations on the office text sequence based on syntactic dependencies and lexical structure information, marking the title segment, body segment, table segment, signature segment, and attachment segment to obtain the structurally annotated sequence; S25. Based on the structural annotation sequence, identify the start position, end position and paragraph to which the field belongs in each office document, and generate a field boundary sequence; S26. Perform syntactic role labeling on the content of each field within the field boundary sequence based on syntactic dependencies to determine the subject role, action role, object role, time role, numerical role, and responsibility role corresponding to the field content; S27. Sequentially associate the field boundary sequence, syntax role annotation results, syntactic dependencies, and lexical structure information to form a file structure sequence.

[0026] In this embodiment, S26 specifically includes: S261. Extract the field start position, field end position, field paragraph and field text content corresponding to each field content within the field boundary sequence, and generate a field content sequence; S262. Based on the start and end positions of the fields, map each field in the field content sequence to a word sequence to determine the field word set corresponding to each field. S263. Based on syntactic dependency relations, perform dependency path extraction operations on each word in the field word set to obtain the word dependency path set between each adjacent word; S264. Identify the lexical type of each lexical based on the lexical dependency path set, and map the identification results to the corresponding fields to generate field dependency features. The lexical types include head words, modifiers, governing words, and governed words. S265. Based on field dependency features, perform role candidate labeling operations on each field to obtain candidate roles for topic, action, object, time, numerical, and responsibility. S266. Based on the paragraph to which the field belongs, lexical structure information, and field dependency features, perform consistency verification operations on each role candidate and generate role verification results; S267. Based on the role verification results, perform role confirmation operations on each field to determine the subject role, action role, object role, time role, numerical role, and responsibility role corresponding to the field content, and bind them to the corresponding fields to generate grammatical role annotation results.

[0027] In this embodiment, S3 specifically includes: S31. Perform positional encoding on the file structure sequence to generate a structure encoding sequence; S32. Perform context semantic modeling operations on the structure encoding sequence using a pre-trained BERT model to generate a context semantic sequence; S33. Based on a preset semantic tag set, perform tag matching operation on the context semantic sequence, and combine it with the field boundary sequence to obtain the candidate tag sequence corresponding to each field; S34. Based on the grammatical role annotation results and syntactic dependencies, perform consistency filtering on each candidate tag sequence to determine the target semantic tag corresponding to each field; S35. Bind each target semantic tag to the corresponding field according to the field order in the file structure sequence to obtain the semantic tag sequence.

[0028] In this embodiment, the pre-training process of the BERT model specifically includes: Collect text data from historical office documents and perform operations such as format standardization, noise character removal, duplicate text removal, and paragraph order reordering to generate a pre-trained text sequence; Sentence segmentation and word segmentation are performed on the pre-trained text sequence to extract sentence boundaries and generate word sequence. Each word in the word sequence is uniquely identified and encoded to generate a word encoding sequence. Based on the sequential position of each word in the word encoding sequence in the pre-trained text sequence, position encoding operation is performed on each word, and sentence segment identification encoding operation is performed on each word according to the sentence boundary to generate a pre-trained sequence; According to the preset mask ratio, target words are randomly selected from the pre-training sequence, and the target words are replaced with preset mask markers to generate a mask sequence. The BERT model then performs prediction calculations on the original words corresponding to the target words based on the mask sequence. The prediction bias is obtained by performing a difference analysis between the prediction calculation results and the pre-trained sequence. Sentence pairs are extracted from the pre-trained text sequence based on sentence boundaries. Sentence pairs consisting of adjacent sentences are marked as positive samples, and sentence pairs consisting of non-adjacent sentences are marked as negative samples, forming a sentence pair sequence. In the BERT model, the inter-sentence relation determination is performed on the sentence pair sequence, and the relation determination result is binary-labeled according to the positive and negative sample labels to obtain the determination bias. Based on the prediction bias and decision bias, the BERT model is updated with parameters, and the random mask, prediction calculation, difference analysis, positive and negative sample construction, inter-sentence relationship determination, decision bias calculation and parameter update operations are performed iteratively until the prediction bias and decision bias reach the preset convergence condition, thus obtaining the pre-trained BERT model.

[0029] In this embodiment, S33 specifically includes: S331. Receive the context semantic sequence and extract the context semantic fragments corresponding to each field based on the field boundary sequence; S332. Perform semantic aggregation operation on each context semantic fragment to generate the field semantic representation corresponding to each field; S333. Read the preset semantic tag set and extract the tag semantic representation corresponding to each semantic tag in the semantic tag set; S334. Perform cosine similarity calculation on the semantic representation of each field and the semantic representation of each label to obtain the label matching value between each field and each semantic label. S335. Semantic tags with matching values ​​greater than a preset matching threshold are taken as candidate tags for the corresponding fields, and the candidate tags are sorted in descending order according to the matching values ​​to generate a candidate tag sequence for each field.

[0030] In this embodiment, S34 specifically includes: S341. Receive the candidate label sequence corresponding to each field, and extract the label matching value corresponding to each candidate label; S342. Construct consistency constraint rules based on the grammatical role annotation results and syntactic dependencies, and perform consistency filtering operations on each candidate label sequence according to the consistency constraint rules to extract candidate labels that satisfy the consistency constraints to form a set of filtered labels. S343. Select the candidate tag with the largest tag matching value from the set of filter tags corresponding to each field, and use it as the target semantic tag for the corresponding field.

[0031] In this embodiment, S4 specifically includes: S41. Perform tag aggregation operation on each field in the file structure sequence according to the semantic tag sequence, extract the set of fields with consistent tags, and perform order constraint processing on the field set in combination with the positional relationship of the fields in the file structure sequence to generate a field aggregation sequence. S42. Perform a field merging operation on each field set in the field aggregation sequence, merging fields with the same semantic label and syntactic dependency relationship into field units, and generating a field merging sequence. S43. Perform contextual semantic association calculation on the field merge sequence in the BERT model, extract the semantic association strength between each field unit, and construct the field association sequence; S44. Filter multiple representative field units that represent the core content of each file, and perform topic identification operation on each representative field unit according to the field association sequence and syntax role labeling results to generate a set of file topics; S45. Based on the field association sequence and syntactic dependency relationship, perform business entity identification operation on each field unit to generate a business entity set; S46. Based on the semantic tag sequence, the syntactic role annotation results and the field association sequence, perform key field filtering operations on each field unit in the field merging sequence to obtain a set of key fields. S47. Combine and associate the file topic set, business entity set, and key field set according to the order in the file structure sequence to generate a semantic feature sequence.

[0032] In this embodiment, S5 specifically includes: S51. Based on semantic feature sequences, perform node matching and retrieval operations in the pre-constructed office knowledge graph according to file topic, business entity and key fields, extract the corresponding file type nodes, business entity nodes, process nodes and archive nodes, and generate a candidate node set. S52. Perform cosine similarity calculation on the semantic feature sequence and the candidate node set to obtain the node similarity corresponding to each candidate node. S53. Perform a sorting and filtering operation on the candidate node set based on node similarity to determine the target node set that matches the semantic feature sequence; S54. Based on the target node set, extract the node connection relationship, business association relationship, process association relationship and archiving association relationship corresponding to each target node in the office knowledge graph, and generate a node association sequence; S55. Perform association path expansion operation on each target node according to the node association sequence to form a candidate matching path consisting of file type nodes, business entity nodes, process nodes and archive nodes. S56. Perform path integrity verification and semantic consistency verification operations on the candidate matching paths to obtain the graph matching path set.

[0033] In this embodiment, the construction process of the office knowledge graph specifically includes: Collect historical office document data, historical workflow data, historical archive data, and organizational permission data, and process them in a unified format to generate a graph to construct a dataset; Perform text extraction, field segmentation, and semantic annotation operations on the graph construction dataset to generate a graph semantic data sequence; Extract file types, business entities, process nodes, and archive nodes from the graph semantic data sequence to generate a graph node set; Based on historical process flow data, historical archive data, and organizational permission data, extract the node connection relationships, business relationships, process relationships, and archive relationships between each graph node to generate a graph relationship set; An office knowledge graph is constructed based on the graph node set and the graph relationship set.

[0034] In this embodiment, S6 specifically includes: S61. Extract file type nodes, business entity nodes, process nodes and archive nodes from the graph matching path set, and generate a file processing sequence according to the path connection order. S62. Determine the file category corresponding to each office file based on the file type node in the file processing sequence, and generate file classification results; S63. Determine the archiving directory, archiving level and archiving identifier of each office document based on the archiving node in the document processing sequence, and generate the classification archiving results; S64. Based on the process nodes in the document processing sequence, determine the corresponding flow nodes, flow order and processing responsibility entities for each office document, and generate the process flow path; S65. Perform workflow status registration, processing node binding, and responsible entity binding operations on each office document according to the workflow path, and generate workflow results; S66. Link and record the document classification results, classification archiving results, and process flow results to form processing feedback data.

[0035] In this embodiment, S7 specifically includes: S71. Extract the file classification results, classification archiving results and process flow results from the processed feedback data, and generate a feedback sample sequence; S72. Associate the feedback sample sequence with the semantic label sequence and semantic feature sequence to generate model update samples; S73. Perform parameter update operations on the BERT model based on the model update samples; S74. Extract the matching relationship between file type nodes, business entity nodes, process nodes and archive nodes based on the feedback sample sequence, and generate the graph update relationship; S75. Perform node relationship update operation on the office knowledge graph based on the graph update relationship.

[0036] refer to Figure 3 An office document processing system based on big data and deep learning, comprising: The data acquisition module is used to collect and preprocess text data from multiple sources of office documents to generate a sequence of office texts. The syntax analysis module is used to perform syntax analysis on the sequence of office texts, extract the syntactic dependencies and lexical structure information of each office text, and perform structure annotation on the sequence of office texts based on the analysis results, identify field boundaries and grammatical roles, and form a document structure sequence. The label matching module is used to perform contextual semantic modeling operations on the file structure sequence through a pre-trained BERT model to form a semantic representation sequence. Based on a preset semantic label set, a label matching operation is performed on the semantic representation sequence to obtain a semantic label sequence. The feature extraction module is used to perform field merging operations on the file structure sequence based on the semantic label sequence, and to perform field semantic association calculations on the merging results in the BERT model to extract the file topic, business entities and key fields to form a semantic feature sequence. The association matching module is used to perform node similarity calculation and target node matching operations in a pre-built office knowledge graph based on semantic feature sequences, and to expand the association paths of each target node to generate a graph matching path set; The document processing module is used to perform classification, archiving, and workflow operations on each office document based on the path set matched by the graph, and to record the classification, archiving, and workflow results to form processing feedback data. The feedback update module is used to update the parameters of the BERT model and the knowledge graph based on the processed feedback data.

[0037] Example 1: To verify the feasibility of this invention in practice, it was applied to a daily office document processing scenario within a large organization. In this scenario, office documents originate from multiple business process nodes and storage locations, and the document types include reports, approvals, contracts, and notices. Different documents differ in format structure, language expression, and field organization. Traditional processing methods mainly rely on manual classification and rule matching. With the continuous increase in the number of documents, problems arise such as inconsistent classification results, incomplete extraction of key fields, and deviations in the selection of workflow paths, leading to a decrease in overall processing efficiency and an increased burden of manual review.

[0038] In this scenario, text extraction and preprocessing operations are first performed on multi-source office documents to convert different format files into a unified office text sequence. Noisy characters, duplicate content, and abnormal paragraphs in the text are cleaned up to ensure that the basic data for subsequent analysis remains consistent.

[0039] Subsequently, the office text sequence is subjected to syntactic analysis. Through syntactic dependency identification and lexical structure analysis, field boundaries, syntactic roles and paragraph structures in the text are marked, thereby converting unstructured text into a file structure sequence with clear structural features.

[0040] Based on this, the pre-trained BERT model is used to perform contextual semantic modeling on the file structure sequence, so that the semantic relationship of different fields in the overall context can be accurately expressed, and the label matching is completed based on the semantic label set to form a semantic label sequence.

[0041] After completing semantic modeling, fields with the same semantic labels and dependencies are merged, and the file topic, business entity and key fields are extracted by combining contextual semantic association calculations to generate a semantic feature sequence.

[0042] Subsequently, the semantic feature sequence is mapped to a pre-built office knowledge graph. Through node matching and path expansion processing, a graph matching path set is formed, enabling the document content to form a stable correspondence with business entities, process nodes, and archive structures.

[0043] Based on the path set of the knowledge graph, files are automatically classified, archived, and processed, maintaining semantic consistency throughout the process while minimizing manual intervention. After processing, the classification and processing results are used as feedback data to update the BERT model parameters and knowledge graph structure, thereby achieving continuous optimization.

[0044] In this implementation process, multiple batches of office documents were selected for comparative testing. The method of this invention was compared and analyzed with traditional keyword matching methods and single deep learning classification methods, and the following data results were obtained: Table 1 Comparison of Office Document Processing Effectiveness

[0045] As shown in Table 1, the method of this invention has good results in document classification, field extraction, process matching, and processing efficiency. The document classification accuracy has increased from 72.4% of the traditional rule-based method to 94.8%, and the key field extraction completeness rate has increased from 68.9% to 93.5%. This indicates that the present invention enhances the ability to identify the content structure and semantic relationships of office documents by combining syntactic analysis with BERT semantic modeling.

[0046] In terms of workflow matching, the matching accuracy of the method of this invention reaches 92.6%, which is higher than the 65.3% of the traditional rule method and the 80.5% of the single deep learning method. In addition, the archiving path matching accuracy of the method of this invention reaches 93.9%, which is higher than the traditional rule method and the single deep learning method. It can be seen that after node matching and associated path expansion based on the office knowledge graph, the correspondence between the document processing results and the actual business process is more accurate.

[0047] In terms of processing efficiency, the average processing time per file of the method of this invention is 1.9 seconds, which is lower than the 5.6 seconds of the traditional rule method and the 3.2 seconds of the single deep learning method. The proportion of manual intervention is also reduced to 8.3%. At the same time, the semantic conflict recognition rate of the method of this invention reaches 91.7%, and the consistency of processing results reaches 94.3%, indicating that the invention improves the processing efficiency while also enhancing the stability of file processing results.

[0048] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An office document processing method based on big data and deep learning, characterized in that, Includes the following steps: S1. Collect and preprocess text data from multiple sources of office documents to generate a sequence of office texts; S2. Perform syntactic analysis on the office text sequence to extract the syntactic dependencies and lexical structure information of each office text, and perform structural annotation on the office text sequence based on the analysis results to identify field boundaries and syntactic roles, forming a file structure sequence; S3. Perform contextual semantic modeling on the file structure sequence using a pre-trained BERT model to form a semantic representation sequence. Perform label matching on the semantic representation sequence based on a preset semantic label set to obtain a semantic label sequence. S4. Perform field merging operation on the file structure sequence based on the semantic label sequence, and perform field semantic association calculation on the merging result in the BERT model to extract the file topic, business entity and key fields to form a semantic feature sequence. S5. Based on semantic feature sequences, perform node similarity calculation and target node matching operations in the pre-constructed office knowledge graph, and expand the associated paths of each target node to generate a graph matching path set; S6. Based on the path set matched by the graph, perform classification, archiving and workflow operations on each office document, and record the classification, archiving and workflow results to form processing feedback data; S7. Update the parameters of the BERT model and the office knowledge graph based on the processed feedback data.

2. The office document processing method based on big data and deep learning according to claim 1, characterized in that, The text data of the multi-source office documents represents the file data to be processed from different storage locations and business process nodes. The preprocessing includes format unification, text extraction, noise character removal, duplicate text removal, paragraph ordering, and metadata standardization.

3. The office document processing method based on big data and deep learning according to claim 1, characterized in that, S2 specifically includes: S21. Perform sentence segmentation and word segmentation operations on each office text in the office text sequence to generate an office sentence sequence and a word sequence. S22. Perform part-of-speech tagging and lexical structure recognition on the word sequence to extract noun phrases, verb phrases, time phrases, numerical phrases and business phrases to form lexical structure information; S23. Perform syntactic dependency analysis based on the office clause sequence and lexical structure information to identify subject-predicate relations, verb-object relations, attributive-head relations, adverbial-head relations and coordinate relations, and form syntactic dependency relations. S24. Perform structural annotation operations on the office text sequence based on syntactic dependencies and lexical structure information, marking the title segment, body segment, table segment, signature segment, and attachment segment to obtain the structurally annotated sequence; S25. Based on the structural annotation sequence, identify the start position, end position and paragraph to which the field belongs in each office document, and generate a field boundary sequence; S26. Perform syntactic role labeling on the content of each field within the field boundary sequence based on syntactic dependencies to determine the subject role, action role, object role, time role, numerical role, and responsibility role corresponding to the field content; S27. Sequentially associate the field boundary sequence, syntax role annotation results, syntactic dependencies, and lexical structure information to form a file structure sequence.

4. The office document processing method based on big data and deep learning according to claim 3, characterized in that, S26 specifically includes: S261. Extract the field start position, field end position, field paragraph and field text content corresponding to each field content within the field boundary sequence, and generate a field content sequence; S262. Based on the start and end positions of the fields, map each field in the field content sequence to a word sequence to determine the field word set corresponding to each field. S263. Based on syntactic dependency relations, perform dependency path extraction operations on each word in the field word set to obtain the word dependency path set between each adjacent word; S264. Identify the lexical type of each lexical based on the lexical dependency path set, and map the identification results to the corresponding fields to generate field dependency features. The lexical types include head words, modifiers, governing words, and governed words. S265. Based on field dependency features, perform role candidate labeling operations on each field to obtain candidate roles for topic, action, object, time, numerical, and responsibility. S266. Based on the paragraph to which the field belongs, lexical structure information, and field dependency features, perform consistency verification operations on each role candidate and generate role verification results; S267. Based on the role verification results, perform role confirmation operations on each field to determine the subject role, action role, object role, time role, numerical role, and responsibility role corresponding to the field content, and bind them to the corresponding fields to generate grammatical role annotation results.

5. The office document processing method based on big data and deep learning according to claim 1, characterized in that, S3 specifically includes: S31. Perform positional encoding on the file structure sequence to generate a structure encoding sequence; S32. Perform context semantic modeling operations on the structure encoding sequence using a pre-trained BERT model to generate a context semantic sequence; S33. Based on a preset semantic tag set, perform tag matching operation on the context semantic sequence, and combine it with the field boundary sequence to obtain the candidate tag sequence corresponding to each field; S34. Based on the grammatical role annotation results and syntactic dependencies, perform consistency filtering on each candidate tag sequence to determine the target semantic tag corresponding to each field; S35. Bind each target semantic tag to the corresponding field according to the field order in the file structure sequence to obtain the semantic tag sequence.

6. The office document processing method based on big data and deep learning according to claim 5, characterized in that, The pre-training process of the BERT model specifically includes: Collect text data from historical office documents and perform operations such as format standardization, noise character removal, duplicate text removal, and paragraph order reordering to generate a pre-trained text sequence; Sentence segmentation and word segmentation are performed on the pre-trained text sequence to extract sentence boundaries and generate word sequence. Each word in the word sequence is uniquely identified and encoded to generate a word encoding sequence. Based on the sequential position of each word in the word encoding sequence in the pre-trained text sequence, position encoding operation is performed on each word, and sentence segment identification encoding operation is performed on each word according to the sentence boundary to generate a pre-trained sequence; According to the preset mask ratio, target words are randomly selected from the pre-training sequence, and the target words are replaced with preset mask markers to generate a mask sequence. The BERT model then performs prediction calculations on the original words corresponding to the target words based on the mask sequence. The prediction bias is obtained by performing a difference analysis between the prediction calculation results and the pre-trained sequence. Sentence pairs are extracted from the pre-trained text sequence based on sentence boundaries. Sentence pairs consisting of adjacent sentences are marked as positive samples, and sentence pairs consisting of non-adjacent sentences are marked as negative samples, forming a sentence pair sequence. In the BERT model, the inter-sentence relation determination is performed on the sentence pair sequence, and the relation determination result is binary-labeled according to the positive and negative sample labels to obtain the determination bias. Based on the prediction bias and decision bias, the BERT model is updated with parameters, and the random mask, prediction calculation, difference analysis, positive and negative sample construction, inter-sentence relationship determination, decision bias calculation and parameter update operations are performed iteratively until the prediction bias and decision bias reach the preset convergence condition, thus obtaining the pre-trained BERT model.

7. The office document processing method based on big data and deep learning according to claim 1, characterized in that, S4 specifically includes: S41. Perform tag aggregation operation on each field in the file structure sequence according to the semantic tag sequence, extract the set of fields with consistent tags, and perform order constraint processing on the field set in combination with the positional relationship of the fields in the file structure sequence to generate a field aggregation sequence. S42. Perform a field merging operation on each field set in the field aggregation sequence, merging fields with the same semantic label and syntactic dependency relationship into field units, and generating a field merging sequence. S43. Perform contextual semantic association calculation on the field merge sequence in the BERT model, extract the semantic association strength between each field unit, and construct the field association sequence; S44. Filter multiple representative field units that represent the core content of each file, and perform topic identification operation on each representative field unit according to the field association sequence and syntax role labeling results to generate a set of file topics; S45. Based on the field association sequence and syntactic dependency relationship, perform business entity identification operation on each field unit to generate a business entity set; S46. Based on the semantic tag sequence, the grammatical role annotation results and the field association sequence, perform key field filtering operations on each field unit in the field merging sequence to obtain a set of key fields. S47. Combine and associate the file topic set, business entity set, and key field set according to the order in the file structure sequence to generate a semantic feature sequence.

8. The office document processing method based on big data and deep learning according to claim 1, characterized in that, S5 specifically includes: S51. Based on semantic feature sequences, perform node matching and retrieval operations in the pre-constructed office knowledge graph according to file topic, business entity and key fields, extract the corresponding file type nodes, business entity nodes, process nodes and archive nodes, and generate a candidate node set. S52. Perform cosine similarity calculation on the semantic feature sequence and the candidate node set to obtain the node similarity corresponding to each candidate node. S53. Perform a sorting and filtering operation on the candidate node set based on node similarity to determine the target node set that matches the semantic feature sequence; S54. Based on the target node set, extract the node connection relationship, business association relationship, process association relationship and archiving association relationship corresponding to each target node in the office knowledge graph, and generate a node association sequence; S55. Perform association path expansion operation on each target node according to the node association sequence to form a candidate matching path consisting of file type nodes, business entity nodes, process nodes and archive nodes. S56. Perform path integrity verification and semantic consistency verification operations on the candidate matching paths to obtain the graph matching path set.

9. The office document processing method based on big data and deep learning according to claim 1, characterized in that, S6 specifically includes: S61. Extract file type nodes, business entity nodes, process nodes and archive nodes from the graph matching path set, and generate a file processing sequence according to the path connection order. S62. Determine the file category corresponding to each office file based on the file type node in the file processing sequence, and generate file classification results; S63. Determine the archiving directory, archiving level and archiving identifier of each office document based on the archiving node in the document processing sequence, and generate the classification archiving results; S64. Based on the process nodes in the document processing sequence, determine the corresponding flow nodes, flow order and processing responsibility entities for each office document, and generate the process flow path; S65. Perform workflow status registration, processing node binding, and responsible entity binding operations on each office document according to the workflow path, and generate workflow results; S66. Link and record the document classification results, classification archiving results, and process flow results to form processing feedback data.

10. An office document processing system based on big data and deep learning, executing the office document processing method based on big data and deep learning as described in any one of claims 1 to 9, characterized in that, include: The data acquisition module is used to collect and preprocess text data from multiple sources of office documents to generate a sequence of office texts. The syntax analysis module is used to perform syntax analysis on the sequence of office texts, extract the syntactic dependencies and lexical structure information of each office text, and perform structure annotation on the sequence of office texts based on the analysis results, identify field boundaries and grammatical roles, and form a document structure sequence. The label matching module is used to perform contextual semantic modeling operations on the file structure sequence through a pre-trained BERT model to form a semantic representation sequence. Based on a preset semantic label set, a label matching operation is performed on the semantic representation sequence to obtain a semantic label sequence. The feature extraction module is used to perform field merging operations on the file structure sequence based on the semantic label sequence, and to perform field semantic association calculations on the merging results in the BERT model to extract the file topic, business entities and key fields to form a semantic feature sequence. The association matching module is used to perform node similarity calculation and target node matching operations in a pre-built office knowledge graph based on semantic feature sequences, and to expand the association paths of each target node to generate a graph matching path set; The document processing module is used to perform classification, archiving and workflow operations on each office document according to the path set matched by the graph, and record the classification and archiving results and workflow results to form processing feedback data; The feedback update module is used to update the parameters of the BERT model and the knowledge graph based on the processed feedback data.

Citation Information

Patent Citations

  • Heterogeneous document structured data extraction system and method based on multi-modal fusion

    CN121658894A

  • Knowledge graph-based context organization method and system

    CN121958573A