An enterprise-level intelligent text proofing system

By utilizing an enterprise-level intelligent text proofreading system and employing a multi-path mutual proofreading model and text context recognition technology, efficient and accurate knowledge text proofreading and deduplication are achieved. This solves the problems of low proofreading efficiency, insufficient accuracy, and poor adaptability in enterprise knowledge management databases, thereby improving the quality and efficiency of knowledge management.

CN120975043BActive Publication Date: 2026-03-31CHINA ORDINS GRP CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies in enterprise knowledge management repositories suffer from low proofreading efficiency, insufficient proofreading accuracy, and poor adaptability, making it difficult to effectively guarantee the quality of knowledge texts and affecting the efficiency and effectiveness of enterprise knowledge management.

Method used

An enterprise-level intelligent text proofreading system is adopted, including a text input and retrieval module, a text context recognition module, a custom rule configuration module, an intelligent text proofreading module, a proofreading result output module, a text deduplication module, and a knowledge update reminder module. It uses a multi-path mutual proofreading model and a text context recognition model for automated proofreading and deduplication. Combined with an enterprise knowledge graph and custom rules, it achieves word-by-word weighted voting and real-time monitoring.

Benefits of technology

It improves the efficiency of knowledge text proofreading, ensures the accuracy and professionalism of proofreading, adapts to the constantly changing and updated text types and business contexts in the enterprise knowledge management base, protects the originality of knowledge assets, and enhances the efficiency of enterprise decision support and business collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975043B_ABST
    Figure CN120975043B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of enterprise-level intelligent text proofreading system, belong to enterprise data management technical field, the technical problem that the quality of the text of the existing method enterprise knowledge management library is difficult to be effectively guaranteed in warehouse is solved.Enterprise knowledge management library is input and retrieved module, for receiving text to be proofread and supporting multi-format input and draft temporary storage;Text context recognition module, utilize the text context recognition model of well-trained identification context category to which the text to be proofread belongs;Custom rule configuration module, for custom rule and dynamically loaded to intelligent text proofreading module;Intelligent text proofreading module, utilize the well-trained multi-pass mutual proofreading model, the context category and custom rule, proofread the text to be proofread, obtain proofreading result;Proofreading result output module, for the word-by-word weighted voting of proofreading result, generate final proofreading text.The high-quality content maintenance of the knowledge text of enterprise knowledge management library is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise data management technology, and in particular to an enterprise-level intelligent text proofreading system. Background Technology

[0002] With the advancement of enterprise informatization, knowledge management repositories have become an important platform for enterprises to store, share, and manage knowledge assets. However, when employees store knowledge texts in the knowledge management repository, the problem of excessive errors is becoming increasingly prominent, seriously affecting the quality of the knowledge management repository and the accuracy of enterprise knowledge transfer.

[0003] Currently, when employees store knowledge text in the knowledge management repository, companies mainly rely on the following proofreading methods:

[0004] Manual proofreading: Companies assign dedicated personnel to manually review and proofread submitted texts. While this method can uncover some deep-seated logical and professional knowledge errors, it has significant drawbacks. Firstly, manual proofreading is inefficient; given the massive amounts of text data available to companies, proofreaders need to expend considerable time and energy, making it difficult to meet the company's need for rapid knowledge updates. Secondly, manual proofreading is highly susceptible to the professional knowledge level and fatigue of the personnel, making it prone to omissions and misjudgments, resulting in some erroneous text entering the knowledge management database.

[0005] General-purpose text proofreading software: Some companies use existing general-purpose text proofreading software on the market to assist in proofreading. This software is typically based on common language rules and dictionaries, and can detect and correct common grammatical errors, spelling mistakes, etc. However, the text in a company's knowledge management repository often possesses high levels of professionalism and specific domain knowledge. General-purpose proofreading software, lacking an understanding of professional knowledge and adaptation to specific domain rules, struggles to accurately identify and correct spelling errors in technical terms, industry-standard formatting errors, and content errors that contradict the company's business logic, resulting in low proofreading accuracy.

[0006] Proofreading Models Based on Traditional Machine Learning: Some advanced enterprises have attempted to build text proofreading models using traditional machine learning algorithms. These models, through learning from large amounts of labeled text data, can discover potential error patterns in the text. However, traditional machine learning models rely on manual design for data feature extraction, making it difficult to fully capture the complex features of text, especially when dealing with diverse expressions and complex semantic relationships in corporate texts, thus limiting proofreading effectiveness. Moreover, traditional machine learning models have poor adaptability to new text types or contexts, requiring extensive feature engineering and model retraining, making it difficult to meet the ever-changing and expanding needs of enterprise knowledge management bases.

[0007] Existing technologies have the following main drawbacks in addressing the problem of excessive errors when employees store text in knowledge management repositories:

[0008] The proofreading process is inefficient and cannot meet the requirements of enterprises for rapid import of large-scale texts into the database.

[0009] Insufficient proofreading accuracy makes it difficult to handle the professionalism and domain-specificity of enterprise texts, resulting in a large number of erroneous texts entering the knowledge management database;

[0010] It has poor adaptability and cannot adapt well to the constantly changing and updated text types and text contexts in the enterprise knowledge management base.

[0011] These deficiencies make it difficult to effectively guarantee the quality of enterprise knowledge management bases, affecting the efficiency and effectiveness of enterprise knowledge management, and consequently having an adverse impact on enterprise decision support, employee training, and business collaboration. Summary of the Invention

[0012] Based on the above analysis, the embodiments of the present invention aim to provide an enterprise-level intelligent text proofreading system to solve the technical problem that the quality of text entering the enterprise knowledge management base is difficult to effectively guarantee using existing methods.

[0013] The objective of this invention is mainly achieved through the following technical solutions:

[0014] This invention provides an enterprise-level intelligent text proofreading system, comprising:

[0015] The text input and retrieval module is used to receive text to be proofread and supports multi-format input and draft saving.

[0016] The text context recognition module uses a trained text context recognition model to identify the context category to which the text to be proofread belongs;

[0017] The custom rule configuration module is used to define custom rules and dynamically load them into the intelligent text proofreading module;

[0018] The intelligent text proofreading module uses a trained multi-path mutual proofreading model, the context category, and custom rules to proofread the text to be proofread and obtain the proofreading result.

[0019] The proofreading result output module is used to perform word-by-word weighted voting on the proofreading results to generate the final proofreading text.

[0020] Furthermore, the system also includes:

[0021] The text deduplication module is used to compare the final proofread text with existing texts in the enterprise knowledge management database at the semantic level to obtain the deduplication comparison result; if the deduplication comparison result is lower than the preset similarity threshold, it is added to the database.

[0022] The data storage module is used to store the text to be proofread, context category, proofreading results, final proofread text, and plagiarism check results;

[0023] The knowledge update reminder module is used to monitor text updates in the enterprise knowledge management repository and automatically push text entry review reminders.

[0024] Furthermore, the trained text context recognition model is obtained through the following training process;

[0025] Construct a text context training sample set consisting of multiple context categories, where the sample label is the context category to which the text belongs;

[0026] A text context recognition model is constructed, which sequentially includes an input layer, a contextual bidirectional convolutional layer, a max pooling layer, and a fully connected classification layer. The text context recognition model is trained using the text context training sample set to obtain a trained text context recognition model.

[0027] Furthermore, the input layer is used to vectorize the text to be proofread, thereby obtaining the word vector matrix corresponding to the text to be proofread;

[0028] The contextual bidirectional convolutional layer is used to convolve the word vector matrix using a forward convolution kernel to obtain a forward feature map; to convolve the inverse matrix of the word vector matrix using a backward convolution kernel to obtain a backward feature map; to obtain gating parameters based on the forward and backward feature maps; and to obtain a fused feature map based on the gating parameters.

[0029] The max pooling layer is used to generate a corresponding fixed-length feature vector by sliding a window through the fused feature map output by each convolutional kernel and selecting the maximum value within the window.

[0030] The fully connected classification layer is used to concatenate the fixed-length feature vectors corresponding to each convolutional kernel into a global feature vector, perform fully connected classification on the global feature vector, output the context classification probability, and obtain the context classification result based on the context classification probability.

[0031] Furthermore, a multi-path cross-checking model for parallel coding is constructed, including parallel first, second, and third cross-checking paths;

[0032] The first cross-proofing path includes, in sequence, a first encoder, a first cross-model attention layer, and a T5 decoder, which are used to proofread the text to be proofread and obtain the first proofreading result.

[0033] The second cross-verification pathway includes, in sequence, a second encoder, a second cross-model attention layer, and a Seq2Seq decoder, which are used to verify the text to be verified and obtain the second verification result.

[0034] The third cross-proofing pathway includes, in sequence, a third encoder, a third cross-model attention layer, and a BERT classifier, which are used to proofread the text to be proofread and obtain the third proofreading result.

[0035] The first encoder includes a parallel BERT encoder and a Seq2Seq encoder; the second encoder includes a parallel BERT encoder and a T5 encoder; and the third encoder includes a parallel Seq2Seq encoder and a T5 encoder.

[0036] Furthermore, the multi-path cross-calibration model is trained through a specific process to obtain a trained multi-path cross-calibration model.

[0037] Construct a specific text proofreading training sample set; the specific text proofreading training sample set includes a general text context training sample set and multiple specific context category training sample sets;

[0038] The multi-path cross-checking model based on parallel coding is pre-trained using a general text context training sample set, and the pre-trained parameters of the multi-path cross-checking model based on parallel coding are saved as initial parameters.

[0039] The initial parameters are loaded, and the training sample set for each specific context category is trained sequentially. The model parameters are iteratively optimized through forward and backward propagation until the loss function converges, thus obtaining the multi-path cross-calibration model parameters based on parallel encoding for multiple specific context categories.

[0040] Furthermore, the proofreading result output module, based on the three sets of proofreading results obtained from the first, second, and third mutual proofreading paths of the intelligent text proofreading module, performs word-by-word voting using dynamic weights, including:

[0041] For each position of the text in the first, second, and third proofreading results, a dynamic weighted word-by-word voting process is conducted, and the text that receives the most weighted votes is selected as the final proofread text.

[0042] Output the voting results for each position of the text in the original order of the text to be proofread, and obtain the proofread text corresponding to the text to be proofread.

[0043] Furthermore, the workflow of the text plagiarism detection module includes:

[0044] The final proofread text is preprocessed to obtain the preprocessed final proofread text;

[0045] Word2Vec is used to convert words in the preprocessed final proofreading text into vector representations; based on the vector representations of words, the text vectors of the final proofreading text are obtained.

[0046] The text vector is compared with the text vector in the knowledge management database of the plagiarism detection company, and the similarity value is calculated as the plagiarism detection comparison result;

[0047] If the plagiarism comparison result is lower than the preset similarity threshold, the text to be proofread will be placed in the pending database entry state.

[0048] Furthermore, the workflow of the knowledge update reminder module includes:

[0049] Scan the knowledge management database according to a preset time cycle, compare text timestamps, version numbers or hash values, and identify whether there are text updates or additions;

[0050] Based on knowledge association rules, identify texts that are related to the updated or added text; wherein, the association includes direct references, indirect associations, or texts on the same topic;

[0051] Integrate key changes to the updated or added text with related text information, and define reminder messages for multi-channel push notifications.

[0052] Furthermore, custom rules for proofreading text can be flexibly configured based on business needs, including text-based keywords, key words, format features, context categories, knowledge graph association rules, as well as user permissions and department permissions.

[0053] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0054] 1. This invention adopts a multi-module collaborative working approach, such as the convenient operation of the text input and retrieval module, the automated processing of the intelligent text proofreading module, and the word-by-word voting mutual proofreading algorithm, which greatly reduces the workload and time cost of manual proofreading, and can quickly proofread a large number of knowledge texts to be added to the database, meeting the needs of enterprises for rapid knowledge updates and large-scale knowledge text processing, and improving the efficiency of knowledge text proofreading.

[0055] 2. This invention accurately identifies the context of a text through a text context recognition module, and combines it with a proofreading method based on multi-path mutual proofreading model fusion and context self-adaptation in an intelligent text proofreading module, as well as proofreading of professional terms, industry standard formats, etc. It fully considers the professionalism and domain specificity of enterprise texts, effectively solves the problem of insufficient proofreading accuracy in existing methods, ensures the accuracy and professionalism of knowledge text content, and improves proofreading accuracy.

[0056] 3. The custom rule configuration module in this invention allows enterprises to flexibly configure proofreading rules according to their own business characteristics, industry standards and text writing habits. At the same time, the intelligent text proofreading module can adaptively adjust the proofreading strategy for different text contexts, so that the system can better adapt to the constantly changing and updated knowledge text types and business contexts in the enterprise knowledge management base, and has universality and scalability, enhancing the system's adaptability.

[0057] 4. The text plagiarism detection module in this invention utilizes advanced text similarity detection algorithms to compare the text to be proofread with existing knowledge text content in the enterprise's knowledge management database. It employs cosine similarity and knowledge graph feature comparison to identify semantically similar content (such as synonym replacement and word order adjustment), quickly identifying potentially plagiarized or over-quoted parts and generating a plagiarism report. This helps enterprises protect the originality of their intellectual property assets, avoid intellectual property issues caused by text duplication, and safeguard the originality of knowledge.

[0058] 5. The knowledge update reminder module in this invention can monitor the dynamics of the enterprise knowledge management base in real time. When a knowledge update is detected, it promptly notifies relevant personnel to review and update the knowledge text to be added to the base and related texts, thereby ensuring the timeliness and effectiveness of the knowledge text in the knowledge management base and improving the efficiency and quality of enterprise decision support, employee training and business collaboration.

[0059] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description

[0060] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0061] Figure 1 This is a schematic diagram of an enterprise-level intelligent text proofreading system module in an embodiment of the present invention;

[0062] Figure 2 This is a schematic diagram of the front-end and back-end separated architecture for knowledge text input, storage, and retrieval in an embodiment of the present invention;

[0063] Figure 3 This is a schematic diagram showing the output of the multi-path mutual verification model and the word-by-word voting mutual verification algorithm in an embodiment of the present invention;

[0064] Figure 4 This is a schematic diagram illustrating how enterprise users store and query the enterprise knowledge management base in an embodiment of the present invention;

[0065] Figure 5 This is a flowchart illustrating the process of an enterprise user applying an enterprise-level intelligent text proofreading system using an enterprise knowledge management base, as described in this embodiment of the invention. Detailed Implementation

[0066] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0067] To address the aforementioned problems, this invention proposes an enterprise-level intelligent text proofreading system, such as... Figure 1 As shown, it includes:

[0068] The text input and retrieval module is used to receive text to be proofread and supports multi-format input and draft saving.

[0069] The text context recognition module uses a trained text context recognition model to identify the context category to which the text to be proofread belongs;

[0070] The custom rule configuration module is used to define custom rules and dynamically load them into the intelligent text proofreading module;

[0071] The intelligent text proofreading module uses a trained multi-path mutual proofreading model, the context category, and custom rules to proofread the text to be proofread and obtain the proofreading result.

[0072] The proofreading result output module is used to perform word-by-word weighted voting on the proofreading results to generate the final proofreading text.

[0073] The system also includes:

[0074] The text deduplication module is used to compare the final proofread text with existing texts in the enterprise knowledge management database at the semantic level to obtain the deduplication comparison result; if the deduplication comparison result is lower than the preset similarity threshold, it is added to the database.

[0075] The data storage module is used to store the text to be proofread, context category, proofreading results, final proofread text, and plagiarism check results;

[0076] The knowledge update reminder module is used to monitor text updates in the enterprise knowledge management repository and automatically push text entry review reminders.

[0077] The following is a description of each module:

[0078] M1, the text input and retrieval module, specifically.

[0079] Enterprise users can upload text files to be imported into the system through the text input and retrieval module interface, and import them into the system for temporary storage.

[0080] Alternatively, enterprise users can directly edit and write their texts online in the text editor of the text input and retrieval module and then save the draft.

[0081] This module also allows for preliminary searches to determine whether the knowledge text to be added to the enterprise knowledge management database already exists. If it does not exist, subsequent addition operations are performed, effectively avoiding duplication of knowledge text.

[0082] like Figure 2 As shown, this system adopts a front-end and back-end separation architecture, and implements input, retrieval, and draft management functions through the collaborative efforts of business logic layer components. It supports multi-format text input and parsing.

[0083] Text file upload supports file types, for example, Word (docx, doc, xls, xlsx, ppt, pptx), PDF, txt, CSV, Markdown, etc.

[0084] Online editing, for example, supports real-time preview of rich text (HTML) and Markdown.

[0085] Text parsing:

[0086] (1) Integration of domestic document parsing engine: Integrating Yongzhong Office SDK to parse Word / PDF content and extract plain text and structured data (such as tables and titles);

[0087] (2) Alternatively, an open-source parsing library can be integrated: Apache Tika (optional) can be used as a backup solution. Docker containerization can be used to ensure that the open-source parsing library is isolated from the domestic environment.

[0088] Performance optimization:

[0089] Large file chunked upload (based on HTTP Range protocol), with a real-time upload progress bar displayed on the front-end interface;

[0090] Asynchronous text parsing task queue (using RocketMQ / Kafka) avoids blocking the main thread.

[0091] Search result filtering: Combines custom business rules of the enterprise knowledge management base (such as department permissions, personnel permissions, file classification, etc.) to search for knowledge text content that enterprise users who have uploaded knowledge text have permission to access.

[0092] Draft saving: Before the text content to be added to the enterprise knowledge management repository is officially submitted, a draft is temporarily saved.

[0093] Metadata storage: Draft content (including editing history) is stored in MinIO in JSON format. MinIO is an open-source, high-performance, distributed object storage system that supports ARM architecture and domestic operating systems (such as Kylin OS).

[0094] MinIO's cross-platform nature allows it to adapt to different operating systems and hardware environments. By using MinIO, enterprises can achieve high-performance object storage for their knowledge management repositories in private cloud, hybrid cloud, or edge computing environments.

[0095] Version control: Each save generates an incremental snapshot, supporting version rollback by time; supports multi-version concurrent control to avoid editing conflicts.

[0096] The knowledge text to be added to the database needs to be proofread before it is officially added to the database. It will then be used as proofread text for subsequent modules.

[0097] M2, the text context recognition module, specifically.

[0098] The text context recognition module uses a trained text context recognition model to identify the context category to which the text to be proofread belongs.

[0099] The input text context recognition module uses deep learning algorithms to automatically and accurately identify the context of the text, such as technical white papers and technical manuals, laying the foundation for subsequent targeted intelligent proofreading and effectively improving proofreading accuracy.

[0100] First, the text to be proofread is subjected to text context recognition, and knowledge text proofreading with different proofreading rules is performed for specific contexts, laying the foundation for solving the insufficient proofreading accuracy in existing technologies.

[0101] The trained text context recognition model is obtained through the following training process;

[0102] Construct a text context training sample set consisting of multiple context categories, where the sample label is the context category to which the text belongs;

[0103] A text context recognition model is constructed, which sequentially includes an input layer, a contextual bidirectional convolutional layer, a max pooling layer, and a fully connected classification layer. The text context recognition model is trained using the text context training sample set to obtain a trained text context recognition model.

[0104] Enterprise users include various personnel such as finance, business, administration, human resources, legal, marketing, project management, product, technology, testing, and customer service. Therefore, the knowledge texts submitted by various enterprise users to the enterprise knowledge management base are divided into multiple knowledge text contexts.

[0105] Construct a text context training sample set consisting of multiple text context categories, where the sample label is the text context category to which the text belongs;

[0106] For example, the text context is divided into 8 major categories, each of which includes corresponding subcategories, resulting in a total of 50 text context categories. These 50 categories include 49 specific text contexts and 1 general text context. In practical applications, text context categories can be added or removed according to specific needs.

[0107] For each of the 50 text context categories, 100 texts were compiled as training samples. The sample labels are shown in Table 1 as text context categories, resulting in a text context training sample set consisting of 5000 training samples. Each text context training sample includes multiple characters and is either a long text or a paragraph.

[0108] Table 1: Text Context Classification

[0109]

[0110]

[0111]

[0112] The text context training sample set is used to train the constructed text context recognition model.

[0113] For example, a text context training sample is described: for example, a sample x of the technical manual category is formed by preprocessing a technical manual text, that is, extracting the text from the technical manual to form text x; the sample label corresponding to sample x is the text context category, namely "technical manual".

[0114] A text context recognition model is constructed, and the model is trained using a text context training sample set, so that the trained text context recognition model has the function of automatically determining the text context category to which the input text to be proofread belongs.

[0115] A text context recognition model is constructed, consisting of an input layer, a contextual bidirectional convolutional layer, a max pooling layer, and a fully connected layer.

[0116] The input layer is used to vectorize the text to be proofread, thereby obtaining the word vector matrix corresponding to the text to be proofread;

[0117] The contextual bidirectional convolutional layer is used to convolve the word vector matrix using a forward convolution kernel to obtain a forward feature map; to convolve the inverse matrix of the word vector matrix using a backward convolution kernel to obtain a backward feature map; to obtain gating parameters based on the forward and backward feature maps; and to obtain a fused feature map based on the gating parameters.

[0118] The max pooling layer is used to generate a corresponding fixed-length feature vector by sliding a window through the fused feature map output by each convolutional kernel and selecting the maximum value within the window.

[0119] The fully connected classification layer is used to concatenate the fixed-length feature vectors corresponding to each convolutional kernel into a global feature vector, perform fully connected classification on the global feature vector, output the context classification probability, and obtain the context classification result based on the context classification probability.

[0120] (1) Input layer: The text to be proofread is converted into a word vector matrix. Where n is the length of the text to be proofread, d is the word vector dimension, and e i (i = 1, 2, ..., n) is the word vector of the i-th character in the text to be proofread after being processed by the input layer;

[0121] (2) Contextual Bidirectional Convolutional Layer: The contextual bidirectional convolutional layer uses convolutional kernels W1 and W2, where W1 is the forward convolutional kernel and W2 is the inverse convolutional kernel. The forward convolutional kernel W1 is used to convolve the word vector matrix E to obtain the forward feature map H1; the inverse convolutional kernel W2 is used to convolve the inverse matrix E of the input layer matrix E. -1 Convolution is performed to obtain the inverse feature map H2.

[0122] The forward feature map H1 is as follows:

[0123] H1 = W1 * E Formula (1)

[0124] Here, * represents the forward convolution operation. c represents the number of output channels of the forward convolution kernel, and k represents the size of the convolution window of the forward convolution kernel.

[0125] The inverse feature map H2 is as follows:

[0126]

[0127] Based on the forward feature map H1 and the inverse feature map H2, the gating parameter g is obtained as follows:

[0128] g=σ(U g [H1; H2]) Formula (3)

[0129] Among them, U g Let g be the gating matrix, and σ be the sigmoid function. The forward feature map H1 and the inverse feature map H2 are fused using the gating parameter g to generate the fused feature map H, as follows:

[0130] H=g⊙H1+(1-g)⊙H2 Formula (4)

[0131] Where ⊙ represents the Hadamard product. Contextual bidirectional convolution solves the problem of semantic fragmentation in the context of traditional TextCNN (text classification model).

[0132] (3) Max Pooling Layer: A pooling layer selects the maximum value within a fixed-size window (e.g., a 2x2 or 3x3 window) as its output by sliding a window across the feature map. This process is called max pooling. For knowledge-based text data, this means that the max pooling layer extracts the most salient feature on the fused feature map output by each convolutional kernel.

[0133] Generating fixed-length feature vectors: Because pooling reduces the size of the feature maps, the fused feature maps output by each convolutional kernel are pooled to form a fixed-length feature vector. This feature vector contains key information about the original knowledge text within its specific context.

[0134] (4) Fully connected classification layer: The fixed-length feature vectors corresponding to each convolution kernel are concatenated into a global feature vector. The global feature vector is fully connected for classification, and the context classification probability is output. The context classification result is obtained based on the context classification probability. The category with the highest context classification probability is selected as the final context classification result.

[0135] The text context recognition model training is as follows:

[0136] The text context recognition model is trained using a text context training sample set, enabling the model to automatically determine the text context category to which the input text belongs.

[0137] Set hyperparameters for learning rate, batch size, and number of iterations;

[0138] We choose the cross-entropy loss function to measure the difference between the text context recognition model's predicted context classification results and the actual sample labels;

[0139] The Adam optimization algorithm is used to update the parameters of the text context recognition model; the model parameters are updated through backpropagation to minimize the cross-entropy loss function.

[0140] The training continues until the loss function converges or the predetermined number of iterations is reached, resulting in a well-trained text context recognition model. The parameters of the trained text context recognition model are then saved for use in subsequent text context recognition.

[0141] The trained text context recognition module takes the text to be proofread as input and outputs the text context category corresponding to the text to be proofread.

[0142] For example, the learning rate is set to 0.001; the batch size is set to 64; and the predetermined number of iterations is 100. In actual use, these settings can be changed according to specific needs.

[0143] Input the text to be proofread into the text context recognition module, which automatically identifies the text context category to which the text to be proofread belongs, so as to provide the text context category for subsequent targeted text proofreading.

[0144] M3, the intelligent text proofreading module, specifically.

[0145] The intelligent text proofreading module is the core of the system in this invention. It integrates natural language processing technology and enterprise knowledge graph to perform comprehensive text proofreading.

[0146] On the one hand, proofreading and detection are carried out through a text context recognition model based on contextual bidirectional convolution, a multi-path mutual proofreading model based on parallel encoding, and a weight-optimized word-by-word voting algorithm;

[0147] On the other hand, by leveraging enterprise expertise and custom rules for knowledge management within the knowledge graph, this module proofreads for spelling errors in technical terms, industry-standard formats, and content that contradicts the enterprise's business logic, ensuring the accuracy and professionalism of the text content. This module significantly improves proofreading accuracy and efficiency, resolving the technical problems of low proofreading efficiency and low proofreading accuracy in existing technologies.

[0148] Existing BERT models include encoders and classifiers; existing Seq2Seq models include encoders and decoders; existing T5 models include encoders and decoders.

[0149] like Figure 3 The multi-path cross-calibration model is shown.

[0150] Construct a multi-path cross-checking model for parallel coding, including parallel first, second, and third cross-checking paths;

[0151] The first cross-proofing path includes, in sequence, a first encoder, a first cross-model attention layer, and a T5 decoder, which are used to proofread the text to be proofread and obtain the first proofreading result.

[0152] The second cross-verification pathway includes, in sequence, a second encoder, a second cross-model attention layer, and a Seq2Seq decoder, which are used to verify the text to be verified and obtain the second verification result.

[0153] The third cross-proofing pathway includes, in sequence, a third encoder, a third cross-model attention layer, and a BERT classifier, which are used to proofread the text to be proofread and obtain the third proofreading result.

[0154] The first encoder includes a parallel BERT encoder and a Seq2Seq encoder; the second encoder includes a parallel BERT encoder and a T5 encoder; and the third encoder includes a parallel Seq2Seq encoder and a T5 encoder.

[0155] The first, second and third cross-checking paths are denoted as [(BERT,Seq2Seq),T5], [(BERT,T5),Seq2Seq] and [(Seq2Seq,T5),BERT], respectively.

[0156] The first encoder includes a parallel encoding BERT encoder and a Seq2Seq encoder;

[0157] The second encoder includes a parallel encoding BERT encoder and a T5 encoder;

[0158] The third encoder includes a parallel Seq2Seq encoder and a T5 encoder.

[0159] The first cross-model attention layer performs weighted fusion based on the importance of the encoding results of the BERT encoder and the Seq2Seq encoder in the first encoder to obtain the weighted first encoded feature; the T5 decoder classifies the first encoded feature to obtain the first proofreading result;

[0160] The second cross-model attention layer performs weighted fusion based on the importance of the encoding results of the BERT encoder and T5 encoder in the second encoder to obtain the weighted second encoded features; the Seq2Seq decoder classifies the second encoded features to obtain the second proofreading result;

[0161] The third cross-model attention layer performs weighted fusion based on the importance of the encoding results of the Seq2Seq encoder and the T5 encoder in the third encoder to obtain the weighted third encoding feature; the BERT classifier classifies the third encoding feature to obtain the third proofreading result.

[0162] Taking the first cross-model proofreading path [(BERT, Seq2Seq), T5] as an example, the text to be proofread is simultaneously input into the BERT encoder and the Seq2Seq encoder for parallel encoding. Then, through the first cross-model attention layer, the weights of the encoding outputs of BERT and Seq2Seq are dynamically assigned. Finally, the output of the cross-model attention layer is input into the T5 decoder for decoding to obtain the proofreading result of the text to be proofread.

[0163] (1) Extract the encoder of the BERT model and the encoder of the Seq2Seq model. The length of the text to be proofread by the two encoders is n. In order to ensure the smooth operation of matrix operations in the subsequent cross-model attention layer and to ensure dimension matching when the encoded features are weighted and fused, the hidden vector dimension of the embedding layer of the BERT encoder and the Seq2Seq encoder is designed to be d-dimensional. That is, the semantic representation of each token of the BERT encoder and the Seq2Seq encoder is designed to be a d-dimensional vector.

[0164] The text to be proofread is simultaneously input into the BERT encoder and the Seq2Seq encoder, which encode it in parallel and independently to extract features from the text. The BERT encoder output is... (Text length n, semantic vector dimension d), the Seq2Seq encoder output is:

[0165] (2) First cross-model attention layer: The significance of the first cross-model attention layer is that it can automatically determine the importance of the encoding results of the two encoders in parallel encoding based on the features of the input text, and perform weighted fusion of the output features of the two encoders based on their importance.

[0166] Independent computation of BERT attention Attn BERT and ERNIE Attention Attn Seq2Seq , respectively represented as

[0167]

[0168] in, The BERT key projection matrix maps the BERT output to attention key vectors. The BERT value projection matrix maps the BERT output to an attention value vector. The Seq2Seq key projection matrix maps the Seq2Seq output to attention key vectors. The Seq2Seq value projection matrix maps the Seq2Seq output to an attention value vector. From a matrix perspective, after matrix operations, the BERT attention... ERNIE's attention Both attention dimensions are consistent and meet the conditions for feature fusion.

[0169] Then, dynamic weight allocation is performed, and the dynamic weight λ is represented as follows:

[0170] λ=σ(W g +b g ) Formula (7)

[0171] Where σ is the Sigmoid activation function; The dynamic weight matrix is ​​used to map eigenvectors to a scalar value; b g It is a dynamic bias term used to adjust the output of the Sigmoid function.

[0172] The weighted encoded features are obtained as follows:

[0173] H weight =λ·Attn BERT +(1-λ)·Attn Seq2Seq Formula (8)

[0174] (3) Extract the decoder and classifier of the T5 model, and convert the output H across the attention layer of the model. weight The input is fed into the T5 decoder, and after decoding and classification, the first proofreading result is output.

[0175] In addition to the above [(BERT,Seq2Seq),T5] pathway, the Seq2Seq encoder and T5 encoder are encoded in parallel, and after passing through the second cross-model attention layer, the BERT classifier is used for output, resulting in the second cross-calibration pathway [(Seq2Seq,T5),BERT].

[0176] Similarly, the BERT encoder and T5 encoder are then encoded in parallel, and after passing through the third cross-model attention layer, the Seq2Seq decoder is used to output the result, resulting in the third cross-calibration path [(BERT, T5), Seq2Seq].

[0177] Since the BERT model does not have a decoder, if parallel encoding is performed using a Seq2Seq encoder and a T5 encoder, the output can be directly obtained from the BERT model's classifier.

[0178] By employing a multi-path cross-correction model, the advantages of different models can be fully utilized, and the correction bias of a single model can be corrected, thereby enhancing the language representation capabilities of the correction model. Simultaneously, it can alleviate the inherent limitations of single models, which are highly sensitive to data distribution and feature selection, making them susceptible to data noise and overfitting. Furthermore, the three cross-correction paths generate three sets of correction results, paving the way for the subsequent application of a weighted, word-by-word voting algorithm.

[0179] For parallel coding, since the language representation ability of a single model is limited and its stability and generalization ability are insufficient, this invention adopts a parallel coding method.

[0180] The multi-path cross-calibration model is trained by a certain process to obtain a trained multi-path cross-calibration model;

[0181] Construct a specific text proofreading training sample set; the specific text proofreading training sample set includes a general text context training sample set and multiple specific context category training sample sets;

[0182] The multi-path cross-checking model based on parallel coding is pre-trained using a general text context training sample set, and the pre-trained parameters of the multi-path cross-checking model based on parallel coding are saved as initial parameters.

[0183] The initial parameters are loaded, and the training sample set for each specific context category is trained sequentially. The model parameters are iteratively optimized through forward and backward propagation until the loss function converges, thus obtaining the multi-path cross-calibration model parameters based on parallel encoding for multiple specific context categories.

[0184] Construct a general text proofreading training sample set, which consists of general text sample data and their labels in a general text context;

[0185] For example, the sample data is "I'm so happy today", and the sample label is "I'm really happy today".

[0186] Using a general text proofreading training sample set and pre-training the three mutual proofreading pathways, the initial parameters of the first, second and third mutual proofreading pathways in the general text context are obtained.

[0187] Forty-nine specific text context proofreading training sample sets were constructed. The initial parameters of the three mutual proofreading pathways under the general text context were placed into these pathways, and fine-tuning training was performed for each specific text context to obtain parameters for the 49 specific text context proofreading scenarios. For the technical manual context, the technical manual text dataset and sample labels were combined to form the technical manual text proofreading training sample set. Similarly, for the 49 specific contexts excluding the general text context and contract text context, corresponding proofreading training sample sets were obtained. Therefore, a total of 49 specific text context proofreading training sample sets were obtained.

[0188] Fine-tuning training for specific text contexts. Since the training of the proofreading context recognition algorithm can identify the type of text context of the input text to be proofread, and each text type has its own unique text features, the performance of the multi-model cross-proofreading module can be optimized under specific contexts by training the neural network parameters for each specific text context.

[0189] The initial parameters of the three-path proofreading model in the general text context are fed into the three-path mutual proofreading model. For the contract proofreading context, the parameters are obtained by training the first path [(BERT,Seq2Seq),T5] using the contract proofreading dataset. Similarly, by training the second path [(BERT, T5), Seq2Seq] and the third path [(Seq2Seq, T5), BERT] using the contract verification dataset, the parameters can be obtained.

[0190] Therefore, the parameter W of the three-way mutual verification module in the context of contract verification... 合同 It can be represented as:

[0191]

[0192] Similarly, for the remaining 48 specific text proofreading contexts, the three-path mutual proofreading module is trained using the corresponding proofreading training sample set until the load loss function converges, thus obtaining the parameters of the three-path mutual proofreading model under the 11 specific text proofreading contexts.

[0193] The joint loss function L is as follows:

[0194] L = L CE +γ1·L con +γ2·L ada +γ3·L reg Formula (10)

[0195] Where γ1, γ2, and γ3 are hyperparameters; L CE L con L ada and L reg These are cross-entropy loss, multipath consistency loss, context adaptation loss, and regularization loss, respectively.

[0196] The cross-entropy loss is used to measure the difference between the predicted results and the actual results of the three pathways, as follows:

[0197]

[0198] in, Let y be the predicted probability of the i-th character in the text to be proofread by the k-th path. i The actual label for the i-th character;

[0199] The multipath consistency loss forces the output distributions of the three pathways to be similar, using symmetric KL (Kullback-Leibler divergence, relative entropy) divergence, as follows:

[0200]

[0201] The context adaptation loss is used to enhance the model's ability to adapt to specific contexts, as follows:

[0202]

[0203] Where, θ general Pre-trained parameters for general context, is the fine-tuning parameter for the s-th specific context, where S is the total number of specific context categories;

[0204] The regularization loss is L2 regularization.

[0205] For example, the values ​​of hyperparameters γ1, γ2, and γ3 need to be adjusted according to task requirements and data characteristics. The recommended initial values ​​and their design logic are shown in Table 2.

[0206] Table 2 Recommended initial values ​​for hyperparameters

[0207]

[0208] The search is performed within the ranges of γ1∈[0.3,0.8], γ2∈[0.05,0.4], and γ1∈[0.005,0.03]. In the early stages of training, γ1 and γ2 are increased to accelerate model convergence; in the later stages of training, γ1 and γ2 are decreased to unleash the model's potential.

[0209] Calculate the squared Euclidean distance between the specific context parameter and the general context parameter to measure the difference between the two.

[0210] Regularization loss is used to prevent overfitting.

[0211] For 49 contexts, 49 sets of model parameters for the three-path mutual calibration model were obtained.

[0212] The intelligent text proofreading module is trained by a multi-path mutual proofreading model based on parallel coding to obtain the parameters of the multi-path mutual proofreading model under the general text context, as well as the parameters under the 49 specific context text proofreading contexts.

[0213] M4, proofreading result output module, specifically.

[0214] The proofreading result output module is used to make decisions based on the results output by the intelligent text proofreading module, and obtain the proofread text result; this module improves the accuracy of text proofreading.

[0215] The proofreading result output module, based on the three sets of proofreading results obtained from the first, second, and third mutual proofreading paths of the intelligent text proofreading module, uses dynamic weights to perform word-by-word voting, including:

[0216] For each position of the text in the first, second, and third proofreading results, a dynamic weighted word-by-word voting process is conducted, and the text that receives the most weighted votes is selected as the final proofread text.

[0217] Output the voting results for each position of the text in the original order of the text to be proofread, and obtain the proofread text corresponding to the text to be proofread.

[0218] A word-by-word voting process, based on weighted optimization of the three sets of proofreading results obtained from the first, second, and third mutual proofreading pathways of the intelligent text proofreading module, includes:

[0219] For each position of the text in the first, second, and third proofreading results, a weighted word-by-word vote is performed, and the text with the most weighted votes is selected as the final proofreading output text. The voting results of each position are output in the original order of the text to be proofread to obtain the proofread text corresponding to the text to be proofread.

[0220] The word-by-word voting mutual proofreading algorithm selects the text with the largest weighted vote at each position in the first, second, and third proofreading results output from the three mutual proofreading paths as the proofread text, and outputs the final proofread text.

[0221] The initial weights of the three channels are all 1.0. After each vote, the weight parameters are optimized through backpropagation. For each position, the weight of the channel that is consistent with the final result is increased, and the weight of the channel that is inconsistent is decreased.

[0222] Since the basic models used in the first, second, and third cross-checking pathways of the input text context recognition module are consistent, the texts obtained by the three cross-checking pathways are usually basically the same. However, due to the superposition of defects in the model itself or the imperfect training of the pathways, there may be occasional cases where the output results of individual pathways have additional typos compared with the other pathways.

[0223] Input the text to be proofread into the text context recognition module to obtain the corresponding context category; then input the text to be proofread into the intelligent text proofreading module of the corresponding context, and the proofreading result output module outputs the final proofreading result text.

[0224] The proofreading output module integrates the first, second, and third proofreading results from the first, second, and third proofreading paths through a word-by-word voting mechanism, and selects the text with the most weighted votes at each position as the final proofreading output, thereby generating a more accurate proofreading text.

[0225] M5, the text plagiarism detection module, specifically.

[0226] The text plagiarism detection module uses advanced text similarity detection algorithms to compare the text to be proofread with existing, officially approved knowledge content in the enterprise's knowledge management database. It quickly identifies potentially plagiarized or over-quoted parts, generates a plagiarism report, and safeguards the originality of enterprise knowledge. This module effectively guarantees the originality of enterprise knowledge.

[0227] The workflow of the text plagiarism detection module includes:

[0228] The final proofread text is preprocessed to obtain the preprocessed final proofread text;

[0229] Extract features from the preprocessed final proofread text and convert it into text vectors;

[0230] Word2Vec is used to convert words in the preprocessed final proofreading text into vector representations; based on the vector representations of words, the text vectors of the final proofreading text are obtained.

[0231] If the plagiarism comparison result is lower than the preset similarity threshold, the text to be proofread will be placed in the pending database entry state.

[0232] For the proofread user-inputted text, perform deduplication processing within the enterprise knowledge management database.

[0233] The text plagiarism detection module mainly consists of the following sub-modules: data preprocessing, feature extraction and vectorization representation, similarity calculation, plagiarism database, and plagiarism result analysis and report generation. Its workflow is as follows:

[0234] (1) Data preprocessing submodule: Cleans and segments the input text, removes irrelevant characters, and extracts the core content of the text;

[0235] (2) Feature extraction and vectorization representation submodule: The preprocessed final proofread text is converted into the corresponding text vector form in order to perform similarity calculation;

[0236] (3) Similarity calculation submodule: compares the vector of the input text with the text vector in the plagiarism detection database and calculates the similarity; for example, cosine similarity is used.

[0237] For example, a similarity threshold of 20% can be set; however, the setting of the similarity threshold depends on the specific application scenario and the requirements for strictness in plagiarism detection. In specific applications, adjustments should be made according to specific needs.

[0238] (4) Deduplication Database Submodule: Stores knowledge text data and its feature vectors in the enterprise knowledge management base for use by the similarity calculation submodule;

[0239] (5) Submodule for analyzing and generating plagiarism detection results: This module analyzes the similarity calculation results, generates a detailed plagiarism detection report, identifies the duplicate parts in the text and their sources, and outputs the duplication rate.

[0240] M6, data storage module, specifically.

[0241] The data storage module securely stores the proofread text and results, while simultaneously building an enterprise text big data resource library to provide a data foundation for knowledge management;

[0242] Both drafts and final versions of the text are stored.

[0243] like Figure 4 As shown. The data storage module adopts a hybrid storage architecture to meet the storage needs of enterprise-level intelligent text proofreading systems for different types of data. It mainly includes the following storage components:

[0244] Object storage systems are used to store large amounts of unstructured text files, such as documents and reports submitted by company employees.

[0245] Relational data storage: Stores structured data, including user information, proofreading task records, system configurations, etc.

[0246] No relational data storage: suitable for storing semi-structured data, such as text metadata, proofreading rules, etc.

[0247] Data writing and reading: Text files are written to object storage using the API (Application Programming Interface) provided by the object storage service. Each file is assigned a unique object key for identification and retrieval. When reading a file, its content is retrieved using the object key.

[0248] Data Management: The object storage system provides lifecycle management capabilities, allowing you to set expiration times for knowledge text files to optimize storage costs and performance.

[0249] M7, Knowledge Update Reminder Module, specifically.

[0250] The knowledge update reminder module monitors the dynamics of the enterprise knowledge management base in real time. When knowledge text is added or updated, it automatically reminds relevant personnel to review and update other related knowledge text content based on custom rules, ensuring the timeliness of knowledge. After the relevant personnel revise and approve the text, the knowledge text is officially added to the base. This achieves real-time maintenance of the timeliness of the enterprise knowledge management base.

[0251] The workflow of the knowledge update reminder module includes:

[0252] Scan the knowledge management database according to a preset time cycle, compare text timestamps, version numbers or hash values, and identify whether there are text updates or additions;

[0253] Based on knowledge association rules, identify texts that are related to the updated or added text; wherein, the association includes direct references, indirect associations, or texts on the same topic;

[0254] Integrate key changes to the updated or added text with related text information, and define reminder messages for multi-channel push notifications.

[0255] The knowledge update reminder module aims to monitor the update dynamics of various knowledge in the enterprise knowledge management base in real time, and when an update is detected, it will quickly notify relevant personnel to review and revise the associated text to ensure the timeliness and accuracy of the knowledge.

[0256] Workflow:

[0257] (1) Update monitoring trigger: Use the periodic scanning polling script theorem to scan the knowledge management base, compare the timestamp, version number or hash value of the text, and identify the updated knowledge text.

[0258] (2) Related text location: Based on the knowledge association rules, find other texts that are related to the updated text, including direct citations, indirect associations, or texts on the same topic.

[0259] (3) Reminder message compilation: Integrate the key modification information of the update text with the detailed information of the related text and compile it into a clear reminder message.

[0260] (4) Multi-channel push: Through enterprise office software, enterprise email or enterprise SMS channels, the reminder will be accurately pushed to the person in charge of the knowledge text or the person in charge of the relevant department.

[0261] For example, when a company updates its product manual, the knowledge update reminder module monitors the changes in real time, locates related technical documents, training materials, and other related texts based on knowledge graph association rules, and generates a notification containing a summary of the updated content and modification suggestions.

[0262] The knowledge update reminder module can promptly convey knowledge update dynamics, helping enterprises maintain the timeliness and accuracy of their knowledge management base, and improve decision-making quality and business collaboration efficiency.

[0263] M8, the custom rule configuration module, specifically.

[0264] The custom rule configuration module meets the personalized needs of enterprises. Enterprise administrators can flexibly configure text proofreading rules according to their own business characteristics, industry standards, and text writing habits.

[0265] Custom rules for proofreading text can be flexibly configured based on business needs, including text-based keywords, key words, format features, context categories, knowledge graph association rules, as well as user permissions and department permissions.

[0266] Overall, this system achieves comprehensive quality control of enterprise documents through multi-module collaboration, improves document quality and knowledge management efficiency, and facilitates the efficient flow of enterprise information and the inheritance of corporate culture.

[0267] Based on the custom rules in this module, the accuracy of proofreading can be effectively improved.

[0268] The enterprise-level intelligent text proofreading system of this invention effectively solves the problem of poor adaptability of existing technologies through the synergistic effect of the custom rule configuration module, the intelligent text proofreading module, and the knowledge update reminder module.

[0269] Enterprises can flexibly configure proofreading rules based on their own business characteristics, industry standards, and text writing habits, and define specific text context rules, including features such as text keywords, format, and content structure, as well as setting user permissions and department permissions.

[0270] Enterprise administrators define specific contextual rules based on the enterprise's own business characteristics and text types. These rules can be based on features such as text keywords, format, and content structure; user permissions, department permissions; and setting context categories; and setting corresponding keywords for each context category.

[0271] like Figure 5 The diagram shows a flowchart of an enterprise-level intelligent text proofreading system applied by enterprise users. This system enables efficient, real-time proofreading and approval of enterprise knowledge texts before their formal submission to the enterprise knowledge management repository, improving proofreading efficiency, accuracy, and real-time performance.

[0272] In summary, the enterprise-level intelligent text proofreading system of this invention has the following beneficial effects:

[0273] 1. This invention adopts a multi-module collaborative working approach, such as the convenient operation of the text input and retrieval module, the automated processing of the intelligent text proofreading module, and the word-by-word voting mutual proofreading algorithm, which greatly reduces the workload and time cost of manual proofreading, and can quickly proofread a large number of knowledge texts to be added to the database, meeting the needs of enterprises for rapid knowledge updates and large-scale knowledge text processing, and improving the efficiency of knowledge text proofreading.

[0274] 2. This invention accurately identifies the context of a text through a text context recognition module, and combines it with a proofreading method based on multi-path mutual proofreading model fusion and context self-adaptation in an intelligent text proofreading module, as well as proofreading of professional terms, industry standard formats, etc. It fully considers the professionalism and domain specificity of enterprise texts, effectively solves the problem of insufficient proofreading accuracy in existing methods, ensures the accuracy and professionalism of knowledge text content, and improves proofreading accuracy.

[0275] 3. The custom rule configuration module in this invention allows enterprises to flexibly configure proofreading rules according to their own business characteristics, industry standards and text writing habits. At the same time, the intelligent text proofreading module can adaptively adjust the proofreading strategy for different text contexts, so that the system can better adapt to the constantly changing and updated knowledge text types and business contexts in the enterprise knowledge management base, and has universality and scalability, enhancing the system's adaptability.

[0276] 4. The text plagiarism detection module in this invention utilizes advanced text similarity detection algorithms to compare the text to be proofread with existing knowledge text content in the enterprise's knowledge management database. It employs cosine similarity and knowledge graph feature comparison to identify semantically similar content (such as synonym replacement and word order adjustment), quickly identifying potentially plagiarized or over-quoted parts and generating a plagiarism report. This helps enterprises protect the originality of their intellectual property assets, avoid intellectual property issues caused by text duplication, and safeguard the originality of knowledge.

[0277] 5. The knowledge update reminder module in this invention can monitor the dynamics of the enterprise knowledge management base in real time. When a knowledge update is detected, it promptly notifies relevant personnel to review and update the knowledge text to be added to the base and related texts, thereby ensuring the timeliness and effectiveness of the knowledge text in the knowledge management base and improving the efficiency and quality of enterprise decision support, employee training and business collaboration.

[0278] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. An enterprise-level intelligent text proofing system, characterized in that, The application relates to a smart text proofreading system and method. The application comprises the following: a text input and retrieval module for receiving a text to be proofread and supporting multi-format input and draft storage; a text context recognition module for recognizing a context category to which the text to be proofread belongs by using a trained text context recognition model; the text context recognition model comprises an input layer, a context bidirectional convolution layer, a maximum pooling layer and a full connection classification layer in sequence; the input layer is used for vectorizing the text to be proofread to obtain a word vector matrix corresponding to the text to be proofread; the context bidirectional convolution layer is used for performing convolution on the word vector matrix by using a forward convolution kernel to obtain a forward feature map and performing convolution on the word vector matrix by using a reverse convolution kernel to obtain a reverse feature map; gate parameters are obtained based on the forward feature map and the reverse feature map, and fused feature maps are obtained based on the gate parameters; the maximum pooling layer is used for sliding a window on each convolution kernel output fused feature map and selecting a maximum value in the window to generate a fixed-length feature vector corresponding to each convolution kernel; the full connection classification layer is used for splicing the fixed-length feature vectors corresponding to each convolution kernel into a global feature vector, performing full connection classification on the global feature vector, outputting a context classification probability and obtaining a context classification result based on the context classification probability; a self-defined rule configuration module is used for defining rules and dynamically loading the rules to a smart text proofreading module; a smart text proofreading module is used for proofreading the text to be proofread by using a trained multi-channel mutual proofreading model, the context category and self-defined rules to obtain a proofreading result; a multi-channel mutual proofreading model with parallel encoding is constructed, comprising a first mutual proofreading channel, a second mutual proofreading channel and a third mutual proofreading channel arranged in parallel; the first mutual proofreading channel comprises a first encoder, a first cross-model attention layer and a T5 decoder in sequence, and the first encoder comprises a BERT encoder and a Seq2Seq encoder with parallel encoding; the first cross-model attention layer is used for weighted fusion based on the importance of the encoding results of the BERT encoder and the Seq2Seq encoder to obtain weighted first encoding features; the T5 decoder is used for decoding the first encoding features to obtain a first proofreading result; the second mutual proofreading channel comprises a second encoder, a second cross-model attention layer and a Seq2Seq decoder in sequence, and the second encoder comprises a BERT encoder and a T5 encoder with parallel encoding; the second cross-model attention layer is used for weighted fusion based on the importance of the encoding results of the BERT encoder and the T5 encoder to obtain weighted second encoding features; the Seq2Seq decoder is used for decoding the second encoding features to obtain a second proofreading result; the third mutual proofreading channel comprises a third encoder, a third cross-model attention layer and a BERT classifier in sequence, and the third encoder comprises a Seq2Seq encoder and a T5 encoder with parallel encoding. The third cross-model attention layer is configured to perform weighted fusion based on importance of encoding results of the Seq2Seq encoder and the T5 encoder, to obtain a third encoding feature after weighting; and the BERT classifier is configured to classify the third encoding feature to obtain a third proofreading result. The proofreading result output module is configured to perform word-by-word weighted voting on the proofreading result to generate a final proofreading text.

2. The system of claim 1, wherein, The system further comprises: A text duplication checking module configured to perform semantic-level similarity comparison between the final proofreading text and existing texts in an enterprise knowledge management library to obtain a duplication checking result; and if the duplication checking result is lower than a preset similarity threshold, the final proofreading text is stored in the enterprise knowledge management library. A data storage module configured to store the text to be proofread, the context category, the proofreading result, the final proofreading text, and the duplication checking result. A knowledge update reminding module configured to monitor text updates in the enterprise knowledge management library and automatically push a text storage review reminding.

3. The system of claim 1, wherein, The trained text context recognition model is obtained through the following training process: A text context training sample set composed of multiple context categories is constructed, and a sample label is a context category to which a text belongs. The text context recognition model is trained using the text context training sample set to obtain the trained text context recognition model.

4. The system of claim 1, wherein, The trained multi-pass mutual proofreading model is obtained through the following process: A specific text proofreading training sample set is constructed, which includes a general text context training sample set and multiple specific context category training sample sets. The multi-pass mutual proofreading model based on parallel encoding is pre-trained using the general text context training sample set, and the pre-trained multi-pass mutual proofreading model parameters based on parallel encoding are saved as initial parameters. The initial parameters are loaded, and each specific context category training sample set is trained in turn. The model parameters are iteratively optimized through forward propagation and back propagation until the loss function converges, to obtain the multi-pass mutual proofreading model parameters based on parallel encoding in multiple specific context categories.

5. The system of claim 4, wherein, The proofreading result output module performs word-by-word voting using dynamic weights based on three groups of proofreading results obtained by the intelligent text proofreading module through the first, second, and third mutual proofreading passes, including: For each position of the first, second, and third proofreading results, word-by-word voting is performed using dynamic weights, and the text with the most weighted votes is adopted as the final proofreading text. The voting results of each position of the text are output in the original order of the text to be proofread, to obtain the proofread text corresponding to the text to be proofread.

6. The system of claim 2, wherein, The text duplication checking module workflow includes: The final proofreading text is preprocessed to obtain a preprocessed final proofreading text. Word2Vec is used to convert words in the preprocessed final proofreading text into vector representations. Based on the vector representations of the words, a text vector of the final proofreading text is obtained. The text vector is compared with text vectors in the duplication checking enterprise knowledge management library to calculate a similarity value as a duplication checking result. If the duplication checking result is lower than a preset similarity threshold, the text to be proofread enters a storage state.

7. The system of claim 2, wherein, The knowledge update reminding module workflow includes: Scanning the knowledge management library according to a preset time period, comparing text timestamps, version numbers or hash values to identify whether there is text update or addition; According to the knowledge association rule, find out the text associated with the updated or added text; wherein, the association includes direct reference, indirect association or same theme text; Integrate the key modification information of the updated or added text and the associated text information, and define the reminder message for multi-channel push message.

8. The system according to any of claims 1-7, characterized in that Based on the business needs, flexibly configure the custom rules for proofreading text, including text-based keywords, key words, format features, context categories, knowledge graph association rules, and user permissions, department permissions.

Citation Information

Patent Citations

  • Text proofreading method based on knowledge base

    CN115293137A

  • Text proofreading method and device, equipment and storage medium

    CN116151228A

  • Multi-feature fusion Chinese patent text classification method based on TRIZ invention principle

    CN117609866A

  • Mass text duplicate checking method, system and equipment in real scene and storage medium

    CN119830895A

  • Weighted voting method-based language disease error correction model fusion method

    CN120317241A