Project text optimization processing method and device

By constructing in-depth analysis of knowledge graphs and relationship extraction, combined with segmented processing and reinforcement learning algorithms, the inadequate semantic coherence evaluation and optimization strategy formulation in project text optimization is solved, and the intelligent optimization of project text is achieved, and professionalism and efficiency are improved.

CN120430294AInactive Publication Date: 2025-08-05ZHEJIANG WANCHUANG HUILI TECHNOLOGY SERVICE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510569882.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the process of project text optimization, the existing technology has problems such as insufficient semantic coherence evaluation, inflexible optimization strategy formulation, and inaccurate template matching, making it difficult to achieve overall optimization and professional improvement of text structure.

Method used

By constructing knowledge graphs and relationship extraction, deep analysis of text structure characteristics is formed, differentiated text template libraries are formed, segmented processing strategies are used to combine part-of-speech annotation and syntactic analysis, recurrent neural networks are used for semantic coherence modeling, and text optimization strategy network is built based on reinforcement learning algorithms to dynamically generate repair strategies.

Benefits of technology

It significantly improves the professionalism and efficiency of project text optimization, can accurately identify semantic break points and perform intelligent optimization, and generate high-quality optimized text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430294A_ABST
    Figure CN120430294A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a project text optimization processing method and device, and the method comprises the steps: deeply analyzing text structure features and semantic features through a knowledge graph and relation extraction, and forming a differentiated text template library; a segmentation processing strategy is adopted, text boundaries are recognized in combination with part-of-speech tagging and syntactic analysis, sequence modeling is conducted on semantic coherence through a recurrent neural network, and semantic breaking points are accurately positioned. A text optimization strategy network is constructed based on a reinforcement learning algorithm, grammar error positions, semantic coherence breaking points and a text template are used as state input, a repair strategy is dynamically generated, and intelligent optimization of a text is achieved. By means of the method, the defects of a traditional technology in the aspects of semantic coherence evaluation, optimization strategy making and the like are effectively overcome, and the professionality and efficiency of project text optimization are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field The present application relates to the field of data processing, and specifically to a method and device for optimizing project text processing. Background Art Existing methods for optimizing project texts have significant shortcomings. Traditional methods are relatively simple in terms of text structure analysis and semantic association identification, lacking in-depth exploration of the logical relationships between technical content, making it difficult to ensure the professionalism and consistency of optimization results.

[0001] Furthermore, existing technologies face bottlenecks in text segment processing and semantic coherence assessment. Most systems focus solely on correcting superficial grammatical errors, neglecting the semantic cohesion between paragraphs, making it difficult to optimize the overall text structure. Template matching also lacks flexibility, hindering optimization effectiveness.

[0002] Existing systems have technical shortcomings in optimizing strategy formulation and execution. They lack intelligent optimization decision-making mechanisms, making it difficult to dynamically adjust repair strategies based on text characteristics. Addressing these issues is crucial for improving the quality and efficiency of project text optimization. Summary of the Invention In response to the problems in the existing technology, this application provides a project text optimization processing method and device, which can effectively solve the shortcomings of traditional technology in semantic coherence evaluation, optimization strategy formulation, etc., and significantly improve the professionalism and efficiency of project text optimization.

[0003] In order to solve at least one of the above problems, the present application provides the following technical solutions: In a first aspect, the present application provides a project text optimization processing method, comprising: Construct a training corpus, extract text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, use a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculate weight coefficients for the knowledge association edges, and cluster the texts in the corpus based on the weight coefficients to form a differentiated text template library, input the text structure features, semantic features, and grammatical rules into a multi-task text analysis model for training, use a cross-entropy loss function to optimize the parameters of the multi-task text analysis model, and obtain the grammatical error recognition accuracy and semantic coherence assessment accuracy of the multi-task text analysis model on a validation set; The project text to be optimized is segmented, sentence boundaries and paragraph structure are identified based on part-of-speech tagging and syntactic analysis, the project text is divided into text segments, and the segments are input into the multi-task text analysis model to perform grammar checking and semantic coherence assessment tasks, obtaining grammatical error position annotations in the text segments and semantic coherence scores between adjacent segments, using a recurrent neural network to perform sequence modeling on the semantic coherence scores, identifying semantic coherence breakpoints, and screening text templates that match the project text structure based on the differentiated text template library; A text optimization strategy network is constructed based on the reinforcement learning algorithm. The grammatical error position, semantic coherence breakpoint and text template are used as state input, and a text repair action sequence is output. According to the action sequence, the project text is subjected to grammatical correction and semantic coherence optimization. The repaired text fragments are reorganized into a complete text. The format of the complete text is verified using a text format standardization check model. The format of the parts that fail the verification is adjusted to obtain the optimized project text.

[0004] Furthermore, the process also includes: establishing a data interface with a project application text library and an external standard library, collecting text data to construct a basic corpus, performing word segmentation and part-of-speech tagging on the text content in the basic corpus, extracting word vector features, topic features, and syntactic tree features to construct a text structure feature matrix, using a word embedding model and a deep semantic model to extract semantic features of the text, and inputting the text structure feature matrix and semantic features into a graph neural network to generate node representations of a knowledge graph; The text in the basic corpus is input into the relationship extraction model to identify the technical elements and their associations in the text, extract the technical structure association features and key point dependency features to construct knowledge association edges, and use the attention mechanism to calculate the weight coefficient of the knowledge association edge. The weight coefficient is input into the spectral clustering algorithm to cluster the text in the corpus, and a differentiated text template library is constructed based on the clustering results.

[0005] Furthermore, the method further includes: constructing a network structure of a multi-task text analysis model, fusing features of outputs from a text structure feature extraction layer, a semantic feature extraction layer, and a grammatical rule encoding layer, performing feature mapping on the fused features using a shared encoder to obtain a text representation vector, inputting the text representation vector into a grammatical error recognition branch and a semantic coherence assessment branch, respectively, and adding a softmax classifier to the output layer of each branch; The multi-task text analysis model is trained based on labeled training samples. The cross-entropy loss of the grammatical error recognition branch and the semantic coherence evaluation branch is calculated. The losses of the two branches are weighted summed to obtain the total loss function. The model parameters are optimized using the stochastic gradient descent method. The grammatical error recognition accuracy and the semantic coherence evaluation accuracy are calculated on the validation set, respectively.

[0006] Furthermore, the process also includes: pre-processing the project text to be optimized, segmenting the text content using a word segmentation model, tagging the word segmentation results with parts of speech, identifying the grammatical structure and dependency relationships of the sentences based on a syntactic analysis tree model, extracting sentence boundary markers based on punctuation information, identifying the start and end positions of paragraph structures using a text segmentation algorithm, and dividing the project text into text segments based on the sentence boundary markers and paragraph structures; The divided text segments are input into the multi-task text analysis model, the word vector features and syntactic features of each text segment are extracted, the grammar checking task is performed based on the feature extraction results to mark grammatical errors, the cosine similarity of the word vectors between adjacent text segments is calculated, and the semantic coherence score is evaluated in combination with the syntactic features.

[0007] Furthermore, the method further includes: constructing a text grammatical error annotation matrix, encoding the position of the grammatical errors annotated in each text segment, calculating semantic coherence scores between adjacent text segments, constructing the semantic coherence scores into a temporal feature sequence, modeling the temporal feature sequence using a long short-term memory network, calculating the state vector and forget gate output of each time step, and identifying the location of the semantic coherence breakpoint based on the change amplitude of the state vector; Calculate the structural feature vector of the project text, encode the structural features of the templates in the differentiated text template library, use cosine similarity to calculate the similarity score between the project text structural feature vector and the template structural feature vector, and select the text template with the highest similarity score as the optimized template.

[0008] Furthermore, the method further includes: constructing a text state vector, concatenating a grammatical error position matrix, a semantic coherence breakpoint position vector, and text template features to obtain a state input, encoding the state input using a policy network, generating a probability distribution of text repair actions, sampling the action space based on a Monte Carlo tree search algorithm, constructing an action decision tree based on the sampling results, and calculating a value score for each action node; The policy network is trained, and the similarity between the optimized text and the original text is used as the reward function. The value estimation of each state-action pair is calculated based on the temporal difference algorithm. The value estimation results are backpropagated to update the policy network parameters, and the output of the trained policy network is converted into a text repair action sequence.

[0009] Furthermore, the method further includes: correcting the marked grammatical error locations according to the text repair action sequence, optimizing the connection of semantic coherence breakpoints using semantic connectives in the text template, adjusting the word order of the repaired text segments based on semantic similarity calculation, and splicing and reorganizing the optimized text segments according to the original text structure sequence to generate a complete text including paragraph markers; The reorganized complete text is input into the format standardization check model, the format feature vector of the text is extracted, and the format elements such as font, font size, indentation, paragraph spacing, etc. of the text are verified for standardization. The positions of format elements that do not meet the standards are identified, and the format elements that fail to pass the verification are automatically adjusted according to the preset format rule template, and the optimized project text that meets the standards is output.

[0010] In a second aspect, the present application provides a project text optimization processing device, comprising: A pre-training module is used to construct a training corpus, extract text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, use a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculate weight coefficients for the knowledge association edges, and cluster the texts in the corpus based on the weight coefficients to form a differentiated text template library, input the text structure features, semantic features, and grammatical rules into a multi-task text analysis model for training, use a cross-entropy loss function to optimize the parameters of the multi-task text analysis model, and obtain the grammatical error recognition accuracy and semantic coherence assessment accuracy of the multi-task text analysis model on a validation set; A text processing module is used to segment the text of the project to be optimized, identify sentence boundaries and paragraph structure based on part-of-speech tagging and syntactic analysis, divide the project text into text segments, input the multi-task text analysis model to perform grammar checking and semantic coherence assessment tasks, obtain grammatical error location annotations in the text segments and semantic coherence scores between adjacent segments, use a recurrent neural network to perform sequence modeling on the semantic coherence scores, identify semantic coherence breakpoints, and screen text templates that match the project text structure based on the differentiated text template library; The text optimization module is used to build a text optimization strategy network based on the reinforcement learning algorithm, take the grammatical error position, semantic coherence breakpoint and text template as state input, output text repair action sequence, perform grammatical error correction and semantic connection optimization on the project text according to the action sequence, reorganize the repaired text fragments into a complete text, use the text format standardization check model to verify the format of the complete text, adjust the format of the part that fails the verification, and obtain the optimized project text.

[0011] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the project text optimization processing method when executing the program.

[0012] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the project text optimization processing method when executed by a processor.

[0013] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instruction, which implements the steps of the project text optimization processing method when executed by a processor.

[0014] It can be seen from the above technical solution that the present application provides a project text optimization processing method and device, which deeply analyzes the text structure features and semantic features through knowledge graphs and relationship extraction to form a differentiated text template library. A segmented processing strategy is adopted, combined with part-of-speech tagging and syntactic analysis to identify text boundaries, and semantic coherence is sequence modeled through recurrent neural networks to accurately locate semantic breakpoints. A text optimization strategy network is constructed based on a reinforcement learning algorithm, and the grammatical error position, semantic coherence breakpoint and text template are used as state inputs to dynamically generate repair strategies and realize intelligent optimization of text. This method effectively solves the shortcomings of traditional technologies in semantic coherence evaluation, optimization strategy formulation, etc., and significantly improves the professionalism and efficiency of project text optimization. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0015] Figure 1 Schematic diagram of the process of optimizing the project text in the embodiment of the present application; Figure 2 This is a structural diagram of a project text optimization processing device in an embodiment of the present application; Figure 3 Schematic diagram of the structure of the electronic device in the embodiment of the present application.

[0016] Reference numerals: Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0017] The acquisition, storage, use, and processing of data in this application's technical solution comply with relevant national laws and regulations.

[0018] Taking into account the problems existing in the prior art, the present application provides a project text optimization processing method and device, which deeply analyzes the text structure features and semantic features through knowledge graphs and relationship extraction to form a differentiated text template library. A segmented processing strategy is adopted, combining part-of-speech tagging and syntactic analysis to identify text boundaries, and a recurrent neural network is used to perform sequence modeling of semantic coherence to accurately locate semantic breakpoints. A text optimization strategy network is constructed based on a reinforcement learning algorithm, and the grammatical error position, semantic coherence breakpoint and text template are used as state inputs to dynamically generate repair strategies and realize intelligent optimization of text. This method effectively solves the shortcomings of traditional technologies in semantic coherence evaluation, optimization strategy formulation, etc., and significantly improves the professionalism and efficiency of project text optimization.

[0019] In order to effectively solve the shortcomings of traditional technologies in semantic coherence evaluation and optimization strategy formulation, and significantly improve the professionalism and efficiency of project text optimization, this application provides an embodiment of a project text optimization processing method, see Figure 1 The project text optimization processing method specifically includes the following contents: Step S101: constructing a training corpus, extracting text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, using a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculating weight coefficients for the knowledge association edges, and clustering the texts in the corpus based on the weight coefficients to form a differentiated text template library, inputting the text structure features, semantic features, and grammatical rules into a multi-task text analysis model for training, using a cross-entropy loss function to optimize the parameters of the multi-task text analysis model, and obtaining the grammatical error recognition accuracy and semantic coherence evaluation accuracy of the multi-task text analysis model on a validation set; Optionally, this embodiment first constructs a data collection mechanism for the system. By establishing a data interface with the project application text library, historical high-quality project application documents are obtained, including different types of application materials such as technology innovation projects and achievement transformation projects. At the same time, it connects to authoritative data sources such as external technical standard libraries and industry specification libraries to collect standardized technical description texts. Distributed crawler technology is used in the collection process to ensure the real-time and integrity of the data. The collected text data is preliminarily cleaned to remove invalid characters, unify the encoding format, and perform document structure analysis to identify different types of content blocks such as titles, texts, and charts.

[0020] This embodiment implements an in-depth feature extraction mechanism. The improved BERT model is used to perform word segmentation and part-of-speech tagging on the cleaned text. The model is pre-trained on technical document corpus and can accurately identify professional terms and technical descriptions. When extracting word vector features, the Word2Vec model is used to train domain-specific word embeddings. The word vector dimension is set to 300, and the contextual semantics of the word is captured through a sliding window. The topic feature extraction adopts the LDA topic model to adaptively determine the number of topics based on the document content. The syntactic tree features are obtained through dependency syntactic analysis, focusing on the logical structure of the technical solution description. These feature combinations constitute the text structure feature matrix, which provides a basis for subsequent analysis.

[0021] This embodiment designs an innovative semantic feature extraction scheme. A bidirectional long short-term memory network is used to extract deep semantic features from text. The network comprises multiple BiLSTM layers, each with a hidden state dimension of 256. To enhance semantic representation capabilities, an attention mechanism is introduced, assigning different weights to text content at different locations. In particular, attention weights are increased for important content such as technological innovations and key technical parameters. Furthermore, a pre-trained BERT model is used to extract context-sensitive semantic representations, capturing long-range semantic dependencies.

[0022] This embodiment implements a complete knowledge graph construction process. The text structure feature matrix and semantic features are input into a graph neural network, and a graph attention network (GAT) is used to learn node representations. The network consists of three GAT layers, each with 8 attention heads and an output dimension of 128. The initial node features include a combination of word vectors, topic distributions, and syntactic features. Through multi-layer feature propagation and attention aggregation, node embeddings that take into account structural and semantic information are generated. Residual connections and layer normalization are introduced to improve the expressive power of the graph.

[0023] This embodiment innovatively designs a relationship extraction model. A relationship classifier is constructed based on a pre-trained language model, and the model adopts an encoder-classifier structure. The encoder uses the RoBERTa model, which has been trained for domain adaptation of technical document data. The classifier contains a multi-layer perceptron to identify the types of associations between technical elements. The model can accurately identify technical dependencies (such as the sequence of method steps, component composition relationships) and key point associations (such as problem-solution, innovation-effect). In order to improve extraction accuracy, a remote supervision data enhancement mechanism is introduced.

[0024] This embodiment implements an accurate weight calculation method. A multi-head attention mechanism is used to calculate the importance weight of the knowledge association edge. The calculation formula is: w_ij = softmax(Q_iK_j^T / √d), where Q_i and K_j are the query vector and key vector of nodes i and j respectively, and d is the vector dimension. The attention score reflects the importance of the association edge in the technical structure. The weight calculation takes into account multiple factors: the importance of the association type, the technical relevance of the node, the integrity of the path, etc. The calculated weights are normalized to ensure the rationality of the weight distribution.

[0025] This embodiment designs an efficient text clustering mechanism. A text similarity matrix is constructed based on the calculated association edge weights. Text clustering is performed using an improved spectral clustering algorithm, which combines structural and semantic features in its similarity calculation. A dynamic threshold is used during the clustering process to determine the number of clusters, ensuring that the text similarity within each cluster meets the required level. The clustering results are used to construct a differentiated text template library, with each cluster corresponding to a specific type of text template. The template library contains standardized description frameworks for different types of projects.

[0026] This example implements a complete model training process. A multi-task text analysis model is constructed, consisting of a shared feature extraction layer and a task-specific output layer. The feature extraction layer uses a Transformer encoder, and the output layer is divided into two branches: grammatical error detection and semantic coherence assessment. The cross-entropy loss function is used during training, with the loss function being: L = α L_ grammar + β L_coherence, where α and β are task weight coefficients. The optimization process uses the Adam optimizer, and the learning rate adopts a cosine annealing strategy. The performance indicators of the two tasks are evaluated on the validation set.

[0027] Through the above technical innovations, this embodiment effectively solves several key problems in the analysis of project application texts: insufficient feature extraction, incomplete knowledge representation, inaccurate template construction, etc. In practical applications, this solution can accurately analyze the structure and semantic features of project texts and build high-quality knowledge graphs and template libraries. It is particularly suitable for the processing of highly professional application documents such as technological innovation projects and achievement transformation projects, and significantly improves the accuracy of text analysis and the precision of template matching. The adaptive characteristics of this solution enable it to handle different types of project application requirements, and through continuous data accumulation and model optimization, it provides a reliable basic support for subsequent text optimization.

[0028] Step S102: Segment the project text to be optimized, identify sentence boundaries and paragraph structures based on part-of-speech tagging and syntactic analysis, divide the project text into text segments, input the multi-task text analysis model to perform grammar checking and semantic coherence assessment tasks, obtain grammatical error position annotations in the text segments and semantic coherence scores between adjacent segments, use a recurrent neural network to perform sequence modeling on the semantic coherence scores, identify semantic coherence breakpoints, and screen text templates that match the project text structure based on the differentiated text template library; Optionally, this embodiment first designs a comprehensive text preprocessing mechanism. For the project text to be optimized, an improved preprocessing process is adopted, including text cleaning, coding unification and format standardization. Special attention is paid to special content such as professional terms, mathematical formulas and chart references in technical documents, and special marking rules are used for processing. Text segmentation adopts a word segmentation model based on deep learning. The model has been specifically trained on technical document corpus and can accurately identify professional terms and compound words. The word segmentation results are tagged with parts of speech through a conditional random field model. The tagging process pays special attention to the recognition accuracy of technical nouns, verbs and conjunctions.

[0029] This embodiment implements an innovative syntactic analysis solution. A dependency syntactic analysis tree model is constructed, and a deep bidirectional LSTM network is used to extract the grammatical features of sentences. The input of the model includes a word vector sequence and part-of-speech tagging information, and the grammatical structure representation of the sentence is obtained through multi-layer feature extraction. In the process of constructing the syntactic tree, special attention is paid to the logical relationships in the description of the technical solution, such as causal relationships, conditional relationships, and progressive relationships. Based on the syntactic analysis results, combined with punctuation and keyword information, the boundary position of the sentence is accurately identified. Special cases, such as boundary identification of multiple nested sentences and long sentences, are handled through rule enhancement.

[0030] This embodiment designs an accurate paragraph structure recognition mechanism. An attention-based text segmentation model is adopted, and the model structure includes two main parts: an encoder and a classifier. The encoder uses a Transformer structure to capture long-distance text dependencies through a self-attention mechanism. The classifier performs binary classification on each position to determine whether it is a paragraph boundary. The model training adopts a multi-task learning framework to simultaneously predict paragraph boundaries and topic coherence. In order to improve recognition accuracy, document structure features are introduced, such as title hierarchy, indentation format and other information. Based on the recognition results, the project text is divided into multiple semantically relatively independent text segments.

[0031] This embodiment implements an efficient text analysis process. The segmented text segments are input into a multi-task text analysis model, which uses a shared encoder and task-specific decoder architecture. The encoder is based on the BERT model and has been trained for domain adaptation on technical document data. The grammar checking task uses a sequence tagging method, using a CRF layer to predict the error type of each word. The semantic coherence assessment task uses an attention mechanism to calculate the semantic relevance between adjacent segments. The evaluation process considers multiple dimensions: vocabulary overlap, topic consistency, and the naturalness of logical transitions.

[0032] This embodiment innovatively designs a semantic coherence modeling scheme. A recurrent neural network based on bidirectional GRU is constructed to model the semantic coherence score sequence. The hidden state dimension of the network is 256, and the retention and update of historical information is controlled by a gating mechanism. The model input is a semantic coherence score vector between adjacent segments, which contains scores of multiple evaluation dimensions. The sequence modeling process highlights important semantic transition points through the attention mechanism. The calculation formula of the attention score is: a_t = softmax(v^T tanh(W_h h_t + W_s s_t)), where h_t is the hidden state, s_t is the score vector, and W_h, W_s and v are learnable parameters.

[0033] This embodiment designs a comprehensive breakpoint identification mechanism. Based on the sequence modeling results, a dynamic programming algorithm is used to identify semantic coherence breakpoints. This algorithm considers local score variations and global text structure to set an adaptive breakpoint threshold. This threshold is dynamically adjusted based on text type and length to avoid over-segmentation. For identified breakpoints, further analysis is performed to determine their causes, such as topic jumps, logical gaps, or incoherent expression, providing a basis for subsequent optimization.

[0034] This embodiment implements an accurate template matching solution. Structural features of the project text are extracted, and a feature vector is constructed containing information such as paragraph organization, logical relationships, and technical description characteristics. At the same time, the templates in the differentiated text template library are structurally encoded, and matched using a similarity calculation method. The similarity calculation uses an improved cosine similarity formula: sim(d,t) = (d·t) / (||d||·||t||+λ), where d is the document feature vector, t is the template feature vector, and λ is a smoothing factor. The matching process considers matching at multiple levels, including structural similarity, topic relevance, and consistency of expression style.

[0035] This embodiment also optimizes the template selection strategy. Based on the preliminary matching results, a multi-round screening mechanism is designed. The first round is based on structural similarity for rough selection, and the templates ranked in the top K in similarity are selected. The second round considers the technical characteristics of the text and analyzes the degree of match between the template and the text in terms of technical description depth, innovative point display, etc. The last round selects the most suitable optimized template based on the project type and application requirements. In order to improve the accuracy of the selection, a feedback mechanism for the template usage effect is established, and the template weight is dynamically adjusted according to the historical application effect.

[0036] Through the above technical innovations, this embodiment effectively solves several key problems in the analysis of project application texts: inaccurate structural recognition, insufficient semantic analysis, inaccurate template matching, etc. In practical applications, this solution can accurately analyze the structural features and semantic characteristics of project texts and identify key positions that need to be optimized. It is particularly suitable for the processing of highly professional application documents such as technological innovation projects and achievement transformation projects, and significantly improves the accuracy of text analysis and the pertinence of optimization suggestions. The adaptive characteristics of this solution enable it to handle different types of project application requirements, and through continuous optimization and improvement, it provides reliable technical support for improving the quality of project application documents.

[0037] Step S103: Construct a text optimization strategy network based on the reinforcement learning algorithm, take the grammatical error position, semantic coherence breakpoint and text template as state input, output a text repair action sequence, perform grammatical error correction and semantic coherence optimization on the project text according to the action sequence, reorganize the repaired text fragments into a complete text, use a text format standardization check model to verify the format of the complete text, adjust the format of the part that fails the verification, and obtain the optimized project text.

[0038] Optionally, this embodiment first designs an innovative state vector construction mechanism. For the grammatical error positions in the project text, a position encoding matrix is used for representation, and the matrix dimension is [sequence length, number of error types]. The vector of each position reflects the specific grammatical error type present there. The semantic coherence breakpoints are marked by binary vectors, and the vector dimension is the same as the number of text fragments, and 1 represents the breakpoint position. The text template features are extracted through a pre-training model, which includes two parts: structural features and semantic features. These three types of features are integrated into a unified state representation through a feature fusion network, which serves as the input of the policy network. The fusion network adopts a multi-layer perceptron structure, and dynamically weights different features through an attention mechanism.

[0039] This embodiment implements a deep policy network architecture. An optimized policy network is constructed based on the Actor-Critic framework. The Actor network is responsible for generating text repair actions, and the Critic network evaluates state-action values. The Actor network uses a Transformer encoder to extract state features, and the decoder generates a probability distribution of action sequences. To enhance the decision-making ability of the model, a multi-head self-attention mechanism is added to the encoder to capture the correlation between state features. The action space contains a variety of basic operations: grammatical correction, semantic conjunction insertion, sentence rearrangement, etc. At each time step, the model selects the optimal repair action based on the current state.

[0040] This embodiment designs an accurate action sampling strategy. An improved Monte Carlo tree search algorithm is used for action exploration, and each node of the search tree represents a state-action pair. The value estimate of the node is calculated by the following formula: Q(s,a) = R(s,a) + γ*V(s'), where R is the immediate reward, γ is the discount factor, and V(s') is the value estimate of the next state. During the search process, the UCB criterion is used to balance exploration and utilization, and the action path with the highest score is selected. In order to improve the search efficiency, a pruning mechanism is introduced to terminate low-value search branches in advance. At the same time, the diversity of sampling is optimized through experience replay.

[0041] This embodiment implements an effective reward design mechanism. The reward function comprehensively considers multiple optimization objectives: grammatical accuracy, semantic coherence, and similarity to the original text. The reward for grammatical error correction is based on the reduction in the number of errors, and the improvement in semantic coherence is calculated based on the elimination of breakpoints. To maintain the original meaning of the text, similarity constraints are set to penalize modifications that deviate too much from the original text. The reward calculation formula is: R = w1 Rg + w2 Rc + w3*Rs, where Rg, Rc, and Rs are grammaticality, coherence, and similarity rewards respectively, and w is the weight coefficient.

[0042] This embodiment innovatively designs a parameter optimization mechanism. It uses the Proximal Policy Optimization (PPO) algorithm for policy network training, employing trust region constraints to prevent excessive policy updates. During training, temporal difference learning is used to estimate state-action values, and importance sampling corrects for distribution bias in offline data. To improve training efficiency, a parallel environment is used for data collection to accelerate policy convergence. Furthermore, gradient clipping is implemented to prevent training instability.

[0043] This embodiment implements a complete text repair process. Based on the action sequence output by the policy network, it first corrects the marked grammatical errors, including spelling correction and grammatical rule repair. Appropriate connectives are then inserted at semantic breakpoints to enhance text coherence. The repaired text segments are then reassembled and sorted based on logical order and semantic relevance. The reassembly process optimizes the segment combination scheme using a dynamic programming algorithm to ensure the generated text has a reasonable structure.

[0044] This embodiment designs a strict formatting and compliance checking mechanism. A formatting and compliance checking model is constructed, which uses a convolutional neural network to extract text formatting features, including font, font size, paragraph spacing, and other factors. The checking process uses a combination of a rule engine and deep learning, ensuring compliance with basic formatting standards while also handling complex formatting decisions. Non-compliant sections are automatically adjusted based on pre-set format templates to ensure the compliance of the final text.

[0045] Through the above technical innovations, this embodiment effectively solves several key problems in project text optimization: inflexible optimization strategies, unstable repair effects, poor format standardization, etc. In practical applications, the solution can intelligently identify and repair problems in the text and generate high-quality optimized text. It is particularly suitable for the optimization of highly professional application documents such as technological innovation projects and achievement transformation projects, and significantly improves the standardization and readability of the text. The adaptive characteristics of the solution enable it to handle different types of text optimization needs, and through continuous optimization and learning, it provides reliable technical support for improving the quality of project application documents.

[0046] From the above description, it can be seen that the project text optimization processing method provided in the embodiment of the present application can deeply analyze the text structure features and semantic features through knowledge graphs and relationship extraction to form a differentiated text template library. A segmented processing strategy is adopted, combining part-of-speech tagging and syntactic analysis to identify text boundaries, and semantic coherence is sequence modeled through recurrent neural networks to accurately locate semantic breakpoints. A text optimization strategy network is constructed based on a reinforcement learning algorithm, and the grammatical error position, semantic coherence breakpoint and text template are used as state inputs to dynamically generate repair strategies and realize intelligent optimization of text. This method effectively solves the shortcomings of traditional technologies in semantic coherence evaluation, optimization strategy formulation, etc., and significantly improves the professionalism and efficiency of project text optimization.

[0047] In one embodiment of the project text optimization processing method of the present application, the following contents may also be specifically included: Step S201: Establish a data interface with the project application text library and the external standard library, collect text data to build a basic corpus, perform word segmentation and part-of-speech tagging on the text content in the basic corpus, extract word vector features, topic features, and syntactic tree features to build a text structure feature matrix, use a word embedding model and a deep semantic model to extract the semantic features of the text, and input the text structure feature matrix and semantic features into a graph neural network to generate node representations of a knowledge graph; Step S202: Input the text in the basic corpus into the relationship extraction model, identify the technical elements and their associations in the text, extract the technical structure association features and key point dependency features to construct knowledge association edges, use the attention mechanism to calculate the weight coefficients of the knowledge association edges, input the weight coefficients into the spectral clustering algorithm to cluster the text in the corpus, and construct a differentiated text template library based on the clustering results.

[0048] In one embodiment of the project text optimization processing method of the present application, the following contents may also be specifically included: Step S301: Constructing a network structure for a multi-task text analysis model, performing feature fusion on the outputs of the text structure feature extraction layer, the semantic feature extraction layer, and the grammatical rule encoding layer, performing feature mapping on the fused features using a shared encoder to obtain a text representation vector, inputting the text representation vector into the grammatical error recognition branch and the semantic coherence assessment branch, respectively, and adding a softmax classifier to the output layer of each branch; Step S302: Train the multi-task text analysis model based on the labeled training samples, calculate the cross entropy loss of the grammatical error recognition branch and the semantic coherence evaluation branch, perform a weighted summation of the losses of the two branches to obtain a total loss function, optimize the model parameters using the stochastic gradient descent method, and calculate the grammatical error recognition accuracy and the semantic coherence evaluation accuracy on the validation set.

[0049] Optionally, this embodiment first designs an innovative multi-task network architecture. Three feature extraction layers are constructed: the text structure feature extraction layer adopts a multi-layer CNN structure with convolution kernel sizes of 3, 4, and 5, respectively, and 128 convolution kernels for each size, and obtains key structural features through maximum pooling; the semantic feature extraction layer uses a bidirectional Transformer structure, which includes 6 encoder layers, 8 attention heads per layer, and a hidden layer dimension of 512 to capture the deep semantic information of the text; the grammatical rule encoding layer adopts an improved LSTM network to convert predefined grammatical rules into vector representations. The hidden layer dimension of the network is 256, and the integration of rule information is controlled by a gating mechanism.

[0050] This embodiment implements a precise feature fusion mechanism. A multi-layer attention network is used to fuse the outputs of the three feature extraction layers. The attention calculation formula is: α = softmax(W[h_s; h_m; h_g]), where h_s, h_m, and h_g represent structural features, semantic features, and grammatical rule features, respectively, and W is a learnable weight matrix. Residual connections are introduced during the fusion process to preserve the original feature information. To enhance the expressiveness of features, positional encoding is added after the fusion layer to capture the positional dependencies of features.

[0051] This embodiment designs an efficient shared encoder structure. A shared encoder is constructed based on the Transformer architecture, which includes a multi-layer self-attention mechanism and a feedforward neural network. Each layer of the encoder contains a multi-head self-attention module, which focuses on different feature subspaces through different attention heads. Position encoding uses sine-cosine functions, and the calculation formula is: PE(pos,2i) = sin(pos / 10000^(2i / d)), PE(pos,2i+1) = cos(pos / 10000^(2i / d)), where pos is the position, i is the dimension index, and d is the model dimension. Through the shared encoder, the model can learn a universal text representation.

[0052] This embodiment implements an innovative branch structure design. The grammatical error recognition branch adopts a sequence tagging architecture and uses a BiLSTM-CRF model to predict the error type for each word. The CRF layer considers the transition probability between labels to improve the coherence of the prediction. The semantic coherence assessment branch adopts a hierarchical attention mechanism, first calculating the attention score at the word level and then aggregating it to obtain a sentence-level coherence representation. The two branches share underlying features but maintain task-specific output layers, achieving a balance between parameter sharing and task separation.

[0053] This embodiment innovatively designs a classifier optimization scheme. A modified softmax classifier is added to the output layer of each branch. The classifier's temperature parameter τ is learnable and is used to adjust the smoothness of the predicted probabilities. The output dimension of the grammatical error recognition branch is equal to the number of error types, and the output dimension of the semantic coherence assessment branch is equal to the number of coherence levels. The classifier output undergoes label smoothing to reduce the risk of overfitting.

[0054] This embodiment implements a systematic method for constructing training samples. A large number of project application documents are collected and training samples are constructed through a combination of expert annotation and rule-based automatic annotation. Grammatical errors include spelling errors, grammatical rule violations, inappropriate word combinations, and other types. Semantic coherence annotation considers multiple dimensions: thematic coherence, logical coherence, and expression coherence. To increase sample diversity, data augmentation techniques are employed, including synonym replacement and sentence structure transformation.

[0055] This embodiment designs a complete loss function calculation mechanism. The loss of the grammatical error recognition branch uses weighted cross-entropy, and the loss function is: L_g = -Σ(w_i * y_i * log(p_i)), where w_i is the error type weight, y_i is the true label, and p_i is the predicted probability. The loss of the semantic coherence assessment branch also uses cross-entropy, but adds a sequential constraint to ensure the monotonicity of the evaluation results. The losses of the two branches are weighted and summed using learnable weight coefficients to obtain the total loss function.

[0056] This embodiment implements an efficient optimization strategy. It uses a modified stochastic gradient descent algorithm for parameter optimization, introducing a momentum term and an adaptive learning rate. A warmup strategy is used to gradually increase the learning rate during the initial training phase, followed by cosine annealing for further adjustment. To prevent vanishing and exploding gradients, gradient clipping is employed. A gradient accumulation mechanism is also implemented to support large-scale training. The performance metrics of the two tasks are monitored separately on the validation set, and an early stopping strategy is employed to prevent overfitting.

[0057] Through the above technological innovations, this embodiment effectively solves several key problems in project text analysis: insufficient feature representation, high task coupling, unstable training, etc. In practical applications, the solution can accurately identify grammatical errors and semantic coherence problems in the text, providing a reliable basis for subsequent optimization. It is particularly suitable for the analysis of highly professional application documents such as technological innovation projects and achievement transformation projects, and significantly improves the accuracy and comprehensiveness of text quality assessment. The multi-task learning framework of this solution enables it to handle multiple related tasks simultaneously, and improves the generalization ability and learning efficiency of the model through parameter sharing and knowledge transfer.

[0058] In one embodiment of the project text optimization processing method of the present application, the following contents may also be specifically included: Step S401: Preprocess the project text to be optimized. Use a word segmentation model to segment the text content, perform part-of-speech tagging on the segmentation results, identify the grammatical structure and dependency relationships of the sentences based on a syntactic analysis tree model, extract sentence boundary markers based on punctuation information, use a text segmentation algorithm to identify the start and end positions of paragraph structures, and divide the project text into text segments based on sentence boundary markers and paragraph structures. Step S402: Input the divided text segments into the multi-task text analysis model, extract the word vector features and syntactic features of each text segment, perform a grammar checking task based on the feature extraction results to mark grammatical errors, calculate the cosine similarity of the word vectors between adjacent text segments, and evaluate the semantic coherence score in combination with the syntactic features.

[0059] Optionally, this embodiment first designs an in-depth text preprocessing mechanism. The project text to be optimized is cleaned and standardized in multiple dimensions, including special character processing, format unification, and encoding conversion. An improved BERT-BiLSTM word segmentation model is used. This model has been specially trained on a technical document dataset and can accurately identify special content such as professional terms, compound words, and technical parameters. The model input includes character-level embedding and positional encoding. Contextual features are extracted through a multi-layer bidirectional LSTM, and finally a CRF layer is used to determine word segmentation boundaries.

[0060] This embodiment implements an innovative part-of-speech tagging solution. A part-of-speech tagger is constructed based on a pre-trained language model. The tagging set has been expanded to include special part-of-speech categories such as technical terms, parameter descriptions, and formula variables. The tagging process uses a conditional random field model, taking into account the dependencies and transition probabilities between words. The model's feature template includes multiple dimensions such as word form, context window, and character-level features. To improve the tagging accuracy of professional terms, a domain dictionary is introduced to assist in judgment.

[0061] This embodiment designs a precise syntactic analysis mechanism. A neural network-based syntactic analysis tree model is constructed, using a deep Transformer structure to extract the grammatical features of sentences. The model includes a multi-layer self-attention mechanism, focusing on different grammatical relationships through different attention heads. The construction of the syntactic tree adopts a transition system approach, predicting the transition action at each step based on the current state and the elements in the stack. The identification of dependency relationships pays special attention to the logical associations in the technical solution, such as causal relationships, conditional relationships, and progressive relationships.

[0062] This embodiment implements a reliable boundary identification method. A sentence boundary identification algorithm is designed by combining punctuation information and grammatical structure features. The algorithm not only considers conventional sentence-end punctuation but also analyzes the nested relationships between paired punctuation marks, such as brackets and quotation marks. For complex and long sentences, the completeness is analyzed using a dependency tree structure to avoid inappropriate segmentation. The criteria for identifying sentence boundaries include multiple dimensions, such as grammatical completeness, semantic independence, and expression coherence.

[0063] This embodiment innovatively designs a paragraph structure recognition mechanism. It employs a hierarchical text segmentation model based on a BiLSTM-CRF architecture. Input features include both text content features and format features. Content features are extracted using a pre-trained language model, while format features include information such as indentation, blank lines, and font. The model outputs a probability distribution for whether each position is a paragraph boundary. To improve recognition accuracy, a topic coherence constraint is introduced to ensure the rationality of paragraph segmentation.

[0064] This embodiment implements an efficient feature extraction process. The divided text segments are input into a multi-task analysis model, and a unified text representation is extracted through a shared encoder. The word vector features are dynamically fused, combining static word vectors and context-related word representations. Syntactic features are extracted from the dependency syntax tree through a graph neural network, including syntactic path features and dependency relationship features. The feature extraction process uses multi-layer feature fusion to build higher-level semantic representations layer by layer.

[0065] This embodiment designs a comprehensive grammar checking mechanism. It identifies grammatical errors based on extracted features and uses sequence labeling to predict the error type for each word. Error types include spelling errors, grammatical rule violations, and inappropriate collocations. The prediction process uses a CRF layer to model transition constraints between labels, ensuring consistent prediction results. For each identified error, information such as its location, type, and severity is recorded.

[0066] This embodiment implements an innovative semantic coherence evaluation scheme. When calculating the semantic similarity between adjacent text fragments, an improved cosine similarity calculation method is adopted. The calculation formula is: sim(v1,v2) = (v1·v2) / (||v1||·||v2||+λ), where v1 and v2 are vector representations of the text fragments, and λ is a smoothing factor. The vector representation combines word vector features and syntactic features, and dynamically weights different features through the attention mechanism. The evaluation process considers multiple dimensions: vocabulary overlap, topic consistency, naturalness of logical transitions, etc.

[0067] Through the above technical innovations, this embodiment effectively solves several key problems in project text analysis: inaccurate text segmentation, insufficient feature extraction, single evaluation indicators, etc. In practical applications, the solution can accurately analyze the structural features and semantic characteristics of project texts and identify key positions that need to be optimized. It is particularly suitable for the processing of highly professional application documents such as technological innovation projects and achievement transformation projects, and significantly improves the accuracy of text analysis and the comprehensiveness of evaluation. The modular design of the solution enables it to handle different types of text analysis needs, and through sufficient feature extraction and multi-dimensional evaluation, it provides a reliable basis for subsequent text optimization.

[0068] In one embodiment of the project text optimization processing method of the present application, the following contents may also be specifically included: Step S501: constructing a text grammatical error annotation matrix, encoding the position of the grammatical errors annotated in each text segment, calculating the semantic coherence scores between adjacent text segments, constructing the semantic coherence scores into a temporal feature sequence, modeling the temporal feature sequence using a long short-term memory network, calculating the state vector and forget gate output of each time step, and identifying the location of the semantic coherence breakpoint based on the change amplitude of the state vector; Step S502: Calculate the structural feature vector of the project text, encode the structural features of the templates in the differentiated text template library, calculate the similarity score between the project text structural feature vector and the template structural feature vector using cosine similarity, and select the text template with the highest similarity score as the optimized template.

[0069] Optionally, this embodiment first designs an innovative error labeling mechanism. A multidimensional matrix is used to represent grammatical errors in text segments, and the matrix dimension is [sequence length, number of error types]. Position encoding uses a relative position representation method, and the calculation formula is: PE(pos,i) = pos / 10000^(2i / d), where pos is the position of the word in the sequence, i is the dimension index, and d is the encoding dimension. Each error type corresponds to an independent channel, and the error position is marked by one-hot encoding. To improve the robustness of the representation, error degree weights are introduced to reflect the severity of different errors.

[0070] This embodiment implements a systemic semantic scoring scheme. When calculating the semantic coherence score between adjacent text segments, multiple evaluation dimensions are comprehensively considered: lexical cohesion, thematic coherence, and the naturalness of logical transitions. The scoring process uses a weighted summation approach, with the weights of each dimension dynamically calculated using an attention mechanism. The score sequence is constructed as a temporal feature, with each time step corresponding to a coherence vector between segments. The vector dimensions contain scores for multiple evaluation indicators.

[0071] This embodiment designs a deep temporal modeling method. An improved LSTM network is used to process the semantic coherence score sequence. The network has a hidden dimension of 256 and contains multiple LSTM layers. The state update formula for each time step is: h_t = tanh(W_h[h_(t-1); x_t] + b_h) f_t = σ(W_f[h_(t-1); x_t] + b_f) Where h_t is the hidden layer state, f_t is the forget gate output, x_t is the input feature, and W and b are learnable parameters. The gating mechanism controls the retention and updating of historical information and captures long-range semantic dependencies.

[0072] This embodiment implements a precise breakpoint identification mechanism. Based on the state vector of the LSTM network, the magnitude of state change between adjacent time steps is calculated. This magnitude of change is measured using the Euclidean distance: d_t = ||h_t - h_(t-1)||_2. An adaptive threshold θ is set. When the magnitude of change exceeds the threshold, the location is marked as a semantic breakpoint. The threshold is determined by considering the overall coherence level of the text and local variation characteristics, using a dynamic adjustment strategy. To improve recognition reliability, a context window constraint is introduced.

[0073] This embodiment innovatively designs a structural feature extraction scheme. The structural features of the project text are vectorized, with feature dimensions encompassing multiple levels: document-level features (such as chapter organization and logical hierarchy), paragraph-level features (such as paragraph function and cohesion), and sentence-level features (such as sentence type and tone). Feature extraction utilizes a hierarchical neural network, generating a unified structural representation through multi-layer feature fusion. To enhance the expressive power of features, a position-aware attention mechanism is introduced.

[0074] This embodiment implements a comprehensive template encoding method. Structural feature encoding is performed on each template in the differentiated text template library. The encoding process uses the same feature extraction framework as the project text to ensure consistency in the feature space. The structural features of the template include not only static organizational structure but also dynamic expression features. The encoding results form a template feature matrix that supports efficient similarity calculation and retrieval. To improve the accuracy of the encoding, template type information is introduced as an auxiliary feature.

[0075] This embodiment designs an efficient similarity calculation mechanism. It uses an improved cosine similarity to calculate the degree of structural match between the project text and the template. The calculation formula is: sim(d, t) = (d·t) / (||d||·||t||+λ), where d is the document feature vector, t is the template feature vector, and λ is the smoothing factor. Feature weights are introduced into the similarity calculation process to highlight the importance of key structural features. To improve computational efficiency, batch computing and parallel processing mechanisms are implemented.

[0076] This embodiment implements an innovative template selection strategy. When screening templates based on similarity scores, a multi-round screening mechanism is adopted. The first round is a rough selection based on structural similarity, selecting the candidate templates with the top K scores. The second round considers the technical characteristics of the text and analyzes the degree of match between the template and the text in terms of the depth of technical description and the display of innovative points. The final round combines the project type and application requirements to select the most suitable optimized template. Constraints are set during the selection process to ensure the rationality of the selection results.

[0077] Through the above technical innovations, this embodiment effectively solves several key problems in project text analysis: inaccurate recognition of semantic breaks, insufficient extraction of structural features, inaccurate template matching, etc. In practical applications, this solution can accurately analyze the semantic coherence problems of project texts, identify key positions that need to be optimized, and select appropriate optimization templates. It is particularly suitable for the processing of highly professional application documents such as technological innovation projects and achievement transformation projects, and significantly improves the accuracy of text analysis and the precision of template matching. The adaptive characteristics of this solution enable it to handle different types of project application requirements, and through continuous optimization and improvement, it provides reliable technical support for improving the quality of project application documents.

[0078] In one embodiment of the project text optimization processing method of the present application, the following contents may also be specifically included: Step S601: Construct a text state vector. The state input is obtained by concatenating the grammatical error position matrix, the semantic coherence breakpoint position vector, and the text template features. The state input is encoded using a policy network to generate a probability distribution of text repair actions. The action space is sampled using a Monte Carlo tree search algorithm. An action decision tree is constructed based on the sampling results, and the value score of each action node is calculated. Step S602: Train the policy network, use the similarity between the optimized text and the original text as the reward function, calculate the value estimation of each state-action pair based on the temporal difference algorithm, backpropagate the value estimation results to update the policy network parameters, and convert the trained policy network output into a text repair action sequence.

[0079] Optionally, this embodiment first designs an innovative state representation mechanism. The grammatical error position matrix, the semantic breakpoint position vector, and the template features are multimodally fused to construct a unified state vector. The error position matrix uses a sparse representation, and each position contains information about the error type and severity. The breakpoint position is marked with a binary vector, where 1 represents the breakpoint position. The template feature contains two parts: structural features and expression features. Feature fusion uses an attention mechanism, and the calculation formula is: F = Attention(We[E; D; T]), where E, D, and T are error features, breakpoint features, and template features, respectively, and We is a learnable weight matrix.

[0080] This embodiment implements a deep policy network structure. The policy network is constructed using an Actor-Critic framework. The Actor network is responsible for generating the probability distribution of repair actions, and the Critic network evaluates state-action values. The Actor network is based on the Transformer architecture and contains multiple layers of encoders and decoders. The encoder processes state inputs and captures the correlation between features through a multi-head self-attention mechanism. The decoder generates the probability distribution of the next action based on the encoding results and the historical action sequence. The network output is normalized through a softmax layer to obtain the probability distribution in the action space.

[0081] This embodiment designs an accurate action sampling strategy. Action exploration is performed based on an improved Monte Carlo tree search algorithm. Each node in the search tree represents a state-action pair. The value estimate of a node is calculated using the following formula: Q(s,a) = R(s,a) + γ V(s'), where R is the immediate reward, γ is the discount factor, and V(s') is the value estimate of the next state. The sampling process uses the UCB criterion to select actions. The criterion formula is: UCB = Q(s,a) + c sqrt(ln(N) / n), where N is the total number of visits, n is the number of action visits, and c is the exploration coefficient.

[0082] This embodiment implements an innovative decision tree construction method. An action decision tree is constructed based on the sampling results, with each path in the tree representing a possible repair sequence. The node's value score comprehensively considers multiple factors, including the feasibility of the action, the repair effect, and its conformance to the template. The score is calculated using a weighted summation method, with the weights of each factor learned through network learning. To improve decision reliability, a path constraint mechanism is introduced to filter out unreasonable action sequences.

[0083] This embodiment innovatively designs a reward calculation mechanism. The similarity between the optimized text and the original text is used as the reward signal. The similarity calculation considers multiple dimensions: lexical similarity, semantic similarity, structural similarity, etc. The similarity calculation formula is: sim = w1 sim_w + w2sim_s + w3*sim_t, where sim_w, sim_s, and sim_t are lexical, semantic, and structural similarities, respectively, and w is the weight coefficient. To prevent excessive modification, a similarity threshold constraint is set.

[0084] This embodiment implements an efficient value estimation method. The value of a state-action pair is calculated using a temporal difference algorithm. The estimation formula is: Q_t = Q_(t-1) + α*(R_t + γ*Q_(t+1) - Q_(t-1)), where α is the learning rate, R_t is the immediate reward, and γ is the discount factor. The estimation process considers long-term benefits, balancing immediate rewards with future benefits using the discount factor. To improve estimation accuracy, an experience replay mechanism is employed, randomly sampling from historical trajectories for learning.

[0085] This embodiment designs a complete parameter update mechanism. It uses a policy gradient method to update network parameters. The gradient calculation formula is: ∇J(θ) = E[∇logπ(a|s;θ)*A(s,a)], where π is the policy function and A is the advantage function. The update process uses an adaptive learning rate, dynamically adjusted based on the gradient variance. To improve training stability, gradient clipping and weight decay are introduced. Multi-process parallel training is also implemented to accelerate convergence.

[0086] This embodiment implements a reliable action sequence generation method. The trained policy network output is converted into a specific repair action sequence. This conversion process utilizes a beam search algorithm to retain multiple high-probability candidate sequences. Each action contains a specific operation type and parameters, such as error correction, semantic supplementation, and structural adjustment. Constraints are set during the generation process to ensure the rationality and enforceability of the action sequence.

[0087] Through the above technical innovations, this embodiment effectively solves several key problems in project text optimization: inflexible repair strategies, inaccurate action decisions, unstable optimization effects, etc. In practical applications, the solution can generate targeted repair strategies based on the specific characteristics of text problems, and achieve intelligent optimization of text quality. It is particularly suitable for the optimization of highly professional application documents such as technological innovation projects and achievement transformation projects, and significantly improves the automation level and optimization effect of text optimization. The adaptive characteristics of the solution enable it to handle different types of text optimization needs, and through continuous learning and optimization, it provides reliable technical support for improving the quality of project application documents.

[0088] In one embodiment of the project text optimization processing method of the present application, the following contents may also be specifically included: Step S701: According to the text repair action sequence, the marked grammatical error locations are corrected, the semantic connectives in the text template are used to optimize the connection of semantic coherence breakpoints, the word order of the repaired text segments is adjusted based on semantic similarity calculation, and the optimized text segments are spliced and reorganized according to the original text structure order to generate a complete text including paragraph markers; Step S702: Input the reorganized complete text into the format standardization check model, extract the format feature vector of the text, perform standardization verification on the text's format elements such as font, font size, indentation, paragraph spacing, etc., identify the positions of format elements that do not meet the standards, automatically adjust the format elements that fail the verification according to the preset format rule template, and output the optimized project text that meets the standards.

[0089] Optionally, this embodiment first designs an accurate grammatical error correction mechanism. According to the correction instructions in the repair action sequence, the marked grammatical error locations are corrected one by one. The error correction process adopts a rule-based and neural network-based approach. The rule base contains common grammatical error patterns and corresponding correction strategies. For complex grammatical problems, a pre-trained language model is used to generate correction suggestions. The model is based on the BERT architecture and is fine-tuned through mask prediction tasks, which improves the accuracy of error correction in technical document scenarios.

[0090] This embodiment implements an innovative semantic cohesion optimization solution. A semantic connective knowledge base is extracted from the text template, containing different types of connectives and their usage scenarios. Connectives are selected based on contextual semantic similarity calculations using the formula: sim(c,s) = cos(v_c, v_s), where v_c is the connective vector and v_s is the contextual semantic vector. To ensure natural cohesion, a connective applicability scoring mechanism is introduced, taking into account both semantic matching and frequency of use.

[0091] This embodiment designs an efficient word order adjustment strategy. It uses an improved semantic similarity algorithm to calculate the strength of association between text segments and construct a segment association graph. Nodes in the graph represent text segments, and edge weights reflect the degree of semantic association. A graph sorting algorithm is used to optimize the segment order to ensure semantic fluency. The sorting process considers multiple constraints, including logical coherence, thematic consistency, and naturalness of expression. A local adjustment mechanism is also introduced to optimize local word order while maintaining the global structure.

[0092] This embodiment implements a reliable text reorganization method. The reorganization strategy is designed based on the structural order of the original text, preserving the overall framework and hierarchical relationship of the document. The reorganization process uses a dynamic programming algorithm to optimize the combination scheme of the fragments. The combination objective function is: F(S) = w1 C(S) + w2L(S) + w3*T(S), where C is the coherence score, L is the logic score, T is the structural integrity score, and w is the weight coefficient. The generation of paragraph markers takes into account semantic integrity and expression habits.

[0093] This embodiment innovatively designs a format feature extraction mechanism. A deep convolutional neural network is constructed to extract text format features. The network includes multiple convolutional and pooling layers. The input features include visual and structural features of the text, such as font attributes, spatial layout, and paragraph structure. The dimensionality of the feature vector is designed to take into account the representation requirements of different format elements. To improve the expressiveness of features, an attention mechanism is introduced to highlight the importance of key format features.

[0094] This embodiment implements a rigorous compliance verification process. Multi-dimensional verification is performed based on a pre-set format specification library, which contains all the format requirements for project application documents. The verification process utilizes a combination of a rules engine and deep learning. The rules engine handles clear format specifications, while the deep learning model handles complex format judgments. Verification results are output as a probability distribution, with the probability value reflecting the degree of format compliance.

[0095] This embodiment designs an intelligent format adjustment mechanism. For format elements that fail verification, an adjustment strategy generator is constructed. This adjustment strategy is based on a preset format rule template, which contains solutions for common formatting issues. The adjustment process adopts a progressive strategy, first addressing basic formatting issues and then optimizing complex formatting elements. Real-time verification is performed after each adjustment to ensure that the adjustment meets the requirements. Format consistency checks are also implemented to ensure uniform formatting across the entire document.

[0096] This embodiment implements an efficient output processing flow. The optimized project text undergoes a final quality check, which includes text completeness, formatting compliance, and accuracy of expression. The output process supports multiple document formats to meet the needs of different application scenarios. A version control mechanism is also implemented to record important modifications during the optimization process, supporting backtracking and comparison. To enhance practicality, modification suggestions are also provided to help users understand the specific content of the optimization.

[0097] Through the above technical innovations, this embodiment effectively solves several key problems in project text optimization: inaccurate grammatical correction, unnatural semantic connection, poor format standardization, etc. In practical applications, this solution can achieve an all-round improvement in text quality and generate standardized and professional project application documents. It is particularly suitable for scenarios with high requirements for document quality, such as technological innovation projects and achievement transformation projects, and significantly improves the readability and standardization of the text. The adaptive characteristics of this solution enable it to handle different types of project application requirements, and through continuous optimization and improvement, it provides reliable technical support for improving the quality of project application documents.

[0098] In order to effectively solve the shortcomings of traditional technologies in semantic coherence evaluation, optimization strategy formulation, etc., and significantly improve the professionalism and efficiency of project text optimization, this application provides an embodiment of a project text optimization processing device for implementing all or part of the content of the project text optimization processing method, see Figure 2 The project text optimization processing device specifically includes the following contents: A pre-training module 10 is used to construct a training corpus, extract text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, use a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculate weight coefficients for the knowledge association edges, and cluster the texts in the corpus based on the weight coefficients to form a differentiated text template library, input the text structure features, semantic features, and grammatical rules into a multi-task text analysis model for training, use a cross-entropy loss function to optimize the parameters of the multi-task text analysis model, and obtain the grammatical error recognition accuracy and semantic coherence evaluation accuracy of the multi-task text analysis model on a validation set; The text processing module 20 is configured to segment the project text to be optimized, identify sentence boundaries and paragraph structure based on part-of-speech tagging and syntactic analysis, divide the project text into text segments, input the segments into the multi-task text analysis model to perform grammar checking and semantic coherence assessment tasks, obtain grammatical error location annotations in the text segments and semantic coherence scores between adjacent segments, perform sequence modeling on the semantic coherence scores using a recurrent neural network, identify semantic coherence breakpoints, and screen text templates that match the project text structure based on the differentiated text template library; The text optimization module 30 is used to construct a text optimization strategy network based on the reinforcement learning algorithm, take the grammatical error position, semantic coherence breakpoint and text template as state input, output a text repair action sequence, perform grammatical error correction and semantic connection optimization on the project text according to the action sequence, reorganize the repaired text fragments into a complete text, use a text format standardization check model to verify the format of the complete text, adjust the format of the part that fails the verification, and obtain the optimized project text.

[0099] From the above description, it can be seen that the project text optimization processing device provided by the embodiment of the present application can deeply analyze the text structure features and semantic features through knowledge graphs and relationship extraction to form a differentiated text template library. A segmented processing strategy is adopted, combining part-of-speech tagging and syntactic analysis to identify text boundaries, and a recurrent neural network is used to perform sequence modeling on semantic coherence to accurately locate semantic breakpoints. A text optimization strategy network is constructed based on a reinforcement learning algorithm, and the grammatical error position, semantic coherence breakpoint and text template are used as state inputs to dynamically generate repair strategies and realize intelligent optimization of text. This method effectively solves the shortcomings of traditional technologies in semantic coherence evaluation, optimization strategy formulation, etc., and significantly improves the professionalism and efficiency of project text optimization.

[0100] From a hardware perspective, in order to effectively address the shortcomings of traditional technologies in semantic coherence assessment and optimization strategy formulation, and significantly improve the professionalism and efficiency of project text optimization, this application provides an embodiment of an electronic device for implementing all or part of the project text optimization processing method. The electronic device specifically includes the following: A processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to transmit information between the project text optimization processing device and related devices such as core business systems, user terminals, and related databases; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., but this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the project text optimization processing method and the embodiments of the project text optimization processing device in the embodiments, the contents of which are incorporated herein and repeated parts are not repeated.

[0101] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0102] In practical applications, part of the project text optimization processing method can be executed on the electronic device side as described above, or all operations can be completed on the client device. The specific selection can be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application does not impose any restrictions on this. If all operations are completed on the client device, the client device may also include a processor.

[0103] The aforementioned client device may include a communication module (i.e., a communication unit) capable of establishing a communication connection with a remote server to facilitate data transmission with the server. The server may include a server at the task scheduling center or, in other implementation scenarios, a server on an intermediate platform, such as a server on a third-party server platform that is communicatively linked to the task scheduling center server. The server may comprise a single computer device, a server cluster consisting of multiple servers, or a distributed server configuration.

[0104] Figure 3 Schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 3 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that the Figure 3 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0105] In one embodiment, the project text optimization processing method function may be integrated into the central processing unit 9100. The central processing unit 9100 may be configured to perform the following control: Step S101: constructing a training corpus, extracting text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, using a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculating weight coefficients for the knowledge association edges, and clustering the texts in the corpus based on the weight coefficients to form a differentiated text template library, inputting the text structure features, semantic features, and grammatical rules into a multi-task text analysis model for training, using a cross-entropy loss function to optimize the parameters of the multi-task text analysis model, and obtaining the grammatical error recognition accuracy and semantic coherence evaluation accuracy of the multi-task text analysis model on a validation set; Step S102: Segment the project text to be optimized, identify sentence boundaries and paragraph structures based on part-of-speech tagging and syntactic analysis, divide the project text into text segments, input the multi-task text analysis model to perform grammar checking and semantic coherence assessment tasks, obtain grammatical error position annotations in the text segments and semantic coherence scores between adjacent segments, use a recurrent neural network to perform sequence modeling on the semantic coherence scores, identify semantic coherence breakpoints, and screen text templates that match the project text structure based on the differentiated text template library; Step S103: Construct a text optimization strategy network based on the reinforcement learning algorithm, take the grammatical error position, semantic coherence breakpoint and text template as state input, output a text repair action sequence, perform grammatical error correction and semantic coherence optimization on the project text according to the action sequence, reorganize the repaired text fragments into a complete text, use a text format standardization check model to verify the format of the complete text, adjust the format of the part that fails the verification, and obtain the optimized project text.

[0106] As can be seen from the above description, the electronic device provided in the embodiment of the present application deeply analyzes the text structure features and semantic features through knowledge graphs and relationship extraction to form a differentiated text template library. A segmented processing strategy is adopted, combining part-of-speech tagging and syntactic analysis to identify text boundaries, and a recurrent neural network is used to perform sequence modeling on semantic coherence to accurately locate semantic breakpoints. A text optimization strategy network is constructed based on a reinforcement learning algorithm, and the grammatical error position, semantic coherence breakpoint and text template are used as state inputs to dynamically generate repair strategies and realize intelligent optimization of text. This method effectively solves the shortcomings of traditional technologies in semantic coherence evaluation, optimization strategy formulation, etc., and significantly improves the professionalism and efficiency of project text optimization.

[0107] In another embodiment, the project text optimization processing device can be configured separately from the central processing unit 9100. For example, the project text optimization processing device can be configured as a chip connected to the central processing unit 9100, and the function of the project text optimization processing method can be implemented through the control of the central processing unit.

[0108] like Figure 3 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Figure 3 In addition, the electronic device 9600 may also include all components shown in Figure 3 For components not shown, reference may be made to the prior art.

[0109] like Figure 3 As shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.

[0110] Memory 9140 can be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It can store the aforementioned failure-related information and also store programs that execute the relevant information. The CPU 9100 can execute the programs stored in memory 9140 to implement information storage or processing.

[0111] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 may be, for example, a keypad or touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.

[0112] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), or SIM card. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is capable of storing additional data. Examples of such memory are sometimes referred to as EPROMs. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs, or processes used by the central processing unit 9100 to execute operations of the electronic device 9600.

[0113] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, images, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various driver programs for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0114] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.

[0115] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless local area network modules. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130, providing audio output via the speaker 9131 and receiving audio input from the microphone 9132, thereby implementing common telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 9130 is coupled to the central processing unit 9100, enabling local recording via the microphone 9132 and playback of stored audio via the speaker 9131.

[0116] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the project text optimization processing method in the above-mentioned embodiments, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements all steps of the project text optimization processing method in the above-mentioned embodiments, where the execution subject is a server or a client. For example, when the processor executes the computer program, the following steps are implemented: Step S101: constructing a training corpus, extracting text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, using a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculating weight coefficients for the knowledge association edges, and clustering the texts in the corpus based on the weight coefficients to form a differentiated text template library, inputting the text structure features, semantic features, and grammatical rules into a multi-task text analysis model for training, using a cross-entropy loss function to optimize the parameters of the multi-task text analysis model, and obtaining the grammatical error recognition accuracy and semantic coherence evaluation accuracy of the multi-task text analysis model on a validation set; Step S102: Segment the project text to be optimized, identify sentence boundaries and paragraph structures based on part-of-speech tagging and syntactic analysis, divide the project text into text segments, input the multi-task text analysis model to perform grammar checking and semantic coherence assessment tasks, obtain grammatical error position annotations in the text segments and semantic coherence scores between adjacent segments, use a recurrent neural network to perform sequence modeling on the semantic coherence scores, identify semantic coherence breakpoints, and screen text templates that match the project text structure based on the differentiated text template library; Step S103: Construct a text optimization strategy network based on the reinforcement learning algorithm, take the grammatical error position, semantic coherence breakpoint and text template as state input, output a text repair action sequence, perform grammatical error correction and semantic coherence optimization on the project text according to the action sequence, reorganize the repaired text fragments into a complete text, use a text format standardization check model to verify the format of the complete text, adjust the format of the part that fails the verification, and obtain the optimized project text.

[0117] As can be seen from the above description, the computer-readable storage medium provided in the embodiment of the present application deeply analyzes the text structure features and semantic features through knowledge graphs and relationship extraction to form a differentiated text template library. A segmented processing strategy is adopted, combining part-of-speech tagging and syntactic analysis to identify text boundaries, and a recurrent neural network is used to perform sequence modeling on semantic coherence to accurately locate semantic breakpoints. A text optimization strategy network is constructed based on a reinforcement learning algorithm, and the grammatical error position, semantic coherence breakpoint and text template are used as state inputs to dynamically generate repair strategies and realize intelligent optimization of text. This method effectively solves the shortcomings of traditional technologies in semantic coherence evaluation, optimization strategy formulation, etc., and significantly improves the professionalism and efficiency of project text optimization.

[0118] The embodiments of the present application also provide a computer program product capable of implementing all steps of the project text optimization processing method in the above-mentioned embodiment, where the execution subject is a server or a client. When the computer program / instructions are executed by a processor, the steps of the project text optimization processing method are implemented. For example, the computer program / instructions implement the following steps: Step S101: constructing a training corpus, extracting text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, using a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculating weight coefficients for the knowledge association edges, and clustering the texts in the corpus based on the weight coefficients to form a differentiated text template library, inputting the text structure features, semantic features, and grammatical rules into a multi-task text analysis model for training, using a cross-entropy loss function to optimize the parameters of the multi-task text analysis model, and obtaining the grammatical error recognition accuracy and semantic coherence evaluation accuracy of the multi-task text analysis model on a validation set; Step S102: Segment the project text to be optimized, identify sentence boundaries and paragraph structures based on part-of-speech tagging and syntactic analysis, divide the project text into text segments, input the multi-task text analysis model to perform grammar checking and semantic coherence assessment tasks, obtain grammatical error position annotations in the text segments and semantic coherence scores between adjacent segments, use a recurrent neural network to perform sequence modeling on the semantic coherence scores, identify semantic coherence breakpoints, and screen text templates that match the project text structure based on the differentiated text template library; Step S103: Construct a text optimization strategy network based on the reinforcement learning algorithm, take the grammatical error position, semantic coherence breakpoint and text template as state input, output a text repair action sequence, perform grammatical error correction and semantic coherence optimization on the project text according to the action sequence, reorganize the repaired text fragments into a complete text, use a text format standardization check model to verify the format of the complete text, adjust the format of the part that fails the verification, and obtain the optimized project text.

[0119] As can be seen from the above description, the computer program product provided by the embodiment of the present application deeply analyzes the structural features and semantic features of the text through knowledge graphs and relationship extraction to form a differentiated text template library. A segmented processing strategy is adopted, combining part-of-speech tagging and syntactic analysis to identify text boundaries, and a recurrent neural network is used to perform sequence modeling of semantic coherence to accurately locate semantic breakpoints. A text optimization strategy network is constructed based on a reinforcement learning algorithm, and the grammatical error position, semantic coherence breakpoint and text template are used as state inputs to dynamically generate repair strategies and realize intelligent optimization of text. This method effectively solves the shortcomings of traditional technologies in semantic coherence evaluation, optimization strategy formulation, etc., and significantly improves the professionalism and efficiency of project text optimization.

[0120] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatuses, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0121] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0122] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0124] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

Claims

1. A project text optimization processing method, characterized in that: The method comprises: Construct a training corpus, extract text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, use a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculate weight coefficients for the knowledge association edges, and cluster the texts in the corpus based on the weight coefficients to form a differentiated text template library, input the text structure features, semantic features, and grammatical rules into a multi-task text analysis model for training, use a cross-entropy loss function to optimize the parameters of the multi-task text analysis model, and obtain the grammatical error recognition accuracy and semantic coherence assessment accuracy of the multi-task text analysis model on a validation set; The project text to be optimized is segmented, sentence boundaries and paragraph structure are identified based on part-of-speech tagging and syntactic analysis, the project text is divided into text segments, and the segments are input into the multi-task text analysis model to perform grammar checking and semantic coherence assessment tasks, obtaining grammatical error position annotations in the text segments and semantic coherence scores between adjacent segments, using a recurrent neural network to perform sequence modeling on the semantic coherence scores, identifying semantic coherence breakpoints, and screening text templates that match the project text structure based on the differentiated text template library; A text optimization strategy network is constructed based on the reinforcement learning algorithm. The grammatical error position, semantic coherence breakpoint and text template are used as state input, and a text repair action sequence is output. According to the action sequence, the project text is subjected to grammatical correction and semantic coherence optimization. The repaired text fragments are reorganized into a complete text. The format of the complete text is verified using a text format standardization check model. The format of the parts that fail the verification is adjusted to obtain the optimized project text.

2. The project text optimization processing method according to claim 1, characterized in that: The method comprises: constructing a training corpus, extracting text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, using a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculating weight coefficients for the knowledge association edges, and clustering the texts in the corpus based on the weight coefficients to form a differentiated text template library, including: Establish a data interface with the project application text library and the external standard library, collect text data to build a basic corpus, perform word segmentation and part-of-speech tagging on the text content in the basic corpus, extract word vector features, topic features, and syntactic tree features to build a text structure feature matrix, use a word embedding model and a deep semantic model to extract the semantic features of the text, and input the text structure feature matrix and semantic features into the graph neural network to generate node representations of the knowledge graph; The text in the basic corpus is input into the relationship extraction model to identify the technical elements and their associations in the text, extract the technical structure association features and key point dependency features to construct knowledge association edges, and use the attention mechanism to calculate the weight coefficient of the knowledge association edge. The weight coefficient is input into the spectral clustering algorithm to cluster the text in the corpus, and a differentiated text template library is constructed based on the clustering results.

3. The project text optimization processing method according to claim 1, characterized in that: The text structure features, semantic features, and grammatical rules are input into a multi-task text analysis model for training, a cross-entropy loss function is used to optimize the parameters of the multi-task text analysis model, and the grammatical error recognition accuracy and semantic coherence evaluation accuracy of the multi-task text analysis model are obtained on a validation set, including: Construct a network structure for a multi-task text analysis model. Fusion is performed on the outputs of the text structure feature extraction layer, semantic feature extraction layer, and grammatical rule encoding layer. A shared encoder is used to map the fused features to obtain a text representation vector. The text representation vector is input into the grammatical error recognition branch and the semantic coherence assessment branch, respectively. A softmax classifier is added to the output layer of each branch. The multi-task text analysis model is trained based on labeled training samples. The cross-entropy loss of the grammatical error recognition branch and the semantic coherence evaluation branch is calculated. The losses of the two branches are weighted summed to obtain the total loss function. The model parameters are optimized using the stochastic gradient descent method. The grammatical error recognition accuracy and the semantic coherence evaluation accuracy are calculated on the validation set, respectively.

4. The project text optimization processing method according to claim 1, characterized in that: The project text to be optimized is segmented, sentence boundaries and paragraph structures are identified based on part-of-speech tagging and syntactic analysis, the project text is divided into text segments, and the segments are input into the multi-task text analysis model to perform grammar checking and semantic coherence evaluation tasks, including: Preprocess the text of the project to be optimized. Use a word segmentation model to segment the text content, perform part-of-speech tagging on the segmentation results, identify the grammatical structure and dependency relationships of the sentences based on the syntactic analysis tree model, extract sentence boundary markers based on punctuation information, use a text segmentation algorithm to identify the start and end positions of the paragraph structure, and divide the project text into text segments based on the sentence boundary markers and paragraph structure. The divided text segments are input into the multi-task text analysis model, the word vector features and syntactic features of each text segment are extracted, the grammar checking task is performed based on the feature extraction results to mark grammatical errors, the cosine similarity of the word vectors between adjacent text segments is calculated, and the semantic coherence score is evaluated in combination with the syntactic features.

5. The project text optimization processing method according to claim 1, characterized in that: The method includes obtaining grammatical error position annotations in text segments and semantic coherence scores between adjacent segments, performing sequence modeling on the semantic coherence scores using a recurrent neural network, identifying semantic coherence breakpoints, and screening a text template that matches the project text structure based on the differentiated text template library, including: Construct a text grammatical error annotation matrix, encode the position of the grammatical errors annotated in each text segment, calculate the semantic coherence scores between adjacent text segments, construct the semantic coherence scores into a temporal feature sequence, model the temporal feature sequence using a long short-term memory network, calculate the state vector and forget gate output for each time step, and identify the location of the semantic coherence breakpoint based on the change amplitude of the state vector; Calculate the structural feature vector of the project text, encode the structural features of the templates in the differentiated text template library, use cosine similarity to calculate the similarity score between the project text structural feature vector and the template structural feature vector, and select the text template with the highest similarity score as the optimized template.

6. The project text optimization processing method according to claim 1, characterized in that: The text optimization strategy network is constructed based on the reinforcement learning algorithm, which takes the grammatical error location, semantic coherence breakpoint and text template as state input and outputs a text repair action sequence, including: Construct a text state vector, concatenate the grammatical error position matrix, the semantic coherence breakpoint position vector, and the text template features to obtain the state input. Use a policy network to encode the state input and generate a probability distribution of text repair actions. Use the Monte Carlo tree search algorithm to sample the action space, construct an action decision tree based on the sampling results, and calculate the value score of each action node. The policy network is trained, and the similarity between the optimized text and the original text is used as the reward function. The value estimation of each state-action pair is calculated based on the temporal difference algorithm. The value estimation results are backpropagated to update the policy network parameters, and the output of the trained policy network is converted into a text repair action sequence.

7. The project text optimization processing method according to claim 1, characterized in that: The step of performing grammatical correction and semantic cohesion optimization on the project text according to the action sequence, reorganizing the repaired text fragments into a complete text, verifying the format of the complete text using a text format standardization check model, and adjusting the format of the parts that fail the verification to obtain the optimized project text includes: According to the text repair action sequence, the marked grammatical errors are corrected, the semantic connectives in the text template are used to optimize the connection of semantic coherence breakpoints, the word order of the repaired text fragments is adjusted based on semantic similarity calculation, and the optimized text fragments are spliced and reorganized according to the original text structure order to generate a complete text including paragraph markers; The reorganized complete text is input into the format standardization check model, the format feature vector of the text is extracted, and the format elements such as font, font size, indentation, paragraph spacing, etc. of the text are verified for standardization. The positions of format elements that do not meet the standards are identified, and the format elements that fail to pass the verification are automatically adjusted according to the preset format rule template, and the optimized project text that meets the standards is output.

8. A project text optimization processing device, characterized in that: The device comprises: A pre-training module is used to construct a training corpus, extract text structure features and semantic features from a project application text library and an external standard library to construct a knowledge graph, use a relationship extraction model to identify technical structure associations and key point dependency relationships in the text to construct knowledge association edges, calculate weight coefficients for the knowledge association edges, and cluster the texts in the corpus based on the weight coefficients to form a differentiated text template library, input the text structure features, semantic features, and grammatical rules into a multi-task text analysis model for training, use a cross-entropy loss function to optimize the parameters of the multi-task text analysis model, and obtain the grammatical error recognition accuracy and semantic coherence assessment accuracy of the multi-task text analysis model on a validation set; A text processing module is used to segment the text of the project to be optimized, identify sentence boundaries and paragraph structure based on part-of-speech tagging and syntactic analysis, divide the project text into text segments, input the multi-task text analysis model to perform grammar checking and semantic coherence assessment tasks, obtain grammatical error location annotations in the text segments and semantic coherence scores between adjacent segments, use a recurrent neural network to perform sequence modeling on the semantic coherence scores, identify semantic coherence breakpoints, and screen text templates that match the project text structure based on the differentiated text template library; The text optimization module is used to build a text optimization strategy network based on the reinforcement learning algorithm, take the grammatical error position, semantic coherence breakpoint and text template as state input, output text repair action sequence, perform grammatical error correction and semantic connection optimization on the project text according to the action sequence, reorganize the repaired text fragments into a complete text, use the text format standardization check model to verify the format of the complete text, adjust the format of the part that fails the verification, and obtain the optimized project text.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the project text optimization processing method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the project text optimization processing method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Protocol conversion and communication adaptation system for optical storage and charging grid-connected device

    CN120729963A

  • Semantic topology calculation method based on theme-knowledge dual drive and related equipment

    CN121706784A