Adaptive perception event element extraction method
By constructing multi-domain data sets and sentence complexity evaluation, dynamically selecting traditional or large models for event elements extraction, solving the problem of insufficient adaptability to sentence complexity in the existing technology, and achieving efficient and accurate event elements extraction.
Patent Information
- Application Number
- CN202510203233.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-18
AI Technical Summary
The existing event factor extraction methods are inaccurate and inefficient when dealing with complex sentence patterns and nested structures. Traditional methods are acceptable when performing simple sentences, but large models are costly and inefficient when dealing with simple sentences, making it difficult to dynamically adapt to different sentence complexities.
Build a multi-domain event element extraction data set, train traditional deep learning models and large models, select the adapted models through sentence complexity evaluation for extraction, including feature extraction, model training, sentence complexity evaluation and result integration, and dynamically select the model to adapt to different sentence complexity.
It realizes efficient and accurate extraction of sentences of different complexity, improves the accuracy and adaptability of event factor extraction, adapts to complex and changeable text data and cross-domain needs, and improves the efficiency and quality of extraction.
Smart Images

Figure CN120337901A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to an event element extraction method with adaptive perception. Background Art
[0002] In natural language processing, event element extraction is an important task, aiming to accurately extract key elements related to events from text, such as time, place, person, trigger word, event type, etc. This technology is of great significance for applications such as information retrieval, knowledge graph construction, and intelligent question answering. With the rapid development of the Internet, the demand for processing massive text data is increasing day by day, posing higher requirements for the accuracy and efficiency of event element extraction. Traditional event element extraction methods mostly rely on predefined rules or templates. This method performs well when dealing with sentences with clear structures and simple grammars, but its performance often deteriorates significantly when facing complex sentence patterns, nested structures, or implicit information.
[0003] In recent years, event element extraction models based on deep learning have gradually become a research hotspot. Among them, small or medium-sized traditional models, such as RoBerta+CRF, etc., perform excellently when dealing with simple sentences due to their high computational efficiency and few parameters; while large models, such as ChatGLM, Qwen, etc., show higher accuracy when dealing with complex sentences containing parallel nested relationships, etc., due to their strong feature learning ability and context understanding ability. However, directly applying large models to process all sentences is not only costly, but also may be too redundant and generate hallucinations when dealing with simple sentences, resulting in low efficiency. Summary of the Invention
[0004] To solve the above problems, the present invention proposes an event element extraction method with adaptive perception, aiming to dynamically and differentially call more suitable models for event element extraction according to the complexity of sentences, so as to improve the accuracy and efficiency of event element extraction.
[0005] To achieve the above object, an event element extraction method with adaptive perception of the present invention includes the following steps:
[0006] Step S1: Data collection and processing, collecting data from open-source resources in multiple fields, and constructing a high-quality multi-field event element extraction dataset after cleaning and annotation;
[0007] Step S2: Event element extraction model training, using the constructed multi-field event element extraction dataset to pre-train and fine-tune two different types of models, the two different types of models being traditional deep learning models and large models;
[0008] Step S3: Sentence complexity assessment. By analyzing the length, syntactic structure, dependency relationship, and lexical richness characteristics of sentences, complexity assessment metrics are constructed. Based on these metrics, the model determines the complexity level of each sentence and classifies it;
[0009] Step S4: Adaptive model perception. According to the results of sentence complexity assessment, an adaptive model is selected. For instances classified as simple sentences, a traditional deep learning model is called for event element extraction. For instances classified as complex sentences, the model switches to a large model for event element extraction;
[0010] Step S5: Result integration and output. After event element extraction is completed, data cleaning, deduplication, formatting, and logical verification operations are performed on the extraction results. Finally, the integrated extraction results are output in a unified standard format.
[0011] Furthermore, the traditional deep learning model selects the RoBerta+CRF model, and the large model selects the Qwen-14B model. The methods of pre-training and fine-tuning include: first pre-training on a large scale of unlabeled data, and then fine-tuning on the selected tasks.
[0012] Furthermore, the training of the event element extraction model adopts cross-validation and early stopping training strategies, and optimizes the model performance by adjusting hyperparameters. The hyperparameters include learning rate, batch size, number of iterations, optimizer type, and weight decay coefficient.
[0013] Furthermore, the sentence complexity assessment includes:
[0014] Step S31: Multi-dimensional feature extraction. Multi-dimensional features include lexical features, syntactic features, semantic depth and complexity of semantic relationships, and pragmatic features. Lexical features include word length, rarity, and word frequency. Syntactic features include sentence length, number of clauses, and diversity of part-of-speech combinations. Pragmatic features include context and communicative purpose;
[0015] Step S32: Feature fusion and quantization. Weighted average or principal component analysis is used to fuse the extracted multi-dimensional features, and the fused features are standardized by linear scaling or normalization methods;
[0016] Step S33: Build a complexity assessment model. Select a random forest or multi-layer perceptron as the machine learning model. Use the extracted and quantified multi-dimensional features as input and the complexity label of the sentence as the output of the model. Use a large-scale sentence dataset with accurately labeled complexity to deeply train the model. Finally, use the trained assessment model to evaluate the sentence complexity and obtain the evaluation results.
[0017] Compared with the prior art, the present invention has the following beneficial technical effects:
[0018] The present invention proposes an adaptive event element extraction method based on sentence complexity evaluation. By constructing a multi-domain event element extraction data set, training an event element extraction model, evaluating sentence complexity, selecting an adaptive model, integrating and outputting results, etc., efficient and accurate extraction of sentences of different complexity is achieved. By introducing adaptive model selection, the language features of the input text can be accurately analyzed, especially quantitative analysis of the key indicator of sentence complexity is carried out. Based on this, the extraction model is dynamically and intelligently selected, thereby improving the accuracy and adaptability of the extraction effect, so as to better adapt to complex and changeable text data and the needs of different fields. This method not only gives full play to the advantages of traditional deep learning models in feature extraction and pattern recognition, but also uses the ability of large models in big data processing to further improve the overall performance of event element extraction, and also enhances the adaptability and robustness of the system to different types of texts through dynamic adjustment strategies. Therefore, the present invention has significant advantages and broad application prospects in complex and changeable text data processing and cross-domain applications. This dynamic adjustment strategy not only improves the efficiency of event element extraction, but also significantly improves the extraction quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 It is a flowchart of an adaptive perception event element extraction method provided by an embodiment of the present invention;
[0021] Figure 2 It is a comparison chart of event element extraction results under different models provided by the embodiments of the present invention. DETAILED DESCRIPTION
[0022] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] The process of an adaptive perception event element extraction method provided by the present invention is as follows: Figure 1As shown below, it includes the following steps:
[0024] Step 1: Construction of the dataset
[0025] Data collection: To improve the generalization ability and adaptability of the model, text data is collected from open-source resources in multiple fields, including news reports, social media posts, academic paper abstracts, etc., to ensure the diversity and representativeness of the data, covering different fields and types of texts. At the same time, some classic event element extraction datasets are collected to expand the content.
[0026] Data cleaning: The collected data is cleaned to remove irrelevant information, noisy data, and duplicate content. At the same time, basic text processing operations such as word segmentation and part-of-speech tagging are performed.
[0027] Data annotation: The text data is annotated with event elements using manual annotation (such as doccano) or semi-automatic annotation. The annotation content includes key information such as the time, location, subject, object, trigger word, etc. of the event occurrence. The annotation process needs to follow certain norms and standards to ensure the accuracy and consistency of the annotation results.
[0028] Data integration: The manually annotated dataset is integrated with the classic dataset, and the integrated dataset is further divided according to the complexity of the event sentences in the text, so that corresponding strategies can be adopted for different types of sentences in the subsequent processing. Through the above steps, a carefully processed multi-field event element extraction dataset is obtained. This dataset covers various types of texts, including news reports, social media posts, academic paper abstracts, etc., ensuring the diversity and representativeness of the data and laying a foundation for subsequent research work.
[0029] Step 2: Training and fine-tuning of the event extraction model
[0030] Model selection: Two different types of models are selected for training and fine-tuning, namely traditional deep learning models (such as RoBerta+CRF) and large models (such as Qwen). Traditional deep learning models are suitable for processing sentences with clear structures and simple grammars, while large models have stronger context understanding capabilities and the ability to handle complex structures.
[0031] Among them, the traditional deep learning model selected the RoBerta+CRF model. The core algorithm principle of RoBERTa is the multi-head self-attention mechanism and the feed-forward neural network based on the Transformer architecture, which can capture the long-range dependencies in the input sequence. It is trained using the methods of pre-training and fine-tuning. First, it is pre-trained on a large amount of unlabeled data, and then fine-tuned on the selected tasks. By calculating the association of query, key, and value matrices, the attention weights are dynamically allocated, combined with the feed-forward neural network to transform features, strengthening semantic understanding. In training, a large-scale dataset is used, increasing the batch size and extending the number of training epochs. A large batch size allows the model to utilize more data, and a long training time helps with convergence. In task design, only the masked language model (MLM) task is retained. Tokens are randomly masked with a 15% probability, where 80% are replaced with "[MASK]", 10% remain the original words, and 10% are replaced with random words to learn the language structure. In addition, the Sentence-Order Prediction (SOP) task can also be selected to predict the sentence order and help understand the logical relationship. In terms of model scale, RoBERTa usually has more layers and hidden units than BERT, which can improve the ability to capture complex semantic syntax.
[0032] Conditional random field (CRF) is an undirected graph model that combines the characteristics of the maximum entropy model and the hidden Markov model and has achieved good results in sequence labeling tasks such as word segmentation, part-of-speech tagging, and named entity recognition. The CRF model assumes that there are only two variables, X and Y, in the Markov random field. X usually represents the observation sequence or input sequence, which is the input data of the model, that is, the information we observe and are given. For example, in the part-of-speech tagging task of natural language processing, X can be each word in the sentence to be tagged with part of speech; in the named entity recognition task, X can be a single character or can also be considered as the embedded vector after processing. Y generally represents the output sequence, tag sequence, or state sequence, which is the result that the model needs to predict or output under the condition of the given input X. Taking the part-of-speech tagging task as an example again, Y is the part of speech corresponding to these words; in named entity recognition, Y is the named entity label corresponding to each element in the input sequence, such as labels representing person names, place names, organization names, etc. X is given, and Y is the output under the condition of the given X. In the linear-chain conditional random field (linear-CRF), X and Y have the same structure, that is, X=(X1,X2,…,Xn), Y=(Y1,Y2,…,Yn). The CRF model calculates the conditional probability of the output sequence y under the given input sequence x by defining the feature function and the corresponding weights. The feature function includes the transition feature defined on the edge and the state feature defined on the node, which respectively depend on the current and previous positions and the information at the current position. By maximizing the conditional probability, the CRF model can learn the optimal sequence labeling result.
[0033] The RoBERTa+CRF model is selected as the traditional deep learning model in the event element extraction algorithm with adaptive perception. This is mainly because the RoBERTa model has powerful language representation capabilities, can deeply understand text semantics and context information, and provides rich feature representations for event element extraction. At the same time, the CRF model performs excellently in sequence labeling tasks, can consider the dependencies between adjacent labels, and effectively handle the relevance and constraints between event elements. Combining the two can give full play to their respective advantages and improve the accuracy and efficiency of event element extraction.
[0034] The large model selected is Qwen-14B, which adopts the Transformer architecture and realizes parallel computing through the self-attention mechanism, greatly improving the training efficiency. At the same time, it also introduces position encoding to retain the position information of the input sequence, ensuring that the model can understand the relative position relationship between words. These technologies enable Qwen-14B to have better sequence modeling capabilities and can show excellent performance when facing complex and changing language tasks.
[0035] Model training and fine-tuning: The selected model is trained using the constructed multi-domain event element extraction dataset. During the training process, training strategies such as cross-validation and early stopping are adopted, and the model performance is optimized by adjusting hyperparameters such as the learning rate and batch size. On the basis of the pre-trained model, fine-tuning is carried out for specific domains or tasks such as the military and political news domain. During the fine-tuning process, a small amount of labeled data is used to further train the model to make it better adapt to the needs of specific domains or tasks.
[0036] To make full use of the dataset and evaluate the generalization ability of the model, the cross-validation method is adopted in the training process. Specifically, the dataset is divided into multiple subsets. Each time, a part of the subsets is selected as the training set, and the remaining part is used as the validation set for multiple trainings and validations. This can not only evaluate the performance of the model on different data subsets but also effectively reduce the model bias caused by improper data division. At the same time, the performance of the model on the validation set is monitored in real-time during the training process. When the loss on the validation set no longer decreases or the accuracy no longer improves, it is considered that the model has started overfitting. At this time, the training should be stopped to avoid further deteriorating the model performance. This strategy is called early stopping and is an effective means to prevent overfitting.
[0037] During the training process, the model performance is optimized by continuously adjusting hyperparameters. The learning rate is one of the most important hyperparameters when training a neural network, which determines the step size for updating weights in each iteration. An overly large learning rate may cause the model to fail to converge, while an overly small learning rate will make the training process slow and prone to getting stuck in local optima. Therefore, we need to conduct experiments to find the learning rate suitable for the current task and dataset. The batch size determines the number of samples used to update weights in each iteration. A larger batch size can accelerate the training process, but may also lead to inaccurate gradient estimation, thus affecting the model's performance. On the contrary, a smaller batch size can improve the accuracy of gradient estimation, but will also increase the instability of training. Therefore, we need to choose an appropriate batch size according to the specific situation. In addition to the learning rate and batch size, there are other hyperparameters that need to be adjusted, such as the number of iterations, the type of optimizer (e.g., Adam, SGD, etc.), the weight decay coefficient, etc. The selection and optimization of these hyperparameters usually require experiments and validation.
[0038] Model performance evaluation: To comprehensively evaluate the performance of each model, we divided the integrated dataset into two subsets of simple sentences and complex sentences according to sentence complexity and conducted performance tests separately. The experimental results show that traditional models perform well in processing simple sentences and can accurately identify and extract key information; while large models demonstrate stronger capabilities in processing complex sentences, being able to capture more detailed information and better understand the context relevance.
[0039] Step 3: Sentence complexity evaluation
[0040] Multi-dimensional Feature Extraction: In text analysis, multi-dimensional feature extraction is a crucial step in understanding the complexity and meaning of sentences. Multi-dimensional features include lexical features, syntactic features, semantic depth and complexity of semantic relations, and pragmatic features. Specifically, lexical features include word length, rarity, and word frequency. Word length is obtained by calculating the character length of each word in a sentence and statistically averaging the lengths; rarity is determined using a lexical database (such as WordNet) to assess the rarity of a word, with values ranging from 0 (common) to 1 (extremely rare); word frequency is obtained from a large corpus (such as Google Ngram) to get the occurrence frequency of a word. In addition, syntactic features cover sentence length, number of clauses, and diversity of part-of-speech combinations. Sentence length is determined by counting the number of characters in a sentence; the number of clauses is identified using a syntactic analysis tool (such as the Stanford Parser); the diversity of part-of-speech combinations is evaluated by calculating the number of different part-of-speech combinations. At the semantic level, feature extraction involves semantic depth and complexity of semantic relations. Semantic depth uses a pre-trained semantic model (such as BERT) to calculate the semantic representation vector of a sentence and is evaluated by the dimension and numerical distribution of the vector; the complexity of semantic relations is revealed by analyzing the semantic relations (such as hyponymy, synonymy, etc.) between words in a sentence. Finally, pragmatic features include context and communicative purpose, where context considers the theme and background information of the text passage or discourse in which the sentence is located, and the communicative purpose determines whether the sentence is declarative, interrogative, imperative, or exclamatory and analyzes its expressed intention.
[0041] Feature Fusion and Quantification: The features extracted above are independent of each other and are difficult to directly compare and comprehensively analyze. Therefore, effective methods need to be adopted to fuse these multi-dimensional features to form a comprehensive feature representation. In the process of feature fusion, methods such as weighted average or principal component analysis (PCA) are used. The weighted average method assigns different weights to different features and performs a weighted average to obtain a comprehensive feature value. This method is simple and easy to implement, but the determination of weights requires relying on experience and professional knowledge. Principal component analysis (PCA), on the other hand, is a more objective and accurate method. It performs a linear transformation on the original features to extract the main feature components, thereby achieving the fusion of multi-dimensional features. The PCA method can effectively remove redundant information between features and retain the most representative feature components, providing a more accurate and effective feature representation for subsequent text analysis.
[0042] After feature fusion, the fused features are further processed to be converted into computable values within the range of 0 to 1, which is achieved through linear scaling or normalization methods. The linear scaling method linearly maps the feature values to the range of 0 to 1 based on the minimum and maximum values of the features. The normalization method, on the other hand, standardizes the feature values by calculating the mean and standard deviation of the features to make them conform to the standard normal distribution. Both methods can effectively convert the fused features into computable numerical forms, facilitating subsequent analysis and comparison.
[0043] Construct a complexity evaluation model: To construct an effective sentence complexity evaluation model, first select appropriate machine learning models, such as Random Forest or Multi-Layer Perceptron (MLP), which have demonstrated excellent capabilities in dealing with complex data relationships. Subsequently, use the multi-dimensional features extracted and quantified previously as inputs, and use the complexity labels of sentences (clearly classified into categories such as simple, medium, and complex) as the output targets of the model. On this basis, use a large-scale sentence dataset with accurately labeled complexities to deeply train the model to ensure that it can accurately capture the subtle differences in sentence complexity.
[0044] For the Random Forest model, key parameters such as n_estimators (i.e., the number of decision trees, set to 100 for example) and max_depth (i.e., the maximum depth of the decision tree, limited to 10 layers for example) are adjusted to find the best balance between the generalization ability of the model and the risk of overfitting. For the Multi-Layer Perceptron model, the configuration of the hidden layer is optimized, for example, setting two hidden layers with 64 and 32 neurons respectively, and finely adjusting the learning rate to 0.01 to improve the training efficiency and final performance of the model. Through the tuning of these parameters, an accurate and efficient sentence complexity evaluation model is constructed.
[0045] Optimization and Validation of Sentence Complexity Evaluation Model: First, a cross-validation strategy is adopted. Specifically, K-fold cross-validation (e.g., K = 5) is implemented. By dividing and validating the training dataset multiple times, the performance of the model under different parameter configurations is evaluated in an unbiased manner, so as to select the optimal model parameters. In addition, to effectively prevent overfitting during model training, regularization techniques are introduced, including L1 and L2 regularization methods. These techniques impose constraints on the model parameters, which helps to reduce the complexity of the model and improve its generalization ability. Subsequently, it enters the model validation stage. A brand-new, unseen test dataset is used to strictly validate the trained evaluation model. During the validation process, a series of evaluation metrics are calculated, including accuracy, recall, and F1 value. These metrics comprehensively measure the performance of the model in identifying sentences of different complexities, providing us with an objective and quantitative evaluation basis, and further guiding the further optimization of the model structure or parameters in order to achieve higher evaluation accuracy and stronger generalization ability.
[0046] Real-time Update and Adaptive Learning: Implement a regular data collection strategy, continuously incorporate new sentence samples, and carefully annotate their complexities. Subsequently, these new data are integrated into the original training dataset to retrain the model, thereby enhancing the model's ability to capture the latest language phenomena. In addition, to address possible performance fluctuations in practical applications, an adaptive learning mechanism is constructed. This mechanism continuously monitors the accuracy metric of the model in the actual scenario. Once it drops below the preset threshold, it automatically triggers the retraining process to promptly respond to performance degradation. Furthermore, an online learning algorithm is introduced, endowing the model with the ability to learn new data features in real time, enabling it to gradually absorb new knowledge and continuously optimize its performance without interrupting the service, thus demonstrating stronger adaptability and robustness when facing the ever-changing text complexity evaluation requirements.
[0047] Sentence Complexity Evaluation: Finally, the trained sentence complexity evaluation model is used to evaluate the sentence complexity to obtain the evaluation result.
[0048] Step 4: Adaptive Model Selection
[0049] Model Selection Strategy Design: Based on the results of sentence complexity evaluation, design the following adaptive model selection strategy: For sentences with lower complexity (such as simple sentences with a single subject and object), select traditional deep learning models for event element extraction, which usually have faster speed and better performance when dealing with simple sentences. For sentences with higher complexity (such as compound sentences and nested sentences), select large models for event element extraction, which can capture richer context information when dealing with complex sentences, thus obtaining better performance.
[0050] Implementation of the model switching mechanism: To implement the model switching mechanism, configuration files or environment variables are used to specify the paths and parameters of different models. During program execution, based on the result of sentence complexity evaluation, the corresponding model paths and parameters are read from the configuration file, and then the corresponding models are loaded and run. Another way to implement model switching is to use dynamic loading technology. During program execution, based on the result of sentence complexity evaluation, the corresponding model is dynamically loaded and run, which can be achieved through Python's importlib module or other dynamic loading libraries.
[0051] Caching mechanism: To improve efficiency, a caching mechanism can also be introduced. For sentences whose complexity has been evaluated, the mapping relationship between their complexity scores and the selected extraction models can be cached for direct use in subsequent processing without having to re-evaluate the complexity. This can be achieved by using Python's functools.lru_cache decorator or other caching libraries.
[0052] Step 5: Result integration and output
[0053] Result integration: Integrate the extraction results output by different models, including operations such as data cleaning, duplicate removal, and formatting, to further improve the accuracy and consistency of the extraction results.
[0054] Result output: Format and output the extraction results according to actual needs. For example, the output results should include key information such as the time, location, and participants of the event, and store and display the extracted event elements in structured text, JSON format, or database form; at the same time, API interfaces or web services can also be provided for other systems or applications to call and integrate, providing high-quality data support for downstream tasks.
[0055] Through the above implementation methods, the present invention realizes the adaptive perception of sentence complexity. The comparison of the extraction results of event elements by different models is as Figure 2 shown. By comparing the extraction results, it can be obtained that the F1 values of the extraction results of traditional models, large models, and the adaptive perception model proposed by the present invention for event sentences with different complexities are 0.83, 0.89, and 0.93 respectively. From this, it can be concluded that the adaptive perception event element extraction method proposed by the present invention can perform high-quality event element extraction on texts with large differences in sentence complexity, effectively improving the accuracy and efficiency of event element extraction, providing a solution for the event element extraction task in the field of natural language processing, not only improving the accuracy and consistency of the extraction results, but also providing high-quality data support for downstream tasks such as event relationship recognition, text summary generation, and event analysis and answering.
[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An event element extraction method with adaptive perception, characterized in that The method includes: Step S1: Data collection and processing. Collect data from open-source resources in multiple fields, and construct a high-quality multi-field event element extraction dataset after cleaning and annotation. Step S2: Event element extraction model training. Use the constructed multi-field event element extraction dataset to pre-train and fine-tune two different types of models, namely traditional deep learning models and large models. Step S3: Sentence complexity assessment. Analyze the length, syntactic structure, dependency relationship, and lexical richness features of sentences to construct complexity assessment metrics. Based on these metrics, the model automatically determines the complexity level of each sentence and classifies it. Step S4: Adaptive model perception. According to the results of sentence complexity assessment, perform adaptive model selection. For instances classified as simple sentences, call the traditional deep learning model for event element extraction. For instances classified as complex sentences, switch to the large model for event element extraction. Step S5: Result integration and output. After completing event element extraction, perform data cleaning, deduplication, formatting, and logical verification operations on the extraction results, and finally output the integrated extraction results in a unified standard format.
2. The method according to claim 1, wherein The traditional deep learning model selects the RoBerta+CRF model, and the large model selects the Qwen-14B model. Pre-training and fine-tuning include: first pre-train on a large scale of unlabeled data, and then fine-tune on the selected tasks.
3. The method according to claim 1, characterized in that, Event element extraction model training adopts cross-validation and early stopping training strategies, and optimizes the model performance by adjusting hyperparameters. The hyperparameters include learning rate, batch size, number of iterations, optimizer type, and weight decay coefficient.
4. The method according to claim 1, wherein Sentence complexity assessment includes: Step S31: Multi-dimensional feature extraction. The multi-dimensional features include lexical features, syntactic features, semantic depth and complexity of semantic relationships, and pragmatic features. Lexical features include word length, rarity, and word frequency. Syntactic features include sentence length, number of clauses, and diversity of part-of-speech combinations. Pragmatic features include context and communicative purpose. Step S32: Feature fusion and quantization. Use weighted average or principal component analysis to perform feature fusion on the extracted multi-dimensional features, and standardize the fused features by linear scaling or normalization methods. Step S33: Construct a complexity assessment model. Select a random forest or multi-layer perceptron as the machine learning model, use the extracted and quantified multi-dimensional features as input, and use the complexity label of the sentence as the output of the model. Use a large-scale sentence dataset with accurately labeled complexity to deeply train the model, and finally use the trained assessment model to evaluate the sentence complexity to obtain the evaluation results.
Citation Information
Cited By
Large model training-oriented multi-dimensional generalizable data generation method and device and software system
CN121189320A
Method and system for automatically and accurately tracking academic activities of scholars based on large language model
CN121616285A
Method and system for automatically and accurately tracking academic activities of scholars based on large language model
CN121616285B