Academic paper intelligent evaluation method and device based on multi-dimensional vector representation

Through the combination of multi-dimensional vector representation and multiple machine learning algorithms, the shortcomings of traditional paper evaluation methods are solved, and in-depth, accurate and efficient evaluation of academic papers is achieved.

CN120337914AActive Publication Date: 2025-07-18BEIJING SANSAN SMART EDUCATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510445332.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Traditional paper evaluation methods cannot deeply understand the content quality and logical structure of the paper, resulting in inaccurate and comprehensive enough, time-consuming and subjective, making it difficult to meet the needs of large-scale academic evaluations.

Method used

A multi-dimensional vector representation method is used to extract vocabulary vectors, part-of-speech vectors and paragraph vectors, combined with deep neural network models and comprehensive scoring models, and integrate multiple machine learning algorithms for academic paper evaluation.

Benefits of technology

It realizes multi-dimensional semantic analysis of academic papers, improves the accuracy and efficiency of evaluation, and can objectively and comprehensively evaluate the content quality and logical structure of the paper.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337914A_ABST
    Figure CN120337914A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, in particular to an academic paper intelligent assessment method and device based on multi-dimensional vector representation, and the method comprises the steps: carrying out the multi-dimensional vector extraction of a to-be-assessed academic paper; wherein the extracted multi-dimensional vector comprises a vocabulary vector, a part-of-speech vector and a paragraph vector; inputting the extracted vocabulary vector, the extracted part-of-speech vector and the extracted paragraph vector into a trained deep neural network model for processing to obtain a preliminary assessment result of the academic paper to be assessed; inputting the to-be-assessed academic paper and the preliminary assessment result into a comprehensive scoring model for processing to obtain a final assessment result of the to-be-assessed academic paper; wherein the comprehensive scoring model integrates a plurality of machine learning algorithms. Therefore, multi-dimensional feature analysis is carried out on the to-be-evaluated academic paper, evaluation is deeper, comprehensive evaluation is carried out on the paper quality in combination with multiple machine learning algorithms, and the accuracy and efficiency of evaluation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of machine learning, and particularly to an intelligent evaluation method and device for academic papers based on multi-dimensional vector representation. Background Art

[0002] Traditional paper evaluation methods usually rely on keyword matching and grammar checking. These methods cannot deeply understand the content quality and logical structure of papers. They often ignore the specific meaning and usage of words in context, resulting in potentially inaccurate and incomplete evaluation results. With the explosive growth of information, the academic field has an increasingly urgent need for efficient and accurate paper evaluation methods. Traditional manual evaluation methods are time-consuming and subjective, and it is difficult to meet the needs of large-scale academic evaluations. Summary of the Invention

[0003] To overcome the deficiencies in the prior art, this application provides an intelligent evaluation method and device for academic papers based on multi-dimensional vector representation, which can achieve multi-dimensional semantic analysis and improve the evaluation efficiency and accuracy.

[0004] In a first aspect, this application provides an intelligent evaluation method for academic papers based on multi-dimensional vector representation. The method includes the following steps:

[0005] Extract multi-dimensional vectors from the academic paper to be evaluated; among them, the extracted multi-dimensional vectors include word vectors, part-of-speech vectors, and paragraph vectors;

[0006] Input the extracted word vectors, part-of-speech vectors, and paragraph vectors into a trained deep neural network model for processing to obtain a preliminary evaluation result of the academic paper to be evaluated;

[0007] Input the academic paper to be evaluated and the preliminary evaluation result into a comprehensive scoring model for processing to obtain a final evaluation result of the academic paper to be evaluated; among them, the comprehensive scoring model integrates multiple machine learning algorithms.

[0008] In a possible implementation manner, the extraction of multi-dimensional vectors from the academic paper to be evaluated includes the following steps:

[0009] Preprocess the academic paper to be evaluated; the preprocessing includes denoising, word segmentation, part-of-speech tagging, and paragraph segmentation;

[0010] Extract word vectors from the preprocessed academic paper to be evaluated based on a word vector model;

[0011] Process the word vectors and part-of-speech tagging information based on a part-of-speech vector model to obtain part-of-speech vectors;

[0012] Process the word vector, part-of-speech vector, and paragraph segmentation information based on the paragraph vector model to obtain the paragraph vector.

[0013] In a possible implementation, the extracted multi-dimensional vector further includes the document topic feature, where the document topic feature is obtained by processing the to-be-evaluated academic paper based on the latent Dirichlet allocation model.

[0014] In a possible implementation, the deep neural network model is a three-layer neural network model composed of an input layer, a hidden layer, and an output layer, and is trained as follows, including the following steps:

[0015] Initialize the weights and biases of the deep neural network;

[0016] Perform forward propagation on the training data to obtain the prediction result; where the training data is a set of word vectors, part-of-speech vectors, and paragraph vectors obtained by extracting multi-dimensional vectors from a large number of collected academic paper samples.

[0017] Calculate the loss value based on the prediction result and the corresponding true label;

[0018] Use the backpropagation algorithm to calculate the gradient values of the loss value with respect to the weights and biases, and adjust the weights and biases based on the gradient values until the deep neural network converges or reaches the preset number of training rounds.

[0019] In a possible implementation, the multiple machine learning algorithms integrated by the comprehensive scoring model include the extreme gradient boosting algorithm, the term frequency-inverse document frequency algorithm, the text ranking algorithm, and the k-nearest neighbor algorithm.

[0020] In a possible implementation, the step of inputting the to-be-evaluated academic paper and the preliminary evaluation result into the comprehensive scoring model for processing to obtain the final evaluation result of the to-be-evaluated academic paper includes the following steps:

[0021] Optimize the preliminary evaluation result based on the extreme gradient boosting algorithm to obtain the optimized preliminary evaluation result;

[0022] Input the to-be-evaluated academic paper into the term frequency-inverse document frequency algorithm and the text ranking algorithm for processing to obtain supplementary information; where the word importance is calculated based on the term frequency-inverse document frequency algorithm and the text structure is analyzed based on the text ranking algorithm, and the obtained word importance and text structure information are used as the supplementary information;

[0023] Calibrate the preliminary evaluation result based on the k-nearest neighbor algorithm to obtain the calibrated preliminary evaluation result;

[0024] Comprehensively evaluate the optimized preliminary evaluation result, supplementary information, and calibrated preliminary evaluation result according to preset rules to obtain the final evaluation result of the academic paper to be evaluated.

[0025] In a possible implementation manner, calibrating the preliminary evaluation result based on the k-nearest neighbor algorithm to obtain a calibrated preliminary evaluation result includes the following steps:

[0026] Calculate the Euclidean distance between the academic paper to be evaluated and each academic paper in the dataset based on the k-nearest neighbor algorithm;

[0027] Select at least one academic paper with the closest Euclidean distance as the nearest neighbor academic paper;

[0028] Calibrate the preliminary evaluation result according to the score of the nearest neighbor academic paper.

[0029] In a second aspect, the present application provides an intelligent evaluation device for academic papers based on multi-dimensional vector representation. The device includes:

[0030] An extraction module for extracting multi-dimensional vectors from the academic paper to be evaluated; among them, the extracted multi-dimensional vectors include word vectors, part-of-speech vectors, and paragraph vectors;

[0031] A preliminary evaluation module for inputting the extracted word vectors, part-of-speech vectors, and paragraph vectors into a trained deep neural network model for processing to obtain a preliminary evaluation result of the academic paper to be evaluated;

[0032] A comprehensive evaluation module for inputting the academic paper to be evaluated and the preliminary evaluation result into a comprehensive scoring model for processing to obtain the final evaluation result of the academic paper to be evaluated; among them, the comprehensive scoring model integrates multiple machine learning algorithms.

[0033] In a third aspect, the present application provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the intelligent evaluation method for academic papers based on multi-dimensional vector representation according to any one of the first aspects are executed.

[0034] In a fourth aspect, the present application provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the intelligent evaluation method for academic papers based on multi-dimensional vector representation according to any one of the first aspects are executed.

[0035] An intelligent evaluation method and device for academic papers based on multi-dimensional vector representation provided by this embodiment extract multi-dimensional vectors for the academic papers to be evaluated; among them, the extracted multi-dimensional vectors include lexical vectors, part-of-speech vectors, and paragraph vectors; input the extracted lexical vectors, part-of-speech vectors, and paragraph vectors into a trained deep neural network model for processing to obtain a preliminary evaluation result of the academic papers to be evaluated; input the academic papers to be evaluated and the preliminary evaluation result into a comprehensive scoring model for processing to obtain a final evaluation result of the academic papers to be evaluated; among them, the comprehensive scoring model integrates a variety of machine learning algorithms. Thus, multi-dimensional feature analysis is performed on the academic papers to be evaluated, making the evaluation more in-depth, and comprehensively evaluating the paper quality by combining a variety of machine learning algorithms, improving the accuracy and efficiency of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 Shows a flowchart of the intelligent evaluation method for academic papers based on multi-dimensional vector representation according to an embodiment of the present application;

[0038] Figure 2 Shows a flowchart of extracting multi-dimensional vectors for the academic papers to be evaluated according to an embodiment of the present application;

[0039] Figure 3 Shows a flowchart of inputting the academic papers to be evaluated and the preliminary evaluation result into a comprehensive scoring model for processing to obtain a final evaluation result of the academic papers to be evaluated according to an embodiment of the present application;

[0040] Figure 4 Shows a schematic structural diagram of the intelligent evaluation device for academic papers based on multi-dimensional vector representation according to an embodiment of the present application;

[0041] Figure 5 Shows a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be understood that the accompanying drawings in this application only serve the purpose of illustration and description, and are not used to limit the protection scope of this application. Additionally, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of this application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. Furthermore, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application.

[0043] In addition, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of this application claimed, but merely represents the selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the protection scope of this application.

[0044] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the subsequently stated features, but does not exclude the addition of other features.

[0045] In view of the technical problems proposed in the background art, this application provides an intelligent evaluation method and device for academic papers based on multi-dimensional vector representation, which can achieve multi-dimensional semantic analysis and improve the evaluation efficiency and accuracy.

[0046] In one embodiment, referring to the accompanying specification Figure 1 , an intelligent evaluation method for academic papers based on multi-dimensional vector representation provided by this application includes the following steps:

[0047] S1. Extract multi-dimensional vectors from the academic paper to be evaluated; among them, the extracted multi-dimensional vectors include word vectors, part-of-speech vectors, and paragraph vectors;

[0048] S2. Input the extracted word vectors, part-of-speech vectors, and paragraph vectors into a trained deep neural network model for processing to obtain a preliminary evaluation result of the academic paper to be evaluated;

[0049] S3. Input the academic paper to be evaluated and the preliminary evaluation result into a comprehensive scoring model for processing to obtain the final evaluation result of the academic paper to be evaluated; wherein, the comprehensive scoring model integrates multiple machine learning algorithms.

[0050] Specifically, in step S1, the present application uses word2vec, pos2ve, and paragraph2vec technologies to capture the complex semantic features of words, sentences, and paragraphs in the academic paper to be evaluated, rather than based on a single keyword or syntactic analysis, making the subsequent evaluation process more in-depth and comprehensive, and capable of accurately evaluating the content quality and logical structure of the academic paper to be evaluated.

[0051] See the attached Figure 2 specification. The multi-dimensional vector extraction for the academic paper to be evaluated includes the following steps:

[0052] S101. Preprocess the academic paper to be evaluated; the preprocessing includes denoising, word segmentation, part-of-speech tagging, and paragraph segmentation.

[0053] S102. Extract word vectors for the preprocessed academic paper to be evaluated based on a word vector model.

[0054] S103. Process the word vectors and part-of-speech tagging information based on a part-of-speech vector model to obtain part-of-speech vectors.

[0055] S104. Process the word vectors, part-of-speech vectors, and paragraph segmentation information based on a paragraph vector model to obtain paragraph vectors.

[0056] In step S101, the preprocessing of the academic paper to be evaluated is mainly to improve the data quality and not affect the accuracy of subsequent multi-dimensional vector extraction. The preprocessing methods include denoising, word segmentation, part-of-speech tagging, and paragraph segmentation. Among them, denoising is mainly used for data cleaning, such as removing obvious formatting errors, garbled characters, special symbols, etc. in the academic paper to be evaluated. Word segmentation can use professional word segmentation tools, such as NLTK (Natural Language Toolkit), to split continuous text into individual words. Part-of-speech tagging can use part-of-speech tagging tools, such as Stanford CoreNLP, spaCy, etc., to label each word with the corresponding part-of-speech tag, such as noun, verb, adjective, adverb, etc., according to the grammatical function and semantic features of the word in the sentence. At the same time, paragraph segmentation is also performed based on punctuation marks or text structure features, and paragraph IDs are marked to provide a good basis for subsequent extraction of paragraph vectors.

[0057] In steps S102 - S104, the lexical vector is the basic structure of the text. A good lexical vector can group semantically similar words together, which helps with subsequent text classification, text clustering, and other operations. Therefore, in step S102, a word vector model (word2vec) is used to extract lexical vectors from the pre - processed academic paper to be evaluated. In the evaluation of the paper, the part - of - speech vector can help judge the rationality of lexical collocations and the correctness of sentence structures. Therefore, in step S103, a part - of - speech vector model (pos2vec) is used to map each part of speech to a low - dimensional vector. In this way, the word sequence in the sentence is converted into a part - of - speech vector sequence, and the words in the sentence are vector - represented from the perspective of parts of speech. In the evaluation of the paper, the paragraph vector can be used to analyze the paragraph coherence, theme consistency, and logical relationships between paragraphs of the paper. Therefore, in step S104, a paragraph vector model (paragraph2vec) is used to process the lexical vectors, part - of - speech vectors, and paragraph segmentation information to obtain paragraph vectors. Among them, the working principles of the word2vec model, pos2vec model, and paragraph2vec model are technical means well - known to those skilled in the art and will not be elaborated here.

[0058] It should be noted that before using the word2vec model, pos2vec model, and paragraph2vec model to extract multi - dimensional vectors from the pre - processed academic paper to be evaluated, the word2vec model, pos2vec model, and paragraph2vec model need to be trained. For example, the paper dataset of students' mock exams is used as training data. After pre - processing, it is input into the word2vec model for training. During the training process, the parameters of the lexical vectors are continuously adjusted to make words with similar semantics closer in the vector space. After training, each word can be mapped to a vector of a fixed length, and these vectors contain the semantic information of the words, providing a basis for subsequent text analysis. Similar to the word2vec model, by training the pos2vec model, the vector representation of each part of speech is obtained. By training the paragraph2vec model, the vector representation of each paragraph is obtained.

[0059] Furthermore, the Latent Dirichlet Allocation (LDA) model can also be used to extract the document topic features of the academic paper to be evaluated. This is because the above-mentioned word2vec model, pos2vec model, and paragraph2vec model mainly focus on the local and middle-level semantic features of the text, while the LDA model focuses on revealing the topic structure of the text from a global level. It can discover the latent topics in the document collection and the distribution of each document on different topics. Combining the lexical vectors, part-of-speech vectors, paragraph vectors with the document topic features can provide a more rich and comprehensive basis for subsequent evaluation.

[0060] In step S2, the extracted multi-dimensional vectors are used as the input of the deep neural network model. Through the weight connections between neurons in each layer and the action of the activation function, these complex features are deeply analyzed and processed to capture the deep patterns and relationships in the academic paper to be evaluated, and then the preliminary evaluation result of the quality of the academic paper to be evaluated is output.

[0061] In one embodiment, the deep neural network model selects a three-layer neural network model (NNM), which consists of an input layer, a hidden layer, and an output layer. Each layer is interconnected through modifiable weights, and the hidden layer units form a net activation through the weighted sum of their inputs. This architecture can effectively capture the complex features and patterns in the paper and understand the content quality and logical structure of the paper. Among them, the deep neural network model is trained in the following way: First, initialize the weights and biases of the deep neural network, for example, initialize them with small random numbers; then perform forward propagation on the training data. The data is weighted and summed in the hidden layer through the weight connections to obtain the net activation value, and then processed through the activation function to introduce non-linearity into the model and enhance the model's learning ability for complex relationships. The processed result continues to be passed to the next layer until the output layer obtains the prediction result. Then, compare the prediction result of the deep neural network with the true labels of the paper samples in the training set (such as the actual score, quality level, etc. of the paper), select an appropriate loss function to calculate the difference between the two, and obtain the loss value; then, based on the calculated loss value, use the backpropagation algorithm to calculate the gradient values of the loss value with respect to the weights and biases. According to the calculated gradient values, use an optimization algorithm to update the weights and biases. In this way, continuously repeat the processes of forward propagation, loss calculation, backpropagation, and parameter update until the deep neural network converges or reaches the preset number of training rounds.

[0062] In step S3, mainly use the preliminary evaluation result output by the deep neural network model as one of the inputs of the comprehensive scoring model integrating multiple machine learning algorithms to evaluate the academic paper to be evaluated from different perspectives and generate the final evaluation result.

[0063] In this application, the multiple machine learning algorithms integrated in the comprehensive scoring model include the Extreme Gradient Boosting (XGBoost) algorithm, the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm, the TextRank algorithm, and the k-Nearest Neighbors (kNN) algorithm. Refer to the attached Figure 3 description. The step of inputting the academic paper to be evaluated and the preliminary evaluation result into the comprehensive scoring model for processing to obtain the final evaluation result of the academic paper to be evaluated includes the following steps:

[0064] S301. Optimize the preliminary evaluation result based on XGBoost to obtain an optimized preliminary evaluation result;

[0065] S302. Input the academic paper to be evaluated into TF-IDF and TextRank for processing to obtain supplementary information; specifically, calculate the lexical importance based on TF-IDF and analyze the text structure based on TextRank, and use the obtained lexical importance and text structure information as the supplementary information;

[0066] S303. Calibrate the preliminary evaluation result based on kNN to obtain a calibrated preliminary evaluation result;

[0067] S304. Comprehensively evaluate the optimized preliminary evaluation result, the supplementary information, and the calibrated preliminary evaluation result according to preset rules to obtain the final evaluation result of the academic paper to be evaluated.

[0068] In step S301, XGBoost is a powerful gradient boosting framework. In this application, it optimizes the preliminary evaluation result output by the deep neural network model. Specifically, define an objective function that includes a loss function and a regularization term. The loss function can choose the mean squared error (MSE) to measure the difference between the preliminary evaluation result and the true score, and the regularization term prevents the model from overfitting. XGBoost constructs decision trees in an iterative manner. In each round of iteration, a new tree is fitted to predict the residuals of the previous round, continuously reducing the value of the loss function. When constructing a decision tree, use a greedy algorithm to traverse the features and split values, select the split point with the largest gain for splitting, and perform a pruning operation after the tree construction is completed to remove the leaf nodes that do not contribute much to the improvement of the model performance, thereby optimizing the preliminary evaluation result.

[0069] In step S302, supplementary information is mainly provided by TF-IDF and TextRank. Among them, TF-IDF calculates the importance of vocabulary, and TextRank analyzes the text structure. Specifically, TF-IDF first calculates the term frequency (TF) to represent the frequency of a word appearing in a paper, then calculates the inverse document frequency (IDF) to measure the general importance of a word, and then multiplies the term frequency (TF) and the inverse document frequency (IDF) to obtain the TF-IDF value of each word in the academic paper to be evaluated, and the importance of the vocabulary is reflected by this TF-IDF value. TextRank constructs a graph model, takes the sentences or words in the academic paper to be evaluated as nodes, and uses methods such as cosine similarity to calculate the similarity between nodes to determine the edges. An iterative algorithm (such as PageRank) is used to calculate the node weights. The sentences with high weights are key sentences, reflecting the logical structure and core ideas of the paper, and the words with high weights are used as keywords, which complement TF-IDF and provide information at the text structure level for evaluation.

[0070] In step S303, first, multi-dimensional vectors, TF-IDF values, etc. are extracted from the academic paper to be evaluated and each academic paper in the dataset according to the above method, and each academic paper is represented as a feature vector. Then, based on kNN, metric methods such as Euclidean distance and cosine distance are used to calculate the distance between the academic paper to be evaluated and each academic paper in the dataset. According to the distance, the nearest k academic papers are selected as the nearest neighbors, and the preliminary evaluation result is calibrated by voting or weighted average according to the scores of the nearest neighbor papers. Among them, the scores of the nearest neighbor papers can be the existing data marked when collecting the dataset. The calibrated preliminary evaluation result is obtained.

[0071] In step S304, the preliminary evaluation result optimized by XGBoost, the supplementary information provided by TF-IDF and TextRank, and the preliminary evaluation result calibrated by kNN are comprehensively considered, and a comprehensive consideration is carried out from multiple dimensions such as the content quality, logical structure, and vocabulary application of the paper. Through the preset scoring rules and evaluation templates, detailed evaluation content of the paper is generated, and an objective and comprehensive final score is given to complete the evaluation of the academic paper to be evaluated. In one embodiment, the preset scoring rules can set different weights for the evaluation indicators in different dimensions and perform quantitative scoring on the performance in each dimension. Furthermore, the comprehensive evaluation of the academic paper to be evaluated is completed, providing detailed evaluation content and an objective final score for the user, helping the user understand the advantages and disadvantages of the paper, and providing a reference for the modification and improvement of the paper.

[0072] It can be seen that an intelligent evaluation method for academic papers based on multi-dimensional vector representation provided by the present application uses a variety of text vector representation technologies for multi-dimensional feature analysis, and combines a variety of machine learning algorithms to comprehensively evaluate the quality of papers, improving the accuracy and efficiency of evaluation.

[0073] Based on the same inventive concept, an intelligent evaluation device for academic papers based on multi-dimensional vector representation is also provided in an embodiment of the present application. Since the principle of solving problems by the device in the embodiment of the present application is similar to that of the above-mentioned intelligent evaluation method for academic papers based on multi-dimensional vector representation in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0074] As shown in the attached Figure 4 description, an intelligent evaluation device for academic papers based on multi-dimensional vector representation provided in an embodiment of the present application includes:

[0075] An extraction module 401, configured to perform multi-dimensional vector extraction on the academic paper to be evaluated; wherein, the extracted multi-dimensional vectors include word vectors, part-of-speech vectors, and paragraph vectors;

[0076] A preliminary evaluation module 402, configured to input the extracted word vectors, part-of-speech vectors, and paragraph vectors into a trained deep neural network model for processing to obtain a preliminary evaluation result of the academic paper to be evaluated;

[0077] A comprehensive evaluation module 403, configured to input the academic paper to be evaluated and the preliminary evaluation result into a comprehensive scoring model for processing to obtain a final evaluation result of the academic paper to be evaluated; wherein, the comprehensive scoring model integrates a variety of machine learning algorithms.

[0078] In one embodiment, the extraction module 401 performs multi-dimensional vector extraction on the academic paper to be evaluated, including: preprocessing the academic paper to be evaluated; the preprocessing includes denoising, word segmentation, part-of-speech tagging, and paragraph segmentation; performing word vector extraction on the preprocessed academic paper to be evaluated based on a word vector model; processing the word vectors and part-of-speech tagging information based on a part-of-speech vector model to obtain part-of-speech vectors; and processing the word vectors, part-of-speech vectors, and paragraph segmentation information based on a paragraph vector model to obtain paragraph vectors.

[0079] In one embodiment, the extracted multi-dimensional vectors further include document theme features, and the extraction module 401 processes the academic paper to be evaluated based on an LDA model to obtain the document theme features.

[0080] In one embodiment, the deep neural network model is a three-layer neural network model composed of an input layer, a hidden layer, and an output layer, and the device further includes:

[0081] A training module, configured to initialize the weights and biases of a deep neural network; perform forward propagation on training data to obtain prediction results; wherein the training data is a set of vocabulary vectors, part-of-speech vectors, and paragraph vectors obtained by performing multi-dimensional vector extraction on a large number of collected academic paper samples; calculate a loss value based on the prediction results and corresponding true labels; use the backpropagation algorithm to calculate the gradient values of the loss value with respect to the weights and biases, and adjust the weights and biases based on the gradient values until the deep neural network converges or reaches a preset number of training epochs.

[0082] In one embodiment, the multiple machine learning algorithms integrated in the comprehensive scoring model include XGBoost, TF-IDF, TextRank, and kNN. The comprehensive evaluation module 403 inputs the academic paper to be evaluated and the preliminary evaluation results into the comprehensive scoring model for processing to obtain the final evaluation results of the academic paper to be evaluated, including: optimizing the preliminary evaluation results based on XGBoost to obtain optimized preliminary evaluation results; inputting the academic paper to be evaluated into TF-IDF and TextRank for processing to obtain supplementary information; wherein calculating the lexical importance based on TF-IDF and analyzing the text structure based on TextRank, and using the obtained lexical importance and text structure information as the supplementary information; calibrating the preliminary evaluation results based on kNN to obtain calibrated preliminary evaluation results; comprehensively evaluating the optimized preliminary evaluation results, supplementary information, and calibrated preliminary evaluation results according to preset rules to obtain the final evaluation results of the academic paper to be evaluated.

[0083] In one embodiment, the comprehensive evaluation module 403 calibrates the preliminary evaluation results based on kNN to obtain calibrated preliminary evaluation results, including: calculating the Euclidean distance between the academic paper to be evaluated and each academic paper in the dataset based on kNN; selecting at least one academic paper with the closest Euclidean distance as the nearest neighbor academic paper; calibrating the preliminary evaluation results according to the scores of the nearest neighbor academic papers.

[0084] An intelligent evaluation device for academic papers based on multi-dimensional vector representation provided by the present application extracts multi-dimensional vectors from the academic papers to be evaluated through an extraction module; among them, the extracted multi-dimensional vectors include lexical vectors, part-of-speech vectors, and paragraph vectors; the preliminary evaluation module inputs the extracted lexical vectors, part-of-speech vectors, and paragraph vectors into a trained deep neural network model for processing to obtain a preliminary evaluation result of the academic papers to be evaluated; the comprehensive evaluation module inputs the academic papers to be evaluated and the preliminary evaluation result into a comprehensive scoring model for processing to obtain a final evaluation result of the academic papers to be evaluated; among them, the comprehensive scoring model integrates multiple machine learning algorithms. Thus, multi-dimensional feature analysis is performed on the academic papers to be evaluated, making the evaluation more in-depth, and comprehensively evaluating the paper quality by combining multiple machine learning algorithms, improving the accuracy and efficiency of the evaluation.

[0085] Based on the same concept of the present invention, the specification appendix Figure 5 As shown, the structure of an electronic device 500 provided by an embodiment of the present application, the electronic device 500 includes: at least one processor 501, at least one network interface 504 or other user interfaces 503, a memory 505, and at least one communication bus 502. The communication bus 502 is used to realize the connection and communication between these components. The electronic device 500 optionally includes a user interface 503, including a display (for example, a touch screen, LCD, CRT, holographic imaging, or projector, etc.), a keyboard or a pointing device (for example, a mouse, a trackball, a touchpad, or a touch screen, etc.).

[0086] The memory 505 may include a read-only memory and a random access memory, and provide instructions and data to the processor 501. A part of the memory 505 may also include a non-volatile random access memory (NVRAM).

[0087] In some embodiments, the memory 505 stores the following elements, executable modules or data structures, or subsets thereof, or extended sets thereof:

[0088] An operating system 5051, including various system programs, used to implement various basic services and process hardware-based tasks;

[0089] An application program module 5052, including various application programs, such as a launcher, a media player, a browser, etc., used to implement various application services.

[0090] In an embodiment of the present application, by invoking the program or instructions stored in the memory 505, the processor 501 is configured to execute the steps in an intelligent evaluation method for academic papers based on multi-dimensional vector representation, and can implement multi-dimensional semantic analysis to improve the evaluation efficiency and accuracy.

[0091] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps in the intelligent evaluation method for academic papers based on multi-dimensional vector representation.

[0092] Specifically, the storage medium can be a general storage medium, such as a removable disk, a hard disk, etc. When the computer program on the storage medium is run, it can execute the above-mentioned intelligent evaluation method for academic papers based on multi-dimensional vector representation.

[0093] In the embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of devices or units can be in an electrical, mechanical or other form.

[0094] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0095] In addition, each functional unit in the embodiments provided by the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0096] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0097] Finally, it should be noted that: The above embodiments are only specific implementation manners of this application, used to illustrate the technical solutions of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: Any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An intelligent evaluation method for academic papers based on multi-dimensional vector representation, characterized in that, The method includes the following steps: Perform multi-dimensional vector extraction on the academic paper to be evaluated; among them, the extracted multi-dimensional vectors include word vectors, part-of-speech vectors, and paragraph vectors; Input the extracted word vectors, part-of-speech vectors, and paragraph vectors into the trained deep neural network model for processing to obtain the preliminary evaluation result of the academic paper to be evaluated; Input the academic paper to be evaluated and the preliminary evaluation result into the comprehensive scoring model for processing to obtain the final evaluation result of the academic paper to be evaluated; among them, the comprehensive scoring model integrates multiple machine learning algorithms.

2. The intelligent evaluation method for academic papers based on multi-dimensional vector representation according to claim 1, wherein The multi-dimensional vector extraction of the academic paper to be evaluated includes the following steps: Preprocess the academic paper to be evaluated; the preprocessing includes denoising, word segmentation, part-of-speech tagging, and paragraph segmentation; Extract word vectors from the preprocessed academic paper to be evaluated based on the word vector model; Process the word vectors and part-of-speech tagging information based on the part-of-speech vector model to obtain part-of-speech vectors; Process the word vectors, part-of-speech vectors, and paragraph segmentation information based on the paragraph vector model to obtain paragraph vectors.

3. The intelligent evaluation method for academic papers based on multi-dimensional vector representation according to claim 2, characterized in that, The extracted multi-dimensional vectors also include document topic features, among which the document topic features are obtained by processing the academic paper to be evaluated based on the Latent Dirichlet Allocation model.

4. The intelligent evaluation method for academic papers based on multi-dimensional vector representation according to claim 1, characterized in that, The deep neural network model is a three-layer neural network model composed of an input layer, a hidden layer, and an output layer, and is trained as follows, including the following steps: Initialize the weights and biases of the deep neural network; Perform forward propagation on the training data to obtain a prediction result; among them, the training data is a set of word vectors, part-of-speech vectors, and paragraph vectors obtained by performing multi-dimensional vector extraction on a large number of collected academic paper samples; Calculate the loss value based on the prediction result and the corresponding true label; Use the backpropagation algorithm to calculate the gradient values of the loss value with respect to the weights and biases, and adjust the weights and biases based on the gradient values until the deep neural network converges or reaches the preset number of training rounds.

5. The intelligent evaluation method for academic papers based on multi-dimensional vector representation according to claim 1, wherein The multiple machine learning algorithms integrated in the comprehensive scoring model include the Extreme Gradient Boosting algorithm, the Term Frequency-Inverse Document Frequency algorithm, the TextRank algorithm, and the k-Nearest Neighbor algorithm.

6. The intelligent evaluation method for academic papers based on multi-dimensional vector representation according to claim 5, wherein The step of inputting the academic paper to be evaluated and the preliminary evaluation result into the comprehensive scoring model for processing to obtain the final evaluation result of the academic paper to be evaluated includes the following steps: Optimize the preliminary evaluation result based on the Extreme Gradient Boosting algorithm to obtain an optimized preliminary evaluation result; Input the academic paper to be evaluated into the Term Frequency-Inverse Document Frequency algorithm and the TextRank algorithm for processing to obtain supplementary information; among them, calculate the lexical importance based on the Term Frequency-Inverse Document Frequency algorithm and analyze the text structure based on the TextRank algorithm, and use the obtained lexical importance and text structure information as the supplementary information; Calibrate the preliminary evaluation result based on the k-Nearest Neighbor algorithm to obtain a calibrated preliminary evaluation result; Comprehensively evaluate the optimized preliminary evaluation result, the supplementary information, and the calibrated preliminary evaluation result according to the preset rules to obtain the final evaluation result of the academic paper to be evaluated.

7. The intelligent evaluation method for academic papers based on multi-dimensional vector representation according to claim 6, characterized in that, Calibrating the preliminary evaluation result based on the k-nearest neighbor algorithm to obtain the calibrated preliminary evaluation result includes the following steps: Calculating the Euclidean distance between the academic paper to be evaluated and each academic paper in the dataset based on the k-nearest neighbor algorithm; Selecting at least one academic paper with the closest Euclidean distance as the nearest neighbor academic paper; Calibrating the preliminary evaluation result according to the score of the nearest neighbor academic paper.

8. An intelligent evaluation device for academic papers based on multi-dimensional vector representation, characterized in that, The device includes: An extraction module for extracting multi-dimensional vectors from the academic paper to be evaluated; among them, the extracted multi-dimensional vectors include lexical vectors, part-of-speech vectors, and paragraph vectors; A preliminary evaluation module for inputting the extracted lexical vectors, part-of-speech vectors, and paragraph vectors into a trained deep neural network model for processing to obtain the preliminary evaluation result of the academic paper to be evaluated; A comprehensive evaluation module for inputting the academic paper to be evaluated and the preliminary evaluation result into a comprehensive scoring model for processing to obtain the final evaluation result of the academic paper to be evaluated; among them, the comprehensive scoring model integrates multiple machine learning algorithms.

9. An electronic device, characterized in that, Including: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the academic paper intelligent evaluation method based on multi-dimensional vector representation according to any one of claims 1 to 7 are executed.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of the academic paper intelligent evaluation method based on multi-dimensional vector representation according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Commodity performance evaluation method through network evaluation test guiding

    CN106296288A

  • Foreign language text evaluation method and apparatus

    CN108280065A

  • Method and apparatus for extracting evaluation relationship of evaluation text information

    CN109117470A

  • Text quality evaluation method and device, equipment and medium

    CN117787284A

  • Textual content evaluation using machine learned models

    EP4163815A1

Cited By

  • Personal information protection evaluation method based on FaaS and deep learning

    CN121659354A