An education evaluation target data processing method and system based on artificial intelligence

Through the evaluation of the density sequence of interest distribution and information binning technology, the education evaluation text is processed, the education evaluation subtext is generated, and the information of interest is identified, which solves the problem of difficulty in accurately capturing detailed information in the existing technology, and achieves more efficient information extraction.

CN119250069BActive Publication Date: 2025-05-06SHENZHEN YUNJUAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411038017.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2025-05-06
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

The existing educational evaluation data processing methods based on artificial intelligence are difficult to accurately capture the detailed information in the educational evaluation text, and ignore the differences and importance of information between different paragraphs, resulting in low accuracy and efficiency of information extraction.

Method used

The educational evaluation text is processed by the method of estimating the distribution density sequence of interest. The sequence of distributed density of interest is generated by filtering samples of different degrees of different degrees. Then it is binned to obtain multiple information paragraphs of binning, and the text is intercepted based on these paragraphs to generate the educational evaluation subtext. Then, each subtext and the original text are identified with information of interest, and the recognition results are combined to improve the accuracy of information extraction.

Benefits of technology

This method can more accurately identify the detailed information in the educational evaluation text, overcome the problem of inaccurate information extraction, and improve the accuracy and efficiency of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119250069B_ABST
    Figure CN119250069B_ABST
Patent Text Reader

Abstract

The present application provides an education evaluation target data processing method and system based on artificial intelligence, which adopts the binned information paragraphs obtained by binning the distribution density sequence of interest, and intercepts the education evaluation sub-text in the education evaluation text for identification. In this way, compared with identifying different paragraphs of the unprocessed education evaluation text separately, the detailed information in the education evaluation text can be identified more accurately, and the first information of interest identification result is used as a supplement to the second information of interest identification result of the education evaluation text, thereby overcoming the problem of inaccurate information extraction caused by the random distribution of the information of interest in the education evaluation text. In other words, in the scheme of the present application, the extraction of information of interest is more accurate to prevent information omission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to an education evaluation target data processing method and system based on artificial intelligence. Background Art

[0002] In the field of education, with the continuous deepening of educational informatization, the processing and analysis of educational evaluation data has become increasingly important. Traditional educational evaluation methods often rely on manual reading and summarizing a large number of evaluation texts. This process is not only time-consuming and laborious, but also easily affected by subjective factors, making it difficult to ensure the accuracy and consistency of the evaluation results. In order to improve the efficiency and quality of evaluation, in recent years, more and more educational institutions and research institutions have begun to explore the use of information technology and artificial intelligence to carry out automated educational evaluation. Existing artificial intelligence-based educational evaluation data processing methods still have some limitations. On the one hand, since educational evaluation texts usually contain a large amount of information, and the distribution of this information in the text is random, it is often difficult to accurately capture all key information points by directly identifying the entire text, especially those detailed information scattered in different paragraphs. On the other hand, existing methods often use a unified recognition strategy to process the entire text, ignoring the differences and importance of information between different paragraphs, resulting in low accuracy and efficiency of information extraction. Summary of the invention

[0003] In view of this, the present application provides an education evaluation target data processing method and system based on artificial intelligence. The technical solution of the present application is implemented as follows:

[0004] On the one hand, the present application provides an education evaluation target data processing method based on artificial intelligence, the method comprising: estimating an interest distribution density sequence of an education evaluation text to obtain an interest distribution density sequence of the education evaluation text, wherein the interest distribution density sequence characterizes the dispersion of the interest information in the education evaluation text, and the interest distribution density sequence estimation includes multiple filtering samplings of different degrees; binning the interest distribution density sequence to obtain a plurality of binned information paragraphs in the interest distribution density sequence; intercepting the education evaluation text according to the plurality of binned information paragraphs to obtain a plurality of education evaluation sub-texts corresponding to each other, wherein each of the education evaluation sub-texts includes a plurality of the interest information;

[0005] Identify the information of interest for each of the educational evaluation sub-texts and the educational evaluation text one by one to obtain a first information of interest identification result corresponding to each of the educational evaluation sub-texts and the educational evaluation text; merge each of the first information of interest identification results to obtain a second information of interest identification result corresponding to the educational evaluation text.

[0006] On the other hand, the present application provides a computer system, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the above method when executing the program.

[0007] The beneficial effects of the present application include at least: the artificial intelligence-based education evaluation target data processing method and system provided in the present application adopts the binned information paragraphs obtained by binning the distribution density sequence of interest, and intercepts the obtained education evaluation sub-text in the education evaluation text for identification. In this way, compared with identifying different paragraphs of the unprocessed education evaluation text separately, the detailed information in the education evaluation text can be more accurately identified, and the first information of interest identification result is used as a supplement to the second information of interest identification result of the education evaluation text, thereby overcoming the problem of inaccurate information extraction due to the random distribution of the information of interest in the education evaluation text. In other words, in the scheme of the present application, the extraction of information of interest is more accurate to prevent information omission. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.

[0009] Figure 1 A schematic diagram of the implementation flow of an artificial intelligence-based education evaluation target data processing method provided in an embodiment of the present application.

[0010] Figure 2 A hardware entity diagram of a computer system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0011] The embodiment of the present application provides an artificial intelligence-based education evaluation target data processing method, which can be executed by a processor of a computer system. The computer system can refer to a server, a laptop, a tablet computer, a desktop computer, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable gaming device), and other devices with data processing capabilities.

[0012] Figure 1 A schematic diagram of the implementation flow of an artificial intelligence-based education evaluation target data processing method provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:

[0013] Step S10: estimate the interest distribution density sequence of the education evaluation text to obtain the interest distribution density sequence of the education evaluation text, wherein the interest distribution density sequence represents the dispersion of the interest information in the education evaluation text, and the interest distribution density sequence estimation includes multiple filtering samplings of different degrees.

[0014] In step S10, the computer system estimates the distribution density sequence of interest for a long text of educational evaluation corresponding to multiple evaluation time points or stages (these texts may cover a comprehensive evaluation of the performance of students or teachers throughout the semester). The purpose of this step is to identify which parts of the text contain the most valuable or frequently mentioned evaluation information, that is, the information of interest, through intelligent analysis technology. Specifically, the computer system preprocesses the long text using natural language processing technology, including word segmentation, stop word removal, stem extraction and other steps for subsequent processing. Subsequently, the system uses multiple filtering sampling methods of varying degrees to construct the distribution density sequence of interest. Here, filtering sampling usually involves convolution operations and pooling operations in machine learning. Taking convolutional neural networks (CNNs) as an example, the system can design a series of convolution layers, each layer using convolution kernels of different sizes to "scan" text data (in this scenario, text data can be regarded as a one-dimensional sequence, and each word or word vector represents an element in the sequence). Each convolution kernel is responsible for capturing phrases or patterns of a specific length in the text, which may be highly correlated with the information of interest. Through convolution operations, the system is able to generate a series of feature maps, each of which represents an abstract representation of the original text at different scales, that is, text features at different levels. Next, these feature maps are refined based on the pooling layer to reduce the spatial size of the data (for text data, it is to reduce the length of the sequence) while retaining the most important information. Pooling operations (such as maximum pooling or average pooling) help reduce the spatial resolution of features, so that the model pays more attention to whether certain features exist rather than their exact location. Through these multiple and gradually deepening convolutions and pooling, one or more interest distribution density sequences are generated. Each sequence represents the distribution of interesting information at different levels of abstraction in the text. The values ​​in these sequences can represent the density of interesting information in a specific text fragment (such as a sentence or paragraph). The higher the value, the more key evaluation information the fragment contains.

[0015] Step S20: binning the distribution density sequence of interest to obtain a plurality of binned information segments in the distribution density sequence of interest.

[0016] In step S20, the computer system further processes the interest distribution density sequence generated in step S10, and divides the sequence into multiple binned information paragraphs through information binning (clustering) technology. These paragraphs represent the areas where the interest information is more concentrated in the text, which are of great value for subsequent educational evaluation analysis.

[0017] Specifically, the computer system analyzes the numerical distribution characteristics of the distribution density sequence of interest and identifies the high-density and low-density areas in the sequence. The purpose of information binning is to divide the density value of the continuous distribution into several discrete intervals, and the density values ​​in each interval are highly similar, while the differences between different intervals are relatively obvious. In the embodiment of the present application, this information binning operation can be implemented by a variety of clustering algorithms, such as K-means clustering, hierarchical clustering or DBSCAN. Here, the K-means clustering algorithm is used as an example for illustration.

[0018] For a distribution density sequence of interest, the computer system can treat each point in the sequence (or each continuous high-density area) as a data sample, and then apply the K-means algorithm for clustering. When the algorithm is executed, the computer system randomly selects K initial cluster centers. Then, for each point in the sequence, the distance to each cluster center is calculated based on its density value, and it is assigned to the cluster to which the nearest cluster center belongs. After all points are assigned, the computer system recalculates the cluster center of each cluster (usually the mean of the density values ​​of all points in the cluster). This process is repeated until the cluster center no longer changes significantly or the preset number of iterations is reached.

[0019] After the information is binned, the computer system will obtain K binned information paragraphs, each of which contains a continuous area with similar density values ​​in the original text. The text in these paragraphs often revolves around one or more specific evaluation topics, so it is of great significance for analyzing the specific performance of students or teachers and identifying key issues in the evaluation.

[0020] Step S30: intercepting the education evaluation text according to the multiple boxed information paragraphs to obtain multiple corresponding education evaluation sub-texts, wherein each education evaluation sub-text includes multiple pieces of information of interest.

[0021] In step S30, the computer system accurately intercepts the original education evaluation text according to the multiple boxed information paragraphs obtained in step S20 to generate a series of corresponding education evaluation sub-texts. This process is automatic and efficient, and is intended to separate independent fragments containing key information from a long evaluation text to facilitate subsequent in-depth analysis. Specifically, the computer system traverses each boxed information paragraph, which has been identified as an area containing high-density information of interest through information boxing technology (such as K-means clustering). For each boxed information paragraph, the system records its starting position and ending position in the original text. Then, using this information, the system intercepts the text of the corresponding area from the original text to generate an independent education evaluation sub-text. Each education evaluation sub-text directly corresponds to a boxed information paragraph, so there is a one-to-one correspondence between them. These sub-texts not only retain the structure and context of the original text, but also focus more on specific evaluation topics or aspects. Since the selection of boxed information paragraphs is based on the density distribution of information of interest, each sub-text naturally contains multiple information points of interest, which are crucial for understanding the evaluation content and identifying key issues.

[0022] In the embodiment of the present application, such an interception operation has significant practical significance. For example, suppose that a comprehensive evaluation report at the end of a semester contains evaluations of multiple aspects of a student, such as learning attitude, class participation, homework completion, etc. Through information binning, the computer system may have divided these evaluation contents into different paragraphs. Subsequently, in step S30, the system will intercept independent sub-texts about learning attitude, class participation, etc. from the comprehensive evaluation report according to the boundaries of each binned information paragraph. These sub-texts can be directly used for further analysis, such as sentiment analysis, keyword extraction, or comparison with other students' evaluations.

[0023] Since the interception operation is based on the precise control of the text position, no additional machine learning models or complex algorithms are required. However, in practical applications, in order to ensure the accuracy of the interception, the computer system may need to combine natural language processing technology (such as word segmentation, part-of-speech tagging, etc.) to assist in determining the text boundaries and the specific location of the information of interest. In addition, in order to verify the validity of the interception results, the system can also perform a quality assessment on the intercepted sub-text, such as checking whether it contains the expected keywords, whether the context of the original text is maintained, etc. Although these steps do not directly belong to the content of step S30, they are of great significance to the optimization and improvement of the entire education evaluation data processing process.

[0024] Step S40: Identify the information of interest for each education evaluation sub-text and education evaluation text one by one, and obtain the first information of interest identification result corresponding to each education evaluation sub-text and education evaluation text.

[0025] In step S40, the computer system performs comprehensive and in-depth identification of information of interest on each intercepted education evaluation subtext and the original education evaluation text. This process is intended to accurately capture the key information points in the text and provide a reliable basis for subsequent data analysis and decision support. Based on this, the computer system can integrate advanced natural language processing (NLP) technology and machine learning models. In the embodiment of the present application, the information of interest includes, for example, but is not limited to positive or negative evaluation vocabulary, key performance indicators (KPIs), specific behavior descriptions, etc. In order to effectively identify this information, the computer system can adopt a variety of methods, such as rule-based matching, statistical models or deep learning models. Taking the deep learning model as an example, the system can use convolutional neural networks (CNN) or recurrent neural networks (RNN, especially its variants LSTM or GRU) to capture semantic features and contextual information in the text. These networks can learn complex representations of text data and learn recognition patterns of information of interest from labeled data through supervised learning. During the recognition process, the system encodes each word or phrase (such as using word embedding technology) and then passes it to the model as input. The model outputs the probability that each word or phrase belongs to the information of interest by analyzing the sequence characteristics of the text. By setting a threshold, the system can filter out high-probability information of interest as the first information of interest recognition result.

[0026] In addition, in order to improve the accuracy and robustness of recognition, the system may also use ensemble learning methods to combine the prediction results of multiple models for comprehensive judgment. For example, a voting mechanism or weighted average method can be used to integrate the outputs of different models to obtain more stable and reliable recognition results.

[0027] In the embodiment of the present application, it is assumed that there is an evaluation text about a student's learning attitude, which contains multiple positive and negative evaluation words. Through the identification of the information of interest in step S40, the computer system can automatically capture these key words and distinguish which words represent positive learning attitudes (such as "diligent" and "serious") and which words reflect negative learning attitudes (such as "lazy" and "inattentive"). This information is of great significance for comprehensively evaluating the student's learning status and formulating personalized education programs.

[0028] Finally, after identifying the information of interest for each educational evaluation subtext and educational evaluation text one by one, the computer system will output the first information of interest identification results corresponding to each of them. These results not only help educational managers quickly understand the core content of the evaluation text, but also provide valuable data resources for subsequent data mining and decision analysis.

[0029] Step S50: Merge each first information of interest recognition result to obtain a second information of interest recognition result corresponding to the education evaluation text.

[0030] In step S50, the computer system integrates the first information of interest recognition results of each educational evaluation sub-text and the original educational evaluation text obtained in step S40 to generate a comprehensive second information of interest recognition result. Specifically, the computer system collects all first information of interest recognition results, which may be presented in different forms, such as keyword lists, sentiment tendency scores, or specific information encodings. In order to merge these results, the system needs a unified way to quantify or compare them. In most cases, if the recognition results are expressed in numerical form (such as sentiment tendency scores), the system can directly use mathematical operations (such as addition, average, etc.) to merge these results. In practical applications, due to the differences in the length, content complexity, and types and quantities of information of interest of different sub-texts and original texts, simple mathematical operations may not be sufficient to accurately reflect the information distribution of the overall evaluation text. Therefore, the computer system may need to adopt a more complex merging strategy.

[0031] A possible merging strategy is the weighted average method. In this method, the system assigns different weights to each sub-text or original text according to its importance in the overall evaluation. The distribution of weights can be based on a variety of factors, such as text length, information density, the authority of the evaluator, etc. The system then applies these weights to the respective first information of interest identification results and calculates the weighted average to obtain the second information of interest identification result. For example, in an embodiment of the present application, it is assumed that there are two sub-texts that evaluate the student's learning attitude and class participation respectively, and the original text is a comprehensive evaluation of the student. The system may believe that learning attitude is more important for the overall evaluation, and therefore assigns a higher weight to the learning attitude evaluation. After obtaining the first information of interest identification results of learning attitude evaluation, class participation evaluation and comprehensive evaluation, the system will perform weighted average on these results according to preset weights, and finally obtain a comprehensive second information of interest identification result.

[0032] Through the merging process of step S50, the computer system can generate a second information of interest identification result that fully reflects the key information of the education evaluation text. This result not only helps education managers quickly understand the overall evaluation situation, but also provides a strong data basis for subsequent data analysis and decision support.

[0033] In one embodiment, step S10, estimating the interest distribution density sequence of the education evaluation text to obtain the interest distribution density sequence of the education evaluation text, includes:

[0034] Step S11: according to the educational evaluation text, the optimized interest distribution density sequence estimation method is used to estimate the interest distribution density sequence, and the interest distribution density sequence of the educational evaluation text is obtained; wherein the optimized interest distribution density sequence estimation method is obtained by optimizing the following process:

[0035] Step S1: using the interest distribution density sequence estimation method to be optimized to estimate the interest distribution density sequence according to the education evaluation sample text, and obtaining the estimated interest distribution density sequence of the education evaluation sample text;

[0036] Step S2: using a cost calculation function according to the estimated interest distribution density sequence of the education evaluation sample text and the prior interest distribution density sequence of the education evaluation sample text, to obtain the training cost of the education evaluation sample text;

[0037] Step S3: Optimize the algorithm parameters of the distribution density sequence estimation method of interest to be optimized according to the training cost of the education evaluation sample text.

[0038] In step S11, the computer system uses a deeply trained and optimized interest distribution density sequence estimation method (specifically a deep neural network algorithm) to process new education evaluation texts. Deep neural networks are known for their powerful nonlinear modeling capabilities and automatic feature learning capabilities, and are very suitable for processing complex, high-dimensional data, such as natural language texts.

[0039] For example, suppose there is a comprehensive education evaluation text set containing evaluations of students by multiple teachers in different subjects over multiple semesters. This text set is not only huge in quantity, but also rich in content, covering multiple dimensions such as students' learning attitude, classroom performance, homework completion, teamwork ability, etc. The goal of the computer system is to automatically extract the distribution density sequence of interesting information in each evaluation text from this text set, so as to conduct more in-depth data analysis and mining later.

[0040] Based on this, the computer system loads a pre-trained deep neural network model. This model is optimized through supervised learning on a large number of annotated educational evaluation sample texts. The structure of the model may include multiple convolutional layers, pooling layers, fully connected layers, etc., which are used to extract high-level features in the text layer by layer, and finally output a distribution density sequence of interest. When a new educational evaluation text is input into the model, the model will automatically process it. For example, the model can first convert the text into a sequence of word vectors, then capture local features (such as phrase or sentence level features) through a convolutional layer, and then downsample through a pooling layer to reduce the amount of calculation and extract the most important features. Subsequently, these features will be passed to the fully connected layer for nonlinear combination and decision-making, and finally generate a distribution density sequence of interest corresponding to the length of the text.

[0041] In this sequence, the value of each element represents the density of interesting information at the corresponding position in the text (which may be at the word, phrase or sentence level). The higher the density value, the more key evaluation information the position contains, and the more important it is for subsequent data analysis and mining.

[0042] During the training phase, in order to obtain a deep neural network model that can accurately estimate the distribution density sequence of interest, the computer system collects a set of annotated educational evaluation sample texts. These sample texts cover a wide range of evaluation content and scenarios to ensure that the trained model has good generalization capabilities. For example, the construction of a training data set is a complex and meticulous process. First, experts in the field of education need to manually annotate a large number of educational evaluation texts, assign labels to key evaluation information in the text (such as positive evaluations, negative evaluations, specific behavior descriptions, etc.), and generate a priori distribution density sequences of interest. This process is not only time-consuming and labor-intensive, but also requires the annotator to have deep professional knowledge and rich experience.

[0043] Once the annotation is completed, the computer system can use these sample texts to train the deep neural network model. In the early stage of training, the model is usually a deep neural network framework that has not been fully adjusted and has a relatively simple structure. This framework contains multiple learnable parameters (such as weights and biases), which will be iteratively updated through the back-propagation algorithm during the training process.

[0044] At the beginning of training, the computer system inputs sample texts into the model one by one. For each sample text, the model first performs forward propagation calculations: converting the text into a sequence of word vectors, extracting features layer by layer through a multi-layer neural network, and finally generating an estimated distribution density sequence of interest. However, in the early stages of training, since the model parameters are randomly initialized, the generated estimated sequence often differs greatly from the prior sequence.

[0045] In order to quantify the difference between the model's estimated results and the actual labeled results, the computer system uses a cost calculation function (also called a loss function) to calculate the training cost. The training cost is an important indicator for measuring model performance, which reflects the size of the model's prediction error on the training data set.

[0046] For example, in the embodiment of the present application, the cost calculation function can be mean square error (MSE), cross entropy loss, etc. Mean square error is applicable to regression problems, which calculates the average of the squares of the differences between the model's estimated value and the actual value. For the task of estimating the distribution density sequence of interest, the mean square error can be a good measure of the difference between the estimated sequence and the prior sequence.

[0047] Assuming that the mean square error is used as the cost function, its calculation formula is:

[0048]

[0049] Where N is the number of sample texts, M is the length of the distribution density sequence of interest in each sample text (i.e., the number of words, phrases, or sentences in the text), and y ij is the prior density of interest at the jth position in the i-th sample text, is the density of interest at the corresponding location estimated by the model.

[0050] During the training process, the computer system calculates the MSE value of each sample text and adds them up to get the total training cost. Then, the system evaluates the performance of the current model based on this total cost and guides the subsequent parameter optimization process.

[0051] According to the training cost calculated in step S2, the computer system will adjust the parameters (such as weights and biases) of the deep neural network model through optimization techniques such as the back propagation algorithm. The back propagation algorithm is an efficient parameter optimization method that guides the update direction of the parameters by calculating the gradient of the cost function with respect to the model parameters. For example, during the training process, each forward propagation calculation is followed by a back propagation calculation. During the back propagation process, the computer will first calculate the gradient of the cost function with respect to the output layer of the model (that is, the partial derivative of the cost function with respect to the estimated distribution density sequence of interest). Then, these gradients will be propagated backward to each layer of the model layer by layer, and the gradient of the parameters of each layer will be calculated. After obtaining the gradient, the computer system can use the gradient descent method or its variants (such as stochastic gradient descent, Adam optimizer, etc.) to update the parameters of the model. The rule for parameter update is usually to subtract a small amount proportional to the gradient (that is, the product of the learning rate and the gradient) from the current parameter value. Through multiple iterations of this process, the model parameters will gradually converge to an optimal solution, minimizing the prediction error of the model on the training data set. It is worth noting that some details need to be paid attention to during the training process, such as overfitting, underfitting, and learning rate adjustment. Overfitting refers to the phenomenon that the model performs well on the training data set but poorly on the test data set; underfitting refers to the phenomenon that the model performs poorly on both the training data set and the test data set. In order to avoid these problems, the computer system can use regularization, early stopping, Dropout and other techniques to limit the complexity of the model and improve its generalization ability; at the same time, the training process of the model can be monitored through the validation set and hyperparameters such as the learning rate can be adjusted in time. After sufficient training, the computer system will obtain an optimized deep neural network model. This model can accurately estimate the distribution density sequence of interest of the newly input educational evaluation text, providing strong support for subsequent data analysis and mining. In practical applications, this model can be integrated into various educational evaluation systems to help educational administrators, teachers and students better understand the evaluation results and make scientific decisions.

[0052] In one embodiment, step S11, using an optimized interest distribution density sequence estimation method to estimate the interest distribution density sequence according to the education evaluation text, to obtain the interest distribution density sequence of the education evaluation text, includes:

[0053] According to the educational evaluation text, the optimized interest distribution density sequence estimation method is used to perform the following steps:

[0054] Step S111: performing filtering sampling of the education evaluation texts for multiple times with different degrees of accuracy to obtain multiple implicit representation sets with different granularities;

[0055] Step S112: performing a fusion operation on multiple implicit representation sets of different degrees to obtain a fused implicit representation set;

[0056] Step S113: extract features from the fused implicit representation set to obtain a distribution density sequence of interest for the education evaluation text.

[0057] In step S111, the computer system performs multiple filtering sampling operations of different degrees on the educational evaluation text, with the purpose of capturing the information features in the text from multiple scales. Filter sampling is a commonly used signal processing technology, which reduces noise or highlights specific frequency components by smoothing or sharpening the signal. In the context of educational evaluation text processing, filtering sampling can be understood as abstracting and simplifying the text to different degrees, thereby obtaining text representations of different granularities (or scales).

[0058] For example, suppose there is a comprehensive evaluation text about a student's learning performance in a semester, which contains evaluations of the student's class participation, homework completion, learning attitude and other aspects. The computer system preprocesses the text, such as word segmentation, stop word removal, etc., and then uses the processed text as input for filtering and sampling. The filtering and sampling operation can be implemented in many ways. One common method is to use the convolution layer in a convolutional neural network (CNN). The convolution layer slides on the text through a sliding window (convolution kernel) and calculates the dot product (or weighted sum) of the convolution kernel and the text in the window at each position to obtain a series of feature values. These feature values ​​can be regarded as implicit representations (or feature maps) of the text at different granularities.

[0059] In order to obtain a multi-granular implicit representation set, the computer system can use multiple convolution kernels of different sizes and steps for filtering sampling. For example, using convolution kernels of sizes 3, 5, and 7, and setting different step sizes (such as 1 or 2), implicit representation sets that capture local, mid-range, and global text features can be obtained respectively. Each element in these implicit representation sets (i.e., feature maps) represents an abstract representation of text at different granularities.

[0060] After obtaining implicit representation sets of multiple granularities, the computer system performs a fusion operation on these representations to generate a comprehensive fused implicit representation set containing more useful information. The fusion operation aims to integrate feature information of different granularities, reduce information redundancy and enhance the generalization ability of the model.

[0061] For example, when fusing multi-granularity implicit representations, the computer system can adopt a variety of strategies. A simple and effective method is feature concatenation. Specifically, the feature maps of different granularities are concatenated in the channel dimension to form a new, higher-dimensional feature map as the fused implicit representation. This method can retain the feature information of all granularities, but it also increases the computational complexity and number of parameters of the model. In order to balance the relationship between model performance and computational efficiency, the computer system can also use the attention mechanism to perform feature fusion. The attention mechanism allows the model to dynamically adjust the weight distribution of features of different granularities during the fusion process, so that the model can pay more attention to the feature information that has a greater impact on the prediction results. For example, when calculating the fused implicit representation, the computer system can assign an attention weight to each granularity feature map, which can be calculated by an additional neural network (such as a self-attention network or a fully connected network). Then, the feature map of each granularity is multiplied by its corresponding attention weight and summed (weighted sum) to obtain the fused implicit representation set. After obtaining the set of fused implicit representations, the computer system further performs feature extraction operations on these representations to generate the final distribution density sequence of interest in the educational evaluation text. Feature extraction is a key technology in machine learning, which aims to extract the most valuable feature information for the prediction task from the original data or preliminary representation. For example, in the scenario of educational evaluation text processing, feature extraction usually involves further convolution, pooling and other operations on the fused implicit representation. These operations can help the model capture more abstract and advanced feature information in the text, such as semantic roles, emotional tendencies, etc.

[0062] Specifically, the computer system can use one or more convolutional layers to perform convolution operations on the fused implicit representation to extract richer local feature information. Then, the convolution result is downsampled through the pooling layer to reduce the size of the feature map and retain the most important feature information. After multiple convolution and pooling operations, the model will eventually output a one-dimensional vector as a distribution density sequence of interest for the educational evaluation text.

[0063] Each element in this one-dimensional vector represents the density of interesting information at the corresponding position in the text (such as word, phrase or sentence level). The higher the density value, the more key evaluation information is contained in the position, and the more important it is for subsequent data analysis and mining. It is worth noting that in the process of feature extraction, the selection and optimization of model parameters have an important influence on the accuracy of the final result. These parameters include the size and number of convolution kernels, the type and step size of the pooling layer, the selection of activation function, etc. In order to obtain the optimal combination of model parameters, the computer system can use cross-validation, grid search and other methods to perform hyperparameter tuning.

[0064] In summary, step S11 achieves effective processing and analysis of educational evaluation texts through a series of operations such as multiple filtering sampling to obtain multi-granularity implicit representations, fusion of multi-granularity implicit representations, and feature extraction to generate distribution density sequences of interest. This process not only makes full use of the rich information in the text, but also improves the accuracy and efficiency of information extraction through the introduction of machine learning technology, thus laying a solid foundation for subsequent data analysis and mining.

[0065] In one embodiment, step S20, binning the distribution density sequence of interest to obtain a plurality of binned information segments in the distribution density sequence of interest, includes:

[0066] Step S21: identifying the sequence unit in the distribution density sequence of interest, and obtaining the probability density function of the sequence unit;

[0067] Step S22: binning the sequence units corresponding to the same distribution law according to the probability density function of the sequence units to obtain a plurality of binning information segments in the distribution density sequence of interest.

[0068] In step S21, the primary task of the computer system is to analyze each sequence unit in the distribution density sequence of interest and estimate its probability density function. The sequence unit here can be a single numerical point in the sequence, or it can be a continuous area (such as a word, a sentence or a text). The probability density function describes the possibility distribution of the values ​​of the sequence unit and is an important basis for information binning. For example, suppose there is an evaluation text about a student's classroom performance in a semester. After processing in step S10, an interesting distribution density sequence is obtained. This sequence is a one-dimensional array, and each element represents the density of the interesting information at the corresponding position in the text (such as each word or sentence). To simplify the explanation, it is assumed that the elements in the sequence are continuous and non-negative.

[0069] In order to estimate the probability density of each sequence unit, the computer system can use the kernel density estimation (KDE) method. KDE is a non-parametric probability density estimation method that smoothes the sample data through a kernel function (such as a Gaussian kernel) to estimate the overall probability density function. In specific implementation, the computer system will perform a weighted summation of the sample points around each sequence unit, and the weight is given by the kernel function. The choice of kernel function and the setting of bandwidth will affect the smoothness and accuracy of the estimation result. For example, selecting a Gaussian kernel and appropriately adjusting the bandwidth parameter can make the estimated probability density function neither too smooth to lose details nor too rough to introduce noise.

[0070] Through the KDE method, the computer system can estimate a probability density value for each unit in the sequence, and then obtain the probability density function of the entire sequence. This function describes the possible distribution of the values ​​of the sequence units and provides a basis for subsequent information binning.

[0071] In step S22, the computer system divides the sequence units with similar distribution laws into the same information box (or paragraph) according to the probability density function obtained in step S21. The purpose of information binning is to divide the continuous density sequence into several discrete paragraphs, and the sequence units in each paragraph have similar information density characteristics. For example, let's continue with the previous classroom performance evaluation text as an example. After obtaining the probability density function of the sequence, the computer system divides the sequence into multiple binned information paragraphs based on this function. The division process can be regarded as a clustering problem: similar sequence units are clustered into groups to form several paragraphs.

[0072] In order to achieve information binning, computer systems can use a variety of clustering algorithms, such as K-means clustering, hierarchical clustering or DBSCAN. However, since the probability density function is continuous and may have a complex shape (such as multimodal distribution), directly applying these traditional clustering algorithms may not work well. Therefore, it is necessary to design a suitable binning strategy based on the characteristics of the probability density function.

[0073] One feasible method is to combine probability density threshold and local maximum value for binning. First, the computer system can set a global density threshold as the basic basis for binning. Then, the boundaries of binning are determined by finding the local maximum points in the probability density function. Local maximum points usually represent significant changes in the information density of the sequence and are natural boundaries for dividing different paragraphs.

[0074] In specific implementation, the computer system can first traverse the entire probability density function and find all local maximum points. Then, the sequence is divided into several initial paragraphs based on these maximum points. Next, for each initial paragraph, the computer system will check the density distribution inside it. If there is an area below the global density threshold in the paragraph, the paragraph may need to be further subdivided; if the density distribution within the paragraph is relatively uniform and above the threshold, the paragraph can be retained as an independent binning information paragraph.

[0075] The binning process is adaptive, that is, the number and boundaries of bins should be determined according to the specific density distribution. Too many bins may lead to excessive fragmentation of information, while too few bins may not accurately reflect the changes in key information in the sequence. Therefore, it is necessary to balance the integrity of information and the independence of paragraphs in the binning process. In addition, in order to avoid overfitting and improve the generalization ability of the model, the computer system can also use methods such as cross-validation to evaluate the effects of different binning strategies and select the optimal binning scheme.

[0076] In summary, step S20 realizes effective division of the distribution density sequence of interest by identifying the probability density function of the sequence unit and performing information binning according to the probability density function. This process not only extracts the concentrated paragraphs of key information in the text, but also provides valuable data structure support for subsequent data analysis and mining.

[0077] In one embodiment, step S40, identifying the information of interest for each education evaluation sub-text and education evaluation text one by one, and obtaining the first information of interest identification result corresponding to each education evaluation sub-text and education evaluation text respectively, includes:

[0078] Complete the following operations for any educational evaluation subtext:

[0079] Step S41: performing implicit representations of different granularities on the education evaluation subtext, and obtaining multiple implicit representations of different granularities corresponding to the education evaluation subtext;

[0080] Step S42: performing preliminary information of interest recognition on multiple implicit representations of different granularities to obtain preliminary information of interest recognition results corresponding to the implicit representation of each granularity;

[0081] Step S43: performing second-order information of interest recognition on each primary information of interest recognition result to obtain a second-order information of interest recognition result corresponding to each primary information of interest recognition result, wherein the recognition accuracy of the second-order information of interest recognition is greater than the recognition accuracy of the primary information of interest recognition;

[0082] Step S44: determining the second-order information of interest recognition result as the first information of interest recognition result.

[0083] In step S41, the primary task of the computer system is to perform multi-granular implicit representations on each educational evaluation sub-text. Implicit representation here refers to converting text data into numerical or vector forms that can be understood and processed by computers for subsequent analysis and recognition. Multi-granularity means that these representations will cover different text units from word level, phrase level to sentence level and even paragraph level, thereby providing more comprehensive text feature information. For example, suppose there is an educational evaluation sub-text about student classroom performance: "Xiao Ming actively participated in discussions in class and put forward several creative ideas." The computer system performs word segmentation on this text and obtains words such as "Xiao Ming", "in class", "actively participated", "discussed", "put forward", "several", "creative", and "ideas". Then, word embedding technology (such as Word2Vec, GloVe, etc.) is used to convert these words into vector representations of fixed dimensions as the most basic word-level implicit representation.

[0084] Furthermore, in order to obtain higher-level implicit representations, the system can use models such as convolutional neural networks (CNN), recurrent neural networks (RNN, especially LSTM or GRU) or Transformer to capture features at the phrase, sentence or even paragraph level. For example, through the convolutional layer and pooling layer of CNN, the system can extract local feature maps from the text, which can be regarded as implicit representations at the phrase level. The RNN or Transformer model can use sequence information to generate implicit representations at the sentence or paragraph level that contain context.

[0085] Ultimately, for a given educational evaluation subtext, the computer system will generate multiple sets of implicit representations with different granularities, which will serve as inputs for subsequent recognition steps.

[0086] In step S42, the computer system processes the multi-granular implicit representation generated in step S41 using a preliminary recognition strategy to preliminarily screen out text units that may contain information of interest. Preliminary recognition is usually based on some relatively simple but efficient feature matching or classification methods, aiming to quickly narrow the search scope of information of interest.

[0087] For example, in the initial recognition process, the computer system can use a rule-based approach or a simple machine learning classifier. For example, the system can predefine a set of keywords or phrases as identifiers of information of interest (such as "active participation", "creative", etc.), and search for these identifiers in the text through a string matching algorithm. In addition, the system can also use pre-trained classifiers (such as support vector machines, logistic regression, etc.) to classify implicit representations to determine whether they belong to the category of information of interest. For the processing of multi-granular implicit representations, the system can perform initial recognition on the representation of each granularity separately. For example, at the word level representation, the system can recognize words such as "active participation" and "creative"; at the phrase level representation, the system may further recognize phrases such as "active participation in discussion" and "creative views". These initial recognition results will provide candidate regions for subsequent second-order recognition.

[0088] In step S43, the computer system further refines and confirms the primary information of interest recognition result obtained in step S42 to improve the accuracy and reliability of recognition. Second-order recognition usually involves more complex machine learning models or deep learning networks, which can use richer contextual information and more sophisticated feature representations to recognize information of interest.

[0089] For example, in the second-order recognition process, the computer system can use deep learning-based object detection or named entity recognition (NER) models. These models can process complex text data and accurately locate the specific location (such as the start and end position) and category (such as positive evaluation, negative evaluation, key behavior, etc.) of the information of interest in the text.

[0090] Taking the object detection model as an example, the system can use the initial recognition results as candidate regions (Region of Interest, RoI), and use deep neural networks (such as Faster R-CNN, YOLO, etc.) to further extract features and classify these regions. The model will output the bounding box of each candidate region and its corresponding category label and confidence score. By setting an appropriate threshold, the system can filter out high-confidence information of interest as the final recognition result.

[0091] For named entity recognition tasks, the system may use LSTM or Transformer-based sequence labeling models (such as BiLSTM-CRF). These models can label each word or token in the text sequence and identify entities belonging to a specific category (such as names of people, places, institutions, and in this scenario, positive and negative reviews, etc.). Through the second-order recognition process, the computer system can more accurately determine the specific location and category of the information of interest.

[0092] In step S44, the computer system determines the second-order information of interest identification results obtained in step S43 as the first information of interest identification results of each educational evaluation subtext. These results not only contain the specific content and category of the information of interest, but also accurately mark their positions in the text, providing valuable information resources for subsequent data analysis and mining.

[0093] Through the synergy of the above four steps, the computer system can efficiently and accurately extract key evaluation information from the education evaluation text. The multi-granularity implicit representation strategy ensures comprehensive coverage of text information; the primary recognition strategy quickly narrows the search scope of information of interest; the second-order recognition strategy further improves the accuracy and reliability of recognition through complex machine learning models. The final first-order information of interest recognition results not only help education managers quickly understand the core content of the evaluation text, but also provide strong data support for subsequent education decisions.

[0094] In one embodiment, implicit representations of different granularities are obtained by executing a shared operator, wherein the shared operator includes a feature extraction module, at least one first shared branch operator and at least one second shared branch operator in the same composition mode; the first shared branch operator and the second shared branch operator both have multiple stage modules connected in sequence. Step S41, performing implicit representations of different granularities on the education evaluation subtext, and obtaining multiple implicit representations of different granularities corresponding to the education evaluation subtext includes:

[0095] Step S411: performing feature extraction using a feature extraction module according to the education evaluation sub-text to obtain a feature extraction result of the education evaluation sub-text;

[0096] Step S412: performing stage processing using a plurality of sequentially connected stage modules in the first shared branch operator according to the feature extraction result, and obtaining a stage processing result of each stage module in the first shared branch operator;

[0097] Step S413: performing stage processing using multiple stage modules in the second shared branch operator according to the feature extraction result and the stage processing result of each stage module in the first shared branch operator, to obtain the stage processing result of each stage module in the second shared branch operator;

[0098] Step S414: The stage processing results of each stage module in the first shared branch operator and the stage processing results of each stage module in the second shared branch operator are used as implicit representations of multiple different granularities corresponding to the education evaluation sub-text.

[0099] As the core component of step S41, the shared operator aims to efficiently generate implicit representations of different granularities by sharing computing resources and feature information. The operator consists of a feature extraction module, a first shared branch operator, and a second shared branch operator. Each part undertakes a specific processing task and works together to achieve a comprehensive analysis of the education evaluation subtext.

[0100] The feature extraction module is a lightweight convolutional network that extracts preliminary low-level features from the original text. These features may include vocabulary-level embedding vectors, character-level n-gram features, or contextual representations obtained through a pre-trained language model. The role of the feature extraction module is to provide a unified input basis for subsequent processing. The first shared branch operator and the second shared branch operator have similar composition methods, but each is responsible for processing feature information at different levels. They both contain multiple sequentially connected stage modules (hierarchical networks), and the size of each stage module (such as convolution kernel size, step size, etc.) is different to meet the needs of feature capture at different granularities. The design concept of the shared branch operator is to reduce computational redundancy and increase the diversity of feature representation by reusing feature extraction results and staged processing.

[0101] In step S411, the computer system uses the feature extraction module to perform preliminary feature extraction on the education evaluation subtext. This process usually involves converting the text into a numerical representation (such as through word embedding) and applying a convolution operation to capture local features.

[0102] For example, suppose there is an evaluation subtext about a student's classroom performance: "Xiao Ming actively participated in discussions in class and demonstrated good teamwork." The computer system converts the text into a series of word embedding vectors, each of which represents the semantic representation of a word in the text. The feature extraction module then applies one or more convolutional layers to these vectors to extract local dependencies and patterns between words. The result of the convolution operation is a feature map that contains the representation of the text in a low-level feature space.

[0103] In step S412, the computer system further processes the feature extraction result using the first shared branch operator. This process is implemented through multiple sequentially connected stage modules, each of which is responsible for converting the input features into a higher-level representation.

[0104] For example, let's continue with the evaluation subtext above. In the first shared branch operator, the feature extraction result is passed as input to the first stage module. This module may be a convolution layer with a larger convolution kernel to capture a wider range of contextual information. After being processed by this module, the output features are passed to the next stage module for further abstraction and refinement. The output of each stage module represents the feature representation of the text at different levels, gradually rising from the vocabulary level to the phrase or sentence level.

[0105] In step S413, the computer system uses the feature extraction result and the stage processing result of the first shared branch operator as input to drive the processing flow of the second shared branch operator. This design aims to fuse feature information from different sources to generate richer and more diverse implicit representations.

[0106] For example, in the second shared branch operator, each stage module receives two inputs: one is the direct output of the feature extraction module; the other is the output of the corresponding stage module in the first shared branch operator. Through some form of fusion operation (such as feature concatenation, weighted sum, etc.), these two inputs are merged into a unified feature vector as the input of the current stage module. This fusion strategy enables the second shared branch operator to simultaneously utilize low-level and high-level feature information to generate a more comprehensive feature representation. As the processing flow deepens, the output of each stage module contains more context-related feature information, thereby supporting more sophisticated identification of information of interest. In step S414, the computer system uses the output of each stage module in the first shared branch operator and the second shared branch operator as multiple implicit representations of different granularities corresponding to the educational evaluation subtext. These representations cover different levels of feature information from low-level to high-level and from local to global.

[0107] Through the processing flow of the above steps, the computer system successfully extracts multiple implicit representations of different granularities from the education evaluation subtext. These representations not only contain rich text feature information but also achieve effective utilization of computing resources through the design of shared operators. In subsequent steps, these implicit representations will be used as input for identification of information of interest to support more accurate and comprehensive information extraction tasks. This multi-granularity implicit representation method in the embodiment of the present application helps to capture the subtle differences and complex patterns in the text, thereby providing education managers with a more in-depth and detailed student performance analysis report.

[0108] In one embodiment, step S412, performing stage processing using a plurality of sequentially connected stage modules in the first shared branch operator according to the feature extraction result to obtain the stage processing result of each stage module in the first shared branch operator, includes:

[0109] Step S4121: using the first stage module in the first shared branch operator to perform stage processing according to the feature extraction result, to obtain the first stage processing result of the first shared branch operator;

[0110] Step S4122: Set the value of s from 1 to G-1 to complete the following operation: use the s+1th stage module in the first shared branch operator to perform stage processing according to the sth stage processing result of the first shared branch operator to obtain the s+1th stage processing result of the first shared branch operator, where G is the number of multiple stage modules connected in sequence.

[0111] In step S4121, the computer system performs preliminary processing on the feature extraction result using the first stage module in the first shared branch operator. This stage module usually has a larger size (such as a larger convolution kernel, step size or receptive field) so as to be able to capture more extensive and significant feature patterns in the text.

[0112] For example, suppose the feature extraction result is a feature map, which contains the representation of the educational evaluation subtext in the low-level feature space. This feature map may be a three-dimensional array, whose dimensions correspond to the number of feature channels, the length of the text sequence, and the feature dimension. In the first stage module, the computer system may use a convolution layer with a larger convolution kernel to perform a convolution operation on the feature map.

[0113] The convolution operation moves on the feature map in a sliding window manner, and calculates the weighted sum of the elements in the window and the convolution kernel at each position (usually with a bias term), and then introduces nonlinearity through an activation function (such as ReLU) to obtain a new feature map as output. This output feature map is the first stage processing result of the first shared branch operator, which contains the text features that have been initially abstracted and refined.

[0114] In step S4122, the computer system adopts an iterative method to pass the processing results of the sth stage module as input to the s+1th stage module for further processing. This process starts from s=1 and continues until s=G-1, where G is the total number of stage modules in the first shared branch operator. The size of each subsequent stage module is usually smaller than the previous stage module in order to capture more fine-grained text features.

[0115] For example, taking the second stage module as an example, the computer system passes the output feature map of the first stage module as input to the second stage module. In the second stage module, a convolution layer with a smaller convolution kernel may be used to further refine the features. Due to the reduction in the size of the convolution kernel, the module is able to capture more local and subtle text feature patterns. After the convolution operation and activation function processing, the second stage module outputs a new feature map as the second stage processing result.

[0116] As s increases, each subsequent stage module performs similar processing with the output of the previous stage as input. Each stage module optimizes the quality of feature representation by adjusting its internal parameters (such as convolution kernel weights and biases). These parameters are usually obtained through supervised learning on a large amount of labeled data.

[0117] In the progressively deeper processing flow, each stage module is dedicated to extracting text features at a specific level. As the size of the stage module gradually decreases and the feature representation gradually becomes abstract, this process can generate a feature hierarchy, which contains feature information of different granularities from low to high, from local to global. The feature hierarchy has significant advantages in the embodiments of the present application. First, it can capture a variety of feature patterns in the text, including features at the vocabulary level, features at the phrase level, and features at the sentence or paragraph level. This multi-level feature representation helps to understand the text content more comprehensively and improve the accuracy of identifying information of interest. Secondly, the feature hierarchy also supports flexible feature selection and combination. In subsequent steps, the computer system can select feature representations of different levels for combination and optimization as needed to adapt to different recognition tasks and data set characteristics. For example, in some cases, more emphasis may be placed on features at the vocabulary level; while in other cases, more emphasis may be placed on features at the sentence or paragraph level. Finally, the feature hierarchy also helps to improve computational efficiency and generalization capabilities. By sharing computing resources and feature information (i.e., feature extraction results and outputs of lower-level stage modules) in the early stages, the computer system can generate rich feature representations without increasing excessive computational burden. At the same time, since lower-level features are usually more general and stable, they also show better generalization performance across different datasets and tasks.

[0118] In summary, step S412 uses multiple sequentially connected stage modules in the first shared branch operator to perform step-by-step in-depth processing on the feature extraction results of the education evaluation subtext, generating a feature hierarchy structure containing multi-level feature information. This structure not only enriches the feature representation of the text but also improves the accuracy and efficiency of subsequent identification of information of interest.

[0119] In one embodiment, step S413, performing stage processing using multiple stage modules in the second shared branch operator according to the feature extraction result and the stage processing result of each stage module in the first shared branch operator to obtain the stage processing result of each stage module in the second shared branch operator, includes:

[0120] Step S4131: superimposing the feature extraction result and the stage processing result of the s-th stage module of the first shared branch operator to obtain a first superposition result;

[0121] Step S4132: performing stage processing using the first stage module in the second shared branch operator according to the first superposition result to obtain the first stage processing result of the second shared branch operator;

[0122] Step S4133: Set the value of s from 1 to G-1 to complete the following operations: superimpose the sth stage processing result of the first shared branch operator and the sth stage processing result of the second shared branch operator to obtain the s+1th superposition result, and use the s+1th stage module in the second shared branch operator to perform stage processing according to the s+1th superposition result to obtain the s+1th stage processing result of the second shared branch operator.

[0123] In step S4131, the computer system superimposes (or fuses) the feature extraction result with the processing result of the s-th stage module of the first shared branch operator. The purpose of this step is to combine the original feature information with the feature representation that has been preliminarily abstracted and refined to form more comprehensive input data for subsequent processing. For example, assume that the feature extraction result is a three-dimensional feature map F, whose dimensions are the number of feature channels C1, the text sequence length L, and the feature dimension D1. The processing result of the s-th stage module of the first shared branch operator is a feature map Fs of the same three-dimensional shape, but the number of feature channels may be different (assuming it is Cs), and the text sequence length and feature dimension may be the same as F or appropriately adjusted to match.

[0124] The superposition operation is usually performed on the feature channel dimension, that is, the feature channels of F and Fs are spliced ​​together to form a new feature map F1. If the number of feature channels of F and Fs is different, appropriate padding or truncation may be required to ensure that they can be spliced ​​on the same dimension. The number of feature channels of the spliced ​​feature map F1 will be C1+Cs, while the text sequence length and feature dimension remain unchanged. The superposition operation can be expressed as a mathematical splicing operation, that is:

[0125] F1 = Concat(F, Fs);

[0126] Among them, Concat represents the concatenation operation, which concatenates F and Fs into a new feature map F1 in the feature channel dimension.

[0127] After obtaining the first superposition result F1, the computer system processes F1 using the first stage module of the second shared branch operator. This stage module may include multiple neural network components such as convolutional layers, activation functions, and pooling layers to further refine and abstract features.

[0128] For example, the first stage module of the second shared branch operator may be a subnetwork containing multiple convolutional layers. These convolutional layers use different convolutional kernels to capture local patterns in the feature map F1 and introduce nonlinearity through activation functions (such as ReLU). After the convolution and activation operations, pooling operations (such as maximum pooling or average pooling) may be required to reduce the size of the feature map and reduce the amount of computation. Assuming that the first stage module contains a convolutional layer and a ReLU activation function, the convolution operation can be expressed as:

[0129] Y = σ(W*F1+b);

[0130] Where Y is the output feature map of the convolution operation, σ is the ReLU activation function, W is the weight matrix of the convolution kernel, * represents the convolution operation, and b is the bias term. The ReLU activation function is defined as:

[0131] σ(x)=max(0,x);

[0132] After the convolution and activation operations, the output feature map obtained is the first stage processing result of the second shared branch operator.

[0133] In step S4133, the computer system uses an iterative method to superimpose the s-th stage processing result of the first shared branch operator with the s-th stage processing result of the second shared branch operator, and pass the superimposed result as input to the s+1-th stage module of the second shared branch operator for processing. This process is iteratively executed starting from s=1 until s=G-1. For example, taking s=1 as an example, after obtaining the 1st stage processing result of the second shared branch operator, the computer system superimposes the 1st stage processing result Fs of the first shared branch operator with the 1st stage processing result (here it is assumed that the output of the second shared branch operator has been appropriately adjusted to match the size of Fs). The superposition operation is similar to that in step S4131 and is also performed on the feature channel dimension. The superimposed result is passed as input to the 2nd stage module of the second shared branch operator for processing. The processing flow of the 2nd stage module is similar to that of the 1st stage module, and may include components such as convolutional layers, activation functions, and pooling layers. The output feature map obtained after processing is the 2nd stage processing result of the second shared branch operator. As s increases, each subsequent stage uses the superposition and processing results of the previous stage as input for similar processing. Each stage module optimizes the quality of feature representation by adjusting its internal parameters (such as convolution kernel weights and biases). These parameters are usually obtained through supervised learning on a large amount of labeled data. In the process of step-by-step superposition and stage processing, the computer system not only fuses feature information from different sources (feature extraction results and processing results of the first shared branch operator), but also gradually refines and abstracts feature representations through the processing of multiple stage modules. This multi-level fusion and processing strategy enables the second shared branch operator to generate a richer and more diverse set of feature representations, providing more choices and possibilities for subsequent interesting information recognition tasks.

[0134] In the embodiment of the present application, by fusing the feature extraction result with the processing result of the first shared branch operator, the computer system can capture multiple feature patterns in the text and generate a more comprehensive feature representation. This comprehensive feature representation helps to more accurately identify interesting information points in the text, such as key evaluation words, phrases or sentences. Secondly, by the way of step-by-step superposition and stage processing, the computer system can gradually refine and abstract feature representations, thereby generating feature sets of different granularities. These feature sets not only include low-level features (such as features at the vocabulary level), but also include high-level features (such as features at the sentence or paragraph level). This multi-granular feature representation set helps to more flexibly adapt to different recognition tasks and data set characteristics, and improve the accuracy and robustness of the recognition results. Finally, the implementation method of step S413 also embodies the design idea of ​​modularity and reusability. Through the combined use of shared operators and stage modules, the computer system can flexibly build and expand feature extraction networks to adapt to different application scenario requirements. This design idea not only reduces the complexity and cost of model design but also improves the versatility and maintainability of the model. In summary, step S413 generates a rich and diverse feature representation set by fusing multi-source feature information, superimposing and processing in stages, which provides strong support for the subsequent task of identifying information of interest. This feature extraction strategy in the embodiment of the present application has important application value and is expected to promote the development and innovation of related technologies.

[0135] In one embodiment, the phase module includes a reference phase module and a moving phase module. Step S4122, using the s+1th phase module in the first shared branch operator to perform phase processing according to the sth phase processing result of the first shared branch operator to obtain the s+1th phase processing result of the first shared branch operator, includes:

[0136] Step S41221: performing stage processing using the s+1th reference stage module in the first shared branch operator according to the sth stage processing result of the first shared branch operator to obtain the s+1th reference stage processing result of the first shared branch operator;

[0137] Step S41222: Use the s+1th moving stage module in the first shared branch operator to perform stage processing according to the s+1th benchmark stage processing result of the first shared branch operator, and use the s+1th moving stage processing result of the first shared branch operator as the s+1th stage processing result of the first shared branch operator.

[0138] In step S41221, the computer system uses the s+1th benchmark stage module in the first shared branch operator to perform stage processing on the sth stage processing result to generate the s+1th benchmark stage processing result. The main function of the benchmark stage module is to perform stable, benchmark feature extraction and transformation on the input features, providing a reliable reference framework for subsequent processing. For example, assume that the sth stage processing result is a feature map Fs, which contains the feature representation of the text at a certain level. The benchmark stage module may include one or more convolutional layers, activation functions, and possible normalization layers (such as batch normalization) for further feature refinement of Fs.

[0139] Specifically, the convolution layer moves on Fs in a sliding window manner, calculates the weighted sum of the elements in the window and the convolution kernel, and introduces nonlinearity through the activation function to generate a new feature map. The activation function usually chooses ReLU (Rectified Linear Unit) because it can effectively suppress negative values ​​and retain positive values, which helps to speed up training and prevent the gradient disappearance problem. The normalization layer is used to adjust the distribution of the feature map to make it more suitable for subsequent processing. After being processed by the benchmark stage module, the s+1th benchmark stage processing result (denoted as Fbs) not only retains the important information in the input feature map Fs, but also introduces new feature representations through operations such as convolution and activation, providing a richer information basis for subsequent processing.

[0140] In step S41222, the computer system then uses the s+1th moving stage module in the first shared branch operator to further process the s+1th reference stage processing result Fbs to generate the s+1th stage processing result (i.e., the final output). Different from the reference stage module, the main function of the moving stage module is to capture and transform the feature graph more flexibly and dynamically to adapt to the complex and changeable information patterns in the text.

[0141] For example, the moving phase module may contain similar components as the baseline phase module (such as convolutional layers, activation functions, etc.), but its parameters (such as weights and biases of convolutional kernels) may be more flexible and changeable. In addition, the moving phase module may also introduce some special operations or layers (such as dilated convolution, deformable convolution, etc.) to increase the model's ability to capture feature maps and flexibility. Dilated convolution increases the receptive field by inserting holes (i.e., zero padding) between convolution kernel elements, allowing the convolution operation to capture a wider range of information without increasing the number of parameters. This operation is particularly suitable for scenarios that need to capture long-distance dependencies. Deformable convolution allows the convolution kernel to be sampled offset on the feature map, thereby more flexibly capturing key information points in the feature map. By learning the offset parameters, deformable convolution can adaptively adjust the sampling position to better adapt to the complex and changeable feature patterns in the text. In the moving phase module, these special operations or layers are combined with components such as convolutional layers and activation functions to further process the s+1 baseline phase processing result Fbs. The processing result of the s+1th stage obtained after processing not only contains richer and more dynamic feature representations, but also can better adapt to the complex and changeable information patterns in the text.

[0142] In step S4122, the reference phase module and the mobile phase module do not work in isolation, but collaborate with each other to jointly capture and refine text features. The reference phase module provides a stable reference framework for the mobile phase module, enabling the mobile phase module to achieve more flexible and dynamic feature capture while maintaining overall stability. The mobile phase module enhances the model's capture capability and adaptability by introducing special operations and layers, enabling the model to better cope with complex and changing information patterns in text.

[0143] This synergy not only improves the model's ability to capture text features, but also enhances the model's robustness and generalization ability. In the embodiments of the present application, this synergy is particularly important. Because educational evaluation texts often contain rich information points and complex contextual relationships, the model needs to have strong feature capture and extraction capabilities to accurately identify the key information points.

[0144] In the embodiment of the present application, step S4122 introduces the combined use of the benchmark phase module and the mobile phase module, and the computer system can realize the multi-dimensional capture and refinement of text features. This multi-dimensional feature representation not only contains rich information points but also has stronger robustness and generalization ability, which helps to improve the accuracy and reliability of subsequent identification of information of interest. Secondly, by adjusting the internal parameters and structures of the benchmark phase module and the mobile phase module (such as convolution kernel size, step size, void rate, offset, etc.), the computer system can flexibly adapt to different educational evaluation text characteristics and recognition task requirements. This flexibility enables the model to be more widely used in various educational evaluation scenarios and achieve good results. Finally, the implementation method of step S4122 also embodies the idea of ​​modular design. By dividing the feature extraction process into multiple reusable modules (such as benchmark phase module and mobile phase module), the computer system can more conveniently expand and optimize the model. This modular design not only reduces the complexity and cost of model design but also improves the versatility and maintainability of the model.

[0145] In summary, step S4122 achieves multi-dimensional capture and refinement of text features by introducing the combined use of the baseline phase module and the mobile phase module.

[0146] In one embodiment, the reference stage module includes a reference frame, a normalization unit, and a forward neural unit. Based on this, step S41221, according to the s-th stage processing result of the first shared branch operator, the s+1-th reference stage module in the first shared branch operator is used to perform stage processing to obtain the s+1-th reference stage processing result of the first shared branch operator, including:

[0147] Step S412211: using a normalization unit to perform normalization processing according to the s-th stage processing result of the first shared branch operator to obtain a first normalization processing result;

[0148] Step S412212: using the reference frame to perform self-saliency recognition according to the normalization processing result to obtain a reference frame recognition result;

[0149] Step S412213: superimposing the s-th stage processing result of the first shared branch operator and the reference frame recognition result to obtain a reference frame superposition result;

[0150] Step S412214: performing normalization processing using a normalization unit according to the reference frame superposition result to obtain a second normalization processing result;

[0151] Step S412215: using the forward neural unit to perform recognition according to the second normalization processing result to obtain a first forward result;

[0152] Step S412216: Superimpose the first forward result and the reference frame superposition result, and use the obtained first forward superposition result as the s+1th reference stage processing result of the first shared branch operator.

[0153] In step S412211, the computer system uses a normalization unit to normalize the sth stage processing result of the first shared branch operator. Normalization is a commonly used technique in deep learning, which aims to accelerate the convergence speed of the model and improve the generalization ability of the model by adjusting the distribution of data. In the embodiment of the present application, normalization is particularly important because educational evaluation texts usually have a long sequence length and complex feature distribution, which can easily lead to gradient vanishing or explosion problems during model training.

[0154] For example, suppose the result of the processing in the sth stage is a feature map Fs, whose dimension is [C, H, W], where C represents the number of feature channels, and H and W represent the height and width of the feature map, respectively. The normalization unit may use batch normalization or layer normalization to process Fs. Taking batch normalization as an example, its processing process can be expressed as:

[0155]

[0156] Among them, μ and σ 2 are the mean and variance of the feature map Fs on each feature channel, ∈ is a very small number (such as 10^-5) to prevent zero division errors, and γ and β are learnable scaling and offset parameters. After batch normalization, the obtained Normalized_Fs is the first normalization result.

[0157] In step S412212, the computer system uses the reference frame to perform self-saliency recognition (or self-attention processing) on ​​the first normalized processing result. The self-attention mechanism is an important breakthrough in the field of deep learning in recent years. It allows the model to dynamically pay attention to the key information points in the sequence when processing sequence data. In the embodiment of the present application, the self-attention mechanism is particularly useful because it can help the model capture the key evaluation words or phrases in the evaluation text. For example, the reference frame can be understood here as a special attention weight generation mechanism. It may include one or more neural network layers (such as fully connected layers, convolutional layers, etc.) for learning the importance weights of each position (or feature channel) from the first normalized processing result. Then, these weights are used to perform weighted summation on the first normalized processing result to obtain the reference frame recognition result. Specifically, the reference frame may first perform dimensionality reduction processing on the first normalized processing result through a convolutional layer (to reduce the amount of calculation and improve generalization ability), and then generate an attention weight matrix through a softmax function. Each row of the matrix represents the attention weight of a position (or feature channel) to all other positions. Finally, by performing a matrix multiplication operation on the weight matrix and the first normalization processing result, the reference frame recognition result can be obtained.

[0158] In step S412213, the computer system superimposes the s-th stage processing result of the first shared branch operator with the reference frame recognition result (or called feature fusion). The superposition operation is a feature fusion method commonly used in deep learning. It generates a new feature map by adding feature maps from different sources at corresponding positions (or multiplying elements by elements, etc.). In an embodiment of the present application, the superposition operation helps to combine the original features with the key features extracted by the self-attention mechanism to generate a richer and more comprehensive feature representation. For example, assume that the reference frame recognition result is a feature map Attn_Fs with the same dimension ([C, H, W]) as the s-th stage processing result of the first shared branch operator. The superposition operation can simply add Attn_Fs to Fs at corresponding positions:

[0159] Fused_Fs[c,h,w]=Fs[c,h,w]+Attn_Fs[c,h,w];

[0160] Among them, Fused_Fs is the result of the superposition of the reference frame. Through the superposition operation, Fused_Fs not only retains the information in the original feature map Fs but also incorporates the key feature information extracted by the self-attention mechanism.

[0161] In step S412214, the computer system re-normalizes the result of the reference frame overlay. The purpose of this step is similar to that of step S412211, which is to accelerate the convergence speed of subsequent processing and improve the generalization ability of the model by adjusting the distribution of data. However, unlike step S412211, the normalization process here is performed after the overlay operation, so it can take into account the impact of the overlay operation on the feature distribution.

[0162] For example, similar to step S412211, the re-normalization process may also adopt methods such as batch normalization or layer normalization. Taking batch normalization as an example, its processing process is the same as the description in step S412211, except that the input data is changed to the reference frame superposition result Fused_Fs. The feature map obtained after the re-normalization process is the second normalization process result.

[0163] In step S412215, the computer system uses the feedforward neural unit to perform recognition processing on the second normalization processing result. The feedforward neural unit is usually a neural network submodule containing multiple fully connected layers (or convolutional layers, etc.), which is used to further abstract and transform the input features. In the embodiment of the present application, the feedforward neural unit can capture more abstract and advanced information patterns in the feature map, thereby providing strong support for the subsequent identification of information of interest.

[0164] For example, a feedforward neural unit may contain multiple fully connected layers (for one-dimensional features) or convolutional layers (for two-dimensional feature maps such as sequences of text embedding vectors). Taking the convolutional layer as an example, the feedforward neural unit may gradually extract high-level features in the feature map by stacking multiple convolutional layers. Each convolutional layer uses different convolution kernel sizes and steps to capture feature patterns of different scales, and introduces nonlinearity through activation functions (such as ReLU). The feature map obtained after processing by multiple convolutional layers is the recognition result of the feedforward neural unit (i.e., the first forward result).

[0165] In step S412216, the computer system performs a final superposition operation on the first forward result and the reference frame superposition result, and outputs the superposition result as the s+1th reference stage processing result of the first shared branch operator. The purpose of this step is to combine the recognition result of the forward neural unit with the original superposition feature to generate a more comprehensive and robust feature representation.

[0166] For example, the final superposition operation is similar to the superposition operation in step S412213, except that the input data is changed to the first forward result and the reference frame superposition result. The feature map obtained by the final superposition operation is the s+1th reference stage processing result of the first shared branch operator. This result not only contains the high-level features extracted by the forward neural unit, but also retains the key information points in the original superposition features, providing strong support for subsequent processing steps.

[0167] Through the detailed analysis and example description of the above steps, it can be seen that the reference stage module in step S41221 realizes the refined extraction and conversion of text features by integrating multiple components such as reference frames, normalization units and forward neural units. This process not only reflects the powerful ability of deep learning in feature extraction, but also demonstrates the unique advantages of self-attention mechanism in sequence data processing.

[0168] In one embodiment, the moving phase module includes a moving frame, a normalization unit, and a forward neural unit. Based on this, step S41222, according to the s+1th reference phase processing result of the first shared branch operator, the s+1th moving phase module in the first shared branch operator is used to perform phase processing to obtain the s+1th moving phase processing result of the first shared branch operator, including:

[0169] Step S412221: using a normalization unit to perform normalization processing according to the s+1th benchmark stage processing result of the first shared branch operator to obtain a third normalization processing result;

[0170] Step S412222: using the moving frame to perform self-saliency recognition according to the third normalization processing result to obtain a moving frame recognition result;

[0171] Step S412223: superimposing the s+1th reference stage processing result of the first shared branch operator and the moving frame recognition result to obtain a moving frame superposition result;

[0172] Step S412224: performing normalization processing using a normalization unit according to the moving frame superposition result to obtain a fourth normalization processing result;

[0173] Step S412225: using the forward neural unit to perform recognition according to the fourth normalization processing result to obtain a second forward result;

[0174] Step S412226: superimpose the second forward result and the moving frame superposition result, and use the obtained second forward superposition result as the (s+1)th moving stage processing result of the first shared branch operator.

[0175] In step S412221, the computer system uses the normalization unit to normalize the processing results of the s+1th benchmark stage. This step is similar to the normalization process in step S412211, which aims to accelerate the convergence speed of the model and improve the generalization ability of the model by adjusting the distribution of data. However, since the input data has been processed by the benchmark stage module, the normalization in this step may focus more on maintaining the stability and consistency of the data.

[0176] For example, assume that the result of the s+1th benchmark stage processing is a feature map Fbs, whose dimensions and features are the same as those described in step S412211. The normalization unit again processes Fbs using methods such as batch normalization or layer normalization to obtain a third normalization processing result Norm_Fbs. The processing process is consistent with the description in step S412211, but the input data becomes Fbs.

[0177] In step S412222, the computer system uses the moving frame to perform self-saliency recognition on the third normalization processing result Norm_Fbs. Unlike the reference frame in the reference stage module, the design of the moving frame is more flexible and can capture key information points that change dynamically in the text. This step realizes dynamic focusing on text features through the self-attention mechanism, which helps the model better understand the text content. For example, the moving frame may contain multiple learnable parameters that define how the attention weights are generated. Similar to the reference frame, the moving frame may also process Norm_Fbs through one or more neural network layers (such as fully connected layers, convolutional layers, etc.) to generate an attention weight matrix. However, unlike the reference frame, the attention weights of the moving frame may be more localized or have specific spatial patterns to adapt to dynamically changing information points in the text.

[0178] Assume that the moving box is processed by a convolution layer to reduce the dimension of Norm_Fbs, and the attention weight matrix Attn_Matrix is ​​generated by the softmax function. Then, Attn_Matrix is ​​multiplied with Norm_Fbs to obtain the moving box recognition result Attn_Fbs. This process realizes the dynamic focus and importance weighting of text features.

[0179] In step S412223, the computer system performs a superposition operation on the s+1th reference stage processing result Fbs and the moving frame recognition result Attn_Fbs. This step is similar to the reference frame superposition in step S412213, but the input data is changed to Fbs and Attn_Fbs. The superposition operation generates a new feature map Fuse_Fbs by adding the two feature maps at corresponding positions (or multiplying them element by element, etc.), which combines the original features and the key features extracted by the self-attention mechanism.

[0180] In step S412224, the computer system normalizes the moving frame superposition result Fuse_Fbs again. This step is similar to the normalization processing in step S412214 and step S412221, and is intended to further adjust the distribution of data and provide stable input for subsequent processing. For example, the normalization unit processes Fuse_Fbs again using batch normalization or layer normalization to obtain the fourth normalization processing result Norm_Fuse_Fbs. The processing process is consistent with the description in the previous step, but the input data becomes Fuse_Fbs.

[0181] In step S412225, the computer system uses the forward neural unit to identify the fourth normalization processing result Norm_Fuse_Fbs. The forward neural unit is a submodule comprising multiple neural network layers, which is used to further abstract and transform the input features. In the embodiment of the present application, the forward neural unit can capture more abstract and advanced information patterns in the feature map, providing strong support for the subsequent identification of information of interest. For example, the forward neural unit may include multiple convolutional layers, activation function layers (such as ReLU), pooling layers, etc. These layers are stacked and combined to form a complex feature extraction and conversion network. Taking the convolutional layer as an example, the forward neural unit may perform a convolution operation on Norm_Fuse_Fbs through multiple convolution kernels, and introduce nonlinearity through the activation function to generate a new feature map. After being processed by multiple convolutional layers and pooling layers, the obtained feature map is the recognition result of the forward neural unit (i.e., the second forward result Fwd_Result).

[0182] In step S412226, the computer system performs a final superposition operation on the second forward result Fwd_Result and the moving frame superposition result Fuse_Fbs, and outputs the superposition result as the processing result of the s+1th moving stage of the first shared branch operator. This step is intended to combine the recognition result of the forward neural unit with the original superposition feature to generate a more comprehensive and robust feature representation. For example, the final superposition operation is similar to the superposition operation in steps S412213 and S412223, and the input data becomes Fwd_Result and Fuse_Fbs. The feature map obtained by the final superposition operation is the processing result of the s+1th moving stage. This result not only contains the high-level features extracted by the forward neural unit, but also retains the key information points in the original superposition features, providing strong support for subsequent processing steps.

[0183] Through the detailed analysis and example description of the above steps, it can be seen that the mobile phase module in step S41222 realizes the dynamic capture and conversion of text features by integrating multiple components such as mobile boxes, normalization units and forward neural units. This process not only reflects the powerful ability of deep learning in feature extraction and conversion, but also demonstrates the unique advantages of the self-attention mechanism in sequence data processing. In the embodiment of the present application, this feature extraction and conversion strategy helps the model to better understand the text content and improve the accuracy and efficiency of identifying information of interest.

[0184] In one embodiment, the preliminary information of interest recognition is obtained by executing at least one first multi-granularity feature representation algorithm and at least one second multi-granularity feature representation algorithm of the same architecture, the number of the first multi-granularity feature representation algorithms is equal to the number of the first shared branch operators, and the number of the second multi-granularity feature representation algorithms is equal to the number of the second shared branch operators. Based on this, step S42 performs preliminary information of interest recognition on multiple implicit representations of different granularities to obtain preliminary information of interest recognition results corresponding to the implicit representation of each granularity, including:

[0185] The following operations are performed for any implicit representation of a granularity corresponding to the educational evaluation subtext:

[0186] Step S421: using the first multi-granularity feature representation algorithm to perform preliminary information of interest recognition according to the s-th stage processing result of the first shared branch operator, to obtain a preliminary information of interest recognition result corresponding to the s-th stage processing result of the first shared branch operator;

[0187] Step S422: using a second multi-granularity feature representation algorithm to perform preliminary information of interest recognition according to the s-th stage processing result of the second shared branch operator, to obtain a preliminary information of interest recognition result corresponding to the s-th stage processing result of the second shared branch operator;

[0188] Step S423: Merge the preliminary information of interest recognition result corresponding to the s-th stage processing result of the first shared branch operator and the preliminary information of interest recognition result corresponding to the s-th stage processing result of the second shared branch operator to obtain the preliminary information of interest recognition result corresponding to the implicit representation of each granularity.

[0189] In this embodiment, the preliminary information of interest recognition is performed by a first multi-granularity feature representation algorithm and a second multi-granularity feature representation algorithm of the same architecture. The number of these algorithms is equal to the number of the first shared branch operator and the second shared branch operator, respectively. This design allows the system to process the output of each shared branch operator independently and specifically, thereby making full use of the advantages of multi-granularity implicit representation.

[0190] In step S421, the computer system uses the corresponding first multi-granularity feature representation algorithm to perform preliminary information of interest identification for the s-th stage processing result of the first shared branch operator. This process aims to extract preliminary information of interest from the output of the first shared branch operator to provide a basis for subsequent processing.

[0191] For example, assume that the first shared branch operator includes three stage modules, which output three implicit representations of different granularities Fs1, Fs2 and Fs3 respectively. For the sth stage (assuming s=2), its output is Fs2. At this time, the computer system will select the first multi-granularity feature representation algorithm A1 corresponding to the second stage of the first shared branch operator (assuming that there are three first multi-granularity feature representation algorithms, corresponding to the three stage modules respectively).

[0192] Algorithm A1 may be a model based on a convolutional neural network (CNN) or a recurrent neural network (RNN), which can further process and analyze Fs2 to identify the information of interest. This recognition may include the extraction of key words, phrases, or specific patterns, depending on the design and training objectives of the algorithm. After processing by algorithm A1, the computer system obtains the preliminary information of interest recognition result IR1_s corresponding to Fs2. This result may be a feature map marked with information of interest, a list containing words of interest, or any other form of data structure, depending on the output design of the algorithm.

[0193] In step S422, the computer system uses the corresponding second multi-granularity feature representation algorithm to perform preliminary information of interest recognition for the s-th stage processing result of the second shared branch operator. This process is similar to step S421, but the processed input data comes from the second shared branch operator.

[0194] For example, continuing with s=2, the second stage output of the second shared branch operator is Fbs2 (note that the naming here may be different from the output of the first shared branch operator to distinguish the outputs of the two branches). The computer system selects the second multi-granularity feature representation algorithm B1 corresponding to the second stage of the second shared branch operator (again assuming that there are three second multi-granularity feature representation algorithms in total).

[0195] Algorithm B1 may have a similar architecture to Algorithm A1 but different parameters and training data to adapt to the characteristics of the second shared branch operator output. It also processes and analyzes Fbs2 to identify the information of interest therein. After processing by Algorithm B1, the computer system will obtain a preliminary information of interest identification result IR2_s corresponding to Fbs2. The form and content of this result may be similar to IR1_s, but it contains the information of interest extracted from the output of the second shared branch operator.

[0196] In step S423, the computer system merges the preliminary information of interest recognition results obtained in step S421 and step S422. This process aims to integrate the recognition results of the two shared branch operators to generate a more comprehensive and accurate preliminary information of interest recognition result. For example, for the case of s=2, the computer system merges IR1_s and IR2_s. The way of merging may depend on the specific form of the recognition result. If the recognition result is in the form of a feature graph, it can be merged by feature concatenation, weighted summation or other fusion strategies. If the recognition result is in the form of a vocabulary list, it can be sorted and filtered by set union operation or according to the importance of the vocabulary. Assuming that IR1_s and IR2_s are both lists containing words of interest, the computer system can obtain the merged recognition result IR_merged_s by seeking the union:

[0197] IR_merged_s=IR1_s∪IR2_s;

[0198] However, in practical applications, directly seeking a union may introduce duplicate words. To avoid this, the computer system may need to first perform a deduplication operation on the two lists before merging them, that is:

[0199] IR_merged_s=(IR1_s\IR2_s)∪(IR2_s\IR1_s)∪(IR1_s∩IR2_s);

[0200] Through the detailed analysis and example explanation of the above steps, we can see how the preliminary interesting information recognition process in step S42 is achieved through the collaborative work of the multi-granularity feature representation algorithm and the shared branch operator. This process not only reflects the powerful ability of deep learning in processing complex text data, but also shows how to improve the accuracy and efficiency of information recognition through parallel processing and result merging.

[0201] In one embodiment, the second-order information of interest recognition is obtained by using a hierarchical algorithm, and the hierarchical algorithm includes a suggestion box generator and multiple downsampling operators. Based on this, step S43 performs second-order information of interest recognition on each primary information of interest recognition result to obtain a second-order information of interest recognition result corresponding to each primary information of interest recognition result, including:

[0202] For any preliminary information of interest identification result, complete the following operations:

[0203] Step S431: generating a suggestion box using a suggestion box generator according to the preliminary information of interest recognition result, and obtaining a suggestion box generation result corresponding to the preliminary information of interest recognition result;

[0204] Step S432: downsampling using multiple downsampling operators according to the suggestion box generation result and the preliminary interesting information recognition result, to obtain a downsampling result corresponding to each downsampling operator;

[0205] Step S433: Determine the down-sampling result as the second-order information of interest recognition result.

[0206] In the embodiment of the present application, although the primary interesting information recognition can capture some key information points in the text, it is often difficult to be accurate to the specific vocabulary, phrase or sentence level because it is based on a relatively simple feature representation and recognition strategy. Therefore, in order to further improve the accuracy and precision of recognition, the computer system introduces a second-order interesting information recognition process. This process aims to refine the primary recognition results through more complex and advanced algorithms to generate more accurate and detailed recognition results.

[0207] In this embodiment, the second-order information of interest recognition is performed using a hierarchical algorithm, which includes a key component, a region proposal network (RPN) and multiple downsampling operators. These components work together to achieve a refined processing of the primary recognition results.

[0208] The suggestion box generator is a neural network component for generating potential regions of interest. It can slide a window on the input feature map and predict whether there is a region of interest (i.e., a suggestion box) within each window. RPN is usually used in conjunction with a convolutional neural network (CNN) to reduce the amount of computation and improve recognition efficiency by sharing convolutional layers. The downsampling operator is an operation used to reduce the dimension of the data, which is usually achieved through pooling or a convolution operation with a stride greater than 1. During the downsampling process, the system selectively retains important information and discards unimportant details, thereby refining and compressing the data. In this embodiment, multiple downsampling operators are used to further process the features within the suggestion box to generate downsampling results of different scales.

[0209] In step S431, the computer system generates a suggestion box using RPN according to the preliminary information of interest recognition result. This step aims to extract potential regions of interest from the preliminary recognition result and provide a target region for subsequent downsampling operations.

[0210] For example, assume that the preliminary information of interest recognition result is a set of text fragments marked with key words or phrases. In this step, the computer system maps these text fragments to the corresponding feature map (if the preliminary recognition result itself is based on the representation of the feature map, this step can be omitted). Then, RPN slides the window on the feature map and predicts a binary label (whether it contains the region of interest) and a bounding box regression parameter (used to adjust the window position and size) for each window. Through post-processing operations such as non-maximum suppression (NMS), the system can filter out the N highest-scoring suggestion boxes from all predicted windows as the final generation results. These suggestion boxes not only contain the location information of the region of interest, but also provide candidate regions for further processing.

[0211] In step S432, the computer system uses multiple downsampling operators to perform downsampling processing according to the suggestion box generation result and the primary interesting information recognition result. This step is intended to generate more detailed and accurate second-order recognition results through refined feature extraction and compression.

[0212] For example, for the feature map area within each proposal box, the computer system performs cropping and scaling operations according to the position and size of the proposal box to ensure the consistency of the input data. Then, the system applies multiple downsampling operators to downsample these areas. The downsampling operators may include pooling layers, convolution layers, or a combination of them at different scales. Each operator extracts features and reduces data dimensions through different strides, kernel sizes, and padding methods. For example, the first downsampling operator may use a 3x3 pooling kernel and a pooling operation with a stride of 2 to reduce the size of the feature map and retain important information; the second downsampling operator may use a convolution layer with a different number of filters to further extract high-level features. After processing by multiple downsampling operators, the system will obtain a series of downsampling results of different scales. These results not only contain the key information in the original feature map, but also remove redundancy and noise through downsampling operations to improve the robustness and accuracy of recognition.

[0213] In step S433, the computer system determines the down-sampling result as the second-order information of interest recognition result and uses it as the input or output of subsequent processing. This step is the last link of the hierarchical algorithm and the summary of the entire second-order recognition process. For example, for the down-sampling result in each suggestion box, the computer system can perform further feature fusion or decoding operations to generate the final recognition result. Feature fusion can be achieved by splicing, weighted summation or attention mechanism, aiming to combine feature information of different scales to form a more comprehensive representation. The decoding operation may involve converting the feature map back to the original data space (such as a text sequence) for subsequent analysis or display. The final generated second-order information of interest recognition result may be a vocabulary, phrase or sentence set marked with precise location and category. These results not only contain the key information in the primary recognition, but also are refined and optimized through the second-order recognition process to make the recognition results more accurate and reliable.

[0214] Through the detailed analysis and example explanation of the above steps, we can see how the second-order interesting information recognition process in step S43 is achieved through the collaboration of hierarchical algorithms and multiple high-level components. This process not only improves the accuracy and precision of recognition, but also demonstrates the powerful ability of deep learning in processing complex text data.

[0215] In one embodiment, step S432, downsampling is performed using multiple downsampling operators according to the suggestion box generation result and the preliminary interesting information recognition result to obtain downsampling results corresponding to each downsampling operator, including:

[0216] Step S4321: downsampling is performed using a first downsampling operator according to the suggestion box generation result and the preliminary interesting information recognition result, to obtain a first downsampling result corresponding to the first downsampling operator;

[0217] Step S4322: Downsampling is performed using the r+1th downsampling operator according to the rth downsampling result and the preliminary information of interest recognition result to obtain the r+1th downsampling result corresponding to the r+1th downsampling operator, wherein r is a positive integer not greater than H-1, H is a positive integer not greater than 2, H is the number of multiple downsampling operators, and the accuracy of the r+1th downsampling result is greater than that of the rth downsampling result.

[0218] In step S4321, the computer system uses the first downsampling operator to perform downsampling processing based on the suggestion box generation result and the preliminary information of interest recognition result. This step is the initial stage of the downsampling process, and its goal is to perform preliminary feature extraction and compression on the potential area of ​​interest, laying the foundation for subsequent more refined processing.

[0219] For example, suppose the suggestion box generator has generated a set of suggestion boxes for a preliminary interesting information recognition result, each of which marks a potential key area in the text. At the same time, the preliminary recognition result may be a list of key words or phrases that correspond to key information points in the text.

[0220] In the first downsampling stage, the computer system can select a relatively simple downsampling operator, such as a pooling layer with a large stride or a convolution layer with a small number of filters. This operator processes the feature map area within the proposal box, reduces the data dimension and extracts key features through pooling or convolution operations. For example, if the feature map is a three-dimensional array [C, H, W] (where C is the number of channels, H is the height, and W is the width), the first downsampling operator can use a 3x3 pooling kernel and a pooling operation with a stride of 2 to generate the first downsampling result. In this way, the size of the feature map will be halved (in height and width), while the number of channels remains unchanged. The processed feature map will contain fewer pixels but retain key information.

[0221] In step S4322, the computer system uses an iterative method to perform further downsampling processing using the r+1th downsampling operator according to the rth downsampling result and the preliminary information of interest recognition result. This process is iteratively executed starting from r=1 until r=H-1, where H is the total number of downsampling operators and H is not greater than 2.

[0222] In each iteration, the r+1th downsampling operator will be more refined and complex than the previous one. This usually means that it will have more filters, smaller step sizes, or more complex nonlinear transformations to capture more detailed information. In this way, the system can gradually improve the accuracy and detail of the downsampling results.

[0223] For example, assuming there are two downsampling operators (ie, H=2), the iteration process will be performed once: from the first downsampling result to the second downsampling result.

[0224] In the second downsampling stage, the computer system can choose an operator that is more complex than the first downsampling operator. For example, it may be a subnetwork containing multiple convolutional layers, each followed by an activation function (such as ReLU) and an optional batch normalization layer. These layers will process the first downsampling results in turn, extracting higher-level features through convolution operations and introducing nonlinearity through activation functions.

[0225] In this process, the preliminary interesting information recognition results may be used as a form of attention mechanism or constraint to guide the processing of the downsampling operator. For example, the system can give higher weights to the feature map regions corresponding to the key words or phrases in the preliminary recognition results when calculating the convolution kernel weights. In this way, the downsampling results will be more inclined to preserve the information of these key areas.

[0226] The final second downsampling result will have higher accuracy and more detailed feature representation than the first downsampling result. It not only contains the key information in the original text, but also removes redundancy and noise through the step-by-step downsampling process, thereby improving the robustness and accuracy of recognition.

[0227] In practical applications, the selection and optimization of downsampling operators is a complex process that requires consideration of multiple factors such as computing resource limitations, recognition accuracy requirements, and data set characteristics. The following are some possible optimization strategies:

[0228] Network architecture design: Design a suitable network architecture according to specific task requirements, including the number and type of downsampling operators and the connection between them. For example, residual connections can be used to avoid the gradient vanishing problem or dense connections can be used to enhance feature propagation.

[0229] Parameter initialization and training: Use appropriate parameter initialization methods (such as Xavier initialization or He initialization) to speed up the training process and avoid gradient explosion or vanishing problems. At the same time, use effective optimization algorithms (such as Adam or RMSprop) to update network parameters to improve training efficiency and stability.

[0230] Regularization and dropout: To prevent overfitting, regularization terms (such as L1 regularization or L2 regularization) can be introduced during training or a dropout layer can be added to the network to randomly discard some neuron outputs. This can enhance the generalization ability of the model and improve its performance on unseen data.

[0231] Attention mechanism: As mentioned above, the initial information of interest recognition results can be used as a form of attention mechanism to guide the downsampling process. By giving higher weights to key areas, the model's ability to capture important information is enhanced.

[0232] Post-processing and decoding: After the downsampling process is completed, the downsampling results can be further post-processed and decoded to generate the final recognition results. For example, non-maximum suppression (NMS) can be used to filter suggestion boxes or conditional random fields (CRF) can be used to optimize sequence labeling results.

[0233] Through the detailed analysis and example explanation of the above steps, we can see how the step-by-step downsampling process in step S432 achieves the refinement and optimization of the primary information of interest recognition results through the collaboration of multiple downsampling operators. This process not only improves the accuracy and detail of recognition, but also demonstrates the hierarchical strategy and step-by-step refinement idea of ​​deep learning when processing complex text information.

[0234] Figure 2 A hardware entity diagram of a computer system provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and the processor 1001 implements the steps in the method of any of the above embodiments when executing the program.

[0235] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. An educational evaluation target data processing method based on artificial intelligence, characterized in that: The method comprises: Predicting the interest distribution density sequence of the education evaluation text to obtain the interest distribution density sequence of the education evaluation text, wherein the interest distribution density sequence represents the dispersion of the interest information in the education evaluation text, and the interest distribution density sequence prediction includes multiple filtering samplings of different degrees; Performing information binning on the distribution density sequence of interest to obtain a plurality of binned information segments in the distribution density sequence of interest, wherein the information binning is clustering the distribution density sequence of interest; The education evaluation text is intercepted according to the multiple boxed information paragraphs to obtain multiple corresponding education evaluation sub-texts, wherein each of the education evaluation sub-texts includes multiple pieces of the information of interest; Identifying information of interest for each of the education evaluation sub-texts and the education evaluation text one by one, and obtaining first information of interest identification results corresponding to each of the education evaluation sub-texts and the education evaluation text; Each of the first information of interest recognition results is merged to obtain a second information of interest recognition result corresponding to the education evaluation text.

2. The method according to claim 1, characterized in that The estimating the interest distribution density sequence of the education evaluation text to obtain the interest distribution density sequence of the education evaluation text includes: According to the educational evaluation text, an optimized interest distribution density sequence estimation method is used to estimate the interest distribution density sequence to obtain the interest distribution density sequence of the educational evaluation text; The optimized estimation method of the distribution density sequence of interest is obtained by optimizing the following process: According to the educational evaluation sample text, the interest distribution density sequence is estimated using the interest distribution density sequence estimation method to be optimized to obtain an estimated interest distribution density sequence of the educational evaluation sample text; Using a cost calculation function according to the estimated interest distribution density sequence of the education evaluation sample text and the prior interest distribution density sequence of the education evaluation sample text, the training cost of the education evaluation sample text is obtained; The algorithm parameters of the interest distribution density sequence prediction algorithm to be optimized are optimized according to the training cost of the education evaluation sample text.

3. The method according to claim 2, characterized in that The method of estimating the distribution density sequence of interest using an optimized distribution density sequence estimation algorithm according to the education evaluation text to obtain the distribution density sequence of interest of the education evaluation text includes: According to the educational evaluation text, the following steps are performed using the optimized interest distribution density sequence estimation method: Performing filtering sampling of the educational evaluation texts at different levels for a plurality of times to obtain a plurality of implicit representation sets with different granularities; Performing a fusion operation on the multiple implicit representation sets of different degrees to obtain a fused implicit representation set; Feature extraction is performed on the fused implicit representation set to obtain a distribution density sequence of interest of the education evaluation text.

4. The method according to claim 1, characterized in that: The performing information binning on the distribution density sequence of interest to obtain a plurality of binned information segments in the distribution density sequence of interest includes: Identifying sequence units in the distribution density sequence of interest to obtain a probability density function of the sequence units; According to the probability density function of the sequence unit, the sequence units corresponding to the same distribution law are binned to obtain a plurality of binning information segments in the distribution density sequence of interest; The identifying of the information of interest for each of the education evaluation sub-texts and the education evaluation text one by one, and obtaining the first information of interest identification results corresponding to each of the education evaluation sub-texts and the education evaluation text, respectively, includes: Complete the following operations for any of the educational evaluation sub-texts: Performing implicit representations of different granularities on the education evaluation subtext to obtain multiple implicit representations of different granularities corresponding to the education evaluation subtext; Performing preliminary information of interest recognition on the plurality of implicit representations of different granularities to obtain preliminary information of interest recognition results corresponding to the implicit representation of each granularity; Performing second-order information of interest recognition on each of the primary-order information of interest recognition results to obtain a second-order information of interest recognition result corresponding to each of the primary-order information of interest recognition results, wherein the recognition accuracy of the second-order information of interest recognition is greater than the recognition accuracy of the primary-order information of interest recognition; The second-order information-of-interest identification result is determined as the first information-of-interest identification result.

5. The method according to claim 4, characterized in that The implicit representations of different granularities are obtained by executing a shared operator, wherein the shared operator includes a feature extraction module, at least one first shared branch operator and at least one second shared branch operator in the same composition mode; the first shared branch operator and the second shared branch operator both have a plurality of stage modules connected in sequence; The implicit representation of the education evaluation subtext at different granularities is performed to obtain a plurality of implicit representations at different granularities corresponding to the education evaluation subtext, including: Using the feature extraction module to extract features according to the education evaluation subtext, to obtain a feature extraction result of the education evaluation subtext; Performing stage processing using a plurality of sequentially connected stage modules in the first shared branch operator according to the feature extraction result, to obtain a stage processing result of each stage module in the first shared branch operator; Performing stage processing using multiple stage modules in the second shared branch operator according to the feature extraction result and the stage processing result of each stage module in the first shared branch operator to obtain the stage processing result of each stage module in the second shared branch operator; The stage processing results of each stage module in the first shared branch operator and the stage processing results of each stage module in the second shared branch operator are used as implicit representations of multiple different granularities corresponding to the educational evaluation sub-text.

6. The method according to claim 5, characterized in that The step of performing stage processing using a plurality of sequentially connected stage modules in the first shared branch operator according to the feature extraction result to obtain a stage processing result of each stage module in the first shared branch operator includes: Performing stage processing using the first stage module in the first shared branch operator according to the feature extraction result to obtain the first stage processing result of the first shared branch operator; The value of s is changed from 1 to G-1 to complete the following operation: according to the s-th stage processing result of the first shared branch operator, the s+1-th stage module in the first shared branch operator is used to perform stage processing to obtain the s+1-th stage processing result of the first shared branch operator, where G is the number of the multiple sequentially connected stage modules; The step of performing stage processing using multiple stage modules in the second shared branch operator according to the feature extraction result and the stage processing result of each stage module in the first shared branch operator to obtain the stage processing result of each stage module in the second shared branch operator includes: Superimposing the feature extraction result and the stage processing result of the s-th stage module of the first shared branch operator to obtain a first superposition result; Performing stage processing using the first stage module in the second shared branch operator according to the first superposition result to obtain a first stage processing result of the second shared branch operator; The value of s is changed from 1 to G-1 to complete the following operations: superimpose the sth stage processing result of the first shared branch operator and the sth stage processing result of the second shared branch operator to obtain the s+1th superposition result, and use the s+1th stage module in the second shared branch operator to perform stage processing according to the s+1th superposition result to obtain the s+1th stage processing result of the second shared branch operator.

7. The method according to claim 6, characterized in that The stage module includes a reference stage module and a moving stage module; the stage processing is performed using the s+1th stage module in the first shared branch operator according to the sth stage processing result of the first shared branch operator to obtain the s+1th stage processing result of the first shared branch operator, including: Performing stage processing using the s+1th reference stage module in the first shared branch operator according to the sth stage processing result of the first shared branch operator to obtain the s+1th reference stage processing result of the first shared branch operator; According to the s+1th benchmark stage processing result of the first shared branch operator, the s+1th moving stage module in the first shared branch operator is used to perform stage processing, and the s+1th moving stage processing result of the first shared branch operator is used as the s+1th stage processing result of the first shared branch operator.

8. The method according to claim 7, characterized in that The reference stage module includes a reference frame, a normalization unit and a forward neural unit; the step of performing stage processing using the s+1th reference stage module in the first shared branch operator according to the sth stage processing result of the first shared branch operator to obtain the s+1th reference stage processing result of the first shared branch operator includes: Performing normalization processing using the normalization unit according to the s-th stage processing result of the first shared branch operator to obtain a first normalization processing result; According to the normalization processing result, the reference frame is used to perform self-saliency recognition to obtain a reference frame recognition result; Superimposing the s-th stage processing result of the first shared branch operator and the reference frame recognition result to obtain a reference frame superposition result; Performing normalization processing using the normalization unit according to the reference frame superposition result to obtain a second normalization processing result; Using the forward neural unit to perform recognition according to the second normalization processing result to obtain a first forward result; Superimposing the first forward result and the reference frame superposition result, and using the obtained first forward superposition result as the s+1th reference stage processing result of the first shared branch operator; The movement phase module includes a movement frame, a normalization unit and a forward neural unit; The step of performing stage processing using the s+1th moving stage module in the first shared branch operator according to the s+1th benchmark stage processing result of the first shared branch operator to obtain the s+1th moving stage processing result of the first shared branch operator includes: Using the normalization unit to perform normalization processing according to the s+1th benchmark stage processing result of the first shared branch operator to obtain a third normalization processing result; Using the moving frame to perform self-saliency recognition according to the third normalization processing result to obtain a moving frame recognition result; Superimposing the s+1th benchmark stage processing result of the first shared branch operator and the moving frame recognition result to obtain a moving frame superposition result; Using the normalization unit to perform normalization processing according to the moving frame superposition result to obtain a fourth normalization processing result; Using the forward neural unit to perform recognition according to the fourth normalization processing result to obtain a second forward result; The second forward result and the moving frame superposition result are superimposed, and the obtained second forward superposition result is used as the (s+1)th moving stage processing result of the first shared branch operator.

9. The method according to claim 6, characterized in that The preliminary information of interest recognition is obtained by performing at least one first multi-granularity feature representation algorithm and at least one second multi-granularity feature representation algorithm of the same architecture, wherein the number of the first multi-granularity feature representation algorithms is equal to the number of the first shared branch operators, and the number of the second multi-granularity feature representation algorithms is equal to the number of the second shared branch operators; the preliminary information of interest recognition is performed on the multiple implicit representations of different granularities to obtain preliminary information of interest recognition results corresponding to the implicit representations of each granularity, including: The following operations are performed for any implicit representation of a granularity corresponding to the educational evaluation subtext: Using the first multi-granularity feature representation algorithm to perform preliminary information of interest recognition according to the s-th stage processing result of the first shared branch operator, to obtain a preliminary information of interest recognition result corresponding to the s-th stage processing result of the first shared branch operator; Using the second multi-granularity feature representation algorithm to perform preliminary information of interest recognition according to the s-th stage processing result of the second shared branch operator, to obtain a preliminary information of interest recognition result corresponding to the s-th stage processing result of the second shared branch operator; Merging the preliminary information of interest recognition result corresponding to the s-th stage processing result of the first shared branch operator and the preliminary information of interest recognition result corresponding to the s-th stage processing result of the second shared branch operator to obtain a preliminary information of interest recognition result corresponding to the implicit representation of each granularity; The second-order information of interest identification is performed using a hierarchical algorithm, the hierarchical algorithm comprising a suggestion box generator and a plurality of downsampling operators; The performing second-order information of interest recognition on each of the primary-order information of interest recognition results to obtain a second-order information of interest recognition result corresponding to each of the primary-order information of interest recognition results includes: For any of the preliminary information of interest identification results, complete the following operations: generating a suggestion box using the suggestion box generator according to the preliminary information of interest recognition result, and obtaining a suggestion box generation result corresponding to the preliminary information of interest recognition result; Performing downsampling using the multiple downsampling operators according to the suggestion box generation result and the preliminary interesting information recognition result, to obtain a downsampling result corresponding to each of the downsampling operators; Determining the downsampling result as the second-order information of interest recognition result; The downsampling is performed using the multiple downsampling operators according to the suggestion box generation result and the preliminary interesting information recognition result to obtain a downsampling result corresponding to each downsampling operator, including: Downsampling is performed using a first downsampling operator according to the suggestion box generation result and the preliminary interesting information recognition result to obtain a first downsampling result corresponding to the first downsampling operator; Downsampling is performed using the r+1th downsampling operator according to the rth downsampling result and the preliminary information of interest recognition result to obtain the r+1th downsampling result corresponding to the r+1th downsampling operator, wherein r is a positive integer not greater than H-1, H is a positive integer not greater than 2, H is the number of the multiple downsampling operators, and the accuracy of the r+1th downsampling result is greater than that of the rth downsampling result.

10. A computer system comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Methods, devices and computer-readable storage media for real-time speech recognition

    US20200219486A1

  • KR20210094324A