Student information input system

By combining pre-trained error correction models with manual proofreading, the problem of frequent spelling errors in student information entry is solved, fast and accurate information proofreading is achieved, and the accuracy and timeliness of student information are ensured.

CN120706414AInactive Publication Date: 2025-09-26WEIFANG NURSING VOCATIONAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510801976.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology has problems with frequent spelling errors and inaccurate proofreading during the student information entry process, resulting in untimely and inaccurate information entry.

Method used

A pre-trained error correction model combined with manual proofreading is used to perform error correction through information embedding, Transformer encoder, and RoBERTa pre-trained model to form a Soft-Mask mechanism and soft mask sequence, combined with logical verification and data entry verification to ensure information accuracy.

Benefits of technology

It achieves fast and accurate proofreading of student information, reduces omissions caused by human proofreading, and improves the accuracy and timeliness of information entry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706414A_ABST
    Figure CN120706414A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information input, in particular to a student information input system, which comprises an error correction module used for performing error correction and labeling on a typed text through a pre-trained error correction model so as to ensure the accuracy of student information; the error correction determination module is used for manually determining whether error correction is carried out or not based on the labeling result; the display module is used for displaying the recorded student information; error correction and labeling are carried out on a typed text through a pre-trained error correction model to ensure the accuracy of the student information, manual confirmation is carried out based on a labeling result, and through a mode of combining the model and manual work, not only is rapid proofreading improved, but also omission is avoided to a certain extent, and the accuracy of the student information is improved. And the accuracy of student information is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information entry, and more particularly to a student information entry system. Background Art

[0002] With the rise of the big data era and the rapid development of mobile Internet technology, traditional paper documents are gradually being replaced by electronic information. Nowadays, electronic text information such as news, e-books, e-mails, and e-newspapers have been deeply integrated into people's daily lives.

[0003] In schools, in order to better understand students' educational status, there is usually a unified education information entry system. Counselors enter students' information. Through the entered information, students' basic information and academic performance information can be understood. In the process of entering student information in text form, some spelling errors will inevitably occur. Therefore, it is particularly important to detect and correct errors in the input text. The existing error correction method mainly relies on manual review and modification word by word. In the case of a large amount of text information for school students, not only a lot of work is required, but also in the proofreading process, due to human factors, omissions and other reasons occur in the proofreading, resulting in inaccurate proofreading. Therefore, one of the problems that need to be solved at present is to quickly and accurately proofread the input text student information and understand the student information in a timely and accurate manner. Summary of the Invention

[0004] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a student information entry system, which helps to quickly proofread student information in text form and to a certain extent avoid omissions and improve the accuracy of student information.

[0005] The present invention provides a student information entry system, comprising: a typing module for typing text for describing student information; The error correction module is used to correct and annotate typed text using a pre-trained error correction model to ensure the accuracy of student information; An error correction determination module is used to manually determine whether to perform error correction based on the annotation results; Display module, used to display the entered student information; Storage module, used to store student information.

[0006] Furthermore, the pre-trained error correction model includes: an information embedding module, a detection network based on a Transformer encoder, and an error correction network based on a RoBERTa pre-trained model; Information embedding module, used to embed input text to generate text embedding; Transformer encoder-based detection network: used to predict the error rate of each character in the text embedding and form a soft-mask mechanism and soft mask sequence; The error correction network based on the RoBERTa pre-trained model is used to extract key information features from different parts of the input mask sequence and perform linear transformations, transforming the feature space into the target space. It then obtains the corrected new words through the Softmax activation function and outputs the final text.

[0007] Furthermore, the system further comprises: a verification module for performing data verification on the manually confirmed text; wherein the data verification includes logic verification and data storage verification; Logical verification includes: checking whether the Chinese character area code is within the predetermined range, whether the year, month, and day data are reasonable, and whether the code is complete; storage verification includes: checking whether the student name and student number form a unique mapping relationship.

[0008] Furthermore, the information embedding module includes: The character embedding layer is used to first insert special symbols [CLS] and [SEP] at the beginning and end of the text to identify the boundaries and special functions of the text; then each character in the text with special symbols inserted is mapped into a character embedding vector of a certain dimension; The paragraph embedding layer uses a two-channel vector representation mechanism to generate a paragraph embedding vector that marks which sentence each character belongs to; Position embedding layer, used to capture the position information of characters in the text to generate position embedding vectors; Fusion layer: fuses the character embedding vector, paragraph embedding vector, and position embedding vector in series to obtain text embedding , .

[0009] Furthermore, the error rate of each character in the predicted text embedding and forming a Soft-Mask mechanism and a soft mask sequence include: Embed text The input is processed by the stacked Transformer encoder and the error probability of each character in the text embedding is calculated through the MLP and Sigmoid activation function. The calculation formula is as follows:

[0010] in, represents the error probability, and Represent weight and bias respectively; Represents the hidden layer state parameters of the Transformer encoder; The calculated error probability as the mask embedding vector weight, The weighted summation is performed as the text embedding weight to form the Soft-Mask mechanism, and its calculation formula is as follows:

[0011] in, represents mask embedding, Represents text embedding.

[0012] Furthermore, the Transformer-based encoder is also used to identify all characters in the text embedding and output a set L of all character labels. ; in, The value of is 0 or 1, where label 1 indicates that the character is wrong and label 0 indicates that it is correct.

[0013] The beneficial effects of the present invention are: The present invention corrects and annotates the typed text through a pre-trained error correction model to ensure the accuracy of student information, and performs manual confirmation based on the annotated results. By combining the model and manual work, it not only helps to improve rapid proofreading and avoid omissions to a certain extent, but also further improves the accuracy of student information. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0015] Figure 1 A block diagram of a student information entry system provided by an embodiment of the present invention; Figure 2 This is a processing flow chart of a pre-trained error correction model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0016] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0017] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, each technical and scientific term used in this embodiment has the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0018] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0019] In the present invention, terms such as "upper", "lower", "left", "right", "front", "back", "vertical", "horizontal", "side", "bottom", etc. indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are relational words determined only for the convenience of describing the structural relationships of the various parts or elements of the present invention, and do not specifically refer to any part or element in the present invention, and should not be understood as limiting the present invention.

[0020] In the present invention, terms such as "fixed connection," "connected," and "connection" should be interpreted broadly to mean a fixed connection, an integral connection, or a detachable connection; a direct connection or an indirect connection through an intermediary. Relevant researchers or technicians in this field may determine the specific meanings of these terms in the present invention based on specific circumstances, and they should not be construed as limitations of the present invention.

[0021] Example 1: like Figure 1 As shown, this embodiment provides a student information entry system, including: a typing module for typing student information in text format, wherein the student information includes: student name, age, home address, contact information, student grades and other information.

[0022] The error correction module is used to correct student information in typed text format through a pre-trained error correction model to ensure the accuracy of the student information.

[0023] The error correction prompt module is used to mark the error correction information based on the error correction results; An error correction determination module is used to manually determine whether to perform error correction based on the marked error correction information; Information display module, used to display the entered student information; Information storage module, used to store student information.

[0024] The verification module is used to perform data verification on the manually confirmed text; wherein, data verification includes logic verification and data storage verification; Logical verification includes: checking for some low-level errors, such as whether the Chinese character area code should be within a certain range, whether the year, month, and day data are reasonable, whether the code is complete, etc. Entry verification includes: checking whether the student name and student ID form a unique mapping relationship; Among them, such as Figure 2 As shown in Figure 1, the pre-trained error correction model includes: an information embedding module, a detection network based on a Transformer encoder, and an error correction network based on a RoBERTa pre-trained model; For ease of understanding and explanation, assume that the input text is: .

[0025] A: Information embedding module, used to embed the input text to generate text embedding. The information embedding module includes character embedding layer, paragraph embedding layer, position embedding layer and fusion layer.

[0026] A1: The character embedding layer is used to: first, insert the special symbols [CLS] and [SEP] at the beginning and end of the text to identify text boundaries and special functions; then, map each character in the text with the special symbols into a character embedding vector of a certain dimension (for example, 768 dimensions).

[0027] A2: The paragraph embedding layer uses a dual-channel vector representation mechanism to generate a paragraph embedding vector that marks which sentence each character belongs to.

[0028] For ease of understanding, the dual-channel vector is represented as the first set of vectors and the second set of vectors. Specifically, the paragraph embedding layer is used to assign the value 0 to all tokens in the first sentence, while the second set of vectors assigns the value 1 to all tokens in the second sentence.

[0029] The dual-channel representation mechanism can distinguish the boundaries between different sentences, thereby better capturing the contextual information in the text.

[0030] A3: Position embedding layer, used to capture the position information of characters in a text sequence to generate a position embedding vector.

[0031] A4: Fusion layer: fuses the character embedding vector, paragraph embedding vector, and position embedding vector in series to obtain text embedding , ,in, Represents characters The text embedding vector is obtained by concatenating the character embedding vector, the position embedding vector and the paragraph embedding, i∈[1,n].

[0032] B: Transformer encoder-based detection network: used to predict the error rate of each character in the text embedding and form a soft-mask mechanism and soft mask sequence.

[0033] Specifically, the detection network based on the Transformer encoder is a classification model. The detection network based on the Transformer encoder includes a stacked Transformer encoder, a feedforward neural network (MLP), and a Sigmoid activation function, and specifically includes the following steps: B1: Embed text The input is processed by the stacked Transformer encoder and the error probability of each character in the text embedding is calculated through the MLP and Sigmoid activation function. The error probability calculation formula is as follows:

[0034] in, represents the error probability, and Represent weight and bias respectively; Represents the hidden layer state parameters of the Transformer encoder.

[0035] Among them, the Sigmoid activation function is used to effectively limit the error probability to the (0,1) interval. More specifically, The closer the value of is to 1, the greater the possibility that the word is considered a typo. The closer the value of is to 0, the less likely the word is to be considered a typo.

[0036] B2: The Transformer encoder is also used to identify all characters in the text embedding and output the set of all character labels L. .

[0037] in, The value of is 0 or 1, where label 1 indicates that the character is wrong and label 0 indicates that it is correct.

[0038] B3: The calculated error probability as the mask embedding vector weight, The weighted summation is performed as the text embedding weight to form the Soft-Mask mechanism, and its calculation formula is as follows:

[0039] in, represents mask embedding, represents text embedding; it can be seen that if the error probability is higher, the soft mask embedding The closer , otherwise the closer .

[0040] Thus, the soft mask sequence is obtained The soft masking mechanism enables the model to learn the internal relationships of sentences more flexibly, improving the understanding and generation quality of text.

[0041] C: An error correction network based on the RoBERTa pre-trained model, which extracts key information features from different parts of the input mask sequence and performs linear transformations, transforming the feature space into the target space. It then uses the Softmax activation function to obtain the corrected new words and outputs the final text. .

[0042] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the terminal embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.

[0043] In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can It can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, other divisions may be used. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not implemented. In addition, the coupling or direct coupling or communication connection shown or discussed between each other may be through some interface, or the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0044] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0045] In addition, it should be noted that the flowcharts in the accompanying drawings show the methods of the embodiments of the present disclosure. In the descriptions corresponding to the flowcharts or block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be performed substantially in parallel, or sometimes in the opposite order, which may depend on the functions involved. Each block in the block diagram and / or flow chart, and the combination of blocks in the block diagram and / or flow chart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or can be implemented using a combination of dedicated hardware and computer instructions.

[0046] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A student information entry system, characterized in that: include: A typing module is used to type texts used to describe student information; The error correction module is used to correct and annotate typed text using a pre-trained error correction model to ensure the accuracy of student information; An error correction determination module is used to manually determine whether to perform error correction based on the annotation results; Display module, used to display the entered student information; Storage module, used to store student information.

2. The student information entry system according to claim 1, characterized in that: The pre-trained error correction model includes: an information embedding module, a detection network based on a Transformer encoder, and an error correction network based on a RoBERTa pre-trained model; Information embedding module, used to embed input text to generate text embedding; Transformer encoder-based detection network: used to predict the error rate of each character in the text embedding and form a soft-mask mechanism and soft mask sequence; The error correction network based on the RoBERTa pre-trained model is used to extract key information features from different parts of the input mask sequence and perform linear transformations, transforming the feature space into the target space. It then obtains the corrected new words through the Softmax activation function and outputs the final text.

3. The student information entry system according to claim 1, characterized in that: The system further comprises: a verification module for performing data verification on the manually confirmed text; wherein the data verification includes logic verification and data storage verification; Logical verification includes: checking whether the Chinese character area code is within the predetermined range, whether the year, month, and day data are reasonable, and whether the code is complete; storage verification includes: checking whether the student name and student number form a unique mapping relationship.

4. The student information entry system according to claim 2, characterized in that: The information embedding module includes: The character embedding layer is used to first insert special symbols [CLS] and [SEP] at the beginning and end of the text to identify the boundaries and special functions of the text; then each character in the text with special symbols inserted is mapped into a character embedding vector of a certain dimension; The paragraph embedding layer uses a two-channel vector representation mechanism to generate a paragraph embedding vector that marks which sentence each character belongs to; Position embedding layer, used to capture the position information of characters in the text to generate position embedding vectors; Fusion layer: fuses the character embedding vector, paragraph embedding vector, and position embedding vector in series to obtain text embedding , .

5. The student information entry system according to claim 2, characterized in that: The error rate of each character in the predicted text embedding and forming a soft-mask mechanism and a soft mask sequence include: Embed text The input is processed by the stacked Transformer encoder and the error probability of each character in the text embedding is calculated through the MLP and Sigmoid activation function. The calculation formula is as follows: in, represents the error probability, and Represent weight and bias respectively; Represents the hidden layer state parameters of the Transformer encoder; The calculated error probability as the mask embedding vector weight, The weighted summation is performed as the text embedding weight to form the Soft-Mask mechanism, and its calculation formula is as follows: in, represents mask embedding, Represents text embedding.

6. The student information entry system according to claim 2, characterized in that: The Transformer-based encoder is also used to identify all characters in the text embedding and output a set L of all character labels. ; in, The value of is 0 or 1, where label 1 indicates that the character is wrong and label 0 indicates that it is correct.