Personalized service-oriented intelligent online evaluation system design method
By designing an intelligent online evaluation system for personalized services, using DSL created by JSON, FreeMarker and Jetbrains MPS, combined with deep learning model and Circle Loss loss function, the support for a variety of questions and the scalability of the system is achieved, solving the shortcomings of the existing online evaluation system, and improving the security and intelligence of the system.
Patent Information
- Application Number
- CN202510075741.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing online evaluation system has insufficient support for various question types and limited scalability, complex operation and chaotic permission management, making it difficult to deal with various question types that may appear in the future.
An intelligent online evaluation system for personalized services was designed, using DSL created by JSON, FreeMarker and Jetbrains MPS, combining deep learning models for feature extraction and model building, using Circle Loss as a loss function, training deep learning models for multi-label classification problems of algorithm tags, and building an online evaluation module to integrate the web page, server and evaluation end.
It has achieved support for a variety of questions, improved the scalability of the system, simplified the operation process, enhanced the security and intelligence of the system, and had high application value.
Smart Images

Figure CN119988220A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software development, and in particular to a design method of an intelligent online evaluation system for personalized services. Background Art
[0002] To support various programming competitions, a powerful online evaluation system (Online Judge) is necessary. The online evaluation system needs to automatically judge the correctness of the solution in real time and ensure that the solution does not exceed the limitations of space and time; OJ provides the function of automatic evaluation, which greatly reduces the burden of manual review and makes the teaching process more efficient. Therefore, it is also widely used in various computer competitions and teaching.
[0003] However, most of the existing OJ platforms are not very scalable in terms of evaluation, and cannot support the evaluation of communication questions, let alone deal with various types of questions that may appear in the future. In addition, they often encounter problems such as complex operations and chaotic authority management when using them. For some open source OJs, it is also difficult to carry out secondary development based on them. Summary of the invention
[0004] Purpose of the invention: The purpose of the present invention is to provide a design method for an intelligent online assessment system for personalized services, a general assessment method based on DSL created by JSON, FreeMarker and Jetbrains MPS, which solves the problems of insufficient support for multiple question types and limited scalability of existing online assessment systems.
[0005] Technical solution: The method for designing an intelligent online evaluation system for personalized services described in the present invention comprises the following steps: (1) Domain-specific languages created using different DSL design methods; (2) Design a sandbox based on Docker and Linux commands; (3) Use deep learning models for feature extraction and model building, use Circle Loss as the loss function, and train deep learning models for multi-label classification of algorithm labels; (4) Build an online evaluation module to integrate the web page, server, and evaluation end.
[0006] Furthermore, in step (1), different DSL design methods include JSON, FreeMarker, and JetbrainsMPS.
[0007] Furthermore, step (1) includes the following steps: (11) Use JSON to describe the test questions and time and space constraints. Use the C++nlohmann::json library to parse them during the evaluation. (12) Generate scripts for test data, use DSL-FreeMarker to describe, use <#list> syntax to describe a large number of regular scripts, and use Java to generate corresponding scripts during evaluation; (13) Use Jetbrains' metalanguage programming system MPS to create Strategy language, conduct concept design, editor design and generator design, and implement the evaluation of traditional questions, constructed questions, interactive questions and communication questions.
[0008] In the concept design, the Concept in Jetbrains MPS is used to describe the properties of a concept and determine the inheritance relationship of the concept. In the editor design, the vertical layout is used to design the syntax of each concept; in the generator design, each concept is mapped to a code snippet and Java class, and the specific code parameters are described by using the properties of the concept in combination with the LOOP macro, COPY_SRC macro, and Collection language; (14) After the DSL design is completed, the conversion from DSL to JAVA code to executable file is performed.
[0009] Furthermore, step (2) includes the following steps: (21) Use Docker to build an image and run it in the image; (22) Use Linux commands to limit time and space; (23) Use seccomp to restrict system calls; (24) Run the code using the nobody user.
[0010] Furthermore, step (3) includes the following steps: (31) Use the Codeforces website to obtain open source website datasets; (32) For multi-label classification problems, BERT is used for title description feature extraction, and deep CNN network is used for code feature extraction; (33) Use Circle Loss as the loss function of the model; the details are as follows: Circle Loss for a single label ;
[0011] in Represents the number of label categories, represents the target category, represents the score of the i-th category. The above formula is an approximation of the max function, and its essence is to make the scores of other categories as small as possible as compared to the target category; if extended to: each target category The scores of the non-target classes should be as large as possible. , that is, the following losses are obtained: ;
[0012] in represents the set of target categories, Represents the set of non-target categories.
[0013] When adding the empty class 0 during prediction, and outputting all labels greater than the empty class as answers, the loss function should be changed to: ;
[0014] in Represents the score of the empty class 0.
[0015] Furthermore, step (4) includes the following steps: (41) Web development based on Vue and Vant component library: Using Vue's responsiveness, component reuse, routing management and Vant component library, we can obtain the user module, question module and competition module of the web page. (42) Based on the Node.js server, use app.get() or app.post() to process various web page requests, make database and evaluation requests as needed, and return the corresponding results; (43) The C++ and Java-based evaluation machine receives the code path, question configuration path, evaluation type, and code type through command line parameters; uses the nlohmann::json library to parse the configuration file config.json and determines the evaluation behavior based on the parameters.
[0016] The intelligent online evaluation system design system for personalized services described in the present invention includes: DSL Design Module: for domain-specific languages created using different DSL design methods; Sandbox module: used to design sandbox based on Docker and Linux commands; Feature building module: used for feature extraction and model building of deep learning models, using Circle Loss as the loss function, and training deep learning models for multi-label classification problems of algorithm labels; Fusion module: used to build an online evaluation module, integrating the web page, server and evaluation end.
[0017] An electronic device described in the present invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the method for designing an intelligent online evaluation system for personalized services described in any one of the items is implemented.
[0018] A storage medium described in the present invention stores a computer program, and when the computer program is executed by a processor, it implements any one of the methods for designing an intelligent online evaluation system for personalized services.
[0019] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: The present invention uses different DSL design methods for different needs, has strong scalability, and is convenient for field experts to describe the evaluation method and for the evaluation machine to understand the evaluation method; in view of the fact that the online evaluation system needs to frequently run unknown programs, a sandbox design based on Docker and Linux commands is proposed to enhance the security of the system. Using deep learning models such as BERT and CNN for feature extraction and model building, and using Circle Loss as the loss function, the deep learning model is trained to perform multi-label classification problems of algorithm labels, and good results have been achieved. An online evaluation system has been constructed, and the integration of the web page, server and evaluation end has been realized to ensure that the above modules can be effectively put into use; the present invention takes into account both the security and intelligence of the system, and therefore has a high application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a system architecture diagram of the present invention.
[0021] Figure 2 Generate a flow chart for the data of the present invention; Figure 3 A command schematic diagram is generated using FreeMarker for the present invention; Figure 4 The inclusion and inheritance relationship diagram of the concepts of the DSL designed for the present invention; Figure 5 A schematic diagram of using Jetbrains MPS to design and run DSL for the present invention; Figure 6 This is a schematic diagram of a code feature extraction model using a deep CNN in the present invention; Figure 7 This is a schematic diagram of the topic description and code submission page of the online evaluation system of the present invention. DETAILED DESCRIPTION
[0022] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0023] like Figure 1As shown, an embodiment of the present invention provides a design and implementation method of an intelligent online evaluation system for personalized services, which includes four steps: design of a general evaluation method for an online evaluation system, design of a sandbox based on Docker and Linux commands, and design of an intelligent algorithm for label prediction. The system is mainly composed of a web page, a server, and an evaluation terminal. The Node.js server acts as a "middle layer" and communicates with the web page, database, and evaluation terminal. The evaluation terminal parses the general evaluation description file designed by DSL design tools such as Jetbrains MPS, and the code is sent to the label prediction algorithm for question label prediction. When evaluating, the evaluation terminal also needs to run the code in Docker.
[0024] S1. Design of a general evaluation method for online evaluation systems S11 directly uses JSON to describe information such as test question descriptions, time and space constraints, etc., which is already the best practice. During the evaluation, the C++ nlohmann::json library is used for parsing, which can manage JSON like an STL container.
[0025] S12 Figure 2 As shown in the figure, in the data generation process, the questioner uploads the generator and script to construct data and obtains the input of the test point. The generator's commands are in the form of a large number of gen 1 2>1.txt commands. In order to facilitate the maintenance of a large number of regular scripts, the existing DSL-FreeMarker is used for description. The <#list> syntax can be used to describe and generate a large number of commands. During the evaluation, since FreeMarker itself is a Java class library, Java can be used to generate the corresponding commands. Figure 3 This is a diagram of using FreeMarker to generate commands. The commented parts are the generated commands.
[0026] Regarding the evaluation method, S13 uses Jetbrains' metalanguage programming system MPS to create the Strategy language, conduct concept design, editor design, and generator design, and can implement the evaluation of traditional questions, construction questions, interactive questions, communication questions, and other question types.
[0027] In conceptual design, the Concept in Jetbrains MPS can be used to describe a concept and determine the inheritance relationship of the concept. In the specific implementation steps, it may include the definition of files, pipes (interactive questions), and the running of programs (user programs, checkers). The former inherits the abstract concept IO, and the latter inherits the abstract concept Executable. Figure 4 It is the inheritance relationship of concepts when using Jetbrains to design the DSL described in the test question.
[0028] In the editor design, taking the top-level concept Strategy as an example, its various steps can be arranged in a vertical layout, which will determine the syntax of the concept. The code is as follows.
[0029] Node cell layout: [- Strategy { name} ( / % Steps % / ) / empty cell: <default> -] In the generator design, first map the top-level concept Strategy to a Java class, and map the remaining concepts to code snippets. Take the file (File concept) as an example, use the LOOP macro and COPY_SRC macro to traverse all steps (Step concept), and use the Collection language to set the loop list as iteration sequence: (genContext, node)->sequence <node<> > { node.Steps.where({it => it.isInstanceOf(File);}); } And set the mapping rule Reduce_File, and map the attributes of the File concept to the corresponding code snippet.
[0030] The mapping of other concepts is also described in a similar way. For example, User and Checker are the execution of a program, which are mapped to Java using ProcessBuilder.
[0031] When the generator is finally completed, you can try to generate Java code. If the program does not report an error, you can make Strategy implement IMainClass as the main class of the Java project and try to run it in MPS, such as Figure 5 shown.
[0032] S2. Sandbox design based on Docker and Linux commands S21Docker uses sandbox technology. When trying to run a program, you can use Docker to build an image and run the program in the image.
[0033] S22In Linux, setitimer can limit the time occupied by a process. If the time exceeds the limit, a Timeout behavior can be set. The setrlimit function can limit the memory usage of a process. Space outside the range cannot be allocated successfully. Time and space restrictions can be set using Linux commands.
[0034] S23 uses seccomp to restrict system calls. The filtering rules are pre-set and enabled without additional context switching. Experiments show that the program running time remains almost unchanged, which reduces a lot of additional overhead compared to ptrace.
[0035] S24 uses the nobody user to run the code. Even if the program bypasses the sandbox rules, it can be protected from greater harm through Linux permission control. In C++, setuid and setgid can be used to implement this.
[0036] S3, label prediction intelligent algorithm S31 used Python crawlers to crawl the basic information of 3697 test questions and part of the AC code of each test question on the Codeforces website, and built a data set for model training. The final data set was divided into training set, validation set and test set at a ratio of 7:1:2, containing 2587, 370 and 740 test questions respectively. For the test question description information, only the main text was taken, and information such as sample input and sample output that was difficult to predict labels was ignored. For the AC code, all comments and some punctuation marks were removed, and the frequency of words appearing in the code was counted.
[0037] S32 This problem can be divided into three tasks: the first is the prediction of difficulty labels, which is equivalent to a regression problem; the second is the prediction of algorithm labels, which is a multi-label problem; and the third is the prediction of test question types, which is equivalent to a three-classification problem.
[0038] S33 For the difficulty prediction problem, the mean square error (MSELoss) is used as the evaluation indicator for this problem. For the multi-label prediction problem, the evaluation indicator is the intersection-over-union ratio (Jaccard Score), that is, the ratio of the intersection and union of the predicted label and the true label. The closer to 1, the better. For the three-classification problem, the accuracy (ACC), precision, recall and F1 score (F1 Score) are used for evaluation. The optimizer of the model is Adam, and the learning rate is 5e-4. The randomization seed of PyTorch is 42, and the number of epochs is 20. The loss function is the sum of MSELoss (difficulty prediction problem), Circle Loss (algorithm label prediction problem) and CrossEntropyLoss (question type prediction problem).
[0039] For the multi-label classification problem, S34 uses BERT to extract title description features, uses a deep CNN network to extract code features, and sends the extracted features to the fully connected layer for further classification. Figure 6 It is the model structure of the deep CNN network.
[0040] S35 uses Circle Loss as the loss function of the model for multi-label classification problems. Specifically, the single-label Circle Loss ;
[0041] in Represents the number of label categories, represents the target category, represents the score of the i-th category. The above formula is an approximation of the max function, and its essence is to make the scores of other categories as small as possible as compared to the target category; if extended to: each target category The scores of the non-target classes should be as large as possible. , that is, the following losses are obtained: ;
[0042] in represents the set of target categories, Represents the set of non-target categories.
[0043] When adding the empty class 0 during prediction, and outputting all labels greater than the empty class as answers, the loss function should be changed to: ;
[0044] in Represents the score of the empty class 0.
[0045] S36 baseline model design: BERT is directly used to extract features from code submissions that have not been preprocessed, and CrossEntropyLoss is used instead of Circle Loss for multi-label classification. The results of the comparative experiments using the proposed model and the baseline model are shown in Tables 1 and 2.
[0046] Table 1 Comparative experiments between the proposed model and the baseline model on difficulty prediction and multi-label prediction problems ;
[0047] Table 2 Comparative experiment of the proposed model and the baseline model on question type prediction ; S4. Design of online evaluation system S41 is a web-based development based on Vue and Vant component library. It uses Vue's responsiveness, component reuse, routing management and convenient Vant component library to implement user modules, question modules, competition modules, etc. on the web. Figure 7 This is the topic description and code submission page for the online evaluation system.
[0048] S42 is developed based on Node.js server, using app.get() or app.post() to process various web page requests, making database and evaluation requests as needed, and returning corresponding results.
[0049] S43 is developed based on C++ and Java evaluation machines, which receive code paths, test configuration paths, evaluation types, and code types through command line parameters. The nlohmann::json library is used to parse the configuration file config.json and determine the evaluation behavior based on the parameters.
[0050] The embodiment of the present invention also provides an intelligent online evaluation system design system for personalized services, including: DSL Design Module: for domain-specific languages created using different DSL design methods; Sandbox module: used to design sandbox based on Docker and Linux commands; Feature building module: used for feature extraction and model building of deep learning models, using Circle Loss as the loss function, and training deep learning models for multi-label classification problems of algorithm labels; Fusion module: used to build an online evaluation module, integrating the web page, server and evaluation end.
[0051] An embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, the method for fusing soil moisture based on XGB-SHAP-RFE feature optimization and SHAP interpretable analysis is implemented according to any one of the items.
[0052] An embodiment of the present invention also provides a storage medium storing a computer program. When the computer program is executed by a processor, it implements any one of the soil moisture fusion methods based on XGB-SHAP-RFE feature optimization and SHAP interpretable analysis.< / default>
Claims
1. A design method for an intelligent online evaluation system for personalized services, characterized in that: The following steps are involved: (1) Domain-specific languages created using different DSL design methods; (2) Design a sandbox based on Docker and Linux commands; (3) Use deep learning models for feature extraction and model building, use Circle Loss as the loss function, and train deep learning models for multi-label classification of algorithm labels; (4) Build an online evaluation module to integrate the web page, server, and evaluation end.
2. The method for designing an intelligent online evaluation system for personalized services according to claim 1, characterized in that: In step (1), different DSL design methods include JSON, FreeMarker and Jetbrains MPS.
3. According to the method for designing an intelligent online evaluation system for personalized services as described in claim 2, it is characterized in that: Step (1) includes the following steps: (11) Use JSON to describe the test questions and time and space constraints. Use the C++nlohmann::json library to parse them during the evaluation. (12) Generate scripts for test data, use DSL-FreeMarker to describe, use <#list> syntax to describe regular scripts, and use Java to generate corresponding scripts during evaluation; (13) Use Jetbrains' metalanguage programming system MPS to create Strategy language, conduct concept design, editor design, and generator design, and implement the evaluation of traditional questions, construction questions, interactive questions, and communication questions; In the concept design, the Concept in Jetbrains MPS is used to describe the properties of a concept and determine the inheritance relationship of the concept. In the editor design, the vertical layout is used to design the syntax of each concept; in the generator design, each concept is mapped to a code snippet and Java class, and the specific code parameters are described by using the properties of the concept in combination with the LOOP macro, COPY_SRC macro, and Collection language; (14) After the DSL design is completed, the conversion from DSL to JAVA code to executable file is performed.
4. The method for designing an intelligent online evaluation system for personalized services according to claim 1, characterized in that: Step (2) includes the following steps: (21) Use Docker to build an image and run it in the image; (22) Use Linux commands to limit time and space; (23) Use seccomp to restrict system calls; (24) Run the code using the nobody user.
5. The method for designing an intelligent online evaluation system for personalized services according to claim 1, characterized in that: Step (3) includes the following steps: (31) Use the Codeforces website to obtain open source website datasets; (32) For multi-label classification problems, BERT is used for title description feature extraction, and deep CNN network is used for code feature extraction; (33) Use Circle Loss as the loss function of the model; the details are as follows: Circle Loss for a single label ; in, Represents the number of label categories, represents the target category, represents the score of the i-th category; the above formula is an approximation of the max function, and its essence is to make the scores of other categories as small as possible as compared to the target category; if extended to: each target category The scores of the non-target classes should be as large as possible. , that is, the following losses are obtained: ; in represents the set of target categories, represents the set of non-target categories; When adding the empty class 0 during prediction, and outputting all labels greater than the empty class as answers, the loss function should be changed to: ; in Represents the score of the empty class 0.
6. The method for designing an intelligent online evaluation system for personalized services according to claim 1, characterized in that: Step (4) includes the following steps: (41) Web development based on Vue and Vant component library: Using Vue's responsiveness, component reuse, routing management and Vant component library, we can obtain the user module, question module and competition module of the web page. (42) Based on the Node.js server, use app.get() or app.post() to process various web page requests, make database and evaluation requests as needed, and return the corresponding results; (43) The C++ and Java-based evaluation machine receives the code path, question configuration path, evaluation type, and code type through command line parameters; uses the nlohmann::json library to parse the configuration file config.json and determines the evaluation behavior based on the parameters.
7. An intelligent online evaluation system design system for personalized services, characterized in that: include: DSL Design Module: for domain-specific languages created using different DSL design methods; Sandbox module: used to design sandbox based on Docker and Linux commands; Feature building module: used to extract features and build models using deep learning models, use Circle Loss as the loss function, and train deep learning models for multi-label classification problems of algorithm labels; Fusion module: used to build an online evaluation module, integrating the web page, server and evaluation end.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, a method for designing an intelligent online evaluation system for personalized services according to any one of claims 1 to 6 is implemented.
9. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a method for designing an intelligent online evaluation system for personalized services according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Competition data processing system and method based on virtual environment
CN110704135A
Dialogue response method and system, electronic equipment and storage medium
CN115203393A
Label classification model training and object screening method and device and storage medium
CN115700550A
Named entity identification method and device based on label prompt and medium
CN116595979A
Text processing method and device, and method and device for training text processing model
WO2025007819A1