Scoring method and device for AI (Artificial Intelligence) competition practical operation questions and computer equipment
By building a pre-built scoring system for automatic scoring of practical questions in AI competitions, the problem of poor consistency under the traditional manual review model is solved, the unification of scoring standards and cost reduction is achieved, and the real-time feedback mechanism is provided, which improves scoring efficiency and accuracy.
Patent Information
- Application Number
- CN202510541476.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-05
AI Technical Summary
The score consistency of AI competition practical questions in the prior art is poor, the traditional manual review model is inefficient, the evaluation is incomplete, and the fairness and consistency cannot be guaranteed.
By building a pre-built scoring system, the data set, code data and target competition models generated by participating users are automatically obtained, and the scoring system is used for comprehensive scoring, including the fusion of code scoring, data scoring and model scoring to ensure that scoring is performed in a standardized environment.
The scoring standards of AI competition practical questions have been achieved, the utilization of scoring resources has been reduced, the scoring cost has been reduced, and the real-time feedback mechanism has been provided, which has improved scoring efficiency and accuracy.
Smart Images

Figure CN120430682A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a scoring method, device, and computer equipment for practical questions in AI competitions. Background Art
[0002] With the rapid development of artificial intelligence technology and the deepening expansion of its industry applications, the demand for talents with full-chain AI engineering practice capabilities continues to grow in fields such as artificial intelligence engineering, intelligent science and technology, and electronic information engineering. To adapt to this trend, higher education institutions and vocational training institutions are widely using AI practice competitions as a core training method, focusing on assessing participants' comprehensive capabilities in AI engineering applications, including data collection, feature engineering, model training and optimization, and system deployment and operation and maintenance.
[0003] The current AI competition evaluation system generally adopts the traditional manual review model. However, during the implementation process, applicants found that traditional technology has at least the problem of poor consistency in scoring of practical questions in AI competitions. Summary of the Invention
[0004] Based on this, the purpose of this application is to solve at least one of the above-mentioned technical defects, especially the technical defect of poor scoring consistency in the prior art. This application provides a scoring method, device and computer equipment for AI competition practical questions.
[0005] In a first aspect, the present application provides a scoring method for AI competition practical questions, the method comprising:
[0006] In response to a contestant's test operation on a current competition practical question, obtaining a plurality of data sets generated by the contestant in a pre-built competition system corresponding to the current competition practical question; wherein the data sets include a training data set;
[0007] Get the code data entered by the participating users;
[0008] Obtaining a target competition model trained by the participating user using the training data set, wherein the target competition model is a model associated with the code data;
[0009] Using a pre-built scoring system, the code data, multiple data sets, and the target competition model are comprehensively scored to obtain the competition scoring results of the participating users;
[0010] Among them, the competition system runs in an operating environment corresponding to a pre-built competition question image, and the scoring system runs in an operating environment corresponding to a pre-built scoring image.
[0011] In one embodiment, a pre-built scoring system is used to comprehensively score the code data, multiple data sets, and the target competition model to obtain the competition scoring results of the participating users, including:
[0012] Using a pre-built scoring system, determine the code score corresponding to the code data, the data score corresponding to the data set, and the model score corresponding to the target competition model;
[0013] The competition scoring results of the participating users are obtained through the fusion of code scoring, data scoring and model scoring.
[0014] In one embodiment, determining a model score corresponding to a target competition model includes:
[0015] Get the evaluation data set uploaded by the user who organized the competition;
[0016] The target competition model is evaluated using the evaluation data set and multiple preset model evaluation indicators to obtain a model score for the target competition model.
[0017] In one embodiment, the data set further includes a validation data set; and determining a data score corresponding to the data set includes:
[0018] Obtaining data annotations corresponding to data in multiple data sets and determining the accuracy of the data annotations;
[0019] Determine the accuracy of the division of the validation data set and the training data set;
[0020] According to the accuracy of labeling and division, the data score corresponding to the data set is obtained.
[0021] In one embodiment, the method further comprises:
[0022] When generating the current stage score, provide real-time feedback to the participating users on the current stage score; the current stage score can be any one of the scores of data score, code score, and target competition model;
[0023] If the participating user modifies the scoring item corresponding to the current stage scoring, obtain the modified scoring item; wherein the scoring item is any item among the data set, code data, and target competition model;
[0024] Leverage pre-built scoring systems to generate updated current-stage scores based on modified scoring items.
[0025] In one embodiment, the method further comprises:
[0026] Determine the data format required by the competition organizer for the current competition practical questions;
[0027] Obtain multiple scoring components written by competition organizers for the current competition practical questions;
[0028] Build a scoring image based on the required data format and multiple scoring components.
[0029] In one embodiment, the method further comprises:
[0030] Obtain data resources of current competition practical questions uploaded by competition organizers;
[0031] Build a competition mirror based on the data resources and the environmental parameters corresponding to the current competition practical questions.
[0032] In one embodiment, the method further comprises:
[0033] Determine the task application resource queue of participating users;
[0034] Get the server load data;
[0035] The resource scheduling results are obtained based on the task application resource queue and server load data; the resource scheduling results are used to schedule computing resources for participating users.
[0036] In a second aspect, the present application provides a scoring device for AI competition practical questions, the device comprising:
[0037] A data set module is configured to obtain, in response to a contestant's test operation on a current contest practical question, a plurality of data sets generated by the contestant in a pre-built contest system corresponding to the current contest practical question; wherein the data sets include a training data set;
[0038] The code acquisition module is used to obtain the code data entered by the participating users;
[0039] A model acquisition module is used to obtain a target competition model trained by a participating user using a training data set, wherein the target competition model is a model associated with the code data;
[0040] The scoring module uses a pre-built scoring system to comprehensively score the code data, multiple data sets, and the target competition model to obtain the competition scoring results of the participating users;
[0041] Among them, the competition system runs in an operating environment corresponding to a pre-built competition question image, and the scoring system runs in an operating environment corresponding to a pre-built scoring image.
[0042] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0043] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0044] The scoring method, device, and computer equipment for AI competition practical questions provided in this application can obtain multiple data sets generated by participating users, input code data, and the target competition model generated and trained by the user by responding to the test operation of the participating users for the current competition practical questions. In this way, the code data, multiple data sets, and target competition model can be comprehensively scored and processed through a pre-built scoring system to accurately obtain the competition scoring results of the participating users. Compared with traditional technologies, this application automatically scores the AI competition practical questions of participating users through a scoring system, thereby unifying the scoring standards and reducing the utilization of scoring resources for AI competition practical questions, thereby reducing the cost of correctly scoring AI competition practical questions. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0046] Figure 1 A scoring system for AI competition practical questions provided in an embodiment of the present application;
[0047] Figure 2 A flowchart of a scoring method for AI competition practical questions provided in an embodiment of the present application;
[0048] Figure 3 A flowchart of the steps for obtaining the contest scoring results of participating users provided for the implementation of this application;
[0049] Figure 4 A flowchart illustrating the steps of determining a model score corresponding to a target competition model provided in an embodiment of the present application;
[0050] Figure 5 A flowchart illustrating the steps of determining a model score corresponding to a target competition model provided in an embodiment of the present application;
[0051] Figure 6 A competition test and scoring method provided in an embodiment of the present application;
[0052] Figure 7 This is a schematic diagram of the structure of a scoring device for AI competition practical questions provided in this application;
[0053] Figure 8 A schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0055] With the rapid development of artificial intelligence technology and the deepening expansion of its industry applications, the demand for talents with full-chain AI engineering practice capabilities continues to grow in fields such as artificial intelligence engineering, intelligent science and technology, and electronic information engineering. To adapt to this trend, higher education institutions and vocational training institutions are widely using AI practice competitions as a core training method, focusing on assessing participants' comprehensive capabilities in AI engineering applications, including data collection, feature engineering, model training and optimization, and system deployment and operation and maintenance.
[0056] The current AI competition evaluation system generally adopts the traditional manual review model, which has the following technical defects when meeting the evaluation needs of new artificial intelligence technologies:
[0057] Lack of review efficiency and real-time feedback mechanism: Manual review requires checking the submitted code and models one by one. Especially when there are many participants, the entire review process may take a lot of time, resulting in the inability to provide feedback during the exam, affecting the participants' opportunities to improve their solutions and restricting the participants' iterative optimization of the model.
[0058] Imperfect quantitative evaluation system: Model accuracy assessment usually requires using predefined test sets to measure the model's generalization ability and stability. Manual review cannot execute these automated processes in a timely and efficient manner, making it difficult to provide rapid feedback on various model indicators. In addition, manual review cannot provide a detailed assessment of the model's performance, efficiency, and resource consumption.
[0059] Scoring consistency risk: Manual reviews have certain limitations when analyzing technical indicators such as code efficiency, scalability, and performance optimization. They cannot quickly and accurately quantify these indicators. Especially for complex models running on large-scale datasets, the reviewer's personal experience and preferences may affect the scoring, resulting in the same solution receiving different scores from different reviewers, making it difficult to ensure fairness and consistency.
[0060] Insufficient system scalability: Traditional practical competitions are bound to the competition questions, and it is impossible to quickly integrate the competition questions and scoring plug-ins according to the customized requirements of the competition questions, and quickly carry out the competition.
[0061] The above-mentioned technical defects have seriously restricted the quality improvement of practical teaching of artificial intelligence and the efficiency of talent training. Therefore, it is urgent to build a scoring method for practical questions of AI competitions with automated and intelligent features to solve the problem of lack of evaluation consistency in the traditional review model.
[0062] In an exemplary embodiment, Figure 1 The present invention provides a scoring system for AI competition practical questions, such as Figure 1 As shown, the system can be used to implement the steps of the scoring method for AI competition practical questions. The system includes: an examination management unit, a scheduling management unit, a candidate data storage unit, a candidate answering unit, an automatic scoring module, and a candidate score statistics module;
[0063] The exam management unit includes: event management module, question management module, scoring plug-in management module, candidate management module, and question scoring verification module;
[0064] The scheduling management unit includes: candidate workspace scheduling module, scoring plug-in scheduling module, cluster resource management module, and system component status monitoring module;
[0065] The candidate data storage unit includes: candidate data initialization module and candidate data backup module;
[0066] The candidate answering unit includes: examination time management module, examination workspace entrance, examination recording monitoring module, and examination results display.
[0067] In an exemplary embodiment, Figure 2 A flowchart of a scoring method for AI competition practical questions provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, a scoring method for practical questions of an AI competition is provided. The method is illustrated by applying a server. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smart phones, and tablet computers. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. In this embodiment, the method includes the following S201 to S204. Among them: the competition system runs in an operating environment corresponding to a pre-built competition question image, and the scoring system runs in an operating environment corresponding to a pre-built scoring image.
[0068] S201. In response to a contestant's test operation on a current competition practical question, obtain multiple data sets generated by the contestant in a pre-built competition system corresponding to the current competition practical question; wherein the data sets include a training data set.
[0069] The current practical competition questions may refer to the AI competition practical questions required for the current exam. The participating users may refer to the examinees. The participating users may also refer to the users who have logged into the exam system on the terminal. The competition system may be the system used to conduct the exam. The competition system runs in an operating environment corresponding to a pre-built competition question image. In this way, it can ensure that the current practical competition questions run smoothly in the preset operating environment. The data set may refer to a data set, such as a training set, a validation set, etc. The training data set refers to the training set of the model, that is, the data set used to train the model.
[0070] For example, when a participating user (candidate) performs operations related to the current practical question in the competition system (such as running code, training a model, submitting results, etc.), the system can automatically trigger the data acquisition mechanism.
[0071] Optionally, the competition system runs based on a pre-built competition image, ensuring that all user operations are performed in a standardized environment to avoid interference from environmental differences. When working on competition problems in this image environment, participating users may dynamically generate multiple datasets through operations such as data preprocessing, feature engineering, and dataset partitioning. For example, the user's core dataset (training set) for model training, validation sets, test sets, or pre-processed derivative datasets.
[0072] In principle, the standardized environment and execution server can automatically capture the data uploaded by participating users, ensuring that the data (especially the training set) generated by participating users when solving AI practical problems is effectively obtained, providing a basis for subsequent scoring and supervision.
[0073] S202: Obtain code data input by the participating user.
[0074] The code data refers to the code used by the participating users to build the target participating model.
[0075] For example, when a contestant completes the code writing for an AI competition practical question and submits or executes an operation, the system can automatically obtain the complete code data entered by the contestant. For example, when a contestant performs operations related to model building in the competition system, the following operations can be used to trigger the server to obtain the contestant's code data: actively submitting the code, such as clicking the pre-set "Submit Model" button on the terminal; running the code in the competition environment, such as executing the training script, starting the model application, etc. The code data of the contestant is automatically saved based on a pre-set periodicity, which can also prevent the accidental loss of the code.
[0076] S203. Obtain a target competition model trained by the participating user using the training data set, wherein the target competition model is a model associated with the code data.
[0077] The target competition model may refer to a competition model generated by a participating user through the code submitted by the participating user.
[0078] For example, after a participating user completes model training based on the training data set, the server can obtain the target competition model generated by the participating user (i.e., the final competition model bound to the user's code). For example, the server can monitor the progress of model training and, upon completion of the training task, automatically obtain the trained target competition model.
[0079] S204: Using a pre-built scoring system, comprehensively score the code data, multiple data sets, and the target competition model to obtain the competition scoring results of the participating users.
[0080] The scoring system runs in an operating environment based on a pre-built scoring image. The scoring image can be a pre-built image used to score practical AI competition questions, and can be automatically built by the system. The competition scoring results can represent the AI competition score of the participating user, which can be a score based on at least three indicators: code data, multiple data sets, and the target competition model.
[0081] For example, the server can invoke a pre-built scoring system to perform a multi-dimensional comprehensive scoring of the code, data set, and model in a standardized environment to generate the final competition score for the participating users. For example, the pre-built scoring system can be run in a pre-built independent runtime environment (e.g., a Docker image).
[0082] Optionally, the system can verify that submitted code and models comply with competition specifications, preventing submissions that violate regulations, such as code security and model structure. The scoring engine can include various model evaluation metrics, such as accuracy, recall, F1-score, AUC, runtime efficiency, model size, resource consumption (memory and computing resources), robustness, and scalability, to comprehensively assess the candidate's model. The scoring engine can also assess the accuracy and rationality of data annotation and the rationality of dataset partitioning (e.g., the allocation of training and validation sets).
[0083] In this embodiment, by responding to the user's test operation on the current competition practical question, it is possible to obtain multiple data sets generated by the user, the input code data, and the target competition model generated and trained by the user. In this way, the code data, multiple data sets, and the target competition model can be comprehensively scored through a pre-built scoring system to accurately obtain the user's competition score. Compared with traditional technologies, this embodiment automatically and correctly scores the user's AI competition practical questions through the scoring system, thereby unifying the scoring standards and reducing the use of scoring resources for AI competition practical questions, thereby reducing the cost of correctly scoring AI competition practical questions.
[0084] In an exemplary embodiment, Figure 3 A flowchart of the steps for obtaining the competition score results of participating users provided for the implementation of this application is as follows: Figure 3 As shown, in Figure 1 Based on this, the steps of the scoring method for AI competition practical questions are exemplified. In step S204, a pre-built scoring system is used to perform comprehensive scoring processing on the code data, multiple data sets, and the target competition model to obtain the competition scoring results of the participating users, including S301 to S302, in which:
[0085] S301. Determine the code score corresponding to the code data, the data score corresponding to the data set, and the model score corresponding to the target competition model using a pre-built scoring system;
[0086] S302: Obtain the competition scoring results of the participating users based on the fusion processing of code scoring, data scoring and model scoring.
[0087] Code scoring refers to the scoring of code data, data scoring refers to the scoring of datasets, and model scoring refers to the scoring of the target competition model.
[0088] For example, the server can independently evaluate the code, dataset, and target competition model, and obtain three sub-item scores. The server can call a pre-built scoring system to score the code data, dataset, and target competition model separately in a standardized environment, obtaining a code score, a data score, and a model score. Furthermore, the server can use a dynamic fusion mechanism to integrate the code score, data score, and model score, ultimately obtaining the competition score results of the participating users. In this way, through sub-item evaluation and a dynamic fusion mechanism, the multi-dimensional performance of the code, data, and model is converted into a quantitative score, which can output fair and interpretable competition score results.
[0089] In this embodiment, through steps S301 to S302, the scores of at least three indicators are integrated, which can reflect the differences of various indicators and comprehensively obtain the competition scoring results of the participating users, thereby ensuring the consistency of the scoring standards of the competition scoring results of each participating user, reducing the utilization of scoring resources, and reducing the scoring cost of AI competition practical questions.
[0090] In an exemplary embodiment, Figure 4 A flow chart of the steps for determining the model score corresponding to the target competition model provided in the embodiment of the present application is as follows: Figure 4 As shown, in Figure 2 and Figure 3 Based on this, the steps of the scoring method for AI competition practical questions are exemplified. In step S301, the model score corresponding to the target competition model is determined, including S401 to S402, in which:
[0091] S401, obtaining a set of evaluation data uploaded by the user organizing the competition;
[0092] S402: Evaluate the target competition model using the evaluation data set and multiple preset model evaluation indicators to obtain a model score for the target competition model.
[0093] The term "competition organizer" can refer to the user who initiates the AI competition practice test or the user who logs into the exam system to solve the test. The evaluation data set can refer to the test data used to evaluate the target competition model. Model evaluation metrics can refer to pre-set evaluation indicators for the target competition model, such as accuracy, recall, F1-score, AUC, operational efficiency, model size, resource consumption (memory, computing resources), robustness, and scalability, to comprehensively evaluate the candidate's model.
[0094] For example, the server can use the target competition model to infer and predict the data in the evaluation data set, obtaining evaluation values corresponding to multiple model evaluation indicators. Furthermore, based on the evaluation values corresponding to the multiple model evaluation indicators, the server can obtain a model score for the target competition model. The server can store the evaluation data by pre-acquiring it, and when applying it, it can directly extract the evaluation data uploaded by the user organizing the competition to evaluate the target competition model.
[0095] Optionally, users who initiate AI competition practice challenges are responsible for designing the challenge, setting the rules, and uploading the evaluation data. They have administrator privileges in the exam system and can configure the challenge environment and scoring logic. The server can quantify model performance and generate a model score based on the evaluation data and pre-set multi-dimensional indicators. For example, the model can be scored by checking whether the model's predictions conflict with the code logic and whether the code data can be used to generate the results.
[0096] In this embodiment, through steps S401 to S402, using evaluation data and executing model evaluation, the model performance is comprehensively evaluated through standardized test data and multi-dimensional indicators. The model can be comprehensively quantitatively evaluated, an accurate model score can be obtained, and the consistency of the model score can be guaranteed, thereby ensuring the consistency of the score of the AI competition practical questions.
[0097] In an exemplary embodiment, Figure 5 A flow chart of the steps for determining the model score corresponding to the target competition model provided in the embodiment of the present application is as follows: Figure 5 As shown, in Figure 2 and Figure 3 Based on this, the steps of the scoring method for AI competition practical questions are exemplified, wherein the data set also includes a verification data set; in step S201, determining the data score corresponding to the data set includes steps S501 to S503, wherein:
[0098] S501: Obtain data annotations corresponding to data in multiple data sets, and determine the accuracy of the data annotations.
[0099] S502: Determine the accuracy of the division of the verification data set and the training data set.
[0100] S503: Obtain a data score corresponding to the data set based on the accuracy of the labeling and the accuracy of the division.
[0101] Data labeling refers to the labeling of specific data within a dataset, for example, code labels. Labeling accuracy refers to the accuracy of the labeling information, for example, by determining whether the labeling contradicts the actual prediction results. Partitioning accuracy refers to the accuracy of the dataset partitioning, for example, by determining whether validation and training data are mixed.
[0102] For example, the server can generate a data score by evaluating the quality of data annotation and the rationality of the data set division, quantifying the compliance and validity of the data set. The server can extract the annotation information (such as classification labels, text annotations, etc.) of all samples in the data set and, using pre-set scoring logic, verify whether the annotations are correct, thereby determining the accuracy of the data annotations. For example, pre-defined rules or model prediction results can be used to reversely infer whether the annotations are inconsistent. For example, a sample may be consistently predicted as category A by multiple models but labeled as category B. The server can also verify whether the data set division meets the requirements of the competition question and determine the percentage of data that meets the requirements, thereby further determining the accuracy of the division.
[0103] The server can calculate the data score corresponding to the data set based on the accuracy of labeling and the accuracy of division through pre-set weight distribution.
[0104] In this embodiment, by objectively checking the quality of annotation and the rationality of data division, the accuracy of data scoring can be ensured, and the consistency of data scoring standards can be ensured, thereby ensuring the consistency of scoring for AI competition practical questions.
[0105] In an exemplary embodiment, the method further includes:
[0106] When generating the current stage score, provide real-time feedback to the participating users on the current stage score; the current stage score can be any one of the scores of data score, code score, and target competition model;
[0107] If the participating user modifies the scoring item corresponding to the current stage scoring, obtain the modified scoring item; wherein the scoring item is any item among the data set, code data, and target competition model;
[0108] Leverage pre-built scoring systems to generate updated current-stage scores based on modified scoring items.
[0109] The current stage score can refer to any one of the three scores: data score, code score, and target competition model. The sub-project can refer to any one of the practical question items in the data set, code data, and target competition model. It is understood that the data score, code score, and target competition model can all be processed and updated using the methods provided in the embodiments of this application.
[0110] For example, for each practical question, after generating the current stage score of the practical question, the current stage score of the practical question can be fed back to the participating user in real time. The participating user can modify the content of the practical question for the current stage score. For example, if the participating user believes that the score given is too low, the content of the practical question can be modified based on the score result. For example, if the participating user is not satisfied with the score of the generated data set, the participating user can modify the data set and regenerate it.
[0111] For the modified scoring practice questions, a pre-built scoring system can be used to regenerate new scores corresponding to the modified scoring practice questions as updated scores for the current stage.
[0112] Optionally, when the initial scoring of any practical project (data set, code data, target competition model) is completed, the system will immediately push the current stage score to the participating users. For example, by highlighting the sub-item scores in the visual user's operation panel, such as data score: 85 / 100; detailed information on the reasons for the deduction can also be provided through a detailed report, such as "15 points deducted from the data score: it was detected that the training set contains 3% of the validation set samples." Participating users can make multiple modifications to any scored practical project within the test time limit, such as adjusting the data partitioning strategy, optimizing the code logic, retraining the model, etc. When the participating user submits the modified project, the system can automatically trigger the local re-evaluation step for the project, and there is no need to re-evaluate the unmodified items.
[0113] In this embodiment, it can ensure that contestants can receive the evaluation results within a few minutes after generating a data set or a model, achieving the effect of real-time scoring; and after the contestants complete their training, the system can quickly process and return the scores and related feedback, helping them to quickly understand the results and make adjustments, thereby improving the uniformity and accuracy of scoring for practical questions in AI competitions.
[0114] In an exemplary embodiment, the method further comprises:
[0115] Determine the data format required by the competition organizer for the current competition practical questions;
[0116] Obtain multiple scoring components written by competition organizers for the current competition practical questions;
[0117] Build a scoring image based on the required data format and multiple scoring components.
[0118] The required data format may refer to the data format required by AI competition practical questions. The scoring component may refer to the component used to score AI competition practical questions. For example, the scoring component may integrate components such as pytest, unittest, and PyLint.
[0119] For example, the server can obtain the data format requirements of the current competition practical problem uploaded by the competition organizer, and can also obtain multiple scoring components written by the competition organizer. The scoring components can be integrated to obtain the scoring components of the competition system. Finally, the server can construct a scoring image of the current AI competition practical problem based on the data format requirements and multiple scoring components.
[0120] Optionally, competition organizers can pre-set structured standards for the data that participants must submit to ensure uniform data processing and scoring. The server can then access a pre-configured modular evaluation tool set for the current AI competition practical problem, which can be used to automate the scoring process. Ultimately, the server can generate a standardized container image containing a complete evaluation environment based on the required data format and scoring components.
[0121] Schematically, the data format of the intelligent scoring image is clarified. For example, the candidate's working directory, test data set, and auxiliary parameters can be required to be input; the details of each scoring unit are output to help identify the candidate's strengths in specific areas; the ease of use of the intelligent scoring image allows the intelligent scoring image to integrate various components such as pytest, unittest, and PyLint according to different exams; in order to ensure the transparency and traceability of the scoring process, the intelligent scoring image can include logging and monitoring tools, and back up the candidate's output data in each scoring process to track every step and decision in each scoring process.
[0122] In this embodiment, by determining the data format requirements and multiple scoring components compiled by the competition organizer for the current competition practical problem, a scoring image can be constructed based on the data format requirements and multiple scoring components. In this way, by pre-specifying the structured standards for contestant submitted data, the competition organizer can ensure the uniformity of data processing and scoring. Furthermore, by clearly defining the data format, integrating modular scoring tools, and building standardized container images, it is possible to achieve automation and uniformity in the scoring of AI practical problem competitions.
[0123] In an exemplary embodiment, the method further comprises:
[0124] Obtain data resources of current competition practical questions uploaded by competition organizers;
[0125] Build a competition mirror based on the data resources and the environmental parameters corresponding to the current competition practical questions.
[0126] For example, competition organizers can upload core resources such as benchmark data sets (such as training data, test data), reference code templates (such as data loading examples), etc. required for practical AI competition questions to the server to obtain a complete competition question data package.
[0127] Based on the competition data package and the environmental parameters pre-set by the competition organizer, the server can use containerization technology to package data resources and dependent environments into standardized competition images to ensure that the development environment of participating users is completely consistent with the evaluation environment.
[0128] Optionally, a standardized process can be developed to upload resources such as competition code packages, training sets, and evaluation sets, allowing for rapid integration of various AI-related competitions. Furthermore, the corresponding Docker images can be dynamically built based on the required dependency environments to ensure that the competitions can be successfully executed in the pre-set operating environment.
[0129] In this embodiment, by pre-packaging the uploaded data resources and environment, the local configuration differences of participating users can be eliminated, the uniformity of the competition tasks can be guaranteed, and thus the consistent scoring of the AI competition practical questions can be guaranteed.
[0130] In an exemplary embodiment, the method further comprises:
[0131] Determine the task application resource queue of participating users;
[0132] Get the server load data;
[0133] The resource scheduling results are obtained based on the task application resource queue and server load data; the resource scheduling results are used to schedule computing resources for participating users.
[0134] The task application resource queue may refer to a task queue of application resources, the load data may refer to the current load of the server, and the resource scheduling result may refer to the result used to schedule resources.
[0135] For example, the server can dynamically obtain the task resource requirements submitted by participating users, such as GPU memory, number of CPU cores, storage space, and other resources, build a priority queue, and clarify the computational scale and urgency of each task. The server can also collect real-time load indicators of the server cluster, such as node computing power utilization, remaining memory, network bandwidth occupancy, and other loads, to obtain server load data. The server can combine the task requirements with the server load status to generate a resource allocation plan and allocate computing nodes or container instances to participating users as needed.
[0136] Optionally, the system scheduling mechanism can prioritize real-time feedback tasks through task queue management, effectively allocating computing resources. Tasks can be assigned to different nodes for parallel execution, ensuring full resource utilization. The scheduling system can monitor server load in real time and dynamically adjust task allocation based on load to avoid overload issues.
[0137] In this embodiment, by applying task resource queues and acquiring server load data, the resource scheduling results for users can be obtained. This allows for real-time task queue management and dynamic adaptation to server load, optimizing computing resource utilization, ensuring system stability and user experience fairness during multi-tasking, and avoiding resource overload issues. This improves system stability and ensures consistent scoring of practical AI competition questions.
[0138] In some specific exemplary embodiments, Figure 6 A competition test and scoring method provided in the embodiment of the present application is as follows: Figure 6 As shown, based on the various embodiments of the scoring method for AI competition practical questions, a detailed exemplary description of the scoring method for AI competition practical questions can be made. The method can specifically include:
[0139] Candidate training set, candidate training set, candidate code, candidate model.
[0140] Automatic scoring module: Dataset scoring: annotation quality assessment and dataset partition assessment; Code scoring: code security check, code logic check, code quality check; Model scoring: specification check, accuracy assessment using test sets, stability assessment, performance assessment, and resource consumption detection.
[0141] Candidate score version management: candidate score statistics management, candidate score version management;
[0142] Candidate answering unit: test result display.
[0143] Illustratively, each method step of each of the above embodiments or a combination of multiple embodiments may be executed through the configuration of each of the above modules or units.
[0144] In some specific exemplary embodiments, the scoring method for AI competition practical questions provided in the embodiments of the present application can timely evaluate the test-taker's output data and quickly generate feedback, thereby encouraging test-takers to continuously explore and innovate. By modularizing the scoring plug-in design, integrating the standardized competition question process, and intelligently adapting the competition question environment, the following technical effects can be achieved:
[0145] By using technical methods to quickly generate feedback on the content produced by the examinees (such as code, models, etc.) during the examination, we not only provide feedback on scores, but also provide a mechanism for multi-faceted feedback such as error prompts, performance suggestions, and resource optimization. This enhances the interactivity of the competition and encourages participants to continue exploring and innovating.
[0146] Through the modular design of test questions and scoring plug-ins, users can quickly integrate test questions and expand scoring plug-ins according to different competition requirements, and support automated evaluation of multiple programming languages, frameworks and technologies.
[0147] Through intelligent scheduling and management solutions, resource allocation is optimized and environmental isolation is achieved, ensuring that the problem workspace and scoring plug-in operate efficiently in their own independent environments. This design improves the stability and compatibility of the problem workspace and scoring plug-in, enabling the system to effectively support the smooth running of large-scale competitions.
[0148] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0149] The following describes the scoring device for AI competition practical questions provided in the embodiments of the present application. The scoring device for AI competition practical questions and the scoring method for AI competition practical questions described above have the same inventive concept, and the implementation solution for solving the problem provided by the device is similar to the implementation solution recorded in the above method. Therefore, the specific limitations in the embodiments of one or more AI competition practical questions scoring devices provided below can refer to the limitations of the AI competition practical questions scoring method described above. The scoring device for AI competition practical questions described below and the scoring method for AI competition practical questions described above can be referenced to each other and will not be repeated here.
[0150] In an exemplary embodiment, Figure 7 This is a schematic diagram of the structure of a scoring device for AI competition practical questions provided in this application, such as Figure 7 As shown, the scoring device 70 for the AI competition practical questions includes: a data set module 710, a code acquisition module 720, a model acquisition module 730, and a scoring module 740, wherein:
[0151] The data set module 710 is configured to obtain, in response to a contestant's test operation on a current contest practical question, multiple data sets generated by the contestant in a pre-built contest system corresponding to the current contest practical question; wherein the data sets include a training data set;
[0152] The code acquisition module 720 is used to obtain the code data input by the participating user;
[0153] The model acquisition module 730 is used to obtain the target competition model trained by the participating user using the training data set, wherein the target competition model is a model associated with the code data;
[0154] Scoring module 740 uses a pre-built scoring system to perform comprehensive scoring on the code data, multiple data sets, and the target competition model to obtain the competition scoring results of the participating users;
[0155] Among them, the competition system runs in an operating environment corresponding to a pre-built competition question image, and the scoring system runs in an operating environment corresponding to a pre-built scoring image.
[0156] In an exemplary embodiment, the scoring module is used to use a pre-built scoring system to determine the code score corresponding to the code data, the data score corresponding to the data set, and the model score corresponding to the target competition model; and obtain the competition scoring results of the participating users based on the fusion processing of the code score, data score and model score.
[0157] In an exemplary embodiment, the scoring module is used to obtain an evaluation data set uploaded by users organizing a competition; use the evaluation data set and multiple preset model evaluation indicators to evaluate the target competition model and obtain a model score for the target competition model.
[0158] In an exemplary embodiment, the dataset further includes a validation dataset. The scoring module is configured to obtain data annotations corresponding to the data contained in the multiple datasets, determine the accuracy of the data annotations, determine the accuracy of the partitioning between the validation dataset and the training dataset, and obtain a data score corresponding to the dataset based on the annotation accuracy and the partitioning accuracy.
[0159] In an exemplary embodiment, the scoring module is also used to provide real-time feedback of the current stage score to the participating users when generating the current stage score; wherein the current stage score is any one of the data score, code score, and target competition model score; when the participating users modify the scoring items corresponding to the current stage score, obtain the modified scoring items; wherein the scoring items are any one of the data set, code data, and target competition model; and use the pre-built scoring system to generate an updated current stage score based on the modified scoring items.
[0160] In an exemplary embodiment, the device also includes a scoring mirror module, which is used to determine the data requirement format of the competition organizing user for the current competition practical questions; obtain multiple scoring components written by the competition organizing user for the current competition practical questions; and construct a scoring mirror based on the data requirement format and multiple scoring components.
[0161] In an exemplary embodiment, the device also includes a competition question mirror module, which is used to obtain data resources of the current competition practical questions uploaded by the competition organization user; and construct a competition question mirror based on the data resources and environmental parameters corresponding to the current competition practical questions.
[0162] In an exemplary embodiment, the device also includes a resource scheduling module, which is used to determine the task application resource queue of the participating user; obtain the load data of the server; obtain the resource scheduling result based on the task application resource queue and the load data of the server; and the resource scheduling result is used to schedule computing resources for the participating user.
[0163] In an exemplary embodiment, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by one or more processors, the one or more processors execute the steps of any method in the above embodiments.
[0164] In an exemplary embodiment, the present application further provides a computer device having a computer program stored therein, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of any of the methods in the above embodiments.
[0165] In an exemplary embodiment, the present application further provides a computer program product, including a computer program, which implements the steps of any one of the methods in the above embodiments when executed by a processor.
[0166] Schematically, as Figure 8 As shown, Figure 8This is a schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. The computer device 800 can be provided as a server. Figure 8 Computer device 800 includes a processing component 802, which further includes one or more processors, and a memory resource represented by memory 801 for storing instructions executable by processing component 802, such as an application. The application stored in memory 801 may include one or more modules, each corresponding to a set of instructions. In addition, processing component 802 is configured to execute the instructions to perform the text recognition method of any of the above embodiments.
[0167] The computer device 800 may further include a power supply component 803 configured to perform power management of the computer device 800, a wired or wireless network interface 804 configured to connect the computer device 800 to a network, and an input / output (I / O) interface 805. The computer device 800 may operate based on an operating system stored in the memory 801, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or the like.
[0168] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0169] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0170] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0171] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referenced to each other.
[0172] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A scoring method for AI competition practical questions, characterized by: The method comprises: In response to a contestant's test operation on a current competition practical question, obtaining a plurality of data sets generated by the contestant in a pre-built competition system corresponding to the current competition practical question; wherein the data sets include a training data set; Obtaining code data input by the participating user; Obtaining a target competition model trained by the participating user using the training data set, wherein the target competition model is a model associated with the code data; Using a pre-built scoring system, comprehensively scoring the code data, the plurality of data sets, and the target competition model to obtain a competition scoring result of the participating user; The competition system runs in an operating environment corresponding to a pre-built competition question image, and the scoring system runs in an operating environment corresponding to a pre-built scoring image.
2. The method according to claim 1, characterized in that The pre-built scoring system is used to perform comprehensive scoring processing on the code data, the plurality of data sets, and the target competition model to obtain the competition scoring results of the participating users, including: Using a pre-built scoring system, determine a code score corresponding to the code data, a data score corresponding to the data set, and a model score corresponding to the target competition model; The competition scoring result of the participating user is obtained by fusion processing of the code scoring, the data scoring and the model scoring.
3. The method according to claim 2, characterized in that Determining a model score corresponding to the target competition model includes: Get the evaluation data set uploaded by the user who organized the competition; The target competition model is evaluated using the evaluation data set and a plurality of preset model evaluation indicators to obtain a model score of the target competition model.
4. The method according to claim 2, characterized in that The data set also includes a validation data set; and determining a data score corresponding to the data set includes: Obtaining data annotations corresponding to the data in the plurality of data sets, and determining the accuracy of the data annotations; Determining the accuracy of the division of the verification data set and the training data set; A data score corresponding to the data set is obtained according to the labeling accuracy and the division accuracy.
5. The method according to claim 2, characterized in that The method further comprises: When generating a current stage score, providing real-time feedback of the current stage score to the participating user; wherein the current stage score is any one of the data score, the code score, and the target competition model score; When the contestant modifies the scoring item corresponding to the current stage scoring, obtaining the modified scoring item; wherein the scoring item is any one of the data set, the code data, and the target competition model; Leverage pre-built scoring systems to generate updated current-stage scores based on modified scoring items.
6. The method according to claim 1, characterized in that The method further comprises: Determine the data format required by the competition organizer for the current competition practical question; Obtaining multiple scoring components compiled by the competition organizing user for the current competition practical question; The scoring image is constructed according to the required data format and the plurality of scoring components.
7. The method according to claim 1, characterized in that The method further comprises: Obtain data resources of current competition practical questions uploaded by competition organizers; The competition question mirror is constructed based on the data resources and the environmental parameters corresponding to the current competition practical question.
8. The method according to claim 1, characterized in that The method further comprises: Determine the task application resource queue of the participating user; Get the server load data; A resource scheduling result is obtained based on the task application resource queue and the load data of the server; the resource scheduling result is used to schedule computing resources for the participating users.
9. A scoring device for AI competition practical questions, characterized by: The device comprises: A data set module is configured to obtain, in response to a contestant's test operation on a current competition practical question, a plurality of data sets generated by the contestant in a pre-built competition system corresponding to the current competition practical question; wherein the data sets include a training data set; A code acquisition module, used to acquire the code data input by the participating user; a model acquisition module, configured to acquire a target competition model trained by the participating user using the training data set, wherein the target competition model is a model associated with the code data; A scoring module, using a pre-built scoring system, performs a comprehensive scoring process on the code data, the plurality of data sets, and the target competition model to obtain a competition scoring result of the participating user; The competition system runs in an operating environment corresponding to a pre-built competition question image, and the scoring system runs in an operating environment corresponding to a pre-built scoring image.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Intelligent student competition system based on number competition big model and virtual engine
CN121766847A
Intelligent management and control method and system for skill competition of industrial workers
CN121836507A