A method for predicting the number of graduates for all higher education levels
By combining LSTM and GRU models with quantization thresholds and individual profiles, the accuracy problem of predicting the number of graduates across all levels of higher education was solved, achieving high-precision prediction of the number of graduates and real-time reflection of adjustments to training programs, thus reducing systematic bias.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-27
- Publication Date
- 2026-07-24
AI Technical Summary
Traditional methods for predicting the number of graduates have failed to accurately match the differences across all levels of higher education and have not made full use of the rigid requirements in the curriculum, resulting in a disconnect between predictions and the actual training process.
By employing an LSTM network and a GRU branch model, combined with quantized thresholds and hierarchical professional-specific data, and correcting through desensitized individual profiles, the initial graduation probability is obtained. The model is then updated in real time to reflect adjustments to the training program, providing high-precision predictions of the number of graduates.
It significantly reduces the risk of systematic bias, enables the early determination of graduation size without relying on ex-post correction, and improves the accuracy and consistency of prediction.
Smart Images

Figure CN122453137A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method for predicting the number of graduates across all levels of higher education. Background Technology
[0002] The full range of higher education levels refers to the four core groups covering my country's higher education system: associate degree (skills-oriented, 2-3 years of study, with a focus on practical training and professional qualification certification), bachelor's degree (comprehensive quality-oriented, 4-5 years of study, with a balance between course learning and practical innovation), master's degree (academic / application-oriented, 2-3 years of study, with academic and professional programs respectively focusing on research capabilities and professional skills), and doctoral degree (cutting-edge research-oriented, flexible 3-5 years of study, with a focus on research innovation and output).
[0003] Significant differences exist between different levels of higher education in terms of duration of study, training objectives, and graduation requirements. Even within the same level, different majors have different training programs that directly impact student mobility. Traditional methods for predicting the number of graduates have not accurately matched the four categories of higher education levels and have not fully utilized the rigid requirements in the training programs (such as curriculum, assessment standards, and graduation conditions), leading to a disconnect between predictions and the actual training process.
[0004] The information disclosed in this background section is intended only to enhance the understanding of the overall background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a method for predicting the number of graduates across all levels of higher education. For scenarios where the number of students enrolled is already determined, this method comprehensively covers four core levels of higher education: associate degree, bachelor's degree, master's degree, and doctoral degree. It integrates key information from professional training programs and provides a differentiated and highly accurate method for predicting the number of graduates that aligns with the actual training logic of universities.
[0006] To achieve the above objectives, this invention provides a method for predicting the number of graduates across all levels of higher education, comprising the following steps: acquiring initial student data across all levels, including associate's, bachelor's, master's, and doctoral levels; inputting the initial student data, quantification thresholds, and level-specific data into a trained prediction model according to the level-specific professional dimension, wherein the prediction model outputs an initial graduation probability corresponding to the level-specific professional dimension, and the quantification threshold is used to indicate the pass standard in the training program; correcting the initial graduation probability according to the level-specific professional dimension using anonymized individual profile data to obtain the predicted number of graduates; and displaying the predicted number of graduates.
[0007] In one embodiment of the present invention, the prediction model is obtained through the following training process: an LSTM network is trained using historical academic year data to obtain a first prediction network for predicting associate degree, bachelor's degree, and master's degree levels; a GRU branch is added to the first prediction network to obtain a dedicated network suitable for doctoral level; wherein, each academic year, the prediction model to be trained is iteratively optimized using the actual number of graduates and the actual implementation of the training program; wherein, the parameters of the prediction model are updated using gradient descent method with the objective function of minimizing the mean square error between the predicted value and the actual value.
[0008] In one embodiment of the present invention, the method includes: extracting quantitative indicators of the training program for the corresponding professional dimension, the indicators including curriculum system, assessment standards, graduation requirements and schedule; and converting the indicators into quantitative thresholds for the corresponding professional dimension.
[0009] In one embodiment of the present invention, the method further includes: in response to receiving adjustment information of the training program, obtaining the adjustment content, automatically identifying the level, major, and stage, and synchronously updating the corresponding quantification.
[0010] In one embodiment of the present invention, the hierarchical professional-specific data refers to dynamic performance data collected separately according to the training level; wherein, for the junior college level, it is the practical training assessment score; for the undergraduate level, it is the pass rate of core courses; for the master's academic level, it is the research foundation score; for the master's professional level, it is the completion rate of practical projects; and for the doctoral level, it is the number of research achievements.
[0011] In one embodiment of the present invention, different dynamic adaptation cycles are set for associate degree, bachelor's degree, master's degree and doctoral degree levels, and performance data is synchronized in real time. Among them, the associate degree level takes the training cycle as the smallest unit, and the assessment results are uploaded after each training cycle is completed. The bachelor's and master's degree levels synchronize the final grades on a semester basis. The doctoral level uses the research nodes of the training program as trigger signals. After the date or result status of the opening, mid-term, pre-defense and submission nodes changes, the number of research results and the node achievement marker factor are updated incrementally.
[0012] In one embodiment of the present invention, the method for extracting the initial student data includes: for associate degree and bachelor's degree programs, extracting the number of students enrolled in the latest semester; if it is the initial enrollment period, extracting the actual number of students enrolled, and supplementing the number of students who registered but did not enroll and the number of students who withdrew in the early stages; for master's degree programs, extracting the number of students enrolled after entering the core training stage, i.e., after the course has ended; if it is the initial enrollment period, supplementing the number of students eliminated during the qualification review, and adjusting the data separately for academic and professional programs; for doctoral degree programs, extracting the number of students enrolled who have passed the thesis proposal; if it is the initial enrollment period, supplementing the number of students eliminated during the qualification assessment and the number of students eliminated due to adjustments in research direction.
[0013] In one embodiment of the present invention, the display of the graduation number prediction results includes: displaying the target year's graduation number, confidence interval, prediction accuracy, and delayed graduation and early graduation at the hierarchical and professional level; and outputting a summary of the number of graduates across all levels and majors, their proportion, and year-on-year change trends.
[0014] In one embodiment of the present invention, displaying the graduation number prediction results includes: listing the core factors and factor weights used in this prediction for each level of major; calculating the contribution value of each factor change to the increase or decrease in the number of graduates; and displaying the graduation number prediction results corresponding to the core factors and factor weights.
[0015] In one embodiment of the present invention, the display of the graduation number prediction results includes: displaying student flow curves at the hierarchical and professional dimensions, wherein the horizontal axis of the curve represents the academic year and the vertical axis represents the number of students; dynamic pie charts and bar charts showing the influence ratio of each core factor and the training program, with the dynamic indicator automatically switching according to the prediction year.
[0016] Compared with existing technologies, the present invention provides a method for predicting the number of graduates across all levels of higher education. By acquiring initial student data across all levels, combining input quantification thresholds with specialized data specific to each level, the model outputs an initial graduation probability. This probability is then further refined using desensitized individual profiles, and the results are displayed. This progressively refined and feedback-driven sequence incorporates enrollment scale, rigid regulations, and individual performance into the same prediction channel. This allows the model to integrate hard institutional constraints with soft data features during the inference stage, thereby determining the graduation scale signal in advance without relying on post-hoc manual correction. This significantly reduces the risk of systemic bias caused by the disconnect between regulations and data in traditional methods. Attached Figure Description
[0017] Figure 1 This is a flowchart of a prediction method according to an embodiment of the present invention;
[0018] Figure 2 This is a schematic diagram of a prediction method according to an embodiment of the present invention. Detailed Implementation
[0019] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings, but it should be understood that the scope of protection of the present invention is not limited to the specific embodiments.
[0020] Unless otherwise expressly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "comprises" shall be understood to include the stated elements or components without excluding other elements or other components.
[0021] like Figure 1As shown, a method for predicting the number of graduates across all levels of higher education, according to a preferred embodiment of the present invention, includes the following steps:
[0022] Step 101: Obtain initial student data for all levels.
[0023] The full range of levels includes: associate degree, bachelor's degree, master's degree and doctoral degree.
[0024] Step 102: According to the hierarchical professional dimension, input the initial student data, quantification threshold and hierarchical professional-specific data into the already trained prediction model.
[0025] The prediction model outputs an initial graduation probability corresponding to the hierarchical professional dimension, where the quantification threshold is used to indicate the pass standard in the training program.
[0026] Step 103: According to the hierarchical professional dimension, the initial graduation probability is corrected using desensitized individual profile data to obtain the predicted number of graduates.
[0027] Step 104: Display the predicted number of graduates.
[0028] The initial student data across all levels refers to the enrollment and registration data for the four levels of higher education (associate degree, bachelor's degree, master's degree, and doctoral degree) in the same academic year, which is the first input source for the prediction model.
[0029] Quantitative thresholds convert the rigid requirements in the training program (number of courses, number of months of practical training, number of failed courses, number of papers) into calculable numerical boundaries for use in the internal logic of the model.
[0030] The data is specific to each level of education and is a dynamic record of performance collected separately. For associate degree programs, it is the practical training assessment score; for bachelor's degree programs, it is the pass rate of core courses; for master's degree programs, it is the academic research foundation score and the completion rate of professional projects; and for doctoral degree programs, it is the number of research achievements. This data is used to differentiate and adjust the mobility probability.
[0031] The anonymized individual profile data, after removing name, student ID, and national ID number, includes academic, research, and skill feature vectors, which are used by the model to make individual-level corrections to the group probability during the secondary calibration phase.
[0032] The initial graduation probability is the group-level graduation likelihood output by the model during a calibration phase. It serves as the baseline probability for subsequent individual corrections and the superposition of rigid rules.
[0033] The method provided in this embodiment acquires initial student data across all levels, combines input quantification thresholds with professional-specific data for each level, and outputs an initial graduation probability. This probability is then corrected a second time using desensitized individual profiles, and the results are displayed. This progressively refined and feedback-driven sequence incorporates enrollment scale, rigid regulations, and individual performance into the same prediction channel. This allows the model to integrate hard institutional constraints with soft data features during the inference stage, thereby determining the graduation scale signal in advance without relying on post-hoc manual correction. This significantly reduces the risk of systemic bias caused by the disconnect between regulations and data in traditional approaches.
[0034] In one optional implementation, the prediction model is obtained through the following training process: an LSTM network is trained using historical academic year data to obtain a first prediction network for predicting associate degree, bachelor's degree, and master's degree levels; a GRU branch is added to the first prediction network to obtain a dedicated network suitable for doctoral level.
[0035] Each academic year, the predictive model to be trained is iteratively optimized based on the actual number of graduates and the actual implementation of the training program.
[0036] The objective function is to minimize the mean square error between the predicted and actual values, and the parameters of the model to be predicted are updated using the gradient descent method.
[0037] LSTM networks, or Long Short-Term Memory networks, are responsible for capturing semester-level time dependencies among associate's, bachelor's, and master's degree students and outputting common flow factor weights.
[0038] The GRU branch, or gated recurrent unit branch, is specifically added for doctoral students to memorize the nonlinear jump characteristics of research milestones.
[0039] Each academic year, the network is retrained using real graduation data, with mean squared error as the objective function. Gradient descent is used to update the network parameters, maintaining the long-term effectiveness of the model.
[0040] By adding GRU branches only to the doctoral level on the existing LSTM channels and iterating in a closed loop with real graduation data each academic year, the model structure retains the advantages of the linear undergraduate-master degree system while adding a dedicated memory path to address the non-linear jumping characteristics of doctoral research pace. The dual-channel difference in the structure allows the same set of parameter files to capture both semester-level patterns and node-level mutations, thereby achieving precise adaptation to different academic system logics while ensuring the overall architecture is consistent, avoiding the high development and synchronization costs of maintaining a separate model for each level.
[0041] An optional implementation method includes: extracting quantitative indicators of the training program for the corresponding professional level, including curriculum system, assessment standards, graduation requirements and schedule; and converting the indicators into quantitative thresholds for the corresponding professional level.
[0042] The quantitative indicators of the training program, and the numerical expression of four types of institutional elements—curriculum system, assessment standards, graduation requirements, and schedule—are the raw materials for generating quantitative thresholds.
[0043] The curriculum system, including the number of core and elective courses and credit requirements specified in the training program, is used to calculate the course category threshold.
[0044] Schedule, with start and end times or milestone dates for each stage, is used to set node class thresholds and drive dynamic adaptation cycles.
[0045] By automatically mapping the curriculum system, assessment standards, graduation requirements, and schedule to quantitative thresholds, any revision of the training program can be quickly translated into boundary conditions that the network can recognize. This external table is managed separately from the network weights, which allows academic staff to update the system independently while ensuring that the model always operates based on the latest boundary conditions. This eliminates the logical mismatch where the system has been changed but the model is still calculating according to the old rules, and shortens the response cycle between system revision and prediction update.
[0046] An optional implementation method further includes: in response to receiving adjustment information of the training program, obtaining the adjustment content, automatically identifying the level, major, and stage, and synchronously updating the corresponding quantification.
[0047] Adjust information, such as notifications from the academic affairs system regarding changes to the curriculum plan, such as adding or deleting courses, modifying credits, and changing deadlines.
[0048] The system automatically identifies and parses the hierarchy, specialty, and stage fields in the information through the model interface, eliminating the need for manual input.
[0049] Synchronous updates write the new quantification thresholds into the external reference table in real time and trigger the model to re-infer, ensuring that the prediction logic is consistent with the latest system.
[0050] By automatically identifying the changed level, major, and stage after receiving adjustment information and updating it synchronously, the system then triggers model re-inference, allowing the system adjustment signal to directly affect the network input without manual remodeling. This automatic identification and synchronization mechanism compresses the "system change → numerical change → input change" chain into a second-level response, thereby ensuring that the decision-making end can obtain the corresponding new graduation scale estimate as soon as the training program is fine-tuned, avoiding a window period.
[0051] One possible implementation method is to use hierarchical professional-specific data, which refers to dynamic performance data collected separately according to the training level.
[0052] For associate degree programs, the results are based on practical training assessments.
[0053] The undergraduate level is measured by the pass rate of core courses.
[0054] For academic master's programs, the score is based on research foundation; for professional master's programs, the score is based on completion of practical projects.
[0055] The doctoral level is determined by the number of research achievements.
[0056] Dynamic performance data refers to the real-time records of grades, assessments, research, and practical activities generated during the teaching process.
[0057] Research foundation points are awarded based on the papers, patents, and projects obtained by master's students, and are used to measure their research progress.
[0058] Practical project completion rate: The percentage of projects completed by master's students in enterprises or training bases is used to measure practical ability.
[0059] By configuring dynamic performance fields that best represent the mobility risk for associate's, bachelor's, master's academic, master's professional, and doctoral programs, and completing field-hierarchy binding at the data entry layer, the model input layer is equipped with hierarchical discrimination function. This binding strategy allows data flows from different educational logics to be diverted to the corresponding feature channels at the entry stage, reducing cross-hierarchical noise interference. This enables accurate feature supply from different channels within the same network under a unified network structure, reducing prediction distortion caused by the mixed use of indicators.
[0060] One optional implementation method is to set different dynamic adaptation periods for associate degree, bachelor's degree, master's degree and doctoral degree levels, and synchronize performance data in real time.
[0061] For associate degree programs, the smallest unit is the practical training cycle, and the assessment results are uploaded after each practical training cycle is completed.
[0062] Both undergraduate and master's programs use semester-end grades concurrently.
[0063] At the doctoral level, the number of research achievements and the milestone achievement marker factor are updated incrementally after the date or achievement status of the project proposal, mid-term, pre-defense, and submission milestones change.
[0064] By setting adaptive cycles (practical training cycle / semester / research node) that are naturally synchronized with the teaching rhythm for each level, and by uploading actual performance and incremental refresh factors immediately when each cycle ends or node changes, the network weights can continuously absorb the latest teaching signals without re-initializing. This cycle-increment mechanism breaks down the "annual major update" into "cycle micro-updates", which not only maintains the model's long-term memory, but also ensures that its short-term sensitivity is always aligned with the actual training progress, thereby continuously outputting the flow probability consistent with the current situation throughout the academic year.
[0065] One optional implementation method for extracting the initial student data includes: for junior college and undergraduate levels, extracting the number of students enrolled in the latest semester; if it is the initial enrollment period, extracting the actual number of students enrolled, and supplementing the number of students who registered but did not enroll and the number of students who withdrew in the early stages.
[0066] For the master's level, extract the number of students enrolled after entering the core training stage, i.e., after the course ends. If it is the initial enrollment period, supplement the number of students eliminated during the qualification review, and make adjustments separately for academic and professional programs.
[0067] For doctoral students, the number of students who have passed the thesis proposal defense is extracted. If it is the early stage of enrollment, the number of students eliminated due to qualification assessment and the number of students eliminated due to adjustment of research direction are added.
[0068] By selecting the most stable and representative calculation basis for each level during the data extraction stage, and simultaneously supplementing the data on those who did not enroll or were disqualified at the beginning of enrollment, the initial number of people entering the network has already eliminated those who dropped out early. This avoids the artificially high error of directly including newly enrolled students into the graduation pool, improves the authenticity of the input data from the source, and ensures that all subsequent probability calculations are based on a reliable basis.
[0069] An optional implementation method is to display the graduation number prediction results, including: displaying the target year's graduation number, confidence interval, prediction accuracy, and delayed graduation and early graduation at the hierarchical and professional level; and summarizing and outputting the overall number of graduates across all levels and majors, their proportion, and year-on-year change trends.
[0070] The confidence interval, given by the Bootstrap method, is used to quantify the uncertainty of the prediction.
[0071] For delayed / early graduation, the model outputs a detailed list of graduates who exceed or fall short of the full study period, which can be used to stagger resource allocation.
[0072] A cross-level overview, which is a school-wide total view that summarizes the number of graduates from junior college to doctoral degree, is used for macro-level decision-making.
[0073] Year-on-year change trend, the direction of increase or decrease compared with the previous year's forecast-actual, is used to assess the effect of policy or system adjustments.
[0074] Prediction accuracy is a relative indicator of how well the predicted value matches the actual value, used to monitor the long-term stability of the model.
[0075] By first outputting detailed information at the hierarchical and professional levels, and then automatically summarizing it into a cross-level overview, and pushing the details and the summary simultaneously through the same interface, the decision-makers can both drill down to specific professional levels to view risks and observe structural changes at the overview level. This detailed-to-general display logic ensures that the same set of calculation results can simultaneously meet the refined management needs of grassroots teaching units and the macro-level resource coordination needs of the university, without the need for additional manual summarization or secondary report development.
[0076] An optional implementation method for displaying the graduation number prediction results includes: listing the core factors and factor weights used in this prediction for each level of major; calculating the contribution value of each factor change to the increase or decrease in the number of graduates; and displaying the graduation number prediction results corresponding to the core factors and factor weights.
[0077] The core factors, such as dropout, leave of absence, and delayed graduation, are institutional variables that directly determine the direction of mobility, and their weights are derived from model training.
[0078] Factor weights represent the coefficients that indicate the magnitude of the impact of each core factor on the number of graduates, and are used to quantify the contribution ratio.
[0079] Contribution value is the absolute amount by which the number of graduates increases or decreases when the value of a certain factor changes.
[0080] Generate traceable impact reports, along with documents listing factors, weights, contribution values, and linguistic explanations, for use as a basis for administrative decision-making.
[0081] By transforming factor weights into descriptions of the direction and magnitude of increases and decreases in the number of graduates after inference, and then outputting them using natural language templates, an automatic translation channel is established between the internal numerical values of the model and the external management language. This mapping mechanism allows users to know what scale of change a policy change brings without understanding the statistical meaning, thereby directly embedding the algorithm output into the context of administrative decision-making and reducing communication costs.
[0082] An optional implementation method for displaying the predicted number of graduates includes: displaying student flow curves at the hierarchical and professional levels, where the horizontal axis of the curve represents the academic year and the vertical axis represents the number of students; dynamic pie charts and bar charts showing the proportion of influence of each core factor and the training program, with the dynamic indicator automatically switching according to the predicted year.
[0083] The student mobility curve, with the horizontal axis representing the academic year and the vertical axis representing the number of students, is used to show the changes in the scale of graduations at each level over time.
[0084] The dynamic pie chart automatically refreshes as the forecast year changes, displaying a circular chart showing the influence percentage of each core factor in real time.
[0085] Dynamic bar charts, updated synchronously with pie charts, are used to compare the contribution of different factors.
[0086] Automatic switching: The dashboard component is redrawn in real time based on the predicted year selected by the user, without the need for manual page refresh.
[0087] A visual dashboard that integrates curves, pie charts, and bar charts into an interactive display interface, supporting one-click export and drill-down.
[0088] By displaying time-series flow curves and real-time percentage pie charts and bar charts in parallel within the same visualization dashboard, and using an automatic switching component to refresh in real time with the predicted year, the time trend and structural percentage are presented in tandem on the same screen. This dual-view design allows observers to grasp the changes in graduation scale over time and the changes in the contribution structure of each factor without switching interfaces, thereby quickly locating key driving factors and shortening the time from observing phenomena to forming decisions.
[0089] For example, please refer to Figure 2 , Figure 2 Exemplary application scenarios of one or more embodiments of the present invention are shown.
[0090] With training level as the core differentiating dimension, professional training program as the key calibration dimension, and two auxiliary dimensions of stage division and student stratification, general modules ensure the framework uniformity, while exclusive modules combined with the training program meet the dual personalized needs of level and major.
[0091] Establish a system of commonalities at different levels, unique characteristics for each major, and rigid factors in the training program: Extract the core factors common to the four training levels (dropout, leave of absence, and extension of graduation), and supplement them with level-specific and major-specific core factors (such as passing the practical training assessment for a certain major in a junior college or passing all core courses for a certain major in a bachelor's degree) in combination with the courses, assessments, and graduation requirements in the professional training program, so as to achieve factor quantification and strong binding with the training program.
[0092] The design level guides the stage division logic to adapt to the training program: the basic stage is divided according to the characteristics of the four levels of study, and then the stage nodes and deduction rules are adjusted according to the progress arrangement in the professional training program (such as the training program stipulates that the second semester of the master's degree is the research / practice stage) to adapt to the training rhythm of different majors.
[0093] Integrating group patterns, individual profiles, and training program requirements to calibrate the model: Based on historical data and training program requirements, a group mobility probability model is constructed. This model is then calibrated using desensitized individual profiles (academic / research / skill characteristics) and further modified by combining rigid rules in the training program (such as not being allowed to graduate if internship requirements are not met), thereby improving the alignment between predictions and actual training.
[0094] Establish a dynamic iteration and training program update mechanism: general modules and special modules are iteratively optimized based on data feedback, and a training program update response channel is established simultaneously. When the training program is adjusted, the parameters of relevant factors and stage rules are automatically updated.
[0095] Please refer to Figure 2 The data collected consists of three core categories: First, the overall enrollment data for all levels must cover basic information such as the actual number of students enrolled, the student composition, and the initial number of students in the target grades across all four levels. Second, level-specific and major-specific data are customized according to the level: for associate degree programs, the focus is on skills assessments and practical training records; for bachelor's degree programs, the focus is on course grades and records of students changing majors; for master's degree programs, data related to research or practice are collected separately for academic and professional programs; and for doctoral degree programs, the focus is on research achievements and grant funding. Third, the professional training program data focuses on extracting quantifiable information such as the curriculum system, assessment standards, graduation requirements, and schedule.
[0096] The data is cleaned and standardized, and statistical standards are unified. The requirements of the training program are transformed into quantitative thresholds, such as ≥10 core courses and ≥6 months of practical training. At the same time, stratified and professionally desensitized individual profiles are constructed to extract the core characteristics of students that meet the requirements of the training program, and privacy protection regulations are strictly followed.
[0097] A comprehensive database covering all levels and disciplines has been established, including a general basic database, level-specific and discipline-specific databases, and a training program database. This enables data classification and storage, as well as cross-database association and retrieval. At the same time, a training program update interface has been established to support the rapid input and parsing of new or revised training programs.
[0098] Clearly define the higher education level and specific major of the predicted subjects, and access the corresponding level and major-specific databases and curriculum databases to ensure data matching. Extract initial input data: for associate degree and bachelor's degree programs, prioritize extracting the latest semester's enrollment numbers; at the beginning of enrollment, extract the actual enrollment numbers and supplement with data on those who registered but did not enroll or those who withdrew initially, and adjust according to the enrollment adaptation period requirements in the curriculum; for master's degree programs, extract the enrollment numbers of those who have entered the core training stage (after course completion), supplement with data on those eliminated during the qualification review at the beginning of enrollment, and adjust according to the differences between academic and professional curriculum programs; for doctoral degree programs, extract the enrollment numbers of those who have passed the thesis proposal, supplement with data on those eliminated during qualification assessments and research direction adjustments at the beginning of enrollment, and adjust according to the preliminary research achievement requirements in the curriculum.
[0099] A comprehensive, multi-level, and multi-disciplinary refined personnel mobility prediction model is developed. First, a quantitative system of personnel mobility factors is constructed, encompassing both level and major. Common core factors include withdrawal, leave of absence, and delayed graduation, with basic probabilities differentiated by level. Quantification rules are deeply integrated into the training program requirements; for example, the dynamic adjustment coefficient for undergraduate withdrawal factors matches the rigid requirement of expulsion for failing ≥4 courses in the training program. Level- and major-specific core factors are customized according to level and major. For associate degree programs, the focus is on skills assessment and practical training duration; for bachelor's degree programs, the focus is on core courses and practical credits; for master's degree programs, research foundation or practical project-related factors are set separately for academic and professional programs; and for doctoral degree programs, factors are based on research achievements and progress milestones. All specific factors are quantified based on the training program. Auxiliary factors are configured as needed and linked to the training program, with interactive effects establishing a level- and major-specific matrix, designed in conjunction with the rigid rules of the training program. Secondly, a four-dimensional hierarchical prediction framework is implemented, dividing the basic stages according to hierarchical differences. For associate degree programs, this includes the enrollment adaptation period, skills training period, and graduation sprint period; for bachelor's degree programs, it includes the enrollment adaptation period, academic stabilization period, and graduation sprint period; for master's degree programs, it includes the course learning period, core training period, and thesis defense period; and for doctoral degree programs, it includes the qualification assessment period, scientific research breakthrough period, and thesis submission and defense period. Then, the stage nodes and deductive logic are adjusted in conjunction with the progress of the professional training program. Finally, a second calibration of student stratification is performed according to the progress of the training program, and the mobility probability is adjusted differently. Finally, dynamic adaptation and model optimization are implemented. For associate degree programs, data is synchronized in real time according to the training cycle; for bachelor's and master's degree programs, according to the semester; and for doctoral degree programs, according to the scientific research nodes of the training program, automatically adjusting factor probabilities. When the training program is adjusted, updated content is entered through a dedicated interface, and the model automatically identifies influencing factors and stage nodes and updates parameters synchronously. The model adapts to factors at each level and in each major in light of changes in the external environment. An LSTM model is generally used for training, while a GRU model is additionally introduced for doctoral programs to adapt to non-linear changes in scientific research progress. Each academic year, factor weights and interaction coefficients are optimized based on actual graduation data and the implementation of the training program.
[0100] The system provides comprehensive, multi-level, and multi-major forecasting results with traceable interpretation. First, it outputs forecasting results by dimension, precisely providing detailed data on the number of graduates in the target year, confidence intervals, forecast accuracy, and details of delayed and early graduations for each level and major. It also provides a comprehensive overview of cross-level and cross-major graduation numbers, percentages, and year-on-year trends. Second, it generates traceable impact reports, quantifying the impact of core factors, curriculum requirements, and dynamic events on the number of graduates at each level and major, such as "a 20% increase in the rate of meeting training duration targets in a certain associate degree program resulted in an increase of 10 graduates." Simultaneously, it conducts cross-level and cross-major comparative analyses to extract differences in flow patterns. Finally, it builds an integrated, multi-level, and multi-major visual dashboard, displaying flow curves, charts showing the impact percentages of core factors and curriculum programs, and generating application reports for each level and major, as well as cross-level coordination reports. These reports are used for specific scenarios such as training resource allocation, course scheduling, and research funding allocation, and for overall university-wide resource coordination.
[0101] The foregoing description of specific exemplary embodiments of the invention is for illustrative and explanatory purposes. These descriptions are not intended to limit the invention to the precise forms disclosed, and it will be apparent that many changes and variations can be made in accordance with the foregoing teachings. The exemplary embodiments were chosen and described in order to explain the specific principles of the invention and its practical application, thereby enabling those skilled in the art to implement and utilize various different exemplary embodiments of the invention, as well as various different choices and variations. The scope of the invention is intended to be defined by the claims and their equivalents.
Claims
1. A method for predicting the number of graduates across all levels of higher education, characterized in that, Includes the following steps: Obtain initial student data for all levels, including associate degree, bachelor's degree, master's degree, and doctoral degree. According to the hierarchical professional dimension, the initial student data, the quantification threshold and the hierarchical professional-specific data are input into the trained prediction model. The prediction model outputs the initial graduation probability corresponding to the hierarchical professional dimension. The quantification threshold is used to indicate the pass standard in the training program. Based on the hierarchical and professional dimensions, the initial graduation probability was corrected using anonymized individual profile data to obtain the predicted number of graduates; The predicted number of graduates is displayed.
2. The method according to claim 1, characterized in that, The prediction model is obtained through the following training process: The LSTM network was trained using historical academic year data to obtain the first prediction network for predicting associate degree, bachelor's degree and master's degree levels. By adding a GRU branch to the first prediction network, a dedicated network suitable for doctoral level is obtained; Each academic year, the prediction model to be trained is iteratively optimized based on the actual number of graduates and the actual implementation of the training program. The objective function is to minimize the mean square error between the predicted and actual values, and the parameters of the model to be predicted are updated using the gradient descent method.
3. The method according to claim 1, characterized in that, The method includes: Quantitative indicators for the training programs at the corresponding professional levels were extracted, including the curriculum system, assessment standards, graduation requirements, and schedule. The indicators are converted into quantitative thresholds for the corresponding professional dimensions.
4. The method according to claim 3, characterized in that, The method further includes: In response to receiving information about adjustments to the training program, the system obtains the adjusted content, automatically identifies the level, major, and stage, and updates the corresponding quantifications simultaneously.
5. The method according to claim 1, characterized in that, Level-specific data refers to dynamic performance data collected separately according to the training level; among which For associate degree programs, the results are based on practical training assessments. At the undergraduate level, the pass rate is the core course pass rate. For academic master's programs, the score is based on research foundation; for professional master's programs, the score is based on completion of practical projects. The doctoral level is determined by the number of research achievements.
6. The method according to claim 5, characterized in that, Different dynamic adaptation cycles are set for associate degree, bachelor's degree, master's degree, and doctoral degree levels, and performance data is synchronized in real time; among them For associate degree programs, the smallest unit is the practical training cycle. Assessment results are uploaded upon completion of each practical training cycle. Both undergraduate and master's programs use semester-end grades concurrently. At the doctoral level, the number of research achievements and the milestone achievement marker factor are updated incrementally after the date or achievement status of the project proposal, mid-term, pre-defense, and submission milestones change.
7. The method according to claim 1, characterized in that, The methods for extracting the initial student data include: For both associate degree and bachelor's degree programs, extract the number of students enrolled in the latest semester. If it is the initial enrollment period, extract the actual number of students enrolled, and supplement the number of students who registered but did not enroll and the number of students who withdrew in the early stages. For the master's level, extract the number of students enrolled after entering the core training stage, i.e., after the course has ended. If it is the initial enrollment period, supplement the number of students eliminated during the qualification review, and make adjustments separately for academic and professional programs. For doctoral students, the number of students who have passed the thesis proposal defense is extracted. If it is the early stage of enrollment, the number of students eliminated due to qualification assessment and the number of students eliminated due to adjustment of research direction are added.
8. The method according to claim 1, characterized in that, The presentation of the predicted number of graduates includes: Displays the target year's number of graduates, confidence intervals, prediction accuracy, and delayed / early graduation rates across different professional levels. The system provides a comprehensive overview of the number of graduates from all levels and majors, including their percentage and year-on-year trends.
9. The method according to claim 1, characterized in that, The presentation of the predicted number of graduates includes: List the core factors and factor weights used in this forecast for each level of specialization; Calculate the contribution of each factor change to the increase or decrease in the number of graduates; Display the predicted number of graduates corresponding to the core factors and factor weights.
10. The method according to claim 1, characterized in that, The presentation of the predicted number of graduates includes: The curves show student mobility across different academic levels and majors, with the horizontal axis representing the academic year and the vertical axis representing the number of students. Dynamic pie charts and bar charts showing the proportion of influence of each core factor and cultivation program, with the dynamic indicator automatically switching according to the predicted year.