Machine learning-based overseas study application admission probability prediction method
By combining a multi-layer neural network model with domain-specific data processing and text semantic analysis, the problems of data utilization and adaptability in study abroad application evaluation are solved, and more accurate and efficient admission probability prediction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING KAOTI EDUCATION TECHNOLOGY CO LTD
- Filing Date
- 2025-11-21
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies for evaluating study abroad applications suffer from insufficient utilization of in-depth data, difficulty in modeling complex relationships, and weak domain generalization ability, resulting in poor consistency and low accuracy of evaluation results. They are also unable to effectively handle unstructured text data and adapt to changes in university majors.
A multi-layer neural network model is adopted, which combines domain-specific data standardization, deep semantic feature extraction of text, and dynamic matching degree feature construction. The model is trained to predict the admission probability through supervised learning and regularization strategies.
It improves the consistency and accuracy of evaluation results, can comprehensively explore the applicant's overall strength, adapts to the changes of different universities and majors, enhances the generalization ability of the model, reduces human interference, and improves evaluation efficiency.
Smart Images

Figure CN121835972A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, and particularly relates to a method for predicting the admission probability of a study abroad application based on machine learning. BACKGROUND
[0002] The evaluation of a study abroad application is a complex decision-making process involving multiple factors. Currently, the common evaluation methods in this field mainly rely on two types of technology: one is qualitative evaluation based on artificial experience, and the other is quantitative prediction based on traditional statistical models.
[0003] In the artificial evaluation method, the study abroad consultant relies on his personal knowledge reserve and past case experience to make subjective judgments on the competitiveness of the applicant. This method has significant limitations: first, the experience of the consultant is difficult to quantify and standardize, and the evaluation standards of different consultants differ greatly, resulting in poor consistency and low reproducibility of the evaluation results; second, the knowledge range of an individual is limited, and it is difficult to fully and accurately grasp the admission preferences and potential laws of thousands of different majors in various institutions around the world.
[0004] To overcome the defects of artificial subjectivity, quantitative methods based on simple statistical models such as linear regression and logistic regression are introduced in existing technologies. This type of method mainly uses standardized test scores (such as GPA, GRE, GMAT, etc.) and other structured data that are easy to quantify to build models. However, such models have inherent technical defects: first, they cannot effectively process and integrate unstructured text data (such as personal statements, recommendation letters) in the application materials. These texts contain a wealth of deep semantic information reflecting the applicant's personal qualities, research potential, and professional matching, but simple models lack the ability to mine such information, resulting in a significant underutilization of data value. Second, linear models assume a simple linear relationship between features and results, and cannot capture and model complex nonlinear interactions between different features. For example, the research experience of an applicant has a high dependence on whether the target major of the applicant focuses on research ability, and this synergistic effect between features cannot be expressed by a linear model. Third, directly applying general machine learning models without considering the specificity of the study abroad application field results in weak model generalization ability. When faced with new institutions, new majors, or changes in enrollment policies, such models will have a significant decline in prediction accuracy and reliability due to the lack of field knowledge in their feature engineering and model structure.
[0005] Therefore, the prior art has obvious deficiencies in depth data utilization, complex relationship modeling and field generalization ability when dealing with the specific problem of studying abroad application. These deficiencies are due to the multi-source heterogeneity of the application data itself (coexistence of structured data and unstructured text), the complexity of the relationship between features, and the strong field dependence. Developing an intelligent prediction method that can deeply integrate domain knowledge, fully exploit the value of heterogeneous data, and accurately model complex feature relationships is of great significance and necessity for improving the objectivity, accuracy and efficiency of studying abroad application evaluation. SUMMARY
[0006] Based on the above purpose, the present application provides a machine learning-based probability prediction method for studying abroad application, comprising the following steps: Data collection and standardization step: collect multi-source heterogeneous data of the applicant, and perform field-specific standardization processing on the non-standardized structured data to generate standardized numerical features; Text deep semantic feature extraction step: performing natural language processing on the unstructured text data in the application documents, using a pre-built studying abroad application field dictionary and a pre-set ability dimension to extract quantitative semantic features; Dynamic matching degree feature construction step: based on the specific requirements of the target university major, comparing and calculating the processed applicant features with the specific requirements to generate dynamic matching degree features representing the fit degree between the applicant and the major; Neural network model processing step: input the standardized numerical features, the quantitative semantic features and the dynamic matching degree features as inputs into a specially trained multi-layer neural network model, which outputs a continuous value representing the probability of admission; Wherein, the special training refers to using historical studying abroad application data and its corresponding admission results to train the multi-layer neural network model in a supervised learning manner, and adopting a regularization strategy during the training process to improve the model generalization ability.
[0007] Preferably, the field-specific standardization processing of non-standardized structured data in the data collection and standardization step comprises: For academic achievements from different countries or regions, according to the official achievement conversion standards or recognized equivalence principles of their respective education systems, query the pre-established multi-education system equivalent achievement comparison database, map various grading or percentage achievements to a unified standardized score scale, and thus obtain the academic achievement standardized value; For extracurricular activities and honor records, according to the duration, the importance of the role assumed, and the level of the awards obtained, a predefined quantitative mapping rule based on expert experience is used to assign weights to different dimensions of each item and perform weighted calculation, and finally a comprehensive extracurricular activity quantitative value is generated.
[0008] Preferably, the quantitative semantic feature extraction step in the text deep semantic feature extraction step includes: For personal statement text, the study abroad application field dictionary is used for scanning and matching, the occurrence frequency and density of the words in the text that match the key words in the dictionary are counted, and they are associated with the total word count of the text, and a personal statement field key word density index reflecting the relevance of the content of the document to the academic field is calculated. For recommendation letter text, first, sentiment analysis is performed to screen out sentences containing positive evaluation, and then words or phrases related to each core ability listed in the preset ability dimension set are identified and extracted in these sentences, and the occurrence frequency of words representing different abilities is counted, and finally a multi-dimensional vector is formed, where the value of each dimension represents the evaluation strength of the recommender on the applicant's ability, i.e. a multi-dimensional recommendation letter ability evaluation vector.
[0009] Preferably, the construction process of the pre-constructed study abroad application field dictionary includes: Through network crawler technology, raw text data is collected from a large number of official recruitment websites of overseas famous universities, professional introduction pages of various colleges, course descriptions and project training target documents; The collected raw text data is automatically cleaned, segmented, tagged, and processed to remove common stop words; Using a keyword extraction algorithm based on term frequency and inverse document frequency, high-frequency and domain-discriminative core words and key phrases are automatically selected from the processed text to form an initial dictionary library; Invite experts in the field of study abroad applications to form a review group to manually review, screen, denoise and supplement the initial dictionary library, add important field terms that cannot be obtained through automatic extraction, and finally form an authoritative and comprehensive study abroad application field dictionary.
[0010] Preferably, the establishment process of the preset ability dimension set includes: Through retrospective analysis of a large number of successful study abroad application cases, a series of soft skills and characteristics frequently considered in admission decisions are summarized; Collect the experiences of experienced admissions officers, study abroad consultants and academic advisors through interviews and questionnaires to collect the ability dimensions they consider essential; The capability list obtained from the above two aspects is fused and de-duplicated to form a preset capability dimension set covering scientific research potential, academic leadership, team cooperation spirit, professional ethics and academic integrity, cross-cultural communication and expression capability.
[0011] Preferably, the dynamic matching degree feature construction step comprises: From the official release channel of the target university and the target major, the latest course introduction, training plan and enrollment requirement text are obtained, and a set of noun keywords for describing the core skills, course themes and expected student background of the major are extracted from the text mining technology to form a major core requirement keyword set; Calculate the text semantic similarity between the applicant's personal statement text and the major core requirement keyword set. The similarity is quantified by calculating the vector angle cosine value of both texts in a high-dimensional vector space, thereby obtaining the professional direction semantic matching degree; Integrate the academic achievement data of the admitted students in the target major in the past several admission cycles, calculate the statistical distribution of the standardized value of the academic achievement, and then calculate the relative percentile or standard score position of the applicant's academic achievement standardized value in the statistical distribution. Through a linear or nonlinear mapping function, the position information is converted into an academic background competitiveness index that is easy for a neural network model to process.
[0012] Preferably, the specialized training in the neural network model processing step comprises: The multi-layer neural network model adopts a feedforward neural network structure. The dimension of the input layer is equal to the total number of all input features. The output layer uses a S-shaped activation function to limit the output value between 0 and 1 to represent the probability. The number of layers and the number of neuron nodes in each layer of the hidden layer are determined through a systematic hyperparameter optimization search process. The training process uses an adaptive moment estimation optimization algorithm to minimize the cross-entropy loss function between the predicted value and the true admission label; The regularization strategy is to apply the dropout method to the output of the hidden layer during the training phase. During the forward propagation process of each training iteration, the output values of a certain proportion of neuron nodes in the hidden layer are temporarily set to zero. The dropout proportion is determined as a hyperparameter together with other structural hyperparameters of the model through the hyperparameter optimization search process.
[0013] Preferably, the systematic hyperparameter optimization search process adopts a grid search or random search strategy: A hyperparameter space that needs to be optimized is preset, including but not limited to: the number of layers of the neural network hidden layer, the number of neuron nodes in each hidden layer, the dropout ratio of the dropout method, the learning rate; selecting a plurality of different combinations of hyperparameters from the hyperparameter space according to a predetermined strategy; For each combination of hyperparameters, initializing a corresponding neural network model, training it using the training dataset, and then evaluating its performance indicators on the independent validation dataset; Selecting the model structure and parameters corresponding to the combination of hyperparameters with the best performance indicators on the validation dataset as the final trained model.
[0014] Preferably, after the neural network model processing step, it further includes: A prediction result interpretability processing step: applying a post-hoc attribution analysis algorithm to calculate the gradient or sensitivity of the final output result of the neural network model with respect to each input feature, and evaluating the importance of each input feature's contribution to this particular prediction result; Outputting the top several input features with the highest importance and their corresponding contribution values, thereby providing the user with key decision factor explanations for this admission probability prediction.
[0015] Preferably, the post-hoc attribution analysis algorithm uses the integrated gradients method: Integrating the gradient of the model output along a path that smoothly changes the model input features from a set of baseline values representing information missing state to the actual feature values of the current applicant; Calculating the cumulative sum of the product of the average gradient change of each input feature along this path and its feature value change, which is the contribution estimate of the feature to the final prediction result; Normalizing the contribution estimates of all features so that the sum of the absolute values of all features' contribution estimates is one, thereby obtaining the relative contribution percentage of each feature.
[0016] Advantages of the present application: 1、The present application can eliminate the interference of human subjective factors through algorithm standardization evaluation process. This method can generate a unified evaluation model by learning a large amount of historical data, so that the evaluation results of different consultants or systems tend to be consistent, ensuring the fairness and consistency of the evaluation process.
[0017] 2、The present application can extract deep semantic information from unstructured data and mine potential features related to the competitiveness of the applicant by introducing advanced natural language processing technology. This technology can comprehensively analyze the application materials, consider factors such as personal traits, research potential, and professional matching degree, greatly improve the value of data utilization, and improve the accuracy of evaluation results.
[0018] 3. This invention, by employing deep learning and other nonlinear modeling techniques, can capture the complex interactive relationships between various features, especially the synergistic effect between the applicant's research experience and the target professional requirements. By deeply integrating different features, the evaluation model can more accurately reflect the applicant's comprehensive strength, avoiding the limitations of simple linear models.
[0019] 4. This invention, by combining professional knowledge in the field of study abroad applications, designs a highly adaptable feature engineering and model structure. It can dynamically adjust evaluation criteria according to changes in different universities, majors, and policies, significantly improving the model's adaptability and predictive accuracy in different scenarios. Through the embedding of domain knowledge, the model's generalization ability is enhanced, effectively responding to future changes and maintaining high predictive accuracy.
[0020] 5. This invention, by comprehensively utilizing multi-source heterogeneous data and combining it with domain knowledge, can more comprehensively and accurately assess the applicant's competitiveness. Simultaneously, the automated assessment method based on machine learning technology greatly improves assessment efficiency, reduces interference from manual operations, and avoids omissions and biases that may exist in traditional assessment processes. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0022] Fig. 1 This is a flowchart of the steps of the method of the present invention; Fig. 2 A flowchart illustrating the steps involved in constructing the dynamic matching degree features in the method of this invention; Fig. 3 The flowchart shows the steps of the post-hoc attribution analysis algorithm of the present invention using the integrated gradient method. Detailed Implementation
[0023] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should also be noted that, to make the embodiments more comprehensive, the following embodiments are the best and preferred embodiments, and those skilled in the art can use other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0024] Please see Figs. 1-3This invention provides a machine learning-based method for predicting the probability of university admissions for international students. First, it collects heterogeneous data from multiple data sources, including structured data (such as GPA, GRE scores, and personal information) and unstructured data (such as personal statements, letters of recommendation, and research experience). During data standardization, structured data undergoes domain-specific standardization, transforming differences in admission requirements and policies across different universities, as well as various evaluation indicators, into a unified standard. For example, GPA grading standards may differ between countries or universities; in such cases, adjustments need to be made according to the target university's standards to ensure all data can be compared and processed under the same standard. This processing step lays the foundation for subsequent feature extraction and model training, guaranteeing data consistency and comparability.
[0025] Furthermore, natural language processing (NLP) techniques are used to process the unstructured text data in application documents. Applicants' personal statements, letters of recommendation, and other textual materials contain a wealth of non-quantitative information about their potential, research interests, academic background, and suitability for the chosen field. By using a pre-constructed dictionary of study abroad application areas and a system of competency dimensions (such as research ability, teamwork, and academic interests), quantitative semantic features that match the target university and program can be extracted. For example, "research experience" in a personal statement can be transformed into a quantitative feature measuring research potential, and "leadership" in a letter of recommendation can be transformed into a feature reflecting an individual's leadership abilities. After this processing, the text data can be transformed into structured features that can be directly used in subsequent machine learning models.
[0026] Furthermore, the specific requirements of the target universities and programs are compared and calculated against the applicant's characteristics. These requirements include academic background, research experience, and language proficiency, while the applicant's characteristics consist of standardized numerical and semantic features extracted in the preceding steps. Through dynamic matching degree calculation, the applicant's various characteristics are compared with the requirements of the target universities, generating a dynamic matching degree feature that characterizes the applicant's fit with the program. This feature not only reflects the applicant's overall strength but also demonstrates the degree of alignment with the specific university's program requirements, thus helping to improve the accuracy of the assessment.
[0027] Standardized numerical features, quantified semantic features, and dynamic matching degree features are input into a specially trained multi-layer neural network model. The model outputs a continuous value representing the probability of admission based on these input features. During training, a supervised learning approach is employed, using historical study abroad application data and corresponding admission results to train the neural network. To improve the model's generalization ability, regularization strategies, such as L2 regularization, are used to prevent overfitting. During training, the neural network automatically identifies and learns the complex relationships between various features and, leveraging the advantages of deep learning, captures the interaction effects between non-linear features through its multi-layer structure, thereby accurately predicting the probability of admission.
[0028] The embodiments of the present invention combine domain knowledge, deep learning technology, and unstructured data processing, which significantly improves the accuracy, stability, and generalization ability of predictions.
[0029] In one possible implementation, due to differences in education systems across different countries or regions, applicants' academic achievements (such as GPA, course grades, etc.) may be assessed using different grading standards (such as letter grades, percentage grades, credit grades, etc.). To address this issue, embodiments of the present invention propose the following standardized steps: First, a database containing grading standards from different education systems needs to be established. This database includes grading conversion rules from different countries, regions, and institutions, as well as grading mappings based on internationally recognized equivalence principles. By querying this database, academic scores from different countries or regions can be converted into a unified, standardized grading scale.
[0030] For academic records from different countries (such as a US GPA or a second-class degree from the UK), the system will convert them into standardized scores according to the corresponding conversion rules. For example, if an applicant is from China and receives a score of 85, the system will look up the Chinese grading standard and convert the score into a standardized score compatible with other countries' grading systems (such as the US 4.0 GPA). This ensures that all applicants' academic records can be evaluated and compared under the same standard.
[0031] Extracurricular activities and honors records represent an applicant's non-academic achievements and typically have a significant impact on the success of study abroad applications. To quantify this non-academic data, this invention proposes a quantitative mapping rule based on expert experience, with the specific implementation steps as follows: The quantification rules are defined based on expert experience, considering multiple dimensions such as the duration of extracurricular activities, the importance of the role played, and the level of awards received. The expert team assigns different weights to different types of extracurricular activities based on these dimensions. For example, activities involving leadership roles may receive higher weights, while activities with general participation may receive lower weights. The duration of the activity is also taken into account, with longer-term activities typically receiving higher scores.
[0032] For each extracurricular activity, the system will calculate a weighted average of the weights of each dimension to arrive at a comprehensive quantitative score. For example, a two-year social practice activity will receive a higher score if the participant held a leadership position and won an award, according to predefined quantitative rules; while a short-term volunteer activity may receive a lower score due to its shorter duration.
[0033] The embodiments of this invention employ a database of equivalent performance across multiple education systems and expert-driven quantitative rules for extracurricular activities, thereby ensuring the uniformity, consistency, and comparability of the data and effectively improving the quality of subsequent model training and the accuracy of predictions.
[0034] In one possible implementation, the personal statement is a crucial document in the study abroad application process, reflecting the applicant's academic interests, motivations, and goals. To ensure the system can accurately understand and evaluate this content, this embodiment of the invention uses a dictionary for study abroad applications to scan and match the personal statement text. The specific steps are as follows: Using a pre-defined dictionary for study abroad applications, which contains keywords related to academic fields (e.g., academic interests, research interests, professional fields, etc.), the system scans the personal statement text to find words that match the keywords in the dictionary.
[0035] The system counts the frequency of each keyword in the personal statement and calculates the density of these keywords. The density calculation formula is the ratio of keyword frequency to the total vocabulary of the text, reflecting the strength of the text's relevance to the academic field.
[0036] Finally, a density index reflecting the relevance of the text to the academic field was calculated. This index helps the model quantify the academic strength of the personal statement text and can be used for subsequent machine learning training.
[0037] Letters of recommendation are another important evaluation tool, typically providing an assessment of an applicant's academic abilities, character, and potential in the eyes of others. This invention's method for processing recommendation letters employs a multi-step process involving sentiment analysis and competency dimension matching. The specific steps are as follows: First, sentiment analysis is performed on the recommendation letter texts to filter out statements containing positive evaluations. The goal of sentiment analysis is to distinguish between positive and negative feedback to ensure that only positive feedback is used in subsequent competency assessments.
[0038] Among the selected positive statements, the system will identify and extract words or phrases related to a preset set of competency dimensions. For example, the competency dimension set may include "leadership," "teamwork," and "innovation," and the system will extract relevant words by matching them with descriptions in the recommendation letter.
[0039] The frequency of words appearing in each competency dimension is statistically analyzed to generate a multi-dimensional vector. The value of each dimension represents the strength of the recommender's evaluation of the applicant's competency in that area. For example, if the recommendation letter mentions the applicant's leadership and teamwork skills multiple times, the values for these two dimensions will be high, reflecting the recommender's strong evaluation of these abilities.
[0040] Finally, the system generates a multi-dimensional recommendation letter capability evaluation vector, where the value of each dimension represents the strength of the recommender's evaluation of the applicant in that dimension. This vector is used as the quantitative feature of the recommendation letter and input into the machine learning model.
[0041] The implementation of deep semantic feature extraction steps helps to improve the accuracy and performance of the college application admission probability prediction model by quantitatively analyzing the content of personal statements and recommendation letters. These beneficial effects make this invention have significant practical application value in the field of college application prediction.
[0042] In one possible implementation, to ensure the constructed dictionary is highly representative and comprehensive, this invention first uses web crawling technology to collect raw text data from the official admissions websites of numerous well-known overseas universities, the program introduction pages of various colleges, course descriptions, and program training objective documents. These documents typically contain a large number of academic terms, course content, and subject-specific terminology, comprehensively reflecting the characteristics of the study abroad application field.
[0043] The collected raw text data often contains messy content and redundant information, therefore it needs to be cleaned and processed. The specific steps are as follows: Automated cleaning: Removes irrelevant punctuation marks, HTML tags, advertising content, and other noisy data, retaining only meaningful text.
[0044] Word segmentation: This process divides continuous text into word units for subsequent semantic analysis.
[0045] Part-of-speech tagging: Each word is tagged with its part of speech to distinguish between different parts of speech such as nouns, verbs, and adjectives, so as to extract professional terms more accurately.
[0046] Remove common stop words: Stop words are words that appear frequently in the text but do not have actual meaning (such as "of", "is", "in" etc.). By removing these words, the core words with information are retained.
[0047] By employing a keyword extraction algorithm based on Term Frequency and Inverse Document Frequency (TF-IDF) on the processed text data, high-frequency core words and key phrases with domain-specific differentiation are automatically selected. The basic principle of the TF-IDF algorithm is to select keywords that are both common and distinguishable across different domains by calculating the frequency of words in the entire corpus and their importance in specific documents. These high-frequency words and phrases constitute the initial dictionary.
[0048] While automated algorithms can effectively extract a large number of useful keywords, some terms that may not be extracted during automated processing could be crucial for understanding the domain. Therefore, this invention invites experts in the field of study abroad applications to form a review panel to manually review and optimize the initial dictionary database. The main tasks of the expert panel include: Manual screening and noise reduction: Remove irrelevant, repetitive, or domain-inappropriate words.
[0049] Add important terms: Include terms or phrases that were not covered in the automatic extraction process but are actually very important in the study abroad application. For example, some subject-specific terms may not have been automatically filtered out because of their low frequency, but they are crucial in the academic field.
[0050] Building an authoritative dictionary: The dictionary, reviewed and optimized by experts, contains all the core terms and key phrases in the field of study abroad applications, ultimately forming an authoritative and comprehensive dictionary for the study abroad application field.
[0051] By combining automation technology with expert review, a high-quality and comprehensive dictionary for the study abroad application field was constructed. This greatly improved the accuracy of machine learning models in understanding the semantics of texts when predicting the probability of admission to study abroad applications, and ultimately enhanced the accuracy and reliability of the prediction results.
[0052] In one possible implementation, to identify the core competency dimensions influencing college admissions, a retrospective analysis of a large number of successful application cases is first required. The aim of this analysis is to summarize the soft skills and traits frequently considered in the application process over the years. Specific steps include: Data collection: Collect information such as personal data, letters of recommendation, academic performance, and interview feedback from successful applicants.
[0053] Data analysis: Using techniques such as statistical analysis and text mining, we conduct in-depth analysis of applicants' non-academic backgrounds (such as personal qualities, leadership, teamwork, etc.).
[0054] Summary: The analysis results reveal the core competency dimensions that influence application decisions, such as research potential, academic leadership, teamwork, professional ethics and academic integrity, and cross-cultural communication and expression skills.
[0055] This retrospective analysis, through the mining of big data, helps to build a preliminary framework of capability dimensions.
[0056] In addition to retrospectively analyzing successful cases, this invention supplements the competency dimension set by incorporating the experience of senior admissions officers, study abroad consultants, and academic mentors. These experts typically possess extensive practical admissions experience and can accurately identify which competencies and traits have a significant impact on actual admissions decisions. Expert feedback is collected primarily through interviews and questionnaires to ensure a broad range of opinions are incorporated. Specific steps include: Interview and Questionnaire Design: Design questionnaires and interview outlines covering all aspects of the study abroad application, especially soft skills and personality traits.
[0057] Expert feedback: We collected opinions and views on the applicants' abilities by communicating with senior admissions officers, study abroad consultants, and academic mentors.
[0058] Prioritization: Experts may provide different dimensions based on years of experience, and finally rank them according to their importance and actual impact.
[0059] By combining the retrospective analysis results with expert feedback, a list of capability dimensions from two sources was obtained. To ensure the simplicity and representativeness of the dimension set, it is necessary to merge and deduplicate, removing duplicate dimensions or items with similar content, ultimately forming a comprehensive, accurate, and non-redundant set of capability dimensions.
[0060] The steps involved in this process include: Deduplication: If the same or similar capability dimensions appear in the retrospective analysis and expert feedback, the most representative description is retained.
[0061] Integration of Capability Dimensions: This involves integrating capability dimensions from different sources to ensure that the final capability set encompasses the applicant's multidimensional characteristics in academic, professional, and personal aspects.
[0062] Verify the completeness of the dimensions: Ensure that the set includes key soft skills such as research potential, academic leadership, teamwork, professional ethics and academic integrity, and cross-cultural communication and expression skills.
[0063] Following the above steps, the final set of pre-defined competency dimensions will include several dimensions that accurately reflect an applicant's potential, personality, and overall qualities. These dimensions include, but are not limited to: Research potential: This reflects the applicant's innovative ability in academic research and their potential to promote the advancement of the discipline.
[0064] Academic Leadership: Assess the applicant’s leadership in an academic environment, including guidance and influence within academic teams.
[0065] Teamwork spirit: Demonstrates the applicant's cooperative spirit and organizational and coordination abilities in team work.
[0066] Professional ethics and academic integrity: This reflects the applicant's ability to maintain high standards of ethics and academic integrity in their academic and professional career.
[0067] Cross-cultural communication and expression skills: Assess the applicant's ability to communicate and express themselves in a multicultural environment, especially their adaptability in an international context.
[0068] This invention establishes a comprehensive and accurate set of preset capability dimensions through backtracking analysis, expert feedback, and fusion deduplication, providing high-quality feature support for machine learning prediction of the probability of admission to study abroad applications.
[0069] In one possible implementation, the first step is to scrape data from the official channels of the target universities and programs (such as university websites, admission brochures, course descriptions, and curriculum plans) to obtain the latest information related to the program. Specific content includes: Course Introduction: This section lists the core and elective courses for the major, helping you understand the professional knowledge system.
[0070] Training Program: Describes the program's training objectives, content, and the skills students should possess.
[0071] Admissions requirements: Clearly state the requirements of the institution for the applicant's academic background, abilities, and other qualifications.
[0072] This information provides the raw data source for subsequent text mining and feature extraction.
[0073] Using text mining techniques, a set of noun keywords related to core skills, course themes, and expected student backgrounds relevant to the major was extracted from the courses, curriculum plans, and admission requirements published by the target universities. This process includes: Text preprocessing: Remove irrelevant words, punctuation marks and other noise, segment words and perform part-of-speech tagging.
[0074] Keyword extraction: Using methods such as TF-IDF (Term Frequency-Inverse Document Frequency) and LDA (Latent Dirichlet Allocation), important noun keywords such as "research ability", "teamwork", and "interdisciplinary knowledge" are extracted from the text.
[0075] Set construction: The extracted keyword set is formed into a "keyword set of professional core requirements", which represents the expectations and requirements of the major for students.
[0076] After constructing the core keyword set for the professional requirements, the next step is to calculate the semantic match between the applicant's personal statement and the keyword set. This process involves the following steps: Text vectorization: Using word embedding techniques (such as Word2Vec and GloVe) to convert the applicant's personal statement text and the set of keywords for the core professional requirements into high-dimensional vectors.
[0077] Semantic similarity calculation: The semantic matching degree between the texts is quantified by calculating the cosine similarity between the applicant's personal statement text vector and the set of professional keywords. Cosine similarity can be calculated using the following formula: ; Here, A and B are the vector representations of the two texts, and the cosine value ranges from -1 to 1. The larger the value, the higher the semantic matching degree.
[0078] Matching quantification: Based on the cosine similarity value, the semantic matching degree between the applicant and the target major is obtained, which is used as input for subsequent prediction models.
[0079] Next, we construct dynamic matching features related to academic performance. To this end, we integrate the academic performance data of students admitted to the target major over several recent admission cycles and perform the following processing: The academic performance of admitted students is standardized by calculating standardized scores, for example, by mapping academic performance to a standard normal distribution using Z-scores or T-scores.
[0080] Calculate the statistical distribution of these standardized academic scores, such as the mean and standard deviation.
[0081] For the current applicant's academic performance, calculate their relative percentile or standard score position in the statistical distribution. That is, determine the applicant's relative position among students admitted in previous years.
[0082] By using a linear or nonlinear mapping function (such as linear regression, SVM regression, etc.), the relative percentile or standard score position of academic performance can be transformed into an academic background competitiveness index that is easy for neural networks to process.
[0083] By integrating technologies such as text mining, academic performance standardization, and mapping functions, it is possible to more accurately and comprehensively assess the match between applicants and their target majors, and provide effective feature inputs for machine learning models, thereby improving the accuracy and practicality of predicting the probability of admission to study abroad applications.
[0084] In one possible implementation, the multilayer neural network employs a feedforward neural network structure, the implementation of which includes the following key elements: Input layer dimension: The dimension of the input layer is equal to the total number of all input features. Input features may include the applicant's academic performance, professional fit, semantic fit of the personal statement text, etc. Each feature corresponds to an input node, and the total number of nodes in the input layer equals the number of features.
[0085] Output layer: The output layer uses a sigmoid activation function (such as the sigmoid function) to constrain the output value between 0 and 1. The output result represents the applicant's probability of admission. The sigmoid function is defined as follows: ; in, The output value of the neural network is limited to between 0 and 1 after passing through the Sigmoid activation function, which conforms to the definition of probability value.
[0086] Hidden layer design: The number of hidden layers and the number of neurons in each layer need to be determined through a hyperparameter optimization search process. More hidden layers result in stronger non-linear mapping capabilities, but may also lead to overfitting. Therefore, a systematic hyperparameter optimization process is used to determine the optimal network architecture, enabling the model to achieve a balance between accuracy and generalization ability.
[0087] The training process employs an adaptive moment estimation optimization algorithm (such as the Adam optimizer) to minimize the cross-entropy loss function between the predicted values and the actual admission labels. The specific process is as follows: The cross-entropy loss function is commonly used in binary classification problems to calculate the difference between the predicted probability and the actual label. The formula is: ; in, It is the actual label of the sample (0 or 1). It is the predicted probability of the model. This refers to the number of samples. By minimizing the cross-entropy loss, the model can gradually adjust the weights and biases to improve prediction accuracy.
[0088] The Adam optimizer combines momentum and adaptive learning rate, enabling it to dynamically adjust the learning rate of each parameter during training, thereby accelerating convergence and finding the optimal solution in high-dimensional space.
[0089] To avoid overfitting, **dropout** is used as a regularization strategy during training. The implementation steps of dropout are as follows: During the forward propagation of each training iteration, a certain percentage of the output values of neurons in the hidden layer are randomly and temporarily set to zero. This forces the neural network to learn more robust feature representations, reduces the model's dependence on certain specific neurons, and thus improves the model's generalization ability.
[0090] The dropout ratio (i.e., the proportion of neurons dropped during each training session) is a hyperparameter that can be adjusted through the hyperparameter optimization process. Typically, the dropout ratio is set between 0.2 and 0.5.
[0091] The dropout ratio, along with other hyperparameters (such as the number of hidden layer nodes, learning rate, etc.), is adjusted through a systematic hyperparameter optimization search process (such as grid search or Bayesian optimization) to find the optimal model configuration.
[0092] The training process of neural network models, through carefully designed structures, optimization algorithms, and regularization strategies, can effectively improve the accuracy of predicting the probability of admission to study abroad applications and the generalization ability of the models, ensuring reliable prediction results in practical applications.
[0093] In one possible implementation, during the hyperparameter optimization process, it is first necessary to predefine a hyperparameter space to be optimized. This hyperparameter space includes several key parameters, as follows: The number of hidden layers in a neural network: This hyperparameter determines the depth of the network, i.e., how many hidden layers the neural network contains. A network that is too shallow may result in insufficient model expressive power, while a network that is too deep may increase the difficulty of training and is prone to overfitting.
[0094] The number of neurons in each hidden layer: The number of neurons in each hidden layer determines the width of the model and affects the representational power of each layer. The combination of the number of layers and the number of nodes has a significant impact on the network's expressive power and needs to be optimized.
[0095] Dropout Ratio: The dropout ratio determines the proportion of neurons randomly dropped during training. This hyperparameter controls the regularization strength and affects the model's generalization ability. Typically, this value is between 0.2 and 0.5.
[0096] Learning rate: The learning rate determines the step size at which the model updates weights during training. A learning rate that is too large may lead to instability in the training process, while a learning rate that is too small may result in slow convergence.
[0097] By determining the range of these key hyperparameters, a space containing these hyperparameters is created as the basis for subsequent searches.
[0098] Once the hyperparameter space is determined, grid search or random search strategies can be used to select different combinations of hyperparameters.
[0099] Grid search refers to an exhaustive search across all predefined combinations of hyperparameters, traversing every combination in the hyperparameter space to find the optimal combination. The advantage of grid search is that it guarantees finding the global optimum, but its disadvantage is high computational cost, especially when the hyperparameter space is large.
[0100] Random search trains by randomly selecting combinations of hyperparameters within the hyperparameter space. This method has lower computational cost compared to grid search, especially in high-dimensional hyperparameter spaces, where random search can find a good hyperparameter combination in a shorter time. Recent studies have shown that random search can find near-optimal solutions more efficiently than grid search in most cases.
[0101] For each selected combination of hyperparameters, initialize a corresponding neural network model and begin training. The training process includes: Create and initialize the neural network model based on the selected hyperparameter configuration. The initialization process includes setting parameters such as the number of hidden layers, the number of nodes, the dropout ratio, and the learning rate.
[0102] The neural network is trained using a training dataset. During training, the model performs forward and backward propagation based on the current combination of hyperparameters, and continuously updates the network's weights and biases using optimization algorithms (such as Adam).
[0103] After training, the model is evaluated using an independent validation dataset. The evaluation process calculates performance metrics on the validation set, such as accuracy, AUC, and F1 score. The hyperparameter combination corresponding to the best-performing model is then selected.
[0104] After multiple evaluations using grid search or random search strategies, the hyperparameter combination that performs optimally on the validation dataset will be selected. The model corresponding to this combination will be the final trained model used for actual admission probability prediction.
[0105] The hyperparameter optimization search process in machine learning-based methods for predicting the probability of university admissions can effectively improve model performance, reduce computational costs, and enhance the model's reliability and generalization ability. By optimizing hyperparameters, more accurate prediction results can be achieved, thereby increasing the application value of the prediction system.
[0106] In one possible implementation, to make the predictions of the neural network model interpretable, a post-hoc attribution analysis algorithm is first applied to calculate the gradient or sensitivity of the model's final output (i.e., the probability of university admission) to each input feature. The specific implementation steps are as follows: By differentiating the output of the neural network model, the impact of each input feature on the prediction result is calculated. The gradient value represents how a small change in the input feature affects the prediction result. A larger gradient value indicates that the feature contributes more to the prediction result; conversely, a smaller gradient value indicates that the feature has a smaller impact on the prediction result.
[0107] Attribution analysis methods can further evaluate the importance of each input feature to the final prediction result based on gradient information. For example, model-agnostic attribution algorithms such as LIME (Locally Interpretable Model-Agnostic Approach) or SHAP (Shapley Weighted Algorithm) can be used to locally interpret the model, thereby measuring the contribution of each feature to a specific prediction result. These methods can provide numerical values for the contribution of each input feature to the model's decision, thus helping to identify features that play an important role in the prediction.
[0108] After calculating the gradient or sensitivity of each feature, the next step is to sort these features and output the top few input features with the highest contribution and their corresponding contribution values. The specific steps are as follows: The contributions of all input features are ranked, and the top few features with the highest contributions are selected. The ranking is usually based on the absolute gradient value of each feature or its Shapley value relative to the prediction result.
[0109] The system presents these features and their contribution information to the user. For example, for a specific prediction of the probability of admission to a study abroad application, the system might indicate that "test scores" contribute 30% to the prediction, while "personal statement quality" contributes 25%. In this way, users can clearly understand which factors play the most crucial role in the model's decision-making process.
[0110] The predictive interpretability processing steps, employing post-hoc attribution analysis algorithms, not only help users understand the core decision factors in the model's predictions but also enhance model transparency, strengthen decision support, and provide a basis for subsequent optimization. This is significant in improving user experience and enhancing model credibility.
[0111] In one possible implementation, the ensemble gradient method is a common approach for interpreting complex machine learning models (especially neural networks). Its core idea is to calculate the contribution of each input feature to the model output along the path from the baseline state (representing the initial state with missing information) to the current actual input features. The specific steps are as follows: First, a baseline value needs to be selected for each input feature. This baseline value typically represents a state where information is missing or the feature is irrelevant. For example, for some numerical features, the baseline value can be 0 or the minimum value of the feature; for text or categorical features, the baseline value can be empty or other meaningless values.
[0112] The key to the ensemble gradient method lies in the smooth transition of feature values from baseline to actual feature values. During computation, the system gradually changes the values of the input features along this path, starting from the baseline and progressively approaching the current feature values of the actual applicant. This change is gradual, with each feature value divided into multiple small steps using a specified step size to ensure a smooth and continuous transition.
[0113] Within each small step, the model calculates the corresponding output gradient, which represents the local sensitivity of each input feature to the output. Then, these gradients are integrated along the entire change path (from the baseline to the actual feature value) to obtain the cumulative contribution of that feature to the final prediction. Specifically, the ensemble gradient method calculates the sum of the products of the gradient of the feature value at each small step and the change in the feature value, i.e.: Contribution ; After calculating the contribution of each feature, normalization is performed. The purpose of normalization is to ensure that the sum of the absolute values of the contributions of all features is 1, thus obtaining the relative percentage contribution of each feature. The normalization formula is: Normalized contribution ; Through this process, the contribution of all features is adjusted to between 0 and 1, and their sum is 1, making it easy to directly compare the relative impact of each feature on the prediction result.
[0114] Finally, the system outputs the relative percentage contribution of all features, showing the impact of each feature on the prediction of the probability of admission to a study abroad application. For example, if the contribution of the "university ranking" feature is 0.4, then it contributes 40% to the final prediction result; if the contribution of the "language score" feature is 0.3, then it contributes 30%, and so on.
[0115] By providing a quantitative contribution of each input feature to the final prediction result, it provides strong support for the interpretability and transparency of machine learning models, thereby improving the credibility, decision support, and compliance of the models.
[0116] The following examples will illustrate this in detail: This invention is applied to the field of study abroad applications, aiming to help students assess their probability of being admitted to their target universities by predicting their application data. This invention uses an ensemble gradient method to interpret the machine learning model, helping students understand the key features that influence their admission probability.
[0117] The prediction system of the present invention consists of the following parts: Input data: student's personal information, academic performance, standardized test scores, quality of recommendation letters, content of personal statement, etc.
[0118] Academic performance (GPA): A real number between [0.0, 4.0] (in this embodiment of the invention, GPA is 3.5).
[0119] TOEFL score: range from 0 to 120 (in this embodiment of the invention, the score is 100).
[0120] GRE score: ranges from 260 to 340 (in this embodiment of the invention, the score is 320).
[0121] Recommendation letter rating: 1 to 5 points (in this embodiment of the invention, the rating is 4).
[0122] Personal statement rating: 1 to 5 (in this embodiment of the invention, the rating is 5).
[0123] Output data: the predicted admission probability (a real value), with a range of [0,1].
[0124] First, the raw data is standardized to ensure that data of different dimensions can be effectively compared and calculated in the model: Standardization method: Use Z-score standardization method: ; in, The original data, The mean of the features, The standard deviation is denoted as . This standardization process is performed for all input features (such as GPA, GRE, etc.).
[0125] In this embodiment of the invention, random forest was chosen as the prediction model because it performs well in multi-feature learning tasks and can handle high-dimensional data.
[0126] Model parameters: Number of decision trees: 100; Maximum depth: No maximum depth limit (using default settings); Minimum number of sample splits: 2.
[0127] Training process: The dataset is used 80% for training and 20% for testing; Use cross-validation to select the optimal hyperparameters; The training process uses the default training algorithm of random forest.
[0128] To interpret the results of the prediction model, the feature contributions are calculated using the ensemble gradient method (IG). The steps of the ensemble gradient method are as follows: Select the minimum value of the feature as the baseline: The baseline for GPA is 0.0; The baseline for TOEFL is 0; The baseline for the GRE is 260; The baseline score for recommendation letters is 1. The baseline for the personal statement is 1.
[0129] When calculating the contribution of each feature, the gradient is computed along the path from the baseline to the actual input. The formula for the ensemble gradient is: ; in, It is the model's prediction function. It is a baseline feature. These are input features. Samples are taken between [0,1].
[0130] Calculate the gradient for each feature (such as GPA, GRE, etc.) to reflect its impact on the final predicted value. Integrating the gradient yields the overall contribution of the feature. Normalize the contribution of each feature so that the sum of the contributions of all features is 1, that is: ; This step ensures that the contribution of each feature can be compared as a percentage.
[0131] In this embodiment of the invention, the predicted admission probability of a student is obtained as 0.75 through training. The contribution of each feature is analyzed using the ensemble gradient method:
[0132] Compared with the traditional SHAP method, the SHAP method yields the following feature contributions: GPA: 28%; GRE: 22%; TOEFL: 12%; Recommendation letter rating: 8%; Personal statement score: 7%; The comparison shows that the ensemble gradient method provides a more intuitive contribution score in terms of interpretability, while the explanation provided by the SHAP method is relatively complex and difficult to understand.
[0133] Using a publicly available study abroad application dataset (in this embodiment of the invention, the dataset contains data from 1000 students), 10 cross-validations were performed, and the results were evaluated using standard precision, recall, and F1 score.
[0134] Experimental results: Accuracy: 0.82; Recall rate: 0.75; F1 score: 0.78.
[0135] Compared with the traditional logistic regression model (accuracy 0.75, F1 score 0.72), the prediction model of this invention has higher accuracy and stronger interpretability. In particular, under the contribution analysis of the integrated gradient method, students can clearly see which features have the greatest impact on admission results.
[0136] By using the ensemble gradient method, this invention provides a more accurate analysis of feature contribution, helping students understand which features have a decisive impact on their admission results. Compared to traditional models, this invention can effectively explain the black-box decision-making process in complex models (such as random forests), thereby increasing students' confidence in the model results.
[0137] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0138] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A machine learning-based method for predicting the probability of admission to study abroad applications, characterized in that, Includes the following steps: Data collection and standardization steps: Collect heterogeneous data from multiple sources from applicants, and perform domain-specific standardization processing on the non-standardized structured data to generate standardized numerical features; The deep semantic feature extraction steps are as follows: Natural language processing is performed on the unstructured text data in the application documents, and quantitative semantic features are extracted using a pre-built dictionary of study abroad application fields and preset ability dimensions. Dynamic matching degree feature construction steps: Based on the specific requirements of the target university's major, the processed applicant features are compared and calculated with the specific requirements to generate dynamic matching degree features that characterize the fit between the applicant and the major; Neural network model processing steps: The standardized numerical features, the quantized semantic features, and the dynamic matching degree features are input into a specially trained multi-layer neural network model, which outputs a continuous value representing the admission probability. The specialized training refers to using historical study abroad application data and their corresponding admission results to train the multi-layer neural network model in a supervised learning manner, and employing regularization strategies during the training process to improve the model's generalization ability.
2. The method for predicting the probability of admission to study abroad applications based on machine learning according to claim 1, characterized in that, The domain-specific standardization process for non-standardized structured data in the data acquisition and standardization steps includes: For academic achievements from different countries or regions, based on the official grade conversion standards or recognized equivalence principles of their respective education systems, a pre-established multi-education system equivalent grade comparison database is queried to map various grade systems or percentage system grades to a unified standardized scoring scale, thereby obtaining standardized academic achievement values. For extracurricular activities and honors, a comprehensive quantitative value for extracurricular activities is generated by assigning weights to different dimensions of each project and performing weighted calculations based on the duration of the activities, the importance of the role played, and the level of the awards. This is done through a predefined quantitative mapping rule based on expert experience.
3. The method for predicting the probability of admission to study abroad applications based on machine learning according to claim 1, characterized in that, The text deep semantic feature extraction step includes extracting quantized semantic features, including: For the personal statement text, the dictionary for the study abroad application field is used for scanning and matching. The frequency and density of words in the text that match the keywords in the dictionary are counted and correlated with the total vocabulary of the text to calculate a keyword density index for the personal statement field that reflects the relevance of the document content to the academic field. For the recommendation letter text, sentiment analysis is first performed to filter out statements containing positive evaluations. Then, words or phrases related to the core abilities listed in the preset ability dimension set are identified and extracted from these statements. The frequency of occurrence of words representing different abilities is counted, and finally a multi-dimensional vector is formed, where the value of each dimension represents the strength of the recommender's evaluation of the applicant's ability in that aspect, i.e., the multi-dimensional recommendation letter ability evaluation vector.
4. The method for predicting the probability of admission to study abroad applications based on machine learning according to claim 3, characterized in that, The construction process of the pre-built dictionary for study abroad application fields includes: Using web crawling technology, raw text data was collected from the official admissions websites of a large number of well-known overseas universities, the professional introduction pages of various colleges, course descriptions and program training objectives documents; The collected raw text data is automatically cleaned, segmented, labeled with parts of speech, and has common stop words removed. A keyword extraction algorithm based on word frequency and inverse document frequency is used to automatically filter out high-frequency core words and key phrases with domain-specific characteristics from the processed text to form an initial dictionary. An expert review panel was invited to manually review, screen, remove noise, and supplement the initial dictionary database, adding relevant but not automatically extracted terms, ultimately forming an authoritative and comprehensive dictionary for the field of study abroad applications.
5. The method for predicting the probability of admission to study abroad applications based on machine learning according to claim 3, characterized in that, The process of establishing the preset capability dimension set includes: By retrospectively analyzing a large number of successful study abroad application cases, we have summarized a series of soft skills and traits that are frequently considered in admission decisions. It draws on the experience of senior admissions officers, study abroad consultants, and academic mentors, and collects what they consider to be the most important dimensions of applicants' abilities through interviews and questionnaires. The ability lists obtained from the above two aspects are merged and deduplicated to form a set of preset ability dimensions that include research potential, academic leadership, teamwork spirit, professional ethics and academic integrity, and cross-cultural communication and expression skills.
6. The method for predicting the probability of admission to study abroad applications based on machine learning according to claim 1, characterized in that, The dynamic matching degree feature construction steps include: Obtain the latest course introductions, training programs and admission requirements from the official channels of the target universities and target majors, and use text mining technology to extract a set of noun keywords that describe the core skills, course themes and expected student backgrounds of the major, thus forming a set of keywords for the core requirements of the major. The semantic similarity between the applicant's personal statement text and the set of keywords for the core requirements of the major is calculated. This similarity is quantified by calculating the cosine of the angle between the two texts in a high-dimensional vector space, thereby obtaining the semantic matching degree of the major direction. By integrating the academic performance data of students admitted to the target major in the past few admission cycles, the statistical distribution of their standardized academic performance values is calculated. Then, the relative percentile or standard score position of the current applicant's standardized academic performance value in the statistical distribution is calculated. This position information is then converted into an academic background competitiveness index that is easy for neural network models to process through a linear or nonlinear mapping function.
7. The method for predicting the probability of admission to study abroad applications based on machine learning according to claim 1, characterized in that, The specialized training in the neural network model processing steps includes: The multilayer neural network model adopts a feedforward neural network structure. The dimension of its input layer is equal to the total number of all input features. The output layer uses a sigmoid activation function to limit the output value between 0 and 1 to represent the probability. The number of hidden layers and the number of neurons in each layer are determined through a systematic hyperparameter optimization search process. The training process uses an adaptive moment estimation optimization algorithm to minimize the cross-entropy loss function between the predicted value and the actual admission label. The regularization strategy involves applying a dropout method to the output of the hidden layer during the training phase. Specifically, during the forward propagation of each training iteration, a certain proportion of the output values of neurons in the hidden layer are randomly set to zero. This dropout proportion is used as a hyperparameter and is determined together with other structural hyperparameters of the model through the hyperparameter optimization search process.
8. The method for predicting the probability of admission to study abroad applications based on machine learning according to claim 7, characterized in that, The systematic hyperparameter optimization search process employs either a grid search or a random search strategy: A hyperparameter space to be optimized is predefined, which includes, but is not limited to: the number of hidden layers in the neural network, the number of neurons in each hidden layer, the dropout ratio of the dropout method, and the learning rate; Select multiple different combinations of hyperparameters from this hyperparameter space according to a predetermined strategy; For each set of hyperparameters, initialize a corresponding neural network model, train it using the training dataset, and then evaluate its performance metrics on an independent validation dataset. The model structure and parameters corresponding to the hyperparameter combination that achieves the best performance on the validation dataset are selected as the final trained model.
9. The method for predicting the probability of admission to study abroad applications based on machine learning according to claim 1, characterized in that, Following the neural network model processing step, the following is also included: The steps for interpretability processing of prediction results are as follows: Apply a post-hoc attribution analysis algorithm to calculate the gradient or sensitivity of the final output of the neural network model to each input feature, and thereby evaluate the importance of each input feature to the specific prediction result. The system outputs the top few input features with the highest importance and their corresponding contribution values, thereby providing users with an explanation of the key decision factors for predicting the admission probability.
10. The method for predicting the probability of admission to study abroad applications based on machine learning according to claim 9, characterized in that, The post-hoc attribution analysis algorithm employs the ensemble gradient method. The input features to the model are smoothly transformed from a set of baseline values representing the missing information state to the actual feature values of the current applicant, and the gradient of the model output is integrated along this path; The cumulative sum of the products of the average gradient change of each input feature along this path and the change of its feature value is calculated. This cumulative sum is the estimated contribution of the feature to the final prediction result. The contribution estimates of all features are normalized so that the sum of the absolute values of the contributions of all features is one, thus obtaining the relative contribution percentage of each feature.