Intelligent recommendation method for professional interests of universities of senior high school students based on big data analysis

Through big data analysis, combined with academic data, behavioral logs and questionnaire surveys, a subject selection rule database and professional ability matrix are built, which solves the problem of single data dimensions in traditional recommendation systems and achieves smarter and more comprehensive talent recommendations.

CN120492731APending Publication Date: 2025-08-15SHANGRAO NORMAL UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510598978.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-10
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing high school college major selection recommendation system ignores students' abilities and interests, has a single data dimension, dynamic recommendation mechanisms, and disconnected career development.

Method used

Based on big data analysis, students' feature vectors are extracted through academic data, student external behavior logs and questionnaire surveys, and subject selection rules database, professional ability matrix and industry demand forecast reports are constructed to make intelligent recommendations.

Benefits of technology

It has achieved comprehensive considerations of academic ability, interest characteristics and industry trends, and provided full-chain innovation covering data collection, feature analysis and effect feedback, which has enhanced the intelligence and practical value of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492731A_ABST
    Figure CN120492731A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent recommendation method for university specialty interests of senior high school students based on big data analysis, and relates to the technical field of university specialty recommendation, and the method comprises the steps: carrying out academic ability analysis and interest feature extraction based on academic data, student external behavior log data and questionnaire survey results, and finally obtaining student feature vectors; and based on college and university department selection requirement data, student feature vectors, professional course outline data and recruitment post data, performing department selection rule base construction, professional ability modeling and industry trend prediction to obtain a department selection rule base, a professional ability matrix and an industry demand prediction report. And performing recommendation matching based on the student feature vector, the subject selection rule base, the professional ability matrix and the industry demand prediction report to obtain a recommendation list. Based on the thinking analysis model, professional recommendation is carried out from student academic ability, student interests, professional features and industry trends, and the problem that a traditional method is single in recommendation dimension is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of university major recommendation, and in particular to an intelligent recommendation method for high school students' university major interests based on big data analysis. Background Art

[0002] Recommendations for high school students' college majors have been very popular in recent years. However, existing recommendation systems for high school students' college majors often only make recommendations from the perspective of employment and score lines, ignoring students' abilities and interests. At the same time, they do not conduct research on the long-term development of students in the industry. There are problems such as a single data dimension, a dynamic recommendation mechanism, and a disconnect between career development.

[0003] Therefore, it is very necessary to design an intelligent recommendation method for high school students' university major interests based on big data analysis. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide an intelligent recommendation method for high school students' university major interests based on big data analysis.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] The present invention provides an intelligent recommendation method for high school students' university major interests based on big data analysis, comprising:

[0007] Step 1: Based on academic data, students' external behavior log data, and questionnaire survey results, academic ability analysis and interest feature extraction are performed to obtain student feature vectors.

[0008] Step 2: Based on university subject selection requirement data, student feature vectors, professional course syllabus data, and recruitment position data, a subject selection rule base is constructed, professional capability modeling is performed, and industry trend forecasting is performed to obtain a subject selection rule base, a professional capability matrix, and an industry demand forecast report;

[0009] Step 3: Perform recommendation matching based on student feature vectors, subject selection rule base, professional ability matrix and industry demand forecast report to obtain a recommendation list.

[0010] Preferably, in step 1, academic ability analysis is performed based on academic data, specifically:

[0011] Collect electronic versions of textbooks, syllabi, past exam questions, and expert annotation data, and pre-process them;

[0012] Based on natural language processing, knowledge points are extracted and entities are recognized on the pre-processed data. Rule matching is performed on the processed data to obtain a list of knowledge points.

[0013] Determine the relationship types of the knowledge point list, annotate the relationships, and build a graph database to obtain a knowledge point-ability mapping table;

[0014] A model is constructed based on the multidimensional item response theory, taking the knowledge point-ability mapping table and student answer records as input, and parameter estimation is performed based on the model to obtain the student ability vector and diagnostic report;

[0015] Based on the historical student ability vectors, a time series ability analysis is performed based on the preset model to obtain a trend chart of student ability changes.

[0016] Preferably, in step 1, interest feature extraction is performed based on the questionnaire survey results, specifically:

[0017] Obtain questionnaire survey results and external behavior log data;

[0018] Conduct explicit interest processing based on student ability vectors and questionnaire data;

[0019] Mining hidden interests based on students’ ability vectors and external behavior log data;

[0020] Fusing the explicit interest processing results and the implicit interest mining results to obtain the fused interest results;

[0021] Based on the fusion of interest results and the trend diagram of student ability changes, interest evolution prediction is performed through a preset model to obtain the interest evolution prediction results.

[0022] Preferably, in step 2, a subject selection rule base is constructed based on the subject selection requirement data of colleges and universities, specifically:

[0023] Obtain the enrollment plans of various universities in previous years after the new college entrance examination, extract the subject selection requirements of each university, conduct rule analysis, perform structured storage based on the analysis results, and generate a subject selection rule library.

[0024] Preferably, in step 2, professional competence modeling is performed based on student feature vectors and professional course syllabus data, specifically as follows:

[0025] Based on Stanford CoreNLP, we extract competency keywords from the professional course outline data, construct a professional competency requirement vector, and then establish a competency dimension mapping table based on the professional competency requirement vector and the student competency vector in the student feature vector, and finally output the professional competency matrix.

[0026] Preferably, in step 2, industry trend prediction is performed based on student feature vectors and recruitment position data, specifically:

[0027] Based on the recruitment data, the job ability requirements are extracted based on the LDA subject model, and the job ability requirement trend is predicted based on the ARIMA model to obtain the industry demand trend of the position. The interest evolution prediction results in the student feature vector are cross-validated with the industry demand trend to obtain an industry demand forecast report.

[0028] Preferably, in step 3, recommendation matching is performed based on the student feature vector and the major feature matrix to obtain a recommendation list, specifically:

[0029] Obtain students' subject selection information and conduct preliminary subject selection screening based on the subject selection information and subject selection rule library;

[0030] Based on the professional ability matrix, student ability vector and integrated interest results, a multi-dimensional matching engine is used to match, calculate the comprehensive score of each major, sort them according to the comprehensive score, and output a recommendation list;

[0031] Based on the diagnostic report and industry demand forecast report, determine whether there are warning prompts for the majors in the recommended list.

[0032] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0033] The present invention provides a method for intelligently recommending high school students' university major interests based on big data analysis. The method comprises: performing academic ability analysis and interest feature extraction based on academic data, student external behavior log data, and questionnaire survey results, ultimately obtaining student feature vectors; constructing a subject selection rule base, modeling professional capabilities, and predicting industry trends based on university subject selection requirement data, student feature vectors, professional course syllabus data, and recruitment position data, obtaining a subject selection rule base, a professional capability matrix, and an industry demand forecast report; performing recommendation matching based on the student feature vectors, the subject selection rule base, the professional capability matrix, and the industry demand forecast report, and obtaining a recommendation list. The present invention is based on a thinking analysis model and recommends majors based on student academic ability, student interests, professional characteristics, and industry trends. It solves the problem of the single recommendation dimension of traditional methods and forms a full-chain innovation covering data collection, feature analysis, intelligent recommendation, and effect feedback, with both theoretical breakthroughs and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 A flow chart of a method provided by an embodiment of the present invention;

[0036] Figure 2 Schematic diagram of the LSTM-MLP model structure;

[0037] Figure 3 Schematic diagram of the local TCN structure;

[0038] Figure 4 Schematic diagram of the global TCN structure. DETAILED DESCRIPTION

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0040] The purpose of this invention is to provide an intelligent recommendation method for high school students' university major interests based on big data analysis, which can solve the problem of the single recommendation dimension of traditional methods, form a full-chain innovation covering data collection, feature analysis, intelligent recommendation, and effect feedback, and has both theoretical breakthroughs and practical application value.

[0041] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0042] Figure 1 A flow chart of the method provided in the embodiment of the present invention is shown in FIG. Figure 1 As shown, the present invention provides an intelligent recommendation method for high school students' university major interests based on big data analysis, including:

[0043] Step 1: Based on academic data, students' external behavior log data, and questionnaire survey results, academic ability analysis and interest feature extraction are performed to ultimately obtain a student feature vector. The student feature vector includes the student ability vector, diagnostic report, student ability change trend chart, fusion interest results, and interest evolution prediction results.

[0044] Step 2: Based on university subject selection requirement data, student feature vectors, professional course syllabus data, and recruitment position data, a subject selection rule base is constructed, professional capability modeling is performed, and industry trend forecasting is performed to obtain a subject selection rule base, a professional capability matrix, and an industry demand forecast report;

[0045] Step 3: Perform recommendation matching based on student feature vectors, subject selection rule base, professional ability matrix and industry demand forecast report to obtain a recommendation list.

[0046] In step 1, academic ability analysis is performed based on academic data, specifically:

[0047] Collect electronic versions of textbooks, teaching syllabuses, past exam questions, and expert annotation data, and pre-process them, specifically:

[0048] First, the data is introduced. The electronic version of the textbook includes the local corresponding version of the textbook PDF / Word document;

[0049] The syllabus is the subject teaching requirements issued by the Ministry of Education;

[0050] Previous exam questions include the college entrance examination questions and answers from the past ten years;

[0051] Expert annotation data is a knowledge point relationship table provided by education experts;

[0052] The preprocessing steps include:

[0053] 1. Text extraction: Use OCR tools (such as Tesseract) to convert PDF textbooks into editable text;

[0054] 2. Structural analysis: parse the textbook catalog, extract chapter titles and sub-chapter titles, and form a preliminary list of knowledge points;

[0055] 3. Data cleaning: remove irrelevant content (such as exercise explanations and illustrations) and retain the description of core knowledge points;

[0056] Based on natural language processing, knowledge points are extracted and entities are recognized on the pre-processed data. Rule matching is performed on the processed data to obtain a list of knowledge points, specifically:

[0057] Use pre-trained models (such as BERT) to perform semantic analysis on the pre-processed data to identify the core concepts in the text;

[0058] Formulate regular expressions to match common knowledge point patterns, such as "definition of XX" and "properties of XX";

[0059] Determine the relationship types of the knowledge point list (for example: forward and backward dependency: the order of learning knowledge points, such as "limit" → "derivative"; association relationship: the co-occurrence of knowledge points in problem solving, such as "trigonometric function" and "geometric proof"; difficulty level: the hierarchical relationship of basic → advanced → extended). Annotate the relationships, such as through automated analysis and statistical analysis of the frequency of co-occurrence of knowledge points in test questions and infer dependencies based on chapter nesting relationships. Alternatively, manual annotation can be performed to ultimately obtain a knowledge point graph database.

[0060] Based on the knowledge point graph database, the knowledge point-ability mapping table is obtained, specifically:

[0061] First, we define the core competencies of a subject. Different subjects have different abilities. Taking mathematics as an example, we define the six core competencies of mathematics: mathematical abstraction, logical reasoning, mathematical modeling, intuitive imagination, mathematical operations, and data analysis.

[0062] The present invention provides a first embodiment, which uses the expert annotation method and the Delphi method to perform multiple rounds of annotation, and determines the subject core literacy corresponding to each knowledge point based on the knowledge point graph database, and finally obtains a knowledge point-competency mapping table;

[0063] The present invention also provides another embodiment, which is implemented through a data-driven method, for example, by extracting a triple of question-knowledge point-ability, and constructing a knowledge point-ability mapping table from the question;

[0064] Other methods may also be used, and the present invention does not limit them, as long as the structured data of knowledge points and abilities can be obtained in the end;

[0065] Obtain students' subject scores and wrong question records;

[0066] The model is constructed based on the multidimensional item response theory, where the model formula is:

[0067]

[0068] Where a ik is the discrimination of question i on ability dimension k, b i The difficulty parameter of question i, θ jk is the ability value of student j on ability dimension k;

[0069] Take the knowledge point-ability mapping table and student answer records as input, perform parameter estimation based on the model, and construct a training matrix, for example:

[0070]

[0071] The solution is based on the EM algorithm, and the student ability vector and diagnosis report are finally output. An example of the student ability vector is:

[0072]

[0073]

[0074] Obtain the historical sequence of students' ability vectors, perform time-series ability analysis based on the preset model, and obtain a trend chart of students' ability changes;

[0075] Among them, the preset model is the improved LSTM model, which is introduced in detail as follows:

[0076] The improved LSTM model is specifically the LSTM-MLP model, and its structural diagram is as follows Figure 2 As shown in Figure 1, the model consists of 1 LSTM layer and 3 MLP layers. The performance of the LSTM combined with MLP is tested with different hyperparameters such as the number of units in each layer, the type of activation function, and the total number of iterations (Table 1). The root mean square error (RMSE), mean absolute error (MAE), and determination coefficient R are introduced. 2 The evaluation index evaluates the LSTM-MLP model with different hyperparameters, and the expression is as follows:

[0077]

[0078] Since there are some extreme outliers in the data sequence, the prediction model is more sensitive to outliers. MAE only considers the absolute value of the error and is more robust. If the model is more sensitive to the overall error size, RMSE is more appropriate. 2 As a commonly used indicator of the prediction effect of deep learning models, it can quantify the prediction accuracy more intuitively.

[0079] Table 1 Different hyperparameters of LSTM-MLP model structure

[0080]

[0081] In step 1, interest features are extracted based on the questionnaire survey results, specifically:

[0082] Obtaining questionnaire survey results and external behavior log data, including the original Holland Career Interest Test answer sheet and customized additional questions on technical preferences, and external behavior log data including students' MOOC learning records and social media behavior (e.g., watching a brief history of quantum physics on a short video platform).

[0083] Based on the student ability vector and questionnaire survey data, explicit interest processing is carried out, specifically:

[0084] The scale was enhanced to add a new dimension of technical propensity, T, and the scoring formula was adjusted to:

[0085]

[0086] Where, α grade Adjusting the coefficients for grade level, the output example is: "Interest Hexagon": {"R": 0.32, "I": 0.85, "A": 0.17, "S": 0.39, "E": 0.28, "T": 0.72}, "Technology Sensitivity": 0.68;

[0087] Based on the student ability vector and external behavior log data, hidden interest mining is carried out, specifically:

[0088] Perform text vectorization and interest intensity calculation on external behavior log data, outputting interest words and related majors, as well as the corresponding relevance;

[0089] Fusing the explicit interest processing results and the implicit interest mining results to obtain the fused interest results;

[0090] Based on the fusion of interest results and the student ability change trend chart, the interest evolution prediction is carried out through the preset model to obtain the interest evolution prediction result;

[0091] Here, the improved TCN model can be used to predict interest evolution, which is introduced in detail as follows:

[0092] The main improvements of the present invention over the traditional TCN model are:

[0093] 1. TCN's causal convolution and dilated convolution correspond to local spatial information and global temporal information respectively, capturing the temporal and spatial correlation of different time series features and improving the efficiency of time series data information extraction;

[0094] 2. Adopt convolutional autoencoder (CAE) encoding and decoding to enhance the robustness of the model after extracting feature information;

[0095] 3. By introducing joint optimization, we combine the advantages of prediction-based models and reconstruction-based models to enhance the accuracy and interpretability of model outlier predictions;

[0096] The specific introduction is:

[0097] In view of the lack of traditional TCN in capturing the dependencies between long-term and short-term time series data, the original TCN is decomposed into two parts: local TCN and global TCN, which can enhance the success rate of capturing abnormal time series data. For example, the input time series x∈R n×k , where n is the maximum length of the timestamp and k is the number of input features. The basic calculation process of a single layer is as follows:

[0098] U=Weight-Normalization(ConvD(x))

[0099] F = Dropout(LeakyRelu(U))

[0100] LeakyRelu=max(0,U)+Leakmin(0,U);

[0101] Where x={x1 (i), x2 (i) ,…,x T (i) There are d model channels, U is the output after convolution and weight normalization, and F is the final output of U optimized by LeakyRelu and Dropout. From the formula of LeakyRelu, we can see that leakage is a very small constant.

[0102] Next, we will introduce the local and global TCN modules.

[0103] Local TCN module: TCN causal convolution is used to extract spatial information as a local time series data processing model. Figure 3 As shown, for the input time series x∈R n×k , in causal convolution, it is based on x1…x T and y1…y T-1 To predict y T , so that y T Close to the actual value;

[0104] The output of each layer in the figure is obtained by combining the input of the previous layer and the input of the previous position. Causal convolution cannot see future data. It is only a one-way structure and a strict time-constrained model. Therefore, it can avoid the leakage of future information and ensure the stability of the model and the integrity of the data.

[0105] The causal convolution is applied to the local information convolution part, and only two adjacent time steps are considered. The convolution kernel size is set to 2 in order to capture the spatial information features. At the same time, in order to make the output time step of the convolution operation consistent with the input time step, zero padding is added on the left side of the convolution (the padding size is: kernel size - 1). Assuming that the input sequence is x = {x1…x T}, the output sequence is y, then the formula of the local TCN model can be expressed as:

[0106]

[0107] Where * represents the convolution operation; t represents the time step of the output time; y[t] is the value of the output sequence at time step t; x[t-k+p] is the value of the input sequence after local convolution offset; w[k] is the weight of the convolution kernel; k represents the size of the convolution kernel. By changing the k value, the convolution kernel can capture feature information at different positions in the input data, thereby realizing feature extraction and representation. p is the number of zero paddings.

[0108] Introduction to global TCN: For the entire global time series, dilated convolution is used to capture global information of the entire time series, such as Figure 4As shown in the figure, the dilated convolution is equivalent to a skip filter, in which each layer is expanded to achieve an exponential expansion of the receptive field. Since the receptive field area is increased, the detection efficiency of the entire time series data is improved;

[0109] Since TCN usually uses one or more convolutional layers as its initial layer, for example, the input sequence is x = {x1…x T}, the convolution kernel is k, the convolution kernel weight is w[k], x[td×k] is the value of the input sequence after global offset; the output sequence is y, the dilation rate is d (d=1, 2, 4), and the formula of dilated convolution can be expressed as:

[0110]

[0111] Where d represents the interval between the values in the convolution kernel. A larger expansion rate leads to a larger receptive field, which can capture more distant dependencies between input sequences. The other parameters have the same meaning as in the local TCN.

[0112] In the global TCN, each time step of the output sequence of this method is obtained by weighted summing the value of the input sequence at different expansion rates with the weight of the convolution kernel. The dilated convolution allows the convolution kernel to span more time steps to capture long-range dependencies, which helps the model effectively capture global temporal information.

[0113] By decomposing TCN, it is possible to capture information correlation in the local and global spatiotemporal domains of the time series.

[0114] After the improved model extracts features from the time series data in time and space, the encoder maps the input data to a lower dimension to represent it, reducing the complexity and computational cost of the model. This method encodes the local TCN and global TCN on spatial features and temporal features respectively, enabling the model to filter out noise and unnecessary information in the input data during the encoding phase. Therefore, the present invention uses a CAE autoencoder for encoding and decoding. The CAE autoencoder will not be introduced in detail in this invention.

[0115] Regarding joint optimization, we first include a prediction-based module. In this module, the fully connected layer is used as a prediction-based model, and the root mean square error is used as the loss function.

[0116] The main purpose of the reconstruction-based module is to extract useful information from the data. The reconstruction model of this method processes the data in batches through a sliding window and contains an RNN encoder. The GRU layer converts the input time series into a fixed-dimensional encoding representation. This encoding representation can then be used to generate a reconstructed target sequence. Therefore, the model learns how to reconstruct normal data during training. For data points that do not match the normal pattern, the reconstruction error is usually large, which can be used for anomaly detection and form multi-task learning with the prediction model, thereby reducing the need for a large amount of labeled data, reducing overfitting of tasks, and improving data efficiency. If the input feature is x, then the feature will satisfy p(x|g), where g∈R d2 Represents the value in the reconstructed latent space, then p(x|g) can represent the probability of observing feature x in the reconstructed space;

[0117] The goal of joint optimization is to find the optimal model parameters of the data distribution reconstruction value x' that is closest to the input feature x, so as to determine the location of abnormal data points or data segments. The posterior density of the true data distribution is: p(g|x)=p(x|g)p(g)p(x), which means that after observing the input feature x, the reconstruction space is true; the boundary data distribution density, that is, the probability that the data is distributed at the edge of the reconstruction space, is expressed as: p(x)=∫p(g)p(x|g)dg. Based on the above information, this method uses negative log-likelihood loss (Negative Log-Likelihood Loss) and KL divergence (Kullback-Leibler divergence) to calculate the corresponding loss of the reconstruction model;

[0118] Corresponding to the goal of joint optimization, at each timestamp, the input time series information x has not only its actual value but also an evaluation value based on prediction and reconstruction. For example, in the obtained time series feature data, the anomaly threshold of the entire time series data is first set, and then the anomaly score S of each feature is calculated. i ,If it exceeds the abnormal threshold, it is an abnormal point or abnormal segment.,This method uses the Epslion function to define a dynamic threshold,which is related to the mean and standard deviation of the error;

[0119] Epsilon=Mean(errors)+z×StandardDeviation(errors);

[0120] Where z is a loop variable used to cyclically calculate different Epsilon anomaly assessment thresholds within a certain range. The formula for calculating the anomaly score of each feature is:

[0121]

[0122] Where μ is used to control and adjust the weight between prediction and reconstruction, and k is the total number of all features. is the predicted value With the actual value x i The square error between the actual value of feature i and the predicted value indicates the degree of deviation, (p(x i )-x i ) 2 It represents the square error between the reconstructed value and the actual value, and indicates the probability of feature i encountering an outlier in the reconstructed model.

[0123] In step 2, based on the college subject selection requirements data, a subject selection rule base is constructed, specifically:

[0124] Obtain the enrollment plans of various universities over the years after the new college entrance examination, extract the subject selection requirements of each university, analyze the rules, store the analysis results in a structured manner, and generate a subject selection rule library;

[0125] You can also refer to local policies and the subject selection requirements of various universities. For example, refer to the Ministry of Education's "Guidelines for Elective Subject Requirements for Undergraduate Admissions in Regular Colleges and Universities" and the new college entrance examination policy documents of various provinces, etc., and set them according to specific needs;

[0126] The present invention provides an example:

[0127]

[0128] In step 2, professional competence modeling is performed based on student feature vectors and professional course syllabus data, specifically:

[0129] Based on Stanford CoreNLP, we extract the ability keywords from the professional course outline data (for example, "data structure": ["logical thinking", "algorithm design"], "analog circuit": ["physical modeling", "experimental design"]), and construct the professional ability requirement vector:

[0130] R major =∑ course ω course ·C course

[0131] Where, ω course is the proportion of course credits, C course For course competency requirements;

[0132] Based on the professional ability requirement vector and the student ability vector in the student feature vector, a capability dimension mapping table is established, and the professional ability matrix is finally output. An example is as follows:

[0133] "computer":{

[0134] "Ability requirements": {"Logical thinking": 0.95, "Algorithm design": 0.88},

[0135] "Associated Courses": ["Data Structures", "Operating Systems"].

[0136] In step 2, industry trend prediction is performed based on student feature vectors and recruitment position data, specifically:

[0137] Based on the recruitment data, the job ability requirements are extracted based on the LDA subject model, and the job ability requirement trend is predicted based on the ARIMA model to obtain the industry demand trend of the position. The interest evolution prediction results in the student feature vector are cross-validated with the industry demand trend to obtain an industry demand forecast report.

[0138] In step 3, recommendation matching is performed based on the student feature vector and the major feature matrix to obtain a recommendation list, specifically:

[0139] Obtain students' subject selection information and conduct preliminary subject selection screening based on the subject selection information and subject selection rule library;

[0140] Based on the professional ability matrix, student ability vector and integrated interest results, a multi-dimensional matching engine is used to match, calculate the comprehensive score of each major, sort them according to the comprehensive score, and output a recommendation list, as shown in the following formula:

[0141] The ability matching formula is:

[0142]

[0143] Where, is the student ability vector, is the professional demand vector;

[0144] The formula for interest relevance is:

[0145] S interest =∑ kw TF-IDF(kw)×time decay factor(t);

[0146] In the formula, kw is the interest keyword, TF-IDF(kw) is the importance of the keyword;

[0147] The comprehensive scoring formula is:

[0148] Score=0.5S academic +0.3S interest +0.2S trend ;

[0149] Where Strend is the trend gain factor, which can be adjusted dynamically according to specific needs;

[0150] Based on the diagnostic report and industry demand forecast report, determine whether there are any warning prompts for the majors in the recommended list, specifically:

[0151] Conduct demand decay detection and competitive pressure analysis based on industry demand forecast reports to determine the industry's risk tags;

[0152] Analyze students' weaknesses based on the diagnostic report and analyze them against their professional abilities. If the weak points are the ones that require the highest professional abilities, then the student is considered high risk.

[0153] If the industry risk label is high risk and the student's weaknesses are also high risk, it is not recommended to apply for this major.

[0154] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0155] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. An intelligent recommendation method for high school students' university major interests based on big data analysis, characterized by: include: Step 1: Based on academic data, students' external behavior log data, and questionnaire survey results, academic ability analysis and interest feature extraction are performed to obtain student feature vectors. Step 2: Based on university subject selection requirement data, student feature vectors, professional course syllabus data, and recruitment position data, a subject selection rule base is constructed, professional capability modeling is performed, and industry trend forecasting is performed to obtain a subject selection rule base, a professional capability matrix, and an industry demand forecast report; Step 3: Perform recommendation matching based on student feature vectors, subject selection rule base, professional ability matrix and industry demand forecast report to obtain a recommendation list.

2. The method according to claim 1, characterized in that In step 1, academic ability analysis is performed based on academic data, specifically: Collect electronic versions of textbooks, syllabi, past exam questions, and expert annotation data, and pre-process them; Based on natural language processing, knowledge points are extracted and entities are recognized on the pre-processed data. Rule matching is performed on the processed data to obtain a list of knowledge points. Determine the relationship types of the knowledge point list, annotate the relationships, and build a graph database to obtain a knowledge point-ability mapping table; A model is constructed based on the multidimensional item response theory, taking the knowledge point-ability mapping table and student answer records as input, and parameter estimation is performed based on the model to obtain the student ability vector and diagnostic report; Based on the historical student ability vectors, a time series ability analysis is performed based on the preset model to obtain a trend chart of student ability changes.

3. The method according to claim 2, characterized in that In step 1, interest features are extracted based on the questionnaire survey results, specifically: Obtain questionnaire survey results and external behavior log data; Conduct explicit interest processing based on student ability vectors and questionnaire data; Mining hidden interests based on students’ ability vectors and external behavior log data; Fusing the explicit interest processing results and the implicit interest mining results to obtain the fused interest results; Based on the fusion of interest results and the trend diagram of student ability changes, interest evolution prediction is performed through a preset model to obtain the interest evolution prediction results.

4. The method according to claim 3, characterized in that In step 2, a subject selection rule base is constructed based on the subject selection requirement data of colleges and universities, specifically: Obtain the enrollment plans of various universities in previous years after the new college entrance examination, extract the subject selection requirements of each university, conduct rule analysis, perform structured storage based on the analysis results, and generate a subject selection rule library.

5. The method according to claim 4, characterized in that In step 2, professional competence modeling is performed based on student feature vectors and professional course syllabus data, specifically: Based on Stanford CoreNLP, we extract competency keywords from the professional course outline data, construct a professional competency requirement vector, and then establish a competency dimension mapping table based on the professional competency requirement vector and the student competency vector in the student feature vector, and finally output the professional competency matrix.

6. The method according to claim 5, characterized in that In step 2, industry trend prediction is performed based on student feature vectors and recruitment position data, specifically: Based on the recruitment data, the job ability requirements are extracted based on the LDA subject model, and the job ability requirement trend is predicted based on the ARIMA model to obtain the industry demand trend of the position. The interest evolution prediction results in the student feature vector are cross-validated with the industry demand trend to obtain an industry demand forecast report.

7. The method according to claim 6, characterized in that In step 3, recommendation matching is performed based on the student feature vector and the major feature matrix to obtain a recommendation list, specifically: Obtain students' subject selection information and conduct preliminary subject selection screening based on the subject selection information and subject selection rule library; Based on the professional ability matrix, student ability vector and integrated interest results, a multi-dimensional matching engine is used to match, calculate the comprehensive score of each major, sort them according to the comprehensive score, and output a recommendation list; Based on the diagnostic report and industry demand forecast report, determine whether there are warning prompts for the majors in the recommended list.

Citation Information

Patent Citations

  • Career development planning system for high school students

    CN106558001A

  • Dynamic subject selection recommendation method and device, terminal and computer readable storage medium

    CN112070639A

  • MOOCs student learning prediction method based on time convolution network

    CN117114931A

  • College entrance examination voluntary reporting auxiliary system and method

    CN119357467A

  • Employment guidance method based on enterprise recruitment data and college entrance examination application filling linkage

    CN119831795A