College employment data rapid analysis system

By designing a rapid analysis system for employment data in colleges and universities, the problem that the existing system cannot deeply analyze employment data is solved, dynamic monitoring and trend analysis of employment data is realized, accurate data support is provided, and graduates' employment competitiveness is enhanced.

CN120217084APending Publication Date: 2025-06-27王恺丽
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510278905.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing university employment data analysis system cannot meet the needs of colleges and universities for in-depth analysis of employment data, especially in dynamic monitoring and trend analysis, which is difficult to reveal the long-term impact of various industries on the future development of students.

Method used

A rapid analysis system for employment data in colleges and universities is designed, including data collection module, data matching module, data preprocessing module and data analysis module. The system can dynamically monitor and analyze employment data, and generate results of the distribution trends, salary distribution, job distribution, job stability and promotion in the employment field through the application of feature extraction, data matching, preprocessing and analysis models.

Benefits of technology

It has achieved rapid collection and processing of employment data. Through dynamic monitoring and trend analysis, it has deeply revealed the impact of various industries on students' future development, provided an accurate basis for colleges and universities to adjust their teaching plans and training plans, and improved graduates' employment competitiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217084A_ABST
    Figure CN120217084A_ABST
Patent Text Reader

Abstract

The invention relates to a rapid college employment data analysis system, which comprises a data acquisition module, a data matching module, a data preprocessing module and a data analysis module, and is characterized in that the data acquisition module is used for receiving to-be-analyzed data from a user side, and the data matching module is used for matching the to-be-analyzed data with a preset data label; the data extraction module is used for extracting a first fusion data set represented by matched data labels, the data preprocessing module is used for preprocessing the to-be-analyzed data and the first fusion data to obtain a second fusion data set, and the data analysis module is used for inputting the second fusion data set into a pre-trained employment analysis model to obtain a data output result. Dynamic monitoring and trend analysis are carried out on employment field distribution, employment salary distribution and employment post distribution of graduates and work stability and work promotion conditions of corresponding work fields through data output results, and colleges and universities are helped to better understand employment prospects and promotion opportunities of the graduates in different industries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data analysis, and particularly relates to a rapid analysis system for college employment data. Background Art

[0002] With the popularization of higher education, the number of college graduates has been increasing year by year, and the employment problem has become one of the key concerns of colleges. Colleges need to understand the employment situation of graduates in order to adjust teaching plans and training programs and improve the employment competitiveness of graduates. However, the existing college employment data analysis systems have certain limitations in function and cannot meet the needs of colleges for in-depth analysis of employment data.

[0003] Currently, most college employment data analysis systems mainly focus on the statistics and display of basic data such as the employment rate, employment regions, and employment industries of graduates, but relatively lack in-depth analysis of the long-term situation in the field of students' work. These systems usually can only provide simple data reports, cannot conduct dynamic monitoring and trend analysis of employment data, and are difficult to reveal the long-term impact of various industries on the future development of students. This makes it difficult for colleges to have an in-depth understanding of industry development trends and career development paths when formulating teaching reforms and talent training programs, and it is difficult to accurately adjust professional settings and course content to adapt to changes in market demand.

[0004] Therefore, it is particularly necessary to develop a rapid analysis system for college employment data that can conduct dynamic monitoring and trend analysis of employment data. Summary of the Invention

[0005] Based on this, the object of the present invention is to provide a rapid analysis system for college employment data, which can conduct dynamic monitoring and trend analysis of college employment data.

[0006] The object of the present invention is achieved by the following solutions:

[0007] A rapid analysis system for college employment data includes:

[0008] A data acquisition module, a data matching module, a data preprocessing module, and a data analysis module. Among them,

[0009] The data acquisition module is used to receive the data to be analyzed from the user terminal. The data to be analyzed includes students' academic information and employment data. The employment data includes employment fields, employment regions, employment durations, employment positions, and salary levels;

[0010] The data matching module is used for:

[0011] Extracting features from the data to be analyzed to obtain data features matching the data to be analyzed. The data features include major features and school name features;

[0012] Match the data features with preset data tags to generate a data matching result, where the data matching result is used to indicate the data tags that match the data features;

[0013] Process the matching result to generate a data extraction instruction and extract a first fusion data set, where the data extraction instruction is used to extract the first fusion data set corresponding to the pre-stored data tags;

[0014] The data preprocessing module is used to preprocess the data to be analyzed and the first fusion data set to obtain a second fusion data set;

[0015] The data analysis module is used to input the second fusion data set into a pre-trained employment analysis model to obtain a data output result, where the data output result is used to indicate the employment field distribution trend, employment salary distribution, employment position distribution, and work stability and work promotion in the corresponding work field.

[0016] In one embodiment, the employment analysis model includes an employment prospect analysis model;

[0017] The data analysis module includes a prospect analysis unit, where the prospect analysis unit is used to input the second fusion data set into the employment prospect analysis model for processing to obtain an employment prospect analysis result, where the employment prospect analysis result is used to indicate the work stability and work promotion in the corresponding field; wherein, the employment prospect analysis model is obtained through the following method:

[0018] Process the first fusion data set based on the logistic regression algorithm to construct the employment prospect analysis model, and the logistic regression algorithm formula is as follows:

[0019]

[0020] z = β0 + β1x1 + β2x2 + … + β n x n

[0021] where P(Y = 1|X) represents the probability that the event Y = 1 (work stability or promotion) occurs given the feature vector X = (x1, x2, …, x n ), β0 is the intercept term, β1, β2, …, β n are the coefficients corresponding to the features x1, x2, …, x n and x i includes employment duration, employment field, employment position, and student academic information.

[0022] In one embodiment, the employment analysis model further includes an employment trend analysis model;

[0023] The data analysis module further includes a trend analysis unit;

[0024] The trend analysis unit is configured to input the second fusion data set into the employment trend analysis model for processing to obtain an employment trend analysis result, which is used to indicate the employment field distribution trend, employment salary distribution, and employment position distribution of graduates. Among them, the employment trend analysis model is constructed based on the following method:

[0025] S421. For the first fusion data set, randomly select K data points as the initial cluster centers, namely C1, C2,..., C n ;

[0026] S422. Calculate the distances between each data point x j and the K cluster centers, and assign it to the cluster where the nearest cluster center is located. The distance calculation formula is as follows;

[0027]

[0028] where x j k is the k-th eigenvalue of the data point x j , C i k is the k-th eigenvalue of the cluster center C1, and m is the number of features;

[0029] S423. For each cluster, calculate the mean of all data points within each cluster and use it as the new cluster center;

[0030] S424. Repeat steps two and three until a preset number of iterations is reached.

[0031] In one embodiment, the data preprocessing module includes a data cleaning unit, a feature extraction unit, and a data fusion unit;

[0032] The data cleaning unit is configured to clean and classify the data to be analyzed and the first fusion data set to obtain an abnormal data set and a first cleaned data set. The abnormal data set is a data set with abnormally deviated values; the first cleaned data set is a data set with values within a certain range;

[0033] The feature extraction unit is configured to extract features from the first cleaned data set to generate the data labels, which are used to indicate the professional features and school name features of the first cleaned data set;

[0034] The data fusion unit is used to perform fusion processing on the first cleaned data set based on data tags to obtain a second fusion data set, and the second fusion data set is used to indicate a data set composed of data subsets corresponding to multiple different data tags after merging the data with the same data tags in the first cleaned data set together.

[0035] In one embodiment, the abnormal data set is obtained by the following method:

[0036] Sort the data to be analyzed and the first fusion data set, and sort them from small to large according to the same characteristics to obtain intermediate data;

[0037] Process the intermediate data based on the interquartile range method to obtain the abnormal value range. The formula for the interquartile range algorithm is as follows:

[0038] IQR = Q3 - Q1

[0039] Lower Bound = Q1 - 1.5 * IQR

[0040] Upper Bound = Q3 + 1.5 * IQR

[0041] Wherein, IQR is the interquartile range, Q3 is the value at the 75% position after sorting the data from small to large, Q3 is the value at the 25% position after sorting the data from small to large, Lower Bound is the lower limit of the value, Upper Bound is the upper limit of the value, and the abnormal value range is that the data value is higher than the upper limit or lower than the lower limit;

[0042] Process the intermediate data based on the abnormal value range, and determine the data with data values higher than the upper limit or lower than the upper limit as the abnormal data set.

[0043] In one embodiment, the data preprocessing module further includes a data modification unit, and the data modification unit is used to send the abnormal data set to the user side and receive the data modification result from the user side, and the data modification result is used to indicate the abnormal data set after correction and confirmation;

[0044] The data cleaning unit is further used to update the first cleaned data set based on the data modification result.

[0045] In one embodiment, the data acquisition module is further used to regularly generate data acquisition instructions based on a preset time and collect historical employment data in the cloud, identity information, and historical academic information of the school employment guidance center based on data crawling technology;

[0046] The data cleaning unit is further configured to process the historical employment data, the identity information, and the historical academic information based on the interquartile range method to obtain a second cleaned data set;

[0047] The data fusion unit is further configured to process the second cleaned data set based on data tags to obtain a first fused data set, where the first fused data set is used to indicate a data set composed of multiple data subsets corresponding to different data tags after merging the data with the same data tags in the second cleaned data set together.

[0048] In one embodiment, a data storage module is further included; the data storage module includes a high-speed storage unit, a low-speed storage unit, and a data classification unit;

[0049] The data classification unit is configured to:

[0050] Process the first fused data set to obtain weight tags for each data, where the weight tags are used to indicate the weights of the data in the first fused data set;

[0051] Identify whether the weight tags of the data in the first fused data set meet a preset high-speed storage condition, and send the first fused data set that meets the high-speed storage condition to the high-speed storage unit for storage; otherwise, send it to the low-speed storage unit for storage, where the high-speed storage condition is that the weight of the weight tag exceeds a preset weight threshold;

[0052] The high-speed storage unit is configured to receive and store the first fused data set from the data classification unit, and is further configured to receive a data extraction instruction from the data matching module and send the corresponding first fused data set to the data matching module;

[0053] The low-speed storage unit is configured to receive and store the first fused data set from the data classification unit, and is further configured to receive a data extraction instruction from the data matching module and send the corresponding first fused data set to the data matching module.

[0054] In one embodiment, the data classification unit is further configured to store the operation log of the first fused data set, where the operation log includes the memory size of the data and the access frequency of the data;

[0055] The weight tag is obtained by the following method:

[0056] Process the first fused data set based on a weight-based data storage algorithm to obtain the weight information of each data in the first fused data set, where the weight calculation formula is as follows:

[0057] p = α1 * f1 + α2 * f2 + α3 * f3

[0058] Wherein, p is the weight of the data, f1 is the access frequency of the data, f2 is the memory size of the data, f3 is the intermediate value of the performance of the high-speed storage unit and the low-speed storage unit, α1, α2, and α3 are the weight coefficients of the data, and α1 + α2 + α3 = 1;

[0059] Process the weight information to obtain a weight label.

[0060] In one embodiment, it further includes a result output module;

[0061] The result output module is used to perform visualization processing on the data output result to generate a visualization chart; and is used to sort out the visualization view and the data output result to generate an employment situation report form.

[0062] The above-mentioned rapid analysis system for college employment data receives students' academic and employment data, such as information on employment fields, regions, durations, etc., through a data collection module, then extracts features from the received data, matches them with preset labels to generate results, and extracts a first fusion dataset accordingly. The data preprocessing module integrates and cleans the data to obtain a second fusion dataset. The data analysis module inputs the data into an employment analysis model and outputs results such as the distribution trend of employment fields and salary distribution. This process not only realizes the rapid collection and processing of employment data, but also deeply reveals the impact of various industries on students' future development through dynamic monitoring and trend analysis. The rapid analysis system for college employment data provides accurate basis for colleges to adjust teaching plans and training programs, and at the same time provides employment data analysis for graduates, helps to enhance the employment competitiveness of graduates, and promotes the close connection between college talent cultivation and market demand.

[0063] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Brief Description of the Drawings

[0064] Figure 1 It is a structural block diagram of a rapid analysis system for college employment data provided by an embodiment of the present application;

[0065] Figure 2 It is a flowchart of the data matching module extracting the first fusion dataset provided by an embodiment of the present application;

[0066] Figure 3a It is a structural block diagram of the data preprocessing module provided by an embodiment of the present application;

[0067] Figure 3b It is a flowchart of obtaining a data anomaly set provided by an embodiment of the present application;

[0068] Figure 4aBlock diagram of the data analysis module provided by the embodiments of the present application;

[0069] Figure 4b Flowchart of constructing an employment trend analysis model provided by the embodiments of the present application;

[0070] Figure 5 Block diagram of the data storage module provided by the embodiments of the present application;

[0071] Figure 6 Block diagram of the result output module provided by the embodiments of the present application. Detailed implementation manners

[0072] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0073] Please refer to Figure 1 , which shows a block diagram of a rapid analysis system for college employment data provided by an exemplary embodiment of the present application. This embodiment is applicable to colleges and universities for rapid analysis of graduate employment data. As Figure 1 shown, the system includes a data acquisition module 110, a data matching module 120, a data preprocessing module 130, and a data analysis module 140.

[0074] Exemplarily, the data acquisition module 110 is used to receive the data to be analyzed from the user terminal. The data to be analyzed includes student academic information and employment data. The employment data includes employment fields, employment regions, employment durations, employment positions, and salary levels.

[0075] Preferably, when the user terminal sends the data to be analyzed including student academic information and employment data (covering employment fields, regions, durations, positions, and salary levels), the data acquisition module 110 parses the data according to the pre-set data format and protocol, accurately separates different types of data, and then stores the parsed data into the corresponding database tables to complete the preliminary collection of the data.

[0076] As Figure 2 shown, the data matching module 120 is used to perform the following steps:

[0077] S221. Extract features from the data to be analyzed to obtain data features matching the data to be analyzed. The data features include major features and school name features.

[0078] Specifically, the data is parsed through a specific algorithm to identify key information such as professional names, codes, and school names, and convert them into structured feature vectors. For example, the Convolutional Neural Network (CNN) algorithm is used to extract features from the data. CNN can automatically extract key features in the data, such as keywords in professional names and unique identifiers of school names. This process gradually extracts high-level features of the data through the operations of convolutional layers and pooling layers, providing accurate feature vectors for subsequent matching. The features extracted by CNN have strong robustness and discrimination, which can effectively improve the accuracy of matching.

[0079] S222. Match the data features with preset data labels to generate a data matching result, which is used to indicate the data label that matches the data features.

[0080] Specifically, by comparing the similarity between the data features and the preset labels, the data label that matches the data features is determined. For example, the cosine similarity algorithm is used to calculate the similarity between the data feature vector and the preset label vector. Cosine similarity measures the similarity between two vectors by calculating the cosine value of the angle between them, and the closer the value is to 1, the higher the similarity. When the similarity exceeds the set threshold, it is considered that the data feature and the data label match successfully. This process can quickly and accurately find the label that matches the data to be analyzed, providing a basis for subsequent data extraction.

[0081] S223. Process the matching result to generate a data extraction instruction and extract the first fusion data set. The data extraction instruction is used to extract the first fusion data set corresponding to the preset data label.

[0082] Preferably, a rule-based extraction algorithm can be used to generate a data extraction instruction according to the matching result. According to the successfully matched data labels, corresponding query statements are constructed to extract data related to these labels from the pre-stored data set to form the first fusion data set. This process ensures that the extracted data is highly relevant to the data to be analyzed, reduces the interference of data with less similarity to data analysis, and provides an accurate data basis for subsequent data analysis.

[0083] The data preprocessing module 130 is used to preprocess the data to be analyzed and the first fusion data set to obtain a second fusion data set. Specifically, the data preprocessing module compares the data to be analyzed with the professional features and school name features in the first fusion data one by one, and integrates other relevant data fields, such as student grades, employment position information, etc., according to the set fusion rules, and finally fuses the two together to form a second fusion data set containing comprehensive and ordered information, providing a data basis for subsequent employment analysis.

[0084] The data analysis module 140 is used to input the second fusion data set into a pre-trained employment analysis model to obtain a data output result, which is used to indicate the distribution trend of employment fields, the distribution of employment salaries, the distribution of employment positions, and the job stability and job promotion in the corresponding work fields.

[0085] Specifically, the employment analysis model is trained based on a large amount of historical data and has powerful data processing and analysis capabilities. On the one hand, the model will deeply mine and analyze various information in the second fusion data set, and use statistical analysis and machine learning algorithms to extract key features and rules from the massive data. For example, by comparing employment data in different time periods, it can accurately predict the development trend of employment fields, clarify which fields are booming, which are tending to be stable or declining; analyze salary data to reveal the changing trends of salary levels in different industries and positions. On the other hand, the model will continuously track the changes in various indicators of employees during the employment process, such as employment duration, job change frequency, etc. By establishing a dynamic monitoring model, it can evaluate job stability in real time; based on employees' promotion records, performance data, etc., analyze the possibility and influencing factors of job promotion, so as to present the distribution trend of employment fields, the distribution of employment salaries, the distribution of employment positions, and the job stability and job promotion in the corresponding work fields in an intuitive and easy-to-understand way, providing valuable decision-making basis for relevant parties such as universities and graduates, and effectively improving the competitiveness of graduates in the workplace.

[0086] In summary, a rapid analysis system for college employment data provided by an embodiment of the present application receives students' academic and employment data, such as information on employment fields, regions, durations, etc., through the data collection module 110, then extracts features from the received data, matches them with preset tags to generate results, and extracts the first fusion data set accordingly. The data preprocessing module 130 integrates and cleans the data to obtain the second fusion data set. The data analysis module 140 inputs the data into the employment analysis model and outputs results such as the distribution trend of employment fields and salary distribution. This process not only realizes the rapid collection and processing of employment data, but also deeply reveals the impact of various industries on students' future development through dynamic monitoring and trend analysis. This rapid analysis system for college employment data provides accurate basis for colleges to adjust teaching plans and training programs, and at the same time provides employment data analysis for graduates, helping to enhance the employment competitiveness of graduates and promoting the close connection between college talent cultivation and market demand.

[0087] Preferably, as Figure 4a shown, the employment analysis model includes an employment prospect analysis model and an employment trend analysis model, and the data analysis module 140 is configured with the following units:

[0088] A foreground analysis unit 410 is configured to input the second fusion dataset into an employment foreground analysis model for processing to obtain an employment foreground analysis result, where the employment foreground analysis result is used to indicate the job stability and job promotion situation in the corresponding field. The employment foreground analysis model is obtained through the following method:

[0089] Process the first fusion dataset based on the logistic regression algorithm to construct an employment foreground analysis model. The formula of the logistic regression algorithm is as follows:

[0090]

[0091] z = β0 + β1x1 + β2x2 + … + β n x n

[0092] where P(Y = 1|X) represents the probability that the event Y = 1 (job stability or promotion) occurs given the feature vector X = (x1, x2, …, x n ). β0 is the intercept term, and β1, β2, …, β n are the coefficients corresponding to the features x1, x2, …, x n , and x i includes employment duration, employment field, employment position, and student academic information.

[0093] Specifically, the employment foreground analysis model can monitor various dynamic data in the employment market in real time by leveraging the continuously updated first fusion dataset and second fusion dataset. By continuously analyzing multi-dimensional features such as employment duration, employment field, employment position, and student academic information, it can understand in real time the employment duration and position promotion situation of graduates in the corresponding field. For schools, it can clearly understand which majors have higher employment stability, and then reasonably adjust the enrollment scale, focus on supporting popular and stable majors, and avoid resource waste. Based on the data on the duration required for promotion from a specific position, it can gain insights into the development potential of each major, optimize the curriculum system, and cultivate talents that better meet market needs and have good career advancement space. For graduates, they can judge the job stability based on the average employment duration of different positions, avoid career choices with greater risks, and can refer to the promotion duration of positions to clearly understand the development potential of different positions, clarify career goals, and formulate a reasonable career development path.

[0094] A trend analysis unit 420 is configured to input the second fusion dataset into an employment trend analysis model for processing to obtain an employment trend analysis result, where the employment trend analysis result is used to indicate the employment field distribution trend, employment salary distribution, and employment position distribution of graduates. As Figure 4b shown, the employment trend analysis model is constructed based on the following method:

[0095] S421. For the first fusion data set, randomly select K data points as the initial clustering centers, namely C1, C2, …, C n ;

[0096] S422. Calculate the distances between each data point x j and the K clustering centers, and assign it to the cluster where the nearest clustering center is located. The distance calculation formula is as follows;

[0097]

[0098] where x j k is the k-th eigenvalue of the data point x j , C i k is the k-th eigenvalue of the clustering center C1, and m is the number of features;

[0099] S423. For each cluster, calculate the mean value of all data points within each cluster and use it as the new clustering center;

[0100] S424. Repeat S422 and S423 until the preset number of iterations is reached.

[0101] Specifically, the employment trend analysis model can identify the distribution patterns and trends in employment data. Through clustering analysis, it can be found which employment fields are the main choices of graduates, which positions have higher salary levels, and the distribution of different employment positions. This information is of great significance for universities to adjust professional settings, optimize the curriculum system, and formulate employment guidance policies. At the same time, for graduates, understanding employment trends helps them better plan their career development paths and improve their employment competitiveness.

[0102] In one embodiment, as Figure 3a shown, the data preprocessing module 130 is configured with the following units:

[0103] The data cleaning unit 310 cleans and classifies the data to be analyzed and the first fusion data set to obtain an abnormal data set and a first cleaned data set. The abnormal data set is a data set with abnormally deviated values; the first cleaned data set is a data set with values within a certain range.

[0104] Preferably, as Figure 3b shown, the abnormal data set is obtained by the following method:

[0105] S311. Organize the data to be analyzed and the first fusion data set, sort them from smallest to largest according to the same features, and obtain intermediate data;

[0106] S312. Process the intermediate data based on the interquartile range method to obtain the abnormal value range. The interquartile range algorithm formula is as follows:

[0107] IQR = Q3 - Q1

[0108] Lower Bound = Q1 - 1.5 * IQR

[0109] Upper Bound = Q3 + 1.5 * IQR

[0110] Among them, IQR is the interquartile range, Q3 is the value at the 75% position after the data is sorted from small to large, Q1 is the value at the 25% position after the data is sorted from small to large, Lower Bound is the lower limit of the value, Upper Bound is the upper limit of the value, and the abnormal value range is that the data value is higher than the upper limit or lower than the lower limit;

[0111] S313. Process the intermediate data based on the abnormal value range, and determine the data with data values higher than the upper limit or lower than the upper limit as the abnormal data set.

[0112] The feature extraction unit 320 extracts features from the first cleaned data set to generate data labels, and the data labels are used to indicate the professional features and school name features of the first cleaned data set.

[0113] The data fusion unit 330 is used to perform fusion processing on the first cleaned data set based on the data labels to obtain the first fusion data set. The first fusion data set is used to indicate a data set composed of multiple different data subsets formed by merging the data with the same data labels in the first cleaned data set together.

[0114] Specifically, the sequential traversal algorithm can be adopted. Using a loop structure, by setting an iterator, starting from the starting position of the data set, in the storage order of the data, with a step size of 1, each data instance in the data set is accessed in turn. During the traversal process, according to the principle of equivalence class partitioning, the data instances carrying the same professional data label are divided into the same equivalence class, and a merge operation is performed on them. The equivalence classes (i.e., data subsets) aggregated by professional labels are integrated to construct the first fusion data set. This data set is classified by major, separating the data of different majors. Each data subset corresponds to a specific major, providing a structured data basis for subsequent data mining, model training and other tasks.

[0115] In the actual usage process, the generation of abnormal data may stem from data entry errors, system failures, the influence of special events, or the extreme performance of individual samples. For example, in the employment data of college graduates, due to the mistakes of staff, the employment salary of a certain graduate may be misrecorded as much higher than the actual value. For these abnormal data, if they can be modified and reused, it will greatly enhance the utilization value of the data. By screening and correcting abnormal data, potential useful information can be mined, thereby improving the accuracy and reliability of employment data analysis.

[0116] Preferably, as shown in Figure 3, the data preprocessing module 130 further includes a data modification unit 340. The data modification unit 340 is configured to send the abnormal data set to the user side and receive the data modification result from the user side. Preferably, the data cleaning unit 310 is further configured to update the first cleaned data set based on the data modification result.

[0117] Preferably, the data modification unit 340 can send the abnormal data set to the user side through a data transmission interface. The user side verifies and modifies the abnormal data set. For example, when the user finds that the employment salary data of a certain graduate is abnormal, after verification, it is modified to the correct value. Then, the user side returns the modified data to the data modification unit 340 through the data transmission interface. After receiving the data modification result, the data modification unit 340 passes it to the data cleaning unit 310. The data cleaning unit 310 updates the first cleaned data set according to these modification results. For example, the data originally marked as abnormal is replaced with the correct data, thereby ensuring the accuracy and integrity of the first cleaned data set and providing a reliable data basis for subsequent data analysis.

[0118] Preferably, in one embodiment, the data acquisition module 110 is further configured to periodically generate data acquisition instructions based on a preset time and collect historical employment data in the cloud, identity information, and historical academic information of the school employment guidance center based on data crawling technology. For example, on the 1st of each month, the system automatically runs the acquisition instructions to collect historical employment data from major recruitment platforms, identity information retained by students in the cloud, and information such as students' academic achievements recorded by the school. On the one hand, it ensures the timeliness of the data, enabling the analysis to keep up with the employment market and students' academic dynamics. On the other hand, the diverse data sources provide rich materials for subsequent in-depth analysis, facilitating employment research and decision-making.

[0119] Preferably, the data cleaning unit 310 is further configured to process the historical employment data, the identity information, and the historical academic information based on the interquartile range method to obtain a second cleaned data set;

[0120] The data fusion unit 330 is further configured to process the second cleaned data set based on data tags to obtain a first fusion data set, where the first fusion data set is used to indicate a data set composed of data subsets corresponding to multiple different data tags after merging the data with the same data tags in the second cleaned data set together.

[0121] In one embodiment, as Figure 5 shown, the college employment data rapid analysis system provided by this application further includes a data storage module 150. Among them, the data storage module 150 is configured with the following units:

[0122] A data classification unit 510, which is used to perform the following functions:

[0123] Function 1: Process the first fusion data set to obtain weight tags for each data, where the weight tags are used to indicate the weights of each data in the first fusion data set.

[0124] Preferably, the weight tags are obtained in the following manner:

[0125] Process the first fusion data set based on a weight-based data storage algorithm to obtain the weight information of each data in the first fusion data set. Among them, the weight calculation formula is as follows:

[0126] p = α1 * f1 + α2 * f2 + α3 * f3

[0127] Where p is the weight of the data, f1 is the access frequency of the data, f2 is the memory size of the data, f3 is the intermediate value of the performance of the high-speed storage unit 520 and the low-speed storage unit 530, α1, α2, and α3 are the weight coefficients of the data, and α1 + α2 + α3 = 1;

[0128] Process the weight information to obtain weight tags.

[0129] Function 2: Identify whether the weight tags of each data in the first fusion data set meet a preset high-speed storage condition, and send the first fusion data set that meets the high-speed storage condition to the high-speed storage unit 520 for storage; otherwise, send it to the low-speed storage unit 530 for storage. The high-speed storage condition is that the weight of the weight tag exceeds a preset weight threshold.

[0130] Specifically, with the continuous growth of the data volume, traditional data storage methods often lead to increased difficulty and slower speed in data reading. Especially, the reading of important data is prone to stagnation due to the interference of other data. To address this issue, the data classification unit 510 of the data storage module 150 of this system processes the first fusion data set to obtain weight tags for each piece of data. These tags, based on a comprehensive consideration of data access frequency, memory size, and storage unit performance, can accurately reflect the importance and access requirements of the data. For example, for data with a high access frequency and small memory occupancy, its weight tag will be higher, indicating that these data are more suitable for storage in the high-speed storage unit 520 to ensure fast reading. The data classification unit 510 identifies whether the data meets the high-speed storage conditions based on the weight tags, sends the qualified data to the high-speed storage unit 520, and sends the remaining data to the low-speed storage unit 530, thus effectively avoiding the problem that important data is stagnated due to the interference of other low-priority data during reading. For example, when a university analyzes the employment data of popular majors, since these data have a high access frequency, they will be preferentially stored in the high-speed storage unit 520 to ensure rapid access when needed, thereby improving the real-time performance and accuracy of data analysis.

[0131] The high-speed storage unit 520 is used to receive and store the first fusion data set from the data classification unit 510, and is also used to receive a data extraction instruction from the data matching module 120 and send the corresponding first fusion data set to the data matching module 120.

[0132] Preferably, the high-speed storage unit 520 can be a static random access memory (SRAM), which is fast and does not require refreshing, and is often used in caches (Cache), which can significantly improve the speed at which the processor accesses data and reduce waiting time; it can also be a dynamic random access memory (DRAM), which is the main component of computer memory, has a fast response speed, can store and read the running programs and data in a timely manner, and ensures the efficient operation of the computer system; it can also be a solid-state drive (SSD), which stores data based on flash memory chips and is equipped with an advanced controller, enabling the data read and write speed to far exceed that of traditional mechanical hard drives. It is widely used in devices such as personal computers and servers to assist the system in starting quickly and programs in loading quickly.

[0133] The low-speed storage unit 530 is used to receive and store the first fusion data set from the data classification unit 510, and is also used to receive a data extraction instruction from the data matching module 120 and send the corresponding first fusion data set to the data matching module 120.

[0134] Preferably, the low-speed storage unit 530 can be a hard disk drive (HDD), which relies on a magnetic head to read and write data on the disk platter. Its characteristics of large capacity and low cost enable it to play a role in scenarios such as data warehouses and surveillance storage, where high read and write speeds are not required but a large amount of storage space is needed. It can also be tape storage, which uses tape as the storage medium, has a large storage capacity and low cost, and is mostly used in scenarios such as data backup and long-term archiving. It can also be optical disc storage, such as common CDs, DVDs, and Blu-ray discs, which record and read data on the disc surface through a laser and have a relatively large storage capacity.

[0135] Through the division of labor and cooperation between the high-speed storage unit 520 and the low-speed storage unit 530, not only is the data reading speed improved, but also the utilization efficiency of storage resources is optimized. The high-speed storage unit 520 focuses on storing important data to ensure fast access to critical data, while the low-speed storage unit 530 is responsible for storing other data, reducing the storage cost. This reasonable storage allocation mechanism enables the system to maintain efficient and stable operation when facing a large amount of data, provides strong support for the rapid analysis of college employment data, helps colleges and universities adjust teaching plans and training programs in a timely manner, and improves the employment competitiveness of graduates.

[0136] In one embodiment, as Figure 6 shown, the college employment data rapid analysis system provided by this application further includes a result output module 160. Among them, the result output module 160 is configured with the following units:

[0137] A visualization unit 610, which is used to perform visualization processing on the data output result to generate a visualization chart.

[0138] Preferably, data visualization technology can be adopted to convert employment data into intuitive charts. For example, a bar chart can be generated to show the comparison of employment rates of different majors in the same field, so as to more intuitively see the employment situation of different majors in different fields. A line chart can also be used to present the changing trend of the employment rate of a certain major in a certain field over time, so as to clearly show the changing trend of employment data over time and help users better understand the dynamic changes in the employment market.

[0139] A report generation unit 620, which is used to organize the visualization view and the data output result to generate an employment situation report form.

[0140] Preferably, the report generation unit 620 can adopt template engine technology. With the help of a pre-designed employment situation report template, it integrates the visual view and the data output result, and accurately fills the key data in the data output result, such as the employment rate of each major, the average salary, the employment position distribution, etc., into the corresponding positions of the template. At the same time, using image embedding technology, the visual view is embedded in the specified area of the report at an appropriate resolution and format, ensuring that the chart is closely combined with the text description, thus greatly improving the readability of data presentation, presenting complex data and charts in a clear and organized manner, and enabling users to quickly understand the employment situation, such as intuitively understanding the core information such as the employment rate of each major and the employment position distribution.

[0141] In summary, a rapid analysis system for college employment data provided by an embodiment of the present application receives students' academic and employment data through the data collection module 110, then extracts the features of the received data, matches them with preset tags to generate results, and extracts the first fusion data set accordingly. The data preprocessing module 130 integrates and cleans the data to obtain the second fusion data set. The data analysis module 140 inputs the data into the employment analysis model and outputs results such as the distribution trend of the employment field and the salary distribution. The result output module 160 converts the output result into a visual icon and an employment situation report, enabling users to more clearly and directly understand the changes in the employment situation of graduates. Through the dynamic monitoring of the employment duration and employment positions and the trend analysis of the employment distribution, colleges and universities can more comprehensively understand the employment situation and future development of graduates, which is conducive to colleges and universities formulating teaching reforms and talent cultivation plans according to the changes in the employment situation of all walks of life, more effectively enhancing the employment competitiveness of graduates, and promoting the close connection between college talent cultivation and market demand.

[0142] The above embodiments only represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.

Claims

1. A rapid analysis system for college employment data, comprising a data acquisition module, a data matching module, a data preprocessing module and a data analysis module, characterized in that: The data collection module is used to receive the data to be analyzed from the user end, the data to be analyzed includes student academic information and employment data, the employment data includes employment field, employment area, employment time, employment position and salary level; The data matching module is used for: Extracting features of the data to be analyzed to obtain data features that match the data to be analyzed, wherein the data features include major features and school name features; Matching the data feature with a preset data tag to generate a data matching result, wherein the data matching result is used to indicate the data tag that matches the data feature; Processing the matching result, generating a data extraction instruction and extracting a first fused data set, wherein the data extraction instruction is used to extract the first fused data set corresponding to the pre-stored data label; The data preprocessing module is used to preprocess the data to be analyzed and the first fused data set to obtain a second fused data set; The data analysis module inputs the second fused data set into a pre-trained employment analysis model to obtain data output results, and the data output results are used to indicate the distribution trend of employment fields, employment salary distribution, employment position distribution, and job stability and job promotion in corresponding work fields.

2. The rapid analysis system for college employment data according to claim 1 is characterized by: The employment analysis model includes an employment prospect analysis model; The data analysis module includes a prospect analysis unit, which is used to input the second fused data set into the employment prospect analysis model for processing to obtain an employment prospect analysis result, and the employment prospect analysis result is used to indicate the job stability and job promotion situation in the corresponding field; wherein the employment prospect analysis model is obtained by: The first fused data set is processed based on a logistic regression algorithm to construct the employment prospect analysis model. The logistic regression algorithm formula is as follows: z=β0+β1x1+β2x2+…+β n x n Where P(Y=1|X) means that given the feature vector X=(x1,x2,…,x n ), the probability of event Y = 1 (job stability or promotion), β0 is the intercept term, β1, β2, …, β n are the features x1, x2, …, x n The corresponding coefficient, x i Including employment duration, employment field, job position and student academic information.

3. The rapid analysis system for college employment data according to claim 2 is characterized in that: The employment analysis model also includes an employment trend analysis model; The data analysis module also includes a trend analysis unit; The trend analysis unit is used to input the second fused data set into the employment trend analysis model for processing to obtain an employment trend analysis result, and the employment trend analysis result is used to indicate the employment field distribution trend, employment salary distribution and employment position distribution of graduates, wherein the employment trend analysis model is constructed based on the following method: S421: For the first fused data set, randomly select K data points as initial cluster centers, namely C1, C2, ..., C n ; S422, calculate each data point x j The distance from K cluster centers is assigned to the cluster where the nearest cluster center is located. The distance calculation formula is as follows; Among them, x j k is the data point x j The kth eigenvalue of i k is the kth eigenvalue of cluster center C1, and m is the number of features; S423, for each cluster, calculating the mean of all data points in each cluster, and taking it as a new cluster center; S424, repeat S422 and S423 until a preset number of iterations is reached.

4. The rapid analysis system for college employment data according to claim 1 is characterized by: The data preprocessing module includes a data cleaning unit, a feature extraction unit and a data fusion unit; The data cleaning unit is used to clean and classify the data to be analyzed and the first fused data set to obtain an abnormal data set and a first cleaned data set, wherein the abnormal data set is a data set with abnormal deviation in values; The first cleaned data set is a data set whose values ​​are within a certain range; The feature extraction unit is used to extract features from the first cleaned data set to generate the data label, where the data label is used to indicate the professional features and school name features of the first cleaned data set; The data fusion unit is used to fuse the first cleaned data set based on the data label to obtain a second fused data set, where the second fused data set is used to indicate a data set consisting of data subsets corresponding to multiple different data labels after merging data with the same data label in the first cleaned data set.

5. The college employment data rapid analysis system according to claim 4 is characterized in that: The abnormal data set is obtained in the following way: Arrange the data to be analyzed and the first fused data set, and sort them from small to large according to the same features to obtain intermediate data; The intermediate data is processed based on the interquartile range method to obtain the abnormal value range. The interquartile range algorithm formula is as follows: IQR=Q3-Q1 LowerBound=Q1-1.5*IQR UpperBound=Q3+1.5*IQR Wherein, IQR is the interquartile range, Q3 is the value at the 75% position after the data is sorted from small to large, Q3 is the value at the 25% position after the data is sorted from small to large, LowerBound is the lower limit of the value, UpperBound is the upper limit of the value, and the abnormal value range is that the data value is higher than the upper limit of the value or lower than the lower limit of the value; The intermediate data are processed based on the abnormal value range, and the data in which the data value is higher than the upper value limit or lower than the upper value limit is determined as the abnormal data set.

6. The rapid analysis system for college employment data according to claim 5 is characterized by: The data preprocessing module further includes a data modification unit, which is used to send the abnormal data set to the user end and receive a data modification result from the user end, wherein the data modification result is used to indicate the abnormal data set after correction and confirmation; The data cleaning unit is further configured to update the first cleaning data set based on the data modification result.

7. The rapid analysis system for college employment data according to claim 6 is characterized by: The data collection module is also used to periodically generate data collection instructions based on a preset time and collect historical employment data, identity information and historical academic information of the school's career guidance center in the cloud based on data crawling technology; The data cleaning unit is further used to process the historical employment data, the identity information and the historical academic information based on the interquartile range method to obtain a second cleaned data set; The data fusion unit is further used to process the second cleaned data set based on the data label to obtain a first fused data set, wherein the first fused data set is used to indicate a data set consisting of data subsets corresponding to multiple different data labels after merging data with the same data label in the second cleaned data set.

8. The college employment data rapid analysis system according to claim 1 is characterized in that: It also includes a data storage module; the data storage module includes a high-speed storage unit, a low-speed storage unit and a data classification unit; The data classification unit is used for: updating the first fused data set based on the second fused data set; Processing the first fused data set to obtain a weight label for each data, wherein the weight label is used to indicate a weight for each data in the first fused data set; Identify whether the weight label of each data in the first fused data set meets a preset high-speed storage condition, and send the first fused data set that meets the high-speed storage condition to a high-speed storage unit for storage; otherwise, send it to a low-speed storage unit for storage, and the high-speed storage condition is that the weight of the weight label exceeds a preset weight threshold; The high-speed storage unit is used to receive and store the first fused data set from the data classification unit, and is also used to receive a data extraction instruction from the data matching module and send the first fused data set corresponding to the data matching module; The low-speed storage unit is used to receive and store the first fused data set from the data classification unit, and is also used to receive a data extraction instruction from the data matching module and send the first fused data set corresponding to the data matching module.

9. The college employment data rapid analysis system according to claim 8 is characterized by: The data classification unit is further used to store an operation log of the first fused data set, wherein the operation log includes a memory size of the data and an access frequency of the data; The weight label is obtained in the following way: The first fused data set is processed by a weight-based data storage algorithm to obtain weight information of each data in the first fused data set, wherein the weight calculation formula is as follows: p=α1*f1+α2*f2+α3*f3 Wherein, p is the weight of the data, f1 is the access frequency of the data, f2 is the memory size of the data, f3 is the intermediate value of the performance of the high-speed storage unit and the low-speed storage unit, α1, α2, α3 are the weight coefficients of the data, and α1+α2+α3=1; The weight information is processed to obtain a weight label.

10. The college employment data rapid analysis system according to claim 1 is characterized in that: Also includes a result output module; The result output module is used to visualize the data output results and generate a visualization chart; and to organize the visualization view and the data output results to generate an employment status report.

Citation Information

Cited By

  • Personalized employment guidance service system based on data analysis

    CN121010483A