Career planning analysis system and method based on multi-modal data fusion
By adopting multimodal data fusion, career interest knowledge graph and improved NLP model in the career planning system, a three-dimensional career portrait of students is generated and personalized career paths are recommended, which solves the problem that the existing system cannot dynamically adapt to changes in students' interests, and improves the analysis accuracy and personalized recommendation effect.
Patent Information
- Application Number
- CN202510307685.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
The existing career planning system lacks the integration of students' multimodal data, and cannot dynamically adapt to changes in students' interests, and lacks analysis depth and accuracy.
The professional interest knowledge graph is pre-constructed by the original data collected based on crawler technology, combined with the improved NLP model and personalized recommendation model, a three-dimensional career portrait of students based on multimodal data fusion is generated, and a personalized career interest matching path is generated through frequent sub-graph mining algorithms.
It has achieved a comprehensive understanding of students' multi-dimensional data, dynamically generated personalized career interest matching paths, adapted to changes in students' interests and abilities, and improved the analysis depth and accuracy.
Smart Images

Figure CN120144872A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data analysis, and in particular relates to a career planning analysis system and method based on multimodal data fusion. Background Art
[0002] With the in-depth development of educational informatization, dynamic data of college students in academic, social, psychological and other aspects are constantly accumulated. However, traditional career planning systems are often based only on single academic performance or limited social interaction data, lacking in-depth mining and analysis of students' comprehensive characteristics. In addition, most existing systems use static evaluation models, which are difficult to adapt to the actual situation of students' interests, abilities and values changing over time, resulting in inaccurate and individuated career guidance. Therefore, how to effectively integrate multi-source heterogeneous data and build a career planning system that can reflect the overall picture of students in real time and intelligently recommend personalized development paths has become an urgent problem to be solved in the current field of educational technology.
[0003] Chinese patent CN116433427A discloses a personalized learning career data portrait method, which includes the following steps: collecting students' learning career data through a learning situation collection system; analyzing the learning situation information through an intelligent analysis system to obtain a learning career data portrait; the personalized learning system provides students with personalized learning resources based on the learning career data portrait. However, the existing method uses a big data analysis engine to iterate and perform correlation analysis on the original learning situation data. The data source is relatively single, and there is a lack of fusion of multimodal data such as students' career interests and external behaviors. It cannot dynamically adapt to changes in students' interests and is insufficient in analysis depth and accuracy. To address the above problems, we propose a career planning analysis system and method based on multimodal data fusion. Summary of the invention
[0004] The purpose of the present invention is to address the shortcomings of the prior art and provide a career planning analysis system and method based on multimodal data fusion, which solves the problems that the existing method uses a big data analysis engine to iterate and perform correlation analysis on the original academic data, the data source is relatively single, and there is a lack of fusion of multimodal data such as students' career interests and external behaviors, and it cannot dynamically adapt to changes in students' interests, and the analysis depth and accuracy are insufficient.
[0005] The present invention is implemented in this way: a career planning analysis method based on multimodal data fusion, the career planning analysis method based on multimodal data fusion comprising:
[0006] Collect the original data of the graph based on crawler technology, parse the original data of the graph and upload the original data of the graph to the storage database, and pre-build the career interest knowledge graph based on the original data of the graph;
[0007] Obtain student-related data, perform normalized fusion processing on the student-related data, parse the student-related data based on an improved NLP model, generate a three-dimensional career portrait of the student based on multi-modal data fusion, visually present the three-dimensional career portrait of the student, and upload the three-dimensional career portrait of the student to the storage database;
[0008] Traverse the storage database, extract a modeling sample set from the storage database, divide the modeling sample set into a training set and a test set, pre-construct an interest recommendation model based on collaborative filtering combined with the ALBERT model, and train the interest recommendation model using the training set and the test set;
[0009] Load the three-dimensional career portrait of the student in real time. The interest recommendation model generates a personalized career interest matching path based on the frequent subgraph mining algorithm combined with the interactive point set to traverse the career interest knowledge graph.
[0010] Preferably, the method for pre-constructing a career interest knowledge graph based on the original graph data includes:
[0011] Use the BeautifulSoup crawler framework to capture the original graph data, and perform abnormal data cleaning processing on the original graph data. Among them, the original graph data includes course management data, job information data, tutor association data, career planning data, and development demand data;
[0012] Based on course attributes, job attributes, and teacher attributes, pre-build the initial architecture of the knowledge graph, and respectively build a data layer and a schema layer based on course attributes, job attributes, and teacher attributes. Extract the original graph data after abnormal data cleaning in the form of entity-attribute-attribute value triples, and fill the structured data of the entity-attribute-attribute value triples associated with the original graph data into the data layer;
[0013] Traverse the data layer of the initial architecture of the knowledge graph, and fill the entity types, attributes, and relationships of the structured data into the schema layer of the initial architecture of the knowledge graph in a bottom-up manner;
[0014] Introduce a career association layer into the initial architecture of the knowledge graph from the perspectives of ability, interest, and value. The career association layer digitally tags the course attributes, job attributes, and teacher attributes of the associated entities based on the perspectives of ability, interest, and value combined with the principal component analysis method. Among them, the digital tags include course weights, job weights, and teacher weights;
[0015] Calculate the association degree between the associated entity and the career association layer based on the fuzzy clustering algorithm, and determine whether the association degree between the associated entity and the career association layer exceeds the preset association threshold. If the association degree between the associated entity and the career association layer exceeds the preset association threshold, extract the corresponding associated entity;
[0016] Load at least one set of associated entities that exceed the preset association threshold, generate an associated entity set, supplement the associated entity set to the career association layer, generate a career interest recommendation path based on the ant colony algorithm, and calculate the comprehensive recommendation degree of the career interest recommendation path;
[0017] Integrate the career association layer including the associated entity set, the career interest recommendation path, and the comprehensive recommendation degree to complete the construction of the career interest knowledge graph.
[0018] Preferably, the association degree between the associated entity and the career association layer is calculated based on the fuzzy clustering algorithm, and is calculated by the following formula:
[0019]
[0020]
[0021] Among them, Sim ij represents the association degree between the associated entity i and the career association layer j. I and J are the numbers of the associated entity i and the career association layer j respectively. |i∩j| represents the interaction quantity between the associated entity i and the career association layer j based on the fuzzy clustering algorithm. α 1 , α 2 , α 3 are respectively the course weight, position weight, and teacher weight of the associated entity. KPD i represents the interaction degree of the associated entity i obtained based on the fuzzy clustering algorithm. M ij is the association matrix between the associated entity i and the career association layer j. m i is the clustering center of the clustering analysis of the associated entity i based on the fuzzy clustering algorithm;
[0022] Generate a career interest recommendation path based on the ant colony algorithm, and calculate the comprehensive recommendation degree of the career interest recommendation path, which is calculated by the following formula:
[0023]
[0024] Among them, L z represents the comprehensive recommendation degree of the career interest recommendation path. K(i), G(i), and S(i) are respectively the course information vector, position information vector, and teacher information vector. A y is the comprehensive heuristic factor based on the ant colony algorithm. L represents all search paths based on the ant colony algorithm. κ represents the ant colony obstacle avoidance factor. β is the ant colony expectation factor. K, G, and S are respectively the number of courses, the number of positions, and the number of teachers.
[0025] Preferably, the method for generating the three-dimensional career portrait of students based on multi-modal data fusion specifically includes:
[0026] Load student-related data, where the student-related data includes students' academic achievements, psychological evaluation results, internship evaluations, and students' basic information;
[0027] Load the NLP model, add a BiLSTM feature extraction layer and a CRF annotation sequence layer to the NLP model, and introduce an adaptive ensemble learning algorithm based on time-stratified sampling to complete the improvement of the NLP model, and output the improved NLP model;
[0028] The improved NLP model performs dimensionality reduction processing on the student-related data, maps the dimensionality-reduced student-related data to the dimensionality-reduced feature space, and the improved NLP model annotates and assigns weight scores to the student-related data in the dimensionality-reduced feature space to obtain a dimensionality-reduced feature space containing student-related data, data annotations, and data weights;
[0029] Pre-construct a career interest spherical model based on ability, interest, and value, traverse the dimensionality-reduced feature space, and fill the student-related data, data annotations, and data weights into the career interest spherical model;
[0030] The BiLSTM feature extraction layer in the NLP model extracts features from the student-related data and data annotations, outputs the predicted labels corresponding to the student-related data, and introduces the predicted labels into the career interest spherical model based on the adaptive ensemble learning algorithm to complete the construction of the student's three-dimensional career portrait and visually present the student's three-dimensional career portrait.
[0031] Preferably, the method for training the interest recommendation model using the training set and the test set specifically includes:
[0032] Load the training set and the pre-constructed interest recommendation model, initialize the parameters of the embedding layer of the interest recommendation model using the Gaussian distribution, initialize the parameters of the hidden layer of the interest recommendation model in combination with the Glorot algorithm, collect samples in the training set based on the with-replacement strategy, and perform parallel iterative training on the interest recommendation model to output at least one set of sub-recommendation models;
[0033] Update the parameter convergence factor and the non-linear activation function of the sub-recommendation model based on the whale optimization algorithm to obtain the sub-recommendation model after updating the parameter convergence factor and the non-linear activation function;
[0034] Calculate the model fitness of the sub-recommendation model using the Gaussian distribution function, obtain the model fitness of at least one set of sub-recommendation models, and select the sub-recommendation model with the minimum model fitness as the interest recommendation model;
[0035] Among them, the model fitness of the sub-recommendation model is calculated by the following formula:
[0036]
[0037] Among them, f(λ t+1 ,σt+1 ) represents the model fitness of the sub-recommendation model, λ t+1 , σ t+1 are respectively the parameter convergence factor and the hyperparameter of the non-linear activation function after the (t + 1)-th iteration update. λ t is the parameter convergence factor after the t-th iteration update. λ 0 is the initial parameter convergence factor, X t represents the number of sub-recommendation models, and Δτ is the dimension of the whale population;
[0038] Load the test set. Using the test set as the input, execute the interest recommendation model, output the test results, calculate the cross-entropy loss between the test results and the true results using the cross-entropy loss function, and determine whether the cross-entropy loss meets the preset loss threshold. If the cross-entropy loss meets the preset loss threshold, output the converged interest recommendation model;
[0039] If the cross-entropy loss does not meet the preset loss threshold, use the Adam optimizer combined with the collaborative filtering algorithm to adjust the learning rate of the interest recommendation model, the parameters of the embedding layer, and the parameters of the hidden layer, and continue to perform parallel iterative training on the interest recommendation model.
[0040] Preferably, the interest recommendation model uses the ALBERT model as the initial model, freezes the feed-forward network in the initial model, replaces the feed-forward network of the initial model with a population-directed learning path network, and introduces a spectral clustering algorithm into the population-directed learning path network. A category adversarial joint learning network is added after the initial model. The category adversarial joint learning network includes a category classifier, a category discriminator, and a graph spectrum scoring device. A frequent subgraph mining algorithm is introduced into the graph spectrum scoring device.
[0041] Preferably, the method for generating a personalized career interest matching path specifically includes:
[0042] Load the student's three-dimensional career portrait in real time. The interest recommendation model constructs at least one group of student-graph knowledge point interaction matrices based on the student's three-dimensional career portrait;
[0043] Calculate the student-graph knowledge point interaction matrices to determine the course interaction degree, teacher interaction degree, and job interaction degree, and integrate the course interaction degree, teacher interaction degree, and job interaction degree to generate an interaction point set;
[0044] The interest recommendation model traverses the career interest knowledge graph based on the frequent subgraph mining algorithm combined with the interaction point set. The category adversarial joint learning network identifies the interest recommendation degree associated with the student in the student-graph knowledge point interaction matrix;
[0045] Taking the interest recommendation degree associated with the student as the index, traverse the career interest knowledge graph, determine the comprehensive recommendation degree that meets the interest recommendation degree associated with the student, index the career interest recommendation path corresponding to the comprehensive recommendation degree, and generate a personalized career interest matching path.
[0046] On the other hand, the present invention also provides a career planning analysis system based on multi-modal data fusion. The career planning analysis system based on multi-modal data fusion includes:
[0047] A knowledge graph construction module, which collects the original graph data based on web crawler technology, parses the original graph data and uploads the original graph data to the storage database, and pre-constructs a career interest knowledge graph based on the original graph data;
[0048] A three-dimensional career portrait module, which is used to obtain student-related data, perform normalization and fusion processing on the student-related data, parse the student-related data based on an improved NLP model, generate a three-dimensional career portrait of the student based on multi-modal data fusion, visually present the three-dimensional career portrait of the student, and upload the three-dimensional career portrait of the student to the storage database;
[0049] A path matching module, which is used to load the three-dimensional career portrait of the student in real time. The interest recommendation model traverses the career interest knowledge graph based on the frequent subgraph mining algorithm combined with the interactive point set, and generates a personalized career interest matching path.
[0050] Preferably, the knowledge graph construction module includes:
[0051] An original data scraping unit, which collects the original graph data based on web crawler technology and performs abnormal data cleaning on the original graph data
[0052] A data parsing unit, which is used to parse the original graph data and upload the original graph data to the storage database;
[0053] A knowledge graph construction unit, which constructs a career interest knowledge graph based on the fuzzy clustering algorithm, the ant colony algorithm, and the original graph data.
[0054] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:
[0055] In the embodiments of the present invention, by performing normalized fusion processing on student-related data, it is possible to comprehensively understand students from multiple dimensions. Moreover, by using multi-modal data fusion, an occupational interest knowledge graph, an improved NLP model, and a personalized recommendation model, a personalized occupational interest matching path can be dynamically generated, which can better adapt to the changes in students' interests and abilities. This overcomes the problems of the existing methods that use a big data analysis engine to perform iterative and correlation analysis on the original learning situation data. The data source is relatively single, lacking the fusion of multi-modal data such as students' occupational interests and external behaviors, being unable to dynamically adapt to the changes in students' interests, and having deficiencies in the depth and accuracy of analysis.
[0056] In the embodiments of the present invention, by using the BeautifulSoup crawler framework to capture the original data of the knowledge graph and performing abnormal data cleaning processing, it is possible to effectively remove noise data and error information, ensuring the quality and integrity of the data. And in the occupational interest knowledge graph, an occupational association layer is introduced, and digital tagging is performed on the attributes of courses, positions, and teachers from three perspectives of ability, interest, and value, enabling multi-dimensional correlation analysis. By calculating the correlation degree between the associated entities and the occupational association layer through a fuzzy clustering algorithm and generating an occupational interest recommendation path based on the ant colony algorithm, it is possible to provide students with accurate and objective personalized learning and career development paths, improving the scientificity and practicality of career planning.
[0057] In the embodiments of the present invention, an occupational interest spherical model is used as the initial architecture of the student's three-dimensional career portrait, and an improved NLP model is introduced. Based on the adaptive ensemble learning algorithm, prediction labels are introduced into the occupational interest spherical model to complete the construction of the student's three-dimensional career portrait. The improved NLP model can capture the bidirectional dependence relationships in the student-related data and consider the global consistency between the annotations by adding a BiLSTM feature extraction layer and a CRF annotation sequence layer. This significantly improves the model's processing ability for multi-modal data and can more accurately extract the features of the student-related data. The adaptive ensemble learning algorithm can introduce prediction labels in real time to optimize the parameters of the occupational interest spherical model, ensuring the accuracy and reliability of the portrait results. The multi-modal data fusion technology can integrate multi-dimensional data such as students' academic achievements, psychological evaluation results, internship evaluations, and basic information. This fusion method not only enriches the data source but also improves the accuracy of data analysis through the information complementation mechanism between different modal data.
[0058] In the embodiments of the present invention, when training an interest recommendation model using a training set and a test set, through Gaussian distribution initialization, Glorot algorithm, parallel iterative training, and whale optimization algorithm, efficient training of the model and parameter optimization are achieved. Based on the ALBERT model, a group-directed learning path network, spectral clustering algorithm, and class adversarial joint learning network are introduced. When training the interest recommendation model using the training set and the test set, efficient training of the model and parameter optimization are realized through Gaussian distribution initialization, Glorot algorithm, parallel iterative training, and whale optimization algorithm, providing a reasonable distribution for the initial parameters of the model, which helps the model converge to a better solution faster in the subsequent training process, avoid falling into local optima, and based on the ALBERT model, introducing a group-directed learning path network, spectral clustering algorithm, and class adversarial joint learning network helps extract potential structural information in the data, provides more accurate feature representations for the model, and further improves the generalization ability of the model. The idea of adversarial learning is introduced, enabling the model to generate both real sample representations and false sample representations different from the real samples while generating real sample representations. This competitive mechanism can enhance the feature learning ability and discriminative ability of the model, improving the discrimination of the model for different categories of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 FIG. is a schematic flowchart of the implementation of a career planning analysis method based on multi-modal data fusion provided by the present invention.
[0060] Figure 2 FIG. shows a schematic flowchart of the implementation of a method for pre-constructing a career interest knowledge graph based on original atlas data.
[0061] Figure 3 FIG. shows a schematic flowchart of the implementation of a method for generating a three-dimensional career portrait of a student based on multi-modal data fusion.
[0062] Figure 4 FIG. shows a schematic flowchart of the implementation of a method for training an interest recommendation model using a training set and a test set.
[0063] Figure 5 FIG. shows a schematic flowchart of the implementation of a method for generating a personalized career interest matching path.
[0064] Figure 6 FIG. shows a schematic structural diagram of a career planning analysis system based on multi-modal data fusion. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0066] Existing methods use a big data analysis engine to perform iterative and correlation analysis on the original learning situation data. The data source is relatively single, lacking the integration of multi-modal data such as students' career interests and external behaviors, and unable to dynamically adapt to the changes in students' interests. There are deficiencies in the depth and accuracy of analysis. To address the above problems, we propose a career planning analysis system and method based on multi-modal data fusion. Briefly, when the method is implemented, first, collect the original graph data based on web crawler technology, pre-construct a career interest knowledge graph based on the original graph data, then parse the student correlation data based on an improved NLP model, generate a three-dimensional career portrait of the student based on multi-modal data fusion, and visually present the three-dimensional career portrait of the student. At the same time, pre-construct an interest recommendation model based on collaborative filtering combined with the ALBERT model, train the interest recommendation model using a training set and a test set, and finally, the interest recommendation model traverses the career interest knowledge graph based on the frequent subgraph mining algorithm combined with the interaction point set to generate a personalized career interest matching path. In the embodiments of the present invention, by performing normalization fusion processing on the student correlation data, it is possible to comprehensively understand the student from multiple dimensions, and the use of multi-modal data fusion, career interest knowledge graph, improved NLP model, and personalized recommendation model can dynamically generate a personalized career interest matching path, which can better adapt to the changes in students' interests and abilities. It overcomes the problems of existing methods that use a big data analysis engine to perform iterative and correlation analysis on the original learning situation data, with a relatively single data source, lacking the integration of multi-modal data such as students' career interests and external behaviors, and being unable to dynamically adapt to the changes in students' interests, and having deficiencies in the depth and accuracy of analysis.
[0067] The embodiments of the present invention provide a career planning analysis method based on multi-modal data fusion. Figure 1 The schematic diagram of the implementation process of the career planning analysis method based on multi-modal data fusion is shown. The career planning analysis method based on multi-modal data fusion specifically includes:
[0068] Step S10, collect the original graph data based on web crawler technology, parse the original graph data and upload the original graph data to the storage database, and pre-construct a career interest knowledge graph based on the original graph data;
[0069] It should be noted that the storage database can be a MatrixOne database, an Activeloop Deep Lake database, or a Neo4j database. The storage database can store the original graph data of multi-source heterogeneity and student-related data. The student-related data refers to various data related to students, including but not limited to learning behavior data (such as grades, learning time, course selection, etc.), career interest data (such as interest test results, club activities participated in, etc.), and external behavior data (such as behavior records on online learning platforms, etc.). These data can comprehensively reflect the learning status and career orientation of students, providing a basis for personalized learning and career planning.
[0070] Step S20: Obtain student-related data, perform normalization and fusion processing on the student-related data, parse the student-related data based on an improved NLP model, generate a three-dimensional career portrait of the student based on multi-modal data fusion, visually present the three-dimensional career portrait of the student, and upload the three-dimensional career portrait of the student to the storage database.
[0071] Step S30: Traverse the storage database, extract a modeling sample set from the storage database, divide the modeling sample set into a training set and a test set, pre-construct an interest recommendation model based on collaborative filtering combined with the ALBERT model, and train the interest recommendation model using the training set and the test set.
[0072] Step S40: Load the three-dimensional career portrait of the student in real time. The interest recommendation model traverses the career interest knowledge graph based on the frequent subgraph mining algorithm combined with the interactive point set to generate a personalized career interest matching path.
[0073] In the embodiment of the present invention, by performing normalization and fusion processing on the student-related data, it is possible to comprehensively understand students from multiple dimensions. Moreover, by using multi-modal data fusion, career interest knowledge graph, improved NLP model, and personalized recommendation model, a personalized career interest matching path can be dynamically generated, which can better adapt to the changes in students' interests and abilities. It overcomes the problems of the existing methods that use a big data analysis engine to perform iterative and correlation analysis on the original learning situation data, the data source is relatively single, lacks the fusion of multi-modal data such as students' career interests and external behaviors, cannot dynamically adapt to the changes in students' interests, and has deficiencies in analysis depth and accuracy.
[0074] The embodiment of the present invention provides a method for pre-constructing a career interest knowledge graph based on the original graph data. Figure 2 The figure shows a schematic implementation flow diagram of the method for pre-constructing a career interest knowledge graph based on the original graph data. The method for pre-constructing a career interest knowledge graph based on the original graph data specifically includes:
[0075] Step S101, use the BeautifulSoup crawler framework to capture the original graph data and perform abnormal data cleaning on the original graph data. The methods of abnormal data cleaning include, but are not limited to, outlier deletion and outlier replacement;
[0076] It should be noted that the present invention uses the BeautifulSoup crawler framework to capture the original graph data. When the BeautifulSoup crawler framework captures the original graph data, it has the advantages of being simple and easy to use, strong fault tolerance, flexible query, and supporting multiple parsers. These characteristics make it perform well in processing multi-modal data, especially suitable for small-scale and medium-scale data crawling tasks. For tasks that require rapid development and processing of complex HTML structures.
[0077] Among them, the original graph data includes, but is not limited to, course management data, job information data, tutor association data, career planning data, development requirement data. The course association data includes course type, question bank data, teaching and research data, and textbook data. The job information data includes, but is not limited to, job type, number of positions, job requirements, job name, main responsibilities, educational requirements, certificate requirements, job difficulty, and job popularity. The tutor association data includes, but is not limited to, tutor name, tutor gender, tutor-student tutoring data, research topic data, remote teaching data, educational background data, work experience data, reward and punishment data, and workload data.
[0078] Step S102, pre-build the initial architecture of the knowledge graph based on course attributes, job attributes, and teacher attributes, and respectively build the data layer and schema layer based on course attributes, job attributes, and teacher attributes. Extract the original graph data after abnormal data cleaning in the form of entity-attribute-attribute value triples, and fill the structured data of the entity-attribute-attribute value triples associated with the original graph data into the data layer;
[0079] It should be noted that the data layer is the foundation of the knowledge graph, storing specific factual information, usually organized in the form of "entity-attribute-attribute value" triples. For example, in this embodiment, the data layer can store a course entity: science-natural science-0.4 (attribute value), a job entity: Chinese language and literature teacher-professor-0.5 (attribute value). The constraints of the schema layer on the data layer enable the knowledge graph to support complex reasoning and association analysis. For example, through the graph, it can be deduced which teachers are suitable for teaching specific courses, or which jobs are most compatible with the students' course backgrounds.
[0080] Step S103, traverse the data layer of the initial architecture of the knowledge graph, and fill the entity types, attributes, and relationships of the structured data into the schema layer of the initial architecture of the knowledge graph in a bottom-up manner;
[0081] It should be noted that the "Vocational Interest Knowledge Graph" is a structured knowledge framework that comprehensively describes the knowledge system in the field of vocational interest through the combination of entities, attributes, and relationships. The entity relationships of structured data include, but are not limited to, the relationships between courses and positions, students and vocational interests, teachers and courses, positions and vocational interests, and students and positions.
[0082] Step S104: Introduce a vocational association layer into the initial architecture of the knowledge graph from the perspectives of ability, interest, and value. The vocational association layer digitally tags the course attributes, position attributes, and teacher attributes of associated entities by combining the principal component analysis method from the perspectives of ability, interest, and value. Among them, the digital tags include course weights, position weights, and teacher weights.
[0083] In the embodiment of the present invention, when the vocational association layer digitally tags the course attributes, position attributes, and teacher attributes of associated entities by combining the principal component analysis method from the perspectives of ability, interest, and value, it can perform weighted fusion based on the principal component analysis method by respectively combining the label weight decay factors, behavior weights, data volumes, and error terms corresponding to the course attributes, position attributes, and teacher attributes, and finally output the digital tags of the course attributes, position attributes, and teacher attributes of the associated entities.
[0084] It should be noted that the introduction of the vocational association layer aims to deeply integrate the courses and teachers in the education field with the position requirements in the vocational field, and establish an association relationship through three dimensions of ability, interest, and value. This design can help students and educators more clearly understand the mapping relationship between courses and occupations, thus facilitating the construction of a personalized vocational interest knowledge graph.
[0085] Step S105: Calculate the association degree between the associated entity and the vocational association layer based on the fuzzy clustering algorithm;
[0086] In this embodiment, the association degree between the associated entity and the vocational association layer is calculated based on the fuzzy clustering algorithm through the following formula:
[0087]
[0088] Among them, Sim ij represents the association degree between the associated entity i and the vocational association layer j. I and J are the numbers of the associated entity i and the vocational association layer j respectively. In this embodiment, the number of associated entities can be 1 - 100, and the number of the vocational association layer can be 10 - 200. |i∩j| represents the interaction number between the associated entity i and the vocational association layer j based on the fuzzy clustering algorithm. α 1 , α 2 , α 3 are respectively the course weight, position weight, and teacher weight of the associated entity. KPD iIt represents the interaction degree of associated entity i obtained based on the fuzzy clustering algorithm, M ij is the association matrix between associated entity i and the occupational association layer j, m i is the clustering center of the clustering analysis of associated entity i based on the fuzzy clustering algorithm.
[0089] Step S106, determine whether the association degree between the associated entity and the occupational association layer exceeds the preset association threshold. In this embodiment, the preset association threshold can be 0.25 - 0.75;
[0090] Step S107, if the association degree between the associated entity and the occupational association layer exceeds the preset association threshold, extract the corresponding associated entity;
[0091] If the association degree between the associated entity and the occupational association layer does not exceed the preset association threshold, it is determined that the association degree between the associated entity and the occupational association layer is low, and the current associated entity is not extracted.
[0092] Step S108, load at least one group of associated entities that exceed the preset association threshold, generate an associated entity set, supplement the associated entity set to the occupational association layer, generate a career interest recommendation path based on the ant colony algorithm, and calculate the comprehensive recommendation degree of the career interest recommendation path;
[0093] In this embodiment, a career interest recommendation path is generated based on the ant colony algorithm, and the comprehensive recommendation degree of the career interest recommendation path is calculated through the following formula:
[0094]
[0095] where, L z represents the comprehensive recommendation degree of the career interest recommendation path, K(i), G(i), S(i) are the course information vector, job information vector, and teacher information vector respectively, A y is the comprehensive heuristic factor based on the ant colony algorithm, L represents all search paths based on the ant colony algorithm. In this embodiment, all search paths can be 1 - 20, κ represents the ant colony obstacle avoidance factor, which can be 0.05, β is the ant colony expectation factor, which can be 0.2 - 0.8, K, G, S are the number of courses, the number of jobs, and the number of teachers respectively.
[0096] Step S109, integrate the occupational association layer including the associated entity set, the career interest recommendation path, and the comprehensive recommendation degree to complete the construction of the career interest knowledge graph.
[0097] In the embodiments of the present invention, by using the BeautifulSoup crawler framework to capture the original data of the knowledge graph and performing abnormal data cleaning, noise data and error information can be effectively removed, ensuring the quality and integrity of the data. Moreover, a career association layer is introduced into the career interest knowledge graph, and digital tagging is performed on the attributes of courses, positions, and teachers from three perspectives: ability, interest, and value, enabling multi-dimensional association analysis. By calculating the association degree between the associated entities and the career association layer through the fuzzy clustering algorithm and generating a career interest recommendation path based on the ant colony algorithm, accurate and objective personalized learning and career development paths can be provided for students, improving the scientificity and practicality of career planning.
[0098] The embodiments of the present invention provide a method for generating a three-dimensional career portrait of students based on multi-modal data fusion. Figure 3 The figure shows a schematic implementation flow diagram of the method for generating a three-dimensional career portrait of students based on multi-modal data fusion. The method for generating a three-dimensional career portrait of students based on multi-modal data fusion specifically includes:
[0099] Step S201, load the student associated data;
[0100] It should be noted that the student associated data includes, but is not limited to, multi-modal data such as students' academic achievements, psychological evaluation results, internship evaluations, and students' basic information. Using the multi-modal data of student associated data to construct a three-dimensional career portrait of students can objectively and accurately generate a three-dimensional career portrait of students. The student associated data realizes comprehensive coverage in learning performance, career interests, potential abilities, and external behaviors.
[0101] Step S202, load the NLP model, add a BiLSTM feature extraction layer and a CRF annotation sequence layer to the NLP model, and introduce an adaptive ensemble learning algorithm based on time-stratified sampling to complete the improvement of the NLP model and output an improved NLP model;
[0102] Step S203, the improved NLP model performs dimensionality reduction processing on the student associated data, maps the dimensionality-reduced student associated data to a dimensionality-reduced feature space, and the improved NLP model performs annotation and weight scoring on the student associated data in the dimensionality-reduced feature space to obtain a dimensionality-reduced feature space containing student associated data, data annotations, and data weights;
[0103] Step S204, pre-construct a career interest spherical model based on ability, interest, and value, traverse the dimensionality-reduced feature space, and fill the student associated data, data annotations, and data weights into the career interest spherical model;
[0104] It should be noted that the spherical model of career interests integrates Holland's RIASEC theory and Prediger's two-dimensional structure model, and introduces "prestige" as the third dimension. This model comprehensively evaluates students' career interests through three dimensions: people / things, data / ideas, and prestige, and can more accurately depict students' career tendencies. When pre-constructing the spherical model of career interests based on ability, interest, and value, the ability dimension reflects students' performance in academics, skills, and practice, including course grades, mastery of professional skills, internship performance, etc.; the interest dimension reflects students' career interest tendencies, usually based on data such as psychological evaluations, career interest tests, and preferences for participating in activities; and the value dimension reflects students' value orientations towards career goals, work environments, career development, etc., including expectations for salary, career achievement, work-life balance, etc.
[0105] In step S205, the BiLSTM feature extraction layer in the NLP model extracts features from the student-related data and data annotations, outputs the predicted labels corresponding to the student-related data, and introduces the predicted labels into the spherical model of career interests based on the adaptive ensemble learning algorithm to complete the construction of the student's three-dimensional career portrait and visually present the student's three-dimensional career portrait.
[0106] In this embodiment, the "student's three-dimensional career portrait" is a comprehensive student feature description model constructed based on three dimensions: ability, interest, and value. By constructing the student's three-dimensional career portrait, one can clearly understand the student's own abilities, interests, and value orientations, thereby helping the student choose the most suitable career path.
[0107] In the embodiment of the present invention, the spherical model of career interests is used as the initial architecture of the student's three-dimensional career portrait, and an improved NLP model is introduced. The predicted labels are introduced into the spherical model of career interests based on the adaptive ensemble learning algorithm to complete the construction of the student's three-dimensional career portrait. The improved NLP model can capture the bidirectional dependencies in the student-related data and consider the global consistency between the annotations by adding a BiLSTM feature extraction layer and a CRF annotation sequence layer. It significantly improves the model's processing ability for multi-modal data and can more accurately extract the features of the student-related data. The adaptive ensemble learning algorithm can introduce the predicted labels in real time to optimize the parameters of the spherical model of career interests and ensure the accuracy and reliability of the portrait results. The multi-modal data fusion technology can integrate multi-dimensional data such as students' academic achievements, psychological evaluation results, internship evaluations, and basic information. This fusion method not only enriches the data sources but also improves the accuracy of data analysis through the information complementary mechanism between different modal data.
[0108] The embodiment of the present invention provides a method for training an interest recommendation model using a training set and a test set. Figure 4The figure shows a schematic implementation process of a method for training an interest recommendation model using a training set and a test set. The method for training an interest recommendation model using a training set and a test set specifically includes:
[0109] Step S301: Load the training set and the pre-constructed interest recommendation model. Initialize the parameters of the embedding layer of the interest recommendation model using a Gaussian distribution, and initialize the parameters of the hidden layer of the interest recommendation model in combination with the Glorot algorithm. Sample the samples in the training set based on the replacement strategy, and perform parallel iterative training on the interest recommendation model to output at least one set of sub-recommendation models.
[0110] In this embodiment, when dividing the modeling sample set into a training set and a test set, the ratio of the training set to the test set can be 4:1 or 3:1, and the modeling sample set can be obtained through a personnel management system, a teaching affairs system, a scientific research management system, a job hunting website, and a faculty database system. Initializing the parameters of the embedding layer using a Gaussian distribution can provide a reasonable initial weight distribution for the model, avoiding the problems of gradient explosion or disappearance caused by too large or too small initial weights. The replacement strategy can ensure the independence of each sampling, increase the diversity of samples, and reduce the overfitting risk caused by insufficient sample quantity.
[0111] Step S302: Update the parameter convergence factor and the non-linear activation function of the sub-recommendation model based on the whale optimization algorithm to obtain the sub-recommendation model after updating the parameter convergence factor and the non-linear activation function.
[0112] In this embodiment, dynamically adjusting the parameter convergence factor and the non-linear activation function through WOA can better adapt to the requirements of different training stages and improve the adaptability and flexibility of the model.
[0113] Step S303: Calculate the model fitness of the sub-recommendation model using the Gaussian distribution function to obtain the model fitness of at least one set of sub-recommendation models, and select the sub-recommendation model with the minimum model fitness as the interest recommendation model; select the sub-recommendation model with the minimum model fitness as the final interest recommendation model to ensure the optimal performance of the model on the training set. Thus, it provides a high-quality model basis for subsequent testing and application.
[0114] Among them, the model fitness of the sub-recommendation model is calculated by the following formula:
[0115]
[0116] Among them, f(λ t+1 ,σ t+1 ) represents the model fitness of the sub-recommendation model, and λ t+1 ,σ t+1 are respectively the parameter convergence factor and the non-linear activation function hyperparameter after the (t + 1)-th iteration update, and λ tis the parameter convergence factor after the t-th iteration update, λ 0 is the initial parameter convergence factor, X t represents the number of sub-recommendation models, and Δτ is the dimension of the whale population; the hyperparameters of the non-linear activation function refer to the parameters used to adjust the behavior of the activation function. These parameters can affect the shape, output range, or gradient characteristics of the activation function. For example, in some activation functions, the hyperparameters can control the "steepness" or saturation value of the activation function. By adjusting these hyperparameters, the training process and performance of the model can be optimized. The parameter convergence factor is a parameter that controls the weight update rate and stability during the iterative process of the optimization algorithm. In swarm intelligence algorithms such as the whale optimization algorithm, the convergence factor is usually used to adjust the global search ability and local search ability of the algorithm. The parameter convergence factor after iterative update will follow a preset rule, and the initial parameter convergence factor is the initial value of the convergence factor when the optimization algorithm starts to iterate. For example, the initial convergence factor is usually set to 2, and the final convergence factor is set to 0.2. The number of sub-recommendation models refers to the number of multiple recommendation models generated during the training process. These models are usually obtained through parallel iterative training and are used to provide diverse candidate models. By selecting the sub-recommendation model with the highest fitness, it can be ensured that the final model has good performance.
[0117] Step S304, load the test set, use the test set as the input, execute the interest recommendation model, output the test result, and calculate the cross-entropy loss between the test result and the true result using the cross-entropy loss function;
[0118] Step S305, determine whether the cross-entropy loss meets the preset loss threshold. In this embodiment, the loss threshold can be 0.05 - 0.15;
[0119] Step S306, if the cross-entropy loss meets the preset loss threshold, output the converged interest recommendation model;
[0120] If the cross-entropy loss does not meet the preset loss threshold, use the Adam optimizer combined with the collaborative filtering algorithm to adjust the learning rate of the interest recommendation model, the parameters of the embedding layer, and the parameters of the hidden layer, and return to step S301 to continue the parallel iterative training of the interest recommendation model.
[0121] In this embodiment, the interest recommendation model uses the ALBERT model as the initial model, freezes the feed-forward network in the initial model, replaces the feed-forward network of the initial model with a group-directed learning path network, introduces a spectral clustering algorithm into the group-directed learning path network, adds a category adversarial joint learning network after the initial model. The category adversarial joint learning network includes a category classifier, a category discriminator, and a graph spectrum scoring device. A frequent subgraph mining algorithm is introduced into the graph spectrum scoring device. The category classifier, the category discriminator, and the graph spectrum scoring device are connected by neural nodes. The initial model consists of an input layer, a hidden layer, and an embedding layer. The hidden layer has 768 layers, and the number of attention heads of the initial model is 16 or 32.
[0122] In the embodiment of the present invention, when training the interest recommendation model using the training set and the test set, through Gaussian distribution initialization, Glorot algorithm, parallel iterative training, and whale optimization algorithm, efficient training of the model and parameter optimization are achieved. Based on the ALBERT model, a group-directed learning path network, a spectral clustering algorithm, and a category adversarial joint learning network are introduced. When training the interest recommendation model using the training set and the test set, efficient training of the model and parameter optimization are achieved through Gaussian distribution initialization, Glorot algorithm, parallel iterative training, and whale optimization algorithm, providing a reasonable distribution for the initial parameters of the model, helping the model converge to a better solution faster in the subsequent training process, avoiding falling into local optima. Based on the ALBERT model, a group-directed learning path network, a spectral clustering algorithm, and a category adversarial joint learning network are introduced, which helps to extract potential structural information in the data, provides a more accurate feature representation for the model, and further improves the generalization ability of the model. The idea of adversarial learning is introduced, enabling the model to generate representations of real samples while also generating representations of fake samples different from the real samples. This competitive mechanism can enhance the feature learning ability and discriminative ability of the model, improving the discrimination of the model for different categories of data.
[0123] The embodiment of the present invention provides a method for generating a personalized career interest matching path. Figure 5 The figure shows a schematic implementation flowchart of the method for generating a personalized career interest matching path. The method for generating a personalized career interest matching path specifically includes:
[0124] Step S401, load the student's three-dimensional career portrait in real time. The interest recommendation model constructs at least one group of student-graph knowledge point interaction matrices based on the student's three-dimensional career portrait.
[0125] In the embodiment of the present invention, constructing a student-graph knowledge point interaction matrix based on the student's three-dimensional career portrait can comprehensively reflect the degree of association between the student and courses, teachers, and positions. This multi-dimensional interaction matrix not only considers the student's academic performance but also combines interests and value orientations.
[0126] Step S402: Calculate the student-graph knowledge point interaction matrix to determine the course interaction degree, teacher interaction degree, and position interaction degree, and integrate the course interaction degree, teacher interaction degree, and position interaction degree to generate an interaction point set;
[0127] It should be noted that by calculating the course interaction degree, teacher interaction degree, and position interaction degree, the association strength between students and different knowledge points can be quantified. This quantification method enables the recommendation system to clarify the students' interests and ability levels in different fields. Thus, it provides an accurate quantitative basis for personalized recommendations, avoiding errors caused by subjective judgments. Integrating the interaction degrees of different dimensions into an interaction point set can provide a clear path starting point for subsequent graph traversal. This integration method not only retains multi-dimensional information but also improves the operability of the data.
[0128] Step S403: The interest recommendation model traverses the career interest knowledge graph based on the frequent subgraph mining algorithm combined with the interaction point set. The category adversarial joint learning network identifies the interest recommendation degree associated with the students in the student-graph knowledge point interaction matrix. The frequent subgraph mining algorithm can identify knowledge points and career paths with high association degrees with students from the career interest knowledge graph. This algorithm can discover hidden patterns and association relationships in the graph;
[0129] Step S404: Using the interest recommendation degree associated with the students as an index, traverse the career interest knowledge graph to determine the comprehensive recommendation degree that meets the interest recommendation degree associated with the students, index the career interest recommendation path corresponding to the comprehensive recommendation degree, and generate a personalized career interest matching path.
[0130] In this embodiment, using the interest recommendation degree associated with the students as an index can quickly locate the career path that best matches the students' interests and abilities. This indexing method not only improves the recommendation efficiency but also ensures the accuracy of the recommendation results.
[0131] The embodiment of the present invention provides a career planning analysis system based on multi-modal data fusion. Figure 6 The structure diagram of the career planning analysis system based on multi-modal data fusion is shown. The career planning analysis system based on multi-modal data fusion specifically includes:
[0132] The knowledge graph construction module 100 collects the original graph data based on the crawler technology, analyzes the original graph data and uploads the original graph data to the storage database, and pre-constructs the career interest knowledge graph based on the original graph data;
[0133] The three-dimensional career portrait module 200 is used to obtain student-related data, perform normalization and fusion processing on the student-related data, parse the student-related data based on an improved NLP model, generate a three-dimensional career portrait of the student based on multi-modal data fusion, visually present the three-dimensional career portrait of the student, and upload the three-dimensional career portrait of the student to the storage database;
[0134] The path matching module 300 is used to load the three-dimensional career portrait of the student in real time. The interest recommendation model generates a personalized career interest matching path based on the frequent subgraph mining algorithm combined with the interactive point set to traverse the career interest knowledge graph.
[0135] In this embodiment, the knowledge graph construction module 100 includes:
[0136] The original data scraping unit 110 collects the original data of the graph based on web crawler technology and performs abnormal data cleaning on the original data of the graph
[0137] The data parsing unit 120 is used to parse the original data of the graph and upload the original data of the graph to the storage database;
[0138] The knowledge graph construction unit 130 constructs a career interest knowledge graph based on the fuzzy clustering algorithm, ant colony algorithm, and original data of the graph.
[0139] It can be understood that the career planning analysis system based on multi-modal data fusion provided by the embodiments of the present invention corresponds to the above-mentioned career planning analysis method based on multi-modal data fusion. For the explanations, examples, beneficial effects, etc. of the relevant content, reference can be made to the corresponding content in the career planning analysis method based on multi-modal data fusion, which will not be elaborated here.
[0140] In summary, the present invention provides a career planning analysis system and method based on multi-modal data fusion. In the embodiments of the present invention, by performing normalization and fusion processing on student-related data, students can be comprehensively understood from multiple dimensions. Moreover, the use of multi-modal data fusion, career interest knowledge graph, improved NLP model, and personalized recommendation model can dynamically generate personalized career interest matching paths, which can better adapt to the changes in students' interests and abilities. It overcomes the problems of the existing methods that use a big data analysis engine to perform iteration and correlation analysis on the original learning situation data, the data source is relatively single, lacks the fusion of multi-modal data such as students' career interests and external behaviors, cannot dynamically adapt to the changes in students' interests, and is insufficient in analysis depth and accuracy.
[0141] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on these embodiments, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of the present invention to be protected. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art can still, without conflict and without making creative efforts, combine, add or delete the features in the embodiments of the present invention according to the circumstances or make other adjustments, so as to obtain different technical solutions that do not essentially depart from the concept of the present invention, and these technical solutions also fall within the scope of the present invention to be protected.
Claims
1. A career planning analysis method based on multimodal data fusion, characterized in that: The career planning analysis method based on multimodal data fusion includes: Collect the original data of the graph based on crawler technology, parse the original data of the graph and upload the original data of the graph to the storage database, and pre-build the career interest knowledge graph based on the original data of the graph; Obtain student-related data, normalize and fuse the student-related data, parse the student-related data based on the improved NLP model, generate a three-dimensional career portrait of the student based on multimodal data fusion, visualize the three-dimensional career portrait of the student, and upload the three-dimensional career portrait of the student to the storage database; Traverse the storage database, extract the modeling sample set from the storage database, divide the modeling sample set into a training set and a test set, pre-build an interest recommendation model based on collaborative filtering combined with the ALBERT model, and use the training set and the test set to train the interest recommendation model; The student's three-dimensional career portrait is loaded in real time, and the interest recommendation model traverses the career interest knowledge graph based on the frequent subgraph mining algorithm combined with the interaction point set to generate a personalized career interest matching path.
2. The career planning analysis method based on multimodal data fusion according to claim 1, characterized in that: The method for pre-constructing a career interest knowledge graph based on original graph data includes: The BeautifulSoup crawler framework is used to crawl the original data of the graph and perform abnormal data cleaning on the original data of the graph. The original data of the graph includes course management data, job information data, tutor-related data, career planning data, and development needs data. Pre-build the initial architecture of the knowledge graph based on course attributes, job attributes, and teacher attributes, and build the data layer and model layer based on course attributes, job attributes, and teacher attributes respectively. Extract the original graph data after abnormal data cleaning in the form of entity-attribute-attribute value triples, and fill the entity-attribute-attribute value triple structured data associated with the original graph data into the data layer; Traverse the data layer of the initial architecture of the knowledge graph, and fill the entity types, attributes, and relationships of the structured data into the model layer of the initial architecture of the knowledge graph in a bottom-up manner; Based on the ability, interest and value perspectives, a career association layer is introduced into the initial architecture of the knowledge graph. The career association layer digitally labels the course attributes, job attributes and teacher attributes of the associated entities based on the ability, interest and value perspectives combined with the principal component analysis method. The digital labels include course weights, job weights and teacher weights.
3. The career planning analysis method based on multimodal data fusion according to claim 2, characterized in that: The method for pre-constructing a career interest knowledge graph based on the original graph data further includes: Calculate the correlation between the associated entity and the occupational correlation layer based on the fuzzy clustering algorithm, determine whether the correlation between the associated entity and the occupational correlation layer exceeds a preset correlation threshold, and if the correlation between the associated entity and the occupational correlation layer exceeds the preset correlation threshold, extract the corresponding associated entity; Load at least one group of related entities exceeding a preset related threshold, generate a related entity set, add the related entity set to the career related layer, generate a career interest recommendation path based on an ant colony algorithm, and calculate the comprehensive recommendation degree of the career interest recommendation path; Integrate the career association layer including the associated entity set, career interest recommendation path, and comprehensive recommendation degree to complete the construction of the career interest knowledge graph.
4. The career planning analysis method based on multimodal data fusion according to claim 3, characterized in that: The correlation degree between the associated entity and the occupational correlation layer is calculated based on the fuzzy clustering algorithm, and is calculated by the following formula: Among them, Sim ij represents the correlation between associated entity i and occupational association layer j, I and J are the number of associated entities i and occupational association layer j, |i∩j| represents the number of interactions between associated entity i and occupational association layer j based on fuzzy clustering algorithm, α1, α2, α3 are the course weight, position weight, and teacher weight of the associated entity, respectively, KPD i represents the interaction degree of associated entity i obtained based on fuzzy clustering algorithm, M ij is the association matrix between the associated entity i and the occupational association layer j, m i is the cluster center of the cluster analysis of the associated entity i based on the fuzzy clustering algorithm; Based on the ant colony algorithm, a career interest recommendation path is generated, and the comprehensive recommendation degree of the career interest recommendation path is calculated using the following formula: Among them, L z represents the comprehensive recommendation degree of the recommended path of career interest, K(i), G(i), S(i) are the course information vector, position information vector, and teacher information vector respectively, A y is the comprehensive heuristic factor based on the ant colony algorithm, L represents all search paths based on the ant colony algorithm, κ represents the ant colony obstacle avoidance factor, β is the ant colony expectation factor, K, G, and S are the number of courses, the number of positions, and the number of teachers, respectively.
5. The career planning analysis method based on multimodal data fusion according to claim 1, characterized in that: The method for generating a three-dimensional career portrait of a student based on multimodal data fusion specifically includes: Load student-related data, including student academic performance, psychological assessment results, internship evaluation, and student basic information; Load the NLP model, add the BiLSTM feature extraction layer and the CRF annotation sequence layer to the NLP model, introduce the adaptive ensemble learning algorithm based on time stratified sampling, improve the NLP model, and output the improved NLP model; The improved NLP model performs dimensionality reduction processing on the student-related data, maps the reduced-dimensional student-related data to the reduced-dimensional feature space, and labels and weights the student-related data in the reduced-dimensional feature space to obtain a reduced-dimensional feature space containing student-related data, data labels, and data weights; Pre-construct a spherical model of career interests based on abilities, interests, and values, traverse the dimension reduction feature space, and fill the student-related data, data annotations, and data weights into the spherical model of career interests; The BiLSTM feature extraction layer in the NLP model extracts features from student-related data and data annotations, outputs the predicted labels corresponding to the student-related data, and introduces the predicted labels into the career interest spherical model based on the adaptive ensemble learning algorithm to complete the construction of the student's three-dimensional career portrait and visualize the student's three-dimensional career portrait.
6. The career planning analysis method based on multimodal data fusion according to claim 5, characterized in that: The method of using a training set and a test set to train an interest recommendation model specifically includes: Load the training set and the pre-built interest recommendation model, use Gaussian distribution to initialize the parameters of the embedding layer of the interest recommendation model, combine the Glorot algorithm to initialize the parameters of the hidden layer of the interest recommendation model, collect samples from the training set based on the replacement strategy, iterate and train the interest recommendation model in parallel, and output at least one set of sub-recommendation models; Based on the whale optimization algorithm, the parameter convergence factor and nonlinear activation function of the sub-recommendation model are updated to obtain the sub-recommendation model after the parameter convergence factor and nonlinear activation function are updated; The model fitness of the sub-recommendation model is calculated by using the Gaussian distribution function, the model fitness of at least one group of sub-recommendation models is obtained, and the sub-recommendation model with the smallest model fitness is selected as the interest recommendation model; Among them, the model fitness of the sub-recommendation model is calculated by the following formula: Among them, f(λ t+1 ,σ t+1 ) represents the model fitness of the sub-recommendation model, λ t+1 ,σ t+1 are the parameter convergence factor and nonlinear activation function hyperparameter after t+1 iteration updates, respectively. t is the parameter convergence factor after the tth iteration update, λ0 is the initial parameter convergence factor, X t represents the number of sub-recommendation models, Δτ is the dimension of the whale population; Load the test set, use the test set as input, execute the interest recommendation model, output the test results, use the cross entropy loss function to calculate the cross entropy loss between the test results and the real results, and determine whether the cross entropy loss meets the preset loss threshold. If the cross entropy loss meets the preset loss threshold, output the converged interest recommendation model; If the cross entropy loss does not meet the preset loss threshold, the Adam optimizer combined with the collaborative filtering algorithm is used to adjust the learning rate, embedding layer parameters, and hidden layer parameters of the interest recommendation model, and the interest recommendation model is continued to be iteratively trained in parallel.
7. The career planning analysis method based on multimodal data fusion according to claim 6, characterized in that: The interest recommendation model uses the ALBERT model as the initial model, freezes the feedforward network in the initial model, adopts a group directed learning path network to replace the feedforward network of the initial model, introduces a spectral clustering algorithm into the group directed learning path network, and adds a category adversarial joint learning network after the initial model. The category adversarial joint learning network includes a category classifier, a category discriminator, and a graph scorer. A frequent subgraph mining algorithm is introduced into the graph scorer.
8. The career planning analysis method based on multimodal data fusion according to claim 7, characterized in that: The method for generating a personalized career interest matching path specifically includes: The student's three-dimensional career portrait is loaded in real time, and the interest recommendation model constructs at least one set of student-atlas knowledge point interaction matrix based on the student's three-dimensional career portrait; Calculate the student-graph knowledge point interaction matrix to determine the course interaction degree, teacher interaction degree, and job interaction degree, and integrate the determined course interaction degree, teacher interaction degree, and job interaction degree to generate an interaction point set; The interest recommendation model is based on the frequent subgraph mining algorithm combined with the interaction point set to traverse the career interest knowledge graph, and the category adversarial joint learning network identifies the interest recommendation degree associated with students in the student-graph knowledge point interaction matrix; Taking the interest recommendation degree associated with the student as the index, traverse the career interest knowledge graph, determine the comprehensive recommendation degree that matches the interest recommendation degree associated with the student, index the career interest recommendation path corresponding to the comprehensive recommendation degree, and generate a personalized career interest matching path.
9. A career planning analysis system based on multimodal data fusion, used to implement the career planning analysis method based on multimodal data fusion as claimed in any one of claims 1 to 8, characterized in that: The career planning analysis system based on multimodal data fusion includes: The knowledge graph construction module collects the original graph data based on crawler technology, parses the original graph data and uploads the original graph data to the storage database, and pre-constructs the career interest knowledge graph based on the original graph data; The three-dimensional career portrait module is used to obtain student-related data, normalize and fuse the student-related data, parse the student-related data based on the improved NLP model, generate the student's three-dimensional career portrait based on multimodal data fusion, visualize the student's three-dimensional career portrait, and upload the student's three-dimensional career portrait to the storage database; The path matching module is used to load students' three-dimensional career portraits in real time. The interest recommendation model traverses the career interest knowledge graph based on the frequent subgraph mining algorithm combined with the interaction point set to generate a personalized career interest matching path.
10. The career planning analysis system based on multimodal data fusion according to claim 9, characterized in that: The knowledge graph construction module includes: The original data capture unit collects the original data of the graph based on crawler technology and performs abnormal data cleaning on the original data of the graph A data analysis unit, used for analyzing the original data of the spectrum and uploading the original data of the spectrum to a storage database; The knowledge graph construction unit constructs a career interest knowledge graph based on fuzzy clustering algorithm, ant colony algorithm, and graph raw data.
Citation Information
Patent Citations
Personalized learning career data portraying method and system
CN116433427A
Cited By
Talent recommendation system for matching local productivity based on vocational education
CN120851469A