Professional recommendation knowledge graph construction method and query system based on large language model
A large-scale language model-based method constructs a major recommendation knowledge graph, addressing the mismatch in current systems by associating majors with personality traits and qualities, providing tailored career recommendations.
Patent Information
- Application Number
- JP2025019468
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-07
- Filing Date
- 2025-02-07
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2045-02-07
Smart Images

Figure 2025121893000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of knowledge graphs, and more particularly to a method and search system for building a major recommendation knowledge graph based on a large-scale language model. [Background technology]
[0002] The college entrance exam is an important process that many people go through, and choosing a major is particularly important when applying. The major you choose will mean different future career paths and will require different personality traits and qualities. Therefore, to determine which major is suitable for you, you need to understand your own personality and the qualities that correspond to the major. You also need to understand which personality traits and qualities you need to improve in order to better adapt to the needs of the major you have chosen.
[0003] However, there is currently a lack of intelligent major recommendation service systems on the Internet that meet this need, and they are unable to maximize the correlation between majors and personality and qualification information, resulting in inconveniences in major selection and career planning. Prior art patent application number CN202010309835.0 discloses a method, apparatus, computer equipment, and storage medium for recommending majors based on academic subjects, which includes: receiving a major recommendation command, the major recommendation command having user identification information, and obtaining user answer data corresponding to the user identification based on the major recommendation command; extracting corresponding subject knowledge characteristics, skill characteristics, thinking ability characteristics, and personality characteristics from the user answer data; inputting the subject knowledge characteristics, skill characteristics, thinking ability characteristics, and personality characteristics into a major recommendation model to obtain an output identification result; obtaining corresponding recommended majors based on the output identification result, and sending the recommended majors to a user terminal corresponding to the user identification information for display. However, this technical solution does not take into account the relationship between the user's own personality characteristics and the personality characteristics required for the major during the recommendation process, so the recommended majors often do not match the user's own personality characteristics. Therefore, how to quickly find the correspondence between majors and personality traits and qualities is an issue that needs to be addressed urgently. Summary of the Invention
[0004] To solve the problems in the prior art, the present invention provides a method and a search system for building a specialty recommendation knowledge graph based on a large-scale language model.
[0005] To achieve the object of the present invention, the following solutions are adopted. In a first aspect, the present invention provides a method for constructing a major recommendation knowledge graph based on a large-scale language model, the method being executed as a computer program in a computer electronic device including a memory for storing a computer program and a processor, the method comprising the steps of: S1: For each major within the recommended range, obtain the standard major name and a description of the personality and qualities required for that major. For each occupation within the recommended range, obtain the standard occupation name and a description of the personality and qualities required for that occupation. For the personality test results that belong to each personality category in the personality test, obtain the standard personality test result name and a description of the personality and qualities corresponding to the personality test result. S2: By setting prompts for input information to the large-scale language model, we build an entity extraction model suitable for completing the task of building a major recommendation knowledge graph. It is necessary to add specific inference examples to the set prompts. At the same time, it is necessary to define the categories of extracted entities and relationships, output each hierarchical concept hierarchically, and specify the output entity attribute format. S3: Using the entity extraction model, extract text entities from the description text of the personality and qualities required for each major within the recommended range, the description text of the personality and qualities required for each occupation corresponding to each profession, and the description text of the personality and qualities corresponding to each personality test result, and from these, extract the personality and quality entities required for each major and the names of occupations suitable for the person, the personality and quality entities required for each occupation, and the personality and quality entities possessed by each personality test result. S4: All majors, all occupations, all extracted personality and quality entities, and all personality test results within the recommended range are treated as nodes in a graph, and relationships between the nodes are added based on the entity extraction results in S3. A multi-stage weighting algorithm is used to obtain the degree of association between majors and personality and qualities, between occupations and personality and qualities, and between personality test results and personality and qualities. Finally, personality and qualities are treated as connections between majors and personality test results, and a knowledge graph for major recommendation is constructed.
[0006] In the first aspect, preferably, the description text of the personality and qualities required for the major, the description text of the personality and qualities required for the occupation, and the description text of the personality and qualities corresponding to the personality test results are all derived from the headword content of the "Encyclopedia" website, and the headword content needs to pre-process the text to meet the input requirements of entity extraction.
[0007] In the first aspect, the personality test is preferably one or more of the MBTI test, the Holland Occupational Personality Test, and the Multiple Intelligences Test.
[0008] In the first aspect, preferably, when constructing an entity extraction model in S2, it is necessary to constantly optimize and adjust the prompts input into the large-scale language model to achieve the entity extraction task, and the optimization and adjustment process is as follows: S21: First, samples are extracted from all written texts from which entity extraction is performed, and target entities in all written text samples are manually marked up to obtain sample data. S22: Add specific inference examples to the prompts of the large-scale language model, define the categories of extracted entities and relations, and require the large-scale language model to output each hierarchical concept hierarchically, while specifying the entity attribute format to be output by the large-scale language model. After that, test the large-scale language model after the prompts are constructed on sample data, and further adjust the prompts based on the test results. After the accuracy of the large-scale language model's extraction of text entities meets the predetermined required standards, use the final prompts in the large-scale language model to construct the entity extraction model.
[0009] In the first aspect, preferably, the description text input into the large-scale language model needs to be marked up in advance using the large-scale language model to mark up the document structure, and the advantage of the large-scale language model in processing large amounts of data is utilized, and a prompt is constructed to allow the large-scale language model to determine the text level in the headword content based on the dependency relationship and perform markup, wherein the main text in the headword content is marked up to level 0, and the Nth level headword in the headword content is correspondingly marked up to level N.
[0010] In the first aspect, preferably, the multi-stage weighting algorithm calculates the relevance and the process of constructing the major recommendation knowledge graph is as follows: S41: For each major, statistics are taken of the first frequency at which different character and quality entities appear in the written text with respect to the character and quality required for the major corresponding to that major, and the first frequency represents the first degree of association between the corresponding major and the character and quality. S42: For each major, determine the name of an occupation suitable for the major from the extracted entities; further, for the occupation suitable for each major, obtain statistics of a second frequency of appearance of different character and quality entities in the description text with respect to the character and quality required for the occupation corresponding to the occupation, and represent a second association degree between the major and the character and quality by the second frequency. S43: For each combination of major and character and quality, the weighted sum of the primary relevance between that major and that character and quality and the secondary relevance between that character and quality and all occupations suitable for that major is calculated to arrive at the final relevance between that major and that character and quality. S44: For each personality test result, statistics are calculated on the third frequency at which different personality and quality entities appear in the written text with respect to the personality and quality corresponding to the personality test result, and the third frequency represents the degree of association between the corresponding personality test result and the personality and quality. S45: All majors, all occupations, all extracted personality and quality entities, and all personality test results are treated as nodes in a graph. Based on the entity extraction results in S3, edge links are established between each major and the occupation suitable for that major, between each major and the personality and qualities required for that major, between each occupation and the personality and qualities required for that occupation, and between personality test results and the personality and qualities possessed by the personality test results. At the same time, weights are assigned to the edges of corresponding nodes based on the final association between the major and the personality and qualities obtained in S43. Weights are assigned to the edges of corresponding nodes based on the association between the personality test results and the personality and qualities obtained in S44, and the personality and qualities are linked to the majors and personality test results, thereby constructing a major recommendation knowledge graph that can recommend majors based on personality test results.
[0011] In the first aspect, preferably, in step S4, each entity, relationship, and the relevance are stored in a Neo4j database, and a Neo4j visualization tool is used to realize a visualized front-end display of the major recommendation knowledge graph based on the Neo4j database.
[0012] In the first aspect, preferably, all standard major names, all standard occupation names, and all standard personality test result names are subjected to spectral clustering to form several major clusters, occupation clusters, and personality test result clusters, which are then further classified and stored in a database.
[0013] In the first embodiment, preferably, the recommended range of majors only includes popular majors, and the popular majors are obtained in the following manner: First, we use the TextRank algorithm to rank all major names that appear in the Encyclopedia entry corpus, which contains a large number of major names. The edge weights between two standard major names in the TextRank algorithm are modified by adopting the similarity of the word vector representations of the two majors. The word vector representation of each major is obtained by encoding the Encyclopedia entry content of that major with Doc2Vec. After that, the majors with the highest set percentage ranking are selected from the ranking results as popular majors and included in the major recommendation range of the knowledge graph.
[0014] In a second aspect, the present invention provides a major recommendation search system based on a large-scale language model and a knowledge graph, including: A personality test module used to obtain the user's personality test results. A knowledge graph module used for storing a major recommendation knowledge graph constructed by the construction method described in any one of the technical solutions of the first aspect. A major recommendation module uses the user's personality test result as a search condition, and recommends all majors that have a shortest path connection with the personality test result from the major recommendation knowledge graph, calculates the product of edge weights of all edges on each shortest path between the personality test result and each recommended major, sums up the products of edge weights of all shortest paths, and then determines the overall relevance between the personality test result and the recommended major, sorts all recommended majors in descending order by overall relevance, and returns them to the user as a recommendation result.
[0015] In a third aspect, the present invention provides computer electronic equipment including a memory and a processor. The memory is used to store computer programs. When the processor executes the computer program, it is used to realize the major recommendation method based on a large-scale language model and knowledge graph described in any of the technical solutions of the first aspect, or to realize the major recommendation search system based on a large-scale language model and knowledge graph described in the technical solution of the second aspect.
[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention allows users to search for and obtain recommended majors and occupations suitable for them based on their personality test results, and by selecting a major, they can obtain search results for occupations, personalities, and qualities that meet their requirements. 2. The present invention utilizes the context learning and thought chain reasoning capabilities of large-scale language models to identify, extract, and supplement relationships between entities, which is more efficient and effective than the general fine-tuned BERT model. 3. This invention not only uses APIs that utilize large-scale language models, but also creates automated scripts to utilize the advantages of large-scale language models in processing large amounts of data to mark the document structure of "encyclopedia" entry data, thereby saving the cost of manual marking and achieving more accurate results. 4. The present invention can use the TextRank algorithm based on Doc2Vec, which replaces the traditional Word2Vec algorithm, taking into account the impact of weighting on text classification and can screen popular subjects. 5. The major recommendation search system based on the large-scale language model and knowledge graph constructed by this invention can perform front-end display based on the knowledge graph, adjust the position of the nodes displayed on the interface, fold and unfold category data according to the classification of personality test results, and display specific attributes by clicking on the nodes and edges of the knowledge graph, which has the effect of making the display more concise and easy to understand at a glance. [Brief explanation of the drawings]
[0017] The invention will now be further described in conjunction with the accompanying drawings. [Figure 1] 1 is a flowchart of a method for building a major recommendation knowledge graph based on a large-scale language model. [Figure 2]1 is a schematic diagram of the modular structure of a major recommendation search system based on a large-scale language model and knowledge graph. [Figure 3] FIG. 1 is a schematic diagram of an exemplary major-personality test result display. [Figure 4] This is a diagram showing exemplary majors, occupations, personality traits, and qualities. DETAILED DESCRIPTION OF THE INVENTION
[0018] Specific embodiments will be described in more detail below with reference to the accompanying drawings. By introducing the present invention through specific embodiments, the features and advantages of the present invention will become more apparent.
[0019] In a preferred embodiment of the present invention, a method for building a major recommendation knowledge graph based on a large-scale language model is provided, and the steps of this method are shown in S1 to S4. The specific implementation method of each step will be described in detail below. S1: For each major within the recommended range, obtain the standard major name for that major and a description of the personality and qualities required for that major. For each occupation within the recommended range, obtain the standard occupation name for that occupation and a description of the personality and qualities required for that occupation. For the personality test results that belong to each personality category in the personality test, obtain the standard personality test result name and a description of the personality and qualities corresponding to the personality test result. Each type of personality test result in the personality test corresponds to one personality category within the personality category range, and after each person takes the personality test, they will obtain one personality category that belongs to each personality category, which will be the personality test result of that test person.
[0020] The text describing the personality and qualities required for the major, the text describing the personality and qualities required for the occupation, and the text describing the personality and qualities corresponding to the personality test results can be collected from various sources as long as the necessary content is available. In an embodiment of the present invention, the text describing the personality and qualities required for the major, the text describing the personality and qualities required for the occupation, and the text describing the personality and qualities corresponding to the personality test results are all derived from the headword content of the "Encyclopedia" website, and the headword content needs to be processed in advance to meet the input requirements of entity extraction.
[0021] In one embodiment of the present invention, information related to majors can be obtained by crawling the headword content of Chinese encyclopedia websites, such as Wikipedia, using crawler technology that crawls websites and collects and stores data based on standard major names. The content includes descriptive text on the characteristics and qualities required for the major. For major name headwords, the major name is typically located in the first-level heading section of a Wikipedia page, such as computer science and technology or software engineering. The description of major information typically includes one or more introductory texts introducing information such as the major content, the characteristics and qualities required for the major, etc. Related occupations suitable for the major are typically located in the second-level heading section of a Wikipedia page, such as "area" or "category." Furthermore, for occupation-related information, crawler technology can be used to crawl the headword content of Chinese Wikipedia websites based on standard occupation names, including descriptive text on the characteristics and qualities required for the occupation.
[0022] The personality test may be any test that can classify an individual's personality or personality traits. In an embodiment of the present invention, the personality test may be one or more of the MBTI test, the Holland Occupational Personality Test, and the Multiple Intelligences Test. Similarly, in one embodiment of the present invention, the test result categories of various personality tests can be crawled based on the names of standard personality test results using crawler technology to crawl the headword content on the Chinese Wikipedia website, and the content includes description text of the personality and qualities corresponding to the personality test results.
[0023] Of course, in addition to crawling data using an online crawler, in another embodiment of the present invention, screening can also be performed using the existing Chinese Wikipedia corpus, the specific method of which is as follows: 1) The original text of Wikipedia's Chinese corpus, which contains some traditional Chinese characters, is first obtained and the traditional characters are simplified using OpenCC. Strings containing major names, occupation names, or personality test result names are used to perform a clause-by-clause match against the information in the corpus, screening out invalid data that does not contain any major names, occupation names, or personality test result names in the Chinese corpus. It is important to note that the strings containing major names should include different names of the same major, not necessarily standard major names. In other words, synonym matching or fuzzy matching methods are used to maximize the sample size. 2) After processing in 1), entity disambiguation is performed on the major-related headings, occupation-related headings, and personality test result-related headings, unifying the naming format of the major names and occupation names, while unifying the format of the personality test result names, for example, unifying "Computer Science" and "Computer Science and Technology" into "Computer Science and Technology." Meanwhile, personality test results, such as the results obtained from the MBTI test (ISTJ, INFP, etc.), are unified according to the standard names of the corresponding personality tests. 3) Because some of the information on majors, occupations, and personality test results is duplicated under different Wikipedia headings, we perform a deduplication process on all information data to remove the duplicated information, obtain a headword dataset corresponding to each of the majors, occupations, and personality test results, and use these as descriptive text for the corresponding personalities and qualities.
[0024] In addition, considering the need to construct a knowledge graph later, all standard major names, all standard occupation names, and all standard personality test result names are spectrally clustered to form several major clusters, occupation clusters, and personality test result clusters, and then each piece of information is classified and stored in the database.
[0025] Furthermore, in practical applications, users' demand for recommendations of majors is often concentrated on popular majors, and unpopular majors are often not included in the recommendation range. Therefore, in another embodiment of the present invention, the major recommendation range only includes popular majors. The popular major screening adopts the TextRank algorithm based on Doc2Vec. Doc2Vec is first used to train word vectors. Compared with the traditional Word2Vec algorithm, it not only considers the semantic information between words, but also the influence of word order on the sentence or text, resulting in a word vector training model that can be used to calculate the similarity and adjacent relationship between each word in the major name, occupation name, and personality test result name, respectively, and to improve the weighted transfer probability matrix of the TextRank algorithm. The following is a detailed description of the popular major acquisition method. The steps are as follows: First, all major names appearing in an encyclopedia headword corpus containing a large number of major names are sorted using the TextRank algorithm, and the edge weight between two standard major names in the TextRank algorithm is modified using the word vector representation similarity between the two majors. The word vector representation of each major is obtained by encoding the encyclopedia headword content of that major using Doc2Vec. After that, the majors with the highest set ratios from the sorting results are screened as popular majors and included in the major recommendation range of the knowledge graph.
[0026] It should be noted that both the TextRank algorithm and the Doc2Vec word vector tool belong to the prior art. The Doc2Vec word vector tool is primarily used to encode the encyclopedia entry content of each major to form corresponding word vectors. The word vectors of different majors can be represented by calculating the cosine distance to represent the similarity between majors, which is then used to improve the weighted transfer probability matrix of the original TextRank algorithm. The encyclopedia entry data containing a large number of major names can be the Wikipedia Chinese corpus, where the first-level titles obtained by screening steps 1) through 3) are used as major names. The TextRank algorithm is first used to sort the major names in the available major information dataset according to their current weights, which can be selected to represent the frequency of occurrence of the major names in the dataset. Then, major names with weights higher than a certain threshold are selected as popular majors and included in subsequent clustering, while the remaining unpopular major names are screened out and not included in subsequent clustering. The formula for calculating the original node weights in the TextRank algorithm is as follows:
[0027]
number
[0028] By converting the word vector using the Doc2Vec word vector tool and using the resulting word vector to calculate the distance between node i and node j, we can consider it as the similarity or proximity between the two major names. i and vertex v j The calculation method for the weighting between the two is as follows:
[0029]
number
[0030] Finally, the edge weights between two standard major names in the TextRank algorithm are modified as follows:
[0031]
number
[0032] The above-mentioned clustering of major names can be achieved by using a spectral clustering method, in one example, by using the spectral clustering method, 92 major classifications, 400 occupation classifications, and 30 personality test result classifications are finally obtained.
[0033] It should also be noted that the occupational set within the above-mentioned recommended scope can be completely obtained through an occupational catalog, and only occupations related to the content of all major headings may be included in the recommended scope, which is specifically determined by actual business needs and data acquisition conditions.
[0034] S2: By setting prompts for the input information to the large-scale language model, an entity extraction model suitable for completing the task of building a major recommendation knowledge graph must be constructed, and specific inference examples must be added to the set prompts. At the same time, the categories of the extracted entities and relationships must be defined, and each hierarchical concept must be output hierarchically. The output entity attribute format must also be specified.
[0035] The large-scale language model (LLM) referred to in the present invention can be realized by adopting a commercial large-scale language model, such as GPT-3.5, GPT-4.0, etc. To facilitate batch processing, it can be realized by directly calling the corresponding API.
[0036] In an embodiment of the present invention, when constructing the entity extraction model in step S2, the prompts input to the large-scale language model are continuously optimized and adjusted to achieve the entity extraction task. The key to using a large-scale language model to extract entities is to utilize three important capabilities of the large-scale language model: in-context learning (ICL), command compliance, and chain-of-thought (CoT). Therefore, the present invention enables human supervision and the addition of prompts, and adds specific inference steps to the concept of live input. That is, first, define entity and relationship categories, then hierarchically output each hierarchical concept, then specify the format of the output entity attributes, and finally obtain an optimized large-scale language model, which then automatically completes the attributes and relationships using the large-scale language model. During this process, the prompts are iteratively optimized based on feedback from the large-scale language model. The optimization and adjustment process is described in detail below. S21: First, samples are extracted from all written texts from which entity extraction is performed, and target entities in all written text samples are manually marked up to obtain sample data. S22: Add specific inference examples into the prompts of the large-scale language model, define the categories of extracted entities and relationships, and require the large-scale language model to output each hierarchical concept hierarchically. At the same time, specify the entity attribute format to be output by the large-scale language model. Then, test the large-scale language model after constructing the prompts on sample data, and further adjust the prompts based on the test results. After the accuracy of the extraction of text entities of the large-scale language model meets the predetermined requirements, use the final prompts in the large-scale language model to construct the entity extraction model.
[0037] Furthermore, the content of encyclopedia entries generally has a clear document structure, and the entities to be extracted are often implied in the document structure. Therefore, in an embodiment of the present invention, before inputting explanatory text for a large-scale language model, it is necessary to use a large-scale language model to perform document structure tagging in advance. Taking advantage of the large-scale language model's ability to process large amounts of data, prompts are structured to allow the large-scale language model to determine the text level in the entry content based on the dependency relationship and perform tagging. The main text in the entry content is tagged as level 0, and the Nth-level title in the entry content is tagged as level N, i.e., the first-level title is level 1, the second-level title is level 2, and so on. For example, in the Wikipedia entry for accounting, "accounting" is level 1, "Accounting can be divided into several fields, including financial accounting, management accounting, government accounting, and cost accounting, and can also be divided into profit accounting and non-profit accounting" is level 0, and "history" is level 2.
[0038] For ease of understanding, the following describes the process of optimizing a prompt "prompt" for a large-scale language model using a concrete example. 1) By adding similar examples in the reasoning process of the thought chain and adding specific reasoning steps to the prompt, redundant or missing results can be optimized. Original Usage Example: Recognize and extract naming entities for the sentence "Computer science has many fields. Some, such as computer graphics, emphasize the computation of specific results, while others, such as the theory of computation, consider the nature of computational problems." Output: "Computer Science", "Computer Graphics", "Computation", "Theory of Computation". Corrected example usage of prompt: The naming entities in the sentence "Computer graphics is the study of digital visual content, involving the synthesis and manipulation of image data. It is closely related to many other areas of computer science, including computer vision, image processing, computational geometry and visualization" are "computer graphics", "computer science", "computer vision", "image processing", and "computational geometry and visualization". What is the naming entity in the sentence "There are many fields in computer science. Some emphasize the computation of specific results, such as computer graphics, while others consider the nature of computational problems, such as theory of computation"? Output: "Computer Science", "Computer Graphics", "Theory of Computation".
[0039] 2) Defining entity and relationship categories in prompts helps large-scale language models return clearer results when applied to knowledge graphs. Original Usage Example: Recognize and extract naming entities for the sentence "Computer graphics is the study of digital visual content, including the synthesis and manipulation of image data. It is closely related to many other areas of computer science, including computer vision, image processing, computational geometry, and visualization." Outputs: "Computer Graphics", "Computer Science", "Computer Vision", "Image Processing", "Computational Geometry and Visualization". Example correction: "Computer graphics is the study of digital visual content, involving the synthesis and manipulation of image data. It is closely related to many other areas of computer science, including computer vision, image processing, computational geometry, and visualization." Recognize this sentence using a naming entity cluster classification and return its dependencies. For example, in "Accounting can be divided into several fields, including financial accounting, management accounting, etc.", the fields of "accounting" are "financial accounting," "management accounting," etc. Output: The domain of "Computer Science" includes "Computer Vision", "Image Processing", "Computational Geometry and Visualization", and "Computer Graphics".
[0040] 3) To realize knowledge graph augmentation support using large-scale language models, the format of the output attributes is specified in prompt. For example, based on the example correction in 2) (used as input), prompt returns a result in the format [$entity1, $entity2, ...] according to the dependency relationships, and requests that the dependency relationships be made explicit. Output: Areas include [Computer Science, Computer Vision], [Computer Science, Image Processing], [Computer Science, Computational Geometry and Visualization], and [Computer Science, Computer Graphics].
[0041] 4) The above is a simple example. After identifying and extracting naming entities from the dataset based on the large-scale language model, relationships between four types of entities, namely, major name, occupation name, personality and qualities, and personality test results, are complemented based on the context learning capability of the large-scale language model.
[0042] After optimizing the prompts for the large-scale language model, a corresponding entity extraction model can be built based on the API of the prompts and the large-scale language model. For example, based on GPT3.5, an API key can be called to create an automatic script to realize an entity extraction model that can perform entity extraction on content with different headwords in bulk.
[0043] In addition, in an embodiment of the present invention, to ensure that the constructed entity extraction model meets the necessary requirements, after the entity extraction model is constructed, a conventional BERT-BiLSTM-CRF structure entity extraction model is further used to perform knowledge entity identification on the major information dataset and the personality test result dataset, to obtain related information for each knowledge entity, and then a comparison and verification is performed between the related information obtained using the conventional BERT-BiLSTM-CRF model and the related information obtained using the large-scale language model. After passing the verification, it can be used to identify naming entities for the remaining introduced major information data.
[0044] S3: Using the entity extraction model, extract text entities from the description text of the personality and qualities required for each major within the recommended range, the description text of the personality and qualities required for each occupation corresponding to each profession, and the description text of the personality and qualities corresponding to each personality test result, and from these, extract the personality and quality entities required for each major and the names of occupations suitable for the person, the personality and quality entities required for each occupation, and the personality and quality entities possessed by each personality test result.
[0045] Similarly, in this step, if the description text of various personalities and qualities is in the form of headwords, the document structure is divided by a large-scale language model, and the headwords are input into the entity extraction model step by step to extract entities.Taking the headword content of major as an example, if the input is "In China, computer science and technology is a second-level major under engineering (first-level classification), and people usually work as back-end development engineers", then computer science and technology and back-end development engineers will be identified and extracted, and the output format will follow the definition above.
[0046] S4: All majors, all occupations, all extracted personality and quality entities, and all personality test results within the recommended range are treated as nodes in a graph, and relationships are added between the nodes based on the entity extraction results in S3. A multi-stage weighting algorithm is used to obtain the degree of association between majors and personality and qualities, between occupations and personality and qualities, and between personality test results and personality and qualities. Finally, personality and qualities are treated as connections between majors and personality test results, and a knowledge graph for major recommendation is constructed. In an embodiment of the present invention, the multi-stage weighting algorithm calculates the relevance and the process of constructing the major recommendation knowledge graph is as follows: S41: For each major, statistics are taken of the first frequency at which different character and quality entities appear in the written text with respect to the character and quality required for the major corresponding to that major, and the first frequency represents the first degree of association between the corresponding major and the character and quality. S42: For each major, determine the name of an occupation suitable for the major from the extracted entities; further, for the occupation suitable for each major, obtain statistics of a second frequency of appearance of different character and quality entities in the description text with respect to the character and quality required for the occupation corresponding to the occupation, and represent a second association degree between the major and the character and quality by the second frequency. S43: For each combination of major and character and quality, the weighted sum of the primary relevance between that major and that character and quality and the secondary relevance between that character and quality and all occupations suitable for that major is calculated to arrive at the final relevance between that major and that character and quality. S44: For each personality test result, statistics are calculated on the third frequency at which different personality and quality entities appear in the written text with respect to the personality and quality corresponding to the personality test result, and the third frequency represents the degree of association between the corresponding personality test result and the personality and quality. S45: All majors, all occupations, all extracted personality and quality entities, and all personality test results are treated as nodes in a graph. Based on the entity extraction results in S3, edge links are established between each major and the occupation suitable for that major, between each major and the personality and qualities required for that major, between each occupation and the personality and qualities required for that occupation, and between personality test results and the personality and qualities possessed by the personality test results. At the same time, weights are assigned to the edges of corresponding nodes based on the final association between the major and the personality and qualities obtained in S43. Weights are assigned to the edges of corresponding nodes based on the association between the personality test results and the personality and qualities obtained in S44, and the personality and qualities are linked to the majors and personality test results, thereby constructing a major recommendation knowledge graph that can recommend majors based on personality test results.
[0047] For ease of understanding, a specific example will be given to illustrate the implementation process of the above multi-stage weighting algorithm. 1) For each major, classification is performed based on the characteristics and qualities extracted by naming entity recognition, and statistics are taken of the frequency of appearance of each characteristic and quality in the major information, and the corresponding relevance (0≦Pi≦1) is calculated, resulting in the following formula:
[0048]
number
[0049] 2) Personality and quality entities are extracted from the entity recognition results of the suitable occupations for each major and the headword (occupation information) content for each occupation, and the frequency with which the personality and qualities appear in the occupation information is tallied. The normalized values are calculated in the same way as above and then used to calculate the second relevance. For example, in the case of a back-end development engineer, an occupation related to software engineering, the normalized values for tenacity are calculated to be 1 (units), responsibility is 0.6, proactivity is 0.4, decision-making ability is 0.2, and communication and collaboration ability is 0.2. The process for calculating the second relevance is as follows: Persistence relatedness: 1 / (1+0.6+0.4+0.2+0.2)=0.417 Responsibility relevance: 0.6 / (1+0.6+0.4+0.2+0.2)=0.25 Active relevance: 0.4 / (1+0.6+0.4+0.2+0.2)=0.167 Decision-making ability relevance: 0.2 / (1+0.6+0.4+0.2+0.2)=0.083 Correlation between communication and cooperation ability: 0.2 / (1+0.6+0.4+0.2+0.2)=0.083
[0050] 3) Since it is necessary to calculate the degree of association between personality test results and majors, statistics are taken of the frequency with which the personality and qualities appear in the entry content corresponding to the personality test results, and these are used to calculate the degree of association between the personality test results and the personality and qualities by calculating the same normalized values as described above. For example, for an ISTJ, the normalized values for tenacity are calculated to be 1, responsibility 0.9, proactivity 0.3, decision-making ability 0.5, and communication and collaboration ability 0.4. The process for calculating the degree of association between personality test results and personality and qualities is as follows: Persistence relatedness: 1 / (1+0.9+0.3+0.5+0.4)=0.326 Responsibility relevance: 0.9 / (1+0.9+0.3+0.5+0.4)=0.290 Active relevance: 0.3 / (1+0.9+0.3+0.5+0.4)=0.097 Decision-making ability relevance: 0.5 / (1+0.9+0.3+0.5+0.4)=0.161 Correlation between communication and cooperation ability: 0.4 / (1+0.9+0.3+0.5+0.4)=0.129
[0051] 4) Because the first and second relevance degrees above reflect the direct and indirect relationships between majors and personality traits and qualities, respectively, it is necessary to combine the two by assigning weights to the first and second relevance degrees. However, because the descriptions of personality traits and qualities in major information are more direct, a weighting of 0.7 is used, while the weighting of personality traits and qualities corresponding to major-related occupations is 0.3. The first relevance degree between a certain major and a certain personality trait and quality and the second relevance degrees between the personality traits and qualities and all occupations that can be engaged in with that major are weighted together to arrive at the final relevance degree between that major and those personality traits and qualities. In the example above, the final weighted relevance degree between majors and personality traits and qualities is as follows: Weighting: 0.7*0.345+0.3*0.417=0.3666 Tenacity: 0.7*0.276+0.3*0.25=0.2682 Responsibility: 0.7*0.138+0.3*0.167=0.1467 Decision making ability: 0.7*0.070+0.3*0.083=0.0739 Communication and cooperation ability: 0.7*0.172+0.3*0.083=0.1453
[0052] It should be noted that the above calculation example only refers to the personality traits and qualities that correspond to occupations related to one major; if there are multiple occupations suitable for a certain major, it is necessary to weight all the secondary associations between all occupations and their personality traits and qualities.
[0053] Finally, after obtaining the weighted relevance corresponding to all combinations of majors and personality traits and qualities from 4) and the relevance corresponding to all combinations of personality test results and personality traits and qualities from 3), all majors, all occupations, all extracted personality and trait entities, and all personality test results are treated as nodes in a graph, and based on the entity extraction results, edge links are established between each major and occupations suitable for that major, edge links are established between each major and the personality traits and qualities required by that major, edge links are established between each occupation and the personality traits and qualities required by that occupation, and edge links are established between personality test results and the personality traits and qualities possessed by that personality test result.
[0054] To recommend majors, the knowledge graph also needs to assign values to edge weights. As mentioned above, values are assigned to the edge weights of corresponding nodes based on the final relevance between majors and personality and qualities, and values are assigned to the edge weights of corresponding nodes based on the relevance between personality test results and personality and qualities. In this way, majors and personality test results can be linked by personality and qualities, and suitable majors can be recommended based on personality test results. The recommendation method is as follows: Using the user's personality test results as a search query, all majors in the major recommendation knowledge graph that have a shortest path connection with the personality test results are recommended majors. The products of the edge weights of all edges on each shortest path between the personality test results and each recommended major are calculated (by multiplying the edge weights of all edges in one path). The sum of the products of the edge weights of all shortest paths is used as the overall relevance between the personality test results and the recommended major, and this overall relevance represents the probability of recommending a major corresponding to the current personality test result.
[0055] From the above example, the overall correlation calculation formula between software engineering majors and ISTJ is as follows:
[0056]
number
[0057] Therefore, after the knowledge graph is completed, a major recommendation search system based on the large-scale language model and the knowledge graph can be built, which can be used to realize major recommendation based on personality test results. As shown in Figure 2, the functional modules of the system include: A personality test module used to obtain the user's personality test results. a knowledge graph module, which is used to store the major recommendation knowledge graph constructed by the construction method; A major recommendation module uses the user's personality test result as a search condition, and recommends all majors that have a shortest path connection with the personality test result from the major recommendation knowledge graph; calculates the product of edge weights of all edges on each shortest path between the personality test result and each recommended major; adds up the products of edge weights of all shortest paths to determine the overall relevance between the personality test result and the recommended major; sorts all recommended majors in descending order by overall relevance; returns the recommendation result to the user; and visualizes the knowledge graph according to business display logic.
[0058] Furthermore, in an embodiment of the present invention, each entity, relationship, and the relevance are stored in a Neo4j database, and the Neo4j visualization tool is used to realize the front-end visualization display of the major recommendation knowledge graph based on the Neo4j database. After the system is fully constructed, when a user inputs a command into the front-end interface, the back-end Neo4j interacts with the command to obtain the corresponding return data from the major recommendation knowledge graph, and the data is displayed on the front-end interface.
[0059] Here, the command input by the user to the front-end interface is a search command or an operation command. In the embodiment of the present invention, the search commands that can be input include: A) By taking a personality test, the personality test results can be obtained, and then the corresponding recommended majors can be obtained through further search. For example, after taking three tests, namely MBTI, Holland Vocational Personality Test, and Multiple Intelligences, based on the analysis results of the major recommendation chart, the test subject will show traits such as resilience to adversity, high initiative, and excellent comprehension ability. Therefore, the recommended majors include computer science and technology, software engineering, and electronic information. B) By searching for the name of your major, you can get the corresponding occupations, personality traits, and qualities. For example, if you enter computer science and technology, the related occupations will be back-end development engineer, algorithm engineer, etc., and the related personality traits and qualities will be tenacity, responsibility, initiative, decision-making and judgment ability, communication and cooperation ability, understanding ability, etc.
[0060] Furthermore, operational commands that can be input in the embodiment of the present invention include: Extract data from Neo4j and visualize the knowledge graph using eCharts relationship diagrams. Collapse and expand knowledge graph nodes in the graph display space and drag the knowledge graph relationship network to adjust the layout. Click on graph nodes or edges to view specific attributes such as relevance.
[0061] In one embodiment of the present invention, Fig. 3 is a diagram illustrating the display effect of the major-personality test results of the present invention, where text information is not important and the display state is shown only by a graph. Fig. 4 is a diagram illustrating the display diagram of the major, occupation, personality and qualities of the present invention, where text information is not important and the display effect is shown only by a graph.
[0062] Similarly, another preferred embodiment of the present invention provides, based on the same inventive concept, a computer electronic device including a memory and a processor corresponding to the subject recommendation method based on the large-scale language model and knowledge graph provided in the above embodiment. The memory is for storing a computer program. The processor is configured to realize the above-mentioned method for recommending a major based on a large-scale language model and a knowledge graph or the above-mentioned system for recommending a major based on a large-scale language model and a knowledge graph when the computer program is executed.
[0063] The logic instructions in the memory may be implemented in the form of a software functional unit and stored in one computer-readable storage medium when sold or used as a separate product. Based on this understanding, the technical solution of the present invention essentially or a part that contributes to the prior art or a part of the technical solution may be embodied in the form of a software product, and the computer software product is stored in one storage medium and includes multiple instructions to cause one computer device (which may be a personal computer, a server, a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention.
[0064] Thus, based on the same inventive idea, another preferred embodiment of the present invention provides a computer-readable storage medium corresponding to the method for recommending a major based on a large-scale language model and a knowledge graph provided by the above embodiment, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the method for recommending a major based on a large-scale language model and a knowledge graph or a search system for recommending a major based on the large-scale language model and a knowledge graph can be realized.
[0065] The storage medium may include a random access memory (RAM) and at least one non-volatile memory (NVM) such as a magnetic disk memory. Note that the storage medium may be any of various media capable of storing program code, such as a USB memory, a removable hard disk, a magnetic disk, or an optical disk.
[0066] The processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It should be noted that the processor may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component.
[0067] It should be noted that the specific operation procedures of the above-described system may refer to the corresponding procedures in the above-described method embodiments, and will not be repeated here for the sake of brevity. This is clearly understood by those skilled in the art. In each embodiment provided in this application, the division of steps or modules in the above-described system and method is merely a logical functional division, and in implementation, other division methods may be used, such as combining or integrating multiple modules or steps, or dividing one module or step.
[0068] The above embodiments are merely preferred aspects of the present invention and are not intended to limit the present invention. Various modifications and variations may be made by those skilled in the art without departing from the spirit and scope of the present invention. Therefore, any technical solutions obtained by adopting equivalent substitution or equivalent conversion methods are included within the scope of protection of the present invention.
Claims
1. A method for constructing a major recommendation knowledge graph based on a large-scale language model, the method being executed as a computer program in a computer electronic device including a memory for storing a computer program and a processor, comprising: S1: For each major within the recommended range, obtain the standard major name for that major and a description of the personality and qualities required for that major; for each occupation within the recommended range, obtain the standard occupation name for that occupation and a description of the personality and qualities required for that occupation; for the personality test results belonging to each personality category in the personality test, obtain the standard personality test result name and a description of the personality and qualities corresponding to the personality test results; S2: By setting prompts in the input information of the large-scale language model, an entity extraction model suitable for completing the task of building a major recommendation knowledge graph is constructed, in which specific inference examples need to be added to the set prompts. At the same time, it is necessary to define the categories of the extracted entities and relationships, and output each hierarchical concept hierarchically, and also to specify the output entity attribute format. S3: Using the entity extraction model, extract text entities for the description text of the personality and qualities required for each major within the recommended range, the description text of the personality and qualities required for each occupation, and the description text of the personality and qualities required for each personality test result, and extract from them the personality and quality entities required for each major, the name of the occupation suitable for the occupation, the personality and quality entities required for each occupation, and the personality and quality entities possessed by each personality test result; S4: All majors, all occupations, all extracted personality and quality entities, and all personality test results within the recommended scope are treated as nodes in a graph, and relationships between the nodes are added based on the entity extraction results of S3. A multi-stage weighting algorithm is used to obtain the association degrees between the major and personality and qualities, between occupations and personality and qualities, and between personality test results and personality and qualities. Finally, the personality and qualities are treated as connections between the major and personality test results, and a knowledge graph for recommending majors is constructed. A method for constructing a major recommendation knowledge graph based on a large-scale language model, comprising the steps above.
2. The description text of the personality and qualities required for the major, the description text of the personality and qualities required for the occupation, and the description text of the personality and qualities corresponding to the personality test results are all derived from the headword content of the "Encyclopedia" website, and the headword content needs to be processed in advance to meet the input requirements of entity extraction; 2. The method for constructing a major recommendation knowledge graph based on a large-scale language model according to claim 1, wherein the personality test is one or more of the MBTI test, the Holland Occupational Personality Test, and the Multiple Intelligence Test.
3. When constructing the entity extraction model in S2, it is necessary to constantly optimize and adjust the prompts input into the large-scale language model to achieve the entity extraction task. The optimization and adjustment process is as follows: S21: First, extract samples from all written texts for which entity extraction is to be performed, and manually mark up the target entities in all written text samples to obtain sample data; S22: Add specific inference examples to the prompts of the large-scale language model, define categories of extracted entities and relationships, and require the large-scale language model to output each hierarchical concept hierarchically, while specifying the entity attribute format to be output by the large-scale language model. Then, test the large-scale language model after constructing the prompts on sample data, and further adjust the prompts based on the test results. After the large-scale language model's extraction accuracy for text entities meets the predetermined requirements, use the final prompts in the large-scale language model to construct the entity extraction model.
4. The method for constructing a major recommendation knowledge graph based on a large-scale language model according to claim 2, characterized in that the description text input into the large-scale language model needs to be marked up in advance using the large-scale language model to form a document structure; utilizing the advantages of the large-scale language model in processing large amounts of data, a prompt is constructed to allow the large-scale language model to determine and mark up the text levels in the entry content according to the subordinate relationships, wherein the main text in the entry content is marked up to level 0, and the Nth-level title in the entry content is marked up to level N correspondingly.
5. The process of calculating the relevance of the multi-stage weighting algorithm and constructing the major recommendation knowledge graph is as follows: S41: For each major, with respect to the characteristics and qualities required for the major corresponding to the major, obtain statistics of first frequencies at which different characteristics and qualities entities appear in the description text, and represent first association degrees between the corresponding major and the characteristics and qualities based on the first frequencies; S42: For each major, determine the name of an occupation suitable for that major from the extracted entities; and for each occupation suitable for that major, calculate statistics of the second frequency of appearance of different character and quality entities in the description text with respect to the character and quality required for the corresponding occupation, and represent a second association degree between the major and the character and quality by the second frequency; S43: For each combination of major and character and quality, calculate a weighted sum of the first relevance between the major and the character and quality, and the second relevance between all occupations suitable for the major and the character and quality, to obtain the final relevance between the major and the character and quality; S44: For each personality test result, calculate statistics of a third frequency at which different personality and quality entities appear in the written text with respect to the personality and quality corresponding to the personality test result, and represent the degree of association between the corresponding personality test result and the personality and quality at the third frequency; S45: The method for constructing a major recommendation knowledge graph based on a large-scale language model as claimed in claim 1, wherein: S45 sets all majors, all occupations, all extracted personality and quality entities, and all personality test results as nodes of a graph; and based on the entity extraction results of S3, establish edge links between each major and the occupation suitable for that major, establish edge links between each major and the personality and qualities required for that major, establish edge links between each occupation and the personality and qualities required for that occupation, and establish edge links between personality test results and the personality and qualities possessed by the personality test results; at the same time, assign weights to the edges of corresponding nodes based on the final association degrees between the majors and the personality and qualities obtained in S43, assign weights to the edges of corresponding nodes based on the association degrees between the personality test results and the personality and qualities obtained in S44, and further link the majors and personality test results with the personality and qualities, thereby constructing a major recommendation knowledge graph that can recommend majors based on the personality test results.
6. 2. The method for constructing a major recommendation knowledge graph based on a large-scale language model according to claim 1, wherein in step S4, each entity, relationship, and the relevance are stored in a Neo4j database, and a Neo4j visualization tool is used to realize a visualized front-end display of the major recommendation knowledge graph based on the Neo4j database.
7. The method for constructing a major recommendation knowledge graph based on a large-scale language model according to claim 6, characterized in that all standard major names, all standard occupation names, and all standard personality test result names are subjected to spectral clustering to form several major clusters, occupation clusters, and personality test result clusters, which need to be further classified and stored in the database.
8. The recommended majors only include popular majors, and the acquisition methods for these popular majors are as follows: First, all the major names appearing in the Encyclopedia entry corpus, which contains a large number of major names, are ranked using the TextRank algorithm. The edge weighting between two standard major names in the TextRank algorithm is modified using the word vector representation similarity between the two majors. The word vector representation of each major is obtained by encoding the Encyclopedia entry content of that major using Doc2Vec. The method for constructing a major recommendation knowledge graph based on a large-scale language model according to claim 7, further comprising: selecting the majors with a set percentage of the highest ranking from the ranking result as popular majors and including them in the major recommendation range of the knowledge graph.
9. A major recommendation search system based on a large-scale language model and a knowledge graph, a personality test module used for obtaining a personality test result of the user; a knowledge graph module used to store a major recommendation knowledge graph constructed according to the construction method of any one of claims 1 to 8; a major recommendation module that uses a user's personality test result as a search condition, and in a major recommendation knowledge graph, recommends all majors that have a shortest path connection with the personality test result, calculates the product of edge weights of all edges on each shortest path between the personality test result and each recommended major, sums up the products of edge weights of all shortest paths to determine the overall relevance between the personality test result and the recommended major, sorts all recommended majors in descending order by overall relevance, returns them to the user as recommendation results, and visualizes the knowledge graph based on business display logic.
10. 1. A computer electronic device comprising: a memory and a processor, The memory is used to store a computer program; The computer electronic device is configured to realize the method for recommending a major based on a large-scale language model and a knowledge graph according to any one of claims 1 to 8, or the system for recommending a major based on a large-scale language model and a knowledge graph according to claim 9, when the computer program is executed.
Citation Information
Patent Citations
Matching measuring apparatus and method and program
JP2018124729A
System for learning word using mobile device and the method thereof
KR1020160113909A
KR20220150173A
Cited By
Enterprise resource recommendation method and device based on scene map, medium and equipment
CN121092598A