Global college information education consultation method and system based on AI large model

By adopting AI big models, multilingual knowledge graphs and artificial intelligence reasoning engines in the global college information education consulting system, the problem of difficulty in integrating global college data and providing personalized consultation in the existing technology is solved, and efficient, accurate and personalized educational information consulting services are achieved.

CN120011504AInactive Publication Date: 2025-05-16SHANGHAI TIANQU YUNQI EDUCATION TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510082371.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively integrate global college data and provide multilingual, highly accurate and personalized educational information consulting services, especially in cross-language query scenarios that deal with multi-table associations and user context.

Method used

The global college information education consulting method based on AI big models is adopted, combining multilingual knowledge graphs, pre-trained big models and artificial intelligence reasoning engines, collects official information from multiple languages ​​in universities around the world, performs preprocessing and feature extraction, forms a knowledge graph, and intent recognition and retrieval through natural language consultation, providing personalized college lists and related information.

Benefits of technology

It has significantly improved the intelligence level of information query in universities around the world, provided comprehensive, efficient, accurate and personalized educational information consulting services, can effectively integrate multilingual data, and supports cross-language query and personalized analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011504A_ABST
    Figure CN120011504A_ABST
Patent Text Reader

Abstract

The invention provides a global college information education consultation method and system based on an AI large model, and relates to the technical field of data processing and artificial intelligence, and the method comprises the steps: collecting global college multi-language official information, carrying out the preprocessing, and constructing a knowledge graph; utilizing the multi-language corpus to finely adjust the pre-training large model; performing intention recognition through natural language consultation, and calling a knowledge graph to retrieve a college list, enrollment standards and associated information; and accurate and personalized education information consultation service is provided through visual display and formatted recommendation. According to the method, the multi-language knowledge graph, the pre-training large model and the artificial intelligence inference engine are combined, so that the intelligent level of global college information query is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing and artificial intelligence technology, and in particular to a global college information education consulting method and system based on an AI big model. Background Art

[0002] As globalization continues to deepen, more and more students and institutions need to retrieve, compare, and analyze information about universities and related majors around the world in order to make more accurate education and career plans. However, in existing technologies, educational information systems are often scattered across countries, regions, or institutions, with inconsistent data structures and lack of support for complex cross-language and cross-disciplinary queries. Specifically:

[0003] (1) Multi-source data is difficult to integrate

[0004] Information on institutions around the world is usually distributed in educational databases in different countries or regions, or stored on websites in different languages. Existing systems often find it difficult to unify and integrate these data, resulting in limited retrieval efficiency and query accuracy.

[0005] (2) Lack of flexible cross-language support

[0006] Education information users often come from all over the world and expect to be able to search and consult in their native language. However, traditional systems often only support one or a few fixed languages, making it difficult to meet the multilingual needs of global users.

[0007] (3) Insufficient personalized in-depth consulting capabilities

[0008] Existing educational information retrieval is mainly based on keywords and simple text searches, and lacks the ability to use deep learning models to understand user context semantics and provide personalized, multi-dimensional analysis results.

[0009] To address the above issues, traditional solutions rely on manual or system-configured query statements (such as SQL) to retrieve data, or simplify the retrieval mode to a fixed template. This makes it difficult to efficiently provide accurate results when dealing with multi-table associations or cross-language query scenarios based on user context.

[0010] Therefore, how to effectively integrate global university data and provide multilingual, highly accurate and personalized education information consulting services has become a major technical challenge in the current field of education information retrieval and artificial intelligence. Summary of the invention

[0011] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide a global school information education consulting method and system based on AI big model, which combines multilingual knowledge graph, pre-trained big model and artificial intelligence reasoning engine to significantly improve the intelligence level of global school information query.

[0012] To achieve the above object, the present invention provides the following solutions:

[0013] A global college information education consulting method based on AI big model, including:

[0014] Collect official information in multiple languages ​​from universities around the world;

[0015] Preprocessing and feature extracting the official information to form a knowledge graph;

[0016] Obtain an open pre-trained large model containing multi-language corpora, and train and fine-tune the open pre-trained large model to obtain a pre-trained language large model;

[0017] Obtain natural language consultation through a front-end page or voice device, and input the natural language consultation into the pre-trained language model for intent recognition to obtain key recognition information;

[0018] Based on the artificial intelligence reasoning engine, the knowledge graph is called to search according to the key identification information to obtain search results; the search results include a list of colleges and universities that meet the conditions and admission standards, as well as the geographical location, language environment, and number of admissions in previous years associated with the colleges and universities;

[0019] Visual display and formatting recommendations are performed based on the search results.

[0020] Preferably, the official information includes: the name of the institution, geographical location, major information, scholarship information and language requirements; the major information includes curriculum setting and duration of study.

[0021] Preferably, the official information is preprocessed and feature extracted to form a knowledge graph, including:

[0022] Deduplication of the official information is performed to obtain deduplication text;

[0023] Based on the contextual relationship, classify the deduplicated text to obtain classified data;

[0024] Based on natural language processing technology, feature extraction is performed on the elements of faculty status, scientific research capabilities, and employment rates described in the classified data to build association relationships;

[0025] A knowledge graph is formed according to the elements and the association relationships.

[0026] Preferably, based on the contextual relationship, the deduplicated text is subjected to text classification to obtain classification data, including:

[0027] Inputting the deduplicated text into a text context feature extraction layer to extract context feature information;

[0028] Inputting the deduplicated text into a global feature extraction layer to extract global feature information;

[0029] Fusing the contextual feature information and the global feature information to obtain a fused text feature;

[0030] The fused text features are input into a trained text classification model to obtain the classification data.

[0031] Preferably, the deduplicated text is input into a text context feature extraction layer to extract context feature information, including:

[0032] Using a preset pre-trained language model to extract initial feature information of the deduplicated text;

[0033] Inputting the initial feature information into a forward gated recurrent unit and a reverse gated recurrent unit;

[0034] The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual relationship feature information.

[0035] Preferably, the outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual feature information, including:

[0036] Using the formula:

[0037]

[0038] The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual feature information; wherein, H t ={H 1 ,H 2 ,…H l} represents the initial feature information input at time t, represents the output of the positive gated recurrent unit at time t, Represents the output of the reverse gated recurrent unit at time t, GRU represents the gated recurrent unit, and G represents the contextual relationship feature information.

[0039] Preferably, the deduplicated text is input into a global feature extraction layer to extract global feature information, including:

[0040] The initial feature information of the deduplicated text is sequentially input into the convolution layer and the pooling layer to obtain the global feature information; wherein the global feature information extraction formula is:

[0041] c i =f(ω·H+b)

[0042]

[0043] In the formula, f is the activation function, ω is the convolution kernel, h is the convolution kernel size, b is the bias value, c is i is the feature vector extracted by the i-th convolutional layer, It represents the value of the features extracted by three different convolution kernels after the maximum pooling layer, and C represents the extracted global feature information.

[0044] Preferably, an open pre-trained large model containing multi-language corpora is obtained, and the open pre-trained large model is trained and fine-tuned to obtain a pre-trained language large model, including:

[0045] Choose an open pre-trained large model with multilingual corpus;

[0046] De-duplication, sentence formatting and segmentation are performed on a preset multi-corpus sample set to obtain a structured fine-tuned corpus; the multi-corpus sample set includes text information on study abroad guides, school comparisons, and professional interpretations;

[0047] Using the structured fine-tuning corpus to supervise the open pre-trained large model, so as to ensure that the vocabulary and expression ability of the model in the field of education are strengthened, and obtain a fine-tuning model for basic education scenarios;

[0048] The basic education scenario fine-tuning model is fine-tuned with small sample instructions using preset multi-language small sample instruction data to ensure that the model understands and answers education information consultation questions expressed in different languages, thereby obtaining a large pre-trained language model enhanced with multiple languages.

[0049] Preferably, based on the artificial intelligence reasoning engine, the knowledge graph is called to search according to the key identification information to obtain search results, including:

[0050] Mapping the key information with the entity and attribute names of the knowledge graph to construct an executable search request;

[0051] Performing a search in the knowledge graph according to the executable search request to match the college entities and associated attributes that meet the conditions to obtain preliminary results;

[0052] The preliminary results are screened and sorted based on preset business rules, and the information on admission standards, geographical location, and number of admissions in previous years is integrated, deduplicated, or merged to obtain the search results; the search results include a list of matching institutions and extended information.

[0053] A global college information education consulting system based on AI big model, including:

[0054] An information collection unit, which is used to collect official information in multiple languages ​​from universities around the world;

[0055] A graph construction unit, used for preprocessing and feature extraction of the official information to form a knowledge graph;

[0056] A model training unit, used to obtain an open pre-trained large model containing multi-language corpora, and train and fine-tune the open pre-trained large model to obtain a pre-trained language large model;

[0057] An intention recognition unit is used to obtain natural language consultation through a front-end page or a voice device, and input the natural language consultation into the pre-trained language large model for intention recognition to obtain key recognition information;

[0058] A retrieval unit, configured to call the knowledge graph to perform retrieval based on the key identification information based on an artificial intelligence reasoning engine to obtain retrieval results; the retrieval results include a list of colleges and universities that meet the conditions and admission standards, as well as geographic locations, language environments, and the number of students admitted in previous years associated with the colleges and universities;

[0059] A visualization unit is used to perform visualization and formatting recommendations based on the search results.

[0060] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0061] The present invention provides a global college information education consultation method and system based on an AI big model, the method comprising: collecting official information in multiple languages ​​of global colleges and universities; preprocessing and feature extraction of the official information to form a knowledge graph; obtaining an open pre-trained big model containing multi-language corpora, and training and fine-tuning the open pre-trained big model to obtain a pre-trained language big model; obtaining natural language consultation through a front-end page or a voice device, and inputting the natural language consultation into the pre-trained language big model for intent recognition to obtain key recognition information; based on an artificial intelligence reasoning engine, calling the knowledge graph for retrieval according to the key recognition information to obtain a retrieval result; the retrieval result includes a list of colleges and universities that meet the conditions and admission standards, as well as the geographical location, language environment, and number of admissions in previous years associated with the colleges and universities; and visual display and formatted recommendation are performed according to the retrieval results. Through the combination of multi-language knowledge graphs, pre-trained big models, and artificial intelligence reasoning engines, users are provided with comprehensive, efficient, accurate, and personalized education information consultation services, which significantly improves the intelligence level of global college information query and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0063] Figure 1 A flow chart of a method provided by an embodiment of the present invention;

[0064] Figure 2 A schematic diagram of the system structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] The purpose of the present invention is to provide a global college information education consultation method and system based on AI big model, which combines multilingual knowledge graph, pre-trained big model and artificial intelligence reasoning engine to significantly improve the intelligence level of global college information query.

[0067] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0068] Figure 1 A flow chart of a method provided by an embodiment of the present invention, such as Figure 1 As shown, the present invention provides a global college information education consultation method based on an AI big model, comprising:

[0069] Step 100: Collect official information in multiple languages ​​from universities around the world;

[0070] Step 200: Preprocessing and feature extraction of official information to form a knowledge graph;

[0071] Step 300: Obtain an open pre-trained large model containing multi-language corpora, and train and fine-tune the open pre-trained large model to obtain a pre-trained language large model;

[0072] Step 400: Obtain natural language consultation through a front-end page or voice device, and input the natural language consultation into a pre-trained language model for intent recognition to obtain key recognition information;

[0073] Step 500: Based on the artificial intelligence reasoning engine, the knowledge graph is called to search according to the key identification information to obtain search results; the search results include a list of colleges and universities that meet the conditions and admission standards, as well as the geographical location, language environment, and number of admissions in previous years associated with the colleges and universities;

[0074] Step 600: Perform visual display and formatting recommendations based on the search results.

[0075] Preferably, the official information includes: the name of the institution, geographical location, major information, scholarship information and language requirements; the major information includes curriculum setting and duration of study.

[0076] Specifically, step 100 in this embodiment includes:

[0077] Step 101: Identify data sources and collect multilingual data

[0078] First, identify the main sources of information on global universities, including official databases of education departments of various countries, official websites of universities, public data released by international education organizations (such as QS rankings, Times rankings, etc.), and third-party education consulting platforms. Develop a multilingual data collection strategy for these sources, covering mainstream languages ​​(such as English, Chinese, French, Spanish, etc.) to ensure data diversity and coverage. At the same time, combine crawler technology and API interfaces to automatically collect structured and unstructured data such as school names, geographic locations, professional information, scholarship information, and language requirements.

[0079] Step 102: Cleaning and normalizing multilingual data

[0080] The collected data may be duplicate, incomplete or inconsistent in format, so data cleaning is required. Duplicate school records are eliminated through deduplication algorithms, and school names and information in different languages ​​are aligned in a unified manner in combination with multilingual translation tools. For example, "University of Oxford" and "Oxford University" are identified as the same entity. In addition, natural language processing techniques (such as entity recognition and regular expressions) are used to extract core fields (such as major names, course settings, academic system, etc.), and the data is stored in a unified structured format (such as JSON or relational database).

[0081] Step 103: Integration and storage of multi-source data

[0082] After cleaning and normalization, the multilingual data is integrated to build a multi-source unified university database. Different language descriptions of the same university are merged through entity alignment technology, and missing information is supplemented (such as scholarship or course information from other sources). Finally, it is stored in the database according to field classification (such as school name, geographical location, major settings, etc.), providing a high-quality data foundation for the subsequent construction of the knowledge graph.

[0083] Preferably, the official information is preprocessed and feature extracted to form a knowledge graph, including:

[0084] Deduplication of the official information is performed to obtain deduplication text;

[0085] Based on the contextual relationship, classify the deduplicated text to obtain classified data;

[0086] Based on natural language processing technology, feature extraction is performed on the elements of faculty status, scientific research capabilities, and employment rates described in the classified data to build association relationships;

[0087] A knowledge graph is formed according to the elements and the association relationships.

[0088] Specifically, step 200 of this embodiment includes:

[0089] Step 201: Data deduplication and normalization

[0090] First, the collected multilingual official information is deduplicated and normalized. String matching algorithms (such as hash algorithms or Jaccard similarity) are used to detect duplicate records, and multilingual translation tools are used to align the same entity described in different languages ​​(such as "University of Oxford" and "Oxford University"). For data with inconsistent formats (such as dates, academic system, etc.), regular expressions or predefined rules are used to standardize the data to ensure the uniformity and accuracy of the data. The deduplicated text data is stored in a structured format (such as JSON or database table), laying the foundation for subsequent classification and feature extraction.

[0091] Step 202: Context-based text classification

[0092] For the deduplicated text data, the contextual relationship is used to classify the text. A pre-trained language model (such as BERT or GPT) is used to perform semantic analysis on the text, and the supervised learning method (such as SVM or classification neural network) is combined to divide the text into different categories (such as faculty situation, scientific research ability, employment rate, etc.). For example, by analyzing the keywords in the text (such as "faculty", "research output", "employment rate") and their contextual semantics, the description content is automatically classified into the corresponding subject category. The classified data is stored in a labeled form to provide a clear semantic range for subsequent feature extraction.

[0093] Step 203: Feature extraction and knowledge graph construction

[0094] Based on natural language processing technology (such as named entity recognition, dependency parsing, etc.), key elements (such as the number of professors in the faculty situation, the number of papers published in the scientific research ability, and the specific percentage of the employment rate) are extracted from the classified data. At the same time, relationship extraction technology (such as rule-based relationship extraction or deep learning model) is used to build the association between elements. For example, "a certain university-scientific research ability-number of papers published" is established as a triple relationship. Finally, the extracted elements and associations are stored in a graph database (such as Neo4j or RDF format) to form a searchable knowledge graph to provide support for subsequent inference engine calls.

[0095] Preferably, based on the contextual relationship, the deduplicated text is subjected to text classification to obtain classification data, including:

[0096] Inputting the deduplicated text into a text context feature extraction layer to extract context feature information;

[0097] Inputting the deduplicated text into a global feature extraction layer to extract global feature information;

[0098] Fusing the contextual feature information and the global feature information to obtain a fused text feature;

[0099] The fused text features are input into a trained text classification model to obtain the classification data.

[0100] Preferably, the deduplicated text is input into a text context feature extraction layer to extract context feature information, including:

[0101] Using a preset pre-trained language model to extract initial feature information of the deduplicated text;

[0102] Inputting the initial feature information into a forward gated recurrent unit and a reverse gated recurrent unit;

[0103] The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual relationship feature information.

[0104] Preferably, the outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual feature information, including:

[0105] Using the formula:

[0106]

[0107] The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual feature information; wherein, H t ={H 1 ,H 2 ,…H l} represents the initial feature information input at time t, represents the output of the positive gated recurrent unit at time t, Represents the output of the reverse gated recurrent unit at time t, GRU represents the gated recurrent unit, and G represents the contextual relationship feature information.

[0108] Preferably, the deduplicated text is input into a global feature extraction layer to extract global feature information, including:

[0109] The initial feature information of the deduplicated text is sequentially input into the convolution layer and the pooling layer to obtain the global feature information; wherein the global feature information extraction formula is:

[0110] c i =f(ω·H+b)

[0111]

[0112] In the formula, f is the activation function, ω is the convolution kernel, h is the convolution kernel size, b is the bias value, c isi is the feature vector extracted by the i-th convolutional layer, It represents the value of the features extracted by three different convolution kernels after the maximum pooling layer, and C represents the extracted global feature information.

[0113] Exemplarily, this embodiment inputs the deduplicated text into a preset pre-trained language model (such as BERT, GPT, etc.), and uses its powerful semantic understanding ability to extract the initial feature information of the text. These features contain the vocabulary, syntax, and semantic information of the text, and are the basis for subsequent feature extraction. The initial feature information is input into the forward and reverse gated recurrent units (GRUs) respectively. The forward GRU extracts features from the beginning to the end of the text, and the reverse GRU extracts features from the end to the beginning of the text. Through bidirectional processing, the dependencies and semantic information between the contexts in the text are captured. The outputs of the forward GRU and the reverse GRU are spliced ​​to form complete contextual relationship feature information. This splicing method can comprehensively consider the correlation between the previous and next semantics in the text, and provide richer context information for classification tasks.

[0114] Furthermore, in this embodiment, the initial feature information of the text is input into the convolution layer, and the local features of the text (such as the pattern of phrases or sentence fragments) are extracted through the convolution kernel. The diversity of convolution kernels (such as convolution kernels of different sizes) can capture features of different granularities in the text. The features extracted by the convolution layer are input into the pooling layer, and the global feature information is extracted through the maximum pooling operation. The pooling layer can effectively reduce the feature dimension while retaining the most significant feature information and enhancing the generalization ability of the model. The feature information after pooling represents the global semantic pattern of the text, providing overall semantic support for the classification task.

[0115] Furthermore, the present embodiment fuses the contextual feature information with the global feature information (such as splicing or weighted fusion) to form a fused text feature. The contextual feature focuses on the local dependencies of the text, and the global feature focuses on the overall semantic pattern. The combination of the two can more comprehensively characterize the text features. The fused text features are input into a trained text classification model (such as a fully connected neural network or classifier) ​​and the classification results are output. The classification model classifies the text into corresponding categories (such as faculty status, scientific research capabilities, employment rate, etc.) based on the semantic information of the fused features.

[0116] This method innovatively combines contextual features and global features, which can capture local semantic dependencies in the text and extract overall semantic patterns. The contextual features capture the semantic associations before and after through bidirectional GRU, and the global features extract the overall semantic structure of the text through convolution and pooling. The fusion of the two enables the classification model to understand the text content more comprehensively. Using bidirectional GRU to extract contextual features can effectively capture the dependencies between the semantics before and after in the text, which is particularly suitable for complex descriptive texts in the field of education (such as descriptions of faculty or scientific research capabilities). Compared with traditional unidirectional RNN or LSTM, bidirectional GRU is more efficient in capturing long-distance dependencies and contextual information. The introduction of convolutional layers and pooling layers enables the model to extract global features of different granularities (such as phrase patterns or sentence structures) from the text. Through the combination of multiple convolution kernels, it is possible to capture diverse semantic patterns in the text and enhance the robustness and generalization ability of the classification model. The fusion of contextual features and global features is the core innovation of this method. Contextual features focus on local dependencies, while global features focus on overall patterns. The combination of the two makes up for the shortcomings of single feature representation, making the classification model more accurate and comprehensive when processing complex texts.

[0117] In summary, this method can efficiently and accurately classify texts by extracting and fusing contextual features and global features, combined with pre-trained language models and deep learning technology. It is particularly suitable for complex multi-language and multi-dimensional text processing tasks in the field of education.

[0118] Preferably, an open pre-trained large model containing multi-language corpora is obtained, and the open pre-trained large model is trained and fine-tuned to obtain a pre-trained language large model, including:

[0119] Choose an open pre-trained large model with multilingual corpus;

[0120] De-duplication, sentence formatting and segmentation are performed on a preset multi-corpus sample set to obtain a structured fine-tuned corpus; the multi-corpus sample set includes text information on study abroad guides, school comparisons, and professional interpretations;

[0121] Using the structured fine-tuning corpus to supervise the open pre-trained large model, so as to ensure that the vocabulary and expression ability of the model in the field of education are strengthened, and obtain a fine-tuning model for basic education scenarios;

[0122] The basic education scenario fine-tuning model is fine-tuned with small sample instructions using preset multi-language small sample instruction data to ensure that the model understands and answers education information consultation questions expressed in different languages, thereby obtaining a large pre-trained language model enhanced with multiple languages.

[0123] Specifically, step 300 of this embodiment includes:

[0124] Step 301: Select open pre-trained large model and pre-processing of multi-corpus sample sets

[0125] First, select an open pre-trained large model that supports multiple languages ​​(such as mBERT, XLM-R or LLaMA, etc.). These models have been pre-trained on multilingual corpora and have cross-language semantic understanding capabilities. Next, perform data preprocessing on multi-corpus sample sets in the field of education (such as study abroad guides, school comparisons, professional interpretations and other text information). Specifically, it includes: deduplication (eliminating duplicate content), sentence formatting (unifying punctuation, language style, etc.), segmentation (dividing long paragraphs of text into short sentences or paragraphs suitable for model training), and annotating the corpus (such as annotating language categories or subject categories). Finally, a structured fine-tuning corpus is obtained, providing a high-quality data foundation for subsequent model fine-tuning.

[0126] Step 302: Supervised fine-tuning to enhance capabilities in the education domain

[0127] The preprocessed structured fine-tuning corpus is used for supervised fine-tuning of the open pre-trained large model. By designing tasks related to the field of education (such as school information classification, professional interpretation generation, admission requirements questions and answers, etc.), the model is trained to enhance its vocabulary, expression and semantic understanding capabilities in the field of education. For example, the model enhances its understanding of specific terms and sentences in the field of education by learning common expressions in the "school comparison" corpus (such as "lower tuition fees" and "higher employment rates"). After fine-tuning, a basic education scenario fine-tuning model is obtained, which can initially cope with consulting tasks in the field of education.

[0128] Step 303: Small sample instruction fine-tuning to enhance multi-language capabilities

[0129] On the basis of fine-tuning the model for basic education scenarios, further fine-tuning is performed using preset multilingual small sample instruction data. The small sample instruction data includes education information consultation questions and their corresponding answers in different languages ​​(such as English, Chinese, French, etc.), covering common consultation scenarios (such as "recommending suitable colleges", "explaining the employment prospects of a certain major", etc.). Through small sample instruction fine-tuning, the model can better understand and answer education questions expressed in different languages, and improve its cross-language generalization ability and answer accuracy. In the end, a large pre-trained language model with multilingual enhancement is obtained, which has the ability to provide accurate education information consultation in a multilingual environment.

[0130] Optionally, step 400 of this embodiment includes:

[0131] Step 401: Input acquisition from the front-end page or voice device

[0132] Users can input natural language inquiries through front-end pages (such as web pages or mobile applications) or voice devices (such as smart speakers or voice assistants). For text input, the front-end page directly receives the text content entered by the user through a form or chat box; for voice input, the voice device converts the voice into text through voice recognition technology (such as Google Speech-to-Text or iFlytek voice recognition). Whether it is text or voice input, a natural language text will eventually be generated as input for subsequent processing.

[0133] Step 402: Preprocessing of natural language text

[0134] The acquired natural language text needs to be preprocessed to improve the model's understanding ability. Preprocessing includes removing extra spaces, normalizing punctuation, spell checking (for text input), and language detection (for multilingual input). For example, for the user input "I want to apply for a computer science major in the United States, please recommend a school", the system will normalize it to "I want to apply for a computer science major in the United States, please recommend a school". For multilingual scenarios, preprocessing also needs to identify the input language (such as Chinese, English, etc.) so that the subsequent model can select the appropriate language processing path.

[0135] Step 403: Natural language consultation intent recognition

[0136] The preprocessed natural language text is input into the pre-trained language model for intent recognition. Pre-trained language models (such as GPT, BERT, etc.) have acquired semantic understanding capabilities in the field of education through fine-tuning. The model will perform semantic analysis on the input text, identify the user's core intent (such as "recommended colleges", "compared majors", "check admission requirements", etc.) and related key information (such as target country, major, language requirements, etc.). For example, for the input "I want to apply for a computer major in the United States", the model will recognize the intent as "recommended colleges" and the key information as "United States" and "computer major".

[0137] Step 404: Extraction and structuring of key identification information

[0138] The results of the model output need to be further extracted and structured to convert the key information into a standardized format that can be used for subsequent retrieval. For example, "United States" and "Computer Science" are mapped to "Target Country: United States" and "Target Major: Computer Science" respectively. If the user input contains ambiguous information (such as "Recommend some good schools"), the system will complete the missing information through default rules or interaction with the user (such as "Which country do you want to apply to?"). Finally, the key information is stored in a structured form to support subsequent knowledge graph retrieval.

[0139] Step 405: Multiple rounds of interaction and dynamic adjustment

[0140] In some cases, the user's natural language consultation may be incomplete or require further clarification (such as "recommended schools" does not specify a country or major). The system dynamically adjusts the recognition results through multiple rounds of interaction with the user. For example, the system can generate prompt questions (such as "Which country's school do you prefer?") and update key information based on the user's answer. In addition, the system can also maintain the coherence of the conversation based on the context. For example, when the user asks questions continuously (such as "recommended schools in the United States" followed by "what scholarships are there"), the system can understand the context and associate the previous and subsequent questions to ensure the accuracy and completeness of the recognition results.

[0141] Through the above steps, this embodiment can accurately identify intent and key information from the user's natural language consultation, and provide high-quality input for subsequent knowledge graph retrieval and recommendation.

[0142] Preferably, based on the artificial intelligence reasoning engine, the knowledge graph is called to search according to the key identification information to obtain search results, including:

[0143] Mapping the key information with the entity and attribute names of the knowledge graph to construct an executable search request;

[0144] Performing a search in the knowledge graph according to the executable search request to match the college entities and associated attributes that meet the conditions to obtain preliminary results;

[0145] The preliminary results are screened and sorted based on preset business rules, and the information on admission standards, geographical location, and number of admissions in previous years is integrated, deduplicated, or merged to obtain the search results; the search results include a list of matching institutions and extended information.

[0146] Specifically, step 500 of this embodiment includes:

[0147] Step 501: Mapping of key information and knowledge graph and construction of search request

[0148] First, based on the key information extracted from the natural language consultation (such as the target country, professional direction, language requirements, etc.), it is semantically mapped with the entity and attribute names in the knowledge graph. For example, "computer science" is mapped to the professional entity in the knowledge graph, and "United States" is mapped to the geographic location attribute. Using the mapping relationship, an executable search request (such as a SPARQL query statement or a graph database query instruction) is constructed to clarify the search conditions (such as filtering the target major or colleges and universities in a specific area). This mapping needs to be combined with the semantic parsing capabilities of the artificial intelligence reasoning engine to ensure that the key information can accurately match the entities and attributes in the knowledge graph.

[0149] Step 502: Perform a search in the knowledge graph to obtain preliminary results

[0150] The constructed search request is input into the knowledge graph for execution, matching the qualified institution entities and their associated attributes (such as admission standards, geographical location, number of admissions in previous years, etc.). The knowledge graph quickly locates relevant information through its triple structure (such as "Institution A-Admission Standards-TOEFL 80") and generates preliminary results. These preliminary results may include multiple records, covering qualified institutions and their detailed attributes, providing a basis for subsequent screening and integration.

[0151] Step 503: Filtering, sorting and information integration based on business rules

[0152] The preliminary results are further processed and filtered, sorted and information integrated using preset business rules. The screening rules may include: excluding institutions that do not meet the minimum admission requirements (such as institutions with insufficient language scores), giving priority to institutions in the region or major preferred by the user; sorting rules can prioritize according to indicators that the user is concerned about (such as tuition fees, employment rates, school rankings, etc.); information integration rules are used to remove or merge duplicate information (such as different admission requirements for the same school). For example, "School A - TOEFL 80 points" and "School A - IELTS 6 points" are merged into a complete record. The implementation of these rules ensures the accuracy and practicality of the search results.

[0153] Exemplarily, the determination of business rules needs to be combined with general standards in the field of education and user needs. For example, admission standards, geographical location, language environment, etc. are core elements of general concern, and the sorting rules can be dynamically adjusted according to user preferences (such as giving priority to colleges with high employment rates or colleges with low tuition fees). In addition, business rules should have dynamic optimization capabilities and be continuously adjusted through user feedback. For example, if a user selects a certain type of college (such as a college with strong scientific research capabilities) multiple times, this embodiment can automatically adjust the rules and give priority to recommending similar colleges. This dynamic optimization mechanism makes business rules more in line with actual needs and improves user experience.

[0154] In summary, this embodiment maps key information to a knowledge graph, constructs a search request, executes the search, and combines business rule screening and integration to ultimately generate a school matching list and extended information that meets user needs, thereby achieving accurate education information consulting services.

[0155] As an optional implementation manner, step 600 of this embodiment includes:

[0156] Step 601: Structuring and data preparation of search results

[0157] After the knowledge graph search is completed, the system will return a list of qualified institutions and their associated attribute information (such as admission standards, geographic location, number of admissions in previous years, etc.). These data are usually stored in a structured format (such as JSON or database records). In order to facilitate visual display and recommendation, the data must first be sorted and optimized, including deduplication, field completion (such as missing geographic location or language environment information), data formatting (such as unifying admission requirements into a score range), etc. The sorted data will be organized according to the user's query intent and preferences (such as sorting by ranking, tuition or employment rate) to provide a clear logical structure for subsequent display.

[0158] Step 602: Design and implementation of visual display

[0159] The search results are displayed graphically on the front-end page to enhance the user's understanding and interactive experience. Visual design includes the following common forms:

[0160] (1) Tabular display: Key information such as school name, geographical location, admission criteria, tuition fees, and employment rate are listed in tabular form to facilitate quick comparison by users.

[0161] (2) Map display: By marking the geographical locations of colleges and universities on an interactive map, users can intuitively understand the distribution of target colleges and universities.

[0162] (3) Graphical presentation: Use bar charts, pie charts, or line charts to display data such as changes in the number of admissions and employment rate trends of colleges and universities to help users analyze and make decisions.

[0163] The front-end page can use modern front-end frameworks (such as React or Vue) combined with data visualization libraries (such as ECharts or D3.js) to achieve dynamic and interactive visualization effects.

[0164] Step 603: Generation of formatted recommendations

[0165] Based on the visual display, the system will generate formatted recommendations based on the user's query intent and preferences. For example, for the query "recommended colleges and universities in the United States for computer science", the system will generate a recommendation list, sorted by ranking or employment rate, with detailed information about each college (such as "College name: Massachusetts Institute of Technology; Location: Boston, USA; Admission criteria: TOEFL 100 or above; Employment rate: 95%"). Recommended content can be displayed in the form of cards, each card containing the core information of the college and related pictures (such as school badges or campus photos) to enhance the user's reading experience.

[0166] Step 604: Implementation of interactive functions

[0167] In order to enhance the user experience, the system provides a variety of interactive functions, allowing users to filter, sort and customize the recommended results. For example, users can dynamically adjust the recommended results by filtering conditions (such as "tuition is less than $30,000" or "employment rate is higher than 90%"); they can also view specific school information by clicking on the marked points on the map, or rearrange the school list by ranking, tuition and other fields through the table header sorting function. In addition, the system supports users to collect or export recommended results (such as generating PDF or Excel files) for subsequent reference and sharing.

[0168] Step 605: Multi-terminal adaptation and responsive design

[0169] To ensure that the display of recommendation results is adapted to different devices, the system adopts responsive design and supports seamless switching between desktop, mobile and tablet devices. For example, on the desktop, users can view tables and maps at the same time, while on the mobile side, the system will automatically adjust the layout and divide the tables and maps into switchable views. In addition, the system supports multi-language display, and users can switch languages ​​(such as Chinese, English, etc.) as needed to ensure the readability and global applicability of the recommended content. Through multi-terminal adaptation and responsive design, the system can provide users with a consistent and smooth user experience.

[0170] Through the above steps, this embodiment can display the search results in an intuitive and clear manner, and generate formatted recommendation content based on user needs, so as to help users make efficient decisions and improve the intelligence and personalization level of education information consulting services.

[0171] like Figure 2 As shown, this embodiment corresponds to the above method and further provides a global college information education consulting system based on an AI big model, including:

[0172] An information collection unit, which is used to collect official information in multiple languages ​​from universities around the world;

[0173] A graph construction unit, used for preprocessing and feature extraction of the official information to form a knowledge graph;

[0174] A model training unit, used to obtain an open pre-trained large model containing multi-language corpora, and train and fine-tune the open pre-trained large model to obtain a pre-trained language large model;

[0175] An intention recognition unit is used to obtain natural language consultation through a front-end page or a voice device, and input the natural language consultation into the pre-trained language large model for intention recognition to obtain key recognition information;

[0176] A retrieval unit, configured to call the knowledge graph to perform retrieval based on the key identification information based on an artificial intelligence reasoning engine to obtain retrieval results; the retrieval results include a list of colleges and universities that meet the conditions and admission standards, as well as geographic locations, language environments, and the number of students admitted in previous years associated with the colleges and universities;

[0177] A visualization unit is used to perform visualization and formatting recommendations based on the search results.

[0178] The beneficial effects of the present invention are as follows:

[0179] (1) The present invention unifies and integrates the information of universities around the world (including school names, geographical locations, major settings, admission requirements, etc.) by constructing a multilingual knowledge graph, thereby solving the problem of information dispersion and inconsistency in the prior art and greatly improving query efficiency and information coverage.

[0180] (2) The present invention uses a pre-trained large model containing multilingual corpora, and enhances its understanding ability in the field of education through fine-tuning, so that the system can support multilingual input and output, meet the educational information consultation needs of global users, and break down language barriers.

[0181] (3) The present invention uses a pre-trained language model to perform deep semantic analysis on the user's natural language consultation, extract key identification information (such as target country, major direction, admission requirements, etc.), and combines it with an artificial intelligence reasoning engine to achieve accurate and efficient knowledge retrieval, providing college and major recommendations that meet user needs.

[0182] (4) The present invention can dynamically adjust search results and recommended content based on user preferences (such as budget, geographical environment, language requirements, etc.) and contextual information, provide personalized education consulting services, and enhance user experience.

[0183] (5) The search results of the present invention are intuitively displayed in multimodal forms such as tables, comparison charts, and maps, providing users with clear and easy-to-understand information about institutions and admission standards, facilitating quick decision-making.

[0184] (6) The present invention uses the association retrieval of the knowledge graph. The system not only provides a list of colleges and universities that meet the requirements, but also displays in-depth information such as the geographical location, language environment, and number of admissions in previous years of the colleges and universities, helping users to evaluate college selection from multiple dimensions.

[0185] (7) The present invention continuously optimizes the model and knowledge graph through user feedback and interaction records, gradually improves the ability to understand and meet user needs, and forms a closed loop of sustainable improvement.

[0186] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0187] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A global college information education consulting method based on AI big model, characterized by: include: Collect official information in multiple languages ​​from universities around the world; Preprocessing and feature extracting the official information to form a knowledge graph; Obtain an open pre-trained large model containing multi-language corpora, and train and fine-tune the open pre-trained large model to obtain a pre-trained language large model; Obtain natural language consultation through a front-end page or voice device, and input the natural language consultation into the pre-trained language model for intent recognition to obtain key recognition information; Based on the artificial intelligence reasoning engine, the knowledge graph is called to search according to the key identification information to obtain search results; the search results include a list of colleges and universities that meet the conditions and admission standards, as well as the geographical location, language environment, and number of admissions in previous years associated with the colleges and universities; Visual display and formatting recommendations are performed based on the search results.

2. The global college information education consulting method based on AI big model according to claim 1 is characterized in that: The official information includes: the name of the institution, geographical location, major information, scholarship information and language requirements; the major information includes curriculum setting and duration of study.

3. The global college information education consulting method based on AI big model according to claim 1 is characterized in that: Preprocessing and feature extraction of the official information to form a knowledge graph includes: Deduplication of the official information is performed to obtain deduplication text; Based on the contextual relationship, classify the deduplicated text to obtain classified data; Based on natural language processing technology, feature extraction is performed on the elements of faculty status, scientific research capabilities, and employment rates described in the classified data to build association relationships; A knowledge graph is formed according to the elements and the association relationships.

4. The global college information education consulting method based on AI big model according to claim 3 is characterized in that: Based on the contextual relationship, the deduplicated text is classified to obtain classification data, including: Inputting the deduplicated text into a text context feature extraction layer to extract context feature information; Inputting the deduplicated text into a global feature extraction layer to extract global feature information; Fusing the contextual feature information and the global feature information to obtain a fused text feature; The fused text features are input into a trained text classification model to obtain the classification data.

5. The global college information education consulting method based on AI big model according to claim 4 is characterized in that: Inputting the deduplicated text into a text context feature extraction layer to extract context feature information includes: Using a preset pre-trained language model to extract initial feature information of the deduplicated text; Inputting the initial feature information into a forward gated recurrent unit and a reverse gated recurrent unit; The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual relationship feature information.

6. The global college information education consulting method based on AI big model according to claim 5 is characterized in that: The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual feature information, including: Using the formula: The outputs of the forward gated recurrent unit and the reverse gated recurrent unit are concatenated to obtain contextual feature information; wherein, H t ={H1,H2,…H l } represents the initial feature information input at time t, represents the output of the positive gated recurrent unit at time t, Represents the output of the reverse gated recurrent unit at time t, GRU represents the gated recurrent unit, and G represents the contextual relationship feature information.

7. The global college information education consulting method based on AI big model according to claim 6 is characterized in that: Inputting the deduplicated text into a global feature extraction layer to extract global feature information includes: The initial feature information of the deduplicated text is sequentially input into the convolution layer and the pooling layer to obtain the global feature information; wherein the global feature information extraction formula is: c i =f(ω·H+b) In the formula, f is the activation function, ω is the convolution kernel, h is the convolution kernel size, b is the bias value, c is i is the feature vector extracted by the i-th convolutional layer, It represents the value of the features extracted by three different convolution kernels after the maximum pooling layer, and C represents the extracted global feature information.

8. The global college information education consulting method based on AI big model according to claim 6 is characterized in that: Obtaining an open pre-trained large model containing multi-language corpora, and training and fine-tuning the open pre-trained large model to obtain a pre-trained language large model, including: Choose an open pre-trained large model with multilingual corpus; De-duplication, sentence formatting and segmentation are performed on a preset multi-corpus sample set to obtain a structured fine-tuned corpus; the multi-corpus sample set includes text information on study abroad guides, school comparisons, and professional interpretations; Using the structured fine-tuning corpus to supervise the open pre-trained large model, so as to ensure that the vocabulary and expression ability of the model in the field of education are strengthened, and obtain a fine-tuning model for basic education scenarios; The basic education scenario fine-tuning model is fine-tuned with small sample instructions using preset multi-language small sample instruction data to ensure that the model understands and answers education information consultation questions expressed in different languages, thereby obtaining a large pre-trained language model enhanced with multiple languages.

9. The global college information education consulting method based on AI big model according to claim 6 is characterized in that: Based on the artificial intelligence reasoning engine, the knowledge graph is called to search according to the key identification information to obtain search results, including: Mapping the key information with the entity and attribute names of the knowledge graph to construct an executable search request; Performing a search in the knowledge graph according to the executable search request to match the college entities and associated attributes that meet the conditions to obtain preliminary results; The preliminary results are screened and sorted based on preset business rules, and the information on admission standards, geographical location, and number of admissions in previous years is integrated, deduplicated, or merged to obtain the search results; the search results include a list of matching institutions and extended information.

10. A global college information education consulting system based on AI big model, characterized by: include: Information collection unit, used to collect official information in multiple languages ​​from universities around the world; A graph construction unit, used for preprocessing and feature extraction of the official information to form a knowledge graph; A model training unit, used to obtain an open pre-trained large model containing multi-language corpora, and train and fine-tune the open pre-trained large model to obtain a pre-trained language large model; An intention recognition unit is used to obtain natural language consultation through a front-end page or a voice device, and input the natural language consultation into the pre-trained language large model for intention recognition to obtain key recognition information; A retrieval unit, configured to call the knowledge graph to perform retrieval based on the key identification information based on an artificial intelligence reasoning engine to obtain retrieval results; the retrieval results include a list of colleges and universities that meet the conditions and admission standards, as well as geographic locations, language environments, and the number of students admitted in previous years associated with the colleges and universities; A visualization unit is used to perform visualization and formatting recommendations based on the search results.

Citation Information

Cited By

  • College entrance examination voluntary consultation system and method based on large language model

    CN121544432A

  • College entrance examination volunteer consultation system and method based on large language model

    CN121544432B

  • Knowledge graph-based cultural and creative education large screen content generation method, system and equipment

    CN121658627A

  • Periodical information retrieval method based on big data driving

    CN121980017A