AI large model training method and system based on education consultation big data

By cleaning, sentiment analysis, sensitive word detection and similarity calculation in the fields of educational consulting and study abroad consulting, fine-tuning and iterative training of AI models, the problem of high professionalism, complex knowledge graph construction, and stiff responses in the existing technology is solved, and efficient and accurate consultation responses are achieved, improving user experience.

CN120011503AInactive Publication Date: 2025-05-16SHANGHAI TIANQU YUNQI EDUCATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510082368.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The response methods for studying abroad consultation in the existing technology have high professional requirements, complex knowledge graph construction, stiff responses, and hallucinations, resulting in poor user experience and high cost.

Method used

By collecting large-scale Q&A data sets in the fields of educational consulting and study abroad consulting, high-quality data cleaning, sentiment analysis, sensitive word detection, similarity calculation, fine-tuning and iterative training, we can improve the performance of AI models in the fields of educational consulting and study abroad consulting.

Benefits of technology

It significantly improves the performance of AI models in the fields of educational consulting and study abroad consulting, improves user experience, reduces consultation costs, ensures the fluency and accuracy of responses, and promotes the continuous optimization and expansion of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011503A_ABST
    Figure CN120011503A_ABST
Patent Text Reader

Abstract

The invention provides an AI large model training method and system based on education consultation big data, and relates to the technical field of natural language processing, and the method comprises the steps of data collection and cleaning, sentiment analysis and sensitive word detection, similarity calculation, model fine tuning and generation effect evaluation. Through the steps, the performance and user experience of the AI model in related fields are remarkably improved, and the method has important application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to an AI big model training method and system based on education consulting big data. Background Art

[0002] In today's globalized world, studying abroad has become an important choice for many students and families. However, the field of overseas study consultation requires extremely high professionalism, and ordinary people often find it difficult to solve related overseas study problems by themselves, so they need to seek help from professional overseas study personnel. However, seeking consulting services from professionals is usually accompanied by high fees, which increases the financial burden on users.

[0003] In the prior art, solutions that use knowledge graph construction technology to achieve online overseas study consultation are gradually emerging. These solutions usually build overseas study knowledge bases based on legal provisions and question-and-answer pairs between users and lawyers, and use deep learning technologies such as semantic understanding, intent recognition, and text matching to build overseas study language understanding modules. Through the knowledge operation module, valid questions are screened out and matched to the knowledge base to obtain answers. However, these prior art overseas study consultation response methods have the following defects:

[0004] (1) High professional requirements: The complexity and professionalism of overseas study consultation issues make it difficult for ordinary users to understand and handle them, leading to an increase in consultation demand.

[0005] (2) Complex knowledge graph construction: The knowledge graph construction method in the existing technology is complex and costly, which limits its promotion in practical applications.

[0006] (3) Stiff responses: Responses based on knowledge graphs and text matching often appear stiff, lack human expression, and provide poor user experience.

[0007] (4) Hallucination problem: The generative large model may output answers that are inconsistent with the actual legal provisions, causing hallucination problems and causing trouble for users.

[0008] Therefore, the existing study abroad consultation response method needs to be improved to reduce the user's consultation cost, improve the fluency and accuracy of the response, and meet the user's actual needs for study abroad consultation. Summary of the invention

[0009] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide an AI large model training method and system based on education consulting big data. Through high-quality data cleaning, sentiment analysis, sensitive word detection, similarity calculation, fine-tuning and iterative training, it can significantly improve the performance of AI models in the fields of education consulting and study abroad consulting, enhance user experience, and promote continuous optimization and expanded application of the model, which has important practical significance and application value.

[0010] To achieve the above object, the present invention provides the following solutions:

[0011] A method for training an AI big model based on education consulting big data, comprising:

[0012] Collecting a large-scale question-and-answer data set in the field of education consulting and study abroad consulting, and cleaning the large-scale question-and-answer data set to remove invalid information and noise data to obtain cleaned data;

[0013] Using natural language processing technology to perform sentiment analysis and sensitive word detection on the cleaned data to identify and mark negative sentiment words and sensitive words, and obtain a question-answering data set marked with negative sentiment words and sensitive words;

[0014] Extracting question-answer pairs from the question-answer data set, and performing similarity calculation based on preset user questions to screen out question-answer pairs with the highest similarity to the user questions; the preset user questions are output by a pre-trained AI artificial intelligence model;

[0015] Based on the question-answer pair with the highest similarity to the user's question, fine-tune the pre-trained AI artificial intelligence model using a deep learning algorithm to obtain a trained AI artificial intelligence model;

[0016] Obtain the response result generated by the trained AI artificial intelligence model;

[0017] The generation effect is evaluated based on the response results, and the trained AI artificial intelligence model is iteratively trained based on the evaluation results.

[0018] Preferably, the large-scale question-answering dataset includes: user questions, expert responses and related scoring data.

[0019] Preferably, the large-scale question-answering dataset is cleaned to remove invalid information and noise data to obtain cleaned data, including:

[0020] Traverse the large-scale question-answering dataset Where Q i Indicates the question asked by the i-th user, A i represents the corresponding answer, and N is the total number of question-answer pairs in the large-scale question-answering dataset;

[0021] Construct a cleaning function; the formula of the cleaning function is: The validity conditions include: content integrity, sentiment analysis, sensitive word detection and repetitiveness check; the expression of content integrity is: len(Q i )>T Q And len(Ai )>T A ; where T Q and T A is a preset threshold, indicating the minimum length of questions and answers; the expression of sentiment analysis is: S(Q i )≥0 and S(A i )≥0; where S(X) represents the sentiment score of text X, ensuring that both the question and the answer have positive sentiment; the expression of sensitive word detection is: Sensitive(Q i )=0 and Sensitive(A i )=0; where Sensitive(X) indicates whether the text X contains sensitive words, and returning 0 indicates that it does not contain sensitive words; the expression for the repeatability check is: Where Dup(Q i ,A j ) indicates a question Q i Q&A with other questions j Repeat, return 0 to indicate no repeat;

[0022] The large-scale question-answering data set is cleaned according to the cleaning function to obtain the cleaned data; the expression of the cleaned data is: D′={(Q i ,A i )|i∈{1,2,…,N}, and C(Q i ,A i )=1}.

[0023] Preferably, natural language processing technology is used to perform sentiment analysis and sensitive word detection on the cleaned data to identify and mark negative sentiment words and sensitive words, and obtain a question-answering data set marked with negative sentiment words and sensitive words, including:

[0024] Use a sentiment analysis model to assign sentiment scores to each question and answer.

[0025] Set a sentiment score threshold and mark texts with scores below the threshold as negative sentiment;

[0026] Use the sensitive word library to detect each question and answer and identify the sensitive words in them;

[0027] Mark the detected sensitive words;

[0028] Based on the sentiment scores and sensitive word labeling results, a question-answering dataset labeled with negative sentiment words and sensitive words is generated.

[0029] Preferably, the question-answer pairs in the question-answer data set are extracted, and similarity calculation is performed according to preset user questions to screen out the question-answer pairs with the highest similarity to the user questions, including:

[0030] Sim(Q u ,Q i )=α·Cosine(V(Q u ),V(Q i ))+β·Semantic(Q u ,Q i )+γ·Length(Q u ,Q i )

[0031] Among them, Q u Ask a question to a preset user; V(Q) represents the vector representation of question Q, Cosine(V(Q u ),V(Q i )) is the cosine similarity, calculate Q u and Q i The similarity between vectors, the value range is between [-1,1], the closer to 1, the more similar. Semantic (Q u ,Q i ) is the semantic similarity, and Q is calculated using the pre-trained language model u and Q i The semantic similarity of the two is in the range of [0,1]. The closer to 1, the more similar the semantics. u ,Q i ) is the length similarity, calculate Q u and Q i The length difference is defined as: The value range is between [0,1]. The closer to 1, the more similar the lengths are. α, β, and γ represent the importance of cosine similarity, semantic similarity, and length similarity, respectively, in the final similarity calculation.

[0032] Preferably, based on the question-answer pair with the highest similarity to the user's question, the pre-trained AI artificial intelligence big model is fine-tuned using a deep learning algorithm to obtain a trained AI artificial intelligence big model, including:

[0033] Convert the question-answer pair with the highest similarity into the model input format to obtain the encoded input data;

[0034] Use deep learning frameworks to load pre-trained AI models;

[0035] The encoded input data is input into the pre-trained AI artificial intelligence big model for fine-tuning training to obtain a trained AI artificial intelligence big model; during the fine-tuning training process, the answer to the question-answer pair with the highest similarity is used as the target output.

[0036] Preferably, the generation effect evaluation indicators include: BLEU, ROUGE and METEOR.

[0037] An AI big model training system based on education consulting big data, including:

[0038] A data collection and cleaning unit, which is used to collect large-scale question-and-answer data sets in the fields of education consulting and overseas study consulting, and clean the large-scale question-and-answer data sets to remove invalid information and noise data to obtain cleaned data;

[0039] An analysis and detection unit, used to perform sentiment analysis and sensitive word detection on the cleaned data using natural language processing technology to identify and mark negative sentiment words and sensitive words, and obtain a question-answering data set marked with negative sentiment words and sensitive words;

[0040] A similarity calculation unit, used to extract question-answer pairs from the question-answer data set, and perform similarity calculation based on preset user questions to screen out question-answer pairs with the highest similarity to the user questions; the preset user questions are output by a pre-trained AI artificial intelligence large model;

[0041] A model fine-tuning unit, configured to fine-tune the pre-trained AI artificial intelligence big model using a deep learning algorithm based on the question-answer pair with the highest similarity to the user's question, to obtain a trained AI artificial intelligence big model;

[0042] A result acquisition unit, used to obtain the response result generated by the trained AI artificial intelligence large model;

[0043] An evaluation unit is used to perform generation effect evaluation according to the response result, and iteratively train the trained AI artificial intelligence large model according to the evaluation result.

[0044] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0045] The present invention provides an AI big model training method and system based on education consulting big data, the method comprising: collecting a large-scale question and answer data set in the field of education consulting and overseas study consulting, and cleaning the large-scale question and answer data set to remove invalid information and noise data to obtain cleaned data; using natural language processing technology to perform sentiment analysis and sensitive word detection on the cleaned data to identify and mark negative sentiment words and sensitive words, and obtain a question and answer data set marked with negative sentiment words and sensitive words; extracting question and answer pairs in the question and answer data set, and performing similarity calculation based on preset user questions to screen out question and answer pairs with the highest similarity to the user questions; the preset user questions are obtained by outputting a pre-trained AI artificial intelligence big model; based on the question and answer pairs with the highest similarity to the user questions, using a deep learning algorithm to fine-tune the pre-trained AI artificial intelligence big model to obtain a trained AI artificial intelligence big model; obtaining a reply result generated by the trained AI artificial intelligence big model; performing generation effect evaluation based on the reply result, and iteratively training the trained AI artificial intelligence big model based on the evaluation result. Through high-quality data cleaning, sentiment analysis, sensitive word detection, similarity calculation, fine-tuning and iterative training, the present invention can significantly improve the performance of AI models in the fields of education consulting and study abroad consulting, enhance user experience, and promote continuous optimization and expanded application of models. It has important practical significance and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0047] Figure 1 A flow chart of a method provided by an embodiment of the present invention;

[0048] Figure 2 A schematic diagram of the system structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] The purpose of the present invention is to provide an AI large model training method and system based on education consulting big data. Through high-quality data cleaning, sentiment analysis, sensitive word detection, similarity calculation, fine-tuning and iterative training, it can significantly improve the performance of AI models in the fields of education consulting and study abroad consulting, enhance user experience, and promote the continuous optimization and expanded application of the model, which has important practical significance and application value.

[0051] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] Figure 1 A flow chart of a method provided by an embodiment of the present invention, such as Figure 1 As shown, the present invention provides an AI big model training method based on education consulting big data, including:

[0053] Step S100: collecting a large-scale question-and-answer data set in the field of education consulting and overseas study consulting, and cleaning the large-scale question-and-answer data set to remove invalid information and noise data to obtain cleaned data;

[0054] Step S200: using natural language processing technology to perform sentiment analysis and sensitive word detection on the cleaned data to identify and mark negative sentiment words and sensitive words, and obtain a question-answering data set marked with negative sentiment words and sensitive words;

[0055] Step S300: extracting question-answer pairs from the question-answer data set, and performing similarity calculation based on preset user questions to screen out question-answer pairs with the highest similarity to the user questions; the preset user questions are output by a pre-trained AI artificial intelligence model;

[0056] Step S400: Based on the question-answer pair with the highest similarity to the user's question, fine-tune the pre-trained AI artificial intelligence model using a deep learning algorithm to obtain a trained AI artificial intelligence model;

[0057] Step S500: Obtain the response result generated by the trained AI artificial intelligence model;

[0058] Step S600: Evaluate the generation effect based on the response result, and iteratively train the trained AI artificial intelligence model based on the evaluation result.

[0059] Preferably, the large-scale question-answering dataset includes: user questions, expert responses and related scoring data.

[0060] Specifically, this embodiment first clarifies the sources of data collection, including online education platforms, overseas study consulting websites, social media, forums, and question-and-answer communities. These platforms usually contain a large number of records of user questions and expert responses. Data can be obtained through API interfaces, web crawlers, or manual collection. Ensure that the selected data sources are authoritative and reliable to improve the quality of the data set.

[0061] Use web crawler tools (such as Scrapy, BeautifulSoup, etc.) to write crawler programs to automatically crawl question and answer data on the selected platform. During the crawling process, focus on user questions, expert responses, and related rating data. In order to ensure the integrity and accuracy of the data, you can set the crawler's crawling frequency and depth to avoid placing too much burden on the target website. The crawled data should be stored in a structured database (such as MySQL, MongoDB, etc.) for subsequent processing and analysis.

[0062] The collected data often contains noise and invalid information, so data cleaning is required. The cleaning process includes removing duplicates, filtering invalid questions (such as meaningless characters, advertisements, etc.), and standardizing the reply format. In addition, the question-answer pairs can be annotated through user rating data to screen out high-quality question-answer pairs. Finally, a structured large-scale question-answer dataset is compiled, which includes user questions, expert replies, and their rating information, laying the foundation for subsequent model training and analysis.

[0063] Preferably, the large-scale question-answering dataset is cleaned to remove invalid information and noise data to obtain cleaned data, including:

[0064] Traverse the large-scale question-answering dataset Where Q i Indicates the question asked by the i-th user, A i represents the corresponding answer, and N is the total number of question-answer pairs in the large-scale question-answering dataset;

[0065] Construct a cleaning function; the formula of the cleaning function is: The validity conditions include: content integrity, sentiment analysis, sensitive word detection and repetitiveness check; the expression of content integrity is: len(Q i )>T Q And len(A i )>T A ; where T Q and T A is a preset threshold, indicating the minimum length of questions and answers; the expression of sentiment analysis is: S(Q i )≥0 and S(A i)≥0; where S(X) represents the sentiment score of text X, ensuring that both the question and the answer have positive sentiment; the expression of sensitive word detection is: Sensitive(Q i )=0 and Sensitive(A i )=0; where Sensitive(X) indicates whether the text X contains sensitive words, and returning 0 indicates that it does not contain sensitive words; the expression for the repeatability check is: Where Dup(Q i ,A j ) indicates a question Q i Q&A with other questions j Repeat, return 0 to indicate no repeat;

[0066] The large-scale question-answering data set is cleaned according to the cleaning function to obtain the cleaned data; the expression of the cleaned data is: D′={(Q i ,A i )|i∈{1,2,…,N}, and C(Q i ,A i )=1}.

[0067] Specifically, the cleaning function of this embodiment not only focuses on the integrity of the data, but also comprehensively considers multiple dimensions such as sentiment analysis, sensitive word detection, and repetitiveness check. This multi-level cleaning standard ensures the high quality of the data set, so that the final cleaned data is not only complete in content, but also emotionally positive, in line with social norms, and avoids potential negative effects. This comprehensive method is relatively rare in the traditional data cleaning process and can effectively improve the practicality and security of the data set.

[0068] In the cleaning process, the combination of sentiment analysis and sensitive word detection reflects a deep understanding of the text content. Sentiment scoring ensures that questions and answers are positive, which can effectively filter out content that may cause negative emotions. At the same time, sensitive word detection further ensures data compliance and avoids the emergence of sensitive topics. This innovative combination makes data cleaning not just a formal processing, but a deep control of data content, which improves the quality and applicability of the data set.

[0069] By introducing a repeatability check mechanism, the cleaning function can effectively identify and remove repeated questions, avoiding data redundancy. This measure not only improves the streamlining of the data set, but also enhances the efficiency of model training, because the model can focus on unique question-answer pairs rather than repeated information during training. This targeted data processing method is of great significance in cleaning large-scale data sets, and can significantly improve the results of subsequent analysis and model training.

[0070] In summary, the cleaning function of this embodiment demonstrates its creativity in the data cleaning process through the combination of comprehensive standards, sentiment and sensitive words, and the introduction of repeatability checks, ensuring the high quality and high applicability of the final cleaned data.

[0071] Preferably, natural language processing technology is used to perform sentiment analysis and sensitive word detection on the cleaned data to identify and mark negative sentiment words and sensitive words, and obtain a question-answering data set marked with negative sentiment words and sensitive words, including:

[0072] Use a sentiment analysis model to assign sentiment scores to each question and answer.

[0073] Set a sentiment score threshold and mark texts with scores below the threshold as negative sentiment;

[0074] Use the sensitive word library to detect each question and answer and identify the sensitive words in them;

[0075] Mark the detected sensitive words;

[0076] Based on the sentiment scores and sensitive word labeling results, a question-answering dataset labeled with negative sentiment words and sensitive words is generated.

[0077] Specifically, this embodiment first selects a suitable sentiment analysis model. Pre-trained sentiment analysis models can be used, such as models based on BERT or LSTM, which perform well in sentiment classification tasks. When sentiment scoring is performed for each question and answer, the model analyzes the sentiment tendency in the text and outputs a sentiment score, which usually ranges from -1 to 1, indicating the intensity of negative to positive sentiment. By sentiment scoring each text, a basis can be provided for subsequent negative sentiment marking.

[0078] After obtaining the sentiment score, you need to set a sentiment score threshold to distinguish between positive and negative sentiment. This threshold can be adjusted according to the characteristics of the dataset and actual needs. For example, the threshold can be set to 0.3, which means that texts with a score lower than 0.3 will be marked as negative sentiment. In this way, questions and answers with poor sentiment can be effectively filtered out, providing clear markings for subsequent data processing.

[0079] The key to sensitive word detection is to build a comprehensive sensitive word library. The library should contain sensitive words related to education consulting and overseas study consulting, including sensitive content in politics, religion, gender, etc. When using the sensitive word library to detect each question and answer, sensitive words in the text can be identified through text matching algorithms (such as regular expressions or string matching). The detected sensitive words will be marked for subsequent processing.

[0080] After completing sentiment analysis and sensitive word detection, the tagging results of these two parts need to be integrated. For each question-answer pair, record its sentiment score, whether it contains negative sentiment tags, and whether it contains sensitive word tags. You can use a structured data format (such as JSON or CSV) to store this information to ensure that the sentiment status and sensitive word status of each question-answer pair can be clearly presented.

[0081] Finally, based on the sentiment scores and sensitive word labeling results, a new question-answering dataset is generated, which contains question-answer pairs labeled with negative sentiment words and sensitive words. This labeled dataset will provide an important reference for subsequent model training and analysis, ensuring that the model can effectively identify and respond to negative sentiment and sensitive content when processing user questions, thereby improving system security and user experience.

[0082] Through the above steps, this embodiment can effectively use natural language processing technology to perform sentiment analysis and sensitive word detection on the cleaned data, and finally obtain a high-quality labeled question and answer dataset.

[0083] Preferably, the question-answer pairs in the question-answer data set are extracted, and similarity calculation is performed according to preset user questions to screen out the question-answer pairs with the highest similarity to the user questions, including:

[0084] Sim(Q u ,Q i )=α·Cosine(V(Q u ),V(Q i ))+β·Semantic(Q u ,Q i )+γ·Length(Q u ,Q i )

[0085] Among them, Q u Ask a question to a preset user; V(Q) represents the vector representation of question Q, Cosine(V(Q u ),V(Q i )) is the cosine similarity, calculate Q u and Q i The similarity between vectors, the value range is between [-1,1], the closer to 1, the more similar. Semantic (Q u ,Q i ) is the semantic similarity, and Q is calculated using the pre-trained language model u and Q i The semantic similarity of the two is in the range of [0,1]. The closer to 1, the more similar the semantics. u ,Q i ) is the length similarity, calculate Q uand Q i The length difference is defined as: The value range is between [0,1]. The closer to 1, the more similar the lengths are. α, β, and γ represent the importance of cosine similarity, semantic similarity, and length similarity, respectively, in the final similarity calculation.

[0086] Specifically, this embodiment first extracts all question-answer pairs from the cleaned question-answer data set. Each question-answer pair includes a user question and a corresponding expert answer. Ensure that the extracted data structure is clear for subsequent processing and calculation. The extracted question-answer pairs will be used to calculate similarity with the preset user questions.

[0087] Use natural language processing technology to convert the preset user questions into vector representations. This can be achieved in a variety of ways, such as using TF-IDF, Word2Vec, or the more advanced BERT model. The vector representation can capture the semantic information of the text and provide a basis for subsequent similarity calculations.

[0088] For each extracted question-answer pair, the vector representation of the question part is calculated, and the similarity is calculated with the vector representation of the preset user question. The similarity calculation includes three main aspects:

[0089] (1) Cosine similarity: By calculating the cosine similarity between question vectors, the directional similarity of two questions is evaluated. The value range of cosine similarity is between -1 and 1. The closer it is to 1, the more semantically similar the two questions are. This method can effectively capture the relative similarity of texts and is particularly suitable for high-dimensional sparse data.

[0090] (2) Semantic similarity: Use a pre-trained language model (such as BERT) to calculate the semantic similarity of questions. This method can take into account the contextual relationship of word meanings and provide a deeper level of semantic understanding. The value of semantic similarity ranges from 0 to 1, and the closer it is to 1, the more semantically similar the two questions are. This method is particularly suitable for dealing with differences caused by synonyms and context changes.

[0091] (3) Length similarity: Calculate the length difference of the questions to evaluate the similarity of the number of words in the questions. The value of length similarity ranges from 0 to 1, and the closer to 1, the more similar the lengths of the two questions are. Length similarity can help identify questions with similar structures, although it does not directly reflect the semantic content.

[0092] In the final similarity calculation, the weights of cosine similarity, semantic similarity, and length similarity can be determined through experiments and cross-validation. These weights can be adjusted based on the model's performance on the validation set to optimize the model's accuracy and response quality. For example, if the model performs well in semantic understanding, the weight of semantic similarity can be increased; if length similarity has a greater impact on the results, its weight can be adjusted accordingly. In this way, it can be ensured that the final similarity calculation can better reflect the true similarity between the user's question and the question-answer pair.

[0093] Finally, based on the calculated similarity values, several question-answer pairs with the highest similarity to the preset user questions are selected. These question-answer pairs will serve as the basis for subsequent response generation to ensure that the model can provide the most relevant answers to user needs.

[0094] Through the above steps, the question-answer pairs in the question-answer dataset can be effectively extracted, and similarity calculations can be performed based on preset user questions to screen out the most relevant question-answer pairs.

[0095] Preferably, based on the question-answer pair with the highest similarity to the user's question, the pre-trained AI artificial intelligence big model is fine-tuned using a deep learning algorithm to obtain a trained AI artificial intelligence big model, including:

[0096] Convert the question-answer pair with the highest similarity into the model input format to obtain the encoded input data;

[0097] Use deep learning frameworks to load pre-trained AI models;

[0098] The encoded input data is input into the pre-trained AI artificial intelligence big model for fine-tuning training to obtain a trained AI artificial intelligence big model; during the fine-tuning training process, the answer to the question-answer pair with the highest similarity is used as the target output.

[0099] Optionally, this embodiment first converts the screened question-answer pairs with the highest similarity to the user's question into a model input format. This process usually includes text preprocessing of questions and answers, such as removing extra spaces, standardizing punctuation, etc. Next, the text is encoded using natural language processing technology to generate an input format that the model can accept. Common encoding methods include using Tokenization to convert text into a vocabulary index, or using more complex embedding techniques (such as Word2Vec or BERT) to generate vector representations. Ultimately, the encoded input data obtained will be prepared in batches to facilitate subsequent model training.

[0100] Use a deep learning framework such as TensorFlow or PyTorch to load a pre-trained AI model. Pre-trained models are usually trained on large-scale datasets and have certain language understanding and generation capabilities. When loading a model, you can select a suitable model architecture and parameter settings to ensure that it is suitable for subsequent fine-tuning training. At this point, the model's weights and structure will be initialized to the pre-trained state, providing a good foundation for fine-tuning.

[0101] The encoded input data is input into the pre-trained AI model for fine-tuning training. During the training process, the answer to the question-answer pair with the highest similarity to the user's question is used as the target output. The loss between the output generated by the model and the target output (such as cross entropy loss) is calculated, and the parameters of the model are updated using the back propagation algorithm. The goal of fine-tuning training is to enable the model to better understand and generate answers related to user questions, thereby improving the performance of the model in specific areas (such as education consulting and study abroad consulting). After the training is completed, the resulting model will be an optimized, trained AI model that can respond to user questions more accurately.

[0102] Preferably, the generation effect evaluation indicators include: BLEU, ROUGE and METEOR.

[0103] Specifically, this embodiment first uses the trained AI artificial intelligence model to generate a series of response results after the model fine-tuning is completed. These response results should be based on a set of preset user questions to ensure that different types of questions and scenarios are covered. The generated response results will be used as the basic data for evaluation and compared with the reference answers. The reference answer can be a standard response provided by an expert, or an answer extracted from a high-quality question and answer dataset.

[0104] In the evaluation of generation effects, it is crucial to choose the right evaluation indicators. Common evaluation indicators include BLEU, ROUGE and METEOR. BLEU (Bilingual Evaluation Understudy) is mainly used to evaluate the degree of n-gram overlap between the generated text and the reference text, and is suitable for tasks such as machine translation. ROUGE (Recall-Oriented Understudy for Gisting Evaluation) focuses on the recall rate and is often used for text summary evaluation. It can measure the similarity between the generated text and the reference text. METEOR (Metric for Evaluation of Translation with Explicit ORdering) combines precision and recall, and takes into account word form changes and synonyms. It is suitable for a variety of text generation tasks.

[0105] Use the corresponding evaluation tools or libraries (such as NLTK, SacréBLEU, etc.) to calculate the BLEU, ROUGE, and METEOR scores of the generated response results. For each generated response, compare it with the corresponding reference answer and calculate the scores of each indicator. These scores will reflect the quality and accuracy of the model when generating text. During the evaluation process, different n-gram values ​​(such as 1-gram, 2-gram, etc.) can be set to obtain more detailed evaluation results.

[0106] After calculating each evaluation indicator, analyze the results. You can use visualization tools (such as Matplotlib or Seaborn) to display the evaluation results graphically, so that you can observe the performance of the model on different types of problems. During the analysis, pay attention to the model's scores on each indicator and identify areas or types of problems with poor performance. This analysis will provide an important basis for subsequent iterative training and help determine the direction of improvement.

[0107] Based on the evaluation results, decide whether to iteratively train the trained AI model. If the scores of certain indicators are lower than the preset threshold, or if the performance is poor on a specific type of question, more relevant data can be collected for retraining. Iterative training can be performed by introducing new question-answer pairs, adjusting model parameters, or optimizing training strategies to improve the generation effect of the model. The iterative training process should continue until the model reaches a satisfactory level in various evaluation indicators, thereby ensuring its effectiveness and reliability in practical applications.

[0108] Corresponding to the above method, such as Figure 2 As shown, this embodiment also provides an AI big model training system based on education consulting big data, including:

[0109] A data collection and cleaning unit, which is used to collect large-scale question-and-answer data sets in the fields of education consulting and overseas study consulting, and clean the large-scale question-and-answer data sets to remove invalid information and noise data to obtain cleaned data;

[0110] An analysis and detection unit, used to perform sentiment analysis and sensitive word detection on the cleaned data using natural language processing technology to identify and mark negative sentiment words and sensitive words, and obtain a question-answering data set marked with negative sentiment words and sensitive words;

[0111] A similarity calculation unit, used to extract question-answer pairs from the question-answer data set, and perform similarity calculation based on preset user questions to screen out question-answer pairs with the highest similarity to the user questions; the preset user questions are output by a pre-trained AI artificial intelligence large model;

[0112] A model fine-tuning unit, configured to fine-tune the pre-trained AI artificial intelligence big model using a deep learning algorithm based on the question-answer pair with the highest similarity to the user's question, to obtain a trained AI artificial intelligence big model;

[0113] A result acquisition unit, used to obtain the response result generated by the trained AI artificial intelligence large model;

[0114] An evaluation unit is used to perform generation effect evaluation according to the response result, and iteratively train the trained AI artificial intelligence large model according to the evaluation result.

[0115] The beneficial effects of the present invention are as follows:

[0116] (1) The present invention ensures the quality of training data by cleaning and labeling data, removing invalid information and noise data, and thereby improving the accuracy of the model's understanding and response to user questions. By screening the question-answer pairs with the highest similarity to the user's questions, the model can better match the user's needs and provide more relevant answers.

[0117] (2) By performing sentiment analysis on question-and-answer data, the model can identify negative sentiment words, thereby taking sentiment factors into account when generating replies and providing more empathetic answers; marking sensitive words can ensure that the generated replies comply with social norms and ethical standards and avoid the appearance of inappropriate content.

[0118] (3) The present invention fine-tunes the model so that it can generate personalized responses based on the user's specific questions, thereby improving the user's satisfaction and experience; the trained model can quickly generate high-quality responses, reduce user waiting time, and improve consultation efficiency.

[0119] (4) By generating effect evaluation and user feedback, the model of the present invention can continuously learn and improve, adapt to changes in user needs, and maintain its effectiveness in the fields of education consulting and study abroad consulting; by utilizing feedback from large-scale question-and-answer data sets, the model can be continuously optimized in practical applications to improve its intelligence level.

[0120] (5) The training method of the present invention is not only applicable to education and study abroad consulting, but can also be extended to question-answering systems in other fields, such as medicine and law, and has broad application potential. With the continuous accumulation of data and iteration of models, the system can form a rich knowledge base and provide users with more comprehensive consulting services.

[0121] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0122] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. An AI big model training method based on education consulting big data, characterized in that: include: Collecting a large-scale question-and-answer data set in the field of education consulting and study abroad consulting, and cleaning the large-scale question-and-answer data set to remove invalid information and noise data to obtain cleaned data; Using natural language processing technology to perform sentiment analysis and sensitive word detection on the cleaned data to identify and mark negative sentiment words and sensitive words, and obtain a question-answering data set marked with negative sentiment words and sensitive words; Extracting question-answer pairs from the question-answer data set, and performing similarity calculation based on preset user questions to screen out question-answer pairs with the highest similarity to the user questions; The preset user questions are obtained by outputting the pre-trained AI artificial intelligence model; Based on the question-answer pair with the highest similarity to the user's question, fine-tune the pre-trained AI artificial intelligence model using a deep learning algorithm to obtain a trained AI artificial intelligence model; Obtain the response result generated by the trained AI artificial intelligence model; The generation effect is evaluated based on the response results, and the trained AI artificial intelligence model is iteratively trained based on the evaluation results.

2. The AI ​​big model training method based on education consulting big data according to claim 1 is characterized in that: The large-scale question-answering dataset includes: user questions, expert replies and related scoring data.

3. The AI ​​big model training method based on education consulting big data according to claim 1 is characterized in that: The large-scale question-answering dataset is cleaned to remove invalid information and noise data to obtain cleaned data, including: Traverse the large-scale question-answering dataset Where Q i Indicates the question asked by the i-th user, A i represents the corresponding answer, and N is the total number of question-answer pairs in the large-scale question-answering dataset; Construct a cleaning function; the formula of the cleaning function is: The validity conditions include: content integrity, sentiment analysis, sensitive word detection and repetitiveness check; the expression of content integrity is: len(Q i ) >T Q And len(A i ) >T A ; where T Q and T A is a preset threshold, indicating the minimum length of questions and answers; the expression of sentiment analysis is: S(Q i )≥0 and S(A i )≥0; where S(X) represents the sentiment score of text X, ensuring that both the question and the answer have positive sentiment; the expression of sensitive word detection is: Sensitive(Q i )=0 and Sensitive(A i )=0; where Sensitive(X) indicates whether the text X contains sensitive words, and returning 0 indicates that it does not contain sensitive words; the expression for the repeatability check is: Where Dup(Q i ,A j ) indicates a question Q i Q&A with other questions j Repeat, return 0 to indicate no repeat; The large-scale question-answering data set is cleaned according to the cleaning function to obtain the cleaned data; the expression of the cleaned data is: D′={(Q i ,A i )|i∈{1,2,…,N}, and C(Q i ,A i )=1}.

4. The AI ​​big model training method based on education consulting big data according to claim 1 is characterized in that: The cleaned data is subjected to sentiment analysis and sensitive word detection using natural language processing technology to identify and mark negative sentiment words and sensitive words, and a question-answering data set marked with negative sentiment words and sensitive words is obtained, including: Use a sentiment analysis model to assign sentiment scores to each question and answer. Set a sentiment score threshold and mark texts with scores below the threshold as negative sentiment; Use the sensitive word library to detect each question and answer and identify the sensitive words in them; Mark the detected sensitive words; Based on the sentiment scores and sensitive word labeling results, a question-answering dataset labeled with negative sentiment words and sensitive words is generated.

5. The AI ​​big model training method based on education consulting big data according to claim 3 is characterized in that: Extract the question-answer pairs in the question-answer data set, and perform similarity calculation based on the preset user questions to screen out the question-answer pairs with the highest similarity to the user questions, including: Sim(Q u ,Q i )=α·Cosine(V(Q u ),V(Q i ))+β·Semantic(Q u ,Q i )+γ·Length(Q u ,Q i ) Among them, Q u Ask a question to a preset user; V(Q) represents the vector representation of question Q, Cosine(V(Q u ),V(Q i )) is the cosine similarity, calculate Q u and Q i The similarity between vectors, the value range is between [-1,1], the closer to 1, the more similar. Semantic (Q u ,Q i ) is the semantic similarity, and Q is calculated using the pre-trained language model u and Q i The semantic similarity of the , the value range is between [0,1], the closer to 1, the more similar the semantics, Length ( Q u ,Q i) For length similarity, calculate Q u and Q i The length difference is defined as: The value range is between [0,1]. The closer to 1, the more similar the lengths are. α, β, and γ represent the importance of cosine similarity, semantic similarity, and length similarity, respectively, in the final similarity calculation.

6. The AI ​​big model training method based on education consulting big data according to claim 1 is characterized in that: Based on the question-answer pair with the highest similarity to the user's question, the pre-trained AI artificial intelligence model is fine-tuned using a deep learning algorithm to obtain a trained AI artificial intelligence model, including: Convert the question-answer pair with the highest similarity into the model input format to obtain the encoded input data; Use deep learning frameworks to load pre-trained AI models; The encoded input data is input into the pre-trained AI artificial intelligence big model for fine-tuning training to obtain a trained AI artificial intelligence big model; during the fine-tuning training process, the answer to the question-answer pair with the highest similarity is used as the target output.

7. The AI ​​big model training method based on education consulting big data according to claim 1 is characterized in that: The indicators for evaluating the generation effect include: BLEU, ROUGE and METEOR.

8. An AI big model training system based on education consulting big data, characterized in that: include: A data collection and cleaning unit, which is used to collect large-scale question-and-answer data sets in the fields of education consulting and overseas study consulting, and clean the large-scale question-and-answer data sets to remove invalid information and noise data to obtain cleaned data; An analysis and detection unit, used to perform sentiment analysis and sensitive word detection on the cleaned data using natural language processing technology to identify and mark negative sentiment words and sensitive words, and obtain a question-answering data set marked with negative sentiment words and sensitive words; A similarity calculation unit, used to extract question-answer pairs from the question-answer data set, and perform similarity calculation based on preset user questions to screen out question-answer pairs with the highest similarity to the user questions; The preset user questions are obtained by outputting the pre-trained AI artificial intelligence model; A model fine-tuning unit, configured to fine-tune the pre-trained AI artificial intelligence big model using a deep learning algorithm based on the question-answer pair with the highest similarity to the user's question, to obtain a trained AI artificial intelligence big model; A result acquisition unit, used to obtain the response result generated by the trained AI artificial intelligence large model; An evaluation unit is used to perform generation effect evaluation according to the response result, and iteratively train the trained AI artificial intelligence large model according to the evaluation result.