Intelligent enrollment question-answering system based on pre-training model

By designing an enrollment intelligent Q&A system based on pre-training models, the lack of AI vertical big models in the field of graduate enrollment in the existing technology is solved, and a high accuracy and low error rate question and answer system is realized, which improves the intelligence and output quality of the system.

CN120104731APending Publication Date: 2025-06-06JINLING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510084679.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively utilize AI vertical large models to provide an intelligent question-and-answer system for graduate enrollment, and there are problems such as information enclosure and management difficulties.

Method used

An intelligent admissions question-and-answer system based on pre-trained models is designed, including a knowledge base module, an extracted converter memory network module, a generative converter memory network module and a feedback module. Through these modules, the knowledge base in the admissions field is built, and the pre-trained model is used to process and optimize the question-and-answer tasks.

Benefits of technology

It achieves higher accuracy and lower error rates, reduces resource utilization, and reduces error rates to less than 3%, while improving the intelligence level of the system and the naturalness and fluency of output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104731A_ABST
    Figure CN120104731A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent enrollment question-answering system based on a pre-training model. The intelligent enrollment question-answering system comprises a knowledge base module, an extraction type converter memory network module, a generation type converter memory network module and a feedback module, wherein the knowledge base module is used for constructing a knowledge base of an enrollment field; the extraction type converter memory network module and the generation type converter memory network module are respectively pre-trained; the extraction type converter memory network module is used for extracting an optimal answer from a given alternative answer set according to historical question and answer records and background knowledge obtained through retrieval; the generative converter memory network module obtains a final output answer according to the optimal answer output by the extraction converter memory network module and the input question; and the feedback module is used for adjusting and optimizing the generative converter memory network module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent question-answering system for enrollment, in particular to an intelligent question-answering system for enrollment based on a pre-training model. Background Art

[0002] This section merely provides background information related to the present disclosure and is not necessarily prior art.

[0003] The admissions offices of various universities have received a large number of personalized admissions inquiries about the school's annual enrollment plans, the adjustment of enrollment majors, the proportion of recommended students, etc. Considering the closedness and confidentiality of information and the accuracy of answers, most schools use human

[0004] However, this method is costly and difficult to manage.

[0005] Artificial Intelligence (AI) (reference: LAUDON KC, LAUDON JP. Essentials of management information systems [M]. Pearson, 2017: 20-45.) is a science that studies and develops theories, methods, technologies and applications for simulating, extending and expanding human intelligence. In 1956, it was invented by scientists such as McCarthy (J.McCarthy), Minsky (MLMinsky), Shannon (CEShannon), Simon (HASimon) and Samuel (ALSamuel) and went through the first wave through symbolic reasoning and handcrafted knowledge. In 1976, the introduction of statistical learning (StatisticalLearning) and neural networks (Neural Networks) ushered in the second wave. In 2006, deep learning (DeepLearning) and representation learning (Representation Learning) were added to develop rapidly again. In 2022, the era of ChatGPT with large models has arrived. Every development has shocked human society.

[0006] The main technologies and modules of AI big models are: Transformer model, multi-head self-attention mechanism, encoder, decoder, proximal strategy optimization technology, human feedback reinforcement learning technology (reference: Cai Rui, Ge Jun, Sun Zhe, et al. A review of the development of AI pre-trained big models [J / OL]. Small Microcomputer Systems: 1-12 [2024-05-11].), which are mainly divided into two categories. One is AI general big models, which can complete tasks in various scenarios without fine-tuning or with a small amount of fine-tuning, such as the GPT series, LLaMA series, BERT, ALBERT, RoBERTa, DeBERTa, etc.; the other is AI vertical big models, which can be trained and optimized in specific industry fields and conduct large-scale deep learning, such as the common big models used in five fields of finance, medicine, law, natural science, and code programming. Compared with AI general big models, AI vertical big models can be deeply trained in specific fields, efficiently handle tasks and problems in the field, and provide more accurate solutions for specific problems in the field.

[0007] With the rapid development of big data technology, natural language processing (NLP) technology, and artificial intelligence technology, how to use the latest technology to replace manual labor with artificial intelligence is an important innovation in management and service. Pre-training model is a common artificial intelligence model, which can learn to obtain contextual semantics on a large amount of unlabeled corpus (reference: Hu Shunbang, Wang Lin, Liu Wuying. Early detection of false news based on pre-training representation and width learning [J / OL]. Journal of Zhengzhou University (Science Edition): 1-6 [2024-04-30].). Recently, many scholars have achieved good results in multiple tasks of natural language processing based on AI pre-training models. Cai Rui et al. introduced and summarized some core technologies of AI pre-training models, studied the development of models, and discussed the limitations and future development of models in various fields. Hu Shunbang et al. proposed an early detection method for false news based on AI pre-training model representation and width learning. The experimental results show that the proposed method has an accuracy rate of more than 80% in detecting false news within 4 hours, which is better than the baseline method. Miao Yi et al. (reference: Miao Yi, Zhang Weifeng, Xu Ling. Video moment retrieval pre-training model based on CLIP [J / OL]. Computer Application Research: 1-8 [2024-05-23].) proposed a video moment retrieval network based on AI pre-training model. The experimental results show that the proposed network effectively learns video visual features and temporal features under the guidance of query text. Yao Chengwei et al. (reference: Yao Chengwei, Chen Chunhui, Chen Mei. AI natural language generation experiment and teaching design for liberal arts students [J]. Experimental Technology and Management, 2024, 41(04): 177-184.) proposed a set of AI natural language generation experiments and teaching plans for liberal arts students. After the experimental plan was implemented in experimental teaching, a set of operational and replicable teaching materials were formed and opened, which is conducive to the development of similar experimental teaching work.

[0008] However, the above existing technologies all have the following defects: 1) Considering the closed and confidential nature of each school's past postgraduate entrance examination information, candidates cannot master the relevant data of multiple schools' past postgraduate entrance examinations at the same time. At the same time, this system also adds expert summary predictions. 2) In the current AI-dominated network, there is no AI vertical large-scale intelligent system for graduate student admissions.

[0009] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0010] Purpose of the invention: The technical problem to be solved by the present invention is to provide an intelligent question-answering system for enrollment based on a pre-trained model in view of the deficiencies in the prior art.

[0011] In order to solve the above technical problems, the present invention discloses an intelligent question-answering system for enrollment based on a pre-training model, comprising:

[0012] A knowledge base module, an extractive transformer memory network module, a generative transformer memory network module and a feedback module; wherein,

[0013] The knowledge base module is used to build a knowledge base in the field of enrollment;

[0014] The extractive transformer memory network module and the generative transformer memory network module are pre-trained respectively;

[0015] The extractive transformer memory network module extracts the best answer from a given set of candidate answers based on historical question and answer records and retrieved background knowledge;

[0016] The generative transformer memory network module obtains a final output answer according to the best answer output by the extractive transformer memory network module and the input question;

[0017] The feedback module is used to adjust and optimize the generative converter memory network module.

[0018] Furthermore, the generative transformer memory network module comprises:

[0019] A first memory network and a generative transformer; wherein the first memory network is used to retrieve background knowledge; and the generative transformer generates text output with reference to the retrieved background knowledge.

[0020] Furthermore, the first memory network includes:

[0021] Transformation encoder, retrieval language model and output module; where,

[0022] The retrieval language model retrieves the best matching text in the knowledge base according to the input question;

[0023] The conversion encoder converts the input multiple answer texts and the context of the question into internal features;

[0024] Update the best answer text using the new question context;

[0025] Use the output module to output the features corresponding to the best answer text.

[0026] Furthermore, the generative converter comprises:

[0027] The generative transformation decoder uses a self-attention layer and a feed-forward neural network to generate the answer text based on the output features of the first memory network.

[0028] Furthermore, the extractive transformer memory network module converts the input multiple answer texts and the context of the question into internal features, including:

[0029] A second memory network and an extractive transformer; wherein the second memory network is used to retrieve knowledge; and the extractive transformer selects the most appropriate feature as output based on the knowledge retrieved by the second memory network.

[0030] Furthermore, the second memory network includes:

[0031] Transformation encoder, retrieval language model and output module; where,

[0032] The retrieval language model retrieves multiple background knowledge texts in the knowledge base according to the input question;

[0033] The conversion encoder converts the input multiple background knowledge texts and the context of the question into internal features;

[0034] Update the text with the best background knowledge using the context of new questions;

[0035] Use the output module to output the best background knowledge text.

[0036] Furthermore, the extraction converter comprises:

[0037] The extractive transformer encoder encodes the best background knowledge text into features as output as the context.

[0038] Furthermore, the retrieval language model uses the RoBERTa-wwm-ext-large model as a benchmark model.

[0039] Furthermore, the feedback module trains and optimizes the generative transformer memory network module according to the final output answer and the user's feedback on the answer.

[0040] Furthermore, the feedback module trains and optimizes the generative transformer memory network module, specifically including:

[0041] Construct a reward model to evaluate the quality of the output of the transformer memory network module based on the user's feedback on the answer, as follows:

[0042] R(a)=f(a,θ)

[0043] Where R(a) is the reward score of answer a, θ is the parameter of the reward model, and f represents the function that evaluates the quality of the generated answer;

[0044] Train the reward model, generate different answers for the input questions, manually score and sort the answers based on their quality, input the scoring data into the reward model for training, and use the trained reward model to generate reward scores for the answers;

[0045] After the generative transformer memory network module generates an answer, the reward model is used to calculate the reward score of the answer to obtain the reward of the answer, and the output strategy of the generative transformer memory network module is adjusted by maximizing the reward.

[0046] Beneficial effects:

[0047] The system takes the pre-trained model as the core, and designs pre-training and fine-tuning based on two types of transformer memory architectures, extractive and generative. Compared with the common test baselines, the accuracy of the system of the present invention reached 85.6% on the A University dataset, which is better than the 70% accuracy of the traditional template-based system. The resource utilization rate was reduced by about 40%, and the error rate was reduced to less than 3%. The experimental results show that the pre-trained model proposed in this paper achieves better performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0049] Figure 1 It is a schematic diagram of the memory architecture of the generative transformer.

[0050] Figure 2 is a schematic diagram of the memory architecture of a decimation converter.

[0051] Figure 3 It is a schematic diagram of the overall framework of the present invention. DETAILED DESCRIPTION

[0052] The overall idea of ​​the present invention is as follows: Considering that the selection of the best knowledge fragment plays an important role in the model effect, the converter memory network architecture is applied to the Chinese training set and the retrieval system is optimized. The retrieval system is optimized by taking advantage of the BERT model, and the application of knowledge discarding increases the generation effect of the decoder. At the same time, the present invention finds that intentionally discarding some knowledge fragments in a certain proportion and not letting the decoder contact them can enhance the decoder's resistance to wrong knowledge fragment selection, so that the system can better understand the context, manage the dialogue state, and generate coherent and natural answers.

[0053] Based on the above ideas, the present invention builds a graduate student intelligent question-answering system based on a pre-trained model for the protection of privacy information and data security, while adding explanations of the decision-making basis and optimizing the algorithm and structure of the model. The details are as follows:

[0054] 1. Model selection and training

[0055] The pre-trained model is the core part of the question-answering system, responsible for understanding the user's input and determining how to respond. The pre-trained model consists of two stages: pre-training and fine-tuning. The pre-training stage is to train the model in an unsupervised or weakly supervised manner on a large-scale corpus, hoping that the model can acquire language-related knowledge, such as syntax and grammar knowledge. The fine-tuning stage is to use the pre-trained model to customize training for certain tasks, so that the pre-trained model "understands" the task better. Taking BERT as an example, the pre-training stage includes two tasks: MLM (Masked Language Model) and NSP (Next Sentence Prediction). The former is similar to "cloze", and the latter is to determine whether the two sentences are adjacent in the original text given two sentences. After BERT pre-training is completed, it can be connected to various types of downstream tasks, such as text classification, sequence labeling, reading comprehension, etc. By fine-tuning on these tasks, better experimental results can be obtained.

[0056] Currently, the pre-training models suitable for question-answering systems are mainly divided into two categories. One is the uni-directional pre-training model (reference: DEVLIN J, CHANG MW, LEE K, et al. Bert: pretraining of deep bidirectional transformers for language understanding [C] / / Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Stroudsburg, France: Association for Computational Linguistics, 2019: 4171-4186.), such as OpenAI's GPT series (GPT-1, GPT-2, GPT-3, etc.). This type of model can only see the previous context of the input sequence during training, but not the subsequent context; the other is the bi-directional pre-training model (reference: CUI Y, CHE W, LIU T, et al. Pre-training with whole word masking for chinese bert [J]. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 2021, 29:3504-3514.), such as Google's BERT, RoBERTa, etc., these models can see the previous and next context of the input sequence at the same time during training. These two types of models have their own strengths. For example, BERT is a bidirectional model and is not suitable for direct text generation, but it can be used to handle dialogue generation tasks by adding a decoder to BERT to form a sequence-to-sequence (Seq2Seq) model. Similarly, GPT is a unidirectional model, which is used for some understanding tasks through fine-tuning based on its ability to generate coherent text.

[0057] Based on the fact that the existing models are not satisfactory, and considering that in addition to communicating with people through fluent and vivid language, it is also necessary to understand the professional knowledge related to graduate admissions, the present invention selects the Transformer Memory Networks model architecture (reference: DINAN E, ROLLER S, SHUSTER K, et al. Wizard of Wikipedia: Knowledge-Powered Conversational Agents [C] / / The Seventh International Conference on Learning Representations. New Orleans, USA: ACM Press, 2019: 1-18.) as the pre-training benchmark architecture, and designs two Transformer Memory Network architectures, one of which is a generative transformer memory architecture (such as Figure 1 ), the other is the extraction converter memory architecture (such as Figure 2 ). The Extractive Transformer Memory model selects the most appropriate response from a series of candidate sentences from the training set, and the Generative model autonomously generates the response. The current conversation context, where the first sentence is the selected topic and the rest are the specific conversation content. For questions that the model has been exposed to, the output of the Extractive Transformer Memory model is more coherent, while for questions that the model has not been exposed to, the Generative Transformer Memory model is more coherent.

[0058] 1.1 Generative Transformer Memory Architecture

[0059] like Figure 1 As shown in Figure 1, the generative transformer memory architecture consists of two parts. One part is the memory network ( Figure 1 The key idea of ​​the memory network is to have an explicit memory component that can be read and written, which is separate from the main computational link of the model. This allows the model to store information for a long time, which is especially useful for tasks that require remembering information from a large amount of text to answer questions (such as question answering tasks). The memory network consists of four parts, such as Figure 1 As shown in the figure, the first part is the retrieval system, which retrieves the top N best matching background knowledge according to the user question. The second part is the dialogue context. The third part is the conversion encoder, which encodes the text provided by the dialogue context and the retrieved background knowledge respectively, and converts the text into an embedding vector. The fourth part is to perform a dot product operation on the embedding vector output by the conversion encoder.

[0060] Figure 1In addition to the memory network in the virtual box, another part is the transformer, including the transformer encoder and the transformer decoder, which are used to provide text expression and generation. The transformer uses a technology called "self-attention mechanism" or "transformer attention mechanism", which allows the model to pay attention to each part of the sequence when processing sequence data. The input and output texts are first converted into embedding vectors, which are used to convert words or characters (depending on the language unit processed by the model) into numerical vectors. Since the Transformer model itself does not contain any loop or convolution operations, it cannot process the order information of the sequence. To solve this problem, the Transformer introduces position encoding, adding a vector representing its position in the sequence to each word at each position. The decoder part uses a self-attention layer and a feedforward neural network to generate the output sequence. The Transformer architecture is easy to parallelize, which allows the model to be trained faster.

[0061] 1.2 Decimation Converter Memory Architecture

[0062] like Figure 2 As shown in the figure, the extractive transformer memory architecture is basically the same as the generative architecture. The difference is that the decoder at the last position in the generative architecture is changed into an encoder (the extractive transformer memory model selects the most suitable one as a reply from a series of alternative sentences from the training set, so it only needs to compile in the alternative sentences; the generative model generates replies autonomously and requires a decoder.), and the role of the encoder is to extract the best matching text from the backup text set (containing 100 reply sentence texts) as a reply.

[0063] 2. System overall framework

[0064] like Figure 3 As shown in the figure, after the two network architectures are tuned separately, the system uses the output of the extractive transformer memory network architecture model as context and inputs it to the generative transformer memory network architecture model together with the user question. That is, the extractive architecture is used to locate the answer fragment first, and then the generative one is used to optimize the expression. This design can improve accuracy. The extractive model can quickly locate relevant information from long texts and reduce the reasoning burden of the generative model. It can also optimize the expression. The generative model optimizes the answer expression according to the context and question, making the output more natural and smooth.

[0065] Embodiment 1:

[0066] like Figure 3As shown in the figure, it is an overall selection and construction method of the intelligent question-answering system for graduate student admissions, which aims to provide a large language model construction method specifically for the intelligent question-answering system for graduate student admissions. Based on user query requests and combined with relevant background knowledge retrieved from the knowledge base, the system first uses an extractive large language model to extract accurate text, and then optimizes the text through a generative large language model. During the operation of the system, a user feedback module is designed to collect user feedback on the generated results and train the reward model accordingly. The generative transformer memory network is fine-tuned through a reinforcement learning algorithm to make its output more in line with user expectations. This method not only significantly improves the intelligence level of the question-answering system, but also ensures the high accuracy, fluency and user satisfaction of the output text, thereby providing high-quality text generation services for the intelligent question-answering system for graduate student admissions. Specifically, it includes the following steps:

[0067] S1: Knowledge base construction

[0068] The construction of the knowledge base consists of two stages: the first is the segmentation and re-editing of the document content, and the second is the establishment of the vector database.

[0069] S11 Phase 1: Document content segmentation reorganization

[0070] This stage is to effectively split and reorganize the original document (for example, the admission brochures and supplementary instructions of various colleges and universities over the years can be collected as the original document) to ensure that each paragraph has a high degree of independence while avoiding unnecessary paragraph coupling. The structure of the reorganized document usually includes:

[0071] -Metadata: Provides information about the document itself, such as author, creation date, modification history, etc.

[0072] - Keywords: Extract the most important keywords in the document. These words help to understand the topic of the document.

[0073] - Cross-references: Create references to other parts within the document so that the model can navigate between different parts.

[0074] -Key functions: Mark out key functions or concepts in the document to help you quickly locate important content.

[0075] -Contents Overview: Provide a clear and focused overview of each paragraph or chapter

[0076] S12 Phase II: Vector Database Establishment

[0077] In this phase, we vectorize the document content and build an efficient vector database to match queries more accurately. This involves:

[0078] S121 Vectorization: Use natural language processing technology (reference: Cui, Y., et al. (2020). "Revisiting Pre-Trained Models for Chinese Natural Language Processing." https: / / arxiv.org / abs / 2004.13922) to convert each paragraph or fragment of the re-edited document into a vector. Each vector represents the semantic information of the paragraph.

[0079] S122 builds a vector database: the vectors of documents are stored in a professional vector database system, such as FAISS, Pinecone or Milvus. The database provides an efficient similarity search function, so that when users query, they can quickly find the document fragment that best matches the semantics of the query content.

[0080] S123 search optimization: Based on vector similarity calculations (such as cosine similarity), it efficiently matches queries with vectors stored in the database and returns the most relevant document fragments.

[0081] The selection and construction method of the S2 generative transformer memory network architecture model is as follows:

[0082] The generative transformer memory architecture consists of two parts. Figure 1As shown in the figure, one part is the memory network, which is used to retrieve knowledge. The key idea of ​​the memory network is to have an explicit memory component that can be read and written. This component is separate from the main computing link of the model, so that the model can store information for a long time, which is especially useful for tasks that need to remember information from a large amount of text to answer questions (such as question answering tasks). The memory network consists of four parts. The first part is the input feature map, which converts the input into internal features; the second part is generalization, which updates the memory with new inputs; the third part is the output feature map, which retrieves relevant memories and produces useful features for predicting answers for a query question; and finally generates output. The other part is the transformer, which is used to provide text expression and generation. The transformer uses a technique called "self-attention mechanism" or "transformer attention mechanism", which allows the model to pay attention to each part of the sequence when processing sequence data. The input and output texts are first converted into embedding vectors, which convert words or characters (depending on the language unit processed by the model) into numerical vectors. Since the Transformer model itself does not contain any loop or convolution operations, it cannot process the order information of the sequence. To solve this problem, Transformer introduces position encoding, adding a vector representing the position of the word in the sequence to each position. The decoder uses a self-attention layer and a feedforward neural network to generate the output sequence. The Transformer architecture is easy to parallelize, which makes the model faster to train.

[0083] The selection and construction methods of the S3 extractive transformer memory network architecture and the generative transformer memory network architecture model are basically the same, such as Figure 2 As shown in the figure, the difference is that the decoder at the last position in the generative architecture is changed into an encoder (the extractive transformer memory model selects the most suitable one as a reply from a series of alternative sentences from the training set, so it only needs to compile in the alternative sentences; the generative model generates replies autonomously and requires a decoder.), and the role of the encoder is to extract the best matching text from the backup text set (containing 100 reply sentence texts) as a reply.

[0084] After the two network architectures are tuned separately, the system uses the output of the extractive transformer memory network architecture model as context and inputs it to the generative transformer memory network architecture model together with the user question. That is, the extractive architecture is used to locate the answer fragment first, and then the generative one is used to optimize the expression. This design can improve accuracy. The extractive model can quickly locate relevant information from long texts and reduce the reasoning burden of the generative model. It can also optimize the expression. The generative model optimizes the answer expression according to the context and question, making the output more natural and smooth. Successfully decoupled tasks: complex tasks are decomposed into two parts: "information screening" and "language generation" for easy tuning.

[0085] The specific implementation steps are as follows:

[0086] 1 The user asks a question, the input format is as follows:

[0087] - Question Q: User questions

[0088] 2. Extractive architecture extracts answer fragments

[0089] Use the extractive approach to extract the most relevant answer fragment from the closest N standard responses:

[0090] 3. Generative models optimize answer expression

[0091] The output of the extractive model and the question are input into the generative architecture to reorganize the language and generate the final answer.

[0092] S4 Feedback Module

[0093] The feedback module uses user feedback to help the output of the generative transformer memory network architecture module be more in line with expectations. It is not suitable for the extractive transformer memory network architecture. Training is performed through user feedback on the model output. The following are the detailed steps:

[0094] S41 trains the reward model, which is used to evaluate the quality of the generative model output based on user feedback. The specific steps are as follows:

[0095] Collect comparative data: For a given input (such as a user's question), generate multiple different responses. Human annotators rank the responses based on their quality, with scores typically ranging from 1 to 5, representing a low to high quality assessment.

[0096] Reward scoring: Build a reward model based on comparison data. This model can score multiple answers generated by the input according to the user's rating. The goal of the reward model is to automatically evaluate the quality of the generated answers by learning these ratings.

[0097] formula:

[0098] R(a)=f(a,θ)

[0099] Where R(a) is the reward score of answer a, θ is the parameter of the reward model, and f represents the function that evaluates the quality of the generated answer.

[0100] S42 reinforcement learning optimized generative architecture

[0101] After training the reward model, the behavior of the generative architecture is further optimized through reinforcement learning. Reinforcement learning enables the generative model to adjust the output strategy based on long-term rewards to improve the model's performance in the task. After generating the answer, the reward model is used to evaluate the output quality and calculate the reward (reward signal). The reward signal indicates the quality of the generated answer, and the model adjusts the output strategy by maximizing the reward. Using Proximal Policy Optimization (PPO), the model's strategy is optimized during the training process so that it can generate higher quality answers in various scenarios.

[0102] Through reinforcement learning, the model can gradually adjust its output so that the answers generated under new inputs can get higher reward scores. As training progresses, the model gradually learns how to generate answers that meet human preferences.

[0103] S43 Reinforcement Learning Feedback Loop

[0104] Through reinforcement learning, the generative model can gradually adjust its output strategy to generate higher quality answers under new inputs. As training progresses, the model gradually learns how to generate answers that meet human preferences through continuous feedback and strategy adjustment.

[0105] Gradual Adjustment: After each answer is generated, the model evaluates the output quality through the reward model and optimizes its generation strategy based on the reward signal.

[0106] Continuous optimization: Over time, the model will continue to optimize itself based on accumulated user feedback, thereby continuously improving the quality of answers.

[0107] Through this feedback system, the generative network architecture can adjust itself according to user feedback. This includes training the reward model, optimizing the generative architecture using reinforcement learning, calculating the rewards and optimizing the strategy. This process enables the model to gradually improve the output quality in practical applications and better meet user needs and expectations.

[0108] Embodiment 2:

[0109] This system uses open domain dialogue technology to implement the intelligent question-answering system for graduate student recruitment, and uses the transformer memory network model for natural language processing. The system mainly includes the following parts: knowledge base, data collection and preprocessing, model training and fine-tuning, and dialogue system construction. Its main task is to establish connections between the relatively independent modules of the system and realize the path from user to model, including user interface design, access control, maximum throughput control, response time control, and timeout control. After the system is running, it will continuously optimize and update the system based on user feedback and system performance, by collecting user feedback, analyzing wrong answers, regularly updating the model, and continuously deploying.

[0110] The overall technical architecture is divided into three parts: data layer, solution layer, and presentation layer.

[0111] The data layer uses big data technology to integrate structured data, data warehouses, file type data, and other unstructured data related to graduate admissions.

[0112] The solution layer uses Haystack and pre-trained models. First, the data generated by the data layer is re-exported and organized in the form of text files. After reading the text files, a Document object that can be read by the natural language processing engine is generated. After certain processing of the files, the documents are stored.

[0113] The presentation layer is mainly used to realize the human-computer interaction of the intelligent question-answering robot. This system aims to realize the research of pre-training models. Therefore, only simple web page interaction is used in the presentation.

[0114] The knowledge base constructed by this system is a self-service network information base, which consists of two parts. One part is information related to graduate student enrollment, such as enrollment brochures, professional scores, and historical enrollment information. The present invention uses Scrapy to store crawled brochures in txt format. Scrapy is an application technology written and implemented in Python for crawling website data and extracting structured data. Scrapy can crawl data in formats such as web pages and PDFs, and can achieve multiple storage methods to obtain results. The other part is Wikipedia data. All Wikipedia project lists can be accessed through https: / / dumps.wikimedia.org / . After downloading the data, the data is cleaned, all HTML tags, reference links, etc. are removed, and only the text part of the article is retained. The knowledge base is an important source of model training data sets and required background knowledge. The knowledge base constructed by this system is built by DokuWiki, an open source software that is good at creating online documents or knowledge bases. All documents stored in pdf, docx, and txt formats are organized in DokuWiki according to "namespaces", and each namespace can contain pages and sub-namespaces. Namespaces are created at the level of year, document type, and content classification.

[0115] The retrieval language model uses RoBERTa-wwm-ext-large as the benchmark model.

[0116] Embodiment 3:

[0117] The present invention conducts an empirical analysis of the proposed system through a specific embodiment, as follows:

[0118] 1 Research samples and data sources and data preprocessing

[0119] The data sets used for pre-training language models mainly come from historical question-and-answer records of graduate student admissions at University A and admissions information on the school's official website, as well as publicly released open source conversation data sets, such as the Douban conversation data set. These data are first pre-processed, including steps such as interception, word segmentation, and noise addition. In order to enable the model to generate or extract the correct answer based on the question, the present invention uses specific domain data to fine-tune the pre-training model.

[0120] The fine-tuning dataset was collected manually and generated by simulating conversations. It is 13M in size and consists of more than 800 conversations. The average conversation length is 252, the longest length is no more than 4 times the average length, and the maximum is no more than 204 words. One question and one answer count as one round, with a maximum of no more than 10 rounds and a minimum of no less than 1 round.

[0121] 2 Experimental results

[0122] The quality of the question answering system is measured by Recall@K (recall rate), F 1 and PPL measurement.

[0123] Recall@K refers to the ratio of the number of relevant results retrieved from the topK results to the number of all relevant results in the database, which measures the recall rate of the retrieval system. The recall rate indicates the ratio of the actual number of positive samples in the predicted positive samples to the total number of positive samples.

[0124] F 1 It is a weighted average of precision and recall. 1 The calculation formula is:

[0125]

[0126] In the case of multi-class or multi-label, this is the F for each class with weights depending on the average parameter 1 Precision reflects the model's ability to distinguish negative samples. The higher the Precision, the stronger the model's ability to distinguish negative samples. Recall reflects the model's ability to recognize positive samples. The higher the Recall, the stronger the model's ability to recognize positive samples. 1 It is a combination of the two. 1 The higher it is, the more robust the model is.

[0127] PPL (Perplexity) is an indicator used to evaluate the quality of a language model. Intuitively, when a very standard, high-quality document that conforms to human natural language habits is given as a test set, the higher the probability that the model generates this text, the smaller the perplexity of the model is, and the better the model is. Its calculation formula is

[0128]

[0129] Where N is the length of the sentence, P(ω i ) is the probability of the i-th word. The first word is P(ω 1 |ω 0 ), and ω 0 It indicates the beginning of the sentence and is a placeholder. From formula (2), we can see that P(ω i |ω 0 ω 1 …ω i-1 ) is larger, the smaller the PPL is.

[0130] Through a large number of experimental tests, the extractive transformer memory model was compared with the common test baseline, and the results are shown in Table 1. The Recall@1 value with memory is significantly improved from 73 to 79.3 compared with the non-memory model. The same rule is also shown for the BoW (Bag-of-Words Memory Network) memory network (reference: PARTHASARATHIP, PINEAU J. Extending NeuralGenerative Conversational Model using External Knowledge Sources [C] / / Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Brussels, Belgium: Association for Computational Linguistics, 2018: 690-695.), where the Recall@1 value increases from 55.1 to 70.3. It can also be concluded that the pre-trained and tuned model performs best.

[0131] Table 1. Comparison of the decimation converter memory model with common test baselines

[0132]

[0133]

[0134] The generative transformer memory model is compared with the common test baseline, and the results are shown in Table 2. The following test premise is that the model does not need to predict the background knowledge sentence, but directly gives the correct knowledge background sentence. At this time, the generative transformer memory model has a significantly lower PPL value than the transformer network, from 40.9 to 23.8, indicating that providing knowledge background sentences can indeed increase the probability of correct prediction. And the auxiliary discarding strategy can indeed further reduce PPL and improve F1 value.

[0135] Table 2 Comparison of generative transformer memory model with common test baselines

[0136]

[0137] The above experimental results are mainly attributed to:

[0138] 1) The generative transformer memory architecture consists of a retrieval system, an encoder, a decoder, and a self-attention layer. The main function of the retrieval system is to retrieve the 7 most matching background knowledge fragments based on the conversation topic. The encoder concatenates these fragments with the conversation history records and encodes them. The self-attention layer selects the most matching background knowledge fragment and inputs it and the historical conversation to the decoder, which finally generates the next sentence of the conversation.

[0139] 2) The strategic advantage of the transformer memory model architecture is that an information retrieval layer is added before the decoder. After the retrieved content and conversation context are input into the encoder, the decoder will pay attention to the previous conversation content and retrieval content when generating vivid and coherent text. That is, the addition of the information retrieval layer allows the model to be exposed to the required professional background knowledge before generating a response. Test results show that the architecture that is exposed to professional knowledge is superior to the transformer architecture that is not exposed to professional knowledge.

[0140] The extractive transformer memory model selects the most appropriate response from a series of candidate sentences from the training set, and the generative model generates the response autonomously. The input of both models is the same, the current conversation context, where the first sentence is the selected topic and the rest is the specific conversation content. For questions that the model has been exposed to, the output of the extractive transformer memory model is more coherent, while for questions that the model has not been exposed to, the generative transformer memory model is more coherent.

[0141] For pre-trained models based on generative and extractive transformer memory architectures, the present invention evaluates model performance, computational efficiency, and system robustness:

[0142] The model performance mainly includes BLEU (Bilingual Evaluation Understudy) score and accuracy. In the verification stage of system development, the BLEU indicator was used to evaluate the similarity between the generated text and the manual answer. The average BLEU score of the fine-tuned model on the verification set was 42.3, which was significantly higher than the 28.7 of the un-fine-tuned model, indicating that the model has improved performance in generating relevant answers. The average BLEU score of the traditional template-based question-answering system is 28.5. This shows that the answers generated by the system are less similar to the manual standard answers, and the answers are mostly in a fixed format, which makes it difficult to cope with the diversity and complexity of the questions. In addition, the present invention evaluates the verification set of historical question-answering data, and the accuracy of the system's answers reaches 85.6%, which is significantly improved compared to the traditional template-based question-answering system (about 70% accuracy).

[0143] Computational efficiency mainly includes two aspects: inference speed and resource utilization. In the test of inference using NVIDIA A100 GPU, the average response time of the system was 1.2 seconds, which is significantly higher than the traditional average response time. In order to improve resource utilization and overall system throughput, this paper uses TensorRT optimization. Experimental results show that the model's video memory occupancy rate is reduced by about 40%, which enables more requests to be processed in parallel under the same hardware conditions.

[0144] The robustness of the system mainly includes two aspects: language diversity processing and error handling. Language diversity processing: The BLEU score of the system in different language environments (including English, Chinese, etc.) is stable at more than 40, demonstrating high robustness and adaptability under multi-language input. Regarding error handling, the present invention uses uncertainty estimation of the model, so that the system can detect and mark high-risk generation results and request manual review when necessary. This mechanism reduces the serious error rate of the system's answers to less than 3%.

[0145] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit, wherein the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, the invention content of the intelligent question-and-answer system for enrollment based on a pre-trained model provided by the present invention and some or all of the steps in each embodiment can be executed. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0146] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of computer programs and their corresponding general hardware platforms. Based on such an understanding, the technical solutions in the embodiments of the present invention can be essentially or partly contributed to the prior art in the form of computer programs, i.e., software products, which can be stored in a storage medium and include several instructions for enabling a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, an MCU or a network device, etc.) to execute the methods described in various embodiments of the present invention or certain parts of the embodiments.

[0147] The present invention provides a concept and method of an intelligent question-answering system for enrollment based on a pre-trained model. There are many methods and approaches to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention. All components not specified in this embodiment can be implemented using existing technologies.

Claims

1. An intelligent question-answering system for enrollment based on a pre-trained model, characterized in that: include: A knowledge base module, an extractive transformer memory network module, a generative transformer memory network module and a feedback module; wherein, The knowledge base module is used to build a knowledge base in the field of enrollment; The extractive transformer memory network module and the generative transformer memory network module are pre-trained respectively; The extractive transformer memory network module extracts the best answer from a given set of candidate answers based on historical question and answer records and retrieved background knowledge; The generative transformer memory network module obtains a final output answer according to the best answer output by the extractive transformer memory network module and the input question; The feedback module is used to adjust and optimize the generative converter memory network module.

2. According to claim 1, an intelligent question-answering system for enrollment based on a pre-training model is characterized in that: The generative transformer memory network module comprises: A first memory network and a generative transformer; wherein the first memory network is used to retrieve background knowledge; and the generative transformer generates text output with reference to the retrieved background knowledge.

3. According to claim 2, an intelligent question-answering system for enrollment based on a pre-training model is characterized in that: The first memory network comprises: Transformation encoder, retrieval language model and output module; where, The retrieval language model retrieves the best matching text in the knowledge base according to the input question; The conversion encoder converts the input multiple answer texts and the context of the question into internal features; Update the best answer text using the new question context; Use the output module to output the features corresponding to the best answer text.

4. According to claim 3, the intelligent question-answering system for enrollment based on a pre-training model is characterized in that: The generative converter comprises: The generative transformation decoder uses a self-attention layer and a feed-forward neural network to generate the answer text based on the output features of the first memory network.

5. According to claim 4, the intelligent question-answering system for enrollment based on a pre-training model is characterized in that: The extractive transformer memory network module converts the input multiple answer texts and the context of the question into internal features, including: A second memory network and an extractive transformer; wherein the second memory network is used to retrieve knowledge; and the extractive transformer selects the most appropriate feature as output based on the knowledge retrieved by the second memory network.

6. The enrollment intelligent question-answering system based on a pre-training model according to claim 5, characterized in that: The second memory network comprises: Transformation encoder, retrieval language model and output module; where, The retrieval language model retrieves multiple background knowledge texts in the knowledge base according to the input question; The conversion encoder converts the input multiple background knowledge texts and the context of the question into internal features; Update the text with the best background knowledge using the context of new questions; Use the output module to output the best background knowledge text.

7. The enrollment intelligent question-answering system based on a pre-training model according to claim 6, characterized in that: The extraction converter comprises: The extractive transformer encoder encodes the best background knowledge text into features as output as the context.

8. The enrollment intelligent question-answering system based on a pre-training model according to claim 7, characterized in that: The retrieval language model uses the RoBERTa-wwm-ext-large model as a benchmark model.

9. The enrollment intelligent question-answering system based on the pre-training model according to claim 8, characterized in that: The feedback module trains and optimizes the generative transformer memory network module according to the final output answer and the user's feedback on the answer.

10. The enrollment intelligent question-answering system based on the pre-training model according to claim 9, characterized in that: The feedback module trains and optimizes the generative transformer memory network module, specifically including: Construct a reward model to evaluate the quality of the output of the transformer memory network module based on the user's feedback on the answer, as follows: R(a)=f(a,θ) Where R(a) is the reward score of answer a, θ is the parameter of the reward model, and f represents the function that evaluates the quality of the generated answer; Train the reward model, generate different answers for the input questions, manually score and sort the answers based on their quality, input the scoring data into the reward model for training, and use the trained reward model to generate reward scores for the answers; After the generative transformer memory network module generates an answer, the reward model is used to calculate the reward score of the answer to obtain the reward of the answer, and the output strategy of the generative transformer memory network module is adjusted by maximizing the reward.