Market supervision public consultation method and system based on pre-trained big language model
Through the market supervision public consultation method based on pre-trained large language model, the problems of inefficiency of traditional consulting services and difficulty in mining demand are solved, and more efficient and accurate consulting services and user satisfaction are achieved.
Patent Information
- Application Number
- CN202510088729.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional market supervision public consulting services are inefficient and expensive, making it difficult to comprehensively and accurately grasp the public's true intentions and needs, and cannot deeply explore users' personalized needs.
The market supervision public consultation method based on pre-trained large language model is adopted. By obtaining text data related to market supervision, the benchmark model GPT-4 is injected to obtain the large language model of market supervision, and the model is further fine-tuned through the market supervision public consultation question and answer data set, the user consultation text is obtained and the answers are generated through the correlation model, and user feedback is continuously collected and the model is retrained.
It improves the efficiency and accuracy of market supervision consulting services, provides more professional and accurate answers, enhances user satisfaction and trust, and ensures long-term effectiveness and adaptability of the system.
Smart Images

Figure CN119988552A_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a market supervision public consultation method and system based on a pre-trained large language model, belonging to the technical field of natural language processing. Background Art
[0002] As an emerging artificial intelligence technology, pre-trained large language models have achieved remarkable results in the field of natural language processing. Pre-trained large language models can capture various basic features of human language through unsupervised learning on massive text data, thereby forming a deep understanding of human language. The advantage of this model is that it can handle the complexity and diversity of natural language and provide strong support for various natural language processing tasks.
[0003] As one of the core businesses of market supervision, reference consulting services have always been a "rigid need" for users, and the demand is becoming increasingly complex. Traditional market supervision public consulting services have failed to achieve the transformation from "passive" to "active" and from "static" to "dynamic", and there are problems such as low efficiency and high processing costs. On the one hand, traditional consulting methods usually require manual sorting and analysis of a large amount of public feedback, which is not only time-consuming and labor-intensive, but also prone to omissions or misunderstandings. On the other hand, due to the diversity and complexity of public feedback, traditional methods often find it difficult to fully and accurately grasp the public's true intentions and needs, and cannot deeply explore the personalized needs of users. Summary of the invention
[0004] The present invention provides a market supervision public consultation method and system based on a pre-trained large language model to solve the above-mentioned problems:
[0005] The present invention proposes a market supervision public consultation method based on a pre-trained large language model, the method comprising:
[0006] Acquire text data related to market supervision, use the large model GPT4 as a benchmark model, and obtain a large language model for market supervision by injecting the text data into the benchmark model;
[0007] Obtain a market supervision public consultation question and answer dataset, and fine-tune the market supervision large language model based on the market supervision public consultation question and answer dataset to obtain a public consultation question and answer model;
[0008] Obtaining a user consultation text, and obtaining a relevance answer through a relevance model according to the user consultation text and a public consultation question-and-answer model;
[0009] Obtain user feedback on answers and retrain the public consultation question-and-answer model.
[0010] Furthermore, text data related to market supervision is obtained, and the large model GPT4 is used as a benchmark model. A pre-trained large language model is obtained by injecting the text data into the benchmark model, including:
[0011] Collect text data related to market regulation, including policy documents, regulatory descriptions, case studies, and authoritative industry news;
[0012] Cleaning the text data, wherein the cleaning includes deleting redundant, irrelevant or erroneous content, standardizing the cleaned data, wherein the standardization includes unifying the date format and currency representation format and segmenting the text and removing stop words, and classifying and labeling the standardized text data, wherein the classification and labeling includes classifying the text data and labeling different types of text and key information keys contained therein, wherein the key information keys include the policy number and implementation date;
[0013] The processed text data is converted into a format suitable for input into the GPT-4 model, including converting the text into a vector form through a word embedding method and organizing and segmenting the text data converted into a vector form according to the format required for GPT-4 model training, wherein the text data in the vector form includes a training set, a validation set, and a test set.
[0014] Fine-tune the pre-trained GPT-4 model using the organized dataset;
[0015] After the model training is completed, its performance is evaluated and optimized. The accuracy of the model, the quality and relevance of the generated text are evaluated by its performance on the validation set and the test set. Based on the evaluation results, the model structure and training process are adjusted or fine-tuned again to obtain a large language model for market supervision.
[0016] Furthermore, a market supervision public consultation question and answer dataset is obtained, and the market supervision large language model is fine-tuned based on the market supervision public consultation question and answer dataset to obtain a public consultation question and answer model, including:
[0017] Obtaining a data set containing public consultation questions and answers related to market regulation, the data set includes historical consultation records, expert answers, and policy explanations, and cleaning the data set, including removing irrelevant information, correcting spelling errors, and dividing questions and answers. After cleaning, the text of the data set is segmented, stop words are removed, and part-of-speech tagging is performed;
[0018] Formatting the processed data, wherein the formatting includes formatting the question and the corresponding answer into a structure that can be understood by the model, that is, the format of "question: [question content] answer: [answer content] answer: [answer content] answer: [answer content] ...";
[0019] Augment the training data using data augmentation techniques, including synonym replacement, back translation, or modifying the question's wording, and split the augmented dataset into training, validation, and test sets.
[0020] Load the market regulation big language model, fine-tune the model by injecting the training set in the question-answering dataset into the market regulation big language, and adjust the configuration settings of the model, including the learning rate, batch size, and training period;
[0021] The performance of the fine-tuned model is evaluated on the test set. The performance of the model is evaluated by indicators such as accuracy, F1 score and BLEU(D) score. The model is further optimized based on the evaluation results. The further optimization includes adjusting the fine-tuning strategy, adding training data and modifying the model architecture to obtain a public consultation question and answer model.
[0022] Further, obtaining a user consultation text, and obtaining a relevance answer through a relevance model according to the user consultation text and the public consultation question-answering model, including:
[0023] Obtain the user's current consultation text and historical consultation texts, and input the current consultation text into the public consultation question-and-answer model to obtain multiple answer documents;
[0024] The topic model LDA is used to extract topics from historical consultation texts and answer documents, and the topic distribution of each document is generated, that is, the probability distribution of each document about each topic;
[0025] Calculate the relevance of the answer document through a relevance model according to the user consultation text and the public consultation question-answering model to obtain a relevance answer;
[0026] Specifically, the correlation model is:
[0027]
[0028] Among them, R(q,d) is the relevance, q is the user consultation text, d is the answer document in the public consultation question and answer model database, V(q) and V(d) are the vector representations of query q and answer document d in the database, respectively, |V(q)| and |V(d| are the vector moduli, Sim C (Q,D) is the similarity function between the user's previous consultation text and the answer document, U(d) is the emergency response factor,
[0029]
[0030] Among them, Q and D are the probability distributions of the user's previous query text and answer document, and M is the intermediate probability distribution.
[0031]
[0032] Among them, i is a different topic, n is the total number of topics,
[0033]
[0034] Furthermore, we obtain user feedback and retrain the public consultation question-answering model, including:
[0035] Receive feedback from the user after each question. If the user is satisfied with the answer, obtain the user's inquiry text and answer document. The inquiry text includes the current inquiry text and historical inquiry texts.
[0036] The consultation text and answer document are input as a data set into the public consultation question-answering model for feedback training.
[0037] The present invention proposes a market supervision public consultation system based on a pre-trained large language model, the system comprising:
[0038] A market supervision model training module is used to obtain text data related to market supervision, and the large model GPT4 is used as a benchmark model to obtain a market supervision large language model by injecting the text data into the benchmark model;
[0039] A public consultation question-answering model training module is used to obtain a market supervision public consultation question-answering dataset, and fine-tune the market supervision large language model based on the market supervision public consultation question-answering dataset to obtain a public consultation question-answering model;
[0040] A module for obtaining a related answer is used to obtain a user consultation text, and obtain a related answer through a relevance model according to the user consultation text and a public consultation question and answer model;
[0041] The feedback module is used to obtain user feedback on answers and retrain the public consultation question and answer model.
[0042] Furthermore, the training market supervision model module includes:
[0043] A text acquisition module for collecting text data related to market supervision, including policy documents, regulatory descriptions, case studies, and authoritative industry news;
[0044] A preprocessing module, configured to clean the text data, wherein the cleaning includes deleting redundant, irrelevant or erroneous content, standardizing the cleaned data, wherein the standardization includes unifying the date format and currency representation format and segmenting the text and removing stop words, and classifying and annotating the standardized text data, wherein the classification and annotation includes classifying the text data and annotating different types of text and key information keys contained therein, wherein the key information keys include policy numbers and implementation dates;
[0045] The data set segmentation module is used to convert the processed text data into a format suitable for input into the GPT-4 model, including converting the text into a vector form through a word embedding method and arranging and segmenting the text data converted into a vector form according to the format required for GPT-4 model training. The text data in vector form includes a training set, a validation set, and a test set.
[0046] The fine-tuning module is used to fine-tune the pre-trained GPT-4 model using the collated dataset;
[0047] The evaluation and optimization module is used to evaluate and optimize the model performance after model training is completed. The accuracy of the model, the quality and relevance of the generated text are evaluated through its performance on the validation set and the test set. Based on the evaluation results, the model structure and training process are adjusted or fine-tuned again to obtain a large language model for market supervision.
[0048] Furthermore, the training public consultation question-answering model module includes:
[0049] A module for preprocessing question and answer datasets, which is used to obtain a dataset containing public consultation questions and answers related to market supervision, including historical consultation records, expert answers, and policy explanations, and to clean the dataset, including removing irrelevant information, correcting spelling errors, and dividing questions and answers. After cleaning, the text of the dataset is segmented, stop words are removed, and part-of-speech tagging is performed;
[0050] A formatting module, used to format the processed data, wherein the formatting includes formatting the question and the corresponding answer into a structure that can be understood by the model, that is, the format of "question: [question content] answer: [answer content] answer: [answer content] answer: [answer content] ...";
[0051] A data augmentation module is used to augment the training data through data augmentation techniques, including synonym replacement, back translation or modifying the wording of the question, and to divide the augmented data set into a training set, a validation set and a test set;
[0052] A question-answer fine-tuning module is used to load the market supervision large language model, fine-tune the model by injecting the training set in the question-answering dataset into the market supervision large language, and adjust the configuration settings of the model, including the learning rate, batch size, and training cycle;
[0053] The adjustment and optimization module is used to evaluate the performance of the fine-tuned model on the test set. The performance of the model is evaluated through indicators such as accuracy, F1 score and BLEU (D) score. The model is further optimized according to the evaluation results. The further optimization includes adjusting the fine-tuning strategy, adding training data and modifying the model architecture to obtain a public consultation question and answer model.
[0054] Furthermore, the module for obtaining the associated answer includes:
[0055] The initial answer acquisition module is used to obtain the user's current consultation text and historical consultation texts, and input the current consultation text into the public consultation question-answering model to obtain multiple answer documents;
[0056] The topic distribution generation module is used to extract topics from historical consultation texts and answer documents through the topic model LDA, and generate the topic distribution of each document, that is, the probability distribution of each document about each topic;
[0057] A correlation calculation module is used to calculate the correlation of the answer document through a correlation model according to the user consultation text and the public consultation question-answering model to obtain a correlation answer;
[0058] Specifically, the correlation model is:
[0059]
[0060] Among them, R(q,d) is the relevance, q is the user consultation text, d is the answer document in the public consultation question and answer model database, V(q) and V(d) are the vector representations of query q and answer document d in the database, respectively, |V(q)| and |V(d| are the vector moduli, Sim C (Q, D) is the similarity function between the user's previous consultation text and the answer document, U(D) is the emergency response factor,
[0061]
[0062] Among them, Q and D are the probability distributions of the user's previous query text and answer document, and M is the intermediate probability distribution.
[0063]
[0064] Among them, i is a different topic, n is the total number of topics,
[0065]
[0066] Furthermore, the feedback module includes:
[0067] The feedback receiving module is used to receive feedback from the user after each question. If the user is satisfied with the answer, the user's inquiry text and answer document are obtained. The inquiry text includes the current inquiry text and the historical inquiry text.
[0068] The feedback training module is used to input the consultation text and answer document as a data set into the public consultation question-answering model for feedback training.
[0069] The beneficial effects of the present invention are as follows: through targeted pre-training and fine-tuning, the model can provide more professional and accurate answers in the field of market supervision; the fine-tuned model better understands queries specific to market supervision, provides highly relevant and specific answers, and increases user satisfaction; by continuously collecting user feedback and retraining the model, the system can adapt to new trends and problems and maintain its long-term effectiveness and accuracy; as the quality of the system's answers improves, the user experience will also improve accordingly, which can enhance user trust and reliance and have a positive impact on the image of market supervision agencies. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0071] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0072] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. The embodiments described are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0074] One embodiment of the present invention provides a market supervision public consultation method based on a pre-trained large language model, the method comprising:
[0075] Acquire text data related to market supervision, use the large model GPT4 as a benchmark model, and obtain a large language model for market supervision by injecting the text data into the benchmark model;
[0076] Obtain a market supervision public consultation question and answer dataset, and fine-tune the market supervision large language model based on the market supervision public consultation question and answer dataset to obtain a public consultation question and answer model;
[0077] Obtaining a user consultation text, and obtaining a relevance answer through a relevance model according to the user consultation text and a public consultation question-and-answer model;
[0078] Obtain user feedback on answers and retrain the public consultation question-and-answer model.
[0079] The working principle and effect of the above technical solution are as follows: First, by collecting text data related to market supervision, such as legal documents, policy guidelines, case analysis, etc., it is used to expand and adjust the pre-trained large model GPT-4. This process covers additional domain-specific training of the basic model to enable it to better understand the context and terminology of market supervision; using the market supervision public consultation question and answer dataset (a set of labeled question and answer pairs), the above-mentioned adjusted market supervision large language model is further fine-tuned. This step trains the model's ability to accurately generate or select appropriate answers in actual question and answer scenarios; in actual applications, when a user submits a consultation text, the public consultation question and answer model is used to analyze and locate the most appropriate answer through the relevance model. The relevance model performs intelligent matching based on the content and context of the user's question to extract the most relevant answer; user feedback is analyzed, and the public consultation question and answer model is retrained using this feedback data, which enables the model to continuously learn and adapt to changing user needs and market environment. Through targeted pre-training and fine-tuning, the model can provide more professional and precise answers in the field of market supervision; the fine-tuned model better understands queries specific to market supervision, provides highly relevant and specific answers, and increases user satisfaction; by continuously collecting user feedback and retraining the model, the system can adapt to new trends and issues and maintain its long-term effectiveness and accuracy; as the quality of the system's answers improves, the user experience will also improve accordingly, which can enhance user trust and reliance and have a positive impact on the image of market regulators.
[0080] A market supervision public consultation method based on a pre-trained large language model obtains text data related to market supervision, takes the large model GPT4 as a benchmark model, and obtains the pre-trained large language model by injecting the text data into the benchmark model, including:
[0081] Collect text data related to market regulation, including policy documents, regulatory descriptions, case studies, and authoritative industry news;
[0082] Cleaning the text data, wherein the cleaning includes deleting redundant, irrelevant or erroneous content, standardizing the cleaned data, wherein the standardization includes unifying the date format and currency representation format and segmenting the text and removing stop words, and classifying and labeling the standardized text data, wherein the classification and labeling includes classifying the text data and labeling different types of text and key information keys contained therein, wherein the key information keys include the policy number and implementation date;
[0083] The processed text data is converted into a format suitable for input into the GPT-4 model, including converting the text into a vector form through a word embedding method and organizing and segmenting the text data converted into a vector form according to the format required for GPT-4 model training, wherein the text data in the vector form includes a training set, a validation set, and a test set.
[0084] Fine-tune the pre-trained GPT-4 model using the organized dataset;
[0085] After the model training is completed, its performance is evaluated and optimized. The accuracy of the model, the quality and relevance of the generated text are evaluated by its performance on the validation set and the test set. Based on the evaluation results, the model structure and training process are adjusted or fine-tuned again to obtain a large language model for market supervision.
[0086] The working principle and effect of the above technical solution are as follows: First, collect text data related to market supervision, such as policy documents, regulatory descriptions, case studies and authoritative industry news. Then clean these data, remove redundant and irrelevant content, standardize the format such as date and currency representation, and perform text segmentation and remove stop words. Classify and annotate the cleaned and standardized data, identify key information keys in the text, such as policy numbers, implementation dates, etc., which helps the model understand and extract important information; convert the processed text data into vector form through word embedding methods, format and divide the data (such as training set, validation set, test set) according to the requirements of the pre-trained model GPT-4, and ensure that the data is suitable for the model training and evaluation process; use the sorted data set to fine-tune the pre-trained GPT-4 model to better adapt to the context and needs of market supervision. The model is continuously adjusted during the training process to optimize learning efficiency and output quality; after the model training is completed, its performance is evaluated on the validation set and test set to detect accuracy, quality and relevance of the generated text; according to the evaluation results, the model structure and training process are adjusted and fine-tuned again as necessary to obtain the best performance. Through precise data processing and model fine-tuning, the large language model will be able to more accurately understand and process professional texts and queries in the field of market supervision, and improve its adaptability and professionalism in this field; optimize the query response quality, and the fine-tuned model can generate more accurate and relevant answers, improving user satisfaction and dependence; the model will continue to improve its performance through continuous evaluation and optimization, ensuring that it develops in sync with market supervision policies and practices; enhance decision-making support capabilities, and the optimized model can provide stronger data support and analytical capabilities for market supervision decisions, assisting the policy formulation and implementation process.
[0087] A market supervision public consultation method based on a pre-trained large language model, obtaining a market supervision public consultation question and answer dataset, and fine-tuning the market supervision large language model based on the market supervision public consultation question and answer dataset to obtain a public consultation question and answer model, including:
[0088] Obtaining a data set containing public consultation questions and answers related to market regulation, the data set includes historical consultation records, expert answers, and policy explanations, and cleaning the data set, including removing irrelevant information, correcting spelling errors, and dividing questions and answers. After cleaning, the text of the data set is segmented, stop words are removed, and part-of-speech tagging is performed;
[0089] Formatting the processed data, wherein the formatting includes formatting the question and the corresponding answer into a structure that can be understood by the model, that is, the format of "question: [question content] answer: [answer content] answer: [answer content] answer: [answer content] ...";
[0090] Augment the training data using data augmentation techniques, including synonym replacement, back translation, or modifying the question's wording, and split the augmented dataset into training, validation, and test sets.
[0091] Load the market regulation big language model, fine-tune the model by injecting the training set in the question-answering dataset into the market regulation big language, and adjust the configuration settings of the model, including the learning rate, batch size, and training period;
[0092] The performance of the fine-tuned model is evaluated on the test set. The performance of the model is evaluated by indicators such as accuracy, F1 score and BLEU(D) score. The model is further optimized based on the evaluation results. The further optimization includes adjusting the fine-tuning strategy, adding training data and modifying the model architecture to obtain a public consultation question and answer model.
[0093] The working principle and effect of the above technical solution are as follows: First, collect a public consultation question and answer dataset containing historical consultation records, expert answers and policy explanations related to market supervision. The data is carefully cleaned, including removing irrelevant information and spelling errors, and accurately matching questions and corresponding answers. Then, the data is processed by word segmentation, stop word removal and part-of-speech tagging to prepare for subsequent model training. The processed data will be formatted into a structure that the model can understand, specifically in the format of "question: [question content] answer: [answer content]...". Use data enhancement techniques such as synonym replacement, back translation or modification of question expressions to expand and enrich the dataset and improve the diversity and coverage during model training. Load the pre-trained large language model GPT-4, and inject the processed question and answer training set data into the model for fine-tuning. During fine-tuning, adjust the model configuration settings, such as learning rate, batch size and training cycle, to ensure that the model adapts to the specific needs of market supervision. Evaluate the performance of the fine-tuned model on the validation set and test set, using performance indicators such as accuracy, F1 score and BLEU(D) score. Based on the evaluation results, the model is further optimized, including adjusting fine-tuning strategies, increasing training data, or improving the model architecture to obtain the final public consultation question-and-answer model. Improve the accuracy and quality of answers. Through data training and fine-tuning specifically for market supervision questions and answers, the model can more accurately understand and answer questions related to market supervision, and improve the relevance and professionalism of the answers; enhance the adaptability and flexibility of the model. Data enhancement and continuous optimization make the model more adaptable to changing query types and expressions, and enhance its ability to respond to new situations; as the quality of questions and answers improves, users are more likely to obtain high-quality and highly relevant answers, thereby improving user satisfaction and trust; through regular performance evaluation and model optimization, ensure that the AI model can continue to adapt to the development of the market supervision field and ensure its long-term effectiveness.
[0094] A market supervision public consultation method based on a pre-trained large language model obtains user consultation text, and obtains a relevance answer through a relevance model according to the user consultation text and a public consultation question-answering model, including:
[0095] Obtain the user's current consultation text and historical consultation texts, and input the current consultation text into the public consultation question-and-answer model to obtain multiple answer documents;
[0096] The topic model LDA is used to extract topics from historical consultation texts and answer documents, and the topic distribution of each document is generated, that is, the probability distribution of each document about each topic;
[0097] Calculate the relevance of the answer document through a relevance model according to the user consultation text and the public consultation question-answering model to obtain a relevance answer;
[0098] Specifically, the correlation model is:
[0099]
[0100] Among them, R(q,d) is the relevance, q is the user consultation text, d is the answer document in the public consultation question and answer model database, V(q) and V(d) are the vector representations of query q and answer document d in the database, respectively, |V(q)| and |V(d| are the vector moduli, Sim C (Q, D) is the similarity function between the user's previous consultation text and the answer document, U(D) is the emergency response factor,
[0101]
[0102] Among them, Q and D are the probability distributions of the user's previous query text and answer document, and M is the intermediate probability distribution.
[0103]
[0104] Among them, i is a different topic, n is the total number of topics,
[0105]
[0106] The working principle and effect of the above technical solution are as follows: after receiving the user's consultation text, the system analyzes the correlation between the text and the question-answer pairs in the knowledge base through the association model to determine the most appropriate answer. By evaluating the semantic similarity function, the similarity function of the user's previous query text and the answer document can provide the user with a personalized answer that meets the user's query requirements and is targeted at the user. In market supervision consultation, the user's personal color is strong, because the user's concern for market supervision policies is often related to personal consumer goods. Therefore, for each answer of the user, the user's historical query text should also be considered. The most relevant answer is provided to the user through a weighted combination of the importance of the current query text and the historical query text. In the application of the topic model, i represents different topics. For each topic, Q(i) is the probability of topic (i) appearing in the probability distribution (Q), and M(i) is the probability of topic (i) appearing in the probability distribution (M). The ∑ in the formula represents the accumulation of all possible topics (i), which means that the contribution of each topic (i) to the total answer relevance is calculated separately, and then these contribution values are added together to get the total relevance between the two distributions. For each corresponding element (i) in the distribution (P) and (Q), the value of (M(i)) is the average of (P(i)) and (Q(i)). This intermediate distribution (M) plays a role in smoothing and neutralizing the difference between the two original distributions during calculation. In the expression N(Q|M), (i) represents a discrete variable or index, which is used to index each element in the probability distribution (Q) and (M). In the context of topic models, this index (i) usually refers to different topics. For example, in the LDA model, a document consists of multiple topics, and each topic consists of multiple words. The topic distribution (Q) and (M) of each document generated by LDA contains the probability that the document belongs to each topic, which is estimated based on the word distribution in the document. When calculating N(Q|M), (Q(i)) and (M(i)) respectively represent the probability of the (i)th topic appearing in two different documents (or two different topic combination models of the same document). The entire summation operation traverses all possible topics (i) and calculates the contribution of each topic to the total relevance. Suppose there are two documents A and B: the topic distribution (Q) of document A may be ([0.2, 0.3, 0.5]), indicating that 20% of document A is topic 1, 30% is topic 2, and 50% is topic 3. The topic distribution (M) of document B may be ([0.1, 0.4, 0.5]), which means that 10% of document B is topic 1, 40% is topic 2, and 50% is topic 3. When calculating N(Q|M), for each topic (i) (here (i) is 1, 2, 3), we will calculate Then add these three values together to get the total topic contribution. User consultation should also consider the popularity of related answer documents and popular searches. U(d) is the emergency response factor, which is the ratio of the number of clicks on this answer document to the maximum number of clicks on all answer documents. It can dynamize user public consultation.
[0107] A market supervision public consultation method based on a pre-trained large language model obtains user feedback on answers and retrains the public consultation question-answering model, including:
[0108] Receive feedback from the user after each question. If the user is satisfied with the answer, obtain the user's inquiry text and answer document. The inquiry text includes the current inquiry text and historical inquiry texts.
[0109] The consultation text and answer document are input as a data set into the public consultation question-answering model for feedback training.
[0110] The working principle and effect of the above technical solution are as follows: after providing the answer, the system will ask the user about his / her satisfaction with the answer. If the user is satisfied, the system will automatically collect the consultation text and answer document in this interaction. The collected consultation text and the corresponding answer document constitute a data pair, which will be organized into a format suitable for input into the public consultation question and answer model. The data set may also undergo preprocessing operations, such as text cleaning (removing meaningless characters or stop words, etc.), word segmentation and vectorization, to ensure data quality and model input standards. The collected and processed data pairs are used to conduct feedback training on the existing public consultation question and answer model. This training is an incremental learning process that aims to use new training samples to adjust the parameters of the model so that the model can better adapt to the actual needs and preferences of users. Model performance improvement: By continuously receiving user feedback and using it to perform incremental training on the model, the model can continuously learn and adapt to the latest user needs and language usage habits, thereby continuously improving its accuracy and user satisfaction; more personalized responses, because the text of the user's historical consultation is included, the model can better understand and predict the query intentions and preferences of specific users, and provide more personalized and relevant answers; feedback training enables the model to be dynamically adjusted and optimized, not limited to offline training at a fixed time point. This dynamism is the key to improving the long-term operating quality of interactive systems such as public consultation services; user participation and satisfaction are improved, allowing users to directly influence the learning and progress of the question-answering system, which can increase users' trust and commitment to the system, thereby improving overall user satisfaction and loyalty.
[0111] A market supervision public consultation system based on a pre-trained large language model, the system comprising:
[0112] A market supervision model training module is used to obtain text data related to market supervision, and the large model GPT4 is used as a benchmark model to obtain a market supervision large language model by injecting the text data into the benchmark model;
[0113] A public consultation question-answering model training module is used to obtain a market supervision public consultation question-answering dataset, and fine-tune the market supervision large language model based on the market supervision public consultation question-answering dataset to obtain a public consultation question-answering model;
[0114] A module for obtaining a related answer is used to obtain a user consultation text, and obtain a related answer through a relevance model according to the user consultation text and a public consultation question and answer model;
[0115] The feedback module is used to obtain user feedback on answers and retrain the public consultation question and answer model.
[0116] The working principle and effect of the above technical solution are as follows: First, by collecting text data related to market supervision, such as legal documents, policy guidelines, case analysis, etc., it is used to expand and adjust the pre-trained large model GPT-4. This process covers additional domain-specific training of the basic model to enable it to better understand the context and terminology of market supervision; using the market supervision public consultation question and answer dataset (a set of labeled question and answer pairs), the above-adjusted market supervision large language model is further fine-tuned. This step exercises the model's ability to accurately generate or select appropriate answers in actual question and answer scenarios;
[0117] In actual applications, when a user submits a consultation text, the public consultation question-and-answer model is used to analyze and locate the most appropriate answer through the relevance model. The relevance model performs intelligent matching based on the content and context of the user's question and extracts the most relevant answer; user feedback is analyzed, and the public consultation question-and-answer model is retrained using this feedback data, which enables the model to continuously learn and adapt to changing user needs and market environments. Through targeted pre-training and fine-tuning, the model can provide more professional and accurate answers in the field of market supervision; the fine-tuned model better understands queries specific to market supervision, provides highly relevant and specific answers, and increases user satisfaction; by continuously collecting user feedback and retraining the model, the system can adapt to new trends and issues and maintain its long-term effectiveness and accuracy; as the quality of the system's answers improves, the user experience will also improve accordingly, which can enhance user trust and reliance and have a positive impact on the image of market regulators.
[0118] A market supervision public consultation system based on a pre-trained large language model, wherein the training market supervision model module comprises:
[0119] A text acquisition module for collecting text data related to market supervision, including policy documents, regulatory descriptions, case studies, and authoritative industry news;
[0120] A preprocessing module, configured to clean the text data, wherein the cleaning includes deleting redundant, irrelevant or erroneous content, standardizing the cleaned data, wherein the standardization includes unifying the date format and currency representation format and segmenting the text and removing stop words, and classifying and annotating the standardized text data, wherein the classification and annotation includes classifying the text data and annotating different types of text and key information keys contained therein, wherein the key information keys include policy numbers and implementation dates;
[0121] The data set segmentation module is used to convert the processed text data into a format suitable for input into the GPT-4 model, including converting the text into a vector form through a word embedding method and arranging and segmenting the text data converted into a vector form according to the format required for GPT-4 model training. The text data in vector form includes a training set, a validation set, and a test set.
[0122] The fine-tuning module is used to fine-tune the pre-trained GPT-4 model using the collated dataset;
[0123] The evaluation and optimization module is used to evaluate and optimize the model performance after model training is completed. The accuracy of the model, the quality and relevance of the generated text are evaluated through its performance on the validation set and the test set. Based on the evaluation results, the model structure and training process are adjusted or fine-tuned again to obtain a large language model for market supervision.
[0124] The working principle and effect of the above technical solution are as follows: First, collect text data related to market supervision, such as policy documents, regulatory descriptions, case studies and authoritative industry news. Then clean these data, remove redundant and irrelevant content, standardize the format such as date and currency representation, and perform text segmentation and remove stop words. Classify and annotate the cleaned and standardized data, identify key information keys in the text, such as policy numbers, implementation dates, etc., which helps the model understand and extract important information; convert the processed text data into vector form through word embedding methods, format and divide the data (such as training set, validation set, test set) according to the requirements of the pre-trained model GPT-4, and ensure that the data is suitable for the model training and evaluation process; use the sorted data set to fine-tune the pre-trained GPT-4 model to better adapt to the context and needs of market supervision. The model is continuously adjusted during the training process to optimize learning efficiency and output quality; after the model training is completed, its performance is evaluated on the validation set and test set to detect accuracy, quality and relevance of the generated text; according to the evaluation results, the model structure and training process are adjusted and fine-tuned again as necessary to obtain the best performance. Through precise data processing and model fine-tuning, the large language model will be able to more accurately understand and process professional texts and queries in the field of market supervision, and improve its adaptability and professionalism in this field; optimize the query response quality, and the fine-tuned model can generate more accurate and relevant answers, improving user satisfaction and dependence; the model will continue to improve its performance through continuous evaluation and optimization, ensuring that it develops in sync with market supervision policies and practices; enhance decision-making support capabilities, and the optimized model can provide stronger data support and analytical capabilities for market supervision decisions, assisting the policy formulation and implementation process.
[0125] A market supervision public consultation system based on a pre-trained large language model, wherein the training public consultation question-answering model module comprises:
[0126] A module for preprocessing question and answer datasets, which is used to obtain a dataset containing public consultation questions and answers related to market supervision, including historical consultation records, expert answers, and policy explanations, and to clean the dataset, including removing irrelevant information, correcting spelling errors, and dividing questions and answers. After cleaning, the text of the dataset is segmented, stop words are removed, and part-of-speech tagging is performed;
[0127] A formatting module, used to format the processed data, wherein the formatting includes formatting the question and the corresponding answer into a structure that can be understood by the model, that is, the format of "question: [question content] answer: [answer content] answer: [answer content] answer: [answer content] ...";
[0128] A data augmentation module is used to augment the training data through data augmentation techniques, including synonym replacement, back translation or modifying the wording of the question, and to divide the augmented data set into a training set, a validation set and a test set;
[0129] A question-answer fine-tuning module is used to load the market supervision large language model, fine-tune the model by injecting the training set in the question-answering dataset into the market supervision large language, and adjust the configuration settings of the model, including the learning rate, batch size, and training cycle;
[0130] The adjustment and optimization module is used to evaluate the performance of the fine-tuned model on the test set. The performance of the model is evaluated through indicators such as accuracy, F1 score and BLEU (D) score. The model is further optimized according to the evaluation results. The further optimization includes adjusting the fine-tuning strategy, adding training data and modifying the model architecture to obtain a public consultation question and answer model.
[0131] The working principle and effect of the above technical solution are as follows: First, collect a public consultation question and answer dataset containing historical consultation records, expert answers and policy explanations related to market supervision. These data are carefully cleaned, including removing irrelevant information and spelling errors, and accurately matching questions and corresponding answers. Then, the data is processed by word segmentation, stop word removal and part-of-speech tagging to prepare for subsequent model training. The processed data will be formatted into a structure that the model can understand, specifically in the format of "question: [question content] answer: [answer content]...". Use data enhancement techniques, such as synonym replacement, back translation or modification of question expression, to expand and enrich the dataset and improve the diversity and coverage during model training. Load the pre-trained large language model GPT-4, and inject the processed question and answer training set data into the model for fine-tuning. During fine-tuning, adjust the model configuration settings, such as learning rate, batch size and training cycle, to ensure that the model adapts to the specific needs of market supervision.
[0132] Evaluate the performance of the fine-tuned model on the validation set and the test set, using performance indicators such as accuracy, F1 score, and BLEU(D) score. Based on the evaluation results, further optimize the model, including adjusting the fine-tuning strategy, adding training data, or improving the model architecture to obtain the final public consultation question-and-answer model. Improve the accuracy and quality of answers. Through data training and fine-tuning specifically for market supervision questions and answers, the model can more accurately understand and answer questions related to market supervision, and improve the relevance and professionalism of the answers; enhance the adaptability and flexibility of the model. Data enhancement and continuous optimization make the model more adaptable to changing query types and expressions, and enhance its ability to respond to new situations; as the quality of questions and answers improves, users are more likely to obtain high-quality, highly relevant answers, thereby improving user satisfaction and trust; through regular performance evaluation and model optimization, ensure that the AI model can continue to adapt to the development of the market supervision field and ensure its long-term effectiveness.
[0133] A market supervision public consultation system based on a pre-trained large language model, wherein the module for obtaining associated answers comprises:
[0134] The initial answer acquisition module is used to obtain the user's current consultation text and historical consultation texts, and input the current consultation text into the public consultation question-answering model to obtain multiple answer documents;
[0135] The topic distribution generation module is used to extract topics from historical consultation texts and answer documents through the topic model LDA, and generate the topic distribution of each document, that is, the probability distribution of each document about each topic;
[0136] A correlation calculation module is used to calculate the correlation of the answer document through a correlation model according to the user consultation text and the public consultation question-answering model to obtain a correlation answer;
[0137] Specifically, the correlation model is:
[0138]
[0139] Among them, R(q,d) is the relevance, q is the user consultation text, d is the answer document in the public consultation question and answer model database, V(q) and V(d) are the vector representations of query q and answer document d in the database, respectively, |V(q)| and |V(d| are the vector moduli, Sim C (Q,D) is the similarity function between the user's previous consultation text and the answer document, U(d) is the emergency response factor,
[0140]
[0141] Among them, Q and D are the probability distributions of the user's previous query text and answer document, and M is the intermediate probability distribution.
[0142]
[0143] Among them, i is a different topic, n is the total number of topics,
[0144]
[0145] The working principle and effect of the above technical solution are as follows: after receiving the user's consultation text, the system analyzes the correlation between the text and the question-answer pairs in the knowledge base through the association model to determine the most appropriate answer. By evaluating the semantic similarity function, the similarity function of the user's previous query text and the answer document can provide the user with a personalized answer that meets the user's query requirements and is targeted at the user. In market supervision consultation, the user's personal color is strong, because the user's concern for market supervision policies is often related to personal consumer goods. Therefore, for each answer of the user, the user's historical query text should also be considered. The most relevant answer is provided to the user through a weighted combination of the importance of the current query text and the historical query text. In the application of the topic model, i represents different topics. For each topic, Q(i) is the probability of topic (i) appearing in the probability distribution (Q), and M(i) is the probability of topic (i) appearing in the probability distribution (M). The Σ in the formula represents the accumulation of all possible topics (i), which means that the contribution of each topic (i) to the total answer relevance is calculated separately, and then these contributions are added together to get the total relevance between the two distributions. For each corresponding element (i) in the distribution (P) and (Q), the value of (M(i)) is the average of (P(i)) and (Q(i)). This intermediate distribution (M) plays a role in smoothing and neutralizing the difference between the two original distributions during calculation. In the expression N(Q|M), (i) represents a discrete variable or index, which is used to index each element in the probability distribution (Q) and (M). In the context of topic models, this index (i) usually refers to different topics. For example, in the LDA model, a document consists of multiple topics, and each topic consists of multiple words. The topic distribution (Q) and (M) of each document generated by LDA contains the probability that the document belongs to each topic, which is estimated based on the word distribution in the document. When calculating N(Q|M), (Q(i)) and (M(i)) respectively represent the probability of the (i)th topic appearing in two different documents (or two different topic combination models of the same document). The entire summation operation traverses all possible topics (i) and calculates the contribution of each topic to the total relevance. Suppose there are two documents A and B: the topic distribution (Q) of document A may be ([0.2, 0.3, 0.5]), indicating that 20% of document A is topic 1, 30% is topic 2, and 50% is topic 3.The topic distribution (M) of document B may be ([0.1, 0.4, 0.5]), which means that 10% of document B is topic 1, 40% is topic 2, and 50% is topic 3. When calculating N(Q|M), for each topic (i) (where (i) is 1, 2, 3), Q(i)log(Q(i)) / (M(i)) is calculated, and then these three values are added to get the total topic contribution. User consultation should also consider the popularity of related answer documents and popular searches. U(D) is the emergency response factor, which is the ratio of the number of clicks on this answer document to the maximum number of clicks on all answer documents. It can make users' public consultation dynamic.
[0146] A market supervision public consultation system based on a pre-trained large language model, wherein the feedback module comprises:
[0147] The feedback receiving module is used to receive feedback from the user after each question. If the user is satisfied with the answer, the user's inquiry text and answer document are obtained. The inquiry text includes the current inquiry text and the historical inquiry text.
[0148] The feedback training module is used to input the consultation text and answer document as a data set into the public consultation question-answering model for feedback training.
[0149] The working principle and effect of the above technical solution are as follows: After providing the answer, the system will ask the user about his / her satisfaction with the answer. If the user is satisfied, the system will automatically collect the consultation text and answer document in this interaction. The collected consultation text and the corresponding answer document constitute a data pair. These data pairs will be organized into a format suitable for input into the public consultation question and answer model. The data set may also undergo preprocessing operations, such as text cleaning (removing meaningless characters or stop words, etc.), word segmentation and vectorization, to ensure data quality and model input standards. The collected and processed data pairs are used to conduct feedback training on the existing public consultation question and answer model. This training is an incremental learning process that aims to use new training samples to adjust the parameters of the model so that the model can better adapt to the actual needs and preferences of users. Model performance improvement: By continuously receiving user feedback and using it to perform incremental training on the model, the model can continuously learn and adapt to the latest user needs and language usage habits, thereby continuously improving its accuracy and user satisfaction; more personalized responses, because the text of the user's historical consultation is included, the model can better understand and predict the query intentions and preferences of specific users, and provide more personalized and relevant answers; feedback training enables the model to be dynamically adjusted and optimized, not limited to offline training at a fixed time point. This dynamism is the key to improving the long-term operating quality of interactive systems such as public consultation services; user participation and satisfaction are improved, allowing users to directly influence the learning and progress of the question-answering system, which can increase users' trust and commitment to the system, thereby improving overall user satisfaction and loyalty.
[0150] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A market supervision public consultation method based on a pre-trained large language model, characterized in that: The method comprises: Acquire text data related to market supervision, use the large model GPT4 as a benchmark model, and obtain a large language model for market supervision by injecting the text data into the benchmark model; Obtain a market supervision public consultation question and answer dataset, and fine-tune the market supervision large language model based on the market supervision public consultation question and answer dataset to obtain a public consultation question and answer model; Obtaining a user consultation text, and obtaining a relevance answer through a relevance model according to the user consultation text and a public consultation question-and-answer model; Obtain user feedback on answers and retrain the public consultation question-and-answer model.
2. According to claim 1, a market supervision public consultation method based on a pre-trained large language model is characterized in that: Obtain text data related to market supervision, use the large model GPT4 as a benchmark model, and obtain a pre-trained large language model by injecting the text data into the benchmark model, including: Collect text data related to market regulation, including policy documents, regulatory descriptions, case studies, and authoritative industry news; Cleaning the text data, wherein the cleaning includes deleting redundant, irrelevant or erroneous content, standardizing the cleaned data, wherein the standardization includes unifying the date format and currency representation format and segmenting the text and removing stop words, and classifying and labeling the standardized text data, wherein the classification and labeling includes classifying the text data and labeling different types of text and key information keys contained therein, wherein the key information keys include the policy number and implementation date; Converting the processed text data into a format suitable for input into the GPT-4 model, including converting the text into a vector form by a word embedding method and arranging and segmenting the text data converted into a vector form according to the format required for GPT-4 model training, wherein the text data in the vector form includes a training set, a validation set, and a test set; Fine-tune the pre-trained GPT-4 model using the organized dataset; After the model training is completed, its performance is evaluated and optimized. The accuracy of the model, the quality and relevance of the generated text are evaluated by its performance on the validation set and the test set. Based on the evaluation results, the model structure and training process are adjusted or fine-tuned again to obtain a large language model for market supervision.
3. According to claim 1, a market supervision public consultation method based on a pre-trained large language model is characterized in that: Obtaining a market supervision public consultation question and answer dataset, and fine-tuning the market supervision large language model based on the market supervision public consultation question and answer dataset to obtain a public consultation question and answer model, including: Obtaining a data set containing public consultation questions and answers related to market regulation, the data set includes historical consultation records, expert answers, and policy explanations, and cleaning the data set, including removing irrelevant information, correcting spelling errors, and dividing questions and answers. After cleaning, the text of the data set is segmented, stop words are removed, and part-of-speech tagging is performed; Formatting the processed data includes formatting the question and the corresponding answer into a structure that the model can understand, i.e., the format of "question: [question content] answer: [answer content] answer: [answer content] answer: [answer content]..."; Augment the training data using data augmentation techniques, including synonym replacement, back translation, or modifying the question's wording, and split the augmented dataset into training, validation, and test sets. Load the market regulation big language model, fine-tune the model by injecting the training set in the question-answering dataset into the market regulation big language, and adjust the configuration settings of the model, including the learning rate, batch size, and training period; The performance of the fine-tuned model is evaluated on the test set. The performance of the model is evaluated by indicators such as accuracy, F1 score and BLEU(D) score. The model is further optimized based on the evaluation results. The further optimization includes adjusting the fine-tuning strategy, adding training data and modifying the model architecture to obtain a public consultation question and answer model.
4. According to claim 1, a market supervision public consultation method based on a pre-trained large language model is characterized in that: Obtaining a user consultation text, and obtaining a relevance answer through a relevance model according to the user consultation text and a public consultation question-and-answer model, including: Obtain the user's current consultation text and historical consultation texts, and input the current consultation text into the public consultation question-and-answer model to obtain multiple answer documents; The topic model LDA is used to extract topics from historical consultation texts and answer documents, and the topic distribution of each document is generated, that is, the probability distribution of each document about each topic; Calculate the relevance of the answer document through a relevance model according to the user consultation text and the public consultation question-answering model to obtain a relevance answer; Specifically, the correlation model is: Among them, R(q,d) is the relevance, q is the user consultation text, d is the answer document in the public consultation question and answer model database, V(q) and V(d) are the vector representations of query q and answer document d in the database, respectively, |V(q)| and |V(d| are the vector moduli, Sim C (Q,D) is the similarity function between the user's previous consultation text and the answer document, U(d) is the emergency response factor, Among them, Q and D are the probability distributions of the user's previous query text and answer document, and M is the intermediate probability distribution. Among them, i is a different topic, n is the total number of topics, 5. According to claim 1, a market supervision public consultation method based on a pre-trained large language model is characterized in that: Obtain user feedback on answers and retrain the public consultation question-answering model, including: Receive feedback from the user after each question. If the user is satisfied with the answer, obtain the user's inquiry text and answer document. The inquiry text includes the current inquiry text and historical inquiry texts. The consultation text and answer document are input as a data set into the public consultation question-answering model for feedback training.
6. A market supervision public consultation system based on a pre-trained large language model, characterized in that: The system comprises: A market supervision model training module is used to obtain text data related to market supervision, and the large model GPT4 is used as a benchmark model to obtain a market supervision large language model by injecting the text data into the benchmark model; A public consultation question-answering model training module is used to obtain a market supervision public consultation question-answering dataset, and fine-tune the market supervision large language model based on the market supervision public consultation question-answering dataset to obtain a public consultation question-answering model; A module for obtaining a related answer is used to obtain a user consultation text, and obtain a related answer through a relevance model according to the user consultation text and a public consultation question and answer model; The feedback module is used to obtain user feedback on answers and retrain the public consultation question and answer model.
7. According to claim 6, a market supervision public consultation system based on a pre-trained large language model is characterized in that: The training market supervision model module includes: A text acquisition module is used to collect text data related to market supervision, including policy documents, regulatory descriptions, case studies, and authoritative industry news; A preprocessing module, configured to clean the text data, wherein the cleaning includes deleting redundant, irrelevant or format-error contents, standardizing the cleaned data, wherein the standardization includes unifying the date format and currency representation format and segmenting the text and removing stop words, and classifying and annotating the standardized text data, wherein the classification and annotating includes classifying the text data and annotating different types of text and key information keys contained therein, wherein the key information keys include policy numbers and implementation dates; A data set segmentation module is used to convert the processed text data into a format suitable for input into the GPT-4 model, including converting the text into a vector form by a word embedding method and arranging and segmenting the text data converted into a vector form according to the format required for GPT-4 model training. The text data in the vector form includes a training set, a validation set, and a test set. The fine-tuning module is used to fine-tune the pre-trained GPT-4 model using the collated dataset; The evaluation and optimization module is used to evaluate and optimize the model performance after model training is completed. The accuracy of the model, the quality and relevance of the generated text are evaluated through its performance on the validation set and the test set. Based on the evaluation results, the model structure and training process are adjusted or fine-tuned again to obtain a large language model for market supervision.
8. According to claim 6, a market supervision public consultation system based on a pre-trained large language model is characterized in that: The training public consultation question-answering model module includes: A module for preprocessing question and answer datasets, which is used to obtain a dataset containing public consultation questions and answers related to market supervision, including historical consultation records, expert answers, and policy explanations, and to clean the dataset, including removing irrelevant information, correcting spelling errors, and dividing questions and answers. After cleaning, the text of the dataset is segmented, stop words are removed, and part-of-speech tagging is performed; A formatting module, used to format the processed data, wherein the formatting includes formatting the question and the corresponding answer into a structure that the model can understand, i.e., the format of "question: [question content] answer: [answer content] answer: [answer content] answer: [answer content]..."; A data augmentation module is used to augment the training data through data augmentation techniques, including synonym replacement, back translation or modifying the wording of the question, and to divide the augmented data set into a training set, a validation set and a test set; A question-answer fine-tuning module is used to load the market supervision large language model, fine-tune the model by injecting the training set in the question-answering dataset into the market supervision large language, and adjust the configuration settings of the model, including the learning rate, batch size, and training cycle; The adjustment and optimization module is used to evaluate the performance of the fine-tuned model on the test set. The performance of the model is evaluated through indicators such as accuracy, F1 score and BLEU (D) score. The model is further optimized according to the evaluation results. The further optimization includes adjusting the fine-tuning strategy, adding training data and modifying the model architecture to obtain a public consultation question and answer model.
9. According to claim 6, a market supervision public consultation system based on a pre-trained large language model is characterized in that: The module for obtaining the associated answer comprises: The initial answer acquisition module is used to obtain the user's current consultation text and historical consultation texts, and input the current consultation text into the public consultation question-answering model to obtain multiple answer documents; The topic distribution generation module is used to extract topics from historical consultation texts and answer documents through the topic model LDA, and generate the topic distribution of each document, that is, the probability distribution of each document about each topic; A correlation calculation module is used to calculate the correlation of the answer document through a correlation model according to the user consultation text and the public consultation question-answering model to obtain a correlation answer; Specifically, the correlation model is: Among them, R(q,d) is the relevance, q is the user consultation text, d is the answer document in the public consultation question and answer model database, V(q) and V(d) are the vector representations of query q and answer document d in the database, respectively, |V(q)| and |V(d| are the vector moduli, Sim C (Q,D) is the similarity function between the user's previous consultation text and the answer document, U(d) is the emergency response factor, Among them, Q and D are the probability distributions of the user's previous query text and answer document, and M is the intermediate probability distribution. Among them, i is a different topic, n is the total number of topics, 10. According to claim 6, a market supervision public consultation system based on a pre-trained large language model is characterized in that: The feedback module comprises: The feedback receiving module is used to receive feedback from the user after each question. If the user is satisfied with the answer, the user's inquiry text and answer document are obtained. The inquiry text includes the current inquiry text and the historical inquiry text. The feedback training module is used to input the consultation text and answer document as a data set into the public consultation question-answering model for feedback training.
Citation Information
Patent Citations
Method and device for recommending answers
CN103425635A
Intelligent questioning and answering method and system
CN108256056A
Chat response method, electronic device and storage medium
CN108491433A
Large language model training method and device and related equipment
CN117390450A
Legal service system of large-scale pre-training language model based on deep learning
CN118445381A
Cited By
Enterprise-level papermaking process question and answer method and system based on large model
CN121051205A
A large model-based enterprise-level papermaking process question and answer method and system
CN121051205B