Knowledge understanding and answer generation method and device for language service
By building an enhanced input data set and target language service model, combined with multiple rounds of dialogue tracking processing, the shortcomings of the existing technology in deep understanding and high professional scenarios in specific fields are solved, and higher accuracy of knowledge understanding and coherence in answer generation are achieved.
Patent Information
- Application Number
- CN202411946513.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-06-03
AI Technical Summary
Existing service platforms or question-and-answer systems based on large language models perform poorly in deep understanding and high professional scenarios in specific fields, resulting in low accuracy of knowledge understanding and insufficient logical and coherence in answer generation.
By obtaining user query and initial language service datasets, preprocessing is performed to obtain query vectors, subject words and domain tags, building enhanced input datasets to enhance model knowledge comprehension ability, building target language service datasets and models, generating initial output answers, and generating final output answers through multiple rounds of dialogue tracking processing.
It improves the accuracy of knowledge understanding and the consistency of answer generation, can handle complex and multi-field user inquiries more accurately, and provides professional and coherent language services.
Smart Images

Figure CN120087462A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method and device for knowledge understanding and answer generation for language services. Background Art
[0002] With the enhancement of computing power and the growth of data volume, large models have demonstrated strong analysis and learning capabilities. Large models not only have powerful transfer learning capabilities, enabling rapid fine-tuning in different tasks and fields, but also have the ability to deeply integrate and process massive amounts of data. Existing service platforms or question-and-answer systems based on large language models mainly use general large language models and vertical domain models. However, general large language models are mainly optimized for broad fields and general tasks, and perform poorly in deep understanding of specific fields and high-professional scenarios, prone to one-sided or incomplete results, and have low knowledge understanding accuracy. Vertical domain large models, on the other hand, focus on specific industry tasks and can handle specific domain knowledge, but the logicality and coherence of the generated answer results are low.
[0003] In summary, the technical problems existing in the related art need to be improved. Summary of the Invention
[0004] Embodiments of the present invention provide a method and device for knowledge understanding and answer generation for language services, effectively improving the knowledge understanding accuracy and answer generation coherence.
[0005] On the one hand, embodiments of the present invention provide a method for knowledge understanding and answer generation for language services, including the following steps:
[0006] Obtain a first user query and an initial language service data set;
[0007] Perform a first preprocessing on the first user query to obtain a query vector, a topic word, and a domain label, where the query vector includes a query word, a query phrase, or a query sentence;
[0008] Construct an enhanced input data set according to the query vector, the topic word, and the domain label, where the enhanced input data set is used to enhance the model's knowledge understanding ability;
[0009] Construct a target language service data set according to the enhanced input data set and the initial language service data set;
[0010] Construct a target language service model according to the target language service data set;
[0011] Generate a first output answer according to the target language service model;
[0012] Perform multi-turn dialogue tracking processing based on the first user query and the first output answer to generate a second output answer.
[0013] In some embodiments, the first preprocessing of the first user query to obtain a query vector, a topic word, and a domain label includes:
[0014] Perform structural partitioning on the first user query to obtain the query word, the query phrase, or the query sentence;
[0015] Perform domain analysis on the first user query according to the candidate domain information of the resource service to obtain the topic word and the domain label.
[0016] In some embodiments, constructing an enhanced input data set according to the query vector, the topic word, and the domain label includes:
[0017] According to the topic word and the domain label, use the representational state transfer interface to connect multiple resource services to obtain an initial knowledge base vector, where the domains of the multiple resource services are different;
[0018] Calculate the cosine similarity according to the query vector and the initial knowledge base vector;
[0019] Perform screening and integration processing on the initial knowledge base vector according to the cosine similarity to obtain a target knowledge base vector;
[0020] Concatenate the target knowledge base vector and the query vector to obtain the enhanced input data set.
[0021] In some embodiments, constructing a target language service data set according to the enhanced input data set and the initial language service data set includes:
[0022] Perform second preprocessing on the initial language service data set to obtain a first language service data set, where the second preprocessing includes data cleaning, data denoising, data tokenization, or data annotation;
[0023] Perform data redundancy removal on the first language service data set to obtain a second language service data set;
[0024] Perform error correction on the second language service data set to obtain a third language service data set;
[0025] Combine the third language service data set and the enhanced input data set to obtain the target language service data set.
[0026] In some embodiments, constructing a target language service model according to the target language service data set includes:
[0027] Obtain an initial language service model;
[0028] Freeze the basic layer parameters of the initial language service model;
[0029] Input the target language service dataset into the initial language service model with frozen basic layer parameters, so that the initial language service model is trained to obtain the target language service model;
[0030] Calculate the continuous word sequence accuracy according to the number of reference translation matches and the total number of words in the candidate translation;
[0031] Calculate the bilingual evaluation substitute score according to the continuous word sequence accuracy, the maximum value of the word sequence, the weight coefficient, and the penalty factor. The bilingual evaluation substitute score is used to evaluate the matching degree between the generated answer and the reference answer;
[0032] Calculate the longest common subsequence metric score according to the total number of words in the reference text and the number of matching words in the longest common subsequence between the generated answer and the reference answer. The longest common subsequence metric score is used to evaluate the relevance between the generated answer and the first user query;
[0033] Set a preset matching degree threshold and a preset relevance threshold;
[0034] If the bilingual evaluation substitute score is less than the preset matching degree threshold or the longest common subsequence metric score is less than the preset relevance threshold, then perform fine-tuning processing on the data sampling frequency, the optimizer learning rate, and the model hyperparameters, and update the target language service model.
[0035] In some embodiments, the generating a first output answer according to the target language service model includes:
[0036] Input the first user query into the target language service model to obtain a reference answer;
[0037] If the reference answer is related to multiple resource services, disassemble the first user query to obtain multiple sub-questions;
[0038] Deduce each sub-question respectively to obtain corresponding sub-question answers;
[0039] Construct the first output answer according to multiple sub-question answers;
[0040] If the first output answer is different from the reference answer, check each sub-question answer to obtain conflict reasoning steps;
[0041] Revise the conflict reasoning step to generate a revised answer, and use the revised answer as the first output answer.
[0042] In some embodiments, the multi-round dialogue tracking process is performed according to the first user query and the first output answer to generate a second output answer, including:
[0043] Generate dynamic buried point information, which is used to record the first user query and the first output answer;
[0044] Obtain a second user query;
[0045] Generate the second output answer according to the second user query and the dynamic buried point information.
[0046] In some embodiments, the calculation of the cosine similarity according to the query vector and the initial knowledge base vector includes:
[0047] Calculate the cosine similarity according to the query vector and the initial knowledge base vector through the cosine similarity calculation formula, and the cosine similarity calculation formula is:
[0048]
[0049] In the formula, Sim is the cosine similarity, Q is the user query, Di is the resource service, v q is the query vector, v d is the initial knowledge base vector.
[0050] On the other hand, an embodiment of the present invention provides a knowledge understanding and answer generation device for language services, including:
[0051] The first module is used to obtain the first user query and the initial language service data set;
[0052] The second module is used to perform a first preprocessing on the first user query to obtain a query vector, a subject word, and a domain label, and the query vector includes a query word, a query phrase, or a query sentence;
[0053] The third module is used to construct an enhanced input data set according to the query vector, the subject word, and the domain label, and the enhanced input data set is used to enhance the model knowledge understanding ability;
[0054] The fourth module is used to construct a target language service data set according to the enhanced input data set and the initial language service data set;
[0055] The fifth module is used to construct a target language service model according to the target language service data set;
[0056] The sixth module is configured to generate a first output answer according to the target language service model;
[0057] The seventh module is configured to perform multi-round dialogue tracking processing according to the first user query and the first output answer, and generate a second output answer.
[0058] On the other hand, an embodiment of the present invention provides a computer device, including:
[0059] At least one processor;
[0060] At least one memory for storing at least one program;
[0061] When the at least one program is executed by the at least one processor, the at least one processor implements the method.
[0062] The beneficial effects of the present invention are as follows:
[0063] In an embodiment of the present invention, first, a first user query and an initial language service data set are obtained, the first user query is subjected to a first preprocessing to obtain a query vector, a topic word, and a domain label, then an enhanced input data set is constructed according to the query vector, the topic word, and the domain label, and a target language service data set is constructed according to the enhanced input data set and the initial language service data set, and then a target language service model is constructed according to the target language service data set, and a first output answer is generated according to the target language service model, and finally, multi-round dialogue tracking processing is performed according to the first user query and the first output answer to generate a second output answer, thereby realizing knowledge understanding and answer generation, and improving the accuracy of knowledge understanding and the coherence of answer generation.
[0064] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the specification and the drawings. Description of the Drawings
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0066] Figure 1 It is a flowchart of a method for knowledge understanding and answer generation for language service according to an embodiment of the present invention;
[0067] Figure 2Schematic diagram of a process for cross - domain semantic retrieval according to an embodiment of the present invention;
[0068] Figure 3 Schematic diagram of a training process of a target language service model according to an embodiment of the present invention;
[0069] Figure 4 Schematic diagram of a process for problem decomposition and multi - round answer generation according to an embodiment of the present invention;
[0070] Figure 5 Schematic diagram of an overall process for cross - domain semantic retrieval and multi - round answer generation according to an embodiment of the present invention;
[0071] Figure 6 Schematic diagram of the structure of a knowledge understanding and answer generation device for language services according to an embodiment of the present invention;
[0072] Figure 7 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed implementation manners
[0073] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application detailed in the appended claims.
[0074] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".
[0075] The terms "at least one", "a plurality", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each one of the corresponding plurality, and any one refers to any one of the plurality.
[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.
[0077] Before elaborating on the embodiments of this application in detail, some nouns and terms involved in the embodiments of this application are first explained, and the nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0078] Language service: Refers to a variety of services provided around social needs, including language knowledge dissemination, language skill cultivation, language technology support, etc. Language services cover aspects such as language education (such as Chinese character learning, traditional culture courses), language translation (such as simplified-traditional conversion services, language translation services), language technology (such as Chinese character information services, voice model examination services), and language culture (such as cultural and artistic translation services, knowledge Q&A), and can meet the actual needs of different groups and society.
[0079] In the related art, with the enhancement of computing power and the explosive growth of data volume, large models have demonstrated unprecedented analysis and learning capabilities, especially in the fields of natural language processing, computer vision, etc. Large models not only have powerful transfer learning capabilities, enabling rapid fine-tuning in different tasks and fields, but also have significant advantages in providing accurate and efficient services due to their ability to deeply integrate and process massive amounts of data. Existing service platforms or question-and-answer systems based on large language models mainly rely on the preliminary application of general large language models and vertical domain models. Among them, typical general large language models, such as the GPT series, BERT, T5, etc., can generate fluent language outputs for open-domain questions. However, these general models are mainly optimized for broad fields and general tasks and often perform poorly in the in-depth understanding and highly professional scenarios of specific fields; vertical domain large models focus on specific industry tasks, such as medical, legal, and historical, and perform better in dealing with domain terms, complex grammar structures, and cultural backgrounds through techniques such as domain fine-tuning and knowledge retrieval. These two types of models have their own focuses. General models are suitable for a wide range of tasks in open domains, while vertical models can meet the complex needs in highly professional scenarios. Existing general large language models mainly rely on large-scale general corpora for training, which enables them to show good language generation capabilities when dealing with open-domain tasks. However, when faced with highly professional fields (such as history, literature, law, etc.), these models often have difficulty fully understanding the professional terms, specific grammar phenomena, or cultural backgrounds within the field, resulting in deficiencies in the accuracy, logic, and context consistency of the generated language responses. In addition, existing general models lack the ability to deeply model and reason about the domain knowledge system, so they are prone to one-sided or incomplete results when answering questions in specific fields. Compared with general large language models, vertical domain large models focus on the in-depth knowledge modeling of specific fields and provide professional support through rich domain corpora and dedicated knowledge graphs. These models combine retrieval-augmented generation (RAG) technology to closely integrate external knowledge bases with language generation, making them more accurate and context-consistent when answering complex professional questions. In addition, vertical models rely on fine-tuning, using domain-specific data to optimize model parameters, enabling them to perform excellently when faced with professional terms, complex grammar, and cultural backgrounds in specific scenarios and be competent for multi-turn conversations and complex knowledge service requirements.
[0080] In the field of language services, the importance of large models in vertical domains is particularly prominent. The platform needs to handle scenarios such as policy interpretation, public service translation, legal text analysis, etc., which not only involve precise conversion between multiple languages but also need to take into account the differences in different cultures and language usage habits. This requirement goes beyond the scope that traditional general models can cover. Currently, in the field of language services, there are still obvious deficiencies in general large language models and vertical domain models, including: (1) Existing retrieval enhancement methods perform poorly in dealing with the fragmentation problems of knowledge in different domains. The lack of an efficient cross-domain semantic retrieval system causes the model to be unable to quickly and accurately obtain the required knowledge resources when facing complex and multi-domain user queries, affecting service quality. Traditional semantic-based service discovery methods are difficult to ensure that the model can quickly and accurately obtain relevant knowledge for Chinese language data resources in different domains, providing scientific references and explanations for dialogue interactions with users. (2) Existing technologies cannot effectively solve the problem of dynamic update of domain knowledge. It is difficult for the model to handle emerging vocabulary, cutting-edge topics, or regulatory changes based on the coherence, logic, and cultural and emotional depth of the generated content, resulting in the generated content not meeting the current language environment or policy requirements and affecting the user experience. (3) Although vertical domain models have been fine-tuned in professional fields, there are still deficiencies in the logic and coherence of the generated results when dealing with complex grammar, industry terms, and cultural contexts. Facing professional terms and industry vocabulary in the Chinese field, it is difficult for the model to effectively understand their rich semantics, unique grammar features, and the underlying cultural background, forming in-depth learning and understanding, so as to more accurately provide users with the required services.
[0081] In view of this, this embodiment realizes high-quality language services by integrating cross-domain knowledge retrieval, large model fine-tuning technology, and an interpretable interaction generation framework. This embodiment focuses on the diverse needs of language services and realizes professional and coherent language generation through model fine-tuning, semantic retrieval, and logical inference, and can flexibly handle complex cross-domain scenarios. At the same time, this embodiment focuses on improving the model's understanding and generation capabilities in professional field language data to meet the needs of the language knowledge field for combining industry terms, policy backgrounds, and cultural emotions, and providing accurate and coherent language services for users through semantic retrieval and logical inference. The present invention overcomes the deficiencies of existing technologies in professional term processing, dynamic knowledge update, interpretable reasoning, etc. by constructing a large language model for vertical domains of language resources, integrating cross-domain professional knowledge and semantic information processing technologies, and realizes an intelligent and efficient language resource service platform. The model has a high accuracy in multi-round interactions.
[0082] A method for knowledge understanding and answer generation for language services provided by an embodiment of the present application relates to the technical field of natural language processing. The method for knowledge understanding and answer generation for language services provided by the embodiment of the present application can be applied to a terminal, can also be applied to a server, or can be software running on a terminal or a server. In some embodiments, the terminal may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, and can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a method for knowledge understanding and answer generation for language services, etc., but is not limited to the above forms.
[0083] The present application can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0084] The following specifically explains the embodiments of the present application with reference to the accompanying drawings:
[0085] Figure 1 is an optional flowchart of a method for knowledge understanding and answer generation for language services provided by an embodiment of the present application, Figure 1 The method in may include but is not limited to steps S101 to S107.
[0086] Step S101, obtain a first user query and an initial language service data set;
[0087] Step S102, perform a first preprocessing on the first user query to obtain a query vector, a topic word, and a domain label, where the query vector includes a query word, a query phrase, or a query sentence;
[0088] Step S103: Construct an enhanced input data set based on the query vector, subject words, and domain labels. The enhanced input data set is used to enhance the model's knowledge understanding ability;
[0089] Step S104: Construct a target language service data set based on the enhanced input data set and the initial language service data set;
[0090] Step S105: Construct a target language service model based on the target language service data set;
[0091] Step S106: Generate a first output answer according to the target language service model;
[0092] Step S107: Perform multi-turn dialogue tracking processing based on the first user query and the first output answer to generate a second output answer.
[0093] Steps S101 to S107 shown in the embodiments of the present application achieve knowledge understanding and answer generation, improving the accuracy of knowledge understanding and the coherence of answer generation.
[0094] In step S101 of some embodiments, the first user query can be obtained through user input information, and the initial language service data set can be obtained through a language service platform. The first user query and the initial language service data set can also be obtained through other means, which are not limited thereto.
[0095] In some embodiments, in step S102, the first preprocessing of the first user query to obtain the query vector, subject words, and domain labels may include, but is not limited to, the following steps:
[0096] Perform a structural division on the first user query to obtain query words, query phrases, or query sentences;
[0097] Perform domain analysis on the first user query according to the resource service candidate domain information to obtain subject words and domain labels.
[0098] In some embodiments, in order to adapt to the dynamically changing knowledge requirements in language services, this embodiment constructs a comprehensive and efficient cross-domain retrieval system based on a vertical domain large model and a retrieval-augmented generation framework, which helps the model better dynamically update the resource library to handle queries that require cutting-edge information. By analyzing language data resources in different fields, rapid and accurate knowledge acquisition is achieved, providing users with scientific and rigorous references and explanations. Among them, the data processing process of the cross-domain retrieval system includes steps S102 - S103. In some embodiments, the first user query can be preprocessed to obtain a query vector, a topic word, and a domain label. Among them, the query vector includes query words, query phrases, or query sentences. The first user query can be first structurally divided to obtain query words, query phrases, or query sentences, and then, based on the candidate domain information of the resource service, the first user query is analyzed in terms of the domain to obtain the topic word and the domain label. Exemplarily, the user query can be structured into query words, query phrases, or query sentences, and the corresponding topic word and domain label are generated based on the candidate domain information of the resource service. These topic words and domain labels will serve as the core input of the unified search interface, ensuring that the model can perform exact matching and semantic understanding for the query. It can be understood that after receiving the user query, the content of the user query is parsed and structured. If there is sufficient knowledge to answer the question, the answer is directly generated; if the question involves multiple fields or requires more in-depth professional knowledge, the cross-domain retrieval module is activated for cross-domain semantic retrieval.
[0099] In some embodiments, in step S103, according to the query vector, the topic word, and the domain label, an enhanced input data set can be constructed, which may include but is not limited to the following steps:
[0100] According to the topic word and the domain label, use the representational state transfer interface to connect multiple resource services to obtain an initial knowledge base vector, and the domains of the multiple resource services are different;
[0101] Calculate the cosine similarity according to the query vector and the initial knowledge base vector;
[0102] According to the cosine similarity, perform screening and integration processing on the initial knowledge base vector to obtain a target knowledge base vector;
[0103] Concatenate the target knowledge base vector and the query vector to obtain an enhanced input data set.
[0104] In some embodiments, an enhanced input data set can be constructed according to the query vector, the topic word, and the domain label, where the enhanced input data set is used to enhance the model's knowledge understanding ability. The process of cross-domain semantic retrieval is as Figure 2As shown, first, according to the subject terms and domain tags, multiple resource services can be connected using the Representational State Transfer (RESTful) interface to obtain an initial knowledge base vector, where the domains of the multiple resource services are different. Then, the cosine similarity is calculated based on the query vector and the initial knowledge base vector. Exemplarily, different domain resource services can be invoked through the Representational State Transfer (RESTful) interface to achieve cross-domain semantic retrieval. The retrieval module docks multiple resource services based on the parsed subject terms and domain tags to ensure seamless docking of the query with the knowledge bases in each domain. To measure the semantic similarity between the user query and the content of the resource service, the cosine similarity can be calculated according to the query vector and the initial knowledge base vector using the cosine similarity calculation formula, where the cosine similarity calculation formula is: In the formula, Sim is the cosine similarity, Q is the user query, Di is the resource service, v q is the query vector, and v d is the initial knowledge base vector. Then, based on the cosine similarity, the initial knowledge base vector is screened and integrated to obtain the target knowledge base vector. Finally, the target knowledge base vector and the query vector are concatenated to obtain the enhanced input dataset. Exemplarily, when the retrieval module returns knowledge (i.e., the initial knowledge base vector) from multiple domains, the vertical domain large model will screen and integrate this multi-source knowledge. The large model will filter out the most relevant results (i.e., the target knowledge base vector) based on the subject terms of the user query and concatenate the target knowledge base vectors into the enhanced input Input = Q + D Filtered , to obtain the enhanced input dataset, where Input is the enhanced input, Q is the user query, and D Filtered is the target knowledge base vector. The enhanced input dataset is re-input into the large model to ensure that the generated answer covers all necessary information, enabling this embodiment to combine cross-domain retrieval with language generation to achieve more intelligent and dynamic knowledge processing and question-answer generation.
[0105] In some embodiments, in step S104, constructing the target language service dataset based on the enhanced input dataset and the initial language service dataset may include, but is not limited to, the following steps:
[0106] Perform a second preprocessing on the initial language service dataset to obtain the first language service dataset. The second preprocessing includes data cleaning, data denoising, data tokenization, or data annotation;
[0107] Remove data redundancy from the first language service dataset to obtain the second language service dataset;
[0108] Correct errors in the second language service dataset to obtain the third language service dataset;
[0109] Combine the third - language service data set and the enhanced input data set to obtain the target - language service data set.
[0110] In some embodiments, to ensure the professionalism and accuracy of the language service model when dealing with specific - domain problems, a fine - tuning technique can be adopted. Through the collection of high - quality corpora, fine - tuning training, and dynamic knowledge update, the model's understanding of Chinese grammar, sentence patterns, and cultural backgrounds can be deepened. Ensure that the model not only understands the surface structure of Chinese but also can insight into the cultural emotions behind it, enabling the model to have learning and reasoning abilities in a multi - task environment, forming the understanding and learning logic route of the language service model. First, the initial language service data set can be subjected to a second pre - processing to obtain the first language service data set. Among them, the second pre - processing includes data cleaning, data denoising, data tokenization, or data annotation. Exemplarily, high - quality professional corpora (i.e., the initial language service data set) can be introduced from a language resource service platform, covering ancient classic literature, modern Chinese literature, scientific and technological literature, policies and regulations, etc., and cooperating with a manually annotated question - and - answer data set to ensure the richness and coverage of the corpora. All data is subjected to corpus cleaning and pre - processing, including denoising, tokenization, and annotation, to ensure compliance with training requirements. Then, data redundancy elimination is performed on the first language service data set to obtain the second language service data set, error correction is performed on the second language service data set to obtain the third language service data set, and the third language service data set and the enhanced input data set are combined to obtain the target - language service data set. Exemplarily, if data redundancy or errors are detected, data redundancy elimination and error correction can be performed on the first language service data set. At the same time, by adding newly released regulations and newly launched resources (i.e., the enhanced input data set) to the corpus in a timely manner to maintain the dynamic update of the corpus, the target - language service data set is obtained, realizing that when the generated answer is insufficient, the latest knowledge (i.e., the enhanced input) can be obtained through the cross - domain retrieval module, ensuring the accuracy and timeliness of the answers generated by the model.
[0111] In some embodiments, in step S105, according to the target - language service data set, constructing the target - language service model may include, but is not limited to, the following steps:
[0112] Obtain the initial language service model;
[0113] Freeze the basic - layer parameters of the initial language service model;
[0114] Input the target - language service data set into the initial language service model with frozen basic - layer parameters, so that the initial language service model is trained to obtain the target - language service model;
[0115] Calculate the continuous word - sequence accuracy according to the number of reference - translation matches and the total number of words in the candidate translation;
[0116] Calculate the bilingual evaluation substitute score based on the continuous word sequence accuracy, the maximum value of the word sequence, the weight coefficient, and the penalty factor. The bilingual evaluation substitute score is used to evaluate the matching degree between the generated answer and the reference answer;
[0117] Calculate the longest common subsequence metric score based on the total number of words in the reference text and the number of matching words in the longest common subsequence between the generated answer and the reference answer. The longest common subsequence metric score is used to evaluate the relevance between the generated answer and the first user query;
[0118] Set a preset matching degree threshold and a preset relevance threshold;
[0119] If the bilingual evaluation substitute score is less than the preset matching degree threshold or the longest common subsequence metric score is less than the preset relevance threshold, then fine-tune the data sampling frequency, the optimizer learning rate, and the model hyperparameters, and update the target language service model.
[0120] In some embodiments, a general initial language service model can be obtained first. Exemplarily, a model pre-trained on a large-scale general corpus can be selected as the initial language service model and used as the basis for subsequent optimization. These models have good natural language understanding and generation capabilities, providing an efficient starting point for fine-tuning. More specifically, when the initial language service model cannot fully cover the domain requirements, enhanced input data from multiple domains can be supplemented through a cross-domain retrieval mechanism. Then, freeze the basic layer parameters of the initial language service model and only fine-tune the high-level parameters related to the domain to avoid the model forgetting general knowledge and maintain its domain adaptability. Then input the target language service data set into the initial language service model with the basic layer parameters frozen, so that the initial language service model is trained to obtain the target language service model. The training process of the target language service model is as Figure 3 shown, and a fine-tuning process of dynamic adjustment is completed based on the BLEU (bilingual evaluation substitute) score and the ROUGE-L (based on the longest common subsequence) evaluation metric. The continuous word sequence accuracy can be calculated according to the number of reference translation matches and the total number of words in the candidate translation. Among them, the calculation formula for the continuous word sequence accuracy is: In the formula, p n is the accuracy of the continuous word sequence n-gram, N h is the number of n-grams that match the reference translation in the candidate translation, N z$n$ is the total number of n-grams in the candidate translation, where an n-gram refers to a sequence of $n$ consecutive words (which can be 1-gram, 2-gram, 3-gram, and 4-gram). Then, based on the consecutive word sequence accuracy, the maximum value of the word sequence, the weight coefficient, and the penalty factor, the bilingual evaluation substitute score is calculated. Here, the bilingual evaluation substitute score is used to evaluate the matching degree between the generated answer and the reference answer. The calculation formula for the bilingual evaluation substitute score is: In the formula, BLEU is the bilingual evaluation substitute score, BP is the penalty factor introduced to prevent overly short outputs, N is the maximum value of the n-gram word sequence (which can take a value of 4), and ω n is the weight coefficient. Then, based on the total number of words in the reference text and the number of words in the longest common subsequence match between the generated answer and the reference answer, the longest common subsequence metric score is calculated. Here, the longest common subsequence metric score is used to evaluate the relevance between the generated answer and the first user query. The calculation formula for the longest common subsequence metric score is: In the formula, ROUGE is the longest common subsequence metric score, N w is the number of words in the LCS (longest common subsequence) match between the generated answer and the reference answer, and N c is the total number of words in the reference text. Finally, a preset matching degree threshold and a preset relevance threshold are set. If the bilingual evaluation substitute score is less than the preset matching degree threshold or the longest common subsequence metric score is less than the preset relevance threshold, then the data sampling frequency, the optimizer learning rate, and the model hyperparameters are fine-tuned, and the target language service model is updated. Exemplarily, the performance monitoring module can be initialized first to calculate the BLEU (bilingual evaluation substitute) score and the ROUGE-L (based on the longest common subsequence) score between the generated answer and the reference answer in real time, and the score calculation and feedback are immediately executed after each batch of training is completed.
[0121] Then, set a performance threshold. For example, when BLEU is lower than 0.7 or ROUGE-L is lower than 0.6, the system automatically marks the task as a low-performance task and records it in the task queue. If it is detected that the task is marked, the system will immediately trigger a fine-tuning mechanism to dynamically adjust the data sampling frequency of the task and the learning rate of the optimizer. For example, the learning rate is reduced to 1e-5 and the sampling frequency is increased. Then, the task is extracted from the task queue and fine-tuning training is preferentially performed. After the training is completed, a new checkpoint is saved. If the training effect is improved, the latest state can be retained; otherwise, it can be rolled back to the previous checkpoint. Finally, the two metrics of BLEU and ROUGE-L of the task are monitored again. If both of these metrics meet the preset requirements, the task mark is removed; if the preset requirements are not met, the model hyperparameters are further adjusted, and the fine-tuning training is repeated to ensure that the task reaches the expected performance. The entire process can achieve dynamic optimization and automatic adjustment of the model by repeatedly performing performance detection, task marking, fine-tuning execution, and checkpoint saving. It can be understood that fine-tuning is triggered according to performance metrics. For example, the learning rate is reduced to a small value (such as 1e-5) to ensure more stable updates, and the increase in data sampling frequency gives the model more training opportunities on this task. That is, based on the performance of the model and the task requirements, some of its parameters (especially the high-level task-related parameters) are further updated through new task data, so that the model can complete the specified task more accurately.
[0122] In some embodiments, in step S106, generating the first output answer according to the target language service model may include, but is not limited to, the following steps:
[0123] Input the first user query into the target language service model to obtain a reference answer;
[0124] If the reference answer is related to multiple resource services, disassemble the first user query to obtain multiple sub-questions;
[0125] Deduce each sub-question separately to obtain the corresponding sub-question answer;
[0126] Construct the first output answer according to multiple sub-question answers;
[0127] If the first output answer is different from the reference answer, check each sub-question answer to obtain conflict reasoning steps;
[0128] Revise the conflict reasoning steps to generate a revised answer, and use the revised answer as the first output answer.
[0129] In some embodiments, after constructing the target language service model, this embodiment improves the large model in generating answers based on language knowledge, which can be used for complex and multi-scenario knowledge Q&A, such as various scenarios like text correction, ambiguity analysis, and traditional-simplified Chinese translation. The model will be able to, according to the user's request, gradually analyze the next execution steps by decomposing the query into sub-questions, and adopt an interpretable generation framework to ensure that the model can not only generate logically clear answers but also enable users to understand the source and logical inference of the answers by showing the generation process, ensuring the depth of logic and cultural sentiment in the field of language knowledge. The question decomposition and multi-round answer generation process are as follows Figure 4 shown. First, the first user query can be input into the target language service model to obtain a baseline answer. If the baseline answer is related to multiple resource services, the first user query is decomposed to obtain multiple sub-questions. Exemplarily, if the baseline answer comes from a single resource in the knowledge service, answer A can be directly generated and output. If the query is complex (i.e., the answer baseline answer comes from multiple resources in the knowledge service), the first user query is decomposed to obtain multiple sub-questions {Q 1 , Q 2 , …, Q n}. Then, each sub-question is deduced separately to gradually deduce and generate intermediate reasoning steps to obtain the corresponding sub-question answers {S 1 , S 2 , …, S n}, and based on the multiple sub-question answers, the first output answer is constructed. Among them, the expression of the first output answer is: In the formula, Answer is the first output answer, S i is the i-th sub-question, that is, the output corresponding to each reasoning step, and n is the total number of sub-questions. Moreover, if there are sub-questions that cannot be answered, cross-domain semantic retrieval can be used to update the input training data set through new knowledge or resources. If the first output answer is different from the baseline answer, each sub-question answer is checked to obtain the conflicting reasoning steps, and the conflicting reasoning steps are revised to generate a revised answer, and the revised answer is used as the first output answer. Exemplarily, the first output answer and the baseline answer can be verified. If the first output answer is the same as the baseline answer, that is, Answer = A, the answer is output and S 1 , S 2 , …, S n and its knowledge source are shown to the user; if the first output answer is different from the baseline answer, that is, there is a contradiction between Answer and A, the answer generation chain can be checked to find inconsistent or conflicting reasoning steps, and they are revised to generate a revised answer containing all the revised results of the sub-questions.
[0130] In some embodiments, in step S107, multi-turn dialogue tracking processing is performed based on the first user query and the first output answer to generate a second output answer, which may include but is not limited to the following steps:
[0131] Generate dynamic logging information for recording the first user query and the first output answer;
[0132] Obtain a second user query;
[0133] Generate a second output answer based on the second user query and the dynamic logging information.
[0134] In some embodiments, dynamic logging information can be generated first, where the dynamic logging information is used to record the first user query and the first output answer. Then, a second user query is obtained, and a second output answer is generated based on the second user query and the dynamic logging information. Exemplarily, after the user input query, dynamic logging can be started to generate dynamic logging information. When the user interacts with the model, the first user query and the first output answer are recorded and used as context for subsequent question answering and generation optimization. If the user has a subsequent query, the second user query is obtained. The multi-turn dialogue of the user is tracked, and the previously generated answer is used as context to participate in subsequent question answering. If the query is unclear, a clarification question can be generated to prompt the user to clarify; if the query is clear, a coherent second output answer is generated. Further, if the user has further questions (i.e., more user queries), the generated answer can be optimized according to the context; if the user is not satisfied, the system records the feedback and adjusts the generation strategy. Further, if a certain type of question appears frequently, it can be marked as high priority and recorded in the dynamic logging, and the model fine-tuning is triggered during idle time to optimize according to the high-priority questions.
[0135] In some embodiments, the overall process of cross-domain semantic retrieval and multi-turn answer generation is as Figure 5 shown. First, the user query can be obtained. If the generated answer cannot be directly obtained using the target language service model, a cross-domain semantic query is performed using the unified retrieval interface to obtain an enhanced input data set, and the target language service model is updated using the enhanced input data set. If the user makes multiple queries, in multi-turn answer generation, the context information can be recorded using dynamic logging, and multi-step inference revision is performed to obtain an interpretable answer, which is returned to the user.
[0136] In some embodiments, this embodiment generates professional, efficient, and timely Q&A for users by retrieving service resources that are continuously updated in knowledge services in different fields. It not only improves the language processing ability of the model in vertical fields but also can dynamically handle new words and unappeared content in a changing language environment, and continuously optimizes the generation process through multi-round interactions. At the same time, this embodiment proposes a comprehension enhancement technology based on dynamic cross-domain retrieval for language knowledge. By enhancing the understanding and retrieval of user queries, it realizes more efficient and accurate query processing and supports the dynamic update of the knowledge base. For the language platform, a response generation technology with multi-step reasoning and dynamic data point optimization is proposed, which improves the interpretability of the responses generated by the model. On this basis, the generation strategy is optimized according to user feedback, and the model is triggered for fine-tuning during idle time to enhance future Q&A capabilities.
[0137] The beneficial effects of implementing the embodiments of the present invention include: The embodiments of the present invention first obtain a first user query and an initial language service data set, perform a first preprocessing on the first user query to obtain a query vector, a topic word, and a domain label, then construct an enhanced input data set according to the query vector, the topic word, and the domain label, and construct a target language service data set according to the enhanced input data set and the initial language service data set. Then, a target language service model is constructed according to the target language service data set, and a first output response is generated according to the target language service model. Finally, multi-round dialogue tracking processing is performed according to the first user query and the first output response to generate a second output response, thereby realizing knowledge understanding and response generation, and improving the accuracy of knowledge understanding and the coherence of response generation. At the same time, based on the existing corpus (language service data set), on the basis of generating professional, accurate, and highly interpretable Chinese responses, the model retrieves and obtains the latest knowledge services by retrieving the dynamically updated corpus, and can continuously absorb, integrate, and innovate knowledge during real-time interaction with users. When the model faces cutting-edge topics, emerging words, or content that has not appeared in the traditional corpus, it can combine the existing knowledge system, and through retrieving knowledge and gradually inferring, generate responses with relevant explanations and corresponding sources for users. This not only ensures the coherence, logic, and cultural sentiment of the generated content but also can meet the users' pursuit of new knowledge, new thinking, and new insights, and truly realizes the frontier generation and dissemination of Chinese language knowledge.
[0138] As Figure 6 shown, the embodiments of the present invention also provide a knowledge understanding and response generation device for language services, including:
[0139] A first module 801, configured to obtain a first user query and an initial language service data set;
[0140] The second module 802 is configured to perform a first preprocessing on the first user query to obtain a query vector, a topic word, and a domain label, where the query vector includes a query word, a query phrase, or a query sentence;
[0141] The third module 803 is configured to construct an enhanced input data set according to the query vector, the topic word, and the domain label, and the enhanced input data set is used to enhance the model's knowledge understanding ability;
[0142] The fourth module 804 is configured to construct a target language service data set according to the enhanced input data set and the initial language service data set;
[0143] The fifth module 805 is configured to construct a target language service model according to the target language service data set;
[0144] The sixth module 806 is configured to generate a first output answer according to the target language service model;
[0145] The seventh module 807 is configured to perform multi-round dialogue tracking processing according to the first user query and the first output answer to generate a second output answer.
[0146] The content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0147] As Figure 7 shown, an embodiment of the present invention further provides a computer device, including:
[0148] At least one processor 901;
[0149] At least one memory 902, configured to store at least one program;
[0150] When at least one program is executed by at least one processor, at least one processor implements Figure 1 the method shown.
[0151] The content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0152] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A knowledge understanding and answer generation method for language service, characterized in that: The following steps are involved: Obtaining a first user query and an initial language service dataset; Performing a first preprocessing on the first user query to obtain a query vector, a subject word, and a domain label, wherein the query vector includes a query word, a query phrase, or a query sentence; Constructing an enhanced input data set according to the query vector, the subject word and the domain label, wherein the enhanced input data set is used to enhance the knowledge understanding capability of the model; Constructing a target language service dataset according to the enhanced input dataset and the initial language service dataset; Building a target language service model according to the target language service dataset; generating a first output answer according to the target language service model; Based on the first user query and the first output answer, multiple rounds of dialogue tracking processing are performed to generate a second output answer.
2. The method according to claim 1, characterized in that The first preprocessing of the first user query to obtain a query vector, a subject word, and a domain label includes: Performing structural division on the first user query to obtain the query word, the query phrase or the query sentence; According to the resource service candidate domain information, a domain analysis is performed on the first user query to obtain the subject word and the domain label.
3. The method according to claim 1, characterized in that: The step of constructing an enhanced input data set according to the query vector, the subject word, and the domain label includes: According to the subject words and the domain labels, a plurality of resource services are connected using a representational state transfer interface to obtain an initial knowledge base vector, wherein the domains of the plurality of resource services are different; Calculating cosine similarity based on the query vector and the initial knowledge base vector; According to the cosine similarity, the initial knowledge base vector is screened and integrated to obtain a target knowledge base vector; The target knowledge base vector and the query vector are concatenated to obtain the enhanced input data set.
4. The method according to claim 1, characterized in that: The step of constructing a target language service dataset according to the enhanced input dataset and the initial language service dataset includes: Performing a second preprocessing on the initial language service data set to obtain a first language service data set, wherein the second preprocessing includes data cleaning, data denoising, data segmentation or data labeling; Eliminating data redundancy from the first language service data set to obtain a second language service data set; performing error correction on the second language service data set to obtain a third language service data set; The third language service dataset and the enhanced input dataset are combined to obtain the target language service dataset.
5. The method according to claim 1, characterized in that: The step of constructing a target language service model according to the target language service dataset includes: Get the initial language service model; Freezing the basic layer parameters of the initial language service model; Inputting the target language service data set into the initial language service model after the basic layer parameters are frozen, so that the initial language service model is trained to obtain the target language service model; Calculate the continuous word sequence accuracy based on the number of reference translation matches and the total number of words in the candidate translation; Calculating a bilingual assessment substitute score according to the continuous word sequence accuracy, the word sequence maximum value, the weight coefficient and the penalty factor, wherein the bilingual assessment substitute score is used to evaluate the matching degree between the generated answer and the reference answer; Calculate a longest common subsequence index score based on the total number of words in the reference text and the number of longest common subsequence matching words between the generated answer and the reference answer, wherein the longest common subsequence index score is used to evaluate the relevance between the generated answer and the first user query; Setting a preset matching degree threshold and a preset correlation threshold; If the bilingual assessment substitute score is less than the preset matching degree threshold or the longest common subsequence index score is less than the preset relevance threshold, the data sampling frequency, optimizer learning rate and model hyperparameters are fine-tuned to update the target language service model.
6. The method according to claim 1, characterized in that Generating a first output answer according to the target language service model includes: Inputting the first user query into the target language service model to obtain a benchmark answer; If the benchmark answer is related to multiple resource services, the first user query is decomposed to obtain multiple sub-questions; Deducing each of the sub-questions separately to obtain the corresponding answer to the sub-question; Constructing the first output answer based on the multiple answers to the sub-questions; If the first output answer is different from the reference answer, each of the sub-question answers is checked to obtain a conflict reasoning step; The conflicting reasoning step is revised to generate a revised answer, and the revised answer is used as the first output answer.
7. The method according to claim 1, characterized in that The performing multi-round dialogue tracking processing according to the first user query and the first output answer to generate a second output answer includes: Generate dynamic embedding information, where the dynamic embedding information is used to record the first user query and the first output answer; Get the second user query; Generate the second output answer based on the second user query and the dynamic embedding information.
8. The method according to claim 3, characterized in that The calculating cosine similarity according to the query vector and the initial knowledge base vector includes: According to the query vector and the initial knowledge base vector, the cosine similarity is calculated by a cosine similarity calculation formula, and the cosine similarity calculation formula is: Where Sim is the cosine similarity, Q is the user query, Di is the resource service, and v q is the query vector, v d is the initial knowledge base vector.
9. A knowledge understanding and answer generation device for language service, characterized in that: include: A first module is used to obtain a first user query and an initial language service dataset; A second module is used to perform a first preprocessing on the first user query to obtain a query vector, a subject word and a domain label, wherein the query vector includes a query word, a query phrase or a query sentence; A third module is used to construct an enhanced input data set according to the query vector, the subject word and the domain label, wherein the enhanced input data set is used to enhance the knowledge understanding ability of the model; A fourth module is used to construct a target language service dataset according to the enhanced input dataset and the initial language service dataset; A fifth module is used to construct a target language service model according to the target language service dataset; A sixth module, configured to generate a first output answer according to the target language service model; The seventh module is used to perform multi-round dialogue tracking processing according to the first user query and the first output answer to generate a second output answer.
10. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 8.
Citation Information
Cited By
Human-computer interaction method and device, electronic equipment and storage medium
CN120596632A