Question and answer data determination method and device, storage medium and electronic equipment

By acquiring multi-dimensional data and multi-role prompts to generate multi-role question-and-answer data, and performing semantic analysis, the problem of low accuracy and efficiency of question-and-answer data in enterprise knowledge bases has been solved, thereby improving the accuracy and efficiency of question-and-answer data and enhancing user experience.

CN121882201APending Publication Date: 2026-04-17BEIJING ZHONGKE JINDEZHU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHONGKE JINDEZHU INTELLIGENT TECH CO LTD
Filing Date
2025-12-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

The accuracy and efficiency of question-and-answer data in existing enterprise knowledge bases are low, resulting in low accuracy and efficiency in answering user inquiries and affecting user experience.

Method used

By acquiring multi-dimensional data and multi-role prompts for the target product, multi-role question-and-answer data is generated, and semantic analysis is performed to determine the target answer data corresponding to the target question data, which is then stored in the question-and-answer knowledge base.

Benefits of technology

It improves the accuracy and efficiency of question-and-answer data, enables structured storage of question-and-answer knowledge, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121882201A_ABST
    Figure CN121882201A_ABST
Patent Text Reader

Abstract

The invention discloses a question and answer data determination method and device, a storage medium and electronic equipment, and relates to the technical field of data processing, and the method comprises the steps: obtaining multi-dimensional data of a target product and multi-role prompt information corresponding to the multi-dimensional data; generating multi-role question and answer data corresponding to the multi-dimensional data based on the multi-role prompt information; performing semantic analysis processing on multi-role answer data in the multi-role question and answer data according to target question data in the multi-role question and answer data to obtain target answer data corresponding to the target question data; and determining the target question data and the target answer data as target question and answer data corresponding to the multi-dimensional data, wherein the target question and answer data is used for being stored in a question and answer knowledge base corresponding to the target product. Compared with the prior art, the method has the advantages that coverage of the question and answer data on the multi-role view angle is achieved, the question and answer data determination accuracy is improved, the question and answer data processing efficiency of the model is improved, and then the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, storage medium and electronic device for determining question and answer data. Background Technology

[0002] In the daily operation of an enterprise, a large amount of product-related questions and answers are generated during customer inquiries and business processing. Enterprises need to store the question and answer data, which includes a large amount of product-related questions and answers, in a knowledge base to support subsequent customer inquiry responses and internal business training.

[0003] Currently, the determination of question and answer data in existing enterprise knowledge bases involves first manually screening valid information, then organizing the screened questions and answers according to a format, and finally having designated personnel check the information one by one to ensure it is correct before entering it into the knowledge base system.

[0004] However, using this method, the knowledge base can only store manually selected and organized question and answer data, resulting in low accuracy and efficiency in answering user inquiries, which in turn affects the user experience. Summary of the Invention

[0005] In view of this, this application provides a question-and-answer data determination method, apparatus, storage medium and electronic device, the main purpose of which is to improve the technical problem in the prior art that if the context length corresponding to the dialogue data increases, the video memory usage will increase linearly, resulting in video memory overflow, which in turn affects the running performance of the model.

[0006] Firstly, this application provides a method for determining question-and-answer data, including: Obtain multi-dimensional data of the target product and multi-role prompts corresponding to the multi-dimensional data; Based on the multi-role prompting information, multi-role question-and-answer data corresponding to the multi-dimensional data is generated; Based on the target question data in the multi-role question and answer data, semantic analysis processing is performed on the multi-role answer data in the multi-role question and answer data to obtain the target answer data corresponding to the target question data. The target question data and the target answer data are determined as the target question-and-answer data corresponding to the multi-dimensional data, and the target question-and-answer data is used to store in the question-and-answer knowledge base corresponding to the target product.

[0007] Optionally, the step of performing semantic analysis on the multi-role answer data in the multi-role question-and-answer data based on the target question data in the multi-role question-and-answer data to obtain the target answer data corresponding to the target question data includes: The target question data is obtained by clustering the multi-role question data in the multi-role question-and-answer data; Based on the target question data, semantic analysis is performed on the multi-role answer data to obtain a set of semantic information corresponding to the multi-role answer data; The target answer data is determined based on the semantic information set and the multi-role answer data.

[0008] Optionally, determining the target answer data based on the semantic information set and the multi-role answer data includes: In the event of semantic conflicts among the semantic information in the semantic information set, the target answer data is determined from the multi-role answer data based on the credibility data corresponding to each of the multi-role answer data. If there are no semantic conflicts in the semantic information set, the multi-role answer data is fused to generate the target answer data.

[0009] Optionally, determining the target question data and the target answer data as the target question-and-answer data corresponding to the multi-dimensional data includes: Semantic matching is performed between the target question-and-answer data and the multi-dimensional data; If the target question-and-answer data matches the multi-dimensional data, the target question-and-answer data is stored in the question-and-answer knowledge base.

[0010] Optionally, after storing the target question-and-answer data in the question-and-answer knowledge base, the process includes: Based on the feedback data of the target question-and-answer data, an optimization strategy for the target question-and-answer data is generated. The target question-and-answer data is adjusted according to the optimization strategy to optimize the question-and-answer data in the question-and-answer knowledge base.

[0011] Optionally, generating multi-role question-and-answer data corresponding to the multi-dimensional data based on the multi-role prompt information includes: Based on the data categories corresponding to the multi-dimensional data, at least one data slice corresponding to the multi-dimensional data is determined; Determine the question-and-answer knowledge content corresponding to the at least one data segment; Based on the multi-role prompting information, question-and-answer data is generated from the question-and-answer knowledge content to obtain the multi-role question-and-answer data.

[0012] Secondly, this application provides a question-and-answer data determination device, comprising: The acquisition module is configured to acquire multi-dimensional data of the target product and multi-role prompt information corresponding to the multi-dimensional data. The generation module is configured to generate multi-role question-and-answer data corresponding to the multi-dimensional data based on the multi-role prompt information; The processing module is configured to perform semantic analysis on the multi-role answer data in the multi-role question-and-answer data based on the target question data in the multi-role question-and-answer data, so as to obtain the target answer data corresponding to the target question data; The determination module is configured to determine the target question data and the target answer data as the target question-and-answer data corresponding to the multi-dimensional data, and the target question-and-answer data is used to store in the question-and-answer knowledge base corresponding to the target product.

[0013] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the question-and-answer data determination method described in the first aspect.

[0014] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the question-and-answer data determination method described in the first aspect.

[0015] Using the above technical solution, this application provides a method, apparatus, storage medium, and electronic device for determining question-and-answer data. This application acquires multi-dimensional data of a target product and multi-role prompt information corresponding to the multi-dimensional data; generates multi-role question-and-answer data corresponding to the multi-dimensional data based on the multi-role prompt information; performs semantic analysis processing on the multi-role answer data in the multi-role question-and-answer data according to the target question data in the multi-role question-and-answer data to obtain the target answer data corresponding to the target question data; and determines the target question data and the target answer data as the target question-and-answer data corresponding to the multi-dimensional data. The target question-and-answer data is used for storage in the question-and-answer knowledge base corresponding to the target product. Compared with existing technologies, this application achieves coverage of multiple perspectives in question-and-answer data by acquiring multi-dimensional data and multi-role prompts of the target product, and generating multi-role question-and-answer data based on the multi-role prompts, thereby improving the accuracy of question-and-answer data determination. By performing semantic analysis on multi-role answer data based on target question data to obtain target answer data, and storing the target question data and target answer data as target question-and-answer data in the question-and-answer knowledge base, this application achieves structured storage of question-and-answer knowledge related to the target product, improves the efficiency of model processing question-and-answer data, and thus enhances the user experience. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a question-and-answer data determination method provided in an embodiment of this application is shown; Figure 2 A flowchart illustrating a question-and-answer data determination method provided in an embodiment of this application is shown; Figure 3 The diagram shows a flowchart of an automated FAQ construction scheme based on multi-view prompting engineering and hierarchical verification provided in an embodiment of this application. Figure 4 This illustration shows a schematic diagram of a question-and-answer data determination device provided in an embodiment of this application; Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0019] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0020] To address the technical problem in current technologies where increasing the context length of dialogue data leads to a linear increase in GPU memory usage, resulting in memory overflow and impacting model performance, this embodiment provides a method for determining question-and-answer data, such as... Figure 1 As shown, the method includes: Step 101: Obtain multi-dimensional data of the target product and the multi-role prompts corresponding to the multi-dimensional data.

[0021] In this embodiment, the target product can be any product or service involved in the business operations of the enterprise. The target product may include business-related products that require the construction of a Frequently Asked Questions (FAQ) knowledge base. For example, the target product in this embodiment may specifically include automobiles, financial products, home improvement services, etc.

[0022] In this application embodiment, multi-dimensional data can be multi-source heterogeneous data related to the target product. Multi-dimensional data can include structured data with a fixed format, such as product manuals and service contracts, or unstructured data that lacks a fixed format but contains a large amount of effective information, such as customer service call transcripts, product launch video texts, and forum discussion content.

[0023] In this embodiment, the multi-role prompt information can be prompt words used to simulate different user roles or perspectives. The setting of different user roles can be flexibly configured based on the actual user scenarios of the target product. For example, the multi-role prompt information in this embodiment may include prompt information for novice users, prompt information for experienced customers, and prompt information for complaining customers. The prompt information for novice users may focus on guiding the generation of Q&A content related to basic product operations and introductory use. The prompt information for experienced customers may focus on guiding the generation of Q&A content related to product function details and advanced application scenarios. The prompt information for complaining customers may focus on guiding the generation of Q&A content related to product fault handling and liability determination.

[0024] In this embodiment, multi-dimensional data can be obtained by batch collection from multiple data sources using automated tools. After collection, data anonymization processing is required to meet data privacy compliance requirements. Anonymization processing can be carried out by combining rule-based matching and machine learning (such as Named Entity Recognition, NER). Alternatively, other feasible anonymization schemes can be used, such as privacy information masking technology based on rule templates. Anonymization processing can be used to automatically identify and remove or replace sensitive information in the data, such as user ID, transaction amount, contact information, etc.

[0025] In this embodiment of the application, the acquired multi-dimensional data can be used to construct an initial credibility score, providing a quantitative basis for subsequent conflict resolution. Specifically, constructing the initial credibility score can involve assigning an initial credibility score to each piece of data based on the authority of its source. For example, a parameter file published on an official website could have a credibility score of 95, historical real dialogue records could have a credibility score of 75, and user forum posts could have a credibility score of 55, etc. It should be noted that the credibility score can be stored as metadata along with the corresponding data, and can be used for conflict resolution decisions as the data flows.

[0026] Step 102: Generate multi-role question and answer data corresponding to multi-dimensional data based on multi-role prompt information.

[0027] In the embodiments of this application, before generating multi-role question-and-answer data based on multi-role prompt information, the acquired multi-dimensional data can be preprocessed. During the preprocessing process, at least one data segment corresponding to the multi-dimensional data can be determined according to the data category corresponding to the multi-dimensional data (such as product function, service process, fault handling, or other topic categories, or structured and unstructured format categories). When dividing the data segments, they can be divided according to semantic boundaries or chapter topics to ensure that each data segment contains relatively independent and complete knowledge information, and the data segments are saved in a standardized format.

[0028] In this embodiment of the application, after preprocessing, the question-and-answer knowledge content corresponding to each data segment can be determined. Specifically, determining the question-and-answer knowledge content corresponding to each data segment can be achieved by traversing all data segments, inputting the current data segment as context data into a large language model (LLM), and simultaneously inputting knowledge point extraction prompts. For example, the knowledge point extraction prompts in this embodiment of the application may include: You are a rigorous domain knowledge analyst in the [target product's domain], skilled at extracting core technical, process, or conceptual information from professional documents; your output must be a precise and concise structured list, and then the LLM generates a list (K_list) of all independent, important, and question-and-answer-compatible knowledge points in the data segment.

[0029] For example, generating question-and-answer data based on multi-role prompting information can employ an iterative multi-role prompting engineering method. For each data segment, an LLM (Local Level Model) can be iteratively invoked, sequentially inputting the content of the current data segment, the list of knowledge points, and different multi-role prompting information into the model, along with input question-generating prompting instructions. For example, the prompting instructions generated based on the multi-role prompting engineering method in this embodiment may include: "Please play the [current role], carefully read the context, and ask five original questions about the [knowledge point] that are most likely to be asked." The questions must represent a real consultation scenario.

[0030] For example, after generating question data, corresponding answer data can be generated and aligned. At this point, LLM can be invoked and preset answer generation prompts can be entered, such as asking an authoritative domain expert to provide a detailed and accurate standard answer to the following question, strictly based on the context. The answer must reflect the original source ID. LLM can generate corresponding answer data based on data fragments and question data, and associate and align question data, answer data, original source ID, and initial trust score to form multi-role question-and-answer data. Each multi-role question-and-answer data can contain fields such as {"question":"…","answer":"…","source_id":"…","trust_score":"[initial trust score]"}.

[0031] Step 103: Based on the target question data in the multi-role question and answer data, perform semantic analysis on the multi-role answer data in the multi-role question and answer data to obtain the target answer data corresponding to the target question data.

[0032] In this embodiment of the application, the target problem data can be a standardized problem selected from each cluster of problems, and the target problem data can cover the core requirements of similar problems in the problem cluster.

[0033] In this embodiment of the application, the target answer data may be answer data that matches the target question data, obtained by filtering or fusing multi-role answer data.

[0034] In this embodiment of the application, semantic analysis processing can be a process of parsing the semantic connotation of text content, extracting key information, and clarifying the logical relationships between information through technical means.

[0035] In this embodiment, semantic analysis can be based on LLM to parse multi-role answer data. The selected LLM can preferably be a model with strong reasoning and information extraction capabilities. Semantic analysis can input each target question data and its corresponding multi-role answer data into the LLM. After LLM analysis, a set of semantic information corresponding to the multi-role answer data can be obtained. The target answer data is determined based on the set of semantic information and the multi-role answer data.

[0036] Step 104: Determine the target question data and target answer data as the target question and answer data corresponding to the multi-dimensional data.

[0037] The target question-and-answer data is used to store the question-and-answer knowledge base corresponding to the target product.

[0038] In this embodiment, to ensure that the quality of the target question data meets the requirements for use in the knowledge base, semantic matching verification can also be performed on the target question data. Semantic matching verification can be an automated verification step through LLM cross-checking to achieve a multi-level quality verification mechanism. Specifically, semantic matching verification can input the target question-and-answer data and the corresponding original multi-dimensional data (data fragments) into an independent LLM instance, and input a preset factual consistency verification prompt, such as "Please check, sentence by sentence, whether the key information (numbers, dates, processes, product models) in the answer can be supported by the original document based on the original knowledge source document ([original document content]) and the final answer ([final answer A]). Please output 'factual consistency' or 'factual conflict'." The LLM can then verify the factual accuracy of the target question-and-answer data sentence by sentence to determine the semantic matching degree between the target question-and-answer data and the original multi-dimensional data.

[0039] For example, if the target question-and-answer data matches the multi-dimensional data, i.e., the LLM cross-check determines that the facts are consistent and the credibility score is higher than the threshold, then the target question-and-answer data is stored in the question-and-answer knowledge base corresponding to the target product. If the LLM cross-check determines that the facts are conflicting or the credibility score is lower than the preset threshold, then the target question-and-answer data is automatically marked as a high-risk question-and-answer set and pushed to domain experts for targeted manual review and revision. If the experts confirm that the question-and-answer data is correct and meets the requirements after review, it is stored in the knowledge base. If there are errors, it is returned to the corresponding stage for reprocessing.

[0040] Compared with existing technologies, the embodiments of this application acquire multi-dimensional data and multi-role prompts of the target product, and generate multi-role question-and-answer data based on the multi-role prompts, thereby achieving coverage of multi-role perspectives in the question-and-answer data and improving the accuracy of question-and-answer data determination. By performing semantic analysis on the multi-role answer data based on the target question data to obtain the target answer data, and storing the target question data and target answer data as target question-and-answer data in the question-and-answer knowledge base, the structured storage of question-and-answer knowledge related to the target product is achieved, improving the efficiency of the model in processing question-and-answer data, and thus enhancing the user experience.

[0041] As an optional approach, when performing semantic analysis on the multi-role answer data in the multi-role question-and-answer data based on the target question data in the multi-role question-and-answer data to obtain the target answer data corresponding to the target question data, the following methods can be used, but are not limited to these: Figure 2 As shown, it includes: Step 201: Cluster the multi-role question data in the multi-role question-and-answer data to obtain the target question data.

[0042] In this embodiment of the application, multi-role problem data can be problem data generated based on multi-dimensional data and multi-role prompt information, corresponding to the perspectives of different user roles. Multi-role problem data can be problem data with similar semantics but different expression forms due to differences in role perspectives (such as novice users focusing on basic operations, experienced users focusing on detailed functions, and complaining users focusing on troubleshooting).

[0043] For the embodiments of this application, clustering of multi-role question data in multi-role question-answering data can be performed using hybrid clustering. The hybrid clustering process can be divided into three stages: The first stage is question vectorization, which can be performed using an embedding model (such as the Qwen3-Embedding-4B model) to process all multi-role question data, converting each question into a computer-recognizable vector form. The vector dimension can be determined according to the model characteristics (such as 768 dimensions). During the question transformation process, the question text can be preprocessed first (such as removing meaningless punctuation, unifying character capitalization, and filtering common stop words). It should be noted that, in addition to the above model, other embedding models with semantic representation capabilities can also be selected, as long as they can achieve effective transformation of question semantics.

[0044] In the embodiments of this application, the second stage of hybrid clustering can be preliminary clustering, in which the K-Means algorithm can be used to group the question vectors. If the distribution of multi-role question data is relatively scattered, alternative solutions such as DBSCAN algorithm and hierarchical clustering algorithm can also be selected. For example, when using hierarchical clustering, the cosine similarity between all question vectors can be calculated first, and then the most similar question clusters can be gradually merged based on the similarity until the average semantic similarity of the questions within the cluster reaches the threshold.

[0045] In this embodiment of the application, the third stage of hybrid clustering can be LLM fine clustering. Since the preliminary clustering is only based on vector similarity grouping, it may not be able to fully capture the deep semantic relationships of the issues. Therefore, LLM needs to be called for further optimization. Specifically, LLM fine clustering can input each issue cluster formed by the preliminary clustering into LLM, along with preset fine clustering prompts. LLM will identify the semantic equivalence between issues based on its understanding of business knowledge, merge redundant issues, and generate target issue data.

[0046] Step 202: Perform semantic analysis on the multi-role answer data based on the target question data to obtain the semantic information set corresponding to the multi-role answer data.

[0047] In this embodiment of the application, the multi-role answer data can be the answer data corresponding to the multi-role question data, and the multi-role answer data can have an original data source ID and an initial credibility score.

[0048] In this embodiment, the semantic information set can be a collection of key information extracted from all multi-role answer data through semantic analysis. For example, the semantic information set in this embodiment may specifically include, but is not limited to, numbers, dates, process steps, concept definitions, and responsibility delineation standards.

[0049] For the embodiments of this application, LLM can first clarify the core requirements of the target question data (such as what the application process of the service is, and the corresponding core requirements are the process steps), and then match the corresponding multi-role answer data. Matching the corresponding multi-role answer data can be done by filtering out the answer data that is directly related to the core requirements from each answer, removing irrelevant statements and extracting the key information of the answer data.

[0050] Step 203: Determine the target answer data based on the semantic information set and multi-role answer data.

[0051] In this embodiment, the target answer data can be determined by analyzing the semantic information set and combining the attributes of the multi-role answer data (such as the initial credibility score). The determination of the target answer data can be divided into two cases: when there is semantic conflict in the semantic information set, the multi-role answer data with the highest initial credibility score is selected as the target answer data based on the credibility data (i.e., the initial credibility score) corresponding to each of the multi-role answer data; when there is no semantic conflict in the semantic information set, the multi-role answer data is fused to generate the target answer data.

[0052] Optionally, when performing the "determining target answer data based on semantic information set and multi-role answer data", the following methods can be used, but are not limited to these: when there is semantic conflict in the semantic information set, the target answer data is determined from the multi-role answer data based on the credibility data corresponding to the multi-role answer data respectively; when there is no semantic conflict in the semantic information set, the multi-role answer data is fused to generate the target answer data.

[0053] In this embodiment of the application, when there is a semantic conflict in the semantic information in the semantic information set, the best option can be selected based on the initial credibility score of the multi-role answer data. The answer with the highest score can be determined as the target answer data by comparing the initial credibility scores of the conflicting answers. At the same time, the conflicting answers with lower scores can be marked as conflict pending review so that they can be manually reviewed if needed later.

[0054] In this embodiment, when there is no semantic conflict in the semantic information set, the semantic conflict may be a factual contradiction or a procedural contradiction in the key information corresponding to different answers. If the LLM determines that all key information is consistent and without contradiction through analysis of the semantic information set, then the multi-role answer data needs to be fused to generate the target answer data. The fusion process can call the LLM with content generation capabilities to input the target question data, all multi-role answer data and the semantic information set into the model, along with fusion prompts. The LLM will integrate the effective content of each answer based on the key information in the semantic information set to form a comprehensive and refined target answer data.

[0055] It should be noted that, in addition to directly comparing credibility scores, alternative conflict resolution algorithms can also be used, such as the game-theoretic fusion method based on DS evidence theory (which converts credibility scores into probability assignments and calculates the probability of the optimal answer through synthesis rules) or the weighted fusion method based on information entropy (which assigns weights according to the uncertainty of the answer information, and the answer with the higher weight is adopted first).

[0056] Optionally, when performing the step of "determining the target question data and target answer data as the target question-and-answer data corresponding to the multi-dimensional data", the following methods can be used, but are not limited to: performing semantic matching between the target question-and-answer data and the multi-dimensional data; and storing the target question-and-answer data in the question-and-answer knowledge base if the matching degree between the target question-and-answer data and the multi-dimensional data meets the matching conditions.

[0057] In this embodiment, the semantic matching process can be to determine the semantic consistency between the target question-and-answer data and the multi-dimensional data. Specifically, semantic matching can be performed by calling an independent LLM instance to perform fact consistency verification. The target question-and-answer data and the corresponding original data fragment content are passed into the LLM model, along with matching degree judgment prompts. For example, based on the original data fragment [data fragment content], please check sentence by sentence whether the key information (numbers, dates, processes, definitions) in the target answer data [target answer data content] has the following conditions: complete match (the original data explicitly supports it); partial match (the original data has relevant statements but requires reasonable deduction); mismatch (the original data does not support it or there is a contradiction).

[0058] In this embodiment, the matching condition can be a condition for determining whether the target question-and-answer data meets the storage requirements. If the target question-and-answer data meets the matching condition, it can be directly stored in the question-and-answer knowledge base corresponding to the target product. The storage format of the knowledge base can be structured. If the target question-and-answer data does not meet the matching condition (e.g., there is mismatched information or the proportion of complete matching is too low), it can be marked as a high-risk question-and-answer set and pushed to domain experts for targeted manual review. The experts conducting the manual review can combine the original data sharding and the verification details of LLM to determine whether there is a deviation in the generation process of the target question-and-answer data (e.g., LLM analysis error) or ambiguity in the original data. If the target question-and-answer data has a generation deviation, it returns to the corresponding step for reprocessing. If the target question-and-answer data has ambiguity in the original data, the data needs to be clarified before generating the target question-and-answer data.

[0059] Optionally, after performing the step of "storing the target question-and-answer data in the question-and-answer knowledge base", the following methods can be used, but are not limited to: generating an optimization strategy for the target question-and-answer data based on feedback data of the target question-and-answer data; and adjusting the target question-and-answer data according to the optimization strategy to optimize the question-and-answer data in the question-and-answer knowledge base.

[0060] In the embodiments of this application, feedback data can be data related to the target question and answer data generated by the user during the use of the question and answer knowledge base. Feedback data can include the user's satisfaction rating of the question and answer results, the user's error correction suggestions, the user's unanswered extended questions, the retrieval frequency of the target question and answer data, etc. Feedback data can be automatically collected through the knowledge base's usage interface (such as a telemarketing system or customer service consultation platform).

[0061] In the embodiments of this application, the collected feedback data can be analyzed using machine learning algorithms (such as classification algorithms and regression algorithms). Through analysis, optimization strategies for the target question-and-answer data can be obtained. Based on the optimization strategies, the target question-and-answer data can be adjusted, thereby optimizing the question-and-answer data in the question-and-answer knowledge base.

[0062] Optionally, when performing the "generating multi-role question-and-answer data corresponding to multi-dimensional data based on multi-role prompt information", the following method can be used, but is not limited to: determining at least one data segment corresponding to the multi-dimensional data according to the data category corresponding to the multi-dimensional data; determining the question-and-answer knowledge content corresponding to at least one data segment; generating question-and-answer data based on the question-and-answer knowledge content based on the multi-role prompt information to obtain multi-role question-and-answer data.

[0063] In this embodiment, the data category can be a type of multi-dimensional data categorized according to knowledge attributes. For example, the data categories in this embodiment may specifically include product function categories, service process categories, policy and regulation categories, fault handling categories, etc.

[0064] In this embodiment of the application, determining data fragments by data category can integrate content belonging to the same category in multi-dimensional data into one or more data fragments. Each data fragment focuses on a single knowledge domain, and the data fragments can be converted into a standardized format (such as {"source_id":"data source identifier","content":"specific content of data fragment","data_type":"data category"}).

[0065] For the embodiments of this application, the question-and-answer knowledge content corresponding to each data shard (i.e., the knowledge in the data shard that can be used to generate question-and-answer pairs) can be determined by calling LLM to pass the data shard content into the model, along with knowledge point extraction prompts. For example, if you are a knowledge analyst for the [domain corresponding to the data category], please extract all independent knowledge points that can be used to generate question-and-answer pairs from the following data shards. Each knowledge point needs to clearly define its core knowledge (such as the operation method of a certain function or the validity period of a certain policy), and output a structured list of knowledge points (including knowledge point ID, knowledge point core, and specific content of the knowledge point).

[0066] For the embodiments of this application, the generation of multi-role question-and-answer data based on multi-role prompt information can adopt an iterative generation strategy: for each data segment's knowledge point list, LLM is called one by one, with each input being a knowledge point, a type of multi-role prompt information, and data segment content, along with a question generation prompt instruction (e.g., please play the [current role], and based on the knowledge point [core knowledge point], combine the data segment content to propose 4-5 real consultation questions that conform to the perspective of that role; the questions must be specific and unambiguous), generating question data for the corresponding role; subsequently, for each generated question data, LLM is called again, with the question, data segment content, and source ID input, along with an answer generation prompt instruction, to generate answer data; finally, the question data, answer data, source ID, initial credibility score, and role identifier can be aligned to form multi-role question-and-answer data.

[0067] As an optional approach, embodiments of this application also provide an automated FAQ construction scheme based on multi-perspective prompting engineering and hierarchical verification, such as... Figure 3 As shown, Figure 3 This paper presents a workflow for an automated FAQ construction scheme based on multi-perspective prompting engineering and hierarchical verification. Figure 3 This approach is designed for general enterprise-level intelligent telemarketing and customer service scenarios. Based on an LLM knowledge enhancement generation architecture, it uses multi-perspective prompting engineering and a hybrid quality verification mechanism to automatically, efficiently, and with high quality construct frequently asked questions (FAQs) from multi-source heterogeneous data. The steps include: Step 1: Data preprocessing and credibility scoring (i.e., the multi-dimensional data in the embodiments of this application); Step 2: Automated FAQ extraction based on multi-perspective question generation (i.e., generating multi-dimensional question and answer data corresponding to multi-role prompt information based on multi-role information in the implementation of this application); Step 3: FAQ standardization, deduplication, and conflict resolution (i.e., in this embodiment, semantic analysis is performed on the multi-role answer data in the multi-role question-and-answer data based on the target question data in the multi-role question-and-answer data to obtain the target answer data corresponding to the target question data). Step 4: Knowledge Base Application and Feedback (i.e., based on the feedback data of the target question-and-answer data in this embodiment of the application, an optimization strategy for the target question-and-answer data is generated; the target question-and-answer data is adjusted according to the optimization strategy to optimize the question-and-answer data in the question-and-answer knowledge base.)

[0068] Compared with existing technologies, the embodiments of this application improve the matching accuracy between target answer data and target question data by determining target answer data based on semantic information sets and multi-role answer data; by fusing multi-role answer data to determine target answer data when semantic information conflicts and when semantic information does not conflict, reasonable acquisition of target answer data under different semantic scenarios is achieved; by semantically matching target question and answer data with multi-dimensional data and storing it in the question and answer knowledge base when the matching degree meets the conditions, the quality of target question and answer data stored in the question and answer knowledge base is improved; by generating optimization strategies based on feedback data of target question and answer data and adjusting target question and answer data according to optimization strategies, continuous optimization of question and answer data in the question and answer knowledge base is achieved.

[0069] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides a question-and-answer data determination device, such as... Figure 4 As shown, the device includes: an acquisition module 31, a generation module 32, a processing module 33, and a determination module 34.

[0070] The acquisition module 31 is configured to acquire multi-dimensional data of the target product and multi-role prompt information corresponding to the multi-dimensional data; The generation module 32 is configured to generate multi-role question and answer data corresponding to multi-dimensional data based on multi-role prompt information; Processing module 33 is configured to perform semantic analysis on multi-role answer data in multi-role question-and-answer data based on target question data in multi-role question-and-answer data, so as to obtain target answer data corresponding to target question data; The determination module 34 is configured to determine the target question data and target answer data as target question and answer data corresponding to multi-dimensional data. The target question and answer data is used to store in the question and answer knowledge base corresponding to the target product.

[0071] In some examples of this embodiment, the processing module 33 is specifically configured to cluster the multi-role question data in the multi-role question-answer data to obtain target question data; perform semantic analysis on the multi-role answer data based on the target question data to obtain a set of semantic information corresponding to the multi-role answer data; and determine the target answer data based on the set of semantic information and the multi-role answer data.

[0072] In some examples of this embodiment, the processing module 33 is further configured to determine the target answer data from the multi-role answer data based on the credibility data corresponding to the multi-role answer data when there is a semantic conflict in the semantic information set; and to perform fusion processing on the multi-role answer data to generate the target answer data when there is no semantic conflict in the semantic information set.

[0073] In some examples of this embodiment, the determining module 34 is specifically configured to perform semantic matching between the target question-and-answer data and the multi-dimensional data; and to store the target question-and-answer data in the question-and-answer knowledge base if the matching degree between the target question-and-answer data and the multi-dimensional data meets the matching conditions.

[0074] In some examples of this embodiment, the determining module 34 is further configured to generate an optimization strategy for the target question-and-answer data based on the feedback data of the target question-and-answer data; and to adjust the target question-and-answer data according to the optimization strategy in order to optimize the question-and-answer data in the question-and-answer knowledge base.

[0075] In some examples of this embodiment, the generation module 32 is specifically configured to determine at least one data segment corresponding to the multi-dimensional data according to the data category corresponding to the multi-dimensional data; determine the question-and-answer knowledge content corresponding to the at least one data segment; and generate question-and-answer data based on the question-and-answer knowledge content using multi-role prompt information to obtain multi-role question-and-answer data.

[0076] It should be noted that other corresponding descriptions of the functional units involved in the question-and-answer data determination device provided in this embodiment can be found in [reference]. Figure 1 and Figure 2 The corresponding descriptions in [the document] will not be repeated here.

[0077] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method shown.

[0078] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0079] like Figure 5 The diagram shown is a hardware structure schematic of an electronic device according to the present invention, comprising: At least one processor 401; and, Memory 402 is communicatively connected to at least one processor 401; wherein, The memory 402 stores instructions that can be executed by at least one processor to enable the at least one processor to perform the question-and-answer data determination method as described above.

[0080] Figure 5 Take a processor 401 as an example.

[0081] The electronic device may also include an input device 404 and an output device 404.

[0082] The processor 401, memory 402, input device 403, and output device 404 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0083] Memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the question-and-answer data determination method in the embodiments of this application, for example, Figure 1 and Figure 2 The method flow is shown. The processor 401 executes various functional applications and communications by running non-volatile software programs, instructions, and modules stored in the memory 402, thereby implementing the question-and-answer data determination method in the above embodiments.

[0084] Memory 402 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the question-and-answer data determination method, etc. Furthermore, memory 402 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 402 may optionally include memory remotely located relative to processor 401, and these remote memories may be connected via a network to the apparatus performing the question-and-answer data determination method. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0085] The input device 403 can receive user clicks and generate signal inputs related to user settings and function control for question-and-answer data determination methods. The output device 403 may include a display device such as a screen.

[0086] One or more modules are stored in memory 402, and when run by one or more processors 401, the question-and-answer data determination method in any of the above method embodiments is executed.

[0087] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0088] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0089] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0090] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the solution of this embodiment, compared with the existing technology, this application embodiment obtains multi-dimensional data and multi-role prompt information of the target product, and generates multi-role question-and-answer data based on the multi-role prompt information, realizing the coverage of question-and-answer data from multiple perspectives and improving the accuracy of question-and-answer data determination; by performing semantic analysis processing on multi-role answer data based on target question data to obtain target answer data, and determining the target question data and target answer data as target question-and-answer data and storing them in the question-and-answer knowledge base, the structured storage of question-and-answer knowledge related to the target product is realized, improving the efficiency of model processing question-and-answer data, and thus improving the user experience; by using semantic information... The system uses information sets and multi-role answer data to determine target answer data, improving the matching accuracy between target answer data and target question data. By fusing multi-role answer data when semantic information conflicts and when there are no semantic conflicts, the system achieves reasonable acquisition of target answer data under different semantic scenarios. By semantically matching target question and answer data with multi-dimensional data and storing it in the question and answer knowledge base when the matching degree meets the conditions, the system improves the quality of target question and answer data stored in the question and answer knowledge base. By generating optimization strategies based on feedback data of target question and answer data and adjusting target question and answer data according to the optimization strategies, the system achieves continuous optimization of question and answer data in the question and answer knowledge base.

[0091] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0092] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for determining question-and-answer data, characterized in that, include: Obtain multi-dimensional data of the target product and multi-role prompts corresponding to the multi-dimensional data; Based on the multi-role prompting information, multi-role question-and-answer data corresponding to the multi-dimensional data is generated; Based on the target question data in the multi-role question and answer data, semantic analysis processing is performed on the multi-role answer data in the multi-role question and answer data to obtain the target answer data corresponding to the target question data. The target question data and the target answer data are determined as the target question-and-answer data corresponding to the multi-dimensional data, and the target question-and-answer data is used to store in the question-and-answer knowledge base corresponding to the target product.

2. The method according to claim 1, characterized in that, The step of performing semantic analysis on the multi-role answer data in the multi-role question-and-answer data based on the target question data in the multi-role question-and-answer data to obtain the target answer data corresponding to the target question data includes: The target question data is obtained by clustering the multi-role question data in the multi-role question-and-answer data; Based on the target question data, semantic analysis is performed on the multi-role answer data to obtain a set of semantic information corresponding to the multi-role answer data; The target answer data is determined based on the semantic information set and the multi-role answer data.

3. The method according to claim 2, characterized in that, The process of determining the target answer data based on the semantic information set and the multi-role answer data includes: In the event of semantic conflicts among the semantic information in the semantic information set, the target answer data is determined from the multi-role answer data based on the credibility data corresponding to each of the multi-role answer data. If there are no semantic conflicts in the semantic information set, the multi-role answer data is fused to generate the target answer data.

4. The method according to claim 1, characterized in that, The step of determining the target question data and the target answer data as the target question-and-answer data corresponding to the multi-dimensional data includes: Semantic matching is performed between the target question-and-answer data and the multi-dimensional data; If the target question-and-answer data matches the multi-dimensional data, the target question-and-answer data is stored in the question-and-answer knowledge base.

5. The method according to claim 4, characterized in that, After storing the target question-and-answer data in the question-and-answer knowledge base, the method further includes: Based on the feedback data of the target question-and-answer data, an optimization strategy for the target question-and-answer data is generated. The target question-and-answer data is adjusted according to the optimization strategy to optimize the question-and-answer data in the question-and-answer knowledge base.

6. The method according to claim 1, characterized in that, The process of generating multi-role question-and-answer data corresponding to the multi-dimensional data based on the multi-role prompt information includes: Based on the data categories corresponding to the multi-dimensional data, at least one data slice corresponding to the multi-dimensional data is determined; Determine the question-and-answer knowledge content corresponding to the at least one data segment; Based on the multi-role prompting information, question-and-answer data is generated from the question-and-answer knowledge content to obtain the multi-role question-and-answer data.

7. A question-and-answer data determination device, characterized in that, include: The acquisition module is configured to acquire multi-dimensional data of the target product and multi-role prompt information corresponding to the multi-dimensional data. The generation module is configured to generate multi-role question-and-answer data corresponding to the multi-dimensional data based on the multi-role prompt information; The processing module is configured to perform semantic analysis on the multi-role answer data in the multi-role question-and-answer data based on the target question data in the multi-role question-and-answer data, so as to obtain the target answer data corresponding to the target question data; The determination module is configured to determine the target question data and the target answer data as the target question-and-answer data corresponding to the multi-dimensional data, and the target question-and-answer data is used to store in the question-and-answer knowledge base corresponding to the target product.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.