AI-assisted large-model intelligent question and answer accuracy improvement method

By building a large-scale voice corpus, dynamic knowledge base and pre-trained language model in the intelligent question-and-answer system, combining logical reasoning and machine learning, optimizing dialogue strategies and sentiment analysis, improving speech recognition and multilingual search capabilities, the shortcomings of the intelligent question-and-answer system in natural language understanding, knowledge representation and update, reasoning capabilities, and improving high accuracy and user satisfaction.

CN120045660APending Publication Date: 2025-05-27CCID CONSULTING CO LTD
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510082010.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing intelligent question-and-answer system has shortcomings in natural language understanding, knowledge representation and update, reasoning ability, language recognition, dialogue ability and knowledge base update mechanism, resulting in low accuracy and user satisfaction.

Method used

By collecting data from multiple sources, establishing a large-scale voice corpus, performing data preprocessing and annotating, building a dynamic knowledge base, introducing external knowledge sources, regularly updating the knowledge base, using large-scale diversified corpus to improve the pre-trained language model, combining logical reasoning and machine learning for knowledge reasoning, optimizing dialogue strategies and sentiment analysis, and improving speech recognition algorithms and multilingual search capabilities.

Benefits of technology

It has achieved the intelligent Q&A system with strong understanding ability, timely data updates, outstanding reasoning ability, strong language recognition and dialogue ability, significantly improving the accuracy of Q&A and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045660A_ABST
    Figure CN120045660A_ABST
Patent Text Reader

Abstract

The invention provides an AI-assisted large-model intelligent question and answer accuracy improvement method, and belongs to the technical field of intelligent question and answer. Comprising the steps of data collection, data cleaning, data labeling, dynamic knowledge base construction, knowledge base updating trigger mechanism establishment, pre-training language model improvement, post-processing language understanding enhancement, combination of rule-based and machine learning reasoning, reasoning result verification and optimization, dialogue strategy optimization, emotion perception and response, multi-language search and language recognition search. And multi-dimensional evaluation index setting and regular evaluation and feedback improvement are realized. The intelligent question and answer accuracy improving method is high in understanding ability, timely in data updating, outstanding in reasoning ability and high in language recognition and dialogue ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent question answering technology, and in particular to a method for improving the accuracy of large-model intelligent question answering assisted by AI. Background Art

[0002] With the continuous development of artificial intelligence, intelligent question-answering technology has gradually matured, but the existing intelligent question-answering system still has the following problems:

[0003] Technical aspects:

[0004] Difficulty understanding natural language: When users input questions, there may be incomplete descriptions, redundancy, word errors, etc. Most existing systems do not handle them in a targeted manner, resulting in biased or incorrect understanding of the questions and returning incorrect answers. In addition, natural language itself is complex, such as polysemy and semantic ambiguity, making it difficult for the system to accurately understand the user's intentions.

[0005] Knowledge representation and update issues: The information coverage of the knowledge base is limited and it is difficult to meet the needs of all natural language questions. At the same time, the knowledge base needs to be constantly updated to adapt to changes in knowledge, but the current knowledge base update mechanism is not perfect enough, resulting in some knowledge being outdated.

[0006] Limited reasoning ability: In intelligent question-answering systems, not all questions can be answered directly using the existing knowledge base. The system also needs to have certain reasoning capabilities. However, current knowledge reasoning technology is not mature enough to handle complex reasoning tasks.

[0007] Application level:

[0008] Poor language recognition ability: Most intelligent question-answering systems have poor language recognition ability, which is manifested in insufficient text reading ability, lack of accurate recognition ability, and difficulties in recognizing synonymous questions.

[0009] Lack of conversational capabilities: Most intelligent question-answering systems do not have conversational capabilities. They lack intelligent functions, are unable to make effective responses based on the context, and are unable to engage in humanized interactions.

[0010] Lack of basic functions: Most intelligent question-answering systems lack basic functions. For example, the proportion of question-answering systems with voice input search capabilities and the ability to recognize English questions is very low. Some intelligent question-answering systems are also unable to handle Chinese input problems well.

[0011] Insufficient knowledge base data and untimely updates: Some knowledge bases lack data, have chaotic structures, and are not updated in a timely manner, resulting in the question-and-answer system being unable to meet user needs well.

[0012] Evaluation level:

[0013] The quality of benchmark tests is not high: There are obvious errors in many problems in the existing datasets, the methods for constructing benchmark tests are also defective, and the evaluation indicators are single.

[0014] Therefore, there is an urgent need in this field for a technical solution that can solve one or more of the above problems.

[0015] The information disclosed in this background section is only intended to enhance the overall understanding of the present invention and should not be regarded as an admission or any form of suggestion that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention

[0016] The object of the present invention is to provide a method for improving the accuracy of intelligent question answering with strong understanding ability, timely data update, outstanding reasoning ability, strong language recognition and dialogue ability.

[0017] To achieve the above object, the present invention provides the following solutions:

[0018] A method for improving the accuracy of large model intelligent question answering assisted by AI, comprising:

[0019] Collect data from multiple sources;

[0020] For data collection related to speech recognition search, establish a large-scale speech corpus covering speech data in different accents, speech rates, and background noise environments;

[0021] Perform data preprocessing, including:

[0022] Remove duplicate data, compare the collected data, and identify and delete duplicate text or speech records;

[0023] Filter low-quality data. For speech data, exclude data with serious noise interference that makes it impossible to accurately identify the content;

[0024] For speech data, annotate the text, accent type, and speech rate corresponding to the speech content, and control the quality of annotation through cross-validation and expert review;

[0025] Based on knowledge graph technology, construct a dynamic knowledge base;

[0026] Introduce external knowledge sources, regularly update the knowledge in the knowledge base, and at the same time verify the new knowledge to ensure its accuracy;

[0027] Set update trigger conditions based on data heat and timeliness;

[0028] When there is a conflict between new data and the knowledge in the existing knowledge base, trigger the update process of the knowledge base, and use methods of logical reasoning and evidence weight evaluation to determine whether to accept the new knowledge, avoiding the introduction of incorrect knowledge;

[0029] Use large-scale and diverse corpora for pre-training;

[0030] Introduce semantic role labeling and dependency parsing tasks as auxiliary tasks in the pre-trained model, and conduct them simultaneously with the pre-training of the language model;

[0031] For users' questions, conduct multi-level semantic analysis; first, conduct lexical analysis to identify the part of speech and morphological changes of each word; then, conduct syntactic analysis to construct the syntactic tree of the sentence and determine the sentence structure; finally, conduct semantic analysis to understand the semantic intention of the sentence based on the pre-trained semantic knowledge base and the results of semantic role labeling;

[0032] Adopt a multi-round dialogue strategy to clarify the user's intention. When there is ambiguity in the initial analysis of the user's question, obtain more accurate user intention through follow-up questions;

[0033] Establish a reasoning system based on logical rules. At the same time, use machine learning techniques to learn a large number of knowledge instances and automatically discover hidden knowledge relationships and reasoning patterns;

[0034] Verify the results obtained from reasoning. By comparing with the existing knowledge in the knowledge base and verifying in a large-scale dataset, if there is a conflict between the reasoning results and the known knowledge or the verification accuracy in the dataset is low, then optimize the reasoning process, adjust the rules or retrain the machine learning model;

[0035] Establish a dialogue state tracking mechanism to record the historical information of the dialogue, including the user's questions, the system's answers and follow-up questions; determine the answering strategy for the next round according to the dialogue state;

[0036] Adopt a dialogue strategy learning algorithm to learn the best dialogue strategies in different scenarios from a large amount of dialogue data;

[0037] Add sentiment analysis to the dialogue, identify the sentiment tendency in the user's question, and adjust the tone and content of the answer according to the sentiment tendency;

[0038] Construct a multilingual index. In the search engine, establish independent indexes for text data in different languages and establish mapping relationships between languages;

[0039] Develop a multilingual semantic matching algorithm. Using cross-lingual word vector technology, map words in different languages to the same semantic space, so as to achieve semantic matching between different languages;

[0040] Improve the speech recognition algorithm by using a deep neural network to enhance the accuracy of speech recognition;

[0041] Seamlessly connect the speech recognition results with the semantic understanding module;

[0042] Incorporate semantic understanding accuracy metrics, measured through a combination of manual and automatic evaluations;

[0043] Set dialogue coherence metrics to evaluate the coherence between responses and questions and the logic between responses in multi-turn conversations;

[0044] For speech recognition searches, set speech recognition accuracy metrics and speech search response metrics;

[0045] Regularly evaluate the performance of the large model, and based on the results of the comprehensive evaluation index system, identify weak performance areas;

[0046] Feed the evaluation results back into each component of the model for targeted optimization.

[0047] Optionally, the establishment of a large-scale speech corpus specifically includes:

[0048] Determine the collection sources, recruit volunteers from different regions to ensure a variety of accents are covered; let the number of regions be n, and the planned number of samples to be collected in each region be m i (i = 1, 2, ···, n), then the total number of collected samples

[0049] Control the speed variation, design different text contents, and let the volunteers read at different speeds; quantify the speed according to the number of words per minute, and let the normal speed be s 0 , the collected speed range is [s min , s max , and samples are collected at regular intervals;

[0050] Use noise synthesis algorithms to simulate background noise environments;

[0051] Perform data preprocessing through endpoint detection algorithms and pre-emphasis algorithms; among them,

[0052] The endpoint detection algorithm is an energy-based endpoint detection algorithm. Let the speech signal be x(n), and the energy By setting an energy threshold T, when E > T, it is determined as the start or end of a speech segment;

[0053] The pre-emphasis algorithm uses the formula y(n) = x(n) - αx(n - 1), where α ranges from 0 to 0.9, x(n) is the original speech signal, and y(n) is the pre-emphasized speech signal;

[0054] After pre-emphasis, the speech signal is divided into short-time frames, and Fourier transform is performed on the short-time frames. By connecting adjacent frames, a good approximation of the signal frequency profile is obtained;

[0055] Multiply each frame by a window function;

[0056] Then, features are extracted by calculating Mel-frequency cepstral coefficients and linear predictive coding coefficients, specifically:

[0057] Perform an N-point FFT on each framed and windowed signal to calculate the spectrum, where N is 256 or 512. After obtaining the magnitude of the FFT, calculate the power spectrum through a formula;

[0058] Calculate the Mel filter bank, and extract frequency bands from the power spectrum through a set of Mel-scale triangular filters;

[0059] Apply the discrete cosine transform to decorrelate the filter bank coefficients and generate a compressed representation of the filter bank;

[0060] Analyze the speech signal through linear prediction, model the speech signal to obtain linear prediction coefficients, and further calculate to obtain linear predictive cepstral coefficients;

[0061] Perform feature fusion, fuse the extracted Mel-frequency cepstral coefficient features and linear predictive cepstral coefficient features to form a new feature vector as the final speech feature vector;

[0062] Let the speech feature vector be v, and map it to a storage location through the hash function h(v); the hash function is: where v i is an element of the feature vector, k is the vector dimension, and m is the number of buckets;

[0063] According to the occurrence frequency f i of different feature values in the speech data, construct a binary tree such that the weighted path length of the tree is minimized, where l i is the coding length of the feature value, thereby realizing the efficient compressed storage of speech data and obtaining a speech corpus.

[0064] Optionally, constructing a dynamic knowledge base based on the knowledge graph technology specifically includes:

[0065] Adopt a method combining rules and machine learning to extract the relationships between entities, and at the same time use the convolutional neural network in deep learning to train the preprocessed data to improve the accuracy and coverage of relationship extraction; for newly emerging entity relationship types, discover them through unsupervised learning algorithms;

[0066] Use graph embedding technology to map entities and relationships in the knowledge graph to a low-dimensional vector space. During the representation learning process, semantic information is added to improve the semantic expression ability of the knowledge graph; adopt the multi-modal knowledge graph representation learning method. When there is multi-modal data, fuse information of different modalities into the representation of the knowledge graph, and jointly map its image features and text description features to a low-dimensional vector space;

[0067] When constructing a knowledge graph from multiple data sources, perform knowledge fusion, align entities with the same name from different sources, and match them through the unique identifier and attribute information of the entities. For the case where there are differences in the attribute values of entities, adopt a credibility evaluation method to determine the final attribute value according to factors such as the reliability of the data source and the data update frequency; integrate knowledge graphs in different domains and establish cross-domain entity relationships to complete the construction of a dynamic knowledge base.

[0068] Optionally, the method of using logical reasoning and evidence weight evaluation to determine whether to accept new knowledge specifically includes:

[0069] When encountering new knowledge, first decompose it into basic propositions and statements;

[0070] Analyze the logical relationships between these propositions;

[0071] Use deductive reasoning and inductive reasoning to test these logical relationships;

[0072] Evaluate the source, diversity, and consistency of evidence;

[0073] Set a threshold for accepting new knowledge according to the rigor of logical reasoning and the evaluation result of evidence weight;

[0074] If the logical relationship is complete, and the evidence mainly comes from authoritative sources and is diverse and consistent, then when these conditions are met by more than 80%, decide to accept the new knowledge.

[0075] Optionally, the method of adopting a dialogue strategy learning algorithm to learn the best dialogue strategies in different scenarios from a large amount of dialogue data specifically includes:

[0076] Data screening and classification, using semantic understanding technology to perform a preliminary scan of a large amount of dialogue data; identify dialogue groups with similar or relevant semantics; for the preliminarily classified dialogue groups, further subdivide them according to the specific characteristics of the scenario;

[0077] Construct a dialogue graph representation, regard each round of dialogue as a node, and construct a directed edge between the nodes according to the order of the dialogue. The weight of the edge is calculated according to the coherence and logical relationship of the dialogue;

[0078] Based on neural network learning, select a graph neural network model suitable for processing graph-structured data, and prepare the model input for the constructed dialogue graph; convert the scene labels into vector form, and fuse them with the node features of the graph neural network through a special embedding layer, so that the model can take into account the scene information during the learning process; with the goal of maximizing the effectiveness of the dialogue, define the loss function as a measure of the degree to which the dialogue deviates from the effective path.

[0079] Model training and optimization: Divide the dialogue graph data into small batches for training. In each batch, ensure that the dialogue data contains different scenarios to balance scene learning; according to the performance of the model on the validation set, adopt an adaptive learning rate adjustment strategy to ensure that the model can converge quickly and avoid overfitting; if there are multiple graph neural network models that perform well in different scenarios, fuse them.

[0080] Dialogue strategy extraction and verification: Analyze the optimal connection paths between nodes in different scenarios from the trained graph neural network model; use an independent test data set to verify the extracted dialogue strategy.

[0081] Compared with the prior art, the present invention has the following beneficial effects:

[0082] The present invention provides a method for improving the accuracy of AI-assisted large model intelligent question answering. Through data collection, data cleaning, data annotation, dynamic knowledge base construction, establishing a knowledge base update trigger mechanism, pre-training language model improvement, post-processing language understanding enhancement, combining rule-based and machine learning reasoning, reasoning result verification and optimization, dialogue strategy optimization, emotion perception and response, multilingual search, language recognition search, setting multi-dimensional evaluation indicators, regular evaluation and feedback improvement, a method for improving the accuracy of intelligent question answering with strong understanding ability, timely data update, outstanding reasoning ability, strong language recognition and dialogue ability is realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0084] Figure 1 It is a schematic flow chart of the method for improving the accuracy of AI-assisted large model intelligent question answering provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0085] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0086] The purpose of the present invention is to provide a method for improving the accuracy of intelligent question answering with strong understanding ability, timely data update, outstanding reasoning ability, and strong language recognition and dialogue ability.

[0087] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0088] Embodiment 1:

[0089] This embodiment provides a method for improving the accuracy of intelligent question answering of a large model assisted by AI, as Figure 1 shown, including:

[0090] Data collection:

[0091] Collect data from multiple sources;

[0092] For the data collection related to speech recognition search, establish a large-scale speech corpus covering speech data under different accents, speech rates, and background noise environments;

[0093] Perform data preprocessing, including:

[0094] Data cleaning:

[0095] Remove duplicate data, compare the collected data, and identify and delete duplicate text or speech records;

[0096] Filter low-quality data. For speech data, exclude data with serious noise interference that makes it impossible to accurately identify the content;

[0097] Data annotation:

[0098] For speech data, annotate the text, accent type, and speech rate corresponding to the speech content, and control the quality of annotation through cross-validation and expert review;

[0099] Dynamic knowledge base construction:

[0100] Based on knowledge graph technology, construct a dynamic knowledge base;

[0101] Establish a knowledge base update trigger mechanism:

[0102] Introduce external knowledge sources, regularly update the knowledge in the knowledge base, and at the same time verify new knowledge to ensure its accuracy;

[0103] Set update trigger conditions based on data popularity and timeliness;

[0104] When new data conflicts with the knowledge in the existing knowledge base, trigger the update process of the knowledge base, and use methods such as logical reasoning and evidence weight evaluation to determine whether to accept the new knowledge and avoid the introduction of incorrect knowledge;

[0105] Improvement of pre-trained language models:

[0106] Use large-scale and diverse corpora for pre-training;

[0107] Introduce semantic role labeling and dependency syntax tasks as auxiliary tasks in the pre-trained model, and conduct them simultaneously with the pre-training of the language model;

[0108] For users' questions, conduct multi-level semantic analysis; first, conduct morphological analysis to identify the part of speech and inflection of each word; then, conduct syntactic analysis to construct the syntactic tree of the sentence and determine the sentence structure; finally, conduct semantic analysis to understand the semantic intention of the sentence based on the pre-trained semantic knowledge base and the results of semantic role labeling;

[0109] Enhancement of post-processing language understanding ability:

[0110] Adopt a multi-turn dialogue strategy to clarify the user's intention. When there is ambiguity in the preliminary analysis of the user's question, obtain a more accurate user intention by asking follow-up questions;

[0111] Combination of rule-based and machine learning reasoning:

[0112] Establish a reasoning system based on logical rules. At the same time, use machine learning technology to learn a large number of knowledge instances and automatically discover hidden knowledge relationships and reasoning patterns;

[0113] Verification and optimization of reasoning results:

[0114] Verify the reasoning results by comparing them with the existing knowledge in the knowledge base and verifying them in a large-scale dataset. If the reasoning results conflict with the known knowledge or the verification accuracy rate in the dataset is low, optimize the reasoning process, adjust the rules or retrain the machine learning model;

[0115] Optimization of dialogue strategies:

[0116] Establish a dialogue state tracking mechanism to record the historical information of the dialogue, including the user's questions, the system's answers and follow-up questions; determine the next answer strategy according to the dialogue state;

[0117] Adopt a dialogue strategy learning algorithm to learn the best dialogue strategies in different scenarios from a large amount of dialogue data;

[0118] Emotion perception and response:

[0119] Add emotion analysis to the dialogue, identify the emotional tendency in the user's question, and adjust the tone and content of the answer according to the emotional tendency;

[0120] Multilingual search:

[0121] Build a multilingual index. In the search engine, establish independent indexes for text data in different languages and establish mapping relationships between languages;

[0122] Develop a multilingual semantic matching algorithm. Utilize cross - language word vector technology to map words in different languages to the same semantic space, thereby achieving semantic matching between different languages;

[0123] Language recognition search:

[0124] Improve the speech recognition algorithm by using a deep neural network to improve the accuracy of speech recognition;

[0125] Seamlessly connect the speech recognition result with the semantic understanding module;

[0126] Setting of multi - dimensional evaluation indicators:

[0127] Add an indicator for semantic understanding accuracy, which is measured by a combination of manual evaluation and automatic evaluation;

[0128] Set an indicator for dialogue coherence to evaluate the coherence between answers and questions and the logic between answers in multi - turn dialogues;

[0129] For language recognition search, set indicators for speech recognition accuracy and speech search response;

[0130] Regular evaluation and feedback for improvement:

[0131] Regularly evaluate the performance of the large - model, and find out the weak links in performance according to the results of the comprehensive evaluation index system;

[0132] Feed back the evaluation results to each component of the model for targeted optimization.

[0133] In one embodiment, the establishment of a large - scale speech corpus specifically includes:

[0134] Determine the collection sources, recruit volunteers from different regions to ensure coverage of multiple accents; for example, draw samples according to the population ratio in different dialect areas. Let the number of regions be n, and the planned number of samples to be collected in each region be m i(i = 1, 2, ···, n), then the total number of collected samples

[0135] Control the speed of speech change, design different text contents, and let volunteers read at different speeds; Quantify the speed of speech according to the number of words per minute, and set the normal speed of speech as s 0 (for example, 150 words per minute), the collected speed range is [s min , s max , collect samples at a certain interval (such as every 10 words / minute);

[0136] Use the noise synthesis algorithm to simulate the background noise environment;

[0137] Perform data preprocessing through the endpoint detection algorithm and the pre-emphasis algorithm; Among them,

[0138] The endpoint detection algorithm is an energy-based endpoint detection algorithm. Let the speech signal be x(n), and the energy By setting the energy threshold T, when E > T, it is determined as the start or end of the speech segment;

[0139] The pre-emphasis algorithm uses the formula y(n) = x(n) - αx(n - 1), where α takes 0 - 0.9, x(n) is the original speech signal, and y(n) is the pre-emphasized speech signal;

[0140] After pre-emphasis, divide the speech signal into short-time frames, perform Fourier transform on the short-time frames, and obtain a good approximation of the signal frequency profile by connecting adjacent frames;

[0141] Multiply each frame by a window function;

[0142] Then extract features by calculating the Mel frequency cepstral coefficients and the linear predictive coding coefficients, specifically:

[0143] Perform N-point FFT on each frame signal after frame division and windowing to calculate the spectrum, where N is 256 or 512. After obtaining the magnitude of the FFT, calculate the power spectrum through the formula;

[0144] Calculate the Mel filter bank, and extract the frequency bands by passing the power spectrum through a set of Mel-scaled triangular filters;

[0145] Apply the discrete cosine transform to decorrelate the filter bank coefficients and generate a compressed representation of the filter bank;

[0146] Analyze the speech signal through linear prediction, model the speech signal to obtain the linear prediction coefficients, and further calculate to obtain the linear prediction cepstral coefficients;

[0147] Perform feature fusion, fuse the extracted Mel-frequency cepstral coefficient features and linear prediction cepstral coefficient features to form a new feature vector as the final speech feature vector;

[0148] Let the speech feature vector be v, and map it to a storage location through the hash function h(v); the hash function is: where v i is an element of the feature vector, k is the vector dimension, and m is the number of buckets;

[0149] According to the occurrence frequency f of different feature values in the speech data i , construct a binary tree to minimize the weighted path length of the tree where l i is the coding length of the feature value, thereby realizing the efficient compression storage of speech data and obtaining a speech corpus.

[0150] In one embodiment, constructing a dynamic knowledge base based on the knowledge graph technology specifically includes:

[0151] Adopt a method combining rules and machine learning to extract the relationships between entities. First, formulate some basic relationship extraction rules. For example, in the sentence "Zhang San works in Li Si's company", the relationship "Zhang San - works in - Li Si's company" can be extracted. At the same time, use the convolutional neural network in deep learning to train the preprocessed data to improve the accuracy and coverage of relationship extraction; for newly emerging entity relationship types, discover them through unsupervised learning algorithms;

[0152] Use graph embedding technology (such as TransE, TransH and other models) to map the entities and relationships in the knowledge graph to a low-dimensional vector space. During the representation learning process, add semantic information to improve the semantic expression ability of the knowledge graph; adopt a multi-modal knowledge graph representation learning method. When there is multi-modal data, fuse the information of different modalities into the representation of the knowledge graph. For example, for an art work entity, map its image feature and text description feature to a low-dimensional vector space together;

[0153] When constructing a knowledge graph from multiple data sources, perform knowledge fusion, align the entities with the same name from different sources, match them through the unique identifier and attribute information of the entities. For the case where there are differences in the attribute values of the entities, adopt a credibility evaluation method to determine the final attribute value according to factors such as the reliability of the data source and the data update frequency; integrate the knowledge graphs in different fields, for example, integrate the knowledge graph in the medical field and the knowledge graph in the biotech field to establish cross-domain entity relationships, thereby completing the construction of the dynamic knowledge base.

[0154] In one embodiment, the method of using logical reasoning and weight of evidence evaluation to determine whether to accept new knowledge specifically includes:

[0155] When encountering new knowledge, first decompose it into basic propositions and statements;

[0156] Analyze the logical relationships between these propositions;

[0157] Use deductive reasoning and inductive reasoning to test these logical relationships;

[0158] Evaluate the source, diversity, and consistency of evidence;

[0159] Set a threshold for accepting new knowledge according to the rigor of logical reasoning and the evaluation results of the weight of evidence;

[0160] If the logical relationship is complete, and the evidence mainly comes from authoritative sources and is diverse and consistent, then when these conditions are met by more than 80%, decide to accept the new knowledge.

[0161] In one embodiment, the method of using the dialogue strategy learning algorithm to learn the best dialogue strategies in different scenarios from a large amount of dialogue data specifically includes:

[0162] Data screening and classification, using semantic understanding technology to conduct a preliminary scan of a large amount of dialogue data; identifying dialogue groups with similar or related semantics; for example, classifying dialogues about product function consultations into one category and distinguishing them from emotional communication dialogues. For the preliminarily classified dialogue groups, further subdivide them according to the specific characteristics of the scenario; for example, in the product function consultation category, further divide it into scenarios such as specific function A consultation and function B consultation.

[0163] Construct a dialogue graph representation, regard each round of dialogue as a node, and the node contains information such as the text content of the dialogue, semantic vector (generated by a pre-trained model), and the role of the dialogue participant (questioner or answerer). According to the order of the dialogue, construct directed edges between the nodes, and the weight of the edge is calculated according to the coherence and logical relationship of the dialogue; for example, the edge weight of a dialogue with good coherence is higher.

[0164] Based on neural network learning, select a graph neural network model suitable for processing graph-structured data to prepare the model input for the constructed dialogue graph; convert the scenario label into a vector form, and fuse it with the node features of the graph neural network through a special embedding layer, so that the model can consider the scenario information during the learning process; define the loss function as a measure of the degree of deviation of the dialogue from the effective path with the goal of maximizing the effectiveness of the dialogue (such as guiding the dialogue towards the direction of solving problems or meeting user needs).

[0165] Model training and optimization: The dialogue graph data is divided into small batches for training. In each batch, ensure that the dialogue data of different scenarios is included to balance scenario learning. According to the performance of the model on the validation set, adopt an adaptive learning rate adjustment strategy to ensure that the model can converge quickly and avoid overfitting. If multiple graph neural network models perform well in different scenarios, fuse them.

[0166] Dialogue policy extraction and verification: From the trained graph neural network model, analyze the optimal connection paths between nodes in different scenarios. Use an independent test data set to verify the extracted dialogue policy.

[0167] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts between the various embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0168] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. An AI-assisted large-model intelligent question-answering accuracy improvement method, characterized in that: include: Collect data from a variety of sources; For data collection related to speech recognition search, a large-scale speech corpus is established, covering speech data in different accents, speaking speeds, and background noise environments; Perform data preprocessing, including: De-duplication, comparing the collected data to identify and delete duplicate text or voice recordings; Filter low-quality data. For voice data, exclude data with severe noise interference that makes it impossible to accurately identify the content; For speech data, the text, accent type, and speaking speed corresponding to the speech content are annotated, and the quality of the annotation is controlled through cross-validation and expert review; Build a dynamic knowledge base based on knowledge graph technology; Introduce external knowledge sources, regularly update the knowledge in the knowledge base, and verify the new knowledge to ensure its accuracy; Set update trigger conditions based on data popularity and timeliness; When new data conflicts with the knowledge in the existing knowledge base, the knowledge base update process is triggered, and logical reasoning and evidence weight evaluation methods are used to determine whether to accept the new knowledge to avoid the introduction of erroneous knowledge; Use large-scale and diverse corpus for pre-training; Introducing semantic role labeling and dependency syntax tasks as auxiliary tasks in the pre-training model, and performing them simultaneously with the pre-training of the language model; For user questions, we conduct multi-level semantic analysis. First, we conduct lexical analysis to identify the part of speech and word form changes of each word. Then, we conduct syntactic analysis to build the sentence syntax tree and determine the sentence structure. Finally, we conduct semantic analysis to understand the semantic intent of the sentence based on the pre-trained semantic knowledge base and semantic role labeling results. Use a multi-round dialogue strategy to clarify user intentions. When preliminary analysis shows that a user's question is ambiguous, use follow-up questions to obtain a more accurate understanding of the user's intentions. Establish a reasoning system based on logical rules. At the same time, use machine learning technology to learn a large number of knowledge instances and automatically discover hidden knowledge relationships and reasoning patterns. Verify the results of reasoning by comparing them with existing knowledge in the knowledge base and verifying them in large-scale data sets. If the reasoning results conflict with known knowledge or the verification accuracy in the data set is low, optimize the reasoning process, adjust the rules or retrain the machine learning model. Establish a dialogue status tracking mechanism to record the historical information of the dialogue, including the user's questions, the system's answers, and follow-up questions; determine the next round of answering strategy based on the dialogue status; Use dialogue strategy learning algorithms to learn the best dialogue strategies in different scenarios from a large amount of dialogue data; Add sentiment analysis to the conversation to identify the emotional tendencies in the user's questions and adjust the tone and content of the answer based on the emotional tendencies; Build multilingual indexes. In the search engine, create independent indexes for text data in different languages ​​and establish mapping relationships between languages. Develop a multilingual semantic matching algorithm, using cross-language word vector technology to map words in different languages ​​to the same semantic space, thereby achieving semantic matching between multiple languages; Improve speech recognition algorithms and use deep neural networks to improve the accuracy of speech recognition; Seamlessly connect speech recognition results with semantic understanding modules; Add semantic understanding accuracy indicators, which are measured by combining manual and automatic evaluation; Set dialogue coherence indicators to evaluate the coherence of answers and questions and the logic between answers in multiple rounds of dialogue; For voice recognition search, set the voice recognition accuracy index and voice search response index; Regularly evaluate the performance of large models and identify weak links in performance based on the results of the comprehensive evaluation index system; The evaluation results are fed back into the various components of the model for targeted optimization.

2. The AI-assisted large-model intelligent question-answering accuracy improvement method according to claim 1 is characterized in that: The establishment of a large-scale speech corpus specifically includes: Determine the source of collection and recruit volunteers from different regions to ensure that multiple accents are covered; let the number of regions be n, and the number of samples planned to be collected in each region be m i (i=1,2,···,n), then the total number of samples collected is Control the change of speech speed, design different text contents, and let volunteers read at different speeds; quantify the speech speed according to the number of words per minute, set the normal speech speed as s0, and the collected speech speed range is [s min ,s max ], collect samples at certain intervals; Use noise synthesis algorithm to simulate background noise environment; Data preprocessing is performed through endpoint detection algorithm and pre-emphasis algorithm; among them, The endpoint detection algorithm is an energy-based endpoint detection algorithm. Let the speech signal be x(n), and the energy By setting the energy threshold T, when E>T, it is determined as the beginning or end of the speech segment; The pre-emphasis algorithm uses the formula y(n)=x(n)-αx(n-1), where α is 0-0.9, x(n) is the original speech signal, and y(n) is the pre-emphasized speech signal; After pre-emphasis, the speech signal is divided into short-time frames, Fourier transform is performed on the short-time frames, and a good approximation of the signal frequency contour is obtained by connecting adjacent frames; Multiply each frame by a window function; Then, the features are extracted by calculating the Mel frequency cepstral coefficients and linear prediction coding coefficients, specifically: Perform N-point FFT on each frame signal after frame division and windowing to calculate the spectrum, where N is 256 or 512. After obtaining the FFT amplitude, calculate the power spectrum using the formula; Calculate the Mel filter bank and extract the frequency bands by passing the power spectrum through a set of Mel-scaled triangular filters; Applying a discrete cosine transform to decorrelate the filter bank coefficients and produce a compressed representation of the filter bank; The speech signal is modeled by linear prediction analysis to obtain linear prediction coefficients, and then the linear prediction cepstrum coefficients are further calculated; Perform feature fusion, fuse the extracted Mel frequency cepstral coefficient features and linear prediction cepstral coefficient features to form a new feature vector as the final speech feature vector; Suppose the speech feature vector is v, and it is mapped to a storage location through a hash function h(v); the hash function is: where v i are the elements of the feature vector, k is the vector dimension, and m is the number of buckets; According to the frequency of occurrence of different feature values ​​in the speech data f i , construct a binary tree so that the weighted path length of the tree Minimum, where l i is the encoding length of the feature value, thereby achieving efficient compression storage of speech data and obtaining a speech corpus.

3. The AI-assisted large-model intelligent question-answering accuracy improvement method according to claim 1 is characterized in that: The construction of a dynamic knowledge base based on the knowledge graph technology specifically includes: A rule-based and machine learning method is used to extract the relationship between entities, and a convolutional neural network in deep learning is used to train the preprocessed data to improve the accuracy and coverage of relationship extraction; new entity relationship types are discovered through unsupervised learning algorithms; Graph embedding technology is used to map entities and relationships in the knowledge graph to a low-dimensional vector space. In the representation learning process, semantic information is added to improve the semantic expression ability of the knowledge graph. A multimodal knowledge graph representation learning method is used. When multimodal data exists, information from different modalities is integrated into the representation of the knowledge graph, and its image features and text description features are mapped to a low-dimensional vector space. When constructing a knowledge graph from multiple data sources, knowledge fusion is performed, entities with the same name from different sources are aligned, and matched through the entity's unique identifier and attribute information. In the case of differences in entity attribute values, a credibility assessment method is used to determine the final attribute value based on factors such as the reliability of the data source and data update frequency. Knowledge graphs from different fields are integrated and cross-domain entity relationships are established to complete the construction of a dynamic knowledge base.

4. The AI-assisted large-model intelligent question-answering accuracy improvement method according to claim 1 is characterized in that: The method of using logical reasoning and weight of evidence assessment to determine whether to accept new knowledge specifically includes: When encountering new knowledge, first break it down into basic propositions and statements; Analyze the logical relationship between these propositions; Use deductive and inductive reasoning to test these logical relationships; Assess the sources, diversity, and consistency of evidence; Set a threshold for accepting new knowledge based on the evaluation of the rigor of logical reasoning and the weight of evidence; If the logical relationship is complete, and the evidence mainly comes from authoritative sources and is diverse and consistent, then the decision to accept new knowledge is made when these conditions are met by more than 80%.

5. The AI-assisted large-model intelligent question-answering accuracy improvement method according to claim 1 is characterized in that: The dialogue strategy learning algorithm is used to learn the best dialogue strategies in different scenarios from a large amount of dialogue data, specifically including: Data screening and classification: using semantic understanding technology to conduct a preliminary scan of a large amount of conversation data; identifying conversation groups with similar or related semantics; and further subdividing the conversation groups after preliminary classification according to the specific characteristics of the scene; Construct a dialogue graph representation, treating each round of dialogue as a node. According to the order of the dialogue, construct directed edges between nodes. The weight of the edge is calculated according to the coherence and logical relationship of the dialogue. Based on neural network learning, a graph neural network model suitable for processing graph structure data is selected to prepare the model input for the constructed dialogue graph; the scene label is converted into a vector form and fused with the node features of the graph neural network through a special embedding layer, so that the model can take scene information into account during the learning process; with the goal of maximizing the effectiveness of the dialogue, the loss function is defined as a measure of the degree to which the dialogue deviates from the effective path; Model training and optimization: Divide the conversation graph data into small batches for training. In each batch, ensure that conversation data from different scenarios are included to balance scenario learning. Adopt an adaptive learning rate adjustment strategy based on the performance of the model on the validation set to ensure that the model can converge quickly and avoid overfitting. If multiple graph neural network models perform well in different scenarios, merge them. Dialogue strategy extraction and verification: Analyze the optimal connection path between nodes in different scenarios from the trained graph neural network model; use an independent test data set to verify the extracted dialogue strategy.

Citation Information

Cited By

  • Non-performing asset cross-scene question and answer framework based on knowledge graph

    CN120216706A

  • Cross-scenario question-answering framework for non-performing assets based on knowledge graph

    CN120216706B

  • Multi-module cooperative intelligent role playing system

    CN120542581A

  • Management method of multi-mode enterprise knowledge base system

    CN120561342A

  • Intelligent text labeling system, method and equipment oriented to standard semantic knowledge extraction

    CN120765857A