Question and answer text generation method based on semantic analysis of AI agent
By using semantic analysis methods to identify and eliminate semantically ambiguous question pairs and build a high-quality FQA corpus, we can solve the accuracy problem caused by semantic ambiguity in the AI intelligent question-answering system and improve the accuracy of question-answering text generation and user trust.
Patent Information
- Application Number
- CN202511095057.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing AI agent question-answering systems fail to effectively identify semantically ambiguous questions in historical question-answering records, resulting in poor accuracy of questions and their answer results in the FQA corpus, affecting the accuracy of question-answer text generation.
Through semantic analysis methods, the semantic change vector between the keyword set of the question pair and the question-answer keyword set is obtained, the semantic abnormality deviation and question-answer matching degree are calculated, the semantically ambiguous question pairs are eliminated, and a high-quality FQA corpus is constructed.
The accuracy of questions and their answer results in the FQA corpus is improved, ensuring the accuracy of AI agent question and answer text generation, reducing semantic ambiguity interference, and increasing users' trust in the question and answer results.
Smart Images

Figure CN120596641B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of semantic analysis, and particularly relates to an AI agent question and answer text generation method based on semantic analysis. BACKGROUND
[0002] An AI agent question and answer system is one of important fields of artificial intelligence application. The AI agent question and answer system can provide uninterrupted service for users, and quickly respond to user queries to improve user query satisfaction. At present, the requirement for AI agent question and answer accuracy is increasing. The AI agent question and answer system answers the question raised by the user in natural language by using accurate and concise natural language, and can provide a reference source of the generated text of the question and answer or explain the question and answer reasoning process, thereby improving the user's trust in the AI agent question and answer result.
[0003] In the prior art, the AI agent question and answer system mostly uses a FQA (Frequently-Asked Questions) question and answer method. The FQA corpus set is established by using the historical question and answer record input mode, and a series of questions and corresponding question and answer results in the FQA corpus set are stored in the AI agent question and answer system, so as to quickly respond to the question and answer request of the user and reduce the waiting time of the user for the question and answer. However, since there may be semantic ambiguity in the historical question and answer record, the prior art does not fully consider the semantic ambiguity of the question in the historical question and answer record to establish the FQA corpus, so that the high-quality question and its question and answer result cannot be effectively identified, resulting in poor accuracy of the question and its question and answer result in the FQA corpus, and affecting the accuracy of the AI agent question and answer text generation. SUMMARY
[0004] In order to solve the above technical problems, the present application provides an AI agent question and answer text generation method based on semantic analysis to solve the existing problems.
[0005] The AI agent question and answer text generation method based on semantic analysis of the present application adopts the following technical scheme:
[0006] One embodiment of the present application provides an AI agent question and answer text generation method based on semantic analysis, comprising the following steps:
[0007] Obtaining each question and answer pair in the historical question and answer record and the question and answer time of each question and answer pair, and extracting the keywords of the question text and the question and answer generation text of each question and answer pair to obtain the question keyword set and the question and answer keyword set of each question and answer pair;
[0008] The semantic change vector of each keyword is obtained by the semantic difference between the keyword in the question keyword set of each question pair and the keyword in the question and answer keyword set, and the semantic abnormal deviation of each keyword in the question keyword set of each question pair is obtained according to the semantic correlation degree between the semantic change vector of each keyword and the keyword frequency in the question and answer keyword set, and the change trend of the elements in the semantic change vector.
[0009] The semantic ambiguity of the question text in each question pair is obtained by the semantic correlation degree between the semantic abnormal deviation of the keyword in the question keyword set of each question pair and the keyword frequency, and the difference between the semantic abnormal deviations of different keywords in the question keyword set, and the question and answer matching degree of each question pair is obtained according to the matching degree between the question keyword set and the question and answer keyword set of each question pair, and the similarity between the question keyword sets of each question pair and its near neighbor question pair.
[0010] The question and answer matching degree is used to eliminate the question pairs in the historical question and answer records to construct the FQA corpus, and the AI agent generates the corresponding question and answer text according to the user's question.
[0011] Preferably, the method for obtaining the semantic change vector of each keyword is as follows:
[0012] The normalized Google distance between each keyword in the question keyword set of each question pair and all keywords in the question and answer keyword set thereof is calculated, and all normalized Google distances are arranged in ascending order according to the keyword frequency in the generated question and answer text, so as to obtain the semantic change vector of each keyword in the question keyword set of each question pair.
[0013] Preferably, the method for obtaining the semantic abnormal deviation of each keyword in the question keyword set of each question pair is as follows:
[0014] The keyword frequency weight vector of the question and answer keyword set of each question pair is extracted, the scalar product between the semantic change vector of each keyword in the question keyword set of each question pair and the keyword frequency weight vector of the question and answer keyword set of each question pair is calculated, and the ratio of the number of positive elements in the first-order difference vector of the semantic change vector to the total number is calculated, and the product of the scalar product and the ratio is taken as the semantic abnormal deviation of each keyword in the question keyword set of each question pair.
[0015] Preferably, the keyword frequency weight vector of the question and answer keyword set of each question pair is extracted by arranging all keyword frequencies in the question and answer keyword set of each question pair in ascending order to form the keyword frequency weight vector of the question and answer keyword set of each question pair.
[0016] Preferably, the method for calculating the semantic ambiguity of the question text in each question pair is as follows:
[0017] wherein, is the semantic ambiguity of the question text in the i-th question pair, is the dot product between the semantic deviation vector of the question keyword set of the i-th question pair and the term frequency weight vector of the question keyword set, is the number of elements in the semantic deviation vector of the question keyword set of the i-th question pair, and are the s-th and s-1-th elements in the semantic deviation vector of the question keyword set of the i-th question pair, respectively.
[0018] Preferably, the semantic deviation of each keyword in the question keyword set of each question pair is arranged in descending order of the term frequency of the keyword in the question text, as the semantic deviation vector of the question keyword set of each question pair, and the term frequency of each keyword in the question keyword set is arranged in descending order, as the term frequency weight vector of the question keyword set of each question pair.
[0019] Preferably, the question pairs with the shortest time interval from each question pair are all regarded as the near neighbor question pairs of each question pair.
[0020] Preferably, the method for calculating the question-answer matching degree of each question pair is as follows:
[0021] wherein, is the question-answer matching degree of the i-th question pair, is the semantic ambiguity of the question text in the i-th question pair, is the Jaccard similarity coefficient between the question keyword set and the question-answer keyword set of the i-th question pair, is the exponential function with a natural constant as the base number, is the question similarity of the i-th question pair.
[0022] Preferably, the method for obtaining the question similarity of each question pair further comprises: calculating the average value of the Jaccard similarity coefficients between the question keyword set of each question pair and the question keyword sets of its near neighbor question pairs, as the question similarity of each question pair.
[0023] Preferably, the step of eliminating the question pairs in the historical question-answer record by using the question-answer matching degree comprises: normalizing the question-answer matching degrees of all question pairs in the historical question-answer record, and if the normalized question-answer matching degree is less than a question-answer matching threshold, the question pair is eliminated from the historical question-answer record.
[0024] The present application has at least the following beneficial effects:
[0025] The present application considers that the prior art does not fully analyze the semantic ambiguity of the question in the historical question and answer record to establish the FQA corpus, so that the high-quality question and its question and answer result cannot be effectively identified, resulting in poor accuracy of the question and its question and answer result in the FQA corpus, affecting the accuracy of the AI agent question and answer text generation. Therefore, in the present application, the strength of the semantic relationship between the keywords in the question text and its question and answer text is considered, and at the same time the abnormal characteristics of the semantic relationship between the keywords with the change of the word frequency are considered, the abnormal deviation characteristics of the semantic relationship between the keywords are accurately measured, which is used for more accurate feature extraction of the semantic ambiguity of the question text in the subsequent;
[0026] Further, the present application accurately measures the semantic ambiguity of the question text in the question pair by focusing on the semantic abnormal deviation characteristics of the keywords with large word frequency and combining the change characteristics of the semantic abnormal deviation, which is beneficial to the subsequent identification of high-quality questions and their question and answer results;
[0027] At the same time, the present application extracts the question and answer matching degree of each question pair in the historical question and answer record by the semantic ambiguity of the question text and the similarity of the questions in the multiple question and answer processes, and effectively identifies the high-quality questions and their question and answer results in the historical question and answer record by using the question and answer matching degree, and then constructs the FQA corpus, improves the accuracy of the question and its question and answer result in the FQA corpus, and avoids affecting the accuracy of the AI agent question and answer text generation. BRIEF DESCRIPTION OF DRAWINGS
[0028] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed in the embodiment or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0029] Figure 1 The step flow chart of the AI agent question and answer text generation method based on semantic analysis provided by the present application. DETAILED DESCRIPTION
[0030] For further elaboration of the technical means and effects taken by the present application to achieve the predetermined object of the application, the specific implementation, structure, features and effects of the AI agent question and answer text generation method based on semantic analysis proposed according to the present application are described in detail as follows in combination with the drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0031] Unless otherwise defined, such as the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the circuit structure, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed or inherent to such article or device. Without more limitation, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the article or device comprising the element. In addition, the term "and / or" used herein includes any and all combinations of one or more related listed items. All technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs.
[0032] The specific scheme of the AI agent question and answer text generation method based on semantic analysis provided by the present application is described in detail below in combination with the drawings.
[0033] The AI agent question and answer text generation method based on semantic analysis provided by one embodiment of the present application, specifically, please refer to Figure 1 , including the following steps:
[0034] Step 1: Obtain each question pair and its question and answer time in the historical question and answer record, and extract the keywords of the question text and the question and answer generation text of each question pair, to obtain the question keyword set and the question and answer keyword set of each question pair.
[0035] Query the historical question and answer record in the last half year through the AI agent question and answer system server, and obtain each question pair and its corresponding question and answer time in the database where the historical question and answer record is located, wherein the question pair includes the question text and the question and answer generation text of the user's question and answer.
[0036] In order to facilitate subsequent more accurate semantic analysis, thereby more effectively identifying high-quality questions and their question and answer results, punctuation marks and meaningless characters in the question text and the question and answer generation text of each question pair are removed by using a regular matching algorithm, and then the removed question text and its corresponding question and answer generation text are respectively input into a keyword extraction algorithm, which can be a TF-IDF algorithm or a TextRank algorithm. In this embodiment, a TF-IDF algorithm is used to obtain a question keyword set of the question text and a question and answer keyword set of the question and answer generation text in each question pair. Both the regular matching algorithm and the TF-IDF algorithm are well-known technologies, and the specific process will not be described again.
[0037] Step 2: Obtain the semantic change vector of each keyword by the semantic difference between the keywords in the question keyword set and the question and answer keyword set of each question pair. According to the semantic correlation between the semantic change vector of each keyword and the keyword frequency in the question and answer keyword set, and combining the change trend of the elements in the semantic change vector, the semantic abnormal deviation of each keyword in the question keyword set of each question pair is obtained.
[0038] In the process of establishing the FQA corpus by using the historical question and answer record input method, the meaning of the words or concepts in the question text of the user question and answer may be ambiguous in the historical question and answer record, thereby affecting the accuracy of the questions and their question and answer results in the established FQA corpus, and further affecting the accuracy of the AI agent question and answer text generation. Therefore, in order to avoid the adverse effects on the accuracy of the AI agent question and answer text generation, it is necessary to accurately extract the semantic ambiguity of the question in the historical question and answer record by using the semantic analysis method.
[0039] The relationship between the keywords in the question text and the keywords in the question and answer generation text is analyzed semantically, the normalized Google distance between the jth keyword in the question keyword set of the ith question pair and all keywords in its question and answer keyword set is calculated, and all normalized Google distances are arranged in order of keyword frequency in the question and answer generation text from small to large. If the keyword frequency in the question and answer generation text is equal, random arrangement is performed, thereby obtaining the semantic change vector of the jth keyword in the question keyword set of the ith question pair. The semantic change vector reflects the change of the semantic relationship between the keywords in the question text and all keywords in the question and answer generation text.
[0040] Generally, the keywords with high frequency in the question-answer generated text often have strong semantic relationship with the keywords in the question text. Therefore, the frequencies of all keywords in the question-answer keyword set of the ith question pair are arranged in descending order, and if the frequencies of the keywords in the question-answer generated text are equal, the keywords are randomly arranged, and the frequency weight vector of the question-answer keyword set of the ith question pair is recorded.
[0041] Further, by using the vector quantity product calculation method, the elements in the semantic change vector are weighted and summed, specifically, the quantity product between the semantic change vector and the frequency weight vector of the question-answer keyword set is calculated, which is recorded as the first quantity product in this embodiment for convenience of distinguishing and describing. The larger the first quantity product is, the weaker the semantic relationship between the keywords in the question text and all the keywords in the question-answer generated text is, and the more likely the question is to have semantic ambiguity, and the quality of the question and its question-answer result is more easily affected.
[0042] At the same time, the first-order difference vector of the semantic change vector is calculated, and the number of positive elements in the first-order difference vector is counted. The ratio of the number of positive elements to the total number is calculated. Under normal circumstances, since the keywords with high frequency in the question-answer generated text often have strong semantic relationship with the keywords in the question text, the semantic change vector should strictly follow the decreasing trend. The higher the ratio is, the less obvious the decreasing trend in the semantic change vector is, and the more significant the semantic change anomaly feature is, that is, the more likely the question is to have semantic ambiguity.
[0043] Therefore, in this embodiment, the product of the ratio corresponding to the jth keyword in the question keyword set of the ith question pair and the first quantity product is recorded as the semantic anomaly deviation of the jth keyword in the question keyword set of the ith question pair, wherein the semantic anomaly deviation reflects the abnormal deviation feature of the semantic relationship between the keywords in the question keyword set and all the keywords in the corresponding question-answer keyword set. The greater the abnormal deviation feature of the semantic relationship is, the more likely the keyword has an inaccurate representation feature, and the more likely the question text with the keyword has semantic ambiguity and is more easily affected by the accuracy of the subsequent FQA corpus construction, and further affects the accuracy of the AI agent question-answer text generation.
[0044] Step 3: The semantic ambiguity of the question text in each question pair is obtained by combining the semantic correlation between the semantic anomaly deviation of the keywords in the question keyword set of each question pair and the word frequency of the keywords, and the difference between the semantic anomaly deviations of different keywords in the question keyword set. The question-answer matching degree of each question pair is obtained based on the matching degree between the question keyword set of each question pair and the question-answer keyword set, as well as the similarity between each question pair and its neighboring question pairs regarding the question keyword set.
[0045] Generally speaking, if the question text given by the user in the historical record contains semantic ambiguity, then the semantic anomaly deviation characteristics of all keywords in the question text are at a higher level and the changes in semantic anomaly deviation are more inconsistent. Therefore, the semantic anomaly deviations of all keywords in the question keyword set of the i-th question pair are arranged in order of the keyword frequency in the question text from small to large. If the keyword frequency in the question-answer generated text is equal, they are randomly arranged and used as the semantic deviation vector of the question keyword set of the i-th question pair.
[0046] At the same time, all keywords in the question keyword set of the i-th question pair are arranged in ascending order according to their frequency of appearance in the question text. If the keyword frequencies in the question-answer generated text are equal, they are randomly arranged to obtain the frequency weight vector of the question keyword set of the i-th question pair.
[0047] Furthermore, the scalar product between the semantic deviation vector and the word frequency weight vector of the question keyword set is calculated, and the semantic anomaly deviation in the semantic deviation vector is weighted and summed using the scalar product calculation method, so that the semantic analysis focuses more on the semantic anomaly deviation characteristics of keywords with higher word frequency, thereby improving the accuracy of subsequent measurement of the semantic ambiguity phenomenon in the question text.
[0048] Through the above analysis, the semantic ambiguity of the question text in the i-th question pair is calculated:
[0049] ;
[0050] Where, is the semantic ambiguity of the question text in the i-th question pair, is the product of the semantic deviation vector of the question keyword set of the i-th question pair and the word frequency weight vector of the question keyword set, is the number of elements in the semantic deviation vector of the question keyword set of the i-th question pair, and They are respectively the sth and s-1th elements in the semantic deviation vector of the question keyword set of the i-th question pair.
[0051] The semantic ambiguity reflects the degree of ambiguity of semantic expression in the historical problem text. The higher the degree of ambiguity of semantic expression, the more likely it is to interfere with the matching between the generated question and answer text and the problem text by the AI agent, that is, the more likely the generated question and answer text cannot solve the problem raised by the user, and the more inappropriate the problem and its question and answer results are for participating in the construction of the FQA corpus.
[0052] Generally, if the problem raised by the user has a large semantic ambiguity feature, the question and answer text given by the AI agent cannot effectively solve the problem raised by the user, at which time the user is more likely to ask the same question multiple times. Therefore, in this embodiment, all question pairs and their corresponding question and answer times in the historical question and answer records are traversed, and the K nearest neighbor question pairs of each question pair are taken as the K nearest neighbor question pairs of each question pair. In this embodiment, K is 30, and the implementer can adaptively take values according to actual conditions.
[0053] Further, the average of the Jaccard similarity coefficients between the question keyword set of the i-th question pair and the question keyword sets of its K nearest neighbor question pairs is calculated, and is taken as the question similarity of the i-th question pair. The greater the question similarity, the more likely the user is to ask the same question multiple times, and the more likely the question and answer text in the question pair cannot accurately solve the problem given in the problem text.
[0054] Therefore, according to the semantic ambiguity of the problem text in each question pair, and in combination with the matching degree between the question keyword set and the question and answer keyword set of each question pair, and the similarity between the question keyword set of each question pair and its nearest neighbor question pair, the question and answer matching degree of each question pair in the historical question and answer records is calculated:
[0055] ;
[0056] In the formula, is the question and answer matching degree of the i-th question pair, is the Jaccard similarity coefficient between the question keyword set and the question and answer keyword set of the i-th question pair, is the exponential function with a natural constant as the base, is the question similarity of the i-th question pair.
[0057] Among them, the question and answer matching degree reflects the influence of semantic ambiguity on the matching of the question and answer result. The smaller the question and answer matching degree, the greater the influence of semantic ambiguity on the question and answer result, the poorer the matching between the question and answer result, the more inappropriate the question and answer result is for participating in the construction of the FQA corpus, and the more it should be excluded when constructing the FQA corpus.
[0058] Step 4: The questions in the historical question and answer records are removed using the question and answer matching degree to construct the FQA corpus, and the AI agent generates corresponding question and answer texts according to the user's question.
[0059] In order to improve the accuracy of the questions and their answers in the FQA corpus, preferably, in this embodiment, the question and answer matching degree of all question pairs in the historical question and answer records is subjected to exponential normalization processing. If the question and answer matching degree after the exponential normalization processing is less than the question and answer matching threshold value, it indicates that the corresponding question pair has a serious semantic ambiguity problem, which affects the accuracy of the question and answer result, and then the question pair is removed from the historical question and answer records. In an embodiment, the value of the question and answer matching threshold value is 0.3.
[0060] Further, all question pairs in the removed historical question and answer records are de-duplicated to ensure the uniqueness of each question and its answer, and all question pairs after de-duplication are standardized to ensure the uniformity of the data format in the question pairs. All standardized question pairs construct the FQA corpus, wherein the construction of the FQA corpus is a known technology, and the specific process is not described again.
[0061] All question pairs in the FQA corpus are stored in the AI agent question and answer system in JSON format, and the FQA corpus is updated and stored every week using the same method to ensure the real-time nature of the FQA corpus. At the same time, the AI agent question and answer system is connected to the large language model through the API interface. The AI agent question and answer system can communicate with the large language model through the API interface. In this embodiment, the large language model connected by the AI agent question and answer system is the Ali Cloud Tongyi Qianwen model. In the actual application process, the implementer can select other large language models.
[0062] When the user asks a question, if the AI agent question and answer system can find a question that matches the user's question text successfully through the text matching algorithm, the answer corresponding to the question is generated as a question and answer text and returned to the user. The text matching algorithm can be Jaccard similarity or Levenshtein edit distance. In this embodiment, Levenshtein edit distance is used for text matching. Using Levenshtein edit distance for text matching is a known technology, and the specific process is not described again.
[0063] If the AI agent question and answer system cannot find a question that matches the user's question text successfully through the text matching algorithm, it means that the current FQA corpus does not have the question asked by the user. The question asked by the user is sent to the large language model. The large language model performs semantic reasoning on the question text asked by the user, and then generates a semantic reasoning question and answer result as a question and answer text and returns it to the user.
[0064] It is to be understood that the phase "one embodiment" or "some embodiments" or the like means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places in the specification are not necessarily all referring to the same embodiment, although it can. Furthermore, the term "comprising" or "containing" or "including" or "having" or the like are used synonymously, unless otherwise indicated.
[0065] It should be noted that the above-mentioned order of the embodiments of the application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above-mentioned description is for a specific embodiment of the present application. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous. At the same time, the size of the serial number of each step in the embodiment does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments in the specification.
[0066] The above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. The AI agent question-answering text generation method based on semantic analysis is characterized by: The following steps are involved: Obtain each question pair and its question and answer time in the historical question and answer records, and extract the keywords of the question text and the question and answer generated text of each question pair to obtain the question keyword set and question and answer keyword set for each question pair; The semantic difference between the keywords in the question keyword set and the question-answer keyword set of each question pair is used to obtain the semantic change vector of each keyword. Based on the semantic correlation between the semantic change vector of each keyword and the word frequency of each keyword in the question-answer keyword set, combined with the change trend of the elements in the semantic change vector, the semantic abnormal deviation of each keyword in the question keyword set of each question pair is obtained. The semantic fuzziness of the question text in each question pair is obtained by combining the semantic correlation between the semantic anomaly deviation of the keywords in the question keyword set of each question pair and the word frequency of the keywords, and the difference between the semantic anomaly deviations of different keywords in the question keyword set. The question-answer matching degree of each question pair is obtained based on the matching degree between the question keyword set of each question pair and the question-answer keyword set, as well as the similarity between each question pair and its neighboring question pairs in terms of the question keyword set. The question-answer matching degree is used to eliminate question pairs in historical question-answer records to build an FQA corpus, and the AI agent is combined to generate corresponding question-answer texts based on the user's questions.
2. The method for generating AI agent question-answering text based on semantic analysis according to claim 1, characterized in that: The method for obtaining the semantic change vector of each keyword is: Calculate the normalized Google distance between each keyword in the question keyword set of each question pair and all keywords in its question-answer keyword set, and arrange all normalized Google distances from small to large according to the word frequency of the keywords in the question-answer generated text, and obtain the semantic change vector of each keyword in the question keyword set of each question pair.
3. The method for generating AI agent question-answer text based on semantic analysis according to claim 1, wherein: The method for obtaining the semantic abnormality deviation of each keyword in the question keyword set of each question pair is: Extract the word frequency weight vector of the question and answer keyword set of each question pair, calculate the scalar product between the semantic change vector of each keyword in the question keyword set of each question pair and the word frequency weight vector of the question and answer keyword set of each question pair, and count the ratio of the number of positive elements in the first-order difference vector of the semantic change vector to the total number of elements, and multiply the scalar product by the ratio as the semantic abnormality deviation of each keyword in the question keyword set of each question pair.
4. The method for generating AI agent question-answer text based on semantic analysis according to claim 1, wherein: The method for extracting the word frequency weight vector of the question and answer keyword set of each question pair is: arranging the word frequencies of all keywords in the question and answer keyword set of each question pair from small to large to form the word frequency weight vector of the question and answer keyword set of each question pair.
5. The method for generating AI agent question-answer text based on semantic analysis according to claim 1, wherein: The calculation method of the semantic ambiguity of the question text in each question pair is: Where, is the semantic ambiguity of the question text in the i-th question pair, is the product of the semantic deviation vector of the question keyword set of the i-th question pair and the word frequency weight vector of the question keyword set, is the number of elements in the semantic deviation vector of the question keyword set of the i-th question pair, and They are respectively the sth and s-1th elements in the semantic deviation vector of the question keyword set of the i-th question pair.
6. The method for generating AI agent question-answer text based on semantic analysis according to claim 5, characterized in that: The semantic anomaly deviations of the keywords in the question keyword set of each question pair are arranged from small to large according to the word frequency of the keywords in the question text, which are used as the semantic deviation vector of the question keyword set of each question pair. The word frequency of all keywords in the question keyword set in the question text is arranged from small to large to form the word frequency weight vector of the question keyword set of each question pair.
7. The method for generating AI agent question-answer text based on semantic analysis according to claim 1, wherein: The multiple question pairs with the closest time interval between the question and answer of each question pair are regarded as the neighboring question pairs of each question pair.
8. The method for generating AI agent question-answer text based on semantic analysis according to claim 1, wherein: The calculation method of the question-answer matching degree of each question pair is: Where, is the question-answer matching degree of the i-th question pair, is the semantic ambiguity of the question text in the i-th question pair, is the Jaccard similarity coefficient between the question keyword set and the question-answer keyword set of the i-th question pair, is an exponential function with a natural constant as base, is the question similarity of the i-th question pair.
9. The method for generating AI agent question-answer text based on semantic analysis according to claim 8, characterized in that: The method for obtaining the question similarity of each question pair further includes: calculating the average value of the Jaccard similarity coefficients between the question keyword set of each question pair and the question keyword sets of its neighboring question pairs as the question similarity of each question pair.
10. The method for generating AI agent question-answer text based on semantic analysis according to claim 1, wherein: The method of using the question-answer matching degree to eliminate question pairs in the historical question-answer records includes normalizing the question-answer matching degrees of all question pairs in the historical question-answer records; if the normalized question-answer matching degree is less than the question-answer matching threshold, the question pair is eliminated from the historical question-answer records.
Citation Information
Patent Citations
Answering system based on artificial intelligence
CN117407515A
Innovation and entrepreneurship coaching question and answer matching method and system based on semantic understanding
CN120256590A