ABI question and answer method and system under condition of no label data

By analyzing the importance and matching of keywords in the ABI question and answer system, building a decision tree and feature matrix, and introducing an attention mechanism, the problem that existing systems cannot handle complex sentences and uncommon expressions is solved, and personalized and context-sensitive answers are achieved.

CN119938877AActive Publication Date: 2025-05-06HENAN ZHONGCHENG INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411957496.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

Existing Q&A systems are unable to handle complex sentence structures or uncommon expressions, and answers lack personalization and context sensitivity.

Method used

By using the method without label data in the ABI question and answer system, the user query content and database files are obtained, the keyword importance and matching degree are analyzed, the decision tree and feature matrix are constructed, and the attention mechanism is introduced to generate personalized and context-sensitive answers.

Benefits of technology

It realizes effective processing of complex sentences and uncommon expression methods, and improves the reference value and personalization of the answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938877A_ABST
    Figure CN119938877A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of knowledge questions and answers, in particular to an ABI questions and answers method and system under the condition of no label data, and the method comprises the following steps: obtaining query content proposed by a user for the same question and files in an ABI management system database; obtaining the importance degree of each keyword in the query keyword set; obtaining the matching degree of each file in the database and the query content; obtaining the core key degree of each keyword in each file in the database; according to the key degree and the importance degree of the core, constructing a feature matrix of query content keywords and database files; obtaining a final related file of the query content; obtaining an overall attention score of each final related file relative to the query content; and screening the answer file according to the total attention score to generate a reply. Complex sentence structures or uncommon expression modes can be processed, effective answers about the query content are provided, and the reference value of the answers is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge question answering, and in particular to an ABI question answering method and system in the case of unlabeled data. Background Art

[0002] In response to the knowledge question-and-answer needs in multiple fields in enterprises, traditional question-and-answer systems are usually based on preset rules and templates. If the questions are beyond the scope of the template, the system's answers may not be of reference value. Therefore, a generative ABI platform can be built to provide support for multiple database types. Administrators or users can configure the corresponding database of the enterprise on this ABI platform to extract information from unlabeled text, generate answers, and understand user intentions. Combined with algorithms such as self-supervised learning, information extraction, and natural language processing, a more intelligent and flexible question-and-answer system based on unlabeled data can be built.

[0003] In the prior art, question-answering systems can only recognize and process limited natural language expressions and cannot understand complex sentence structures or uncommon expressions. If the user's question does not contain the system's preset keywords or sentence patterns, the system may not be able to provide an effective answer, or if the question exceeds the template scope, the system's answer may not have reference value. At the same time, the answers are standardized answers based on preset templates and rules, which are applicable to common questions and lack personalization or context sensitivity. Summary of the invention

[0004] To solve the above problems, the present invention provides an ABI question-answering method and system in the absence of labeled data.

[0005] The ABI question-answering method and system of the present invention in the case of unlabeled data adopts the following technical solution:

[0006] An embodiment of the present invention provides an ABI question-answering method in the case of unlabeled data, the method comprising the following steps:

[0007] Obtain the query content raised by the user on the same question and the files in the ABI management system database; obtain the keywords in the query content;

[0008] Obtain a query keyword sequence and a query keyword set; obtain the importance of each keyword in the query keyword set according to the position of the keyword in the query keyword sequence and the frequency of the keyword in the query content; obtain the matching degree of each file in the database with the query content according to the importance of the keyword and the number of occurrences of the keyword in the database file;

[0009] Obtain the rarity of each keyword in the query keyword set in each file in the database; obtain the core criticality of each keyword in each file in the database based on the matching degree, rarity and importance of the keyword; construct a feature matrix of the query content keywords and database files based on the core criticality and importance; establish a decision tree based on the importance and feature matrix to obtain the final relevant files of the query content;

[0010] Obtain the keyword vector of the query content and the keyword vector of each final relevant file; obtain the overall attention score of each final relevant file relative to the query content based on the query content keyword vector, the keyword vector of the final relevant file, the importance, the core criticality and the feature matrix; filter the answer files of the query content based on the overall attention score, and generate a reply to the query content.

[0011] Furthermore, the query content includes the most recent query question and historical query questions.

[0012] Furthermore, the acquisition of the query keyword sequence and the query keyword set; obtaining the importance of each keyword in the query keyword set according to the position of the keyword in the query keyword sequence and the frequency of the keyword in the query content, includes the following specific steps:

[0013] Arrange all keywords in the query content in chronological order to obtain a query keyword sequence; take a set of all different keywords in the query content as a query keyword set;

[0014]

[0015] Where, X i is the maximum sequence number corresponding to the ith keyword in the query keyword set in the query keyword sequence; N is the number of all keywords in the query keyword sequence; a i is the timing parameter of the ith keyword in the query keyword set. If the ith keyword appears in the most recent query question, the timing parameter is TH1. If the ith keyword appears in the historical query question, the timing parameter is TH2. TH1>TH2; F i is the frequency of the ith keyword in the query keyword set appearing in the query content; sigmoid() is the sigmoid function used for normalization processing; Z i is the importance of the i-th keyword in the query keyword set.

[0016] Furthermore, the method of obtaining the degree of match between each file in the database and the query content according to the importance of the keyword and the number of occurrences of the keyword in the database file includes the following specific steps:

[0017]

[0018] Where M is the number of keywords in the query keyword set; Z a is the importance of the ath keyword in the query keyword set; F j,a is the number of a-th keyword appearing in the j-th file in the database; P j is the matching degree between the jth file in the database and the query content.

[0019] Furthermore, the step of obtaining the rarity of each keyword in the query keyword set in each file in the database and obtaining the core criticality of each keyword in each file in the database according to the matching degree, rarity and importance of the keyword includes the following specific steps:

[0020] By using the TF-IDF statistical method, the TF-IDF value corresponding to the j-th file in the database of the b-th keyword in the query keyword set is obtained, and the TF-IDF value is used as the rarity of the j-th file in the database of the b-th keyword in the query keyword set;

[0021]

[0022] Where P j is the matching degree between the jth file in the database and the query content; TD b,j To query the rarity of the bth keyword in the keyword set in the jth document in the database; TD j is the average rarity of all keywords in the query keyword set in the jth document in the database; sigmoid() is the sigmoid function used for normalization processing; H j,b is the core criticality of the bth keyword in the jth file in the database.

[0023] Furthermore, the construction of a feature matrix of query content keywords and database files according to the core criticality and importance includes the following specific steps:

[0024] The product of the core criticality of the bth keyword in the jth file in the database and the importance of the bth keyword in the query keyword set is used as the feature value corresponding to the jth file and the bth keyword in the feature matrix to be constructed;

[0025] The feature values ​​corresponding to each file and each keyword in the feature matrix to be constructed are filled in according to the corresponding positions to obtain the feature matrix of the query content keywords and database files. The rows in the feature matrix to be constructed represent different files in the database, and the columns represent different keywords in the query content. The position where the rows and columns intersect is where the feature values ​​should be filled in.

[0026] Furthermore, the specific steps of establishing a decision tree according to the importance and feature matrix to obtain the final relevant files of the query content are as follows:

[0027] The keyword corresponding to the maximum importance value is used as the root node of the decision tree; the file whose characteristic value corresponding to the root node keyword is greater than the preset first threshold is used as the relevant file of the first-level branch node, and other keywords whose characteristic value corresponding to the relevant files of the first-level branch node is greater than the preset first threshold are used as the branch nodes of the first level; the branch nodes of other levels are established in sequence until all the keywords are traversed to obtain the decision tree, and the file whose characteristic value corresponding to the keyword of the leaf node of the last level of the decision tree is greater than the preset first threshold is used as the final relevant file of the query content.

[0028] Furthermore, the step of obtaining the query content keyword vector and the keyword vector of each final related file includes the following specific steps:

[0029] The TF-IDF method is used for the query content to convert all different keywords in the query content into numerical vector representations to obtain the query content keyword vector; the TF-IDF method is used for the f-th final related file to convert different keywords belonging to the query content into numerical vector representations to obtain the keyword vector of the f-th final related file.

[0030] Furthermore, the overall attention score of each final relevant file relative to the query content is obtained according to the query content keyword vector, the keyword vector, importance, core criticality and feature matrix of the final relevant file, including the following specific steps:

[0031]

[0032] Where Q is the query content keyword vector; K f is the keyword vector of the fth final relevant document; K f T is the transposed vector of the keyword vector of the fth final relevant file; Z1 is the importance vector of the query content keyword, and the specific method for obtaining the importance vector is as follows: according to the order of the keywords in the query content keyword vector, the importance of all different keywords in the query content is arranged in rows to obtain the importance vector of the query content keyword; H1 fis the core criticality vector of the keywords in the fth final relevant file. The specific method for obtaining the core criticality vector is as follows: according to the order of the keywords in the query content keyword vector, the core criticality of all different keywords in the fth final relevant file is arranged in rows to obtain the core criticality vector of the keywords in the fth final relevant file; ‖‖ is the binary norm, which is used to obtain the absolute value; sigmoid() is the sigmoid function, which is used for normalization processing; R f is the eigenvalue vector of the keyword in the fth final related file, and the specific method for obtaining the eigenvalue vector is as follows: in the feature matrix, the vector composed of the eigenvalues ​​of the row where the fth final related file is located is used as the eigenvalue vector of the keyword in the fth final related file; S f is the overall attention score of the fth final relevant document relative to the query content.

[0033] The present invention also proposes an ABI question-answering system in the case of unlabeled data, comprising a memory and a processor, wherein the processor executes a computer program stored in the memory to implement the steps of the aforementioned method.

[0034] The beneficial effects of the technical solution of the present invention are as follows: according to the present invention, the importance of keywords in the query keyword set and the matching degree of the files in the database with the query content can be accurately analyzed in the ABI question-answering system process, so as to construct a suitable decision tree and introduce an attention mechanism, which can handle more complex sentence structures or uncommon expressions, provide effective answers to the query content, improve the reference value of the answers, and enhance the personalization of the answers or the sensitivity of the context. When determining the importance of keywords in the query keyword set, the importance of each keyword in the query keyword set is determined by analyzing the position of the keywords in the query keyword sequence and the frequency of the keywords in the query content, thereby improving the scientificity, accuracy and objectivity of the importance of the keywords. When determining the matching degree of the files in the database with the query content, the matching degree is determined by analyzing the importance of the keywords and the number of occurrences of the keywords in the database files, thereby improving the scientificity, accuracy and objectivity of the matching degree. When establishing a decision tree, the decision tree is determined by analyzing the rarity of the keywords combined with the matching degree and importance, thereby improving the rationality of obtaining the final relevant files of the query content when establishing the decision tree, and providing a data basis for the subsequent attention mechanism. When determining the overall attention score of the final relevant documents relative to the query content, the query content keyword vector, the keyword vector of the final relevant documents, the importance, the core criticality and the feature matrix are fully considered to determine the overall attention score, thereby improving the comprehensiveness and accuracy of the overall attention score. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0036] Figure 1 A flowchart of the steps of an ABI question-answering method for unlabeled data provided by one embodiment of the present invention;

[0037] Figure 2 A feature matrix to be constructed for an ABI question-answering method in the absence of labeled data is provided in one embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the ABI question-and-answer method and system for unlabeled data proposed by the present invention, its specific implementation method, structure, features and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0039] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0040] The following is a detailed description of the specific scheme of the ABI question-and-answer method and system for unlabeled data provided by the present invention in conjunction with the accompanying drawings.

[0041] See also Figure 1 , which shows a flowchart of the steps of an ABI question-answering method in the case of unlabeled data provided by an embodiment of the present invention, the method comprising the following steps:

[0042] Step S001: Obtain the query content raised by the user on the same question and the files in the ABI management system database; obtain the keywords in the query content.

[0043] It should be noted that the main purpose of this embodiment is to respond to the query content raised by the user for the same question, and the query content is multiple query questions raised by the user. Before starting the analysis, first obtain relevant data.

[0044] Specifically, the query content raised by the user on the same question and the files in the ABI management system database are obtained.

[0045] It should be noted that the query content raised by the user for the same question includes the most recent query question and historical query questions, and the historical query questions include multiple query questions. In this embodiment, all query questions are collectively referred to as query content. Obtaining the query content and the file in the ABI management system database is an existing method and will not be described in detail.

[0046] Specifically, obtain the keywords in the query content as follows:

[0047] Use the TF-IDF (term frequency-inverse document frequency) extraction strategy to obtain keywords in the query content.

[0048] It should be noted that the existing method of obtaining the keywords in the query content through TF-IDF is not described in detail in this embodiment.

[0049] At this point, the keywords in the query content are obtained.

[0050] Step S002, obtain a query keyword sequence and a query keyword set; obtain the importance of each keyword in the query keyword set according to the position of the keyword in the query keyword sequence and the frequency of the keyword in the query content; obtain the matching degree between each file in the database and the query content according to the importance of the keyword and the number of occurrences of the keyword in the database file.

[0051] It should be noted that when a user asks a query, the ABI system needs to first search for documents with relevant content from the company's relevant database. In this process, due to the information explosion era, it is time-consuming and expensive to obtain labeled data, and company-related data is constantly increasing and expanding. Therefore, unlabeled data is used to build the ABI answer system. There is a large amount of information data in the database. To ensure that the query is answered accurately, the keywords in the query will first search for related files in the database, and then search again in combination with the query content above.

[0052] It should be noted that there is a post-placement rule for important content in the word order of Chinese natural language, with adjectives placed in front. The keywords at the end of the query are the core. At the same time, the closer the keywords in the query content before the current query are to the current query, the greater their importance and reference value. At the same time, the greater the frequency of occurrence, the greater their importance and reference value. For example, the query question is: "Application of machine learning in medicine", and the keyword sequence obtained according to the TF-IDF method is: "1: machine learning, 2: medical, 3: application", and the answer focuses on how to apply it. Therefore, by analyzing the position of the keywords in the query keyword sequence and the frequency of their occurrence in the query content, the importance of each keyword in the query keyword set is obtained.

[0053] Specifically, obtain the query keyword sequence and query keyword set as follows:

[0054] All keywords in the query content are arranged in chronological order to obtain a query keyword sequence; a set consisting of all different keywords in the query content is taken as a query keyword set.

[0055] Furthermore, the importance of each keyword in the query keyword set is obtained according to the position of the keyword in the query keyword sequence and the frequency of the keyword in the query content, as follows:

[0056]

[0057] Where, X i is the maximum sequence number corresponding to the ith keyword in the query keyword set in the query keyword sequence; N is the number of all keywords in the query keyword sequence; a i is the timing parameter of the ith keyword in the query keyword set. If the ith keyword appears in the most recent query question, the timing parameter is TH1. If the ith keyword appears in the historical query question, the timing parameter is TH2. TH1>TH2. In this embodiment, TH1=2 and TH2=1 are used for description. F i is the frequency of the ith keyword in the query keyword set appearing in the query content; sigmoid() is the sigmoid function used for normalization processing; Z i is the importance of the i-th keyword in the query keyword set.

[0058] It should be noted that in the above formula Indicates the position of the i-th keyword in the query keyword sequence. The later the position, the The larger the value, the higher the importance. i It further emphasizes whether the ith keyword appears in the most recent query question. If the ith keyword appears in the most recent query question, it also means that the later it appears, the higher its importance; the greater the frequency of the ith keyword appearing in the query content, the more important it is.

[0059] It should be noted that the importance of each keyword in the query keyword set has been obtained above, and the matching of each file in the ABI management system database with the query content is analyzed by combining the importance and the number of occurrences of the keyword in each file in the database. If the importance of the keyword in the query keyword set is greater and it appears more in the files in the database, it means that the current file matches the query content better.

[0060] Specifically, according to the importance of the keyword and the number of occurrences of the keyword in the database file, the matching degree between each file in the database and the query content is obtained, as follows:

[0061]

[0062] Where M is the number of keywords in the query keyword set; Z a is the importance of the ath keyword in the query keyword set; F j,a is the number of a-th keyword appearing in the j-th file in the database; P j is the matching degree between the jth file in the database and the query content.

[0063] It should be noted that Z a It indicates the importance of the a-th keyword in the query keyword set. If the importance of the a-th keyword is high and the number of a-th keyword appearing in the j-th file in the database is also large, it means that the j-th file in the database has a better match with the a-th keyword in the query content. By analyzing all the keywords in the query keyword set, the overall match between the j-th file in the database and the query content is obtained. The higher the match, the more helpful the current file is to the query content.

[0064] At this point, the matching degree between each file in the database and the query content is obtained.

[0065] Step S003, obtaining the rarity of each keyword in the query keyword set in each file in the database; obtaining the core criticality of each keyword in each file in the database based on the matching degree, rarity and importance of the keyword; constructing a feature matrix of the query content keywords and database files based on the core criticality and importance; establishing a decision tree based on the importance and feature matrix to obtain the final relevant files of the query content.

[0066] It should be noted that, according to the degree of match between the database file and the query content, a reference can be provided for the ABI question-and-answer system from the perspective of keywords. In order to improve the accuracy of the answer, the relevance of the file content also needs to be considered. The decision tree model can make decisions based on the input features (i.e., the keywords of the query content), so that the model can identify which features can best distinguish relevant files from irrelevant files. In the process of building a decision tree, the model will automatically select the features that can best improve the classification accuracy. The keywords extracted from the query content will be used to determine whether the file is relevant and evaluate the importance of the features, thereby giving priority to features that are highly relevant to the query content and finding more accurate relevant database files.

[0067] It should be further explained that when constructing a decision tree model, it is first necessary to determine the eigenvalues ​​in the feature matrix of query content keywords and database files, and then construct the feature matrix to establish a decision tree.

[0068] Specifically, the rarity of each keyword in the query keyword set in each file in the database is obtained as follows:

[0069] The specific method for obtaining the rarity is as follows: by using the TF-IDF (term frequency-inverse document frequency) statistical method, the TF-IDF value corresponding to the j-th file in the database for the b-th keyword in the query keyword set is obtained, and the TF-IDF value is used as the rarity of the j-th file in the database for the b-th keyword in the query keyword set. It should be noted that obtaining the TF-IDF value corresponding to the j-th file in the database for the b-th keyword in the query keyword set by using the TF-IDF (term frequency-inverse document frequency) statistical method is an existing method of the TF-IDF (term frequency-inverse document frequency) statistical method, which will not be repeated in this embodiment.

[0070] Furthermore, according to the matching degree, rarity and importance of the keyword, the core criticality of each keyword in each file in the database is obtained; according to the core criticality and importance, the eigenvalues ​​corresponding to each keyword in each file in the feature matrix to be constructed are obtained, as follows:

[0071]

[0072] Where P j is the matching degree between the jth file in the database and the query content; TD b,j To query the rarity of the bth keyword in the keyword set in the jth document in the database; TD j is the average rarity of all keywords in the query keyword set in the jth document in the database; sigmoid() is the sigmoid function used for normalization processing; H j,b is the core criticality of the bth keyword in the jth file in the database.

[0073] The product of the core criticality of the bth keyword in the jth file in the database and the importance of the bth keyword in the query keyword set is used as the feature value corresponding to the jth file and the bth keyword in the feature matrix to be constructed.

[0074] It should be noted that It indicates the rarity of the bth keyword in the jth file based on the matching degree between the jth file and the query content, reflecting the core criticality of the bth keyword in the jth file. By analyzing the importance of the bth keyword, the eigenvalue corresponding to the jth file and the bth keyword in the feature matrix to be constructed is comprehensively obtained. The larger the eigenvalue, the more important the keyword in the file.

[0075] Furthermore, a feature matrix of query content keywords and database files is constructed based on the feature values, as follows:

[0076] The feature values ​​corresponding to each file and each keyword in the feature matrix to be constructed are filled in according to the corresponding positions to obtain the feature matrix of the query content keywords and database files. The rows in the feature matrix to be constructed represent different files in the database, and the columns represent different keywords in the query content. The position where the rows and columns intersect is where the feature values ​​should be filled in.

[0077] It should be noted that, in order to better understand the feature matrix, please refer to Figure 2 , Figure 2 This is the feature matrix to be constructed in this embodiment. The rows in the matrix represent different files in the database, namely DOC1, DOC2..., and the columns represent different keywords in the query content, namely keyword 1, keyword 2..., and the position where the rows and columns intersect is where the feature values ​​should be filled in.

[0078] Furthermore, a decision tree is established based on the importance and feature matrix to obtain the final relevant files of the query content, as follows:

[0079] The keyword corresponding to the maximum importance value is used as the root node of the decision tree; the file whose characteristic value corresponding to the root node keyword is greater than the preset first threshold is used as the relevant file of the first-level branch node, wherein the preset first threshold is specifically 0.5; other keywords whose characteristic values ​​corresponding to the files related to the first-level branch nodes are greater than the preset first threshold are used as the branch nodes of the first level; the branch nodes of other levels are established in turn until all the keywords are traversed to obtain the decision tree, and the file whose characteristic value corresponding to the keyword of the leaf node of the last level of the decision tree is greater than the preset first threshold is used as the final relevant file of the query content.

[0080] It should be noted that keywords that have been used as nodes will no longer participate in building a decision tree.

[0081] At this point, the final relevant files of the query content are obtained.

[0082] Step S004, obtain the keyword vector of the query content and the keyword vector of each final relevant file; obtain the overall attention score of each final relevant file relative to the query content based on the query content keyword vector, the keyword vector of the final relevant file, the importance, the core criticality and the feature matrix; filter the answer files of the query content according to the overall attention score, and generate a reply to the query content.

[0083] It should be noted that in order to ensure that the ABI question-answering system can output the most accurate answer, it is necessary to combine the attention mechanism to further determine the reference degree of the database file to the query content. The attention mechanism is widely used in the field of natural language processing. By dynamically assigning weights to different elements of the input sequence, the model can selectively input the most important information. The attention mechanism is used in the question-answering system to dynamically assign weights to keywords to ensure that the database files with strong relevance are the most important reference files for answering query questions. Before starting the analysis, you first need to obtain the query content keyword vector and the keyword vector of each file in the database, and then analyze the similarity between the two to obtain the attention weight.

[0084] Specifically, the query content keyword vector and the keyword vector of each final related file are obtained as follows:

[0085] The TF-IDF method is used for the query content to convert all different keywords in the query content into numerical vector representations to obtain the query content keyword vector; the TF-IDF method is used for the f-th final related file to convert different keywords belonging to the query content into numerical vector representations to obtain the keyword vector of the f-th final related file.

[0086] It should be noted that obtaining the query content keyword vector and the keyword vector of each final related file by using the TF-IDF method is an existing method of the TF-IDF method, which will not be described in detail in this embodiment.

[0087] Furthermore, based on the query content keyword vector, the keyword vector, importance, core criticality and feature matrix of the final relevant files, the overall attention score of each final relevant file relative to the query content is obtained, as follows:

[0088]

[0089] Where Q is the query content keyword vector; K f is the keyword vector of the fth final relevant document; K f Tis the transposed vector of the keyword vector of the fth final relevant file; Z1 is the importance vector of the query content keyword, and the specific method for obtaining the importance vector is as follows: according to the order of the keywords in the query content keyword vector, the importance of all different keywords in the query content is arranged in rows to obtain the importance vector of the query content keyword; H1 f is the core criticality vector of the keywords in the fth final relevant file. The specific method for obtaining the core criticality vector is as follows: according to the order of the keywords in the query content keyword vector, the core criticality of all different keywords in the fth final relevant file is arranged in rows to obtain the core criticality vector of the keywords in the fth final relevant file; ‖‖ is the binary norm, which is used to obtain the absolute value; sigmoid() is the sigmoid function, which is used for normalization processing; R f is the eigenvalue vector of the keyword in the fth final related file, and the specific method for obtaining the eigenvalue vector is as follows: in the feature matrix, the vector composed of the eigenvalues ​​of the row where the fth final related file is located is used as the eigenvalue vector of the keyword in the fth final related file; S f is the overall attention score of the fth final relevant document relative to the query content.

[0090] It should be noted that ‖Z1·Q‖ means taking the importance of the query content keyword as the weight to obtain the modulus of the query content keyword vector, ‖H1 f ·K f ‖ means taking the core criticality of the keywords in the final relevant files as weights to obtain the modulus of the keyword vector of the final relevant files. It represents the cosine similarity between the keyword vectors in the query content and the final relevant files. The importance of the query content keywords and the core criticality of the keywords in the files are used as weight calculation files and query content respectively. When the formula is normalized, the larger the similarity is, the greater the attention weight is. By combining the eigenvalue vector, the overall attention score of each final relevant file relative to the query content is obtained.

[0091] Further, the answer documents of the query content are filtered according to the overall attention score, and the answer to the query content is generated, as follows:

[0092] An attention score threshold is preset, and the final relevant files with an overall attention score greater than the preset attention score threshold are used as answer files for the query content. Otherwise, they are not used as answer files for the query content. All answer files are integrated into a large language model (LLM for short), and a deep learning model GPT is used to generate natural language answer outputs based on the answer files. The generated natural language answers are displayed to users on the ABI management system, and users are provided with links to answer files related to the query content.

[0093] Through the above steps, an ABI question-answering method without labeled data is completed.

[0094] Another embodiment of the present invention provides an ABI question-answering system in the case of unlabeled data, the system comprising a memory and a processor, and when the processor executes a computer program stored in the memory, the processor performs the following operations:

[0095] Obtain the query content and files in the ABI management system database raised by the user for the same question; obtain the keywords in the query content; obtain the query keyword sequence and the query keyword set; obtain the importance of each keyword in the query keyword set according to the position of the keyword in the query keyword sequence and the frequency of the keyword in the query content; obtain the matching degree of each file in the database with the query content according to the importance of the keyword and the number of occurrences of the keyword in the database file; obtain the rarity of each keyword in the query keyword set in each file in the database; obtain the core criticality of each keyword in each file in the database according to the matching degree, rarity and importance of the keyword; construct a feature matrix of the query content keywords and database files according to the core criticality and importance; establish a decision tree according to the importance and feature matrix to obtain the final relevant files of the query content; obtain the query content keyword vector and the keyword vector of each final relevant file; obtain the overall attention score of each final relevant file relative to the query content according to the query content keyword vector, the keyword vector of the final relevant file, the importance, the core criticality and the feature matrix; filter the answer files of the query content according to the overall attention score, and generate a reply to the query content.

[0096] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An ABI question-answering method for unlabeled data, characterized in that: The method comprises the following steps: Obtain the query content raised by the user on the same question and the files in the ABI management system database; obtain the keywords in the query content; Obtain a query keyword sequence and a query keyword set; obtain the importance of each keyword in the query keyword set according to the position of the keyword in the query keyword sequence and the frequency of the keyword in the query content; obtain the matching degree of each file in the database with the query content according to the importance of the keyword and the number of occurrences of the keyword in the database file; Obtain the rarity of each keyword in the query keyword set in each file in the database; obtain the core criticality of each keyword in each file in the database based on the matching degree, rarity and importance of the keyword; construct a feature matrix of the query content keywords and database files based on the core criticality and importance; establish a decision tree based on the importance and feature matrix to obtain the final relevant files of the query content; Obtain the keyword vector of the query content and the keyword vector of each final relevant file; obtain the overall attention score of each final relevant file relative to the query content based on the query content keyword vector, the keyword vector of the final relevant file, the importance, the core criticality and the feature matrix; filter the answer files of the query content based on the overall attention score, and generate a reply to the query content.

2. According to the ABI question-answering method in the case of unlabeled data as described in claim 1, it is characterized in that: The query content includes the most recent query question and historical query questions.

3. According to claim 2, the ABI question-answering method for unlabeled data is characterized in that: The step of obtaining a query keyword sequence and a query keyword set and obtaining the importance of each keyword in the query keyword set according to the position of the keyword in the query keyword sequence and the frequency of the keyword in the query content includes the following specific steps: Arrange all keywords in the query content in chronological order to obtain a query keyword sequence; take a set of all different keywords in the query content as a query keyword set; Where, X i is the maximum sequence number corresponding to the ith keyword in the query keyword set in the query keyword sequence; N is the number of all keywords in the query keyword sequence; a i is the timing parameter of the ith keyword in the query keyword set. If the ith keyword appears in the most recent query question, the timing parameter is TH1. If the ith keyword appears in the historical query question, the timing parameter is TH2. TH1>TH2; F i is the frequency of the i-th keyword in the query keyword set appearing in the query content; sigmoid() is the sigmoid function, which is used for normalization processing; Z i is the importance of the i-th keyword in the query keyword set.

4. According to claim 1, the ABI question-answering method for unlabeled data is characterized in that: The method of obtaining the degree of match between each file in the database and the query content according to the importance of the keyword and the number of occurrences of the keyword in the database file includes the following specific steps: Where M is the number of keywords in the query keyword set; Z a is the importance of the ath keyword in the query keyword set; F j,a is the number of a-th keyword appearing in the j-th file in the database; P j is the matching degree between the jth file in the database and the query content.

5. According to the ABI question-answering method in the case of unlabeled data as described in claim 1, it is characterized in that: The method of obtaining the rarity of each keyword in the query keyword set in each file in the database and obtaining the core criticality of each keyword in each file in the database according to the matching degree, rarity and importance of the keyword includes the following specific steps: By using the TF-IDF statistical method, the TF-IDF value corresponding to the j-th file in the database of the b-th keyword in the query keyword set is obtained, and the TF-IDF value is used as the rarity of the j-th file in the database of the b-th keyword in the query keyword set; Where P j is the matching degree between the jth file in the database and the query content; TD b,j To query the rarity of the bth keyword in the keyword set in the jth document in the database; TD j is the average rarity of all keywords in the query keyword set in the jth document in the database; sigmoid() is the sigmoid function used for normalization processing; H j,b is the core criticality of the bth keyword in the jth file in the database.

6. According to the ABI question-answering method in the case of unlabeled data as described in claim 5, it is characterized in that: The specific steps of constructing a feature matrix of query content keywords and database files according to the core criticality and importance are as follows: The product of the core criticality of the bth keyword in the jth file in the database and the importance of the bth keyword in the query keyword set is used as the feature value corresponding to the jth file and the bth keyword in the feature matrix to be constructed; The feature values ​​corresponding to each file and each keyword in the feature matrix to be constructed are filled in according to the corresponding positions to obtain the feature matrix of the query content keywords and database files. The rows in the feature matrix to be constructed represent different files in the database, and the columns represent different keywords in the query content. The position where the rows and columns intersect is where the feature values ​​should be filled in.

7. According to the ABI question-answering method in the case of unlabeled data as claimed in claim 1, it is characterized in that: The specific steps of establishing a decision tree according to the importance and feature matrix to obtain the final relevant files of the query content are as follows: The keyword corresponding to the maximum importance value is used as the root node of the decision tree; the file whose characteristic value corresponding to the root node keyword is greater than the preset first threshold is used as the relevant file of the first-level branch node, and other keywords whose characteristic value corresponding to the relevant files of the first-level branch node is greater than the preset first threshold are used as the branch nodes of the first level; the branch nodes of other levels are established in sequence until all the keywords are traversed to obtain the decision tree, and the file whose characteristic value corresponding to the keyword of the leaf node of the last level of the decision tree is greater than the preset first threshold is used as the final relevant file of the query content.

8. According to the ABI question-answering method in the case of unlabeled data as claimed in claim 1, it is characterized in that: The specific steps of obtaining the query content keyword vector and the keyword vector of each final related file are as follows: The TF-IDF method is used for the query content to convert all different keywords in the query content into numerical vector representations to obtain the query content keyword vector; the TF-IDF method is used for the f-th final related file to convert different keywords belonging to the query content into numerical vector representations to obtain the keyword vector of the f-th final related file.

9. According to claim 6, the ABI question-answering method in the case of unlabeled data is characterized in that: The overall attention score of each final relevant file relative to the query content is obtained according to the query content keyword vector, the keyword vector of the final relevant file, the importance, the core criticality and the feature matrix, including the following specific steps: Where Q is the query content keyword vector; K f is the keyword vector of the fth final relevant document; K f T is the transposed vector of the keyword vector of the fth final relevant file; Z1 is the importance vector of the query content keyword, and the specific method for obtaining the importance vector is as follows: according to the order of the keywords in the query content keyword vector, the importance of all different keywords in the query content is arranged in rows to obtain the importance vector of the query content keyword; H1 f is the core criticality vector of the keywords in the fth final relevant file. The specific method for obtaining the core criticality vector is as follows: according to the order of the keywords in the query content keyword vector, the core criticality of all different keywords in the fth final relevant file is arranged in rows to obtain the core criticality vector of the keywords in the fth final relevant file; ‖‖ is the binary norm, which is used to obtain the absolute value; sigmoid() is the sigmoid function, which is used for normalization processing; R f is the eigenvalue vector of the keyword in the fth final related file, and the specific method for obtaining the eigenvalue vector is as follows: in the feature matrix, the vector composed of the eigenvalues ​​of the row where the fth final related file is located is used as the eigenvalue vector of the keyword in the fth final related file; S f is the overall attention score of the fth final relevant document relative to the query content.

10. An ABI question-answering system for unlabeled data, the system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is executed by a processor, the steps of an ABI question-and-answer method for unlabeled data are implemented as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Query result matching degree calculation method and device

    CN111221943A

  • Data query method and device, storage medium and electronic equipment

    CN116756290A

  • Battery cell and battery cell manufacturing method

    KR1020250082669A

  • Connecting natural and security language in the embedding space for better threat hunting and incident response

    US20240427879A1