An enterprise knowledge question answering system based on a large language model

Through the enterprise knowledge question and answer system based on the large language model, the problems of inaccurate user privacy leakage and information extraction are solved, information security and responses are achieved efficient and accurate, and the overall effect of enterprise knowledge question and answers is improved.

CN119691134BActive Publication Date: 2025-09-02HANGZHOU AI ASSISTANT INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510194353.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-09-02
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

In the existing enterprise knowledge Q&A system, the risk of user privacy information leakage is high, the information extraction is inaccurate, and the reference reply set sorting and screening mechanism is lacking, resulting in inaccurate and efficient reply.

Method used

An enterprise knowledge question and answer system based on large language models is adopted, including user question module, information extraction module, large language model module and knowledge reply module. Through privacy protection, feature keyword analysis, feature knowledge graph construction and large language model iteration, accurate reference reply sets are generated and knowledge value sorted.

Benefits of technology

Ensure the security of user information, improve the accuracy of information extraction and the accuracy of reply, and improve the efficiency and accuracy of enterprise knowledge Q&A.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691134B_ABST
    Figure CN119691134B_ABST
Patent Text Reader

Abstract

The present invention discloses an enterprise knowledge question-answering system based on a large language model, which relates to the technical field of text data processing and analysis. By establishing a user consultation node to collect enterprise knowledge consultation questions corresponding to users, and protecting the privacy of the collection environment of the user consultation node, characteristic analysis is performed on the enterprise knowledge consultation questions of each user, and a number of characteristic keywords of each enterprise knowledge consultation question are obtained. A characteristic knowledge graph of the several characteristic keywords belonging to each enterprise knowledge consultation question is established, and a large language model for analyzing enterprise knowledge consultation questions is constructed. The characteristic knowledge graph of each enterprise knowledge consultation question is input into the large language model to obtain a reference answer set for each enterprise knowledge consultation question, the reference answer set for each enterprise knowledge consultation question is ranked by knowledge value, and the final question answer set required for the enterprise knowledge consultation question is extracted from the reference answer set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text data processing and analysis, and in particular to an enterprise knowledge question-answering system based on a large language model. Background Art

[0002] Existing enterprise knowledge question-and-answer services often face the following challenges: The security of the consulting environment is difficult to guarantee when users ask questions, potentially leading to the leakage of user privacy during the collection process. Traditional information extraction methods often fail to accurately capture the core features of enterprise knowledge consultation questions raised by users, resulting in inaccurate analysis and answers. Furthermore, the reference answer sets generated by current large language models for processing enterprise knowledge consultation questions lack effective sorting and filtering mechanisms, resulting in potentially inaccurate and inefficient answer sets provided to users. Therefore, it is necessary to provide a new enterprise knowledge question-and-answer system. Summary of the Invention

[0003] In order to solve the above problems, the purpose of the present invention is to provide an enterprise knowledge question answering system based on a large language model.

[0004] The object of the present invention can be achieved by the following technical solutions: an enterprise knowledge question-answering system based on a large language model, comprising a cloud platform, wherein the cloud platform is communicatively connected to a user question module, an information extraction module, a large language model module, and a knowledge answer module;

[0005] The user question module is used to establish several user consultation nodes, collect the enterprise knowledge consultation questions corresponding to the user through each user consultation node, and protect the privacy of the collection environment of the user consultation node;

[0006] The information extraction module is used to perform feature analysis on each user's enterprise knowledge consulting question, obtain several feature keywords corresponding to each enterprise knowledge consulting question, and establish a feature knowledge graph corresponding to the several feature keywords belonging to each enterprise knowledge consulting question;

[0007] The large language model module is used to construct a large language model for analyzing enterprise knowledge consulting questions, and iterate the accuracy of the large language model. Then, the feature knowledge graph of each enterprise knowledge consulting question is input through the final large language model to obtain a reference answer set for each enterprise knowledge consulting question;

[0008] The knowledge answer module is used to sort the reference answer set of each enterprise knowledge consulting question by knowledge value, and extract the final question answer set required for the enterprise knowledge consulting question from the reference answer set.

[0009] Furthermore, the process of collecting the enterprise knowledge consultation questions corresponding to the user through each user consultation node includes:

[0010] Establish several user consultation nodes and number them as i, i = 1, 2, 3, ..., n, where n is a natural number greater than 0, assign a user consultation node to each user who needs enterprise knowledge consultation, obtain the user code of each user who needs enterprise knowledge consultation, let the user edit the enterprise knowledge consultation questions they need to consult, set the data collection period corresponding to each user consultation node, and in their respective data collection period, the user consultation node assigned to each user collects the enterprise knowledge consultation questions of the corresponding user.

[0011] Furthermore, the process of privacy protection for the collection environment of the user consultation node includes:

[0012] The data collection period when the user consultation node collects enterprise knowledge consultation is divided into several interval periods. In each interval period, it is determined whether the user consultation node has collected privacy-related data. If so, the collection environment in the corresponding interval period is marked as a privacy-to-protect environment. If not, the collection environment in the corresponding interval period is marked as a safe environment.

[0013] When the collection environment is a privacy-unprotected environment, the privacy-related data in the enterprise knowledge consulting problem is decomposed into several private data fragments. Differential privacy technology is used to protect the privacy of the privacy-unprotected environment. Noise fragments are added to the several private data fragments to convert the privacy-related data into protected data.

[0014] When the collection environment is a safe environment, no operation is performed.

[0015] Furthermore, the process of performing feature analysis on each user's enterprise knowledge consulting question and obtaining several feature keywords corresponding to each enterprise knowledge consulting question includes:

[0016] Perform text cleaning on each user's enterprise knowledge consulting question, and then convert each user's enterprise knowledge consulting question into a corresponding information paradigm text. Each user's information paradigm text is divided into several word sequences, and the word sequence parts of speech are marked. The word sequence parts of speech include nouns, verbs, and adjectives.

[0017] Count the word frequency of each word sequence and record it as TF(t, d). Set up a corpus that stores the information paradigm texts of all users. Get the inverse document frequency of each word sequence and record it as IDF(t). Get the word value weight of each word sequence in the information paradigm text to which it belongs and record the word value weight as Value(t). Then we have:

[0018] ;

[0019] Set the threshold for distinguishing whether word sequences of different parts of speech can be used as feature keywords;

[0020] The corresponding distinction thresholds for nouns, verbs and adjectives are denoted as val1, val2 and val3 respectively, and the sentence sequences whose part of speech is noun and whose sentence value weight Value(t)≥val1, the sentence sequences whose part of speech is verb and whose sentence value weight Value(t)≥val2, and the sentence sequences whose part of speech is adjective and whose sentence value weight Value(t)≥val3 are marked as the characteristic keywords in the information paradigm text to which the current sentence sequence belongs, thereby obtaining several characteristic keywords corresponding to each user's enterprise knowledge consulting question.

[0021] Furthermore, the process of establishing a feature knowledge graph corresponding to several feature keywords of each enterprise knowledge consulting question includes:

[0022] Construct several initial blank knowledge graphs, assign a blank knowledge graph to each user's corresponding enterprise knowledge consulting question, and perform entity recognition, relationship extraction, and attribute calibration on several feature keywords corresponding to each enterprise knowledge consulting question in the corresponding blank knowledge graph;

[0023] Through entity recognition, the consulting object type corresponding to each enterprise knowledge consulting question is located. The consulting object types include enterprise product consulting, enterprise service project consulting, enterprise department situation consulting, enterprise operation process consulting, and enterprise-related policy consulting;

[0024] Extract the relationships between all characteristic keywords in each enterprise knowledge consultation question, and then obtain all the consultation items corresponding to each enterprise knowledge consultation question. Each consultation item is used to consult on enterprise-related issues corresponding to a consultation object type.

[0025] The content of attribute calibration is: obtain the information description attributes and information formats of all the inquiry items corresponding to each enterprise knowledge inquiry question. If the information description attributes and information formats meet the preset standard description attributes and standard information formats respectively, no operation is performed. Otherwise, the information description attributes of the inquiry items are converted to standard description attributes, or the information formats of the inquiry items are converted to standard information formats.

[0026] After several feature keywords corresponding to the blank knowledge graph assigned to each enterprise knowledge consulting question complete entity recognition, relationship extraction and attribute calibration, the corresponding blank knowledge graph is converted into the final feature knowledge graph.

[0027] Furthermore, the process of building a large language model to analyze enterprise knowledge consulting problems and iterating the accuracy of the large language model includes:

[0028] Obtain historical response data for enterprise knowledge consulting questions, set a sample division ratio, and divide the historical response data into test sample data, training sample data, and verification sample data according to the sample division ratio. Use artificial intelligence technology to construct a large language model in the initial state for analyzing enterprise knowledge consulting questions. The large language model in the initial state includes a question input layer, a model optimization layer, and a response output layer.

[0029] The historical response data includes standard response texts for several inquiry items corresponding to enterprise knowledge consulting questions at historical time nodes. Test sample data is input into the question input layer of the large language model in the initial state, and the response output layer outputs the real-time response text corresponding to the test sample data. The real-time response texts and the standard response texts of the inquiry items included in the enterprise knowledge consulting questions are compared and analyzed to obtain the model prediction accuracy of the large language model in the initial state, which is recorded as γ. An iteration threshold is set, which is recorded as η.

[0030] If γ ≥ η, no operation is performed;

[0031] If γ<η, perform accuracy iteration on the large language model in the initial state;

[0032] Then the large language model in the initial state is constructed into the final large language model.

[0033] Furthermore, the process of inputting the characteristic knowledge graph of each enterprise knowledge consultation question into the final large language model to obtain the reference answer set for each enterprise knowledge consultation question includes:

[0034] The characteristic knowledge graph of the enterprise knowledge consulting questions corresponding to each current user is input into the final large language model, and the final large language model outputs the real-time reply text of several question consultation items included in each enterprise knowledge consulting question, and uses the real-time reply text as the reference reply of the corresponding enterprise knowledge consulting question, and integrates all the reference replies of each enterprise knowledge consulting question as their respective reference reply sets.

[0035] Furthermore, the process of ranking the reference answer sets for each enterprise knowledge consulting question by knowledge value and extracting the final question answer set required for the enterprise knowledge consulting question from the reference answer sets includes:

[0036] The reference answers in the reference answer set corresponding to the enterprise knowledge consultation questions are grouped according to the consultation object type, thereby generating several sets of value-to-be-ranked answer sets of different consultation object types. Each value-to-be-ranked answer set records a number of reference answers corresponding to the same consultation object type;

[0037] Merge the same reference responses in each set of responses to be ranked into one reference response, and use the frequency of this reference response as its corresponding value weight, thereby generating reference responses with several value weights in each set of responses to be ranked;

[0038] Sort all reference answers in each set of value-to-be-ranked answer sets from low to high according to their value weights, and then convert each set of value-to-be-ranked answer sets into corresponding value-increasingly ranked answer sets. Take the reference answer with the highest value weight in each set of value-increasingly ranked answer sets as the final question answer for the corresponding consulting object type in the value-increasing ranked answer sets, merge the final question answers corresponding to each consulting object type, and then extract the final question answer set required for the corresponding enterprise knowledge consulting question from the reference answer set.

[0039] Compared with the existing technology, the beneficial effects of the present invention are: through the privacy protection mechanism of the user question module, the information security of the user during the consultation process is ensured, the leakage of user privacy is effectively prevented, and the security of the user consultation environment is improved; the information extraction module can perform feature analysis on the user's enterprise knowledge consultation questions, extract feature keywords, and construct a feature knowledge graph, thereby providing accurate data input for subsequent problem analysis and answers; the large language model module constructs a more complete large language model through precision iteration to improve the analysis and understanding capabilities of enterprise knowledge consultation questions, so that the generated reference answer set is more in line with the user's actual consultation needs; the knowledge answer module can sort the reference answer set by knowledge value, and extract the final question answer set that best meets the user's needs, which improves the accuracy and efficiency of enterprise knowledge questions and answers to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a schematic diagram of the present invention. DETAILED DESCRIPTION

[0041] like Figure 1 As shown, an enterprise knowledge question-answering system based on a large language model includes a cloud platform, wherein the cloud platform is communicatively connected to a user question module, an information extraction module, a large language model module, and a knowledge answer module;

[0042] The user question module is used to establish several user consultation nodes, collect the enterprise knowledge consultation questions corresponding to the user through each user consultation node, and protect the privacy of the collection environment of the user consultation node;

[0043] The information extraction module is used to perform feature analysis on each user's enterprise knowledge consulting question, obtain several feature keywords corresponding to each enterprise knowledge consulting question, and establish a feature knowledge graph corresponding to the several feature keywords belonging to each enterprise knowledge consulting question;

[0044] The large language model module is used to construct a large language model for analyzing enterprise knowledge consulting questions, and iterate the accuracy of the large language model. Then, the feature knowledge graph of each enterprise knowledge consulting question is input through the final large language model to obtain a reference answer set for each enterprise knowledge consulting question;

[0045] The knowledge answer module is used to sort the reference answer set of each enterprise knowledge consulting question by knowledge value, and extract the final question answer set required for the enterprise knowledge consulting question from the reference answer set.

[0046] It should be further explained that, in the specific implementation process, several user consultation nodes are established, and each user consultation node collects the user's corresponding enterprise knowledge consultation questions, and the privacy protection process of the user consultation node collection environment includes:

[0047] Establish several user consultation nodes and number them. Let the number be i, then i = 1, 2, 3, ..., n, where n is a natural number greater than 0. Assign a user consultation node to each user who needs enterprise knowledge consultation.

[0048] Obtain the user code corresponding to each user who needs enterprise knowledge consultation. The user code serves as the user's unique identity authentication identifier. The user edits the enterprise knowledge consultation questions he needs to consult, sets the data collection period corresponding to each user consultation node, and in each data collection period, the user consultation node assigned to each user collects the corresponding user's enterprise knowledge consultation questions;

[0049] The data collection period when the user consultation node collects enterprise knowledge consultation is divided into several interval periods. In each interval period, it is determined whether the user consultation node has collected privacy-related data. If so, the collection environment in the corresponding interval period is marked as a privacy-to-protect environment. If not, the collection environment in the corresponding interval period is marked as a safe environment.

[0050] When the collection environment is a privacy-unprotected environment, the privacy-related data in the enterprise knowledge consulting problem is decomposed into several private data fragments. Differential privacy technology is used to protect the privacy of the privacy-unprotected environment. Noise fragments are added to the several private data fragments to convert the privacy-related data into protected data.

[0051] When the collection environment is a safe environment, no operation is performed.

[0052] Specifically, for each user consultation node, the operation of converting the privacy-related data into protected data is as follows:

[0053] Set the segmentation length of privacy-related data, and record the segmentation length as L 切分 ;

[0054] The data length of the privacy-related data in the user consultation node is recorded as L;

[0055] According to L and L 切分 , the privacy-related data is divided into N privacy data segments, and N is expressed as follows:

[0056] N=Ceiling(L / L 切分 );

[0057] Among them, Ceiling() is a rounding-up function, which is used to round up the calculation result in (). That is, when the calculation result in () is an integer, the integer is directly taken; when the calculation result in () is a decimal, the calculation result is rounded up to an integer;

[0058] Sequentially connect several private data segments, identify the connection point between each two private data segments as a position to be inserted, set a random function, use the random function to generate position numbers corresponding to the positions to be inserted in the current number of private data segments, and add noise segments at the corresponding positions to be inserted, thereby converting the privacy-related data into corresponding protected data;

[0059] It should be noted that by converting privacy-related data into protected data, user privacy leakage is effectively avoided.

[0060] It should be further explained that, in the specific implementation process, the process of performing feature analysis on each user's enterprise knowledge consulting question and obtaining several feature keywords corresponding to each enterprise knowledge consulting question includes:

[0061] Perform text cleaning on each user's enterprise knowledge consulting question, remove useless characters and stop words in the enterprise knowledge consulting question after text cleaning, and then convert each user's enterprise knowledge consulting question into the corresponding information paradigm text;

[0062] It should be noted that useless characters refer to characters that do not carry semantic information in enterprise knowledge consulting questions and may cause interference, usually including punctuation marks, mathematical symbols, special characters, numbers, and space symbols; stop words refer to words that appear frequently in enterprise knowledge consulting questions but contribute little to the meaning of the text. Different languages ​​and cultural backgrounds have different stop word lists, usually including several prepositions in English and function words, common pronouns, and conjunctions in Chinese;

[0063] Divide each user's information paradigm text into several word sequences;

[0064] Mark the parts of speech of word sequences, which include nouns, verbs, and adjectives;

[0065] Count the word frequency of each word sequence and record the word frequency as TF(t, d). The calculation formula is as follows:

[0066] ;

[0067] Set up a corpus that stores the information paradigm texts of all users. Obtain the inverse document frequency corresponding to each word sequence. The inverse document frequency is recorded as IDF(t). The calculation formula is as follows:

[0068] ;

[0069] Obtain the sentence value weight of each sentence sequence in the information paradigm text to which it belongs, and record the sentence value weight as Value(t). The calculation formula is as follows:

[0070] ;

[0071] Set the threshold for distinguishing whether word sequences of different parts of speech can be used as feature keywords;

[0072] The corresponding discrimination thresholds of nouns, verbs and adjectives are denoted as val1, val2 and val3 respectively, and val1, val2 and val3 are all real numbers between 0 and 1;

[0073] The sentence sequence whose part of speech is noun and whose sentence value weight Value(t)≥val1 is marked as the characteristic keyword in the information paradigm text to which the current sentence sequence belongs, the sentence sequence whose part of speech is verb and whose sentence value weight Value(t)≥val2 is marked as the characteristic keyword in the information paradigm text to which the current sentence sequence belongs, and the sentence sequence whose part of speech is adjective and whose sentence value weight Value(t)≥val3 is marked as the characteristic keyword in the information paradigm text to which the current sentence sequence belongs, thereby obtaining several characteristic keywords corresponding to each user's enterprise knowledge consulting question.

[0074] The nouns, verbs and adjectives with Value(t)<val1, Value(t)<val2 and Value(t)<val3 in each information paradigm text are marked as low-value words, and all low-value words are integrated as auxiliary information texts of the corresponding information paradigm text.

[0075] It should be noted that the word value weight is used to reflect the relative importance of the word sequence in the corresponding information paradigm text. That is, the higher the value of the word value weight, the more important the current word sequence is in the corresponding information paradigm text, and the higher its value in measuring the text features of the current information paradigm text.

[0076] It should be further explained that, in the specific implementation process, the process of establishing a feature knowledge graph corresponding to several feature keywords of each enterprise knowledge consulting question includes:

[0077] Construct several initial blank knowledge graphs, assign a blank knowledge graph to each user's corresponding enterprise knowledge consulting question, and perform entity recognition, relationship extraction, and attribute calibration on several feature keywords corresponding to each enterprise knowledge consulting question in the corresponding blank knowledge graph;

[0078] Locate the consulting object type corresponding to each enterprise knowledge consulting question through entity recognition;

[0079] The types of consulting objects include consulting on enterprise products, enterprise service projects, enterprise department conditions, enterprise operation process and enterprise-related policies;

[0080] The specific contents are as follows:

[0081] Set up corresponding feature word sets for enterprise product consultation, enterprise service project consultation, enterprise department situation consultation, enterprise operation process consultation and enterprise related policy consultation. The feature word sets record their corresponding mapping keywords, and the mapping keywords are used to match the corresponding feature keywords;

[0082] If the similarity between the characteristic keywords in the enterprise knowledge consultation question and the word segments of the mapped keywords in a certain characteristic word set exceeds the preset similarity threshold, the consultation object type corresponding to the current characteristic word set is marked to the current enterprise knowledge consultation question;

[0083] Otherwise, no action is taken;

[0084] The calculation of the segment similarity between the feature keyword and the mapping keyword is as follows:

[0085] Decompose the feature keyword and the mapping keyword into corresponding character items, obtain the number of character items in the feature keyword that are the same as the mapping keyword, and record this number as S; obtain the number of character items decomposed from the mapping keyword, and record this number as S';

[0086] Then, based on S and S', we get the segment similarity corresponding to the current feature keyword, and record the segment similarity as λ 相似 , then λ 相似 =S / S`, word segment similarity λ 相似 The value range of is (0, 1), and the similarity threshold is recorded as Yb;

[0087] If λ 相似 >Yb, then the consulting object type of the mapping keyword corresponding to the feature word set is marked to the enterprise knowledge consulting question to which the current feature keyword belongs. If λ 相似 ≤Yb, no operation is performed;

[0088] Then, the consulting object type corresponding to each enterprise knowledge consulting question is obtained;

[0089] Extract the relationships between all characteristic keywords in each enterprise knowledge consultation question, and then obtain all the consultation items corresponding to each enterprise knowledge consultation question. Each consultation item is used to consult on enterprise-related issues corresponding to a consultation object type.

[0090] The specific contents are as follows:

[0091] Set the matching expressions corresponding to several question consultation items. The matching expressions are composed as follows:

[0092] Feature keywords whose part of speech is verb + (relation matching mode 1) + feature keywords whose part of speech is adjective + (relation matching mode 2) + feature keywords whose part of speech is noun;

[0093] The first relationship matching mode is used to match the relationship between the feature keywords whose part of speech is a verb and the feature keywords whose part of speech is an adjective, thereby generating a verb-adjective phrase item, and the second relationship matching mode is used to match the relationship between the feature keywords whose part of speech is an adjective and the feature keywords whose part of speech is a noun, thereby generating a noun-adjective phrase item;

[0094] Combine the corresponding verb-adjective phrase items and adjective-noun phrase items under the same expression, and then obtain the corresponding question consultation items for each enterprise knowledge consultation question. The question consultation items are used to provide consultation on a specific type of question related to enterprise knowledge;

[0095] The content of attribute calibration is as follows: obtain the information description attributes and information formats of all the inquiry items corresponding to each enterprise knowledge inquiry question. If the information description attributes and information formats meet the preset standard description attributes and standard information formats respectively, no operation is performed. Otherwise, the information description attributes of the inquiry items are converted to standard description attributes, or the information formats of the inquiry items are converted to standard information formats.

[0096] After the entity recognition, relationship extraction and attribute calibration of several feature keywords corresponding to the blank knowledge graph assigned to each enterprise knowledge consulting question are completed, the corresponding blank knowledge graph is converted into the final feature knowledge graph. Each feature knowledge graph is used to perform consulting operations for several question consulting items corresponding to one enterprise knowledge consulting question.

[0097] It should be further explained that, in the specific implementation process, the process of building a large language model to analyze enterprise knowledge consulting problems and iterating the accuracy of the large language model includes:

[0098] Obtain historical response data for enterprise knowledge consultation questions, set a sample division ratio for the historical response data, and divide the historical response data into test sample data, training sample data, and verification sample data according to the sample division ratio;

[0099] By using artificial intelligence technology, a large language model in the initial state is constructed to analyze enterprise knowledge consulting questions. The large language model in the initial state includes a question input layer, a model optimization layer, and a response output layer.

[0100] The historical response data consists of standard response texts for enterprise knowledge consultation questions at historical time nodes, including several question consultation items, and the standard response texts record the response details data of the corresponding question consultation items;

[0101] Input the test sample data into the question input layer of the large language model in the initial state, and the answer output layer outputs the real-time answer text corresponding to the test sample data. Compare and analyze the real-time answer text and the standard answer text of the question consultation items included in the enterprise knowledge consultation questions, and then obtain the model prediction accuracy of the large language model in the initial state. The model prediction accuracy is recorded as γ. The expression formula of the model prediction accuracy is as follows:

[0102] γ=d1 / d2;

[0103] Among them, d1 is the number of times the large language model in the initial state correctly answers the question consultation items in the input enterprise knowledge consultation question, and d2 is the total number of question consultation items in the enterprise knowledge consultation question input by the large language model in the initial state;

[0104] Set an iteration threshold and record the iteration threshold as η;

[0105] If γ ≥ η, no operation is performed;

[0106] If γ<η, perform accuracy iteration on the large language model in the initial state;

[0107] The accuracy iteration process involves inputting some training sample data into the model optimization layer to train the large language model in its initial state, and then inputting validation sample data into the model input layer to obtain the real-time model fitting rate, model generalization rate, and model normalization rate.

[0108] If the model fitting rate, model generalization rate, and model normalization rate are within their respective safe numerical ranges, the final large language model is constructed. Otherwise, the amount of training sample data is continued to increase, and the accuracy of the large language model in the initial state is continued to be iterated until the model fitting rate, model generalization rate, and model normalization rate are within their respective safe numerical ranges.

[0109] It should be further explained that, in the specific implementation process, the process of inputting the characteristic knowledge graph of each enterprise knowledge consultation question through the final large language model to obtain the reference answer set for each enterprise knowledge consultation question includes:

[0110] The characteristic knowledge graph of the enterprise knowledge consulting questions corresponding to each current user is input into the final large language model, and the final large language model outputs the real-time reply text of several question consultation items included in each enterprise knowledge consulting question, and uses the real-time reply text as the reference reply of the corresponding enterprise knowledge consulting question, and integrates all the reference replies of each enterprise knowledge consulting question as their respective reference reply sets.

[0111] It should be further explained that, in the specific implementation process, the process of ranking the reference answer set for each enterprise knowledge consulting question by knowledge value and extracting the final question answer set required for the enterprise knowledge consulting question from the reference answer set includes:

[0112] The reference answers in the reference answer set corresponding to the enterprise knowledge consultation questions are grouped according to the consultation object type, thereby generating several sets of value-to-be-ranked answer sets of different consultation object types. Each value-to-be-ranked answer set records a number of reference answers corresponding to the same consultation object type;

[0113] Merge the same reference responses in each set of responses to be ranked into one reference response, and use the frequency of this reference response as its corresponding value weight, thereby generating reference responses with several value weights in each set of responses to be ranked;

[0114] Sort all reference answers in each set of value-to-be-ranked answer sets from low to high according to their value weights, and then convert each set of value-to-be-ranked answer sets into corresponding value-increasingly ranked answer sets. Take the reference answer with the highest value weight in each set of value-increasingly ranked answer sets as the final question answer for the corresponding consulting object type in the value-increasing ranked answer sets, merge the final question answers corresponding to each consulting object type, and then extract the final question answer set required for the corresponding enterprise knowledge consulting question from the reference answer set.

[0115] It should be noted that the final answer to the question records the answers to the corresponding enterprise knowledge consulting questions raised by the user, including enterprise product consultation, enterprise service project consultation, enterprise department situation consultation, enterprise operation process consultation and enterprise-related policy consultation, which is used to help users efficiently provide enterprise knowledge answers, improve the user's understanding of enterprise-related knowledge, and the accuracy of enterprise knowledge answers.

[0116] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. An enterprise knowledge question answering system based on a large language model, including a cloud platform, characterized in that: The cloud platform is communicatively connected to a user question module, an information extraction module, a large language model module, and a knowledge answer module; The user question module is used to establish several user consultation nodes, collect the enterprise knowledge consultation questions corresponding to the user through each user consultation node, and protect the privacy of the collection environment of the user consultation node; The information extraction module is used to perform feature analysis on each user's enterprise knowledge consulting question, obtain several feature keywords corresponding to each enterprise knowledge consulting question, and establish a feature knowledge graph corresponding to the several feature keywords belonging to each enterprise knowledge consulting question; The large language model module is used to construct a large language model for analyzing enterprise knowledge consulting questions, and iterate the accuracy of the large language model. Then, the feature knowledge graph of each enterprise knowledge consulting question is input through the final large language model to obtain a reference answer set for each enterprise knowledge consulting question; The knowledge answer module is used to sort the reference answer set of each enterprise knowledge consulting question by knowledge value, and extract the final question answer set required for the enterprise knowledge consulting question from the reference answer set; The process of establishing several user consultation nodes and collecting the enterprise knowledge consultation questions corresponding to users through each user consultation node includes: Establish several user consultation nodes and number them as i, i = 1, 2, 3, ..., n, where n is a natural number greater than 0. Assign a user consultation node to each user who needs enterprise knowledge consultation. Obtain the user code of each user who needs enterprise knowledge consultation. The user edits the enterprise knowledge consultation questions they need to consult. Set the data collection period corresponding to each user consultation node. During the respective data collection period, the user consultation node assigned to each user collects the enterprise knowledge consultation questions of the corresponding user. The process of privacy protection for the collection environment of the user consultation node includes: The data collection period when the user consultation node collects enterprise knowledge consultation is divided into several interval periods. In each interval period, it is determined whether the user consultation node has collected privacy-related data. If so, the collection environment in the corresponding interval period is marked as a privacy-to-protect environment. If not, the collection environment in the corresponding interval period is marked as a safe environment. When the collection environment is a privacy-unprotected environment, the privacy-related data in the enterprise knowledge consulting problem is decomposed into several private data fragments. Differential privacy technology is used to protect the privacy of the privacy-unprotected environment. Noise fragments are added to the several private data fragments to convert the privacy-related data into protected data. When the collection environment is a safe environment, no operation is performed.

2. The enterprise knowledge question answering system based on a large language model according to claim 1, characterized in that: The process of performing feature analysis on each user's enterprise knowledge consulting question and obtaining several feature keywords corresponding to each enterprise knowledge consulting question includes: Perform text cleaning on each user's enterprise knowledge consulting question, and then convert each user's enterprise knowledge consulting question into a corresponding information paradigm text. Each user's information paradigm text is divided into several word sequences, and the word sequence parts of speech are marked. The word sequence parts of speech include nouns, verbs, and adjectives. Count the word frequency of each word sequence and record it as TF(t, d). Set up a corpus that stores the information paradigm texts of all users. Get the inverse document frequency of each word sequence and record it as IDF(t). Get the word value weight of each word sequence in the information paradigm text to which it belongs and record the word value weight as Value(t). Then we have: ; Set the threshold for distinguishing whether word sequences of different parts of speech can be used as feature keywords; The corresponding distinction thresholds for nouns, verbs and adjectives are denoted as val1, val2 and val3 respectively, and the sentence sequences whose part of speech is noun and whose sentence value weight Value(t)≥val1, the sentence sequences whose part of speech is verb and whose sentence value weight Value(t)≥val2, and the sentence sequences whose part of speech is adjective and whose sentence value weight Value(t)≥val3 are marked as the characteristic keywords in the information paradigm text to which the current sentence sequence belongs, thereby obtaining several characteristic keywords corresponding to each user's enterprise knowledge consulting question.

3. The enterprise knowledge question answering system based on a large language model according to claim 2, characterized in that: The process of establishing a feature knowledge graph corresponding to several feature keywords for each enterprise knowledge consulting question includes: Construct several initial blank knowledge graphs, assign a blank knowledge graph to each user's corresponding enterprise knowledge consulting question, and perform entity recognition, relationship extraction, and attribute calibration on several feature keywords corresponding to each enterprise knowledge consulting question in the corresponding blank knowledge graph; Through entity recognition, the consulting object type corresponding to each enterprise knowledge consulting question is located. The consulting object types include enterprise product consulting, enterprise service project consulting, enterprise department situation consulting, enterprise operation process consulting, and enterprise-related policy consulting; Extract the relationships between all characteristic keywords in each enterprise knowledge consultation question, and then obtain all the consultation items corresponding to each enterprise knowledge consultation question. Each consultation item is used to consult on enterprise-related issues corresponding to a consultation object type. The content of attribute calibration is: obtain the information description attributes and information formats of all the inquiry items corresponding to each enterprise knowledge inquiry question. If the information description attributes and information formats meet the preset standard description attributes and standard information formats respectively, no operation is performed. Otherwise, the information description attributes of the inquiry items are converted to standard description attributes, or the information formats of the inquiry items are converted to standard information formats. After several feature keywords corresponding to the blank knowledge graph assigned to each enterprise knowledge consulting question complete entity recognition, relationship extraction and attribute calibration, the corresponding blank knowledge graph is converted into the final feature knowledge graph.

4. The enterprise knowledge question answering system based on a large language model according to claim 3, characterized in that: The process of building a large language model to analyze enterprise knowledge consulting problems and iterating the accuracy of the large language model includes: Obtain historical response data for enterprise knowledge consulting questions, set a sample division ratio, and divide the historical response data into test sample data, training sample data, and verification sample data according to the sample division ratio. Use artificial intelligence technology to construct a large language model in the initial state for analyzing enterprise knowledge consulting questions. The large language model in the initial state includes a question input layer, a model optimization layer, and a response output layer. The historical response data includes standard response texts for several inquiry items corresponding to enterprise knowledge consulting questions at historical time nodes. Test sample data is input into the question input layer of the large language model in the initial state, and the response output layer outputs the real-time response text corresponding to the test sample data. The real-time response texts and the standard response texts of the inquiry items included in the enterprise knowledge consulting questions are compared and analyzed to obtain the model prediction accuracy of the large language model in the initial state, which is recorded as γ. An iteration threshold is set, which is recorded as η. If γ ≥ η, no operation is performed; If γ<η, perform accuracy iteration on the large language model in the initial state; Then the large language model in the initial state is constructed into the final large language model.

5. The enterprise knowledge question answering system based on a large language model according to claim 4, characterized in that: The process of inputting the characteristic knowledge graph of each enterprise knowledge consultation question into the final large language model to obtain the reference answer set for each enterprise knowledge consultation question includes: The characteristic knowledge graph of the enterprise knowledge consulting questions corresponding to each current user is input into the final large language model, and the final large language model outputs the real-time reply text of several question consultation items included in each enterprise knowledge consulting question, and uses the real-time reply text as the reference reply of the corresponding enterprise knowledge consulting question, and integrates all the reference replies of each enterprise knowledge consulting question as their respective reference reply sets.

6. The enterprise knowledge question answering system based on a large language model according to claim 5, characterized in that: The process of ranking the reference answer set for each enterprise knowledge consulting question by knowledge value and extracting the final question answer set required for the enterprise knowledge consulting question from the reference answer set includes: The reference answers in the reference answer set corresponding to the enterprise knowledge consultation questions are grouped according to the consultation object type, thereby generating several sets of value-to-be-ranked answer sets of different consultation object types. Each value-to-be-ranked answer set records a number of reference answers corresponding to the same consultation object type; Merge the same reference responses in each set of responses to be ranked into one reference response, and use the frequency of this reference response as its corresponding value weight, thereby generating reference responses with several value weights in each set of responses to be ranked; Sort all reference answers in each set of value-to-be-ranked answer sets from low to high according to their value weights, and then convert each set of value-to-be-ranked answer sets into corresponding value-increasingly ranked answer sets. Take the reference answer with the highest value weight in each set of value-increasingly ranked answer sets as the final question answer for the corresponding consulting object type in the value-increasing ranked answer sets, merge the final question answers corresponding to each consulting object type, and then extract the final question answer set required for the corresponding enterprise knowledge consulting question from the reference answer set.

Citation Information

Patent Citations

  • Task-based dialogue system and implementation method thereof

    CN116911312A

  • Intelligent dialogue method, system and equipment based on large model and local knowledge base

    CN116992005A

  • Question answering method and device based on artificial intelligence, electronic equipment and storage medium

    CN117725188A

  • Electric power customer service system

    CN117891915A

Cited By

  • Intelligent question answering system based on language large model

    CN122045430A