Artificial intelligence-based user knowledge acquisition method and device, equipment and medium
By deploying an intelligent interactive platform on the server and utilizing task scheduling strategies and multimodal learning technology, the problem of low efficiency in knowledge data acquisition in existing technologies has been solved. This enables intelligent knowledge data acquisition based on user scenarios, improving retrieval efficiency and intelligence.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN CITY ZHITONG INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies cannot retrieve relevant knowledge data from the server based on the user's current knowledge acquisition scenario, and mainly rely on simple dialogue and question-and-answer methods, resulting in low retrieval efficiency and a lack of innovative exploration of new document processing technologies and knowledge extraction methods.
By deploying an intelligent interaction platform on the server, user interaction task information is obtained. A temporary knowledge engine library is obtained from the knowledge engine using a preset task scheduling strategy. Combined with multimodal learning and self-supervised learning techniques, document parsing, domain knowledge enhancement, structured segmentation, and document vectorization data acquisition are performed to achieve knowledge data acquisition in intelligent learning and intelligent tutoring scenarios.
It enables more accurate acquisition of target knowledge data based on user input data and interaction task type, expands the ways of acquiring knowledge data, and improves retrieval efficiency and intelligence.
Smart Images

Figure CN121599072B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to user knowledge acquisition methods, devices, equipment and media based on artificial intelligence. Background Technology
[0002] As enterprises grow in technology and organizational scale, they accumulate knowledge documents and store them on internal servers or data platforms for employees to access and consult. Currently, when storing knowledge documents on servers or data platforms, limited information such as document name, uploader, and upload time is typically extracted. When users log in to the server or data platform to retrieve specific knowledge data or documents, they can enter multiple keywords as search criteria. However, this keyword-based search method requires traversing all knowledge data on the server or data platform, resulting in low efficiency in retrieving search results.
[0003] Following this, knowledge acquisition methods based on artificial intelligence emerged. Enterprises draw experts from various professional and managerial fields to extract, label, and extract knowledge from documents generated during production and management processes, transforming unstructured data into structured data and storing it in the enterprise's digital system. Then, traditional knowledge extraction methods, such as NLP (Natural Language Processing), OCR (Optical Character Recognition), BERT (Bidirectional Encoder Representations from Transformers), keyword standardization, rule engines, and knowledge graphs, are used to organize, aggregate, classify, and build knowledge graphs. Finally, enterprise employees retrieve the necessary knowledge data through keyword searches and conditional queries. However, this method of accumulating and subsequently applying enterprise knowledge data has the following drawbacks:
[0004] 1) Enterprises need to purposefully and systematically accumulate documents in a unified format during the production management process so that experts can quickly label them and extract knowledge, thereby completing the acquisition of knowledge data. The above method lacks innovative exploration of new document processing technologies or knowledge extraction methods, such as combining multimodal learning, self-supervised learning or generative artificial intelligence technologies for document processing or knowledge extraction.
[0005] 2) The knowledge acquisition method mainly relies on strategies such as similarity matching and keyword routing. It cannot obtain relevant knowledge data from the server according to the user's current knowledge acquisition scenario. Instead, it can only obtain relevant knowledge data through simple dialogue and question-and-answer methods. Summary of the Invention
[0006] This invention provides a user knowledge acquisition method, apparatus, device, and medium based on artificial intelligence, aiming to solve the problem in the prior art that it is impossible to acquire relevant knowledge data from the server according to the user's current knowledge acquisition scenario, and that relevant knowledge data can only be acquired through simple dialogue and question-and-answer methods.
[0007] In a first aspect, embodiments of the present invention provide a user knowledge acquisition method based on artificial intelligence, comprising:
[0008] In response to a user interaction command sent by a user terminal, the system obtains current interaction task information corresponding to the user interaction command; wherein the current interaction task information includes at least current user authorization information and current interaction task type.
[0009] Based on the preset task scheduling strategy and the current interactive task information, the corresponding knowledge engine temporary library is obtained from the locally pre-built knowledge engine;
[0010] Obtain current user input data corresponding to the current interaction task information, and obtain current target knowledge data from the knowledge engine temporary library, the knowledge engine, or the server backend based on the current user input data and the current interaction task type; wherein, the preset interaction task type set corresponding to the current interaction task type includes at least intelligent learning scenario type and intelligent tutoring scenario type, and the current interaction task type is one of the preset interaction task type set;
[0011] If a smart interaction end command corresponding to the current interaction task information is detected, the current user input data and the current target knowledge data are saved to the user data storage space corresponding to the current interaction task information.
[0012] Secondly, embodiments of the present invention also provide an artificial intelligence-based user knowledge acquisition device, which includes a unit that performs the method described in the first aspect above.
[0013] Thirdly, embodiments of the present invention also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect above.
[0014] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the method described in the first aspect above.
[0015] This invention provides a user knowledge acquisition method, apparatus, device, and medium based on artificial intelligence. The method includes: responding to a user interaction command sent by a user terminal, acquiring current interaction task information corresponding to the user interaction command; wherein the current interaction task information includes at least current user authorization information and current interaction task type; acquiring a corresponding knowledge engine temporary library from a pre-built knowledge engine according to a preset task scheduling strategy and the current interaction task information; acquiring current user input data corresponding to the current interaction task information, and acquiring current target knowledge data from the knowledge engine temporary library, the knowledge engine, or the server backend according to the current user input data and the current interaction task type; wherein the preset set of interaction task types corresponding to the current interaction task type includes at least intelligent learning scenario type and intelligent tutoring scenario type, and the current interaction task type is one of the preset set of interaction task types; if an intelligent interaction end command corresponding to the current interaction task information is detected, the current user input data and the current target knowledge data are saved to the user data storage space corresponding to the current interaction task information. This invention enables more accurate acquisition of target knowledge data based on the current user input data and the interaction scenario corresponding to the current interaction task type during intelligent interaction between the user terminal and the server, expanding the methods for acquiring knowledge data. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram illustrating an application scenario of the user knowledge acquisition method based on artificial intelligence provided in an embodiment of the present invention.
[0018] Figure 2 A flowchart illustrating the user knowledge acquisition method based on artificial intelligence provided in an embodiment of the present invention;
[0019] Figure 3 A schematic diagram of a sub-process of the user knowledge acquisition method based on artificial intelligence provided in an embodiment of the present invention;
[0020] Figure 4This is a schematic diagram of another sub-process of the user knowledge acquisition method based on artificial intelligence provided in an embodiment of the present invention;
[0021] Figure 5 This is a schematic diagram of another sub-process of the user knowledge acquisition method based on artificial intelligence provided in an embodiment of the present invention;
[0022] Figure 6 This is another sub-process diagram of the user knowledge acquisition method based on artificial intelligence provided in the embodiments of the present invention;
[0023] Figure 7 A schematic block diagram of an artificial intelligence-based user knowledge acquisition device provided in an embodiment of the present invention;
[0024] Figure 8 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0027] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0028] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0029] Please also refer to Figure 1 and Figure 2 ,in Figure 1 This is a schematic diagram illustrating a scenario of the user knowledge acquisition method based on artificial intelligence according to an embodiment of the present invention. Figure 2This is a flowchart illustrating the user knowledge acquisition method based on artificial intelligence provided in an embodiment of the present invention. Figure 1 As shown, the user knowledge acquisition method based on artificial intelligence provided in this embodiment of the invention is applied to server 10, and user terminal 20 is communicatively connected to server 10. Figure 2 As shown, the method includes the following steps S110-S140.
[0030] S110. In response to a user interaction command sent by a user terminal, obtain the current interaction task information corresponding to the user interaction command.
[0031] The current interaction task information includes at least the current user authorization information and the current interaction task type.
[0032] In this embodiment, the technical solution is described with the server as the execution entity. An intelligent interaction platform is deployed on the server. After logging into the intelligent interaction platform by entering user login information (such as user registration account and user password) through a user terminal, the user can acquire corresponding knowledge based on the user's current usage needs (such as querying the knowledge base, AI tutors, AI practice, etc., where AI stands for Artificial Intelligence).
[0033] For example, taking a user who is an employee of a company and the intelligent interaction platform is deployed on the company's server, when a user has a need to acquire, learn, or master company-related knowledge, they can log in to the intelligent interaction platform using their user login information (if it is a new user's first login, they need to complete user registration before logging in). After successful verification, they can initiate the current interaction task request in the user interaction interface corresponding to the intelligent interaction platform displayed on the user's terminal. More specifically, there is an initial dialog box on this user interaction interface. In this initial dialog box, the user can initiate an initial task request by text input, voice input, etc. (e.g., the user enters "I now need to inquire about or learn relevant knowledge in the XX field" in text format). This initial task request, i.e., the user interaction instruction, is sent from the user terminal to the server (such as the knowledge engine in the intelligent interaction platform specifically on the server) in the form of a request content conforming to the Web API (Web API is a web application programming interface) specification. At the same time, the request content includes the current user authorization information corresponding to the user's login information (such as the de-identified user unique ID, the user's job information within the company, the user's company skill tags, etc.). Once the initial task request is received by the server and the request content is parsed and identified, the server can obtain current interactive task information, which includes at least the current user authorization information and the current interactive task type.
[0034] S120. Based on the preset task scheduling strategy and the current interactive task information, obtain the corresponding knowledge engine temporary library from the locally pre-built knowledge engine.
[0035] In this embodiment, once the current user's current interaction task information is determined in the server, in order to interact with the user on knowledge content more quickly, a Web API scheduling request corresponding to the current interaction task information can be initiated from the task scheduling center in the server to the executor in the server. The executor then calls the knowledge storage interface of the knowledge engine according to the Web API scheduling request, thereby obtaining a temporary knowledge engine library for the current user to perform intelligent interaction.
[0036] In one embodiment, before step S110 or step S120, the method further includes:
[0037] If a currently uploaded document file is detected and it is determined that the currently uploaded document file is a non-duplicate document file, then the document character layout, formula content, table content and chart content in the currently uploaded document file are obtained and detected based on the preset file parsing strategy to obtain the current document extraction result;
[0038] Based on a preset file reconstruction strategy, the currently uploaded document file is enhanced with domain knowledge, structured fragmentation, and document vectorization data acquisition to obtain the current file reconstruction result.
[0039] After establishing a mapping relationship between the current document extraction result, the current file reconstruction processing result, and the currently uploaded document file, all are stored in the knowledge base to update the knowledge base.
[0040] In this embodiment, when building the knowledge base in the knowledge engine on the server, the data source for its construction is document files uploaded by multiple users. For example, taking the server extracting knowledge from a currently uploaded document file to form one of the knowledge files in the knowledge base as an example, the server's knowledge base already stores knowledge data corresponding to other uploaded document files. When the server receives the currently uploaded document file, it needs to parse and process the file. When parsing the currently uploaded document file, optical character recognition technology can be used to obtain the document character layout (the chapter table of contents can be obtained because these chapter names and section names use a different font size than the main text), formula content (i.e., specifically locate and obtain all the formulas inserted in the currently uploaded document file, and count the application scenarios corresponding to each formula), table content (i.e., specifically locate and obtain all the tables inserted in the currently uploaded document file, generally each table will have its figure number marked above or below), and chart content (i.e., specifically locate and obtain all the charts inserted in the currently uploaded document file. The difference between charts and tables is that charts display other schematic diagrams that are not data tables, such as product structure diagrams, etc. Similarly, each chart will have its figure number marked above or below).
[0041] After the file parsing of the uploaded document is completed, the extracted document result can be obtained. Then, the file structure of the uploaded document can be reconstructed, specifically through domain knowledge enhancement, structured segmentation, and document vectorization. Domain knowledge enhancement specifically involves parsing the uploaded document's format (e.g., determining if it's PDF, Word, TXT, etc.), then performing word segmentation (e.g., dictionary-based segmentation) to obtain segmentation results including multiple keywords. Next, some keywords in the segmentation results are standardized and converted based on a pre-defined domain keyword library (where domain keywords are generally represented by Chinese terms, e.g., convolutional neural networks are represented by Chinese terms rather than the English abbreviation CNN). Finally, the document's domain theme and core intent can be identified to obtain a deeper understanding of the domain semantics. The document, after standardization based on the domain keyword library, constitutes the current domain knowledge enhancement result. The uploaded document file is structurally segmented, for example, by dividing the content based on the document's character layout information, such as into a table of contents, summary, and chapters, resulting in the current structured segmentation. When retrieving document vectorization data from the uploaded document file, its summary can be obtained first (if the document does not have a directly recorded summary, it can be generated based on the BERT model, where BERT stands for Bidirectional Encoder Representations from Transformers, a bidirectional language model based on transformers). Then, the semantic model corresponding to this summary is obtained as the document vectorization data. After knowledge extraction and document reconstruction processing of the uploaded document file, the obtained document extraction results, document reconstruction results, and the uploaded document file are mapped and stored in the knowledge base, enabling continuous updates to the knowledge base.
[0042] Of course, upon detecting a currently uploaded document file, a duplicate upload check can be performed first. This involves determining whether the currently uploaded document file is a previously uploaded file. If it is confirmed to be a previously uploaded file, a duplicate upload notification is sent to the user's terminal (this notification confirms whether to continue uploading the file or choose a different document file). If it is confirmed that the currently uploaded document file is not a previously uploaded file, it indicates that the duplicate upload has passed the check, and a successful upload notification can be sent to the user's terminal. This duplicate upload check effectively avoids duplicate processing of identical document files and the acquisition of duplicate knowledge data, reducing the consumption of server system resources.
[0043] In one embodiment, before step S110 or step S120, the method further includes:
[0044] If a currently uploaded document file is detected and it is determined that the currently uploaded document file is a non-duplicate document file, then knowledge extraction and knowledge graph construction are performed on the currently uploaded document file based on the pre-trained knowledge extraction model to obtain the current knowledge graph data;
[0045] Obtain the stored knowledge graph data and merge the current knowledge graph data into the stored knowledge graph data to update the stored knowledge graph data.
[0046] In this embodiment, the knowledge engine on the server can store not only the knowledge base but also knowledge graph-related data extracted from document files. Specifically, it performs knowledge extraction and knowledge graph construction on the currently uploaded document file based on a pre-trained knowledge extraction model. The pre-trained knowledge extraction model can be the BERT model, which can extract entities from the currently uploaded document file. Then, a relation classification model (such as a pre-trained language model, more specifically, the BERT model) is used to extract whether there are any relationships between the previously extracted entities. At this point, multiple triples are obtained, where each triple includes an entity pair and the relationship between the two entities in the entity pair. Finally, an attribute extraction model (such as the BERT model, more specifically, a classifier is connected to the output of the BERT model to realize the process of first identifying entities and candidate attribute names, then classifying and judging whether each candidate attribute name has a corresponding attribute value, and finally extracting the attribute value) can be used to extract attributes from the currently uploaded document file and supplement the corresponding attribute data (such as attribute names and attribute values) into the corresponding entities in the above-mentioned multiple triples, thereby completing the construction of the current knowledge graph data.
[0047] To integrate the current knowledge graph data into the existing knowledge graph data stored in the knowledge engine, the matching results between newly extracted entities in the current knowledge graph data and entities in the existing knowledge graph data are obtained. Entity nodes in the current knowledge graph data that share the same entity as those in the existing knowledge graph data are then integrated. Furthermore, the relationships and attributes of these fusionable entity nodes in the current knowledge graph data are also integrated into the existing knowledge graph data, resulting in the updated existing knowledge graph data and enabling continuous updates to the knowledge base. It is evident that the above method for knowledge extraction from document files effectively combines self-supervised learning techniques, making knowledge acquisition more intelligent.
[0048] S130. Obtain the current user input data corresponding to the current interaction task information, and obtain the current target knowledge data from the knowledge engine temporary library, the knowledge engine, or the server backend according to the current user input data and the current interaction task type.
[0049] The preset set of interactive task types corresponding to the current interactive task type includes at least intelligent learning scenario type and intelligent tutoring scenario type, and the current interactive task type is one of the preset set of interactive task types.
[0050] In this embodiment, after the current interaction task information generated by the user's operation using the user terminal is sent to the server, it can indicate the specific usage scenario of the user on the intelligent interaction platform. At this time, the intelligent interaction platform can switch to the corresponding current target interaction interface according to the current interaction task information. Then, on the current target interaction interface displayed on the user terminal, the user can further input current user input data (such as text data, voice data, image data, etc.) for subsequent knowledge data acquisition. After the server performs text recognition on the current user input data, it can be used as an initial request, initial frequently asked questions text, or initial non-frequently asked questions text to obtain the corresponding target knowledge data from the server's knowledge engine.
[0051] In the example above, the current user input data is used as an example, where the user enters a sentence of dialogue data in the current target interaction interface. The target knowledge data is then retrieved from the corresponding area based on the current user input data. However, as long as the current round of dialogue in the current target interaction interface has not ended, the target knowledge data is retrieved in the same way for each current user input data. Furthermore, when retrieving the target knowledge data for each current user input data in the corresponding area, the contextual semantics of the current user input data can be fully considered for comprehensive dialogue data analysis. More specifically, this can be achieved through contextual understanding based on server-side models (such as large language models), dynamic knowledge updates in multi-turn dialogues, or dialogue optimization based on reinforcement learning, thereby providing more accurate target knowledge data.
[0052] In one embodiment, such as Figure 3 As shown, in a first specific embodiment of step S130, step S130 includes:
[0053] S1311. If it is determined that the current interaction task type in the current interaction task information is an intelligent learning scenario type, then switch to the first current interaction interface corresponding to the intelligent learning scenario type.
[0054] S1312. Obtain the current user input data entered on the first current interactive interface;
[0055] S1313. Perform intent recognition on the current user input data according to the pre-deployed intent recognition agent to obtain the current intent recognition result; wherein, the intent recognition agent is equipped with an intent recognition model.
[0056] S1314. Determine the current target data acquisition area based on the current interaction type corresponding to the current intent recognition result, and acquire the current target knowledge data from the current target data acquisition area based on the current user input data and the preset knowledge data acquisition strategy.
[0057] In this embodiment, if the server determines that the current interaction task type in the current interaction task information is an intelligent learning scenario type, it means that the user needs to activate the AI tutor scenario of the intelligent interaction platform. At this time, the user can no longer stay on the main interaction page of the intelligent interaction platform, but switch to the corresponding first current interaction interface according to the current interaction task type. More specifically, if the current interaction task type is determined to be an intelligent learning scenario type, the user switches to the first current interaction interface corresponding to the intelligent learning scenario type. The first current interaction interface has a user dialog box. The user inputs current user input data in this dialog box using methods such as text input, voice input, or image import. As long as the current user input data is not text type data, it will be extracted by the speech recognition model or image recognition model in the server to obtain the current user input text data corresponding to the current user input data. If the current input data is text type data, no text recognition is required; it can be directly used as the current user input text data.
[0058] To enable faster knowledge data acquisition on the server's intelligent interaction platform, an intent recognition agent, a follow-up question agent, a question-answering agent, and a restricted word filtering agent can be pre-deployed in the first current interactive interface corresponding to the intelligent learning scenario type. After acquiring the current user input text data, the restricted word filtering agent (which specifically deploys a word segmentation model and a preset restricted word library) detects the presence of restricted words in the current user input text data. If the restricted word filtering agent determines that restricted words exist in the current user input text data, a preset fallback strategy is triggered to end the current interactive dialogue with the user. If the restricted word filtering agent determines that there are no restricted words in the current user input text data, the intent recognition agent performs intent recognition on the current user input data to obtain the current intent recognition result. Of course, after the intent recognition agent performs intent recognition on the current user input data, if it determines that the current intent recognition result is not any of the common question-answering type, knowledge question-answering type, or specified route type, the follow-up question agent can be activated to initiate a new round of dialogue with the user to more clearly obtain the user's explicit intent to acquire knowledge. After the intent recognition agent performs intent recognition on the current user input data, if it is determined that the current intent recognition result is one of the common question answering type, knowledge question answering type, and specified route type, then the current target data acquisition area (such as one of the knowledge engine temporary library, the knowledge engine, or the server backend) can be determined according to the current interaction type more specifically corresponding to the current intent recognition result, and the current target knowledge data can be acquired from the current target data acquisition area according to the current user input data and knowledge data acquisition strategy.
[0059] In one embodiment, such as Figure 4 As shown, step S1314 includes:
[0060] S13141. If it is determined that the current intent recognition result is an intent recognition result of a common question answer type, then the knowledge engine or the knowledge engine temporary library is used as the current target data acquisition area, and the first candidate knowledge data with a similarity exceeding a preset similarity threshold between the current user input data and the knowledge data acquisition strategy is obtained from the current target data acquisition area, and the first candidate answer data corresponding to the first candidate knowledge data is used as the current target knowledge data.
[0061] S13142. If it is determined that the current intent recognition result is a knowledge question answering type intent recognition result, then the knowledge engine or the knowledge engine temporary library is used as the current target data acquisition area. The current thinking chain decomposition result corresponding to the current user input data is obtained according to the knowledge data acquisition strategy. The thinking chain reply result corresponding to each current thinking chain data in the current thinking chain decomposition result is obtained from the current target data acquisition area and the current target knowledge data is formed.
[0062] S13143. If it is determined that the current intent recognition result is the intent recognition result of the specified route type, then the current route keyword corresponding to the current user input data is obtained, and the server backend is used as the current target data acquisition area, and the current target knowledge data corresponding to the previous route keyword is obtained from the current target data acquisition area according to the knowledge data acquisition strategy.
[0063] In this embodiment, referring to the example above, after determining that there are no restricted words in the current user input text data through the restricted word filtering agent, if the intent recognition agent performs intent recognition on the current user input data and determines that the current intent recognition result is a frequently asked question type (i.e., FAQ, which stands for Frequently Asked Questions), then the knowledge engine or the knowledge engine temporary library can be used as the current target data acquisition area. Then, according to the knowledge data acquisition strategy, the similarity (e.g., cosine similarity) between each knowledge data in the current target data acquisition area and the current user input data is calculated, and a first candidate knowledge data whose similarity to the current user input data exceeds a preset similarity threshold is determined. The knowledge engine or the knowledge engine temporary library stores a first candidate answer data that is a question-answer data pair with the first candidate knowledge data. At this time, the first candidate answer data is displayed as the current target knowledge data on the first current interactive interface for the user to view through the user terminal. Of course, there are still cases where the similarity between each knowledge data in the current target data acquisition area and the current user input data does not exceed the preset similarity. In this case, the user can be prompted to input rewritten data for the current user input data on the first current interactive interface (which can be regarded as a question rewriting operation), so that the user re-enters the updated current user input data. The current user input data re-entered by the user can refer to the processing process of step S13142, and the current user input data is treated as a complex request that needs to be broken down into thought chains for subsequent knowledge acquisition.
[0064] If, after the intent recognition agent performs intent recognition on the current user input data, and determines that the current intent recognition result is a knowledge question-and-answer type (this differs from common question-and-answer types; it indicates that the user needs to ask a more complex question and obtain a corresponding answer, and the answer cannot be determined by simply matching the similarity between the current user input data and the knowledge engine or the knowledge engine's temporary library), then the knowledge engine or the knowledge engine's temporary library can be used as the current target data acquisition area. Then, according to the knowledge data acquisition strategy, the current user input data is first decomposed into a thought chain to obtain a current thought chain decomposition result containing multiple current thought chain data. Then, the thought chain response result corresponding to each current thought chain data in the current thought chain decomposition result is obtained from the current target data acquisition area, and these responses are combined to form the current target knowledge data. This method allows users with complex knowledge question-and-answer needs to quickly obtain the current target knowledge data.
[0065] If the intent recognition agent performs intent recognition on the current user input data and determines that the current intent recognition result is a specified route type, such as the current user input data including the route keyword "check current user points", then the server backend is directly used as the current target data acquisition area. Then, according to the knowledge data acquisition strategy, the current target knowledge data corresponding to the previous route keyword is obtained from the current target data acquisition area. The obtained current target knowledge data (such as the user points corresponding to the user) can be filled into a preset reply text template according to a fixed pattern to obtain "Your current points are XX1 points" to update the current target knowledge data and display it on the first current interaction interface.
[0066] In one embodiment, such as Figure 5 As shown, step S13142 involves obtaining the current thought chain decomposition result corresponding to the current user input data according to the knowledge data acquisition strategy, and obtaining the thought chain response result corresponding to each current thought chain data in the current thought chain decomposition result from the current target data acquisition area, including:
[0067] S131421. Decompose the current user input data into a thought chain according to the thought chain decomposition sub-strategy in the knowledge data acquisition strategy to obtain the current thought chain decomposition result including multiple current thought chain data.
[0068] S131422. Call the pre-deployed question-answering agent locally, input the current thought chain decomposition result as a prompt word into the question-answering agent, and obtain the thought chain reply result corresponding to each current thought chain data in the current thought chain decomposition result.
[0069] In this embodiment, when the current user input data is decomposed into a thought chain using the thought chain decomposition sub-strategy, the current user input data is first decomposed and planned using a Transformer model that incorporates a key-value pair attention mechanism (this process is regarded as a sequence-to-sequence generation task), resulting in multiple intermediate statements. These intermediate statements correspond to multiple current thought chain data, and each current thought chain data corresponds to a question to be answered.
[0070] Subsequently, the decomposition results of the current thought chain, including multiple current thought chain data, are input into the question-answering agent. In specific implementations, the question-answering agent can employ a large language model. By performing knowledge retrieval for each current thought chain data within the knowledge within the large language model (i.e., stored in the knowledge engine or its temporary library) (e.g., referring to the similarity matching retrieval process of commonly used question-answering types in the knowledge base), the thought chain response result corresponding to each current thought chain data can be obtained.
[0071] In specific implementations, the question-answering agent can also embed multiple toolkits to expand its functionality. For example, the multiple toolkits include a knowledge retrieval toolkit, a course toolkit, and a live training course toolkit (where the course toolkit and the live training course toolkit interact with the question-answering agent through the Model Context Protocol). When the knowledge retrieval toolkit is called in the large language model to retrieve knowledge from the current thought chain data, the current thought chain data can be retrieved in multiple ways. For example, a first type of response result is obtained by matching the current thought chain data as a frequently asked question type with the knowledge base in the knowledge engine or the knowledge engine's temporary library; a second type of response result is obtained by matching the current structured segmentation result of the current thought chain data with the structured segmentation results of each knowledge data in the knowledge base, that is, determining the text structure similarity between the current structured segmentation result and the structured segmentation results of each knowledge data in the knowledge base, and selecting the knowledge data with the highest similarity to the current structured segmentation result as the second type of response result; a third type of response result is obtained by matching the current summary extraction result of the current thought chain data with the summary information of each knowledge data in the knowledge base, that is, determining the semantic similarity between the current summary extraction result and the summary extraction results of each knowledge data in the knowledge base, and selecting the knowledge data with the highest similarity to the current summary extraction result as the third type of response result; a fourth type of response result is obtained by matching the current word embedding vector of the current thought chain data with the semantic vector of each knowledge data in the knowledge base, that is, determining the vector similarity between the current word embedding vector and the semantic vector of each knowledge data in the knowledge base, and selecting the knowledge data with the highest similarity to the current word embedding vector as the fourth type of response result. Once at least four types of response results are obtained through the knowledge retrieval toolkit, they can be sorted in descending order based on the similarity of each type of response result, and the top three response results can be selected as the thought chain response results displayed in the user interaction interface corresponding to the question answering agent.
[0072] When retrieving information about corresponding courses or live training courses from the current thought chain data using the course toolkit and live training course toolkit in the large language model, the process involves determining whether the current thought chain data contains keywords for watching course or live training course videos. If such keywords are found, the search criteria are directly used: the course ID or live training course ID corresponding to the keywords. The retrieved course viewing link or live training course participation link is then obtained from the server. This retrieved link is displayed in the user interface corresponding to the question-and-answer agent for subsequent access by the user. The course knowledge data associated with the course viewing link can be viewed once or repeatedly by the user at any time; the live training course course data associated with the live training course participation link requires the user to view it within the corresponding live training course time slot. Therefore, this method allows for the presentation of the desired target knowledge data to the user in multiple ways.
[0073] In one embodiment, such as Figure 6 As shown, in a second specific embodiment of step S130, step S130 includes:
[0074] S1321. If it is determined that the current interaction task type in the current interaction task information is a smart coaching scenario type, then switch to the second current interaction interface corresponding to the smart coaching scenario type.
[0075] S1322. Start the pre-built digital human model and establish the association between the digital human model and the knowledge base;
[0076] S1323. If the current user input data entered by the user on the second current interactive interface is detected, the current target knowledge data is obtained from the knowledge base through the knowledge data acquisition strategy in the digital human model and sent to the user terminal.
[0077] S1324. If a user inputs an end-interaction command on the second current interaction interface, then return to the initial interface of the second current interaction interface.
[0078] In this embodiment, if the server determines that the current interaction task type in the current interaction task information is an intelligent tutoring scenario, it means that the user needs to activate the AI tutoring scenario of the intelligent interaction platform. At this time, the user can no longer stay on the main interaction page of the intelligent interaction platform, but switch to the corresponding second current interaction interface according to the current interaction task type. More specifically, if the current interaction task type is determined to be an intelligent tutoring scenario, the user switches to the second current interaction interface corresponding to the intelligent tutoring scenario type. In this second current interaction interface, a digital human model can be constructed. The knowledge base linked to its backend is the knowledge base in the server (more specifically, it can be linked to the knowledge engine temporary library in the knowledge base corresponding to the current tutoring scenario tag in the current interaction task information). The digital human model's facial image and voice style can adopt the system's default facial image and voice style, or multiple options for facial image and voice style can be provided on the second current interaction interface for the user to click and select. Before detecting the end-of-interaction command, the digital human model continuously collects the user's input data (mainly voice input) on the second current interaction interface. Based on the knowledge base or temporary knowledge engine library, it determines the next standard voice dialogue data for the current user input data in the current practice scenario. After multiple rounds of voice interaction between the user and the digital human model, a complete practice session for the current practice scenario can be completed. Therefore, through this intelligent practice method using a digital human model, users can engage in real-time voice interaction practice to acquire the required knowledge data.
[0079] In one embodiment, the method further includes the following after step S1324:
[0080] Acquire multiple current user input data and corresponding current target knowledge data saved through multiple rounds of interaction with the digital human model, and assemble them into current intelligent coaching comprehensive data according to the data collection time sequence;
[0081] Obtain standard intelligent coaching comprehensive data corresponding to the current intelligent coaching comprehensive data from the knowledge base, and determine the semantic similarity between the current intelligent coaching comprehensive data and the standard intelligent coaching comprehensive data, so as to serve as the current round of coaching evaluation result corresponding to the current intelligent coaching comprehensive data.
[0082] In this embodiment, after a user engages in multiple rounds of voice dialogue with a digital human model in an intelligent coaching scenario and the dialogue ends, the current user input data and corresponding current target knowledge data, arranged sequentially according to the data collection time, can be obtained to form the current intelligent coaching comprehensive data. To determine the user's mastery of the dialogue script for the current coaching scenario (such as standard introductory scripts for a product, such as those for automobiles, finance, or electronics), standard intelligent coaching comprehensive data corresponding to the current intelligent coaching comprehensive data and the current coaching scenario can be obtained from the knowledge base. Then, both the current intelligent coaching comprehensive data and the standard intelligent coaching comprehensive data are converted into corresponding semantic vectors, and the semantic similarity between the two semantic vectors is calculated. This semantic similarity serves as the current round of coaching evaluation result corresponding to the current intelligent coaching comprehensive data. Furthermore, suggestions for user coaching improvement can be provided based on the differences between the current intelligent coaching comprehensive data and the standard intelligent coaching comprehensive data. Therefore, the above method enables the rapid acquisition of intuitively displayed dialogue evaluation results after intelligent semantic interaction between the digital human model and the user in an intelligent coaching scenario.
[0083] S140. If an intelligent interaction end command corresponding to the current interaction task information is detected, the current user input data and the current target knowledge data are saved to the user data storage space corresponding to the current interaction task information.
[0084] In this embodiment, if a user clicks the "End Intelligent Interaction" button on the current user interaction interface of the intelligent interaction platform, an intelligent interaction end command is triggered. At this time, the server obtains the current user input data and the current target knowledge data from the user's complete intelligent interaction process and stores them in the user data storage space corresponding to the user's current interaction task information in the server, thus completing the data storage process after the intelligent interaction ends. The historical interaction data stored in the user data storage space related to the user can serve as the user's context data in the intelligent interaction platform to influence the output of subsequent intelligent interaction dialogues.
[0085] It is evident that the implementation of this method enables more accurate acquisition of target knowledge data during the intelligent interaction between the user terminal and the server, based on the current user input data and the interaction scenario corresponding to the current interaction task type, thus expanding the methods for acquiring knowledge data.
[0086] Figure 7 This is a schematic block diagram of a user knowledge acquisition device based on artificial intelligence provided in an embodiment of the present invention. Figure 7As shown, corresponding to the above-described AI-based user knowledge acquisition method, the present invention also provides an AI-based user knowledge acquisition device 100. This AI-based user knowledge acquisition device 100 includes a unit for executing the above-described AI-based user knowledge acquisition method. Please refer to... Figure 7 The AI-based user knowledge acquisition device 100 includes: a current interactive task acquisition unit 110, a temporary library construction unit 120, a current target knowledge data acquisition unit 130, and an interactive data storage unit 140.
[0087] The current interaction task acquisition unit 110 is used to acquire current interaction task information corresponding to the user interaction command sent by the user terminal in response to the user interaction command.
[0088] The current interaction task information includes at least the current user authorization information and the current interaction task type.
[0089] In this embodiment, the technical solution is described with the server as the execution entity. An intelligent interaction platform is deployed on the server. After logging into the intelligent interaction platform by entering user login information (such as user registration account and user password) through a user terminal, the user can acquire corresponding knowledge based on the user's current usage needs (such as querying the knowledge base, AI tutors, AI practice, etc., where AI stands for Artificial Intelligence).
[0090] For example, taking a user who is an employee of a company and the intelligent interaction platform is deployed on the company's server, when a user has a need to acquire, learn, or master company-related knowledge, they can log in to the intelligent interaction platform using their user login information (if it is a new user's first login, they need to complete user registration before logging in). After successful verification, they can initiate the current interaction task request in the user interaction interface corresponding to the intelligent interaction platform displayed on the user's terminal. More specifically, there is an initial dialog box on this user interaction interface. In this initial dialog box, the user can initiate an initial task request by text input, voice input, etc. (e.g., the user enters "I now need to inquire about or learn relevant knowledge in the XX field" in text format). This initial task request, i.e., the user interaction instruction, is sent from the user terminal to the server (such as the knowledge engine in the intelligent interaction platform specifically on the server) in the form of a request content conforming to the Web API (Web API is a web application programming interface) specification. At the same time, the request content includes the current user authorization information corresponding to the user's login information (such as the de-identified user unique ID, the user's job information within the company, the user's company skill tags, etc.). Once the initial task request is received by the server and the request content is parsed and identified, the server can obtain current interactive task information, which includes at least the current user authorization information and the current interactive task type.
[0091] The temporary library construction unit 120 is used to obtain the corresponding knowledge engine temporary library from the locally pre-built knowledge engine according to the preset task scheduling strategy and the current interactive task information.
[0092] In this embodiment, once the current user's current interaction task information is determined in the server, in order to interact with the user on knowledge content more quickly, a Web API scheduling request corresponding to the current interaction task information can be initiated from the task scheduling center in the server to the executor in the server. The executor then calls the knowledge storage interface of the knowledge engine according to the Web API scheduling request, thereby obtaining a temporary knowledge engine library for the current user to perform intelligent interaction.
[0093] In one embodiment, the AI-based user knowledge acquisition device 100 further includes:
[0094] The document extraction unit is used to detect the document character layout, formula content, table content and chart content in the currently uploaded document file based on a preset file parsing strategy if the currently uploaded document file is detected and it is determined that the currently uploaded document file is a non-duplicate document file, and obtain the current document extraction result.
[0095] The file reconstruction processing unit is used to perform domain knowledge enhancement, structured fragmentation, and document vectorization data acquisition on the currently uploaded document file based on a preset file reconstruction and rebuilding strategy, so as to obtain the current file reconstruction processing result.
[0096] The first knowledge base update control unit is used to establish a mapping relationship between the current document extraction result, the current file reconstruction processing result and the currently uploaded document file, and store them in the knowledge base to update the knowledge base.
[0097] In this embodiment, when building the knowledge base in the knowledge engine on the server, the data source for its construction is document files uploaded by multiple users. For example, taking the server extracting knowledge from a currently uploaded document file to form one of the knowledge files in the knowledge base as an example, the server's knowledge base already stores knowledge data corresponding to other uploaded document files. When the server receives the currently uploaded document file, it needs to parse and process the file. When parsing the currently uploaded document file, optical character recognition technology can be used to obtain the document character layout (the chapter table of contents can be obtained because these chapter names and section names use a different font size than the main text), formula content (i.e., specifically locate and obtain all the formulas inserted in the currently uploaded document file, and count the application scenarios corresponding to each formula), table content (i.e., specifically locate and obtain all the tables inserted in the currently uploaded document file, generally each table will have its figure number marked above or below), and chart content (i.e., specifically locate and obtain all the charts inserted in the currently uploaded document file. The difference between charts and tables is that charts display other schematic diagrams that are not data tables, such as product structure diagrams, etc. Similarly, each chart will have its figure number marked above or below).
[0098] After the file parsing of the uploaded document is completed, the extracted document result can be obtained. Then, the file structure of the uploaded document can be reconstructed, specifically through domain knowledge enhancement, structured segmentation, and document vectorization. Domain knowledge enhancement specifically involves parsing the uploaded document's format (e.g., determining if it's PDF, Word, TXT, etc.), then performing word segmentation (e.g., dictionary-based segmentation) to obtain segmentation results including multiple keywords. Next, some keywords in the segmentation results are standardized and converted based on a pre-defined domain keyword library (where domain keywords are generally represented by Chinese terms, e.g., convolutional neural networks are represented by Chinese terms rather than the English abbreviation CNN). Finally, the document's domain theme and core intent can be identified to obtain a deeper understanding of the domain semantics. The document, after standardization based on the domain keyword library, constitutes the current domain knowledge enhancement result. The uploaded document file is structurally segmented, for example, by dividing the content based on the document's character layout information, such as into a table of contents, summary, and chapters, resulting in the current structured segmentation. When retrieving document vectorization data from the uploaded document file, its summary can be obtained first (if the document does not have a directly recorded summary, it can be generated based on the BERT model, where BERT stands for Bidirectional Encoder Representations from Transformers, a bidirectional language model based on transformers). Then, the semantic model corresponding to this summary is obtained as the document vectorization data. After knowledge extraction and document reconstruction processing of the uploaded document file, the obtained document extraction results, document reconstruction results, and the uploaded document file are mapped and stored in the knowledge base, enabling continuous updates to the knowledge base.
[0099] Of course, upon detecting a currently uploaded document file, a duplicate upload check can be performed first. This involves determining whether the currently uploaded document file is a previously uploaded file. If it is confirmed to be a previously uploaded file, a duplicate upload notification is sent to the user's terminal (this notification confirms whether to continue uploading the file or choose a different document file). If it is confirmed that the currently uploaded document file is not a previously uploaded file, it indicates that the duplicate upload has passed the check, and a successful upload notification can be sent to the user's terminal. This duplicate upload check effectively avoids duplicate processing of identical document files and the acquisition of duplicate knowledge data, reducing the consumption of server system resources.
[0100] In one embodiment, the AI-based user knowledge acquisition device 100 further includes:
[0101] The current knowledge graph acquisition unit is used to perform knowledge extraction and knowledge graph construction on the current uploaded document file based on a pre-trained knowledge extraction model if the currently uploaded document file is detected and it is determined that the currently uploaded document file is a non-duplicate document file, thereby obtaining the current knowledge graph data.
[0102] The second knowledge base update control unit is used to acquire the stored knowledge graph data and merge the current knowledge graph data into the stored knowledge graph data to update the stored knowledge graph data.
[0103] In this embodiment, the knowledge engine on the server can store not only the knowledge base but also knowledge graph-related data extracted from document files. Specifically, it performs knowledge extraction and knowledge graph construction on the currently uploaded document file based on a pre-trained knowledge extraction model. The pre-trained knowledge extraction model can be the BERT model, which can extract entities from the currently uploaded document file. Then, a relation classification model (such as a pre-trained language model, more specifically, the BERT model) is used to extract whether there are any relationships between the previously extracted entities. At this point, multiple triples are obtained, where each triple includes an entity pair and the relationship between the two entities in the entity pair. Finally, an attribute extraction model (such as the BERT model, more specifically, a classifier is connected to the output of the BERT model to realize the process of first identifying entities and candidate attribute names, then classifying and judging whether each candidate attribute name has a corresponding attribute value, and finally extracting the attribute value) can be used to extract attributes from the currently uploaded document file and supplement the corresponding attribute data (such as attribute names and attribute values) into the corresponding entities in the above-mentioned multiple triples, thereby completing the construction of the current knowledge graph data.
[0104] To integrate the current knowledge graph data into the existing knowledge graph data in the knowledge engine, the matching results of newly extracted entities in the current knowledge graph data and entities in the existing knowledge graph data can be obtained. Entity nodes in the current knowledge graph data that are the same as those in the existing knowledge graph data are then integrated. The relationships and attributes of these fusionable entity nodes in the current knowledge graph data are also integrated into the existing knowledge graph data, thereby obtaining the updated existing knowledge graph data and enabling continuous updates to the knowledge base.
[0105] The current target knowledge data acquisition unit 130 is used to acquire the current user input data corresponding to the current interaction task information, and to acquire the current target knowledge data from the knowledge engine temporary library, the knowledge engine, or the server backend according to the current user input data and the current interaction task type.
[0106] The preset set of interactive task types corresponding to the current interactive task type includes at least intelligent learning scenario type and intelligent tutoring scenario type, and the current interactive task type is one of the preset set of interactive task types.
[0107] In this embodiment, after the current interaction task information generated by the user's operation using the user terminal is sent to the server, it can indicate the specific usage scenario of the user on the intelligent interaction platform. At this time, the intelligent interaction platform can switch to the corresponding current target interaction interface according to the current interaction task information. Then, on the current target interaction interface displayed on the user terminal, the user can further input current user input data (such as text data, voice data, image data, etc.) for subsequent knowledge data acquisition. After the server performs text recognition on the current user input data, it can be used as an initial request, initial frequently asked questions text, or initial non-frequently asked questions text to obtain the corresponding target knowledge data from the server's knowledge engine.
[0108] In one embodiment, as a first specific embodiment of the current target knowledge data acquisition unit 130, the current target knowledge data acquisition unit 130 is specifically used for:
[0109] If it is determined that the current interaction task type in the current interaction task information is an intelligent learning scenario type, then switch to the first current interaction interface corresponding to the intelligent learning scenario type;
[0110] Obtain the current user input data entered on the first current interactive interface;
[0111] The current user input data is processed by a pre-deployed intent recognition agent to identify the intent and obtain the current intent recognition result; wherein, the intent recognition agent is equipped with an intent recognition model.
[0112] The current target data acquisition area is determined based on the current interaction type corresponding to the current intent recognition result, and the current target knowledge data is acquired from the current target data acquisition area based on the current user input data and the preset knowledge data acquisition strategy.
[0113] In this embodiment, if the server determines that the current interaction task type in the current interaction task information is an intelligent learning scenario type, it means that the user needs to activate the AI tutor scenario of the intelligent interaction platform. At this time, the user can no longer stay on the main interaction page of the intelligent interaction platform, but switch to the corresponding first current interaction interface according to the current interaction task type. More specifically, if the current interaction task type is determined to be an intelligent learning scenario type, the user switches to the first current interaction interface corresponding to the intelligent learning scenario type. The first current interaction interface has a user dialog box. The user inputs current user input data in this dialog box using methods such as text input, voice input, or image import. As long as the current user input data is not text type data, it will be extracted by the speech recognition model or image recognition model in the server to obtain the current user input text data corresponding to the current user input data. If the current input data is text type data, no text recognition is required; it can be directly used as the current user input text data.
[0114] To enable faster knowledge data acquisition on the server's intelligent interaction platform, an intent recognition agent, a follow-up question agent, a question-answering agent, and a restricted word filtering agent can be pre-deployed in the first current interactive interface corresponding to the intelligent learning scenario type. After acquiring the current user input text data, the restricted word filtering agent (which specifically deploys a word segmentation model and a preset restricted word library) detects the presence of restricted words in the current user input text data. If the restricted word filtering agent determines that restricted words exist in the current user input text data, a preset fallback strategy is triggered to end the current interactive dialogue with the user. If the restricted word filtering agent determines that there are no restricted words in the current user input text data, the intent recognition agent performs intent recognition on the current user input data to obtain the current intent recognition result. Of course, after the intent recognition agent performs intent recognition on the current user input data, if it determines that the current intent recognition result is not any of the common question-answering type, knowledge question-answering type, or specified route type, the follow-up question agent can be activated to initiate a new round of dialogue with the user to more clearly obtain the user's explicit intent to acquire knowledge. After the intent recognition agent performs intent recognition on the current user input data, if it is determined that the current intent recognition result is one of the common question answering type, knowledge question answering type, and specified route type, then the current target data acquisition area (such as one of the knowledge engine temporary library, the knowledge engine, or the server backend) can be determined according to the current interaction type more specifically corresponding to the current intent recognition result, and the current target knowledge data can be acquired from the current target data acquisition area according to the current user input data and knowledge data acquisition strategy.
[0115] In one embodiment, the step of determining the current target data acquisition area based on the current interaction type corresponding to the current intent recognition result, and acquiring the current target knowledge data from the current target data acquisition area according to the current user input data and a preset knowledge data acquisition strategy, includes:
[0116] If the current intent recognition result is determined to be an intent recognition result of a common question answer type, then the knowledge engine or the knowledge engine temporary library is used as the current target data acquisition area, and the first candidate knowledge data with a similarity exceeding a preset similarity threshold between the current user input data and the knowledge data acquisition strategy is obtained from the current target data acquisition area, and the first candidate answer data corresponding to the first candidate knowledge data is used as the current target knowledge data.
[0117] If the current intent recognition result is determined to be a knowledge question answering type intent recognition result, then the knowledge engine or the knowledge engine temporary library is used as the current target data acquisition area. The current thought chain decomposition result corresponding to the current user input data is obtained according to the knowledge data acquisition strategy. The thought chain reply result corresponding to each current thought chain data in the current thought chain decomposition result is obtained from the current target data acquisition area and the current target knowledge data is formed.
[0118] If the current intent recognition result is determined to be the intent recognition result of the specified route type, then the current route keyword corresponding to the current user input data is obtained, and the server backend is used as the current target data acquisition area. The current target knowledge data corresponding to the previous route keyword is obtained from the current target data acquisition area according to the knowledge data acquisition strategy.
[0119] In this embodiment, referring to the example above, after determining that there are no restricted words in the current user input text data through the restricted word filtering agent, if the intent recognition agent performs intent recognition on the current user input data and determines that the current intent recognition result is a frequently asked question type (i.e., FAQ, which stands for Frequently Asked Questions), then the knowledge engine or the knowledge engine temporary library can be used as the current target data acquisition area. Then, according to the knowledge data acquisition strategy, the similarity (e.g., cosine similarity) between each knowledge data in the current target data acquisition area and the current user input data is calculated, and a first candidate knowledge data whose similarity to the current user input data exceeds a preset similarity threshold is determined. The knowledge engine or the knowledge engine temporary library stores a first candidate answer data that is a question-answer data pair with the first candidate knowledge data. At this time, the first candidate answer data is displayed as the current target knowledge data on the first current interactive interface for the user to view through the user terminal. Of course, there are still cases where the similarity between each knowledge data in the current target data acquisition area and the current user input data does not exceed the preset similarity. In this case, the user can be prompted to input rewritten data for the current user input data on the first current interactive interface (which can be regarded as a question rewriting operation), so that the user re-enters the updated current user input data. The current user input data re-entered by the user can refer to the processing process of step S13142, and the current user input data is treated as a complex request that needs to be broken down into thought chains for subsequent knowledge acquisition.
[0120] If, after the intent recognition agent performs intent recognition on the current user input data, and determines that the current intent recognition result is a knowledge question-and-answer type (this differs from common question-and-answer types; it indicates that the user needs to ask a more complex question and obtain a corresponding answer, and the answer cannot be determined by simply matching the similarity between the current user input data and the knowledge engine or the knowledge engine's temporary library), then the knowledge engine or the knowledge engine's temporary library can be used as the current target data acquisition area. Then, according to the knowledge data acquisition strategy, the current user input data is first decomposed into a thought chain to obtain a current thought chain decomposition result containing multiple current thought chain data. Then, the thought chain response result corresponding to each current thought chain data in the current thought chain decomposition result is obtained from the current target data acquisition area, and these responses are combined to form the current target knowledge data. This method allows users with complex knowledge question-and-answer needs to quickly obtain the current target knowledge data.
[0121] If the intent recognition agent performs intent recognition on the current user input data and determines that the current intent recognition result is a specified route type, such as the current user input data including the route keyword "check current user points", then the server backend is directly used as the current target data acquisition area. Then, according to the knowledge data acquisition strategy, the current target knowledge data corresponding to the previous route keyword is obtained from the current target data acquisition area. The obtained current target knowledge data (such as the user points corresponding to the user) can be filled into a preset reply text template according to a fixed pattern to obtain "Your current points are XX1 points" to update the current target knowledge data and display it on the first current interaction interface.
[0122] In one embodiment, the step of obtaining the current thought chain decomposition result corresponding to the current user input data according to the knowledge data acquisition strategy, and obtaining the thought chain response result corresponding to each current thought chain data in the current thought chain decomposition result from the current target data acquisition area, includes:
[0123] The current user input data is decomposed into a thought chain according to the thought chain decomposition sub-strategy in the knowledge data acquisition strategy, and the current thought chain decomposition result including multiple current thought chain data is obtained.
[0124] Call the locally pre-deployed question-answering agent, input the current thought chain decomposition result as a prompt word into the question-answering agent, and obtain the thought chain response result corresponding to each current thought chain data in the current thought chain decomposition result.
[0125] In this embodiment, when the current user input data is decomposed into a thought chain using the thought chain decomposition sub-strategy, the current user input data is first decomposed and planned using a Transformer model that incorporates a key-value pair attention mechanism (this process is regarded as a sequence-to-sequence generation task), resulting in multiple intermediate statements. These intermediate statements correspond to multiple current thought chain data, and each current thought chain data corresponds to a question to be answered.
[0126] Subsequently, the decomposition results of the current thought chain, including multiple current thought chain data, are input into the question-answering agent. In specific implementations, the question-answering agent can employ a large language model. By performing knowledge retrieval for each current thought chain data within the knowledge within the large language model (i.e., stored in the knowledge engine or its temporary library) (e.g., referring to the similarity matching retrieval process of commonly used question-answering types in the knowledge base), the thought chain response result corresponding to each current thought chain data can be obtained.
[0127] In specific implementations, the question-answering agent can also embed multiple toolkits to expand its functionality. For example, the multiple toolkits include a knowledge retrieval toolkit, a course toolkit, and a live training course toolkit (where the course toolkit and the live training course toolkit interact with the question-answering agent through the Model Context Protocol). When the knowledge retrieval toolkit is called in the large language model to retrieve knowledge from the current thought chain data, the current thought chain data can be retrieved in multiple ways. For example, a first type of response result is obtained by matching the current thought chain data as a frequently asked question type with the knowledge base in the knowledge engine or the knowledge engine's temporary library; a second type of response result is obtained by matching the current structured segmentation result of the current thought chain data with the structured segmentation results of each knowledge data in the knowledge base, that is, determining the text structure similarity between the current structured segmentation result and the structured segmentation results of each knowledge data in the knowledge base, and selecting the knowledge data with the highest similarity to the current structured segmentation result as the second type of response result; a third type of response result is obtained by matching the current summary extraction result of the current thought chain data with the summary information of each knowledge data in the knowledge base, that is, determining the semantic similarity between the current summary extraction result and the summary extraction results of each knowledge data in the knowledge base, and selecting the knowledge data with the highest similarity to the current summary extraction result as the third type of response result; a fourth type of response result is obtained by matching the current word embedding vector of the current thought chain data with the semantic vector of each knowledge data in the knowledge base, that is, determining the vector similarity between the current word embedding vector and the semantic vector of each knowledge data in the knowledge base, and selecting the knowledge data with the highest similarity to the current word embedding vector as the fourth type of response result. Once at least four types of response results are obtained through the knowledge retrieval toolkit, they can be sorted in descending order based on the similarity of each type of response result, and the top three response results can be selected as the thought chain response results displayed in the user interaction interface corresponding to the question answering agent.
[0128] When retrieving information about corresponding courses or live training courses from the current thought chain data using the course toolkit and live training course toolkit in the large language model, the process involves determining whether the current thought chain data contains keywords for watching course or live training course videos. If such keywords are found, the search criteria are directly used: the course ID or live training course ID corresponding to the keywords. The retrieved course viewing link or live training course participation link is then obtained from the server. This retrieved link is displayed in the user interface corresponding to the question-and-answer agent for subsequent access by the user. The course knowledge data associated with the course viewing link can be viewed once or repeatedly by the user at any time; the live training course course data associated with the live training course participation link requires the user to view it within the corresponding live training course time slot. Therefore, this method allows for the presentation of the desired target knowledge data to the user in multiple ways.
[0129] In one embodiment, as a second specific embodiment of the current target knowledge data acquisition unit 130, the current target knowledge data acquisition unit 130 is specifically used for:
[0130] If it is determined that the current interaction task type in the current interaction task information is the intelligent coaching scenario type, then switch to the second current interaction interface corresponding to the intelligent coaching scenario type;
[0131] Launch the pre-built digital human model and establish the association between the digital human model and the knowledge base;
[0132] If the user input data entered by the user on the second current interactive interface is detected, the current target knowledge data is obtained from the knowledge base through the knowledge data acquisition strategy in the digital human model and sent to the user terminal.
[0133] If a user inputs an end-interaction command on the second current interaction interface, the system returns to the initial interface of the second current interaction interface.
[0134] In this embodiment, if the server determines that the current interaction task type in the current interaction task information is an intelligent tutoring scenario, it means that the user needs to activate the AI tutoring scenario of the intelligent interaction platform. At this time, the user can no longer stay on the main interaction page of the intelligent interaction platform, but switch to the corresponding second current interaction interface according to the current interaction task type. More specifically, if the current interaction task type is determined to be an intelligent tutoring scenario, the user switches to the second current interaction interface corresponding to the intelligent tutoring scenario type. In this second current interaction interface, a digital human model can be constructed. The knowledge base linked to its backend is the knowledge base in the server (more specifically, it can be linked to the knowledge engine temporary library in the knowledge base corresponding to the current tutoring scenario tag in the current interaction task information). The digital human model's facial image and voice style can adopt the system's default facial image and voice style, or multiple options for facial image and voice style can be provided on the second current interaction interface for the user to click and select. Before detecting the end-of-interaction command, the digital human model continuously collects the user's input data (mainly voice input) on the second current interaction interface. Based on the knowledge base or temporary knowledge engine library, it determines the next standard voice dialogue data for the current user input data in the current practice scenario. After multiple rounds of voice interaction between the user and the digital human model, a complete practice session for the current practice scenario can be completed. Therefore, through this intelligent practice method using a digital human model, users can engage in real-time voice interaction practice to acquire the required knowledge data.
[0135] In one embodiment, the current target knowledge data acquisition unit 130 is further specifically used for:
[0136] Acquire multiple current user input data and corresponding current target knowledge data saved through multiple rounds of interaction with the digital human model, and assemble them into current intelligent coaching comprehensive data according to the data collection time sequence;
[0137] Obtain standard intelligent coaching comprehensive data corresponding to the current intelligent coaching comprehensive data from the knowledge base, and determine the semantic similarity between the current intelligent coaching comprehensive data and the standard intelligent coaching comprehensive data, so as to serve as the current round of coaching evaluation result corresponding to the current intelligent coaching comprehensive data.
[0138] In this embodiment, after a user engages in multiple rounds of voice dialogue with a digital human model in an intelligent coaching scenario and the dialogue ends, the current user input data and corresponding current target knowledge data, arranged sequentially according to the data collection time, can be obtained to form the current intelligent coaching comprehensive data. To determine the user's mastery of the dialogue script for the current coaching scenario (such as standard introductory scripts for a product, such as those for automobiles, finance, or electronics), standard intelligent coaching comprehensive data corresponding to the current intelligent coaching comprehensive data and the current coaching scenario can be obtained from the knowledge base. Then, both the current intelligent coaching comprehensive data and the standard intelligent coaching comprehensive data are converted into corresponding semantic vectors, and the semantic similarity between the two semantic vectors is calculated. This semantic similarity serves as the current round of coaching evaluation result corresponding to the current intelligent coaching comprehensive data. Furthermore, suggestions for user coaching improvement can be provided based on the differences between the current intelligent coaching comprehensive data and the standard intelligent coaching comprehensive data. Therefore, the above method enables the rapid acquisition of intuitively displayed dialogue evaluation results after intelligent semantic interaction between the digital human model and the user in an intelligent coaching scenario.
[0139] The interactive data storage unit 140 is used to save the current user input data and the current target knowledge data to the user data storage space corresponding to the current interactive task information if an intelligent interaction end command corresponding to the current interactive task information is detected.
[0140] In this embodiment, if a user clicks the "End Intelligent Interaction" button on the current user interaction interface of the intelligent interaction platform, an intelligent interaction end command is triggered. At this time, the server obtains the current user input data and the current target knowledge data from the user's complete intelligent interaction process and stores them in the user data storage space corresponding to the user's current interaction task information in the server, thus completing the data storage process after the intelligent interaction ends. The historical interaction data stored in the user data storage space related to the user can serve as the user's context data in the intelligent interaction platform to influence the output of subsequent intelligent interaction dialogues.
[0141] It is evident that the embodiments implementing this device can more accurately acquire target knowledge data based on the current user input data and the interaction scenario corresponding to the current interaction task type during the intelligent interaction between the user terminal and the server, thus expanding the methods for acquiring knowledge data.
[0142] The aforementioned AI-based user knowledge acquisition device can be implemented as a computer program, which can, for example... Figure 8 It runs on the computer device shown.
[0143] Please see Figure 8 , Figure 8 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. This computer device integrates any of the artificial intelligence-based user knowledge acquisition devices provided in the embodiments of the present invention.
[0144] See Figure 8 The computer device 400 includes a processor 402, a memory, and a network interface 405 connected via a system bus 401. The memory may include a storage medium 403 and internal memory 404.
[0145] The storage medium 403 may store an operating system 4031 and a computer program 4032. The computer program 4032 includes program instructions that, when executed, cause the processor 402 to perform an artificial intelligence-based user knowledge acquisition method.
[0146] The processor 402 is used to provide computing and control capabilities to support the operation of the entire computer device.
[0147] The internal memory 404 provides an environment for the computer program 4032 in the storage medium 403 to run. When the computer program 4032 is executed by the processor 402, the processor 402 can execute the above-mentioned user knowledge acquisition method based on artificial intelligence.
[0148] This network interface 405 is used for network communication with other devices. Those skilled in the art will understand that... Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0149] The processor 402 is used to run the computer program 4032 stored in the memory to implement the above-mentioned user knowledge acquisition method based on artificial intelligence.
[0150] It should be understood that, in this embodiment of the invention, the processor 402 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0151] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0152] Therefore, the present invention also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the aforementioned artificial intelligence-based user knowledge acquisition method.
[0153] The storage medium can be any computer-readable storage medium that can store program code, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0154] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0155] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0156] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0157] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0158] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A user knowledge acquisition method based on artificial intelligence, characterized in that, include: In response to a user interaction command sent by a user terminal, the system obtains current interaction task information corresponding to the user interaction command; wherein the current interaction task information includes at least current user authorization information and current interaction task type. Based on the preset task scheduling strategy and the current interactive task information, the corresponding knowledge engine temporary library is obtained from the locally pre-built knowledge engine; Obtain current user input data corresponding to the current interaction task information, and obtain current target knowledge data from the knowledge engine temporary library, the knowledge engine, or the server backend based on the current user input data and the current interaction task type; wherein, the preset interaction task type set corresponding to the current interaction task type includes at least intelligent learning scenario type and intelligent tutoring scenario type, and the current interaction task type is one of the preset interaction task type set; If a smart interaction end command corresponding to the current interaction task information is detected, the current user input data and the current target knowledge data are saved to the user data storage space corresponding to the current interaction task information. Before the step of obtaining current interaction task information corresponding to a user interaction command sent by a user terminal, or before the step of obtaining the corresponding knowledge engine temporary library from a locally pre-built knowledge engine according to a preset task scheduling strategy and the current interaction task information, the method further includes: If a currently uploaded document file is detected and it is determined that the currently uploaded document file is a non-duplicate document file, then the document character layout, formula content, table content and chart content in the currently uploaded document file are obtained and detected based on the preset file parsing strategy to obtain the current document extraction result; Based on a preset file reconstruction strategy, the currently uploaded document file is enhanced with domain knowledge, structured fragmentation, and document vectorization data acquisition to obtain the current file reconstruction result. After establishing a mapping relationship between the current document extraction result, the current file reconstruction processing result, and the currently uploaded document file, all are stored in the knowledge base to update the knowledge base; The step of obtaining the current target knowledge data from the knowledge engine temporary library, the knowledge engine, or the server backend based on the current user input data and the current interaction task type includes: If it is determined that the current interaction task type in the current interaction task information is an intelligent learning scenario type, then switch to the first current interaction interface corresponding to the intelligent learning scenario type; Obtain the current user input data entered on the first current interactive interface; The current user input data is processed by a pre-deployed intent recognition agent to identify the intent and obtain the current intent recognition result; wherein, the intent recognition agent is equipped with an intent recognition model. The current target data acquisition area is determined based on the current interaction type corresponding to the current intent recognition result, and the current target knowledge data is acquired from the current target data acquisition area based on the current user input data and the preset knowledge data acquisition strategy. The step of obtaining the current target knowledge data from the knowledge engine temporary library, the knowledge engine, or the server backend based on the current user input data and the current interaction task type includes: If it is determined that the current interaction task type in the current interaction task information is the intelligent coaching scenario type, then switch to the second current interaction interface corresponding to the intelligent coaching scenario type; Launch the pre-built digital human model and establish the association between the digital human model and the knowledge base; If the user input data entered by the user on the second current interactive interface is detected, the current target knowledge data is obtained from the knowledge base through the knowledge data acquisition strategy in the digital human model and sent to the user terminal. If a user inputs an end-interaction command on the second current interaction interface, the system returns to the initial interface of the second current interaction interface.
2. The method according to claim 1, characterized in that, The step of determining the current target data acquisition area based on the current interaction type corresponding to the current intent recognition result, and acquiring the current target knowledge data from the current target data acquisition area according to the current user input data and the preset knowledge data acquisition strategy, includes: If the current intent recognition result is determined to be an intent recognition result of a common question answer type, then the knowledge engine or the knowledge engine temporary library is used as the current target data acquisition area, and the first candidate knowledge data with a similarity exceeding a preset similarity threshold between the current user input data and the knowledge data acquisition strategy is obtained from the current target data acquisition area, and the first candidate answer data corresponding to the first candidate knowledge data is used as the current target knowledge data. If the current intent recognition result is determined to be a knowledge question answering type intent recognition result, then the knowledge engine or the knowledge engine temporary library is used as the current target data acquisition area. The current thought chain decomposition result corresponding to the current user input data is obtained according to the knowledge data acquisition strategy. The thought chain reply result corresponding to each current thought chain data in the current thought chain decomposition result is obtained from the current target data acquisition area and the current target knowledge data is formed. If the current intent recognition result is determined to be the intent recognition result of the specified route type, then the current route keyword corresponding to the current user input data is obtained, and the server backend is used as the current target data acquisition area. The current target knowledge data corresponding to the previous route keyword is obtained from the current target data acquisition area according to the knowledge data acquisition strategy.
3. The method according to claim 2, characterized in that, The step of obtaining the current thought chain decomposition result corresponding to the current user input data according to the knowledge data acquisition strategy, and obtaining the thought chain response result corresponding to each current thought chain data in the current thought chain decomposition result from the current target data acquisition area, includes: The current user input data is decomposed into a thought chain according to the thought chain decomposition sub-strategy in the knowledge data acquisition strategy, and the current thought chain decomposition result including multiple current thought chain data is obtained. Call the locally pre-deployed question-answering agent, input the current thought chain decomposition result as a prompt word into the question-answering agent, and obtain the thought chain response result corresponding to each current thought chain data in the current thought chain decomposition result.
4. The method according to claim 1, characterized in that, After the step of returning to the initial interface of the second current interactive interface if an end-interaction command is detected by the user on the second current interactive interface, the method further includes: Acquire multiple current user input data and corresponding current target knowledge data saved through multiple rounds of interaction with the digital human model, and assemble them into current intelligent coaching comprehensive data according to the data collection time sequence; Obtain standard intelligent coaching comprehensive data corresponding to the current intelligent coaching comprehensive data from the knowledge base, and determine the semantic similarity between the current intelligent coaching comprehensive data and the standard intelligent coaching comprehensive data, so as to serve as the current round of coaching evaluation result corresponding to the current intelligent coaching comprehensive data.
5. The method according to any one of claims 1-4, characterized in that, Before the step of obtaining current interaction task information corresponding to a user interaction command sent by a user terminal, or before the step of obtaining the corresponding knowledge engine temporary library from a locally pre-built knowledge engine according to a preset task scheduling strategy and the current interaction task information, the method further includes: If a currently uploaded document file is detected and it is determined that the currently uploaded document file is a non-duplicate document file, then knowledge extraction and knowledge graph construction are performed on the currently uploaded document file based on the pre-trained knowledge extraction model to obtain the current knowledge graph data; Obtain the stored knowledge graph data and merge the current knowledge graph data into the stored knowledge graph data to update the stored knowledge graph data.
6. A user knowledge acquisition device based on artificial intelligence, characterized in that, It includes a unit for performing the AI-based user knowledge acquisition method as described in any one of claims 1-5.
7. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the user knowledge acquisition method based on artificial intelligence as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions, which, when executed by a processor, can implement the user knowledge acquisition method based on artificial intelligence as described in any one of claims 1-5.
Citation Information
Patent Citations
Enabling communication with uniquely identifiable objects
US20210288927A1