Method, device and equipment for generating reply information
By identifying ambiguities in user requests and outputting clarifying guidance information in a natural language data interaction system, the problem of the system's inability to accurately understand user intent is solved, enabling efficient and accurate response generation, and improving user experience and system intelligence.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-31
AI Technical Summary
Natural language data interaction systems may fail to accurately understand the user's true intent, leading to incorrect or meaningless results that negatively impact user experience and decision-making quality.
By acquiring users' natural language data processing requests, ambiguity is identified, and when ambiguity is identified, clarification guidance information is output to guide users to clarify the target content. After obtaining the clarification information, a response is generated.
While maintaining a low interaction burden, it dynamically resolves semantic ambiguity, significantly improves the accuracy and efficiency of request responses, and enhances the practicality and intelligence of natural language data interaction systems.
Smart Images

Figure CN121766435A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data processing technology, and in particular to a method for generating response information. This specification also relates to an apparatus for generating response information, a computing device. Background Technology
[0002] In fields such as data analysis, business intelligence, and database querying, allowing users to submit data processing requests to the system using natural language (such as "Please check Beijing's sales figures for last month") has become a key technology for improving efficiency. However, the inherent ambiguity and polysemy of natural language often lead to ambiguous content in user requests; for example, "sales figures" may refer to "net sales figures" or "gross sales figures." If the system cannot accurately understand the user's true intent, it will directly result in returning incorrect or meaningless results, severely damaging the user experience and the quality of decision-making.
[0003] Therefore, there is a need to provide a method for generating response information that can eliminate ambiguity. Summary of the Invention
[0004] In view of this, one or more embodiments of this specification provide methods, apparatus, and devices for generating response information to solve the problem that existing natural language data interaction systems return incorrect or meaningless results because the system cannot accurately understand the user's true intention.
[0005] According to a first aspect of one or more embodiments of this specification, a method for generating response information is provided, comprising:
[0006] Obtain the user's data processing request in natural language for the target dataset;
[0007] Ambiguity identification is performed on the data processing request;
[0008] If ambiguity is identified in the data processing request, clarification guidance information is output; the clarification guidance information is used to guide the user to clarify the target content in the data processing request that causes ambiguity.
[0009] Obtain the clarification information input by the user regarding the target content;
[0010] Based on the data processing request and the clarification information, a response is output.
[0011] According to a second aspect of one or more embodiments of this specification, an apparatus for generating response information is provided, comprising:
[0012] The data processing request acquisition module is used to acquire users' data processing requests in natural language for the target dataset;
[0013] An ambiguity identification module is used to identify ambiguities in the data processing request.
[0014] The clarification guidance information output module is used to output clarification guidance information if the data processing request is found to be ambiguous; the clarification guidance information is used to guide the user to clarify the target content that causes ambiguity in the data processing request;
[0015] The clarification information acquisition module is used to acquire the clarification information input by the user regarding the target content;
[0016] The response information output module is used to output response information based on the data processing request and the clarification information.
[0017] According to a third aspect of one or more embodiments of this specification, a computing device is provided, including a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor, when executing the computer instructions, implements the steps of the method for generating response information.
[0018] One embodiment of this specification can achieve at least the following beneficial effects: by performing ambiguity identification on a user's data processing request in natural language form for a target dataset, and outputting clarification guidance information to guide the user to clarify the content causing ambiguity in the data processing request when ambiguity is identified, thereby obtaining the clarification information input by the user, and then generating and outputting response information based on the data processing request and the clarification information, thus, through intelligent ambiguity identification and accurate clarification guidance on the user's natural language data processing request, semantic ambiguity can be dynamically resolved while maintaining a low interaction burden, significantly improving the accuracy and efficiency of request response. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A schematic diagram illustrating an application scenario of a method for generating response information according to an embodiment of this specification is shown.
[0021] Figure 2 A flowchart illustrating a method for generating response information according to an embodiment of this specification is shown;
[0022] Figure 3This illustration shows a schematic diagram of a graphical user interface displaying clarification guidance information including clarification examples, according to an embodiment of this specification.
[0023] Figure 4 This illustration shows a diagram of a graphical user interface displaying clarification guidance information containing clarification examples, according to another embodiment of this specification.
[0024] Figure 5 This specification illustrates a schematic diagram of a graphical user interface displaying knowledge source prompts according to an embodiment of the present specification;
[0025] Figure 6 This specification illustrates a schematic diagram of an editing interface provided in an embodiment of the present specification, which displays an editing interface in a graphical user interface for editing the personal knowledge entries on which the generated response information is based;
[0026] Figure 7 This specification illustrates a flowchart of a data processing system responding to a user's data processing request in a practical application scenario, according to an embodiment of this specification.
[0027] Figure 8 This specification illustrates an embodiment corresponding to... Figure 2 A schematic diagram of a device for generating response information;
[0028] Figure 9 A structural block diagram of a computing device provided according to an embodiment of this specification is shown. Detailed Implementation
[0029] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0030] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0031] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “an,” “an,” “the,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification includes any or all possible combinations of one or more associated listed items.
[0032] The terms “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in the process, method, product, or apparatus that includes said elements is not excluded.
[0033] Although the terms "first," "second," etc., may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first," without departing from the scope of one or more embodiments of this specification. Ordinal numbers such as "first," "second," etc., do not necessarily indicate order; often they are used to facilitate the distinction of objects. For example, "first server" and "second server" usually refer to two servers. To distinguish these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.
[0034] Depending on the context, the word "if" as used here can be interpreted as "when," "when," or "in response to determination."
[0035] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.
[0036] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.
[0037] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation entry points shall be provided for users to choose to authorize or refuse.
[0038] The following explains the terms and concepts used in one or more embodiments of this specification.
[0039] Large Language Models (LLMs) are deep learning models trained on massive amounts of text data, enabling them to generate natural language text or understand the meaning of language text. LLMs can provide in-depth knowledge and language production on a wide range of topics through training on large datasets. They learn patterns and structures of natural language through large-scale unsupervised training, mimicking human language cognition and generation processes to some extent. LLMs employ a similar Transformer architecture and pre-training objectives as smaller models, with the main differences being increased model size, training data, and computational resources. Compared to traditional Natural Language Processing (NLP) models, LLMs better understand and generate natural text, while also exhibiting some logical thinking and reasoning abilities. LLMs possess in-context learning capabilities, enabling them to learn complex patterns in language and perform a wide range of tasks, including text summarization, translation, sentiment analysis, and multi-turn dialogue. For example, LLM can include the GPT series, T5 (Text-to-Text Transfer Transformer) model, PaLM model, BERT, LLaMA (Large Language Model MetaAI) model, Tongyi Qianwen model, Bailing model, etc.
[0040] In natural language data interaction systems, users often submit data processing requests in unstructured natural language (e.g., "What was the sales volume in Beijing last month?"). The system needs to parse these requests into executable data queries or analysis operations. However, due to the inherent semantic ambiguity, terminology diversity, and contextual dependencies in natural language, user expressions often contain ambiguities. For example, field references may be unclear (which indicator corresponds to "sales volume"), dimension values may be non-standard (does "Beijing" refer to "Beijing Municipality"), or the task logic may be complex (multi-step dependencies).
[0041] Traditional solutions mainly rely on two approaches: one is a static prompt mechanism, which provides general examples or help text in a fixed position on the interface. However, this approach requires users to actively click the prompt button, resulting in a fragmented interaction process. Furthermore, the preset examples cannot dynamically adapt to different query scenarios, leading to limited guidance effectiveness. The other approach is centralized knowledge base management, which requires users to jump to a separate configuration page to manually maintain field mappings or value normalization rules. This approach has high operational costs and a high learning curve, and rule updates require manual triggering to take effect, making it difficult to achieve immediate feedback and continuous optimization.
[0042] To address the shortcomings of related technologies, this specification proposes a method for generating response information for data processing requests in natural language form. This method can accurately identify ambiguous content in user requests and intelligently guide users to clarify, thereby generating accurate responses. This method overcomes the limitations of traditional static prompts or centralized knowledge bases, achieving high-precision, adaptive ambiguity clarification while maintaining a low interaction burden. It significantly improves the accuracy and efficiency of question-and-answer sessions, enhancing the practicality and intelligence of natural language data interaction systems.
[0043] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0044] Figure 1 This diagram illustrates an application scenario of a method for generating response information according to an embodiment of this specification.
[0045] like Figure 1 As shown, users can interact with terminal device 10 through a graphical user interface, and terminal device 10 can communicate with server 20 through a network.
[0046] Specifically, a user can issue a data processing request in natural language form for a target dataset through a graphical user interface (GUI). Then, a data processing system deployed on at least one of the terminal device 10 and server 20 can acquire the data processing request, perform ambiguity identification on the request, and if ambiguity is identified, output clarification guidance information to guide the user to clarify the target content causing the ambiguity in the data processing request. This clarification guidance information can be displayed to the user through the GUI of the terminal device 10. Furthermore, the clarification information input by the user regarding the target content can be acquired through the GUI, and based on the data processing request and the clarification information, a response to the data processing request can be output. This response information can also be displayed to the user through the GUI of the terminal device 10.
[0047] The method for generating response information provided in the embodiments of this specification can be applied to a data processing system that can run on a corresponding computing device. For example, the computing device may include a smart wearable device, smartphone, tablet computer, laptop computer, desktop computer, in-vehicle computer, server, or a server cluster or cloud computing service center composed of multiple servers, etc., and this specification does not specifically limit it in this regard.
[0048] like Figure 1 As shown, the method for generating response information provided in the embodiments of this specification can be applied to at least one of terminal device 10 or server 20. Figure 1 The terminal device 10 may include, but is not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Figure 1 The server 20 in the context can include, but is not limited to, any device, equipment, platform, or device cluster with computing and processing capabilities. In cases such as... Figure 1 In the application scenario shown, server 20 can connect to one or more terminal devices 10 via LAN connection, WAN connection, Internet connection or other types of data network.
[0049] Furthermore, the data processing system used in the method for generating response information provided in the embodiments of this specification can be based on a large language model. For example, it can be based on a large language model to perform ambiguity identification on the data processing request, generate clarification guidance information, and generate response information. Therefore, in practical applications, considering the large number of model parameters of the large language model and the limited computing resources of the terminal device 10, the method for generating response information provided in the embodiments of this specification can be mainly applied to the server 20, but is not limited to it. When the operating resources of the terminal device 10 can meet the deployment and operating conditions of the large language model, the method for generating response information provided in the embodiments of this specification can be performed on the terminal device 10.
[0050] The solution based on the embodiments in this specification effectively solves the parsing error problem caused by ambiguous user expressions or non-standard terminology by embedding a proactive ambiguity identification and interactive clarification mechanism into the natural language data processing flow. Specifically, after receiving a user's natural language data processing request for a target dataset, the system does not directly perform parsing but first identifies ambiguity in the request. When ambiguity is detected, it proactively outputs clarification guidance information to guide the user to clarify the target content causing the ambiguity, and generates a response based on the user's feedback. This mechanism avoids the shortcomings of traditional systems that blindly guess or directly report errors when faced with ambiguous requests, significantly improving the accuracy and reliability of the response results and also improving processing efficiency. At the same time, since the clarification guidance is triggered only when ambiguity is identified and focuses on specific target content, it ensures necessary human-computer collaboration while avoiding unnecessary interactive interference, thus maintaining a good user experience while ensuring parsing accuracy.
[0051] In the embodiments of this specification, a method for generating response information is provided. This application also relates to an apparatus for generating response information. A computing device will be described in detail in the following embodiments.
[0052] Figure 2 A flowchart illustrating a method for generating response information according to an embodiment of this specification is shown.
[0053] From a programming perspective, the entity executing the process can be a program hosted on a server or terminal device. It can be understood that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.
[0054] like Figure 2 As shown, the process may include the following steps:
[0055] Step 202: Obtain the user's data processing request in natural language form for the target dataset.
[0056] In this specification, "user" can refer to an entity that interacts with the methods or data processing systems provided in the embodiments of this specification. In the embodiments of this specification, the "user" in step 202 can represent an operator with data processing needs who initiates a request via natural language; this could be a data analyst, business personnel, or any end-user.
[0057] A dataset can represent a structured collection of data, typically organized in a specific way to facilitate computer storage, management, access, and analysis. In the embodiments of this specification, the specific form of the dataset in step 202 can be a database (such as a relational database like MySQL or a non-relational database like MongoDB), a table in a data warehouse, a file in a data lake, a data collection returned by an API interface with clearly defined fields, etc. The term "target" is used as a distinguishing prefix to emphasize that the dataset is a specific, predefined, and structured data source defined by the current session or request, distinguishing it from other unrelated datasets.
[0058] Natural language processing (NLP) can represent unstructured text expressed by users in everyday spoken or written language, without adhering to machine-readable formats such as SQL, JSON, and GraphQL. NLP information offers high flexibility but may contain ambiguity. An example of a NLP data processing request is: "What were the best-selling products in Beijing last month?" or "Compare the sales performance of high-profit products in Beijing and Shanghai in the first half of the year, and then predict the trend for the next quarter."
[0059] Data processing can include, but is not limited to, operations such as data querying (data retrieval), data aggregation (summation, averaging), data filtering, grouping, sorting, simple analysis (comparison, trend analysis), and visualization. In the embodiments of this specification, the data processing request may be directed to the target dataset. The data processing request can represent an instruction or request from a user to the system, expressing a desire to perform one or more of the aforementioned data processing operations.
[0060] In practical applications, step 202 can be implemented through an intelligent interaction module integrated into a data analysis platform or business intelligence (BI) tool. For example, a natural language input box or voice input button can be presented on the application front end, thereby receiving text strings or voice streams input by the user. Then, the received user input information, along with necessary contextual information (such as user identification, current session identification, and optionally, target dataset identification), can be encapsulated into a structured request object. This request object is then sent to the request parsing and processing engine on the backend server via a network request (such as HTTP / HTTPS protocol), thereby triggering subsequent steps.
[0061] Step 204: Perform ambiguity identification on the data processing request.
[0062] Ambiguity can refer to the fact that, given the context of a target dataset and conventional business logic, a data processing request in natural language form may correspond to multiple reasonable but potentially different computer interpretations that result in different data processing operations or outputs. Alternatively, ambiguity can refer to a semantic mismatch or polysemy between the natural language description and the structure of the target dataset, causing the system to be unable to uniquely determine its corresponding data operation logic.
[0063] In practical applications, the types of ambiguity involved in the embodiments of this specification may include, but are not limited to, semantic ambiguity, detailed requirement ambiguity, generalized requirement ambiguity, and referential ambiguity. Semantic ambiguity can refer to ambiguity caused by synonyms or vague field / table names. For example, ambiguity arises from similar expressions such as "sales revenue" and "turnover," or "customer" and "buyer." Detailed requirement ambiguity can refer to ambiguity caused by missing parameters in a data processing request. For example, "the top 5% of products" is ambiguous as it doesn't specify whether the sorting criterion is "by sales revenue" or "by sales volume." Generalized requirement ambiguity can refer to overly specific data processing requests that may require a broader interpretation. For example, a limited single date could be changed to a time range query; failure to generalize might result in empty results or ignored data. Referential ambiguity can refer to ambiguity caused by context-dependent entity / condition bindings that do not clearly refer to any preceding entity. For example, "find the orders of Beijing customers and calculate their total amount" is ambiguous as it is unclear whether "they" refers to "Beijing customers" or "orders."
[0064] In one or more embodiments of this specification, a specific method for ambiguity identification of data processing requests is further provided.
[0065] Specifically, the ambiguity identification of the data processing request may include: calculating the ambiguity quantization score corresponding to the data processing request; if the ambiguity quantization score is greater than or equal to a preset threshold, then the data processing request is determined to be ambiguous.
[0066] Further, calculating the ambiguity quantification score corresponding to the data processing request may specifically include: determining the factor score of the data processing request on at least one preset ambiguity factor; the preset ambiguity factor includes at least one of a field matching factor, a referential clarity factor, and a parameter integrity factor; and calculating the ambiguity quantification score based on each of the factor scores. Each preset ambiguity factor may correspond to one factor score.
[0067] Furthermore, the factor score of the field matching factor is determined based on the degree of matching between the query terms in the data processing request and the data pattern information in the target dataset; the data pattern information includes attribute fields and the dimension values corresponding to each attribute field. In other words, the field matching factor can be used to measure the degree of matching between the query terms in the data processing request and the data pattern information in the dataset targeted by the data processing request.
[0068] In practical applications, data schema information can represent the structured metadata of the target dataset, including: a list of attribute fields, each with a unique identifier and semantic name; and the value domain corresponding to each attribute field, which is the set of values that the field is allowed to take (i.e., the value domain). For example, the value domain of the "City" field includes "Beijing" and "Shanghai".
[0069] Furthermore, the factor score of the pronoun referencing clarity factor is determined based on the degree of clarity of the referent of the pronoun in the data processing request. In other words, the pronoun referencing clarity factor can be used to measure the degree of clarity of the referent of the pronoun in the data processing request.
[0070] Furthermore, the factor score of the parameter integrity factor is determined based on the degree of integrity of the parameters required to execute the data processing request. In other words, the parameter integrity factor can be used to measure the degree of integrity of the parameters required to execute the data processing request.
[0071] As an example, a higher match rate results in a higher match factor score, indicating less ambiguity. A higher degree of certainty regarding the referent leads to a higher referential factor score, indicating less ambiguity. Higher parameter completeness results in a higher completeness factor score, indicating less ambiguity. The actual factor score magnitude and the trend of ambiguity change are related to the definition of the match factor score and are not limited to this example.
[0072] Based on at least some embodiments of the scheme in this specification, the system (such as a large language model) can determine whether clarification is needed by calculating an ambiguity quantification score. This score takes into account factors such as field matching degree, pronoun reference, and missing key parameters, and initiates clarification when the threshold is reached.
[0073] In an optional embodiment, calculating the factor score corresponding to the field matching degree factor may include: identifying one or more query terms in the data processing request; matching each query term with the attribute fields and corresponding dimension values in the target dataset; determining whether each query term successfully matches the corresponding attribute field or dimension value; and calculating the factor score of the query matching degree factor based on the proportion of successfully matched query terms to the total number of query terms. Here, attribute fields represent columns / fields in a database / table. Dimension values represent the specific values of the attribute fields.
[0074] Specifically, for each query term, it can be determined whether there are matching attribute fields or dimension values, and a corresponding matching status can be generated; the matching status includes matched or unmatched; then, based on the ratio of the number of query terms marked as matched to the total number of query terms, the factor score of the field matching degree factor is determined.
[0075] For example, if three query terms are matched, and two of them can be successfully mapped to attribute fields or dimension values in the dataset, while one cannot be successfully mapped to an attribute field or dimension value in the dataset, then the factor score of the field matching factor can be determined as 2 / 3.
[0076] Further, for each query term, determining whether the query term successfully matches a corresponding attribute field or dimension value can specifically include: if the target dataset contains an attribute field or dimension value identical to the query term, then a successful match is determined; in practical applications, the query term can be marked as matched. Otherwise, the similarity between the query term and the relevant attribute fields and dimension values in the target dataset is calculated; if the similarity is greater than or equal to a preset similarity threshold, then a successful match is determined; in practical applications, the query term can be marked as matched. Otherwise, a failed match is determined; in practical applications, the query term can be marked as unmatched.
[0077] Optionally, the similarity may include edit distance-based similarity. Specifically, the edit distance between the query term and the candidate element is calculated; based on the edit distance, a text similarity score is determined.
[0078] The edit distance may include Levenshtein distance, etc.
[0079] As an example, the text similarity score determined based on edit distance can be calculated using the following formula:
[0080]
[0081] Among them, similarity represents the similarity, edit_distance represents the edit distance, t is the query term, and e is the candidate attribute field or dimension value.
[0082] Optionally, in addition to the edit distance, other similarity calculation methods can also be considered, such as cosine similarity, Jaccard similarity, etc.
[0083] In actual application, the preset similarity threshold can be set according to expert experience or adjusted dynamically according to historical user clarification behaviors.
[0084] Furthermore, in the case of successful similarity matching, the matching result can be used to achieve attribute disambiguation or value disambiguation. Among them, disambiguation is used to indicate selecting the one with the highest possibility from multiple candidates.
[0085] Attribute disambiguation: If the query term matches multiple attribute fields, compare the first matching degree between the query term and each matching attribute field; select the target attribute field based on the first matching degree. Further, the attribute field with the highest first matching degree can be selected as the target attribute field for the query term.
[0086] Among them, the attribute disambiguation is used to identify field names with the same semantics but different expressions. Specific cases of attribute disambiguation can be the matching of "sales amount" and "sales revenue".
[0087] Value disambiguation: If the query term matches multiple dimension values of the same attribute field, compare the second matching degree between the query term and each matching dimension value; select the target dimension value based on the second matching degree. Further, the dimension value with the highest second matching degree can be selected as the target dimension value for the query term.
[0088] Among them, the value disambiguation is used to identify dimension values with the same semantics but different expressions under the same dimension. In actual application, the value disambiguation can be applied to at least one of the following scenarios: normalization of different expressions of the same entity, including normalization of abbreviations and full names; normalization of different writing forms of the same entity; normalization of entity expressions containing specific suffixes or prefixes. The applicable scenarios of value disambiguation are not limited to this. Specific cases of value disambiguation can be the matching of "Beijing" and "Beijing Municipality".
[0089] Based on the embodiments of this specification, in practical applications, the process of ambiguity identification for data processing requests can be implemented by a large language model. Specifically, the data processing request in natural language form, the data pattern information of the target dataset, and ambiguity analysis rules can be provided to the large language model through prompt words. Thus, the large language model can determine whether the data processing request is ambiguous based on the ambiguity analysis rules and the data pattern information. Furthermore, if the large language model determines that the data processing request is ambiguous, it can also generate and output clarification guidance information to guide the user to clarify the ambiguous content.
[0090] Based on at least some embodiments of this specification, by intelligently assessing the degree of ambiguity and adaptively processing it according to the degree of ambiguity, it is possible to avoid unnecessary interference with clear requests and insufficient clarification of complex ambiguities.
[0091] Based on at least some embodiments of this specification, a highly accurate clarification scheme is proposed. Specifically, upon receiving a user's data processing request, the system does not simply trigger clarification, but first performs fine-grained ambiguity identification, quantifies the degree of ambiguity by combining multiple dimensions such as field matching degree, clarity of reference, and parameter completeness, and dynamically decides whether clarification is needed based on a preset threshold, thereby avoiding excessive intervention in low-risk requests.
[0092] In one or more embodiments of this specification, knowledge information in a knowledge base can be used to eliminate ambiguity in data processing requests (i.e., disambiguation). This will be described in detail below.
[0093] Specifically, before performing ambiguity identification on the data processing request, the process may further include: retrieving knowledge information from a knowledge base; the knowledge base includes at least one of a public knowledge base and a personal knowledge base corresponding to the user's user identifier. Correspondingly, performing ambiguity identification on the data processing request may specifically include: performing ambiguity identification on the data processing request based on the retrieved knowledge information.
[0094] In practical applications, before calculating the ambiguity quantization score corresponding to the data processing request, knowledge information can be retrieved from the knowledge base; based on the knowledge information, the ambiguity quantization score corresponding to the data processing request can be calculated.
[0095] Further, calculating the ambiguity quantification score corresponding to the data processing request based on the knowledge information specifically includes: based on the knowledge information, adjusting the factor score of the target query term on the field matching degree factor or the reference clarity factor, so as to adjust the ambiguity quantification score of the data processing request. Assume that the higher the scores of the field matching degree factor and the reference clarity factor, and the lower the ambiguity quantification score, the smaller the ambiguity. Then, based on the knowledge information, usually the factor score of the target query term on the field matching degree factor or the reference clarity factor can be increased, thereby reducing the ambiguity quantification score of the data processing request. Thus, it is beneficial to reduce excessive clarification, that is, to reduce unnecessary clarification operations of the user.
[0096] In practical applications, multiple pieces of knowledge can be associated and used for a single query simultaneously. Further, the multiple pieces of knowledge can include common knowledge entries and / or personal knowledge entries.
[0097] In an optional embodiment, retrieving knowledge information from the knowledge base specifically may include: according to the user identifier and the query term in the data processing request, retrieving personal knowledge entries related to the query term and contextually relevant from the personal knowledge base; and / or, according to the user identifier, querying personal knowledge entries that are contextually irrelevant from the personal knowledge base. Among them, personal knowledge can represent the custom term interpretation rules of the user during the interaction with the system (during the data processing request session), and is used to optimize the semantic parsing accuracy of the current request and subsequent requests provided by the user.
[0098] Among them, the contextually relevant personal knowledge entries need to match both the user identifier and the query term to be recalled; the contextually irrelevant personal knowledge entries only need to match the user identifier to be recalled. In practical applications, for example, the contextually irrelevant personal knowledge entries can reflect the information query habits of the user.
[0099] In addition, in the embodiments of this specification, according to the specific form of the clarification information, personal knowledge entries can be classified into: explanatory type and metric type. Among them, the explanatory type is usually in the form of natural language, and can further include filtering types (such as "repurchase users refer to users with an order quantity greater than 2"), synonym types (such as "Beijing = Jing"), and user habit types (global default filtering conditions). The metric type is usually an executable calculation logic (such as a calculation formula), for example, "sales amount = unit price × quantity".
[0100] In practical applications, the personal knowledge entries used during the current data processing request session can include historical personal knowledge entries from the personal knowledge base, or new personal knowledge entries generated based on clarification information input by the user during the current data processing request session.
[0101] In an optional embodiment, retrieving knowledge information from the knowledge base may specifically include: retrieving public knowledge entries from the public knowledge base that match the query terms in the data processing request. The public knowledge base may represent a pre-defined standardized terminology interpretation rule base applicable to all users.
[0102] Based on the solutions in the embodiments of this specification, personal knowledge and public knowledge can be used simultaneously. Furthermore, when using both personal and public knowledge simultaneously, a knowledge usage priority can be set. For example, personal knowledge can be prioritized to improve the accuracy of ambiguity judgment and response, thereby enhancing the user experience.
[0103] In an optional embodiment, during the process of acquiring knowledge information, for a certain query term, it can first be retrieved from a personal knowledge base. If no matching result is found in the personal knowledge base, it can then be retrieved from a public knowledge base.
[0104] Specifically, before retrieving public knowledge entries matching the query terms from the public knowledge base, the process may further include: retrieving personal knowledge entries matching the query terms from the personal knowledge base. The retrieval of public knowledge entries matching the query terms from the public knowledge base specifically includes: if no personal knowledge entries matching the query terms are found, then retrieving public knowledge entries matching the query terms from the public knowledge base.
[0105] In an optional embodiment, during the acquisition of knowledge information, retrieval can be performed simultaneously from both personal and public knowledge bases. Subsequently, during the reasoning process, the data processing system (such as a large language model) can determine, based on the context, which retrieved knowledge should be prioritized. For example, personal knowledge can be prioritized.
[0106] Specifically, the retrieved knowledge information includes target public knowledge entries and target personal knowledge entries that match the target query terms in the data processing request; the calculation of the ambiguity quantification score corresponding to the data processing request may specifically include: calculating the ambiguity quantification score corresponding to the data processing request based on the target personal knowledge entries.
[0107] In practical applications, knowledge information can be retrieved from the database before the large language model performs ambiguity identification on the data processing request. Furthermore, before using the large language model to identify ambiguity in the data processing request, the data processing request in natural language form, the data pattern information of the target dataset, ambiguity analysis rules, and retrieved indication information can be provided to the large language model via prompt words. Thus, the large language model can analyze the data processing request based on the data pattern information and the retrieved knowledge information, according to the ambiguity analysis rules, to determine whether the data processing request is ambiguous. Further, if the large language model determines that the data processing request is ambiguous, it can also generate and output clarification guidance information to guide the user to clarify the ambiguous content.
[0108] Based on at least some embodiments of this specification, a multi-granularity knowledge association architecture is proposed, which distinguishes between personal knowledge bases and public knowledge bases, and supports mixed invocation of context-dependent and context-independent knowledge. While protecting user privacy, it also takes into account the sharing of general rules, and ensures that the parsing results conform to user habits through a priority strategy (personal knowledge takes precedence over public knowledge).
[0109] Step 206: If ambiguity is identified in the data processing request, clarification guidance information is output; the clarification guidance information is used to guide the user to clarify the target content that causes ambiguity in the data processing request.
[0110] Among them, clarification guidance information can refer to human-computer interaction messages that the system actively generates and outputs after detecting ambiguity, in order to guide users to provide specific clarification content.
[0111] Target content can refer to specific segments or their corresponding semantic units within the text constituting a data processing request that are identified as ambiguous. The ambiguous target content, also known as the target content to be clarified, can include one or more pieces of content in practical applications.
[0112] In one or more embodiments of this specification, the output clarification guidance information may specifically include: identifying the target content in the data processing request that causes ambiguity; outputting clarification guidance information for the target content; and the clarification guidance information including clarification examples for the target content.
[0113] The clarification example is used to guide the user to input clarification information that meets the system's parsing requirements.
[0114] Optionally, the clarification example can be generated based on the attribute fields or dimension values of the target dataset. Optionally, the clarification example can be generated based on the target content. In practical applications, the model can generate clarification examples in real time according to the context (including data processing requests, knowledge information, etc.) to guide users to supplement clarification information. Optionally, the clarification example can be extracted from historical clarification information. Optionally, the clarification example can be obtained by retrieving relevant examples from a knowledge base.
[0115] Optionally, the clarification example is in natural language form. The clarification information is in natural language form.
[0116] Furthermore, the clarification guidance information presents the clarification example in an interactive manner. Optionally, the clarification example can be displayed as an editable template; alternatively, the clarification example can be displayed as guidance information in a fill-in-the-blank format.
[0117] For example, if a user asks, "What is the sales amount?", the system might not be able to determine which field or dimension value in the dataset corresponds to "sales amount." Therefore, a clarification guide can be generated, which includes a statement like, "Please explain {sales amount}." Furthermore, the clarification guide can include clarification examples, such as "sales amount," "total order amount," or "gross profit." In practical applications, one or more clarification examples can be provided.
[0118] In practical applications, the steps of determining the ambiguous target content in the data processing request and determining the clarification guidance information can be performed by a large language model (a device with a large language model deployed).
[0119] Figure 3 This diagram illustrates a graphical user interface displaying clarification guidance information containing clarification examples, as provided in one embodiment of this specification.
[0120] like Figure 3 As shown, in response to the data processing request "query the access behavior of high-value users," the system (large language model), after understanding the intent, determines that secondary clarification is needed. Specifically, it determines that the potentially ambiguous target content, "high-value users," needs secondary clarification. Therefore, clarification guidance information can be output to the user, such as... Figure 3 The area enclosed by the dashed line.
[0121] like Figure 3 As shown, this clarification guidance information is used to guide users to clarify the "high-value user" distinction. Further, as an example, in... Figure 3 The clarification guidance message includes the example of "users who placed orders with a quantity greater than 2". In practice, there can be multiple clarification examples.
[0122] Furthermore, users can specify a region (such as...) Figure 3 The user can enter clarification information in the text box area (which displays a clarification example), for example, "customers who placed orders in the last month". Optionally, the user can also trigger the "+Personal Knowledge" control in the clarification guidance information to actively select knowledge from their personal knowledge base. This provides high user flexibility and helps the system (large language model) generate more accurate response information.
[0123] Furthermore, the system (large language model) can execute data processing logic based on the clarification information already provided by the user, combined with the data processing request, to output response information.
[0124] Figure 4 A schematic diagram is shown illustrating a clarification guidance message containing clarification examples in a graphical user interface, according to another embodiment of this specification.
[0125] like Figure 4 As shown, in response to the data processing request "query the access behavior of high-value users," the system (large language model), after understanding the intent, determines that secondary clarification is needed. Specifically, it determines that the target content of "high-value users" and "access behavior" needs secondary clarification, which may cause ambiguity. Therefore, clarification guidance information can be output to the user, such as... Figure 4 The area enclosed by the dashed line.
[0126] like Figure 4 As shown, this clarification guidance information is used to guide users to clarify the concepts of "high-value users" and "access behavior." Further, as an example, in... Figure 4 The clarification guidance information includes an example for "high-value users" ("users who have placed more than 2 orders") and an example for "access behavior" ("users who have opened the homepage within the last week"). In practice, there can be multiple clarification examples for each target content.
[0127] Furthermore, users can specify a region (such as...) Figure 4 The system displays clarification examples in text boxes. Users can enter clarification information in these areas. For example, for "high-value users," the clarification information could be "customers who placed orders in the last month," and for "access behavior," it could be "ordering behavior in the last year." Optionally, users can also trigger the "+Personal Knowledge" control in the clarification guidance information to actively select knowledge from their personal knowledge base. This provides high user flexibility and helps the system (large language model) generate more accurate response information.
[0128] Furthermore, the system (large language model) can execute data processing logic based on the clarification information already provided by the user, combined with the data processing request, to output response information.
[0129] Figure 3 and Figure 4 This is just an example of how to display clarification guidance information and obtain clarification information from users. In actual applications, it can also be displayed or interacted with by users in other ways.
[0130] Based on the embodiments of this specification, the solution can receive diverse clarification information in natural language form input by the user, thereby covering diverse business scenarios, matching diverse user expression habits, and having strong generalization ability.
[0131] In one or more embodiments of this specification, the data processing request may include one or more processing tasks.
[0132] Specifically, if the data processing request includes multiple processing tasks, then the ambiguity identification of the data processing request may specifically include: performing ambiguity identification on at least one of the multiple processing tasks.
[0133] Furthermore, if at least two of the multiple processing tasks are identified as ambiguous, the clarification guidance information is used to guide the user to clarify the ambiguous target content in the at least two tasks in parallel. That is, multiple ambiguities in multiple tasks can be clarified simultaneously.
[0134] In an optional embodiment, the output clarification guidance information may specifically include: decomposing the data processing request into multiple processing tasks through intent recognition and task planning; identifying multiple target contents that cause ambiguity in the multiple processing tasks; the multiple target contents coming from at least a portion of the multiple processing tasks; and outputting clarification guidance information to guide the user to clarify the multiple target contents.
[0135] Based on the solutions implemented in this specification, for complex queries involving multiple tasks (such as data retrieval and analysis), all items requiring clarification from multiple tasks can be presented collectively before task execution, rather than being clarified step-by-step by task. This reduces the number of user interactions and improves processing efficiency. Of course, in practical applications, it is also possible to choose to process them separately, for example, clarifying the first task first and then clarifying the second task.
[0136] For example, if a user enters "Analyze the sales and profits of Beijing and Shanghai last month", the system can identify four tasks / parameters: two cities (Beijing and Shanghai) and two indicators (sales and profits). It can then generate multiple fill-in-the-blank items on one interface for the user to clarify in parallel.
[0137] In an optional embodiment, the plurality of processing tasks includes a first processing task and a second processing task dependent on the first processing task; the output clarification guidance information may specifically include: outputting first clarification guidance information for a first target content in the first processing task; and after parsing the first processing task based on the first clarification guidance information input for the first clarification guidance information, outputting second clarification guidance information for a second target content in the second processing task.
[0138] More specifically, the output clarification guidance information may include: decomposing the data processing request into multiple processing tasks with dependencies through intent recognition and task planning; the multiple processing tasks include a first processing task and a second processing task dependent on the first processing task; determining a first target content that causes ambiguity in the first processing task; outputting first clarification guidance information to guide the user to clarify the first target content; acquiring the first clarification information input by the user for the first target content; parsing the first processing task based on the first clarification information and determining a second target content that causes ambiguity in the second processing task; outputting second clarification guidance information to guide the user to clarify the second target content; and acquiring the second clarification information input by the user for the second target content.
[0139] The ambiguity identification of the second processing task can rely on the parsing result of the first processing task.
[0140] Based on the solutions in the embodiments of this specification, for multiple tasks with clear dependencies, information can be clarified according to the dependencies, and the order of node clarification can be dynamically adjusted according to the task relationship, reducing repeated interactions and improving processing efficiency.
[0141] For example, if a user enters "Calculate the total revenue first, then tell me how much it has increased year-on-year", the system can identify the task chain, first clarify "total revenue" and calculate the total revenue (first task), and after the user answers "total revenue refers to the sum of the sales of products A and B", the system can then clarify the comparison benchmark of "year-on-year" to calculate the year-on-year growth (second task).
[0142] Based on at least some embodiments of this specification, a task dependency-aware clarification optimization mechanism is proposed for complex multi-task requests. This mechanism can identify the logical dependencies between processing tasks and output clarification guidance in layers as needed. After resolving ambiguities in the upstream task, the downstream dependent task is clarified a second time. This can improve the effectiveness of obtaining clarification information, effectively avoid redundant interactions, and make the entire clarification process more intelligent, smooth, and in line with user cognitive logic.
[0143] Step 208: Obtain the clarification information input by the user regarding the target content.
[0144] The clarification information can be used to represent supplementary, corrective, or confirmatory natural language content provided to eliminate specific ambiguities. The clarification information is targeted at the target content and is used to eliminate its ambiguity. The solution based on the embodiments of this specification supports user input of clarification information in free text format, offering high flexibility.
[0145] In one or more embodiments of this specification, an end-to-end closed-loop mechanism for intelligent clarification and knowledge accumulation is further provided, significantly improving the accuracy, efficiency, and user experience of natural language data processing. This is described in detail below.
[0146] The knowledge accumulation process can be used to identify and transform clarification information provided by users regarding ambiguous target content in data processing requests into structured knowledge stored in a personal knowledge base. In practical applications, the clarification information provided by users during the clarification process can be accumulated by the system as personal knowledge for use in processing current and subsequent data requests.
[0147] Specifically, after obtaining the clarification information input by the user regarding the target content, the process may further include: generating a personal knowledge entry based on the target content and the clarification information; storing the personal knowledge entry in a personal knowledge base corresponding to the user's user identifier; the personal knowledge base is used to assist in ambiguity identification when processing the user's subsequent data processing requests. Specifically, it assists in at least one of ambiguity resolution, reference resolution, or parameter completion.
[0148] Furthermore, generating personal knowledge entries based on the target content and the clarification information may specifically include: constructing personal knowledge entries in the form of key-value pairs, using the target content as the key and the clarification information as the value.
[0149] In this scenario, context-dependent knowledge can be formed. This knowledge is stored in key-value pairs, where the key is the query term and the value is the clarification information corresponding to the query term. In practical applications, the clarification information may include explanations or calculation rules for the query term. This context-dependent knowledge can be retrieved based on both the user identifier and the query term.
[0150] In practical applications, optionally, the query terms and clarification information can be directly stored in a structured format without large language model processing. Alternatively, based on the target content and clarification information, a large language model can be used to extract the semantic relationship between them in a structured manner to generate personal knowledge entries. For example, the Key can be the optimized content of the query terms, and the Value can be the optimized content of the clarification information. Optimizing the query terms and / or clarification information using a model helps improve the standardization of the formatted stored personal knowledge, thereby improving the accuracy of subsequent system retrieval and recall of personal knowledge when executing data processing requests.
[0151] Furthermore, generating personal knowledge entries based on the target content and the clarification information may specifically include: forming generalized personal knowledge entries based on the target content, the clarification information, and the user's historical clarification pairs; wherein, the historical clarification pairs include historical target content that caused ambiguity in the user's historical data processing requests and historical clarification information regarding the historical target content. In practical applications, the generalized personal knowledge entries can be used to reflect generalized user expression habits or cross-session semantic information, etc.
[0152] In this scenario, context-free knowledge entries can be created, reflecting the user's query habits. This context-free personal knowledge can then be retrieved based on the user's identifier.
[0153] In practical applications, prompt words can be used to request large language models to summarize and condense data processing requests and corresponding clarification information into personal knowledge.
[0154] Alternatively, a data processing request may contain multiple ambiguous target contents (such as "Zhang San's performance in 2023" may involve multiple ambiguities such as who "Zhang San" refers to and what indicators "performance" refers to). The system can clarify multiple questions at once and then integrate this clarification information to generate more comprehensive personal knowledge.
[0155] In one or more embodiments of this specification, the method may further include: determining a knowledge type identifier for the personal knowledge entry; the knowledge type identifier including context-dependent or context-independent types; the context-dependent personal knowledge entry being referenced when processing a subsequent data processing request from the user containing matching query terms; the context-independent personal knowledge entry being referenced when processing any subsequent data processing request from the user; and then, storing the knowledge type identifier in association with the personal knowledge entry.
[0156] On the one hand, the knowledge type identifier can be used to indicate the applicable conditions of the knowledge. Context-dependent types, such as "sales amount = sales revenue," can be used only when the data processing request contains "sales amount." Context-independent types, such as "default time range = the last 30 days," can be applied to all queries, specifically, to all queries by the current user. In practical applications, context-independent personal knowledge can include things like default time granularity and commonly used metrics.
[0157] On the other hand, the knowledge type identifier can be used to indicate the recall conditions for the corresponding personal knowledge entries. Context-dependent knowledge entries require both the user identifier and the query term to be matched for recall. Context-independent knowledge entries only require a match of the user identifier for recall.
[0158] From another perspective, the knowledge type identifier is used to indicate its application scenario, including explanatory and metric types. Specifically, the clarification information in the explanatory knowledge can be in natural language form, used to explain the semantics related to query terms; the clarification information in the metric knowledge can be in computational logic form, used to explain the computational rules related to the computational query terms.
[0159] Furthermore, after storing the personal knowledge entry in the personal knowledge base corresponding to the user's user identifier, the process may further include: when processing a new data processing request from the user, querying the personal knowledge base; if the new data processing request contains the target content in the personal knowledge entry, then parsing the new data processing request based on the clarification information corresponding to the target content to generate new response information. In this process, it is unnecessary to output clarification guidance information for the target content again, avoiding repeated clarifications by the user regarding the same or similar terms in different sessions.
[0160] Based on the solutions implemented in this specification, users can reuse ambiguous content in data processing requests with only one clarification, reducing the cost of repeated dialogues and minimizing resource waste.
[0161] As can be seen, the solution based on the embodiments of this specification realizes a closed loop of the disambiguation scheme of "ambiguity clarification -> knowledge accumulation -> automatic disambiguation", which not only improves the accuracy of current data processing, but also completes the intelligent accumulation of personal knowledge to further improve the efficiency and accuracy of subsequent data processing.
[0162] Furthermore, based on the solutions in the embodiments of this specification, considering that different datasets may correspond to different business domains, in order to further improve the accuracy of the application of personal knowledge, when storing personal knowledge, the dataset identifier of the dataset can be associated with the personal knowledge entry for storage; and when querying personal knowledge, the dataset identifier can also be referenced at the same time for knowledge retrieval.
[0163] Similarly, it is understandable that, for public knowledge bases, considering that different datasets may correspond to different business domains, in order to further improve the accuracy of the application of public knowledge, when querying public knowledge, the dataset identifier can also be referenced for knowledge retrieval and recall.
[0164] Based on at least some embodiments of this specification, the system can automatically store the clarification information entered by the user during the query process as the user's corresponding personal knowledge. Therefore, the user does not need to manually maintain the knowledge base independently of the query process, resulting in high convenience and low user operating costs. Furthermore, as the user continuously uses the methods provided in the embodiments of this specification, they can continuously update their personal database. By automatically transforming real-time clarification information into reusable knowledge, a continuous optimization loop of "getting smarter with use" is achieved, realizing "learning from interaction," "learning while using," and "updating as you use." This greatly improves the timeliness and accuracy of the natural language data processing system, and enhances its intelligence level and user experience.
[0165] Based on at least some embodiments of this specification, ambiguity clarification and knowledge accumulation are deeply integrated to form a self-reinforcing closed loop of "ambiguity clarification → knowledge accumulation → automatic disambiguation": each clarification action of a user can be structured into a personal knowledge entry and stored in the user's exclusive knowledge base. Subsequent requests from the same or similar users can directly call the structured personal knowledge entry for automatic parsing (i.e., to achieve disambiguation based on knowledge information) without the need for repeated clarification. This not only improves the efficiency of a single interaction, but also realizes the continuous accumulation and reuse of personalized knowledge.
[0166] Step 210: Based on the data processing request and the clarification information, output the response information.
[0167] The response information refers to the comprehensive output result generated by the system after understanding and clarifying the user's intent, through the execution of corresponding data processing operations, which aims to fulfill the user's request.
[0168] In practical applications, the specific presentation format of the response information can be diverse. For example, the response information may include structured data, such as tables, lists, key-value pairs (JSON), etc. For example, the response information may include analysis report / summary text, such as one or more paragraphs of natural language text summarizing key findings, conclusions, or predictions from the data analysis. For example, the response information may include visual charts, such as charts, graphs, dashboards, etc., to intuitively display data relationships and trends. For example, the response information may include executable code / instructions, such as generated SQL query statements or Python scripts, for user review or further use. In practical applications, the response information may include one or more of the above forms, or other forms not listed. Optionally, the specific format of the response information can be determined based on the specific data processing request.
[0169] Further, step 210 may specifically include processing the data in the target dataset based on the data processing request and the clarification information to generate and output response information. The data processing request represents the original data processing intent, and the clarification information is used to explain / clarify the user's original data processing intent, so that the system can execute the data processing logic more accurately.
[0170] Based on the solutions in the embodiments of this specification, when there is semantic ambiguity in the user's data processing request, the system actively triggers an interactive supplementary information acquisition mechanism. The overall process is different from the post-correction mode of traditional solutions. By placing clarification before the actual data processing is performed, it can significantly reduce the number of interaction rounds and invalid processing. For example, it can avoid invalid processing caused by the system's misunderstanding of the user's data processing request.
[0171] In practical applications, during the process of generating response information for data processing requests, specifically, before outputting the response information, ambiguity identification can be performed, and if ambiguity is identified, the user's input clarification information can be obtained. Furthermore, the user can output clarification guidance information during the model's thinking process and obtain the user's input clarification information, thus ensuring a smooth interaction process.
[0172] In one or more embodiments of this specification, knowledge source prompts may also be output during the output of response information (including before, during, or after the output of response information).
[0173] Specifically, the method for generating response information may further include: if the generation of the response information depends on at least one personal knowledge entry, then outputting knowledge source prompt information.
[0174] In practical applications, knowledge source prompts can be displayed in the user interface; these prompts indicate that the generation of the response depends on at least one personal knowledge entry.
[0175] In practical applications, the at least one personal knowledge entry may include personal knowledge entries included in the recalled knowledge information; it may also include personal knowledge entries generated based on the clarification information input by the user. In practical applications, the at least one personal knowledge entry may include the matched personal knowledge entry and / or context-independent personal knowledge entries.
[0176] Furthermore, the knowledge source prompt information can be displayed in association with the response information. Optionally, the knowledge source prompt information can be displayed before the response information is displayed.
[0177] Figure 5 This specification illustrates a schematic diagram of a graphical user interface displaying knowledge source prompts, provided in one embodiment of the specification.
[0178] like Figure 5 Continue to use Figure 3 For example, in response to the data processing request "query the access behavior of high-value users," the system (large language model), after understanding the intent, determines that the target content "high-value users," which may cause ambiguity, needs secondary clarification, and obtains the clarification information "customers who placed orders in the last month" input by the user. Optionally, in response to the user's input clarification information, the user's input information "high-value users = customers who placed orders in the last month" can be displayed in the graphical user interface.
[0179] Furthermore, the system (large language model) can execute data processing logic based on the clarification information already provided by the user, combined with the data processing request, to output response information. Based on the embodiments of this specification, during the output of response information, knowledge source prompts can be generated and displayed in the graphical user interface. For example... Figure 5 As shown, the information displayed may indicate that "1 piece of personal knowledge has been associated." However, in practical applications, the display of knowledge source information may vary beyond this method. Figure 5 In the form of.
[0180] In one or more embodiments of this specification, the method for generating response information may further include: obtaining the user's editing operation on at least one personal knowledge entry corresponding to the knowledge source prompt information; in response to the editing operation, determining the edited personal knowledge entry; and regenerating response information for the data processing request based on the edited personal knowledge entry.
[0181] Optionally, before obtaining the user's editing operation on the at least one personal knowledge entry corresponding to the knowledge source prompt information, the process may further include: displaying a knowledge viewing control in association with the knowledge source prompt information; obtaining the user's trigger operation on the knowledge viewing control; and, in response to the trigger operation, displaying the at least one personal knowledge entry on which the generation of the reply information depends.
[0182] Optionally, in response to a user action on the knowledge source prompt information, an editing interface for at least one personal knowledge entry may be provided; based on the editing action received through the editing interface, the edited personal knowledge entry is determined; and based on the edited personal knowledge entry, a response to the data processing request is regenerated. Optionally, the user action includes a viewing action; in response to the viewing action, the content of the referenced personal knowledge entry is displayed.
[0183] Based on the embodiments described in this specification, by displaying knowledge source prompts and at least one associated personal knowledge entry, transparent and reliable decision-making basis can be provided during interaction, thereby improving the user experience.
[0184] Furthermore, the editing operation includes at least one of the knowledge selection operation and the knowledge update operation.
[0185] The knowledge selection operation can be used to redetermine the personal knowledge applicable to the current data processing request from the at least one personal knowledge entry, so as to achieve selective use of personal knowledge and improve flexibility.
[0186] The knowledge update operation can be used to modify the clarification information corresponding to one or more personal knowledge entries among the at least one personal knowledge entry, so as to realize the dynamic update of personal knowledge.
[0187] Furthermore, if the editing operation includes a knowledge update operation, the method for generating response information may further include: in response to the knowledge update operation, storing the edited personal knowledge entry into the personal knowledge base for use in processing the user's subsequent data processing requests.
[0188] Based on the embodiments in this specification, in practical applications, if a user clarifies the same target content multiple times with different results, the corresponding entry in the personal knowledge base can be updated according to the user's last updated content, thereby realizing the dynamic updating of personal knowledge.
[0189] Figure 6 This specification illustrates a schematic diagram of an editing interface provided in an embodiment of which displays an editing interface in a graphical user interface for editing the personal knowledge entries on which the generated response information is based.
[0190] Continue Figure 5 For example, if a user performs a viewing action (such as clicking on the information) on the knowledge source prompt "1 piece of personal knowledge has been associated," the following can be displayed: Figure 6 The interface shown is for editing at least one personal knowledge entry. Specifically, as... Figure 6 The area within the dashed box is the editing interface area.
[0191] In the editing interface area, users can edit at least one associated personal knowledge entry. For example, they can perform knowledge selection operations, such as... Figure 6 As shown, knowledge selection can be achieved by manipulating the current selection control on the knowledge bar. Similarly, knowledge update operations can be performed, such as... Figure 6 As shown, knowledge can be updated by modifying the clarification information in a knowledge entry. When a knowledge update operation is performed, the updated knowledge can overwrite the original knowledge and be stored in the personal knowledge base.
[0192] In addition, in the continued use Figure 3 In cases where the user clarifies a specific target content, although... Figure 5 and Figure 6 The example shown only associates one piece of personal knowledge, but in practical applications, more than one piece of personal knowledge can be associated. When associating multiple pieces of personal knowledge, one of them can be... Figure 3 In this scenario, the clarification information entered by the user corresponds to the personal knowledge, and other pieces of personal knowledge can be other personal knowledge in the user's personal knowledge base.
[0193] While this specification provides method steps as described in the embodiments or flowcharts using one or more examples, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many, and does not represent the only possible execution order. The order of some steps may be adjusted according to actual needs, or some steps may be omitted. When the claims involve method steps, adjustments to the order of such steps, or parallel execution between steps, are also within the scope of protection of the claims.
[0194] Figure 2The proposed method effectively addresses parsing errors caused by ambiguous user expressions or non-standard terminology by embedding proactive ambiguity identification and interactive clarification mechanisms into the natural language data processing workflow. Specifically, upon receiving a user's natural language data processing request for a target dataset, the system does not directly perform parsing but first identifies ambiguity in the request. When ambiguity is detected, it proactively outputs clarification guidance information, guiding the user to clarify the target content causing the ambiguity, and regenerates a response based on the user's feedback. This mechanism avoids the shortcomings of traditional systems that blindly guess or directly report errors when faced with ambiguous requests, significantly improving the accuracy and reliability of the response results. Furthermore, since the clarification guidance is triggered only when ambiguity is identified and focuses on specific target content, it ensures necessary human-computer collaboration while avoiding unnecessary interactive interference, thus maintaining a good user experience while ensuring parsing accuracy.
[0195] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they have not been described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.
[0196] According to the above explanation, Figure 7 This specification illustrates a flowchart of a data processing system responding to a user's data processing request in a practical application scenario, according to an embodiment of this specification.
[0197] like Figure 7 As shown, in step 702, the user's data processing request is obtained.
[0198] In practical applications, users can access data processing tools (such as data analysis tools) and input data processing requests in natural language. Optionally, users can input data processing requests via text or voice, or they can select from a set of options provided by the system.
[0199] Typically, the data processing request is targeted at a specific dataset, so the user can select or specify the target dataset.
[0200] Step 704: Determine if clarification is needed.
[0201] Specifically, the data processing request can be ambiguously identified; if the data processing request is found to be ambiguous, it is determined that clarification is required, and step 706 can be executed; if the data processing request is not found to be ambiguous, it is determined that clarification is not required, and step 708 can be executed.
[0202] Step 706: If clarification is required, determine whether clarification information has been obtained.
[0203] Specifically, when clarification is required, clarification guidance information can be output to guide the user to clarify the ambiguous target content in the data processing request. Afterwards, if clarification information is obtained, step 708 can be executed; if no clarification information is obtained, the data processing flow can optionally be interrupted, or the system can optionally infer predicted clarification information for the target content and generate predicted response information.
[0204] Step 708: Generate response information.
[0205] Step 710: After obtaining the clarification information, the personal knowledge base can be updated based on the clarification information. Specifically, personal knowledge entries can be generated and stored in the personal knowledge base based on the target content and the corresponding clarification information.
[0206] Based on the embodiments of this specification, an evolvable, interpretable, and highly available natural language data interaction paradigm is constructed. Based on the embodiments of this specification, a natural language interaction method based on active learning and dynamic clarification is proposed. This method can accurately identify ambiguous content in user requests and intelligently guide users to clarify, thereby generating accurate responses. This method breaks through the limitations of traditional static prompts or centralized knowledge bases. By automatically transforming real-time clarification interactions into reusable knowledge, it achieves a continuous optimization loop of "getting smarter with use." It not only significantly improves the accuracy and efficiency of single question-and-answer sessions but also intelligently handles complex requests containing multiple tasks or complex dependencies, providing transparent and reliable decision-making basis during interactions, ultimately achieving the technical effect of collaborative evolution of user experience and system intelligence.
[0207] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods.
[0208] Figure 8 This specification illustrates an embodiment corresponding to... Figure 2 A schematic diagram of the structure of a device for generating response information.
[0209] like Figure 8 As shown, the device may include:
[0210] The data processing request acquisition module 802 is used to acquire the user's data processing request in natural language form for the target dataset;
[0211] Ambiguity identification module 804 is used to identify ambiguities in the data processing request;
[0212] The clarification guidance information output module 806 is used to output clarification guidance information if the data processing request is found to be ambiguous; the clarification guidance information is used to guide the user to clarify the target content that causes ambiguity in the data processing request.
[0213] The clarification information acquisition module 808 is used to acquire the clarification information input by the user regarding the target content;
[0214] The response information output module 810 is used to output response information based on the data processing request and the clarification information.
[0215] based on Figure 8 The embodiments of this specification also provide some specific implementation schemes of the method, which are described below.
[0216] Optionally, the ambiguity identification module 804 is specifically used to: calculate the ambiguity quantization score corresponding to the data processing request; if the ambiguity quantization score is greater than or equal to a preset threshold, then determine that the data processing request is ambiguous.
[0217] Optionally, calculating the ambiguity quantification score corresponding to the data processing request may specifically include: determining the factor score of the data processing request on at least one preset ambiguity factor; the preset ambiguity factor includes at least one of a field matching factor, a referential clarity factor, and a parameter integrity factor; and calculating the ambiguity quantification score based on each of the factor scores.
[0218] Optionally, the factor score of the field matching factor is determined based on the degree of matching between the query terms in the data processing request and the data pattern information in the target dataset; the data pattern information includes attribute fields and the dimension values corresponding to each attribute field; the factor score of the referential clarity factor is determined based on the degree of clarity of the referent of the pronoun in the data processing request; the factor score of the parameter completeness factor is determined based on the degree of completeness of the parameters required to execute the data processing request.
[0219] Optionally, determining the factor score of the data processing request on the field matching factor specifically includes: identifying one or more query terms in the data processing request; matching each query term with the attribute fields and corresponding dimension values in the target dataset; determining whether each query term has successfully matched the corresponding attribute field or dimension value; and calculating the factor score of the query matching factor based on the proportion of successfully matched query terms to the total number of query terms.
[0220] Optionally, for each query term, determining whether the query term successfully matches the corresponding attribute field or dimension value specifically includes: if there is an attribute field or dimension value in the target dataset that is the same as the query term, then it is determined to be a successful match; otherwise, the similarity between the query term and the relevant attribute field and dimension value in the target dataset is calculated, and if the similarity is greater than or equal to a preset similarity threshold, then it is determined to be a successful match; otherwise, it is determined to be a failed match.
[0221] Optionally, the device for generating response information is further configured to: retrieve knowledge information from a knowledge base; the knowledge base includes at least one of a public knowledge base and a personal knowledge base corresponding to the user's user identifier; the ambiguity identification module 804 is specifically configured to: perform ambiguity identification on the data processing request based on the retrieved knowledge information.
[0222] Optionally, retrieving knowledge information from the knowledge base specifically includes: retrieving context-dependent personal knowledge entries that match the query terms from the personal knowledge base based on the user identifier and the query terms in the data processing request; and / or, retrieving context-independent personal knowledge entries from the personal knowledge base based on the user identifier.
[0223] Optionally, retrieving knowledge information from the knowledge base specifically includes: retrieving public knowledge entries from the public knowledge base that match the query terms in the data processing request.
[0224] Optionally, before retrieving public knowledge entries matching the query terms from the public knowledge base, the method further includes: retrieving personal knowledge entries matching the query terms from the personal knowledge base; specifically, retrieving public knowledge entries matching the query terms from the public knowledge base includes: if no personal knowledge entries matching the query terms are found, then retrieving public knowledge entries matching the query terms from the public knowledge base.
[0225] Optionally, after obtaining the clarification information input by the user for the target content, the method further includes: generating a personal knowledge entry based on the target content and the clarification information; storing the personal knowledge entry in a personal knowledge base corresponding to the user's user identifier; the personal knowledge base is used to assist in ambiguity identification when processing the user's subsequent data processing requests.
[0226] Optionally, generating a personal knowledge entry based on the target content and the clarification information specifically includes: constructing a key-value pair personal knowledge entry using the target content as the key and the clarification information as the value.
[0227] Optionally, generating a personal knowledge entry based on the target content and the clarification information specifically includes: forming a generalized personal knowledge entry based on the target content, the clarification information, and the user's historical clarification pairs; wherein the historical clarification pairs include the ambiguous historical target content in the user's historical data processing requests and historical clarification information regarding the historical target content.
[0228] Optionally, the apparatus for generating response information is further configured to: determine a knowledge type identifier for the personal knowledge entry; the knowledge type identifier includes context-dependent or context-independent types; the context-dependent personal knowledge entry is referenced when processing the user's subsequent data processing request containing matching query terms; the context-independent personal knowledge entry is referenced when processing any subsequent data processing request of the user; and associate the knowledge type identifier with the personal knowledge entry for storage.
[0229] Optionally, after storing the personal knowledge entry in the personal knowledge base corresponding to the user's user identifier, the method further includes: querying the personal knowledge base when processing a new data processing request from the user; if the new data processing request contains the target content in the personal knowledge entry, then parsing the new data processing request based on the clarification information corresponding to the target content to generate new response information.
[0230] Optionally, the apparatus for generating response information is further configured to: if the generation of the response information depends on at least one personal knowledge entry, then output knowledge source prompt information.
[0231] Optionally, the apparatus for generating response information is further configured to: acquire the user's editing operation on the at least one personal knowledge entry corresponding to the knowledge source prompt information; in response to the editing operation, determine the edited personal knowledge entry; and based on the edited personal knowledge entry, regenerate response information for the data processing request.
[0232] Optionally, before obtaining the user's editing operation on the at least one personal knowledge entry corresponding to the knowledge source prompt information, the method further includes: displaying a knowledge viewing control in association with the knowledge source prompt information; obtaining the user's trigger operation on the knowledge viewing control; and, in response to the trigger operation, displaying the at least one personal knowledge entry on which the generation of the reply information depends.
[0233] Optionally, the editing operation includes a knowledge update operation; the device for generating response information is further configured to: in response to the knowledge update operation, store the edited personal knowledge entry into the personal knowledge base for use in processing the user's subsequent data processing requests.
[0234] Optionally, the clarification guidance information output module 806 is specifically used to: determine the target content that causes ambiguity in the data processing request; output clarification guidance information for the target content; and the clarification guidance information includes clarification examples for the target content.
[0235] Optionally, the data processing request includes multiple processing tasks; the ambiguity identification module 804 is specifically used to: perform ambiguity identification on at least one of the multiple processing tasks.
[0236] Optionally, if at least two of the plurality of processing tasks are identified as ambiguous, the clarification guidance information is used to guide the user to clarify the target content that causes ambiguity in the at least two tasks in parallel.
[0237] Optionally, the plurality of processing tasks includes a first processing task and a second processing task dependent on the first processing task; the output clarification guidance information specifically includes: outputting first clarification guidance information for a first target content in the first processing task; and after parsing the first processing task based on the first clarification information input for the first clarification guidance information, outputting second clarification guidance information for a second target content in the second processing task.
[0238] It is understood that the modules mentioned above refer to computer programs or program segments used to perform one or more specific functions. Furthermore, the distinction between these modules does not imply that the actual program code must also be separate.
[0239] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0240] The above is an illustrative scheme of an apparatus for generating response information according to this embodiment. It should be noted that the technical solution of this apparatus for generating response information and the technical solution of the method for generating response information described above belong to the same concept. For details not described in detail in the technical solution of the apparatus for generating response information, please refer to the description of the technical solution of the method for generating response information described above.
[0241] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.
[0242] Figure 9 A structural block diagram of a computing device provided according to an embodiment of this specification is shown.
[0243] The computing device 900 includes:
[0244] Memory 910 and processor 920;
[0245] The memory 910 is used to store computer programs / instructions, and the processor 920 is used to execute the computer programs / instructions, which, when executed by the processor 920, implement the steps of the method for generating response information.
[0246] Specifically, the components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.
[0247] The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0248] In one embodiment of this specification, the above-described components of the computing device 900 and Figure 9 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 9 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0249] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 can also be a mobile or stationary server.
[0250] The processor 920 executes the computer instructions to implement the steps of the method for generating response information.
[0251] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the method for generating response information described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the method for generating response information described above.
[0252] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the method for generating response information as described above.
[0253] The above is an illustrative embodiment of a computer-readable storage medium according to this invention. It should be noted that the technical solution of this storage medium and the technical solution of the method for generating response information described above belong to the same concept. Details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the method for generating response information described above.
[0254] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method for generating response information described above.
[0255] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the method for generating response information described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the method for generating response information described above.
[0256] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the apparatus and device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The apparatus, device and method provided in the embodiments of this specification are corresponding to each other, and therefore the apparatus and device also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatus and device will not be repeated here.
[0257] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0258] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0259] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0260] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0261] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0262] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, embodiments of this specification can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0263] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0264] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0265] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0266] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0267] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0268] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0269] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0270] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. A method for generating response information, comprising: Obtain the user's data processing request in natural language for the target dataset; Ambiguity identification is performed on the data processing request; If the data processing request is found to be ambiguous, a clarification guide message is output. The clarification guidance information is used to guide the user to clarify the ambiguous target content in the data processing request; Obtain the clarification information input by the user regarding the target content; Based on the data processing request and the clarification information, a response is output.
2. The method as described in claim 1, wherein the ambiguity identification of the data processing request specifically includes: Calculate the ambiguity quantization score corresponding to the data processing request; If the ambiguity quantification score is greater than or equal to a preset threshold, then the data processing request is determined to be ambiguous.
3. The method as described in claim 2, wherein calculating the ambiguity quantization score corresponding to the data processing request specifically includes: Determine the factor score of the data processing request on at least one preset ambiguity factor; The preset ambiguity factor includes at least one of the following: field matching degree factor, reference clarity factor, and parameter integrity factor; The ambiguity quantification score is calculated based on the scores of each factor.
4. The method of claim 3, wherein, The factor score of the field matching degree factor is determined based on the degree of matching between the query terms in the data processing request and the data pattern information in the target dataset; the data pattern information includes attribute fields and the dimension values corresponding to each attribute field. The factor score of the referential clarity factor is determined based on the degree of clarity of the referent of the pronoun in the data processing request; The factor score of the parameter integrity factor is determined based on the degree of integrity of the parameters required to execute the data processing request.
5. The method as described in claim 3, wherein determining the factor score of the data processing request on the field matching factor specifically includes: Identify one or more query terms in the data processing request; Each of the query terms is matched with the attribute fields and the corresponding dimension values in the target dataset. Determine whether each of the aforementioned query terms successfully matches the corresponding attribute field or dimension value; The factor score of the query matching factor is calculated based on the proportion of successfully matched query terms to the total number of query terms.
6. The method as described in claim 5, for each query term, determining whether the query term successfully matches the corresponding attribute field or dimension value, specifically includes: If the target dataset contains attribute fields or dimension values that are the same as the query terms, then it is considered a successful match; Otherwise, calculate the similarity between the query term and the relevant attribute fields and dimension values in the target dataset. If the similarity is greater than or equal to a preset similarity threshold, it is determined to be a successful match. Otherwise, it is determined that the match was unsuccessful.
7. The method of claim 1, further comprising, before performing ambiguity identification on the data processing request: Retrieve knowledge information from the knowledge base; The knowledge base includes at least one of a public knowledge base and a personal knowledge base corresponding to the user's user identifier; The ambiguity identification of the data processing request specifically includes: Based on the retrieved knowledge information, ambiguity identification is performed on the data processing request.
8. The method as described in claim 7, wherein retrieving knowledge information from the knowledge base specifically includes: Based on the user identifier and the query terms in the data processing request, retrieve context-sensitive personal knowledge entries from the personal knowledge base that match the query terms; and / or, Based on the user identifier, query context-independent personal knowledge entries from the personal knowledge base.
9. The method as described in claim 7, wherein retrieving knowledge information from the knowledge base specifically includes: Based on the query terms in the data processing request, retrieve public knowledge entries from the public knowledge base that match the query terms.
10. The method of claim 9, further comprising, before retrieving public knowledge entries from the public knowledge base that match the query terms: Retrieve personal knowledge entries that match the query terms from the personal knowledge base; The step of retrieving public knowledge entries from the public knowledge base that match the query terms specifically includes: If no personal knowledge entry matching the query term is found, then a public knowledge entry matching the query term is retrieved from the public knowledge base.
11. The method of claim 1, further comprising, after obtaining the clarification information input by the user regarding the target content: Based on the target content and the clarification information, generate personal knowledge entries; The personal knowledge entries are stored in a personal knowledge base corresponding to the user's user identifier; the personal knowledge base is used to assist in ambiguity identification when processing the user's subsequent data processing requests.
12. The method of claim 11, wherein generating a personal knowledge entry based on the target content and the clarification information specifically includes: Using the target content as the key and the clarification information as the value, construct a personal knowledge entry in the form of a key-value pair.
13. The method of claim 11, wherein generating a personal knowledge entry based on the target content and the clarification information specifically includes: Based on the target content, the clarification information, and the user's historical clarification pairs, a generalized personal knowledge entry is formed; wherein, the historical clarification pairs include the ambiguous historical target content in the user's historical data processing requests and the historical clarification information regarding the historical target content.
14. The method of claim 11, further comprising: Determine the knowledge type identifier for the individual knowledge entry; the knowledge type identifier includes context-dependent or context-independent types. The context-dependent personal knowledge entries are referenced when processing the user's subsequent data processing requests containing matching query terms; the context-independent personal knowledge entries are referenced when processing any of the user's subsequent data processing requests. The knowledge type identifier is associated with and stored with the personal knowledge entry.
15. The method of claim 11, further comprising, after storing the personal knowledge entry in the personal knowledge base corresponding to the user's user identifier: When processing a new data processing request from the user, the personal knowledge base is queried; If the new data processing request contains the target content in the personal knowledge entry, then based on the clarification information corresponding to the target content, the new data processing request is parsed to generate new response information.
16. The method of claim 1, further comprising: If the generation of the response information depends on at least one personal knowledge entry, then a knowledge source prompt message is output.
17. The method of claim 16, further comprising: Obtain the user's editing operation on at least one personal knowledge entry corresponding to the knowledge source prompt information; In response to the editing operation, the edited personal knowledge entry is determined; Based on the edited personal knowledge entries, a response to the data processing request is regenerated.
18. The method of claim 17, further comprising, before obtaining the user's editing operation on the at least one personal knowledge entry corresponding to the knowledge source prompt information: A knowledge viewing control is displayed in association with the knowledge source prompt information; Obtain the user's trigger operation on the knowledge viewing control; In response to the triggering operation, the at least one personal knowledge entry on which the generation of the response information depends is displayed.
19. The method of claim 17, wherein the editing operation includes a knowledge update operation; the method further includes: In response to the knowledge update operation, the edited personal knowledge entry is stored in the personal knowledge base.
20. The method of claim 1, wherein the output of clarification guidance information specifically includes: Determine the ambiguous target content in the data processing request; Output clarification and guidance information for the target content; The clarification guidance information includes clarification examples for the target content.
21. The method of claim 1, wherein the data processing request includes multiple processing tasks; the ambiguity identification of the data processing request specifically includes: Ambiguity identification is performed on at least one of the plurality of processing tasks.
22. The method of claim 21, wherein if at least two of the plurality of processing tasks are identified as ambiguous, the clarification guidance information is used to guide the user to clarify the target content causing the ambiguity in the at least two tasks in parallel.
23. The method of claim 21, wherein the plurality of processing tasks includes a first processing task and a second processing task dependent on the first processing task; The output clarification guidance information specifically includes: Output the first clarification guidance information for the first target content in the first processing task; After parsing the first processing task based on the first clarification information input for the first clarification guidance information, the second clarification guidance information for the second target content in the second processing task is output.
24. An apparatus for generating response information, comprising: The data processing request acquisition module is used to acquire users' data processing requests in natural language for the target dataset; An ambiguity identification module is used to identify ambiguities in the data processing request. The clarification guidance information output module is used to output clarification guidance information if the data processing request is found to be ambiguous. The clarification guidance information is used to guide the user to clarify the ambiguous target content in the data processing request; The clarification information acquisition module is used to acquire the clarification information input by the user regarding the target content; The response information output module is used to output response information based on the data processing request and the clarification information.
25. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 23.