Method and system for realizing data question and answer and computer readable medium
By identifying the indicator entities and other entities in user input, determining business scenarios and generating query requests, the accuracy and flexibility of data queries in the customer operation industry are solved, and an efficient and accurate data query experience is achieved.
Patent Information
- Application Number
- CN202510026532.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The prior art is difficult to provide accurate and flexible data queries in the passenger operation industry, especially when dealing with special expressions and calculation indicators in specific fields, which affects the accuracy and user experience of query results.
By identifying the metric entities and other entities in user input, determining the business scenarios associated with them, and generating request messages for data queries, thereby achieving efficient query of customer operation industry data.
It improves the accuracy and efficiency of data queries, enables users to obtain data in the customer operation industry in a convenient way, reduces users' need to understand data storage, and enhances operational flexibility.
Smart Images

Figure CN119961401A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of passenger transportation, and more particularly, to a method, system, and computer-readable medium for implementing data question answering in the field of passenger transportation. Background Art
[0002] With the development of technology, more and more data query methods have emerged, such as business intelligence (BI) report search, automatic conversion of search statements, intelligent dialogue robots, general query applications, etc.
[0003] Although BI report search technology can achieve a certain degree of query on the data in the report, it is unable to provide data information in a targeted manner according to the user after the report is determined and developed, and has poor flexibility. In addition, when there are a large number of reports within the enterprise, it is impossible to quickly locate the required report due to the cumbersome process of querying and opening the report. Although the automatic conversion technology of search statements can generate search statements based on natural language, the generated statements sometimes have errors, especially the technology cannot correctly process special information such as conditions or indicators contained in natural language, which greatly affects the accuracy of the query results. Intelligent conversational robot services can achieve relatively general data query, but cannot provide customized query requirements for specific fields. For example, for the aviation industry, it cannot parse some calculation indicators that are not in the table structure expression of the data table, nor can it understand special expressions in specific fields. Although general query applications can support a certain degree of data query, users are required to pre-specify the data table and then generate query statements based on the pre-specified data table, which requires users who query data to understand the storage of data.
[0004] The above query method, at least due to its universality, is difficult to understand special expressions in a specific passenger transport industry, and thus is difficult to provide accurate data query in the passenger transport industry. Moreover, at least because the user needs to understand the storage of data, it provides a lot of burden to the user and lacks flexibility.
[0005] Therefore, it is hoped to provide a method for better data query in the passenger transport industry, so that users can query data in the passenger transport industry efficiently in a convenient manner. Summary of the invention
[0006] A brief overview of the disclosure is given below in order to provide a basic understanding of some aspects of the disclosure. However, it should be understood that this overview is not an exhaustive overview of the disclosure. It is not intended to identify the key or important parts of the disclosure, nor is it intended to limit the scope of the disclosure. Its purpose is simply to give some concepts of the disclosure in a simplified form as a prelude to a more detailed description given later.
[0007] According to one aspect of the present disclosure, a method for implementing data question and answer is provided, the method comprising: receiving user input from a user; in response to determining that the user input is related to a data query, identifying an indicator entity contained in the user input, the indicator entity being used to indicate an indicator in the passenger transport industry; based on the identified indicator entity, determining a business scenario associated with a data source containing the indicated indicator; based on the determined business scenario, identifying one or more other entities from the user input, the other entities including at least one of a time condition entity indicating an occurrence time of the indicator entity and an area condition entity indicating an occurrence area of the indicator entity; based on all entities identified from the user input, generating a request message for data query; and displaying data obtained in response to the request message.
[0008] According to another aspect of the present disclosure, a system for implementing data question and answer is provided, the system comprising: a memory storing computer executable instructions; and a processor coupled to the memory, wherein the computer executable instructions, when executed by the processor, cause the processor to perform the above method.
[0009] According to another aspect of the present disclosure, a system for implementing data question and answer is provided, the system comprising components for executing the steps of the above method.
[0010] According to another aspect of the present disclosure, a non-transitory computer-readable medium is provided, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor is caused to perform the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The foregoing and other features and advantages of the present disclosure will become apparent from the following description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are incorporated herein and form a part of the specification, and are further used to explain the principles of the present disclosure and enable those skilled in the art to make and use the present disclosure. Among them:
[0012] Figure 1 A software architecture diagram of a data question-answering system according to an embodiment of the present disclosure is shown.
[0013] Figure 2 A structural schematic diagram of a data question and answer system according to an embodiment of the present disclosure is shown.
[0014] Figure 3 A flowchart of a method for implementing data question and answer according to an embodiment of the present disclosure is shown.
[0015] Figure 4 An example of a processing method of a question-answering system when an indicator entity has multiple interpretations according to an embodiment of the present disclosure is shown.
[0016] Figure 5 A flowchart of entity maintenance according to an embodiment of the present disclosure is shown.
[0017] Fig. 6A An example of defining an indicator through prompts according to an embodiment of the present disclosure is shown.
[0018] Figure 6B The SQL generation model according to the embodiment of the present disclosure is shown. Fig. 6A Screenshot of the metrics learned using the .
[0019] Figure 7 A flowchart of another method for implementing data question answering according to an embodiment of the present disclosure is shown.
[0020] Fig. 8A A screenshot showing the display of the question answering system when a user enters query text.
[0021] Figure 8B A screenshot showing content presented by the question-answering system through memory.
[0022] Figure 8C A screenshot showing an example of identifying the correct metrics using a fine-tuned LLM (Large Language Model).
[0023] Fig.8D Screenshot showing an example of a question-answering system asking a user to clarify the intent of a query.
[0024] Fig. 8E Shows when the user selects Fig.8D Screenshot of the Q&A system display when selecting options in .
[0025] Fig.8F A screenshot showing the display of the question-answering system when a specific year is specified.
[0026] Figure 8G A screenshot showing the display of the question-answering system when a specific event is specified.
[0027] Fig. 9 A timing diagram of data query according to an embodiment of the present disclosure is shown.
[0028] Fig.10 A timing diagram of knowledge base query according to an embodiment of the present disclosure is shown.
[0029] Fig.11 A block diagram showing an example structure of an information processing device that can be employed in the technology according to an embodiment of the present disclosure.
[0030] Note that in the embodiments described below, sometimes the same reference numerals are used in common between different drawings to represent the same parts or parts with the same functions, and their repeated descriptions are omitted. In some cases, similar numbers and letters are used to represent similar items, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0031] For ease of understanding, the position, size, range, etc. of each structure shown in the drawings and the like may not represent the actual position, size, range, etc. Therefore, the present disclosure is not limited to the position, size, range, etc. disclosed in the drawings and the like. DETAILED DESCRIPTION
[0032] Various exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0033] The following description of at least one exemplary embodiment is in fact merely illustrative and is in no way intended to limit the present disclosure and its application or use. That is, the structures and methods herein are shown in an exemplary manner to illustrate different embodiments of the structures and methods in the present disclosure. However, those skilled in the art will appreciate that they merely illustrate exemplary ways of the present disclosure that can be implemented, rather than exhaustive ways. In addition, the drawings need not be drawn to scale, and some features may be enlarged to illustrate the details of specific components.
[0034] In addition, technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.
[0035] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0036] First reference Figure 1 , describing the software architecture diagram of the data question and answer system according to an embodiment of the present disclosure.
[0037] The data question-and-answer system supports natural language understanding (NLU) and can give corresponding answers in the form of visualization or text after automatically analyzing user input. The question-and-answer system can include user-side devices (such as mobile phones, tablets, computers, etc.), and can also include various servers and databases on the network side. These devices give answers to questions raised by users through interaction. The application layer, service layer, and data layer can be deployed in the question-and-answer system. The application layer can support query requests from mobile terminals and PC terminals (including Web and desktop applications) and return corresponding query results. The application layer can include an independent permission management console to set different permissions for different users and / or stored data. For example, the permission management console can support atomic-level data permission management. For example, for users in the Shanghai department, they can be set to only view data from the Shanghai departure airport.
[0038] The service layer may include AskData API service, semantic parsing service, data search service and knowledge base service. AskData API (application programming interface) service is an API service for AskData. AskData is a platform that provides data query and analysis, including many existing platforms. The API service provides a unified query interface for the front end (mobile terminal or Web or desktop application) of data question and answer query through a message format protocol. The semantic parsing service is used to parse natural language to identify the target of natural language and the entities contained in natural language. The entity in this article refers to the important information extracted from natural language, including indicator entities that semantically represent indicators (also known as metrics), time condition entities that semantically represent time points or time periods (such as yesterday, the previous week, etc.), regional condition entities that semantically represent regions (such as starting from Shanghai, Shanghai to Beijing, etc.), query method entities that semantically represent query methods (such as by day, by month, respectively, highest, lowest, year-on-year, month-on-month), etc. The data search service is used to generate database statements (such as structured query languages (SQL) corresponding to different types of databases) for data queries, and summarize and render the query results. The knowledge base service is used to perform retrieval augmentation generation (RAG) queries for questions other than data queries to return relevant results for unstructured data.
[0039] The data layer can include application database, computing database, full-text index and vector database. The application database can store data set for AskData application, such as accessible users, permissions of each user, data dictionary, etc. The computing database can store the data source required for data question and answer, and the data source can be used to determine the data table where the data is stored. The full-text index can be used for document recall of RAG, and can also be used for query suggestion prompts of AskData. The vector database can be used for document recall of RAG, and can also be used to assist in parsing the entities contained in the user input.
[0040] The model layer can include a large model group and a model training module. The large model group can call appropriate models to work together, such as using LLM-HUB (which is used for unified integrated scheduling of internal LLM) and the API for calling LLM to call related models to implement corresponding functions. For example, an existing semantic parsing model can be called to analyze the semantics of natural language, an existing RAG model can be called to generate answers for non-data queries, and so on. The model training module can fine-tune the existing LLM, such as for specific users, specific expressions of specific industries, RAG summarization of specific documents, etc., to fine-tune the existing LLM to better suit the application scenario.
[0041] exist Figure 2 FIG. 2 shows a schematic diagram of the structure of a data question answering system according to an embodiment of the present disclosure. Figure 2 As shown, the client of the application layer (including applications or web pages on smartphones, laptops, desktop computers, tablets, etc.) can communicate with the AskData API service via RESTFUL / SSE using the AskData message protocol to perform data searches, etc. The permission management console of the application layer can communicate with the AskData API service via RESTFUL to set user permissions and data attributes in the application database.
[0042] When the AskData API service of the service layer receives a message corresponding to a question from the client, the AskData API service analyzes whether the target of the message is to perform a data query. When the target is a data query, the API service communicates with the semantic parsing service through RESTFUL. The API service sends the user input to the semantic parsing service so that the semantic parsing service performs semantic classification based on the user input, and parses out the entities contained in the user input, and can determine the relevant intent based on the entity. For example, the intent of an entity can indicate the data source corresponding to the entity. During the parsing process, the semantic parsing service can call functions for performing text similarity matching, synonym matching, and vector query to determine the entities in the user input. When there are resources such as GPU (graphics processing unit), the semantic parsing service can also call LLM to determine the entity. For example, the semantic parsing service can parse the entity by means of the semantic recognition LLM called by LLM-HUB in the model layer using GPU computing reasoning to strengthen the semantic error correction of the user input (for example, when there are typos in the user input). After the semantic parsing service parses out the entity and its intent (for example, the corresponding data source), the entity and query intent are returned to the API service. The API service can communicate with the application database and computing database in the data layer through JDBC, and / or communicate with the data interface in the data layer through RESTFUL / SOAP, so as to obtain corresponding data from the computing database and / or data interface when it is determined based on the application database according to the user's attributes and / or the attributes of the data that the user can query related data, and perform operations on the obtained data to meet the user's query requirements.
[0043] For example, when a user enters a query text, the query text is submitted to the API service. The API service extracts the intent permission code or scene permission code that the user can access and the permission code of the entity involved in the query text from the application database based on the user ID, and submits these permission codes to the semantic parsing service. The semantic parsing service determines whether the user needs to make clarification based on these permission codes (described below), and directly gives the semantic parsing result of the query text if clarification is not required, and gives the semantic parsing result of the query text based on the user's clarification content if clarification is required. The semantic parsing service feeds back the semantic parsing results including scenes, intentions, and entities to the API service, and the API service then determines whether to allow the user to make data queries based on the user's other permission codes (for example, a user can only query data from airports departing from Beijing). When other permission codes are not met, the user is denied data query. When other permission codes are met and it is determined that the user is allowed to make data queries, query statements or interface parameters are generated based on the query text, and the query results are rendered to present them to the user.
[0044] When the target is not a data query (for example, the user input is unstructured data, and the request is to answer the cause of the matter, the processing steps, etc.), the API service communicates with the knowledge base service through RESTFUL / SSE to perform a knowledge base retrieval query. The API service sends the user input to the knowledge base service so that the knowledge base service performs RAG based on the user input to perform a knowledge base query. When performing a knowledge base query, the knowledge base service accesses the knowledge base LLM, embedding (EMBEDDING) model, reranking (Reranking) model, etc. in the model layer to convert the user input into relevant features, and performs full-text indexing and / or vector search in the document / vector library based on the relevant features to find and generate results that meet the user input.
[0045] Although the above description uses the API service to determine whether the user input is a data query as an example, those skilled in the art can understand that after obtaining the user input, the application on the client can determine whether the user input is a data query by itself, and indicate whether it is a data query in the message provided to the API service.
[0046] Next, the data question and answer process for implementing data query is described in detail. Figure 3 A flow chart of a method 300 for implementing data question and answer according to an embodiment of the present disclosure is shown in FIG. The method can be executed by a data question and answer system or application. Through the execution of the method, more accurate data query can be provided for the passenger transport industry by means of more accurate identification of entities, so that the efficiency of data question and answer can be improved.
[0047] In S310 , a user input is received from a user.
[0048] A user can input a question to an information processing device such as a mobile phone, tablet, or computer that can access the question-answering system or application to request the question-answering system or application to return a corresponding answer. The question can be a request for data query. The question can be input in the form of text or voice, and the information processing device converts the voice into text.
[0049] In S320 , in response to determining that the user input is related to the data query, an indicator entity included in the user input is identified, where the indicator entity is used to indicate an indicator in the passenger transport industry.
[0050] According to an embodiment of the present disclosure, when it is identified that the user input contains expressions related to indicators in the passenger transport industry, it is determined that the user input is related to data query. For example, if the user input includes "number of passengers", "occupancy rate", "flight destination", "aircraft model usage time", "income from XX to YY", etc., it can be determined that the user input is related to data query. Determining whether the user input is related to data query is a pre-classification of the target of the user input, which corresponds to making a coarse-grained judgment on the query statement to determine whether it is a data question answering or a knowledge base question answering, so that, for example, Figure 2 The user input is sent to the semantic parsing service or knowledge base service.
[0051] As mentioned above, the entity in this article refers to the important information in the query message input by the user. The indicator entity in the user input can be an entity that semantically indicates an indicator in the passenger transport industry. The indicator entity can indicate specific indicators, such as "passenger number", "passenger load factor", "passenger-kilometers", etc., or it can indicate general indicators, such as "operating status", "flight status", etc. These general indicators can correspond to more specific indicator data in the data table, for example, "flight status" can correspond to more specific "flight mileage", "flight time", etc.
[0052] In order to identify indicator entities from user input, the indicator entities contained in the user input can be first determined based on text similarity matching, synonym matching and vector query. For example, the indicator can be set in advance, and when the user input contains content that matches the preset indicator, the relevant indicator entity is determined. However, in some cases, there may be typos in the user input, which affects the implementation of text similarity matching, synonym matching and vector query, so that the indicator entity cannot be found. At this time, the indicator entity contained in the user input can be identified according to the indicator recognition model based on machine learning. For example, the existing model for identifying specific content can be fine-tuned so that the model can identify the correct content based on the typos. For example, the typo indicator as input and the corresponding correct indicator as output can be used as training samples to fine-tune the model. In this way, even if there are typos in the user input, the correct indicator entity can be identified from it to achieve enhanced semantic error correction.
[0053] In S330 , based on the identified indicator entity, a business scenario associated with a data source containing the indicated indicator is determined.
[0054] After identifying the indicator entity, the data source containing the indicator indicated by the indicator entity can be further determined. The data source stores the data corresponding to the indicator entity. The query intent of the indicator entity can be determined through the indicator indicated by the indicator entity. The data source can indicate a business scenario, and multiple data sources can be associated with the business scenario. Since the query intent can be determined through the indicator entity, that is, what type of data you want to query, the data source can be determined. Since the data source is associated with the business scenario, the relevant data can be further queried based on one or more data sources associated with the business scenario, thereby realizing cross-data source queries, avoiding users from repeatedly querying different operations required by different data sources, and trying to ensure the integrity of the query data.
[0055] For example, when the indicator entity is "number of passengers", the relevant data sources can be determined in the database, such as the "operation table" containing the "number of passengers", the "revenue table" containing the "number of passengers", and / or the "plan table" containing the "number of passengers". These data sources belong to the operational business scenario together. After the operational business scenario is determined, the scenario may be associated with more data sources, which helps to query the data more comprehensively. For another example, when the indicator entity is "number of employees", it can be determined in the database that the relevant data source is the "human resources table" containing the "number of employees", and the associated business scenario is determined to be the internal management business scenario through the "human resources table", and more data related to the entity contained in the user input may be discovered.
[0056] Generally, the query intent indicated by the indicator entity corresponds to a data source and a business scenario, but in some cases, the indicator entity may be a general expression of the indicator. In this case, the query intent indicated by the indicator entity may correspond to multiple data sources. In this case, the associated business scenario can be determined through the general indicator entity (for example, the business scenario is determined through at least one related data source), and based on the multiple specific items in the business scenario corresponding to the indicator entity and contained in the associated data source, the indicator entity is decomposed into corresponding specific indicators so that these indicators are included in the displayed results. For example, when the indicator entity is "operating status", it can be determined that the business scenario related to it is an operating scenario. Then, the specific items "passenger transportation", "passenger ticket sales" and "financial performance" can be determined from the data source associated with the scenario. These specific items correspond to "operating status", so that the indicator entity "operating status" is concretized as "passenger transportation", "passenger ticket sales" and "financial performance", and the corresponding query results are displayed in the results. These specific items represent the query intent of the general indicator entity, and these query intents can be combined and connected in series in the form of a linked list to form an intent chain.
[0057] According to an embodiment of the present disclosure, when the identified indicator entity has multiple interpretations, the exact meaning of the indicator entity can be inferred based on the context of the user input and / or the user's permissions. For example, when the user input contains "number of people", if the user only has the authority to query the company's human resources, then it can be inferred based on the user's permissions that the "number of people" refers to the "number of employees". When inference cannot be made, the user can be requested to clarify the meaning of the indicator entity to determine the exact meaning of the indicator entity. For example, continuing with the above example, the user can be requested to clarify whether the "number of people" refers to the "number of employees" or the "number of passengers". After determining the exact meaning of the indicator entity, the business scenario corresponding to the data source containing the indicated indicator can be determined accordingly.
[0058] exist Figure 4 FIG. 2 shows an example of a processing method of a question-and-answer system or a question-and-answer application when an indicator entity has multiple interpretations according to an embodiment of the present disclosure. Figure 4 As shown in (1), it is assumed that two groups of permission configuration data are stored in the application database: one group is the query permission for operating indicators, which supports querying passenger indicators, passenger indicator dates, and financial indicators for specific flight segments; the other group is the query permission for human performance, which supports querying the personnel structure of a specific department.
[0059] exist Figure 4 In (2), the user is an administrator with the permissions (or role permissions) to query operating indicators and human performance. When the user enters the query "number of people in September", the question-answering system recognizes that the user may want to query the passenger indicator summary or the personnel structure, but cannot determine which type of number of people the user wants to query. Therefore, the question-answering system will infer the user's intention. When the question-answering system cannot infer the specific meaning based on the user's permissions and the context of the user's input (including entities in the user's input), the question-answering system will request clarification from the user, such as asking the user to choose "number of carriers" or "number of employees" in the form of a dialogue. After the user responds, the question-answering system can replace the ambiguous content with the content clarified by the user to query the results based on the clarified intention.
[0060] exist Figure 4 In (3), the user is an administrator with the rights to query operating indicators and human performance. When the user enters the query "the number of people from Beijing to Shanghai in September", the question-answering system recognizes that the user may want to query the passenger transport indicator summary or the personnel structure, but cannot determine which type of number the user wants to query. Therefore, the question-answering system infers the user's intention. Based on the "Beijing to Shanghai" included in the context of the user's input, the question-answering system can determine that the user wants to query the passenger transport indicator summary in the operating indicators at this time, so it queries the results based on the passenger transport indicator summary without asking the user for further clarification.
[0061] exist Figure 4 In (4), the user has the permission to query operating indicators. When the user enters the query "number of people in September", the question-answering system recognizes that the user may want to query the passenger transport indicator summary or the personnel structure, but cannot determine which type of number of people the user wants to query. Therefore, the question-answering system will infer the user's intention. Based on the user's permission to query only operating indicators, the question-answering system determines that the user can only query the passenger transport indicator summary, and then queries the results based on the passenger transport indicator summary without asking the user for further clarification.
[0062] exist Figure 4 In (5), the user has the permission to query human performance. When the user enters the query "the number of people from Beijing to Shanghai in September", the question-answering system recognizes that the user may want to query the passenger transport indicator summary or the personnel structure, but cannot determine which type of number the user wants to query. Therefore, the question-answering system will infer the user's intention. Based on the user's permission to only query human performance, the question-answering system determines that the user can only query the personnel structure, and then queries the results based on the personnel structure without requiring the user to make further clarification. At this time, the entity "Beijing to Shanghai" that is not related to human performance is discarded.
[0063] return Figure 3 In S340, based on the determined business scenario, one or more other entities are identified from the user input, and the other entities include at least one of a time condition entity indicating the time when the indicator entity occurs and a region condition entity indicating the region where the indicator entity occurs.
[0064] According to an embodiment of the present disclosure, one or more other entities corresponding to the items are identified from the user input based on the items contained in the data source associated with the business scenario. Specifically, the data source (such as a data table) associated with the business scenario may store one or more items (for example, origin, destination, time, number of passengers, revenue, etc.), and each item may have related data. After the items are determined based on the data source corresponding to the business scenario, the entities corresponding to these items can be identified by matching the items with the content input by the user. For example, when the items in the data source include "from XX to YY" and "time", the "Beijing to Shanghai" regional condition entity and the "September" time condition entity can be identified from the user query "number of people from Beijing to Shanghai in September".
[0065] Other entities identified here may include time condition entities indicating the occurrence time of the indicator entity involved in the user input and / or area condition entities indicating the occurrence area of the indicator entity, and may also include query mode entities used to indicate how the user wants to process the indicator, such as "year-on-year", "month-on-month", "list", "proportion", etc. In addition, other entities identified here may also have the problem of multiple interpretations similar to the indicator entity, so the same Figure 4 Similar approaches allow users to clarify or infer based on context and / or authority.
[0066] According to the embodiments of the present disclosure, a query mode can be identified based on the semantics of the user input. For example, the user can request to list different data, or request to sum some data, or request to compare some data, such as determining year-on-year / month-on-month growth, etc.
[0067] For example, the user input may include expressions related to "year-on-year" and "month-on-month". Since the present disclosure is applied to the passenger transport industry, the way of calculating year-on-year and month-on-month is different from the usual quarterly and annual calculation methods. Specifically, the year-on-year and month-on-month calculations of the present disclosure are aligned based on weeks or events to better adapt to the characteristics of the passenger transport industry. For example, if the user input contains expressions related to year-on-year, the year-on-year is determined based on the same week of the previous 12 months. For another example, if the user input contains expressions related to month-on-month, the month-on-month is determined based on the same week of the previous week. For another example, if the user input contains expressions related to events, the year-on-year is determined based on the time corresponding to the same event last year. The events here are, for example, the Spring Festival, the May Day holiday, the National Day, etc. If the time determined as the basis for year-on-year and / or month-on-month in the above manner does not meet the predetermined conditions, the time is further adjusted forward by one or more years. For example, due to special reasons such as the epidemic of infectious diseases in a certain year, the passenger data is inaccurate, then the year is adjusted forward to the most recent year without the epidemic of infectious diseases.
[0068] Through the above processing, the user input can be fully parsed to understand the semantics of the user input, thereby facilitating more targeted and accurate queries in the question-answering system.
[0069] In S350 , a request message for data query is generated according to all entities identified from the user input.
[0070] According to an embodiment of the present disclosure, a request message that conforms to the semantics of the user input can be generated according to the indicator entity identified in S320 and other entities identified in S340. After identifying the indicator entity, time condition entity, regional condition entity, query method entity, etc., since different data sources may have different query statements or query requirements, the corresponding request message can be generated according to whether the data source is accessed through a database statement or through an interface call. For example, when the data source associated with the business scenario is accessed through a database statement, the SQL of the database can be generated according to the entity obtained by parsing. Since the user input involves relevant query methods, the generated SQL can be combined to obtain processing corresponding to the query method. Of course, the relevant data can also be obtained through SQL first, and then the data can be processed corresponding to the query method. For another example, when the data source associated with the business scenario is accessed through an interface call, the interface parameters for calling the specified interface can be generated according to the entity obtained by parsing. The interface parameters are related to the specification of the interface, and the relevant information in the user input can be filled in according to the specification.
[0071] Although there are currently methods such as NL2SQL that convert natural language into SQL, these methods do not involve operations such as identifying indicator entities and identifying other entities based on business scenarios based on indicator entities as in the embodiments of the present disclosure. Therefore, their degree of parsing or understanding of natural language is not as accurate as the embodiments of the present disclosure, resulting in the possibility of errors in the SQL statements they generate, and the inability to achieve accurate queries. When querying data, these methods generally display the generated SQL content at the same time and allow users to determine the accuracy of the query results, making these methods not available to a wider range of business users. The embodiments of the present disclosure fine-tune the existing SQL generation model so that the generated SQL meets the definition of indicators in the passenger transport industry, so that the data table fields can be appropriately combined according to the content that the user wants to query to obtain the correct indicators.
[0072] According to the embodiments of the present disclosure, entity maintenance can be performed before the user performs a data question and answer query. The question and answer system or application according to the embodiments of the present disclosure can implement a preparation process for maintaining the question and answer entities. First, a confirmed correct SQL fragment can be generated for each entity by selecting a data source and defining indicators. This part of the operation can be completed by the large language model in the form of a prompt, and the management and maintenance personnel can confirm or adjust the generated content to ensure the accuracy of the output content in the question and answer query. For example, in Figure 5 A flowchart of entity maintenance according to an embodiment of the present disclosure is shown in FIG.
[0073] In 5.1, entity maintenance begins. The entities here include indicator entities and condition entities (such as time condition entities and regional condition entities). In 5.2, the data source is selected. Since SQL is involved here, the data source is the data source corresponding to the database, from which data can be extracted through SQL.
[0074] When the entity is an indicator entity, go to 5.4. In 5.4, filter the indicator entity field, for example, set which text corresponds to the indicator entity. In 5.4.1, define the indicator, for example, define what indicator form the indicator entity is in the data source, which may be a specific item in the data source, or some kind of operation between specific items (such as calculating the ratio). Fig. 6A An example of defining an indicator using the prompt method is shown in FIG. Figure 6B The existing LLM (here is the SQL generation model) is based on Fig. 6A If the learned indicator is still incorrect, the operator needs to make further modifications. In 5.4.2, use LLM to generate SQL fragments. In 5.4.3, confirm the SQL fragments.
[0075] When the entity is a conditional entity, go to 5.3. In 5.3, filter the conditional entity fields, for example, set which texts correspond to the conditional entity. In 5.3.1, define the type of the conditional entity to determine whether it is a time condition or a regional condition, so as to correspond to different items of the data source. In 5.3.2, determine the enumeration value of the conditional entity. Then, in 5.2, summarize the information of 5.3.2 and 5.4.3 to form SQL fragment mapping information to generate the corresponding SQL language.
[0076] return Figure 3 , in S360 , the data obtained in response to the request message is displayed.
[0077] After the question-answering system obtains data related to the entity in the user input through a database query statement or an interface call instruction, it can present this data to the user in response to the user's data query.
[0078] Based on the above technical solution, by determining relevant business scenarios based on indicator entities related to the passenger transport industry, and identifying other entities in the user input based on the business scenarios, the user's query intention can be more accurately identified based on the user input, thereby more accurately returning data corresponding to the query intention, allowing users to conveniently query data in the passenger transport industry by providing natural language, which can not only simplify user operations and enhance operational flexibility, but also provide accurate query data for the passenger transport industry, thereby improving data query efficiency and enhancing user experience.
[0079] In addition, in some cases, the user input may not be related to the data query. For example, the user input does not contain any statements related to indicators in the passenger transport industry. At this point, the question-answering system can determine that the user input is not a data query. In this case, retrieval-augmented generation (RAG) can be used to return an answer corresponding to the user input. For example, information related to the user's question can be retrieved from an external or internal document database, knowledge base, etc., and the retrieved information can be combined with the generative model of the language model architecture to generate an answer using the retrieved content as additional context.
[0080] exist Figure 7 A flowchart of another method 700 for implementing data question answering according to an embodiment of the present disclosure is shown in FIG.
[0081] In 1.0, users input natural language to allow the question-answering system to obtain user query text. For example, users can input the query request "The number of passengers from two airports in Shanghai to Beijing's Daxing Airport from the 5th to the 25th of last month, excluding FM flights and 787 aircraft" in the form of voice or text. Fig. 8A ] A screenshot of the display of the question-answering system when the user enters the query text is shown in FIG.
[0082] In 1.1, the question-answering system pre-classifies the target of the user input to make a coarse-grained judgment on the query text, thereby determining whether the user input is a data query (also known as data question answering) or a non-data query (also known as knowledge base question answering).
[0083] If it is determined in 1.1 that the user input is a data query, proceed to 1.2. In 1.2., the question-answering system parses the user's query intent, that is, parses the semantics of the user input. The specific process of semantic parsing is shown in 2.1-2.5.
[0084] In 2.1, the context information of the previous conversation can be maintained by memorizing the conversation within a certain period of time, so as to query the relevant content for the semantic analysis of the current user input. The context information may supplement the semantics of the user input, such as supplementing the entities involved in the user input, which may reduce the triggering of the above-mentioned clarification of the meaning of the entity. Figure 8B A screenshot showing content presented by the question-answering system through memory.
[0085] In 2.2, the question-and-answer system can extract indicator entities from user input. The disclosed embodiment can classify natural language entities into multiple categories, among which the business scenario of data query can be determined through the indicator entity (as shown in 2.3). When extracting indicator entities, the full-text index of synonyms and professional vocabulary can be used to match the indicator entity. If the indicator entity is not matched, the existing vector search and RERANKING threshold scheme can be used to match the indicator entity. If the indicator entity is still not matched, enhanced recognition can be performed through the intervention of LLM. Different from conventional semantic understanding, the disclosed embodiment utilizes LLM for enhanced recognition, which can improve the accuracy of entity recognition and enhance error correction capabilities.
[0086] Here, LLM can be an enhanced indicator extractor implemented by SFT (supervised fine-tuning) and FewShot (small sample learning) solutions. For example, when the user enters an incorrect word or spelling, it may not be possible to match the entity through the full-text index and word vector. In this case, LLM can be used to enhance the recognition ability of the entity. Figure 8C A screenshot of an example of using the fine-tuned LLM to identify the correct indicator is shown in . In this example, the user input an incorrect indicator entity "Gram Seat Rate", which the question-answering system correctly identified as "Passenger Seat Rate".
[0087] When fine-tuning an existing indicator recognition model (for example, which can be implemented by a general large language model) through FewShot's Prompt, an example of Prompt can be as follows: [{'role':'system','content':'You are a named entity recognition extractor. Your task is to extract "indicator" information from the sentence. Here are some learning examples'},{'role':'user','content':'Occupancy rate and revenue on January 1'},{'role':'assistant','content':'Indicators: Occupancy rate, revenue'},{'role':'user','content':'Aircraft utilization in December'},{'role':'assistant','content':'Indicators: Aircraft utilization'}].
[0088] Since multiple indicators with some identical text can be stored in the database, when the content input by the user may correspond to multiple indicators, the question-answering system can determine what data the user wants to query through the context, etc., to clarify the user's query intention. However, the question-answering system may not be able to infer the user's query intention in some cases, so a request can be sent to the user to let the user clarify the query intention. Fig.8D A screenshot of an example of a question-answering system asking a user to clarify the intent of a query is shown in FIG. Fig. 8EWhen the user selects Fig.8D A screenshot of the Q&A system display when checking the "Number of Carriers" field.
[0089] The clarification mechanism can determine the user's exact query intent based on the various types of information input by the user. If the exact meaning of the entity input by the user can be determined based on the user's permission to only query a certain type of data, then the user's query intent can be determined without going through the clarification mechanism, and the clarification mechanism will not be triggered. In some cases, the exact meaning of the indicator entity input by the user can be determined by the information of other entities in the context contained in the user input (for example, entities specific to the query intent contained in the user input), and the clarification mechanism can be skipped to directly determine the user's query intent. Unlike the clarification function of other data question-and-answer systems, the disclosed embodiment can make a comprehensive judgment based on the context of the user input (for example, content that provides additional information and / or entities specific to the query intent) and / or the user's permissions to infer the user's query intent, thereby reducing the triggering of the clarification mechanism and improving query efficiency.
[0090] In 2.3, the relevant business scenarios can be determined through the extracted indicator entities. The current data question-and-answer system does not determine the scenarios, so it is impossible to understand the indicator entities from the more general semantics that the indicator entities may involve, which may make it difficult to fully discover the relevant data sources. By determining the business scenarios, the question-and-answer application can be made more convenient for business users to use without the need for users to understand the data storage method.
[0091] Other existing data question and answer applications require users to know in advance what data sources there are and what indicators can be calculated based on the data sources. For example, when a user wants to query various indicators such as revenue, number of passengers, average fare, etc., the user needs to indicate these indicators separately, and the data across data sources may need to be queried multiple times, which will increase the use cost of business users. According to the question and answer application of the embodiment of the present disclosure, the business scenario can be determined according to the indicator entity, so as to combine different query intentions in the business scenario according to the query needs of the user. For example, when the more general or vague indicator entity (for example, an indicator entity that cannot directly correspond to an item in any data source) input by the user corresponds to the "operating status" scenario, the more general or vague indicator entity can be decomposed into a combination of various data such as "passenger transport" + "ticket sales" + "financial performance", and the query intentions represented by these data can be combined in series in a linked list to form an intention chain, thereby avoiding the user having to enter precise indicator entities and adding burden to them. In this way, compared with other question and answer applications, users can focus more on familiarizing themselves with the business rather than familiarizing themselves with the data, thereby improving the user's usage efficiency.
[0092] In 2.4, according to the user's query intent, for example, according to each item in the data source corresponding to the indicator entity, the time condition entity, the regional condition entity, the query method entity, etc. are extracted from the user input. Taking "the number of passengers from two airports in Shanghai to Beijing's Daxing Airport from the 5th to the 25th of last month, excluding FM flights and 787 models" as an example, the corresponding data source can be determined through the indicator entity "number of passengers", "the 5th to the 25th of last month" can be parsed into [yyyymmdd, yyyymmdd] according to the time parsing rule engine for easy query, "Shanghai Airport" can be parsed into "SHA, PVG", Beijing Daxing Airport can be parsed into "PKX", and "excluding FM flights and 787 models" can be parsed into the exclusion items of the corresponding flight entity and model entity. The parsed form can be related to the item form in the data source, so that it is convenient to use related expressions to query the data source. In order to determine the multiple types of entities contained in the user input, rule parsing, similarity matching and / or syllable matching can be used. These methods are well known to those skilled in the art and will not be repeated here.
[0093] In 2.5, if the query method is determined to include year-on-year and month-on-month analysis by semantically parsing the user input, the corresponding year-on-year and month-on-month analysis is generated according to the semantics. In other data question-and-answer applications, year-on-year and month-on-month analysis can only meet the year-on-year and month-on-month analysis of a specific month or quarter. However, the passenger transport industry is not suitable for this rule. For example, the total number of days in February and March is different, and the number of passengers will be quite different. If the user specifies any time, such as March 1 to March 5, the same 1 to 5 days as last year will also be different because of whether the weekend is included. According to the data question-and-answer application of the embodiment of the present disclosure, the rule of week alignment ("DayOfWeek" rule, referred to as DOW rule) can be adopted to ensure that the year-on-year and month-on-month calculation is more in line with the special circumstances of the passenger transport industry. The specific implementation method can be: year-on-year is implemented by [T-364, T-364], and month-on-month is implemented by [T-7, T-7], where T is a specific date. In addition, due to special reasons such as the epidemic of infectious diseases, the data of recent years will be distorted compared with the same period last year, so it can be corrected according to the rules before the epidemic of infectious diseases occurs.
[0094] Based on the indicators that users want to query, you can determine how to set the year-on-year comparison. For example, for the above-mentioned operating data, you can set the year-on-year comparison according to the above-mentioned DOW rules, while for data that is not affected by time, such as sales and personnel, you do not need to use the DOW rules, but use the usual annual or quarterly time alignment.
[0095] In addition, semantic analysis can also determine the time period that the user wants to compare. For example, the user can specify a comparison of a specific year. Fig.8FA screenshot of the question-and-answer system display when a specific year is specified is shown in . The question-and-answer system can support users to enter a specific time range. In this case, the DOW rule cannot be used, but the comparison day range can be reduced or increased based on the number of days queried. The question-and-answer system can also support corresponding event ratios, such as "Spring Festival Travel," "Labor Day Holiday," "National Day Holiday," etc. Because the length of holidays varies from year to year, the comparison time will also be leveled according to the user's query intent and the rules of the event. Figure 8G A screenshot of the display of the question-and-answer system when a specific event is specified is shown in FIG.
[0096] Through the execution of 1.2, the semantics of the user input can be identified, thereby determining the query intent and entities. For example, if the user inputs "yesterday's Shanghai to Beijing revenue, number of people, passenger load factor, seat revenue, passenger revenue", the following can be identified through semantics:
[0097] => Query Intent: Carrier Data
[0098] => Time condition entity: Yesterday->2024.09.09
[0099] =>Regional condition entity: Shanghai to Beijing->SHA / PVG-PEK / PKX
[0100] =>Indicator entity:
[0101] Revenue -> Passenger revenue
[0102] Number of people->Number of passengers
[0103] Load Factor
[0104] Seat revenue->Seat-kilometer revenue
[0105] Passenger Revenue -> Revenue Per Passenger Kilometers
[0106] Corresponding to the implementation method of data query, it can be converted into the following computer language:
[0107] =>Query intent: Carrier data->Data source, table=flight_data_reportl
[0108] => Time condition: 2024.09.09->flight_dt='2024-09-09'
[0109] =>Regional condition: SHA / PVG-PEK / PKX->orig IN('SHA','PVG')AND dest IN('PEK','PKX')
[0110] =>Indicators:
[0111] Income->SUM(earning)
[0112] Number of people->SUM(psg)
[0113] Passenger load factor->SUM(psg_km) / SUM(seats_km)
[0114] Seats earned->SUM(earning) / SUM(seats_km)
[0115] Customer revenue->SUM(earning) / SUM(psg_km)
[0116] After completing the semantic analysis of the user input, in 1.3, a permission check is performed to determine whether the user has the permission to query the desired data. Permissions can be subdivided into field values that support permission verification. For example, users can be restricted to querying data for only a few airports. If the user queries for airports beyond the queryable range, the Q&A system will directly terminate the subsequent data query.
[0117] If the user's query intention is information from a data table, it means that the data source to be queried needs to be accessed through a database statement, then go to 1.4.1. If the user's query intention is information called through an interface, it means that the data source to be queried needs to be accessed through an interface call, then go to 1.5.1. For example, the supported database according to the embodiment of the present disclosure may be MySQL / PostgreSQL, and the supported interface may be a Restful interface.
[0118] In 1.4.1, the table structure information of the data can be obtained according to the query intent. In 1.4.2, the information of the indicator entity can be mapped to the SQL statement of the corresponding database. In 1.4.3, the information of the condition entity can be mapped to the SQL statement of the corresponding database. In 1.4.4, the query method can be determined according to the semantics of the user input. In 1.4.5, the final query statement can be combined.
[0119] Regarding the query method, the default query method is the "overall aggregation and summary" of the data, for example, "income" means "total income". When the semantics of the user input include expressions such as "by", "per", and "based on", grouping and aggregation are performed according to the corresponding fields based on the obtained table structure information, such as "every day", "monthly", etc. When the semantics of the user input include ranking expressions, such as "the 10 highest-income routes", after extracting the entity "income", "income" will be used as the sorting rule, 10 as the TopN entity, and the route as the aggregate entity. In addition, the query method can also be to list specific data details, such as "what are the flight numbers that departed from Shanghai last week", which corresponds to "DISTINCT flight numbers".
[0120] In 1.5.1, you can obtain interface information based on the query intent. In 1.5.2, you can map the corresponding interface parameters based on the information of the indicator entity. In 1.5.3, you can map the corresponding interface parameters based on the information of the condition entity. In 1.5.4, you can generate an interface call instruction based on the entity splicing interface parameters to call the specified interface.
[0121] In 1.6.1, the query results can be aligned and spliced. For example, when the query intent is "overall aggregation summary" or "aggregation by a certain field" in the same-year and month-on-month sense, the year-on-year and month-on-month data can be queried multiple times and related operations can be performed according to the same-year and month-on-month rules of semantic analysis, and an "intention chain" can be formed. Similar processing can be applied to the query percentage. For example, the data obtained from multiple queries can be spliced into the final result according to the same indicators and dimensions, and the same-year and month-on-month data (such as growth rate, growth value, base point change, etc.) or percentage data (such as percentage, etc.) can be calculated based on the indicator information. The "intention chain" involved in the scenario obtained through semantic analysis can also be queried multiple times, and the query results can be finally spliced. The calculation results of year-on-year, month-on-month and percentage can be output in the same result table as the indicator data. Multiple queries can be made from multiple data sources according to the intention chain, and the query results can be combined.
[0122] In 1.7, the obtained data can be converted into display data, such as EChartsOptions data, according to the query method. For example, if the corresponding chart can be generated according to the query data (for example, a pie chart for percentage, a horizontal bar chart for TopN, and a line chart for date aggregation), then the JSON data of ECharts Option is generated.
[0123] In 1.8, the final result can be rendered in a specific format (such as AskData Message format). For example, the question-answering system can support rendering methods such as ordinary text, rich text, lists, data tables, graphics (ECharts data compatible), Markdown, chapters, etc. Using the Message format as output can facilitate different client presentations, thereby improving the user experience.
[0124] In 1.9, the final result obtained by rendering can be output to the client (for example, output through an application, a web page, etc.), so that the user can obtain the result corresponding to the query statement.
[0125] exist Fig. 9 A timing diagram of data query according to an embodiment of the present disclosure is shown in FIG.
[0126] like Fig. 9As shown in FIG, the user initiates a query request in natural language from the client. Next, the user query request is classified, and if it is determined that the user query request is a data query, the data query process is entered. Figure 2 The AskData API service shown in the figure receives the query request from the client again and communicates with the semantic analysis service (such as Figure 2The semantic parsing service shown in the figure) collaborates to perform semantic parsing of natural language, such as parsing natural language into indicator entities and their query intent, business scenarios, various conditional entities, etc. For unrecognizable user input content, the semantic parsing service sends it to the fine-tuned semantic LLM for enhanced parsing to perform semantic enhancement error correction. Here, text similarity matching, regular matching of specific entities (such as flights), synonym matching, and vector query can be used to determine entities, and entity classification can be determined, such as conditional entities, metric or indicator entities, query method entities (by day, by month, respectively, highest, lowest), etc. The exact meaning of the entity in the user input can be inferred based on the parsing of the entity to clarify its intent. For entities whose intent cannot be confirmed, a clarification mechanism can be implemented using multiple rounds of dialogue. For example, relative time range entities that cannot be handled by LLM can be parsed according to specific query intents (such as yesterday, last Wednesday, this year's Spring Festival travel, and a specified lunar calendar). For example, a user has the authority to query all data. When he enters "the number of people last week", he cannot judge from the query statement whether it is "the number of passengers last week" or "the total number of employees in the company as of last week", so the ambiguous query statement requires the user to make a second confirmation clarification; if the meaning can be judged by other entities in the query statement, such as "the number of people from Shanghai to Beijing last week", it can be inferred that it is a query of the number of passengers, so the data query intention is directly determined and the corresponding items in the data source are matched. For another example, industry-specific information (such as flight number abbreviations, city transfer airports, etc.) can be parsed and mapped to determine the corresponding specific items in the data table. The semantic parsing service can return the query intention and each entity obtained thereby to the AskData service, and the AskData service sends these intentions and entities and user attributes (such as user ID) to the authority configuration server to enable it to determine whether the user can query the relevant data. When the user has the query authority, the AskData service can interact with the data source or the data storage device called through the interface to query the corresponding data, for example, generate SQL statements based on the intention matching and the corresponding entities to query the data. For example, if the user input specifies a year-on-year calculation or an event cycle ratio calculation (such as "the ratio of this year's Spring Festival travel revenue to that of 23 years"), the corresponding year-on-year query is executed. If the user's query intent can be parsed as a combination of intent chains, multiple queries are performed according to different intents to obtain the corresponding data, thereby performing multiple queries with multiple intent combinations based on the scenario. The AskData service can combine the query results obtained in this way and return the results to the client through, for example, a unified AskData service message protocol. The client renders the returned results and presents them to the user.
[0127] exist Fig.10, a timing diagram of a knowledge base query according to an embodiment of the present disclosure is shown in FIG. Since the knowledge base query performed after determining that it is not a data query according to the user input in the present disclosure embodiment is basically the same as the prior art, it will not be described in detail. A person skilled in the art can understand the following based on the above description and Fig.10 The diagram is sufficient to fully understand the timing of knowledge base queries.
[0128] like Fig.10 As shown, the knowledge base service, embedding model and vector indexing / full-text indexing module pre-process the document, including document classification, document segmentation, and indexing by full text and vector according to permission requirements. The user initiates a query request from the client in natural language. Then, the user query request is classified, and the knowledge base query process is entered when it is determined that the user query request is not a data query. The AskData service receives the query request from the client again, and performs a knowledge base query in combination with other modules in combination with the user's query document permissions. For example, full text and vectors can be applied to achieve multi-way recall. The recalled content can be evaluated by the model to achieve secondary matching screening and find closer content. Then, for the closest content, the knowledge base can be used to summarize the content of the result through LLM execution. Finally, the generated results can be returned to the client in SSE mode through, for example, the unified AskData service Message protocol. The client renders the returned results and presents them to the user.
[0129] According to the above technical solution provided by the embodiment of the present disclosure, effective analysis specific to the passenger transport industry can be performed for user input for data query, so that data corresponding to the query intent can be returned more accurately. The question-and-answer system or question-and-answer application according to the embodiment of the present disclosure provides users with an easy-to-understand query entry, simplifies user operations, and allows users to query by simply entering natural language without being familiar with various reports, greatly reducing the burden on users and improving the efficiency of data query. In addition, by processing query adaptation for special expressions of the customer industry, the question-and-answer application is more in line with the characteristics of the passenger transport industry, and the accuracy of data query in the passenger transport industry is increased. The question-and-answer application according to the embodiment of the present disclosure is an enhanced data analysis tool that does not require understanding of the composition behind the data, does not require professional terms, and has almost no learning cost. It uses text expression to accurately identify specific nouns in the industry field, so that data can be quickly obtained. Through the unified abstraction of result rendering, it can be conveniently applied to various clients and can be embedded in any business system to support agile decision-making. Its flexible query combination can reduce the development and maintenance investment of customized reports in enterprises, and can greatly improve the ability to quickly implement data query requirements. In addition, the modular design of the system is replicable and supports simple and small amount of customized development and training, so that the method according to the embodiment of the present disclosure can be quickly implemented in the business field.
[0130] The above describes various exemplary methods and the like according to the embodiments of the present disclosure. It should be understood that the steps and / or operations of each method can also be combined with each other in any appropriate order, so as to similarly implement more or less operations than described. It should be understood that the machine-readable storage medium or the machine-executable instructions in the program product according to the embodiments of the present disclosure can be configured to perform operations corresponding to the above embodiments. When referring to the above embodiments, the embodiments of the machine-readable storage medium or the program product are clear to those skilled in the art, so they are not repeatedly described. The machine-readable storage medium and the program product for carrying or including the above-mentioned machine-executable instructions also fall within the scope of the present disclosure. Such storage media may include, but are not limited to, floppy disks, optical disks, magneto-optical disks, memory cards, memory sticks, and the like.
[0131] In addition, it should be understood that the above series of processes and devices can also be implemented by software and / or firmware. In the case of being implemented by software and / or firmware, from a storage medium or a network to a computer with a dedicated hardware structure, such as Fig.11 The information processing apparatus 1300 shown installs the programs constituting the software, and when the various programs are installed, the information processing apparatus can execute various functions and the like. Fig.11 : is a block diagram showing an example structure of an information processing device that can be employed in the technology according to an embodiment of the present disclosure.
[0132] exist Fig.11 In the embodiment, a central processing unit (CPU) 1301 executes various processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage section 1308 to a random access memory (RAM) 1303. In the RAM 1303, data required when the CPU 1301 executes various processes and the like is also stored as needed.
[0133] The CPU 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. To the bus 1304, an input / output interface 1305 is also connected.
[0134] The following components are connected to the input / output interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet.
[0135] A drive 1310 is also connected to the input / output interface 1305 as needed. A removable medium 1311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1310 as needed so that a computer program read therefrom is installed into the storage section 1308 as needed.
[0136] In the case where the above-described series of processing is realized by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 1311 .
[0137] Those skilled in the art should understand that such storage media is not limited to Fig.11 The removable medium 1311 shown has a program stored therein and is distributed separately from the device to provide the program to the user. Examples of the removable medium 1311 include magnetic disks (including floppy disks (registered trademark)), optical disks (including compact disk read-only memory (CD-ROM) and digital versatile disks (DVD)), magneto-optical disks (including minidiscs (MD) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be the ROM 1302, a hard disk included in the storage section 1308, or the like, in which the program is stored and distributed to the user together with the device containing them.
[0138] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, devices, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transient storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139] One or more exemplary embodiments of the present disclosure are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0140] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, the present disclosure does not exclude that with the development of computer technology in the future, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0141] Although one or more embodiments of the present disclosure provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment).
[0142] The terms "comprises", "includes" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, product or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, product or apparatus. In the absence of further restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or apparatus comprising the elements. For example, if the words "first", "second" and the like are used to indicate names, they do not indicate any particular order.
[0143] For the convenience of description, the above devices are described in various modules according to their functions. Of course, when implementing one or more embodiments of the present disclosure, the functions of each module can be implemented in the same or more software and / or hardware, or the module implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0144] The present disclosure is described with reference to the flowchart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the process and / or box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the function specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.
[0145] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a particular manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device that implements the functions specified in one or more flows of a flowchart and / or one or more blocks of a block diagram. These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operating steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows of a flowchart and / or one or more blocks of a block diagram.
[0146] Those skilled in the art will appreciate that one or more embodiments of the present disclosure may be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of the present disclosure may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0147] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0148] The same or similar parts between the various embodiments of the present disclosure can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiment. In the description of the present disclosure, the description of the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present disclosure, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in the present disclosure and the features of the different embodiments or examples without contradiction.
[0149] In addition, when used in the present disclosure, the words "herein," "above," "below," "hereunder," "above," and words of similar meaning shall refer to the present disclosure as a whole rather than to any particular portion of the present disclosure. Furthermore, unless expressly stated otherwise or understood otherwise in the context of use, conditional language used herein, such as "may," "might," "for example," "such as," and the like, is generally intended to express that certain embodiments include, while other embodiments do not include, certain features, elements, and / or states. Thus, such conditional language is generally not intended to imply that one or more embodiments require features, elements, and / or states in any way, or whether such features, elements, and / or states are included or performed in any particular embodiment.
[0150] The above description is only an example of one or more embodiments of the present disclosure, and is not intended to limit one or more embodiments of the present disclosure. For those skilled in the art, one or more embodiments of the present disclosure may have various changes and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the scope of the claims.
Claims
1. A method for implementing data question answering, comprising: receiving user input from a user; In response to determining that the user input is related to a data query, identifying an indicator entity included in the user input, where the indicator entity is used to indicate an indicator in the passenger transport industry; Based on the identified indicator entity, determining a business scenario associated with a data source containing the indicated indicator; According to the determined business scenario, one or more other entities are identified from the user input, wherein the other entities include at least one of a time condition entity indicating an occurrence time of the indicator entity and a region condition entity indicating an occurrence region of the indicator entity; Generate a request message for data query based on all entities identified from the user input; as well as Data obtained in response to the request message is displayed.
2. The method according to claim 1, further comprising: In response to identifying that the user input includes a statement related to indicators in the passenger transportation industry, it is determined that the user input is related to a data query.
3. The method according to claim 1, wherein: Identifying the indicator entity included in the user input includes: Determine the indicator entity contained in the user input according to text similarity matching, synonym matching and vector query; and When the indicator entity cannot be determined based on text similarity matching, synonym matching, and vector query, the indicator entity included in the user input is identified based on an indicator identification model based on machine learning.
4. The method according to claim 1, wherein: According to the determined business scenario, identifying one or more other entities from the user input includes: According to the items contained in the data source associated with the business scenario, the one or more other entities corresponding to the items are identified from the user input.
5. The method according to claim 1, further comprising at least one of the following: If the user input includes a year-on-year expression, the year-on-year is determined based on the same week in the previous 12 months; If the user input includes a year-on-year expression, determining the year-on-year based on the same week of the previous week; and If the user input includes an event-related expression, the year-on-year difference is determined based on the time corresponding to the same event last year.
6. The method according to claim 5, further comprising: If the period used as the basis for year-on-year and / or quarter-on-quarter comparisons does not meet the predetermined conditions, the period will be further adjusted forward by one or more years.
7. The method according to claim 1, further comprising: Determine whether the user has the authority to query data related to the indicator entity; as well as In the case where the user does not have the authority to query the data related to the indicator entity, the request message is not generated.
8. The method according to claim 1, wherein: The request message generated for data query includes: The corresponding request message is generated according to whether the data source associated with the business scenario is accessed through database statements or through interface calls.
9. The method according to claim 8, further comprising: When a data source associated with the business scenario is accessed through a database statement, a structured query language (SQL) of the database is generated according to all entities identified from the user input.
10. The method according to claim 9, further comprising: Fine-tune the machine learning-based SQL generation model to make the generated SQL conform to the definition of indicators in the passenger transport industry.
11. The method according to claim 8, further comprising: When a data source associated with a business scenario is accessed through an interface call, interface parameters for calling a specified interface are generated based on all entities identified from the user input.
12. The method according to claim 1, wherein: Based on the identified indicator entity, determining a business scenario associated with a data source containing the indicated indicator includes: When the identified indicator entity has multiple interpretations, the exact meaning of the indicator entity is inferred according to the context of the user input and / or the user's authority to determine the business scenario according to the exact meaning of the indicator entity.
13. The method according to claim 12, further comprising: When the exact meaning of the indicator entity cannot be inferred according to the context of the user input and / or the user's authority, the user is requested to clarify the meaning of the indicator entity to determine the exact meaning of the indicator entity according to the clarification.
14. The method according to claim 1, further comprising: When at least one of the one or more other entities identified has multiple interpretations, the exact meaning of the entity with multiple interpretations is inferred according to the context of the user input and / or the user's authority to generate a request message according to the exact meaning of the entity.
15. The method according to claim 14, further comprising: When the exact meaning of an entity with multiple interpretations cannot be inferred according to the context of the user input and / or the user's authority, the user is requested to clarify the meaning of the entity to determine the exact meaning of the entity.
16. The method according to claim 1, further comprising: When the indicator entity corresponds to multiple projects in a business scenario, the indicator entity is decomposed into indicator entities corresponding to the multiple projects.
17. The method according to claim 1, further comprising: In response to determining that the user input is not relevant to the data query, an answer corresponding to the user input is returned using a retrieval enhancement generation (RAG).
18. A system for implementing data question answering, comprising: a memory storing computer executable instructions; as well as A processor is coupled to the memory, wherein the computer executable instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-17.
19. A system for implementing data question answering, comprising components for performing the steps of the method according to any one of claims 1-17.
20. A non-transitory computer readable medium storing computer executable instructions, which when executed by a processor cause the processor to perform the method according to any one of claims 1-17.
Citation Information
Patent Citations
Text processing method and device, electronic equipment and storage medium
CN113743115A
Entity identification method, system and equipment for operator business performance index based on large model, and medium
CN119150866A
Information query method and device, computer equipment, storage medium and program product
CN119166769A
Entity and attribute resolution in conversational applications
US20150286747A1
Systems and methods for adaptive question answering
US20200279001A1