Data query method and device, equipment and medium
By obtaining the keyword sequence and pointer position of the query question, and using a multi-layered nested knowledge base and processing model to generate data query statements, the problem of irrelevant answers in artificial intelligence data queries is solved, and more efficient and accurate data query results are achieved.
Patent Information
- Application Number
- CN202511750919.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-06
AI Technical Summary
In existing technologies, artificial intelligence data query methods often provide irrelevant answers, making it difficult to meet actual needs in terms of accuracy and efficiency.
By obtaining the keyword sequence of the query question, determining the pointer position of the keyword, performing word extraction and verification, using a multi-layered nested knowledge base for retrieval, generating a data query statement, and finally processing and generating the query results in a pre-built processing model.
It improves the accuracy and efficiency of data retrieval, ensuring that query results more accurately meet user needs.
Smart Images

Figure CN121478831A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of big data technology, and in particular to a data query method, apparatus, device, and medium. Background Technology
[0002] In today's information age, the scale of data accumulated by enterprises and institutions continues to expand. Traditional query methods that rely on fixed keywords or manual screening often fail to locate the required information in a timely and accurate manner due to the complexity of data dimensions and the hidden correlation logic.
[0003] Against this backdrop, artificial intelligence, with its advantages in algorithms such as deep learning and natural language processing, can automatically analyze data relationships, understand semantic needs, and efficiently handle massive data retrieval tasks. It has now become a common means of data querying, greatly improving the efficiency and accuracy of extracting effective information from massive amounts of data.
[0004] However, current methods for querying data using artificial intelligence often result in irrelevant answers, which seriously hinders the widespread production and application of artificial intelligence in the field of data querying. Summary of the Invention
[0005] This disclosure provides a data query method, apparatus, device, and medium to improve the accuracy of data query results output for query questions.
[0006] According to one aspect of this disclosure, a data query method is provided, comprising:
[0007] Obtain the query question, and obtain the keyword sequence of the domain in which the query question belongs;
[0008] Based on the keyword sequence, the pointer positions of the keywords in the query question are determined. Based on the pointer positions of the keywords, word truncation and verification are performed on the query question to determine the word segmentation corresponding to the query question.
[0009] Based on the word segmentation corresponding to the query question, a search is performed in a pre-constructed multi-layered nested knowledge base to obtain retrieval knowledge. The multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network. The retrieval knowledge includes at least one of business knowledge and query statement knowledge.
[0010] The retrieved knowledge is processed using a pre-built processing model to generate data query statements;
[0011] The data query is performed based on the data query statement to obtain the data query results.
[0012] According to another aspect of this disclosure, a data query apparatus is provided, the apparatus comprising:
[0013] The query question acquisition module is used to acquire query questions and to acquire a keyword sequence in the domain of the query question.
[0014] The word segmentation determination module is used to determine the pointer position of the keyword in the query question based on the keyword sequence, perform word truncation and verification on the query question based on the pointer position of the keyword, and determine the word segmentation corresponding to the query question;
[0015] The knowledge retrieval module is used to retrieve knowledge from a pre-built multi-layered nested knowledge base based on the word segmentation corresponding to the query question. The multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network. The retrieval knowledge includes at least one of business knowledge and query statement knowledge.
[0016] The data query statement generation module is used to process the retrieved knowledge through a pre-built processing model to generate data query statements;
[0017] The data query result acquisition module is used to perform data queries based on the data query statement and obtain data query results.
[0018] According to another aspect of this disclosure, an electronic device is provided, the electronic device comprising:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data query method described in any embodiment of this disclosure.
[0022] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions for causing a processor to execute and implement the data query method described in any embodiment of this disclosure.
[0023] According to another aspect of this disclosure, a computer program product is provided, which, when executed by a processor, implements the data query method as described in any of the embodiments of this disclosure.
[0024] This embodiment of the disclosure obtains a query question and a keyword sequence in the domain of the query question; determines the pointer position of the keywords in the query question based on the keyword sequence; performs word segmentation and verification on the query question based on the pointer position of the keywords to determine the word segmentation corresponding to the query question; retrieves the retrieved knowledge based on the word segmentation corresponding to the query question in a pre-built multi-layered nested knowledge base, wherein the multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network; the retrieved knowledge includes at least one of business knowledge and query statement knowledge; processes the retrieved knowledge through a pre-built processing model to generate a data query statement; performs a data query based on the data query statement to obtain data query results. By determining the word segmentation corresponding to the query question and using a pre-built multi-layered nested knowledge base, the data query results output for the query question are more accurate, improving the accuracy of the data query.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this disclosure and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart of a data query method according to an embodiment of this disclosure;
[0028] Figure 2 This is a schematic diagram of the keyword sequence in an embodiment of this disclosure;
[0029] Figure 3 This is a schematic diagram of the frequency marker list in the embodiments of this disclosure;
[0030] Figure 4 This is a schematic diagram of the keyword pointer list in an embodiment of this disclosure;
[0031] Figure 5 This is a schematic diagram of the entity association network of the query statement in the embodiments of this disclosure;
[0032] Figure 6 This is a schematic diagram of the business entity association network in the embodiments of this disclosure;
[0033] Figure 7This is a schematic diagram of a multi-layered nested knowledge base in an embodiment of this disclosure;
[0034] Figure 8 This is a flowchart of a data query method according to an embodiment of this disclosure;
[0035] Figure 9 This is a schematic diagram of the processing model framework in the embodiments of this disclosure;
[0036] Figure 10 This is a schematic diagram of the structure of a data query device according to an embodiment of this disclosure;
[0037] Figure 11 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0038] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0039] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0040] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0041] Figure 1This is a flowchart illustrating a data query method provided in an embodiment of the present disclosure. This embodiment is applicable to situations where natural language is input into a processing model to obtain a data query statement and then a data query result. This method can be executed by the data query device in this embodiment, which can be implemented in software and / or hardware. This device can be integrated into electronic devices such as computer equipment, servers, mobile terminals, or processors. Figure 1 As shown, the method specifically includes the following steps:
[0042] S110, obtain the query question, and obtain the keyword sequence of the domain in which the query question is located.
[0043] In this embodiment, the query question can be specifically understood as a question posed by the user in natural language that requires data querying to obtain an answer. Its core purpose is to obtain specific data from the data source associated with the underlying knowledge base. The keyword sequence can be specifically understood as a sequence that is strongly related to the domain of the query question and can accurately summarize the core business entities of that domain. It is the foundation for subsequently locating key information in the query question and achieving accurate word segmentation and knowledge base retrieval.
[0044] Specifically, the process involves acquiring user queries input using natural language and obtaining keyword sequences related to the domain of those queries. The construction of a keyword sequence for any domain includes: acquiring relevant documents for that domain; merging the summaries of these documents into a general summary text; reading the summary document into summarization software; segmenting the content at the "character" level according to a segmentation pattern; calculating the frequency of occurrence at the character level; selecting a basic dictionary; identifying modal particles and third-person pronouns to create a "useless" word list, which can be manually adjusted; and removing "useless" words from the word frequency list, defining the words in the top 50% by frequency as the keyword sequence for that domain. For example, a user raises a question related to computer data and randomly downloads m relevant documents from CNKI (China National Knowledge Infrastructure) for each of the n disciplines: "Electric Power Industry," "Finance," "Internet Technology," "Secondary Education," and "Computer Software and Applications." These documents could be core journals or papers. The user then merges the abstracts of the n×m documents to form a general abstract text. The abstract document is then read into abstract processing software, which segments the content by "character" granularity and counts the frequency of occurrence of each character. An internet thesaurus is selected as the underlying dictionary, and modal particles and third-person pronouns are identified to form a "useless" word list. This list of useless words can be manually adjusted. The "useless" words are then removed from the word frequency list, and the words in the top 50% of the word frequency ranking are defined as the keyword sequence for that field.
[0045] S120, determine the pointer position of the keyword in the query question based on the keyword sequence, perform word truncation and verification on the query question based on the pointer position of the keyword, and determine the word segmentation corresponding to the query question.
[0046] In this embodiment, the pointer position of the keyword can be specifically understood as the specific character position coordinates of each character in the keyword sequence within the natural language text of the query question. Typically, the first character of the query question text is used as the starting point, and positions are marked sequentially with numerical codes. The pointer position of the keyword is the basis for subsequent precise word segmentation of the query question. The word segmentation corresponding to the query question can be specifically understood as, based on the pointer position of the keyword, after word segmentation and dictionary verification of the query question, the resulting word segmentation accurately matches the domain semantics and query requirements. It can completely preserve the core information of the domain and is the key to connecting natural language queries with underlying knowledge base retrieval.
[0047] Specifically, the pointer positions of keywords in the query question are determined based on the keyword sequence. Then, based on these pointer positions, word segmentation and verification are performed to determine the corresponding word segmentation. This method accurately pinpoints the core semantics of the query question, avoiding the segmentation errors caused by polysemy and ambiguity in traditional word segmentation. This results in word segmentation results that better match the user's true query intent, laying a solid foundation for subsequent keyword matching and semantic understanding.
[0048] Optionally, based on the above embodiments, determining the pointer positions of keywords in the query question based on the keyword sequence includes: matching the text in the query question with the keywords in the keyword sequence to determine the keywords included in the query question; forming a keyword pointer list by the keyword positions in the query question, wherein the keyword pointer list includes the position information of each keyword in the query question in the query question.
[0049] In this embodiment, the keywords included in the query question can be specifically understood as matching the text in the query question with the keywords in the keyword sequence. The keywords in the keyword sequence appearing in the query question are the core carriers for subsequently locating query needs and extracting key semantics. The keyword pointer list can be specifically understood as a structured list formed by organizing the keywords matched in the query question and their specific location information in the query question text. This list not only records which characters are keywords, but also clarifies the position of each keyword in the query question, providing an intuitive and operable position index for subsequent accurate keyword location and word extraction.
[0050] Specifically, match the words in the query problem with the keywords in the keyword sequence to determine the keywords in the keyword sequence that appear in the query problem. Combine the occurrence frequencies of the keywords that appear in the query problem to form a keyword pointer list for the keyword positions in the query problem. The keyword pointer list includes the position information of each keyword in the query problem in the query problem. Exemplarily, the obtained keyword sequence in the field is "dui", "zhong", "wei", "cong", "deng", "shang", "yi", "xia", "xue", "dao", and "xi". The schematic diagram of the keyword sequence is as shown in Figure 2 Figure. Suppose there is existing text "very like the duicuo book issued from the middle class of the school", create a frequency marking list corresponding to the keywords based on the keyword sequence, and create a keyword pointer list based on the frequency marking list. The frequency marking list is to respectively count the frequencies of its appearance in the text in the order of the keyword sequence. It can be seen that "dui" appears 2 times in the text, "zhong" appears 1 time, "wei" appears 0 times, "cong" appears 1 time, "deng", "shang", and "yi" each appear 0 times, and "xia", "xue", "dao", and "xi" each appear 1 time. The schematic diagram of the frequency marking list is as shown in Figure 3 Figure. The keyword pointer list is to respectively mark the character position serial numbers in the text pointed to by the pointers in the order of the pointers. Taking the keyword "dui" as an example, it has two pointer positions, respectively pointing to the 1st and 10th character positions in the text. The schematic diagram of the keyword pointer list is as shown in Figure 4 Figure.
[0051] Optionally, based on the pointer positions of the keywords on the basis of the above embodiments, perform word truncation and verification on the query problem to determine the word segmentation corresponding to the query problem, including: determining the word truncation range based on the pointer positions of the keywords; performing word truncation within the word truncation range by using a sliding window with at least one window length to obtain at least one to-be-verified word; performing word verification on the to-be-verified word, and determining the word segmentation corresponding to the keyword based on the words that pass the verification; determining the word segmentation corresponding to the query problem based on the word segmentation corresponding to the keyword.
[0052] In this embodiment, the word truncation range can be specifically understood as a character range defined by referencing the pointer position of the keyword and combining it with the text of the query question, used to extract complete and valid words. This range is not limited to the position of the keyword itself, but is appropriately extended to the non-keyword characters adjacent to the keyword, avoiding semantic breaks caused by only truncating keyword fragments. At the same time, the extension boundary is limited by domain common sense and text logic to prevent the truncation range from being too large and introducing irrelevant information. The sliding window can be specifically understood as setting fixed length intervals of different lengths within the defined word truncation range. This interval moves position by position from the start position to the end position of the truncation range, like a sliding window. For example, when the window length is 3, words with three characters can be truncated. Its core function is to cover possible collocations of keywords through multi-length combination truncation methods, providing a more comprehensive candidate range for subsequent verification and screening of accurate word segmentation. The words to be verified can be specifically understood as the character combinations obtained after traversing and truncating the word truncation range through the sliding window, which have not been confirmed by dictionary matching. These character combinations are potential effective word segmentation candidates. They may contain complete semantic units that meet the domain requirements, or they may contain incomplete semantics, be irrelevant to the domain, or be incorrectly expressed. They need to be screened through subsequent verification steps before it can be determined whether they are used as the final word segmentation.
[0053] Specifically, initializing the sliding window involves setting its starting position S and length L. The initial window length is... The i-th key pointer points to the position. The longest word length M in the underlying dictionary is used as the maximum window length. Get the position of the i-th key pointer from the list of key pointers. Set the window length to L, the number of window slides to n, and the window slide step size to... .
[0054] If the window's starting position Then according to the start and end positions Word extraction is performed by searching for the word in the underlying dictionary. If the word can be found in the underlying dictionary, it is considered an appropriate word, and it is then labeled as a triple {i, , L}, where the first i represents the order of the keyword, the second represents the starting position of the sliding window, and the third L represents the length of the sliding window.
[0055] If the window's starting position ,by To adjust the sliding step, slide the sliding window in the opposite direction, according to the start and end positions. Perform word extraction, if or (Empty character) stops the sliding and stops the word extraction record. For each word extraction, a search is performed in the underlying dictionary. If the word can be found in the underlying dictionary, it is considered an appropriate word and is labeled as a triple {i, , L}. Wherein, window length ,in As the window length increases, if Continue to judge S and Based on the positional relationships, words can be extracted.
[0056] If the window's starting position ,and According to the start and end positions Word extraction is performed by searching for the word in the underlying dictionary. If the word can be found, it is considered an appropriate word, and it is then labeled as a triple {i, ...} , L}; if ,and Starting from position (S, S+L-1), the window is first translated in a positive direction, according to the start and end positions. Perform word extraction, if or (Empty character) will stop the slide and stop word extraction records. Next, by position... Starting from the beginning position, shift the sliding window negatively, according to the start and end positions. Perform word extraction, if or (Empty character) will stop the slide and stop word extraction records; if ,and Then based on position Starting from the beginning position, first move the window forward, according to the start and end positions. Perform word extraction, if or (An empty character) will stop the sliding motion and halt word extraction. The window length... , As the window length increases, if Continue to judge S and The positional relationship is used to extract words. Let i = i + 1, and continue the above process. If i > the total number of keywords, then stop the loop iteration.
[0057] Optionally, based on the above embodiments, determining the word segmentation corresponding to the keyword based on the successfully verified words includes: when there are multiple successfully verified words corresponding to any keyword, the word with the longest length is taken as the word segmentation corresponding to the keyword; wherein, the word segmentation corresponding to the query question includes the word segmentation corresponding to the keyword and the word segmentation corresponding to non-keywords, and the word segmentation corresponding to non-keywords includes the non-keywords themselves.
[0058] In this embodiment, the longest word can be understood as the word with the most characters among multiple successfully verified words. Such words typically more comprehensively carry the domain information in the query, avoiding semantic fragmentation caused by selecting short words and ensuring a higher degree of matching between the final word segmentation and the query requirements. Keyword-related word segmentation can be understood as the longest word selected from multiple successfully verified words after sliding window truncation and dictionary verification for each matched keyword in the query. This is a core semantic unit that completely covers the keyword and its surrounding related semantics and conforms to domain specifications. Non-keyword-related word segmentation can be understood as words in the query that do not form a single word from the keyword sequence and can completely cover the meaning of the query.
[0059] Specifically, when multiple words successfully validated for each keyword are used, their lengths are compared, and the longest word is selected as the word segment for that keyword at that pointer position. Using the longest word as the word segment for the keyword maximizes the integration of related semantics around the keyword, avoiding query intent deviations caused by overly fine semantic splitting, and making the word segmentation more closely aligned with the domain's business expression. The remaining characters in the query text that are not matched from the keyword sequence are considered non-keywords, and the word segmentation corresponding to the keyword is the non-keyword itself. Non-keywords are segmented directly using themselves, preserving the complete context of the query question and providing a foundation for logical connections when generating subsequent data query statements.
[0060] S130, based on the word segmentation corresponding to the query question, a search is performed in a pre-constructed multi-layered nested knowledge base to obtain search knowledge, wherein the multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network; the search knowledge includes at least one of business knowledge and query statement knowledge.
[0061] In this embodiment, the pre-built multi-layered nested knowledge base can be specifically understood as a pre-built knowledge base formed by nesting business entity association networks and query statement entity association networks through multi-layered relationships. The business entity association network can be specifically understood as a network built based on the relationships between entities, using entities from the core business scenarios of the domain as base nodes. The query statement entity association network can be specifically understood as a network built based on the relationships between entities, using entities extracted from historically valid query statements as base entity nodes. Essentially, it utilizes historical query experience to reduce retrieval costs and improve accuracy. Retrieval knowledge can be specifically understood as knowledge directly related to the query requirements, extracted from the business entity association network and query statement entity association network after retrieving the corresponding word segmentation from the multi-layered nested knowledge base. Its core purpose is to provide business logic support and query experience reference for subsequent generation of data query statements and acquisition of target data. Specifically, it includes at least one type of knowledge: business knowledge and query statement knowledge. Business knowledge is retrieval knowledge from the business entity association network, and query statement knowledge is retrieval knowledge from the query statement entity association network.
[0062] Specifically, based on the word segmentation corresponding to the query question, a search is performed in a pre-built multi-layered nested knowledge base that includes both business entity association networks and query statement entity association networks to obtain retrieval knowledge. The retrieval knowledge includes at least one of business knowledge and query statement knowledge. The multi-layered nested knowledge base helps users retrieve query statement knowledge based on prior experience, improving overall retrieval efficiency.
[0063] Optionally, based on the above embodiments, the retrieval knowledge is obtained by searching a pre-built multi-layered nested knowledge base based on the word segmentation corresponding to the query question, including: matching the word segmentation corresponding to the query question in historical query statements; if a similar query statement corresponding to the query question is matched, searching the business entity association network based on the first entity and first entity relationship in the query statement entity association network of the similar query statement to obtain a second entity and second entity relationship, and using the first entity, the first entity relationship, the second entity and the second entity relationship as the retrieval knowledge; if no similar query statement corresponding to the query question is matched, searching the multi-layered nested knowledge base based on the word segmentation corresponding to the query question to obtain the retrieval knowledge.
[0064] In this embodiment, historical query statements can be specifically understood as query statements previously initiated by users or the system in past data query scenarios that are related to the current business domain. These statements form the core data foundation for quickly matching similar queries and shortening the retrieval path. The first entity can be specifically understood as an entity related to the matched similar historical query statements and existing in the query statement entity association network. The first entity relationship can be specifically understood as the relationship between entities of similar historical query statements extracted from the query statement entity association network. The second entity can be specifically understood as an entity retrieved in the business entity association network based on the first entity and first entity relationship of similar query statements in the query statement entity association network. The second entity relationship can be specifically understood as the relationship between entities retrieved in the business entity association network based on the first entity and first entity relationship of similar query statements in the query statement entity association network.
[0065] Specifically, based on the word segmentation corresponding to the query question, historical query statements are matched to find similar query statements with the same word segmentation as the current query question. When a similar query statement is found, the first entity and its relationship in the entity association network of the query statement are obtained. Based on the first entity and its relationship, the second entity and its relationship are searched throughout the entire business entity association network. These first entity, its relationship, and its relationship are then used as retrieval knowledge. If a similar query statement is found for each query question, the corresponding word segmentation is directly searched from the entire multi-layered nested knowledge base to obtain retrieval knowledge. Prioritizing matching similar historical queries before extending the search not only utilizes the relationships between entities, significantly shortening the retrieval path, but also relies on the multi-layered nested knowledge base to ensure that the retrieval knowledge aligns with actual business needs, reducing the risk of bias.
[0066] Optionally, based on the above embodiments, the construction process of the multi-layer nested knowledge base includes: obtaining historical query statements, extracting entities from the historical query statements, forming multiple triples based on the extracted entities, and constructing the query statement entity association network based on the multiple triples; obtaining entities of business scenarios and the relationships between entities, and constructing the business entity association network based on the entities of business scenarios and the relationships between entities; and embedding the query statement entity association network into the business entity association network to obtain a multi-layer nested knowledge base.
[0067] In this embodiment, the entities in the business scenario can be specifically understood as entities that support business operations from actual business scenarios, have clear business meanings and attributes, and are the basic nodes for constructing the business entity association network. They directly correspond to the business side's definition of data, usage logic, and relationships.
[0068] Specifically, historical query statements are retrieved, and entity extraction is performed on these statements. Multiple triples are then formed based on the extracted entities, and a query statement entity association network is constructed. For example, the historical query statement retrieved is: "selectt.stat_month,user_id,count( as cnt_rwx,
[0069] From (select t.stat_month, t.user_id, t.prodid, t.apply_date,t.start_date, t.end_date from aidata1.view_selfprod_and_prod_pack T) T,
[0070] (select distinct stat_month, cast(t.commodity_id as string)commodity_id, t.commodity_name, t.fee / 1000 Fee from aidata1.ads_dtl_boss_commodity_mon t where commodity_id in ('2400000557', '2400000558', '2000011705')) T2,
[0071] where T.prodid=T2.commodity_id,
[0072] and T.stat_month=T2.stat_month,
[0073] and T.end_date>T.last_day,
[0074] Group by t.stat_month,user_id".
[0075] Based on historical query statements, entity extraction can be performed: Table entities include: `aidata1.view_selfprod_and_prod_pack`; `aidata1.ads_dtl_boss_commodity_mon`; Dimension entities include: `stat_month`, `user_id`; Metric entities include: `prodid`, `apply_date`, `start_date`, `end_date`, `commodity_id`, `commodity_name`, `Fee`. The relationships between these entities are only hierarchical, forming multiple triple arrays, for example: `{stat_month, hierarchical, aidata1.view_selfprod_and_prod_pack}; {user_id, hierarchical, aidata1.view_selfprod_and_prod_pack}`. Based on these triples, a query statement entity association network is constructed. If the task name in this example is: Monthly User Product Order Total Amount, such as... Figure 5 The example query statement shows the entity association network.
[0076] Obtain the entities and relationships between them within the business scenario, and construct a business entity relationship network based on these entities and their relationships. For example, the entities and relationships within the business scenario could be: {Revenue (business term), Synonym, CHBN Revenue (business term)}, {das.app_ld_index_view_day (database table), Attribution, Total Bill Revenue (metric)}, etc. The relationships between entities mainly include two types: synonym relationships and attribution relationships. Define the business entity relationship. as follows:
[0077]
[0078] in, , These are the weight matrices for the fully interconnected entity network, where the diagonal matrix elements are... as well as They are 0 respectively. and Representing entities respectively and entity The synonym and attribution relationships between them are defined as follows: if the value is 1, it indicates that a relationship exists; otherwise, it does not exist. and The weights are set in a 5:3 ratio, for example. and These values can be set to 1 and 0.6 respectively. Based on this weighting ratio, for any two business entities, the specific relationship category between them can be calculated. A business entity relationship network can be constructed based on the formula shown below:
[0079]
[0080] A diagram illustrating the network of business entity relationships is shown below. Figure 6 As shown in the diagram, the query entity association network is embedded into the business entity association network to obtain a multi-layered nested knowledge base. An exemplary diagram of the multi-layered nested knowledge base is shown below. Figure 7 As shown.
[0081] S140, The retrieved knowledge is processed through a pre-built processing model to generate a data query statement.
[0082] In this embodiment, the pre-built processing model can be specifically understood as a pre-built model that can convert natural language into data query statements. The data query statement can be specifically understood as a structured query instruction generated by the processing model based on retrieval knowledge, which can be directly executed in the business database. Its core function is to transform the natural language requirements into machine-recognizable and database-responsive operation statements, ultimately achieving the extraction of target business data.
[0083] Specifically, the retrieved knowledge is processed through a pre-built processing model to generate data query statements. This pre-built model transforms entities and relationships within the retrieved knowledge into syntactically correct, logically rigorous, and executable data query statements, avoiding data retrieval errors caused by unfamiliarity with business rules or grammatical mistakes during manual conversion, thus ensuring the accuracy of data extraction.
[0084] S150, perform a data query based on the data query statement to obtain the data query result.
[0085] In this embodiment, the data query result can be specifically understood as the result returned after executing the generated data query statement in the corresponding business database, which directly matches the original query requirement. This data conforms to the business rules defined in the retrieval knowledge and corresponds to the query question.
[0086] Specifically, data is retrieved from the corresponding database based on the query statement to obtain the results. This method allows for rapid data location and result retrieval, significantly reducing the time cost of data extraction.
[0087] This embodiment of the disclosure obtains a query question and a keyword sequence in the domain of the query question; determines the pointer position of the keywords in the query question based on the keyword sequence; performs word segmentation and verification on the query question based on the pointer position of the keywords to determine the word segmentation corresponding to the query question; retrieves the retrieved knowledge based on the word segmentation corresponding to the query question in a pre-built multi-layered nested knowledge base, wherein the multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network; the retrieved knowledge includes at least one of business knowledge and query statement knowledge; processes the retrieved knowledge through a pre-built processing model to generate a data query statement; performs a data query based on the data query statement to obtain data query results. By determining the word segmentation corresponding to the query question and using a pre-built multi-layered nested knowledge base, the data query results output for the query question are more accurate, improving the accuracy of the data query.
[0088] Figure 8 This is a flowchart of a data query method provided in an embodiment of the present disclosure. This embodiment also provides the training process of the processing model, such as... Figure 8 As shown, the method specifically includes the following steps:
[0089] S211, Obtain the pre-trained processing model, wherein the pre-trained processing model corresponds to the initial model parameters.
[0090] In this embodiment, the pre-trained processing model can be understood as a model based on large-scale general data, which has been trained in the early stages and possesses the basic ability to understand business requirements and map data logic. The pre-trained processing model can be an LLM (Large Language Model), whose core function is to provide a basic framework for subsequent targeted training. In its initial state, it already includes basic capabilities such as general language understanding, entity recognition, and query statement generation. The pre-trained processing model has initial model parameters, which are the starting point for subsequent fine-tuning and optimization of the model based on specific retrieval knowledge and business rules. This can significantly reduce the training cost and cycle of the specific task of generating data query statements from retrieval knowledge. Specifically, the initial model parameters can be understood as the model parameters of the pre-trained processing model after completing training on large-scale general data.
[0091] Specifically, a pre-trained processing model is obtained, which corresponds to the initial model parameters. The initial model parameters provide a high-quality starting point for subsequent task adaptation. Fine-tuning based on these parameters can lead to faster convergence to the optimal solution for the task, reduce the number of training iterations, and improve model deployment efficiency.
[0092] S212, Obtain sample data, which includes sample questions and query statement tags.
[0093] In this embodiment, the sample data can be specifically understood as a dataset constructed for training the processing model, including sample questions and query statement labels. The sample questions can be understood as natural language questions raised by users in actual business scenarios; the query statement labels are the standard data query statements that can be directly executed in the business database corresponding to the sample questions. For example, one sample question in the sample data could be "How many models have been built in each city?", and the corresponding query statement label could be "SELECT T.AREA_ID, COUNT (DISTINCT T.MODEL_ID) CNT_MODEL FROM N_YANGJ_0705_AB_INFO_02_1 T GROUP BY T.AREA_ID;".
[0094] Specifically, sample data needs to be obtained, which must clearly contain two types of core information: sample questions and query statement tags. Sample questions are natural language queries actually initiated by users, while query statement tags are the standard query statements corresponding to the questions, providing high-quality data support for subsequent model training.
[0095] S213, perform semantic analysis on the sample question to obtain multiple recombined questions of the sample question, and form parallel sample data based on the multiple recombined questions and the query statement tags.
[0096] In this embodiment, the multiple recombined problems of the sample problem can be understood as multiple problems obtained by semantically preserving rewriting the original sample problem, which have the same core requirement as the original problem but different expressions. These recombined problems do not change the semantics of the original sample problem; they are generated only by adjusting sentence structure, replacing synonyms, and adjusting word order. The purpose is to enrich the diversity of samples, allowing the processing model to learn that different expressions of the same requirement correspond to the same query statement label. Parallel sample data can be understood as using the query statement label corresponding to the original sample problem as the core, and pairing this label with multiple recombined problems generated based on the original sample problem to form multiple sets of sample data. This avoids the model misjudging the requirement due to differences in expression, and at the same time, the training efficiency is improved by using multiple sets of parallel samples, strengthening the model's ability to adapt to diverse expressions.
[0097] Specifically, semantic analysis is performed on the sample questions, and the sample questions are recombined while ensuring that the semantics remain unchanged, resulting in multiple recombined questions. Parallel sample data is formed based on multiple recombined questions and query statement tags, which can expand the amount of sample data, solve the problem of scarce labeled data, and cover the diverse expression habits of users, allowing the model to learn different expression methods and improve its adaptability to flexible questioning in real-world scenarios.
[0098] S214, Based on the parallel sample data, the parameters of multiple low-rank adaptive models are updated to obtain the first parameter adjustment matrix of the multiple low-rank adaptive models.
[0099] In this embodiment, multiple low-rank adaptive models can be understood as an infrastructure based on a pre-trained processing model, using multiple models constructed with a fine-tuning method for parallel processing of parallel sample data. The first parameter adjustment matrix can be understood as a set of low-rank matrices learned by each of the multiple low-rank adaptive models after receiving parallel sample data for training, used for fine-tuning the pre-trained processing model.
[0100] Specifically, the pre-trained processing model is configured with a gating network. This gating network, based on the number of low-rank adaptation models, randomly shuffles and reassembles the sample questions into n parts, which are then transmitted to multiple low-rank adaptation models. Each low-rank adaptation model updates its LLM parameters according to the data transmitted from the gating system, resulting in the first parameter adjustment matrices for multiple low-rank adaptation models. This set of matrices constitutes the first parameter adjustment matrix, and the dimensions of the first parameter adjustment matrices output by all low-rank adaptation models remain consistent, meeting the requirements for subsequent processes.
[0101] S215, perform matrix fusion on the first parameter adjustment matrices of the multiple low-rank adaptation models to obtain a fusion matrix, and decompose the fusion matrix to satisfy the set rank condition to obtain the decomposed second parameter adjustment matrix.
[0102] In this embodiment, the fusion matrix can be specifically understood as a single matrix formed by integrating the first parameter adjustment matrices obtained from training multiple low-rank adaptive models. Its core function is to summarize the training results of multiple sets of parallel sample data. The fused matrix can integrate the association patterns between different expression forms and query statements, avoiding the limitations of single model training. The second parameter adjustment matrix can be specifically understood as a low-rank matrix obtained by performing low-rank decomposition on the fusion matrix. It enables lightweight fine-tuning of the pre-trained model, which not only continues the low-resource consumption advantage of low-rank fine-tuning technology, but also ensures that the model is adaptable to diverse scenarios.
[0103] Specifically, the first parameter adjustment matrices of multiple low-rank adaptive models are fused to obtain a fusion matrix. This fusion matrix is then decomposed to satisfy a set rank condition, yielding the decomposed second parameter adjustment matrix. Optionally, the set rank condition can be that the rank of the second parameter adjustment matrix is much smaller than the rank of the initial model parameters. For example, the first parameter adjustment matrix is... They are then merged to generate a fusion matrix. Its specific representation is as follows:
[0104]
[0105] For fusion matrix The parameter is decomposed into several matrices with smaller parameters, namely the second adjustment parameter matrix. , making It satisfies the following formula:
[0106]
[0107] in, These are the initial model parameters. The rank average is much smaller than Rank.
[0108] S216, Update the initial model parameters of the pre-trained processing model based on the second parameter adjustment matrix to obtain the trained processing model.
[0109] In this embodiment, the trained processing model can be understood as a model obtained by training the initial model parameters of the pre-trained processing model through the second parameter adjustment matrix. It has the advantages of low-rank fine-tuning and integrates the training rules of multiple sets of parallel samples, and can accurately output the data query statement corresponding to the query question.
[0110] Specifically, the initial model parameters of the pre-trained processing model are updated based on the second parameter adjustment matrix to obtain the trained processing model. For example, if the initial model parameters of a certain layer of the pre-trained processing model are... The second parameter adjustment matrix is The final parameters of the trained processing model are calculated as follows:
[0111]
[0112] S217, obtain the query question, and obtain the keyword sequence of the domain in which the query question is located.
[0113] S218, Based on the keyword sequence, determine the pointer position of the keyword in the query question, and perform word truncation and verification on the query question based on the pointer position of the keyword to determine the word segmentation corresponding to the query question.
[0114] S219, based on the word segmentation corresponding to the query question, a search is performed in a pre-constructed multi-layered nested knowledge base to obtain search knowledge, wherein the multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network; the search knowledge includes at least one of business knowledge and query statement knowledge.
[0115] S220, The retrieved knowledge is processed through a pre-built processing model to generate a data query statement.
[0116] S221, Perform a data query based on the data query statement to obtain the data query result.
[0117] This embodiment of the disclosure obtains a pre-trained processing model, which corresponds to initial model parameters; obtains sample data, which includes sample questions and query statement tags; performs semantic analysis on the sample questions to obtain multiple recombined questions, and forms parallel sample data based on the multiple recombined questions and the query statement tags; updates the parameters of multiple low-rank adaptive models based on the parallel sample data to obtain a first parameter adjustment matrix of the multiple low-rank adaptive models; performs matrix fusion on the first parameter adjustment matrices of the multiple low-rank adaptive models to obtain a fusion matrix, and decomposes the fusion matrix to satisfy a set rank condition to obtain a decomposed second parameter adjustment matrix; and updates the initial model parameters of the pre-trained processing model based on the second parameter adjustment matrix to obtain a trained processing model. The process involves: obtaining a query question and a keyword sequence in the domain of the query question; determining the pointer positions of keywords in the query question based on the keyword sequence; performing word segmentation and verification on the query question based on the pointer positions of the keywords to determine the word segmentation corresponding to the query question; retrieving retrieval knowledge from a pre-built multi-layered nested knowledge base based on the word segmentation corresponding to the query question, wherein the multi-layered nested knowledge base includes nested business entity association networks and query statement entity association networks; the retrieval knowledge includes at least one of business knowledge and query statement knowledge; processing the retrieval knowledge through a pre-built processing model to generate a data query statement; performing a data query based on the data query statement to obtain data query results; and updating the initial model parameters through a low-rank fine-tuning method, which not only reduces the demand for computing resources but also maintains the performance of the processing model.
[0118] Based on the above embodiments, an optional example is provided, which can be used in scenarios where users input query questions into a processing model and obtain query data.
[0119] First, we construct the processing model used to query the data. The framework of this processing model is as follows: Figure 9 As shown, the original initial model parameters of the LLM model are first locked and saved, denoted as . Set up a gated network that fine-tunes the number of nodes in parallel based on a low-rank adaptive model. X random sample data points are shuffled and reorganized into n parallel sample data points, which are then input into the low-rank adaptive model. Each low-rank adaptive model will then use the data transmitted from the gating system... The LLM model is fine-tuned separately, assuming the generated first parameter adjustment matrices are respectively The dimensions of each fine-tuning matrix are consistent; the first parameter adjustment matrix... Perform fusion to generate a fusion matrix. ; for the fusion matrix Decompose it into several matrices with smaller parameters. , making It satisfies the following formula:
[0120]
[0121] in The rank average is much smaller than If the rank is zero, then the final output will be:
[0122]
[0123] according to The initial model parameters of the pre-trained processing model are updated to obtain the trained processing model. A vector fusion strategy is introduced to improve the model resource consumption problem. After introducing the vector fusion strategy, three new vectors are introduced and soft merging is performed. They restructured the key and value activations in self-attention, where each multi-head attention block only requires vector and router updates, thus greatly reducing resource consumption.
[0124] After the processing model is built, the user's input query is obtained. For example, if the user's query is: "Statistically count the number of users who ordered the product each month for the past four months," then the query question and its corresponding keyword sequence are obtained. The pointer positions of the keywords in the query question are determined based on the keyword sequence. Word segmentation and validation are performed based on these pointer positions to determine the corresponding word segmentation. The word segmentation is then used to retrieve information from a pre-built multi-layered nested knowledge base. This retrieved knowledge is then processed by the pre-built processing model to generate a data query statement. Finally, a data query is performed based on this statement to obtain the data query results.
[0125] Figure 10 This is a schematic diagram of a data query device provided in an embodiment of the present disclosure. This embodiment is applicable to situations where natural language is input into a processing model to obtain a data query statement and data query results. The device can be implemented using software and / or hardware, and can be integrated into any device that provides data query functionality, such as… Figure 10 As shown, the data query device specifically includes: a query question acquisition module 310, a word segmentation determination module 320, a retrieval knowledge acquisition module 330, a data query statement generation module 340, and a data query result acquisition module 350.
[0126] The query question acquisition module 310 is used to acquire query questions and to acquire a keyword sequence in the domain of the query question.
[0127] The word segmentation determination module 320 is used to determine the pointer position of the keyword in the query question based on the keyword sequence, perform word truncation and verification on the query question based on the pointer position of the keyword, and determine the word segmentation corresponding to the query question;
[0128] The knowledge retrieval module 330 is used to retrieve knowledge from a pre-built multi-layered nested knowledge base based on the word segmentation corresponding to the query question. The multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network. The retrieval knowledge includes at least one of business knowledge and query statement knowledge.
[0129] The data query statement generation module 340 is used to process the retrieved knowledge through a pre-built processing model to generate a data query statement;
[0130] The data query result acquisition module 350 is used to perform data queries based on the data query statement and obtain data query results.
[0131] This embodiment of the disclosure obtains a query question and a keyword sequence in the domain of the query question; determines the pointer position of the keywords in the query question based on the keyword sequence; performs word segmentation and verification on the query question based on the pointer position of the keywords to determine the word segmentation corresponding to the query question; retrieves the retrieved knowledge based on the word segmentation corresponding to the query question in a pre-built multi-layered nested knowledge base, wherein the multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network; the retrieved knowledge includes at least one of business knowledge and query statement knowledge; processes the retrieved knowledge through a pre-built processing model to generate a data query statement; performs a data query based on the data query statement to obtain data query results. By determining the word segmentation corresponding to the query question and using a pre-built multi-layered nested knowledge base, the data query results output for the query question are more accurate, improving the accuracy of the data query.
[0132] Based on the above embodiments, optionally, the word segmentation determination module 320 is used to: match the text in the query question with the keywords in the keyword sequence to determine the keywords included in the query question; and form a keyword pointer list by the positions of the keywords in the query question, wherein the keyword pointer list includes the position information of each keyword in the query question in the query question.
[0133] Based on the above embodiments, optionally, the word segmentation determination module 320 is used to: determine the word truncation range based on the pointer position of the keyword; perform word truncation within the word truncation range based on a sliding window of at least one window length to obtain at least one word to be verified; perform word verification on the word to be verified, and determine the word segmentation corresponding to the keyword based on the successfully verified word; and determine the word segmentation corresponding to the query question based on the word segmentation corresponding to the keyword.
[0134] Based on the above embodiments, optionally, the word segmentation determination module 320 is used to: when there are multiple successfully verified words corresponding to any keyword, take the word with the longest length as the word segment corresponding to the keyword; wherein, the word segmentation corresponding to the query question includes the word segmentation corresponding to the keyword and the word segmentation corresponding to non-keywords, and the word segmentation corresponding to non-keywords includes the non-keywords themselves.
[0135] Based on the above embodiments, optionally, the knowledge retrieval module 330 is configured to: match the word segmentation corresponding to the query question in historical query statements; if a similar query statement corresponding to the query question is matched, based on the first entity and first entity relationship in the query statement entity association network of the similar query statement, search in the business entity association network to obtain a second entity and a second entity relationship, and use the first entity, the first entity relationship, the second entity and the second entity relationship as the retrieval knowledge; if no similar query statement corresponding to the query question is matched, search in the multi-layer nested knowledge base based on the word segmentation corresponding to the query question to obtain the retrieval knowledge.
[0136] Based on the above embodiments, optionally, the knowledge retrieval module 330 is further configured to: acquire historical query statements, extract entities from the historical query statements, form multiple triples based on the extracted entities, and construct the query statement entity association network based on the multiple triples; acquire entities of business scenarios and the relationships between entities, and construct the business entity association network based on the entities of business scenarios and the relationships between entities; and embed the query statement entity association network into the business entity association network to obtain a multi-layer nested knowledge base.
[0137] Optionally, based on the above embodiments, the device further includes a processing model training module, configured to: acquire a pre-trained processing model, the pre-trained processing model corresponding to initial model parameters; acquire sample data, the sample data including sample questions and query statement tags; perform semantic analysis on the sample questions to obtain multiple recombined questions of the sample questions, and form parallel sample data based on the multiple recombined questions and the query statement tags; update the parameters of multiple low-rank adaptive models based on the parallel sample data to obtain a first parameter adjustment matrix of the multiple low-rank adaptive models; perform matrix fusion on the first parameter adjustment matrices of the multiple low-rank adaptive models to obtain a fusion matrix, decompose the fusion matrix to satisfy a set rank condition to obtain a decomposed second parameter adjustment matrix; and update the initial model parameters of the pre-trained processing model based on the second parameter adjustment matrix to obtain a trained processing model.
[0138] The above-described products can perform the methods provided in any embodiment of this disclosure, and have the corresponding functional modules and beneficial effects for performing the methods.
[0139] Figure 11 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0140] like Figure 11 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0141] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0142] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data querying methods.
[0143] In some embodiments, the data query method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data query method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the data query method by any other suitable means (e.g., by means of firmware).
[0144] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0145] Computer programs used to implement the methods of this disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0146] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0148] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0149] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0150] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0151] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the data query method according to any embodiment of this disclosure.
[0152] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0153] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A data query method, characterized in that, include: Obtain the query question, and obtain the keyword sequence of the domain in which the query question belongs; Based on the keyword sequence, the pointer positions of the keywords in the query question are determined. Based on the pointer positions of the keywords, word truncation and verification are performed on the query question to determine the word segmentation corresponding to the query question. Based on the word segmentation corresponding to the query question, a search is performed in a pre-constructed multi-layered nested knowledge base to obtain retrieval knowledge. The multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network. The retrieval knowledge includes at least one of business knowledge and query statement knowledge. The retrieved knowledge is processed using a pre-built processing model to generate data query statements; The data query is performed based on the data query statement to obtain the data query results.
2. The method according to claim 1, characterized in that, Determining the pointer positions of keywords in the query question based on the keyword sequence includes: Match the text in the query question with the keywords in the keyword sequence to determine the keywords included in the query question; The keyword positions in the query question are formed into a keyword pointer list, which includes the position information of each keyword in the query question.
3. The method according to claim 1, characterized in that, Based on the pointer position of the keyword, the query question is trunculated and verified to determine the word segmentation corresponding to the query question, including: The word extraction range is determined based on the pointer position of the keyword; A sliding window of at least one window length is used to truncate words within the word truncation range to obtain at least one word to be verified. The words to be verified are validated, and the word segments corresponding to the keywords are determined based on the successfully validated words; the word segments corresponding to the query questions are determined based on the word segments corresponding to the keywords.
4. The method according to claim 3, characterized in that, Based on successfully verified words, the word segmentation corresponding to the keyword is determined, including: If there are multiple successfully verified words corresponding to any of the aforementioned keywords, the word with the longest length will be used as the word segment corresponding to the keyword. The word segmentation corresponding to the query question includes word segmentation corresponding to the keyword and word segmentation corresponding to non-keyword, and word segmentation corresponding to non-keyword includes the non-keyword itself.
5. The method according to claim 1, characterized in that, Based on the word segmentation corresponding to the query question, a search is performed in a pre-built multi-layered nested knowledge base to obtain the retrieved knowledge, including: The word segmentation corresponding to the query question is matched in the historical query statements; When a similar query statement is matched to the query question, based on the first entity and first entity relationship in the query statement entity association network of the similar query statement, a search is performed in the business entity association network to obtain the second entity and second entity relationship, and the first entity, the first entity relationship, the second entity and the second entity relationship are used as the search knowledge; If no similar query statement is found that corresponds to the query question, the search is performed in the multi-layered nested knowledge base based on the word segmentation corresponding to the query question to obtain the retrieved knowledge.
6. The method according to claim 1 or 5, characterized in that, The construction process of the multi-layered nested knowledge base includes: Obtain historical query statements, extract entities from the historical query statements, form multiple triples based on the extracted entities, and construct the entity association network of the query statements based on the multiple triples. Obtain the entities in the business scenario and the relationships between them, and construct the business entity association network based on the entities in the business scenario and the relationships between them; The entity association network of the query statement is embedded into the business entity association network to obtain a multi-layered nested knowledge base.
7. The method according to claim 1, characterized in that, The training process of the processing model includes: Obtain a pre-trained processing model, wherein the pre-trained processing model corresponds to the initial model parameters; Obtain sample data, which includes sample questions and query statement tags; Semantic analysis is performed on the sample question to obtain multiple recombined questions of the sample question, and parallel sample data is formed based on the multiple recombined questions and the query statement tags; Based on the parallel sample data, the parameters of multiple low-rank adaptive models are updated to obtain the first parameter adjustment matrix of the multiple low-rank adaptive models. The first parameter adjustment matrices of the multiple low-rank adaptive models are fused to obtain a fusion matrix. The fusion matrix is then decomposed to satisfy a set rank condition to obtain the decomposed second parameter adjustment matrix. The initial model parameters of the pre-trained processing model are updated based on the second parameter adjustment matrix to obtain the trained processing model.
8. A data query device, characterized in that, include: The query question acquisition module is used to acquire query questions and to acquire a keyword sequence in the domain of the query question. The word segmentation determination module is used to determine the pointer position of the keyword in the query question based on the keyword sequence, perform word truncation and verification on the query question based on the pointer position of the keyword, and determine the word segmentation corresponding to the query question; The knowledge retrieval module is used to retrieve knowledge from a pre-built multi-layered nested knowledge base based on the word segmentation corresponding to the query question. The multi-layered nested knowledge base includes a nested business entity association network and a query statement entity association network. The retrieval knowledge includes at least one of business knowledge and query statement knowledge. The data query statement generation module is used to process the retrieved knowledge through a pre-built processing model to generate data query statements; The data query result acquisition module is used to perform data queries based on the data query statement and obtain data query results.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data query method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the data query method according to any one of claims 1-7.