Data query method and device, equipment and storage medium
By constructing query hints and leveraging the language understanding capabilities of large models, the problem of low accuracy in traditional data query methods is solved, resulting in more accurate data query results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-14
AI Technical Summary
Traditional data query methods often result in low accuracy due to the complexity of user-input queries and the presence of technical jargon in enterprise data.
By processing natural language query statements, query hints are constructed to instruct the large model to generate query conditions that conform to business constraints. By leveraging the powerful language understanding and result generation capabilities of the large model, target query conditions that match the query intent are determined.
It improves the accuracy of data query results, ensures that query conditions conform to internal enterprise terminology and user intent, reduces ambiguous interpretations, and generates more accurate query results.
Smart Images

Figure CN121858601A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a data query method, apparatus, device and storage medium. Background Technology
[0002] With the rapid development of data center technology, the amount and complexity of data in databases are constantly increasing. The rationality of data query methods is directly related to the operational efficiency and user experience of the enterprise where the database is located.
[0003] Currently, data query methods in related technologies mainly rely on single keyword matching or keyword-based retrieval through general search engines. However, due to the complexity of user-input queries and the presence of numerous technical terms within enterprise data, traditional data search methods often struggle to achieve accurate results, resulting in low accuracy. Summary of the Invention
[0004] The purpose of this application is to provide a data query method, apparatus, device, and storage medium, which aims to improve the accuracy of data query results.
[0005] To achieve the above objectives, this application adopts the following technical solution: In a first aspect, this application provides a data query method, including responding to a received natural language query statement, determining multiple query keywords based on the natural language query statement, and constructing query hint information based on the query keywords; the query hint information is used to represent the business constraint rules associated with each query keyword; inputting the query hint information and the natural language query statement into a large model, and determining target query conditions that match the query intent of the natural language query statement based on the business constraint rules associated with each query keyword through the large model; and performing data query in the database based on the target query conditions.
[0006] Based on the aforementioned technical means, this application processes natural language query statements to construct query suggestion information that describes business constraint rules for query keywords. This information instructs a large model to generate query conditions that conform to the business constraint rules for query keywords and match the query intent of the natural language query statement. The large model possesses powerful language understanding and result generation capabilities, enabling it to more accurately understand the user's query intent through the query suggestion information and output query conditions that conform to the business constraint rules and match the query intent. The query conditions generated by the large model can retrieve more accurate data query results in the database. Therefore, this application can improve the accuracy of data query results.
[0007] In one possible approach, multiple query keywords are determined based on a natural language query statement, including: extracting multiple initial keywords from the natural language query statement; for each initial keyword, determining the similarity between the initial keyword and each preset keyword in a preset keyword set; and selecting preset keywords whose similarity meets preset conditions as the query keywords corresponding to the initial keywords.
[0008] One possible approach is to further include determining the initial keyword as the query keyword when the similarity between the initial keyword and each preset keyword does not meet the preset conditions.
[0009] In one possible approach, data retrieval in the database based on target query conditions includes: inputting the target query conditions and natural language query statements into an intent classification model; determining the word accuracy and intent coverage of the target query conditions through the intent classification model; redetermining query keywords if the word accuracy is less than a first threshold or the intent coverage is less than a second threshold; and performing data retrieval in the database based on the target query conditions if the word accuracy is greater than the first threshold and the intent coverage is greater than the second threshold.
[0010] In one possible approach, the target query conditions include: keywords, business constraint rules, and logical connectors; data querying in the database based on the target query conditions includes: inputting keywords, business constraint rules, and logical connectors into a standardized format conversion program to obtain a query statement in a preset format; and performing data querying in the database based on the query statement in the preset format.
[0011] One possible approach includes: acquiring an industry corpus and an enterprise corpus; fine-tuning the initial large model based on the industry corpus and the enterprise corpus; and determining the initial large model as the large model if the fine-tuned initial large model meets preset conditions.
[0012] In one possible approach, the industry corpus includes at least one of the following: business constraint rules for at least one industry, industry terminology mapping rules for at least one industry, and multimodal data for at least one industry; the enterprise corpus includes at least one of the following: internal terms of at least one enterprise, abbreviations of internal terms, and alternative names of internal terms.
[0013] Secondly, this application provides a data query apparatus, comprising: a determining module, a receiving module, an input module, and a query module; the receiving module is configured to respond to receiving a natural language query statement; the determining module is configured to determine multiple query keywords based on the natural language query statement; the determining module is also configured to construct query prompt information based on the query keywords; the query prompt information is used to represent the business constraint rules associated with each query keyword; the input module is configured to input the query prompt information and the natural language query statement into a large model; the determining module is also configured to determine, through the large model, target query conditions matching the query intent of the natural language query statement based on the business constraint rules associated with each query keyword; and the query module is configured to perform data querying in the database based on the target query conditions.
[0014] Thirdly, this application provides an electronic device including a memory and a processor; the memory and the processor are coupled; the memory is used to store instructions executable by the processor; when the processor executes the instructions, it performs the methods described in the first aspect and any possible implementation thereof.
[0015] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect and any possible implementation thereof.
[0016] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run computer programs or instructions to implement the methods described in the first aspect and any possible implementation thereof.
[0017] Sixthly, this application provides a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods described in the first aspect and any possible implementation thereof.
[0018] The technical problems that the data query device, electronic device, computer storage medium, chip or computer program product can solve and the technical effects it can achieve can be found in the technical problems and effects solved in the first aspect above, and will not be repeated here. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the structure of a data query system provided in an embodiment of this application; Figure 2 A schematic diagram of a data query device provided in an embodiment of this application; Figure 3 A flowchart illustrating a data query method provided in this application embodiment; Figure 4 A flowchart illustrating a method for determining query keywords provided in this application embodiment; Figure 5 A flowchart illustrating another method for determining query keywords provided in this application embodiment; Figure 6 A flowchart illustrating a large model training method provided in this application embodiment; Figure 7 A structural diagram of a data query device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a data query device provided in an embodiment of this application. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0023] In embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, article, or apparatus that includes that element.
[0024] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0025] In the description of this specification, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0026] With the rapid development of data center technology, the amount and complexity of data in databases are constantly increasing. The rationality of data query methods is directly related to the operational efficiency and user experience of the enterprise where the database is located.
[0027] Currently, traditional data query methods mainly rely on single keyword matching or keyword-based retrieval through general search engines. However, due to the complexity of user-input queries and the presence of numerous technical terms in enterprise data, traditional data search methods often struggle to achieve accurate results, resulting in low accuracy.
[0028] In view of this, this application provides a data query method that processes natural language query statements to construct query hints that describe business constraint rules for query keywords. This enables a large model to generate query conditions that conform to the business constraint rules for query keywords and match the query intent of the natural language query statement. The large model possesses powerful language understanding and result generation capabilities, allowing it to more accurately understand the user's query intent through the query hints and output query conditions that conform to the business constraint rules and match the query intent. The query conditions generated by the large model can retrieve relatively accurate data query results in the database. Therefore, this application can improve the accuracy of data query results.
[0029] The data query method provided in this application can be applied to a data query system. Please refer to [link / reference]. Figure 1 The data query system may include: a data terminal 101, a data query device 102, and a data storage device 103; the data terminal 101 and the data query device 102 are connected in communication.
[0030] The data terminal 101 can be an electronic device such as a personal computer (PC), a laptop computer, a mobile device, a tablet computer, or a laptop computer. This application embodiment does not limit the specific form of the electronic device.
[0031] The data terminal 101 is used to display a human-computer interaction interface, which includes human-computer interaction controls. Users can input natural language query statements on the human-computer interaction page, and the data terminal sends the natural language query statements to the data query device for querying.
[0032] The data query device 102 can be an electronic device such as a personal computer (PC), laptop computer, mobile device, tablet computer, or laptop computer. This application embodiment does not limit the specific form of the electronic device. Alternatively, the data query device 102 can also be a server, or a server cluster consisting of multiple servers. In some implementations, the server cluster can be a distributed cluster server. This application embodiment does not impose any limitations in this regard.
[0033] The data query device 102, in response to a received natural language query statement, determines multiple query keywords based on the natural language query statement and constructs query suggestion information based on the query keywords. The query suggestion information represents the business constraint rules associated with each query keyword. The query suggestion information and the natural language query statement are input into a large model, which, based on the business constraint rules associated with each query keyword, determines target query conditions that match the query intent of the natural language query statement. Data is then queried in the database based on the target query conditions.
[0034] The data storage device 103 is used to deploy a database, which can store and manage the data in the database. The data storage device 103 can be an electronic device with local storage capabilities that integrates storage media such as hard disk drives, solid-state drives, or hybrid hard drives, or it can be a stand-alone storage server, such as a network-attached storage device or a storage area network device. This application embodiment does not impose any limitations on this.
[0035] To facilitate the explanation of the solution provided in this application, the following description will focus on some of the software components provided in this application. Please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram of a data query device 20 provided in an embodiment of this application. Figure 2 As shown, the data query device 20 may include a terminal layer 201, an application server layer 202, a large model service layer 203, a data middleware layer 204, and a front-end display layer 205. The terminal layer 201 is a software component in the data terminal 101, while the application server layer 202, the large model service layer 203, the data middleware layer 204, and the front-end display layer 205 are software components in the data query device 102.
[0036] In one possible implementation, terminal layer 201 is used to obtain natural language query statements and send them to application server layer 202.
[0037] In one possible implementation, the application server layer 202, in response to receiving a natural language query statement, determines multiple query keywords based on the natural language query statement, constructs query suggestion information based on the query keywords, and sends the query suggestion information to the large model service layer 203. The query suggestion information is used to represent the business constraint rules associated with each query keyword.
[0038] In one possible implementation, the large model service layer 203 is used to deploy the large model and also to input query prompts and natural language query statements into the large model. Based on the business constraint rules associated with each query keyword, the large model determines the target query conditions that match the query intent of the natural language query statement and sends the target query conditions to the data platform layer 204.
[0039] In one possible implementation, the data middleware layer 204 is used to perform data queries in the database based on the target query conditions, determine the query results, and send the query results to the front-end presentation layer 205.
[0040] In one possible implementation, the front-end presentation layer 205 is used to visualize the query results, generate query results in a front-end renderable format, send the front-end renderable query results to the terminal layer 201, and display the query results in the terminal layer 201.
[0041] In one possible implementation, terminal layer 201 can be a web browser, a mobile application (e.g., a computer and smartphone), an in-vehicle system, or a smart wearable device. Terminal layer 201 can establish a Transmission Control Protocol / Internet Protocol (TCP / IP) connection with application server layer 202 to send natural language query statements in Hypertext Transfer Protocol (HTTP) request format.
[0042] In one possible implementation, the application server layer 202 can be an agent module deployed on a high-performance physical server cluster, comprising: a prompt word generation unit, a large model invocation unit, and a result verification unit. The agent module connects to the large model service layer 203 via HTTP or HTTPS protocols. Specifically, the prompt word generation unit is used to construct query prompt information, the large model invocation unit is used to invoke the large model, and the result verification unit is used to verify and optimize the parsed results.
[0043] In one possible implementation, the large model service layer 203 can be a cluster of graphics processing unit (GPU) servers running large models.
[0044] In one possible implementation, the data platform layer 204 can be a distributed database server cluster, which can connect to the application server layer 202 via a representational state transfer application programming interface (RESTful API) protocol. The data platform layer 204 can store structured data across all dimensions and respond to query requests through the database interface.
[0045] In one possible implementation, the front-end presentation layer 205 can connect to the data middleware layer 204 via the HTTP protocol, and can render the query results as a list, card or filter component, display it on the terminal layer, and support users to sort by age, job level, and perform secondary filtering.
[0046] Furthermore, the actions, terms, etc., involved in the various embodiments of this application can be referenced interchangeably without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are merely examples, and other names may be used in specific implementations without limitation.
[0047] It should be noted that the system architecture and application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems. For example, the data terminal 101, the data query device 102, and the data storage device 103 can be separate devices or different functional modules of the same device.
[0048] The data query method provided in this application embodiment can be applied to the aforementioned data query device, specifically to the processor of the data query device. Please refer to... Figure 3 The data query methods provided in this application specifically include S301-S303.
[0049] S301. In response to receiving a natural language query statement, determine multiple query keywords based on the natural language query statement, and construct query prompt information based on the query keywords.
[0050] The query suggestion information is used to indicate the business constraint rules associated with each query keyword.
[0051] Natural language queries are requests made in natural language to retrieve or manipulate data from data storage devices. They allow non-technical users to input queries in natural language, enabling them to retrieve data from databases without needing to master Structured Query Language (SQL) or complex query syntax.
[0052] For example, if a user's data query requirement is to find individuals whose project participation experience is the Changhe Project, whose job level is deputy department head, whose age limit is under 40, and whose personnel attribute is cadre, the natural language query statement would be: "Query cadres under 40 years old who have participated in the internal 'Changhe Project' and whose job level is deputy department head."
[0053] Query keywords are core words or phrases in a natural language query that are directly related to data retrieval or manipulation. Query keywords express the user's query intent, target data object, filtering conditions, and operation type. For example, the query keywords extracted from the query "querying cadres under 40 years old who participated in the internal 'Changhe Plan' and held the rank of deputy department head" would be "Changhe Plan," "deputy department head," "under 40 years old," and "cadre."
[0054] Query hints are instructions or statements that guide the large model to parse based on the business constraint rules associated with each query keyword. Query hints clearly define the parsing objective and guide the large model to generate expected parsing results. For example, a query hint based on "querying cadres under 40 years old who participated in the internal 'Changhe Plan' and hold the rank of deputy department head" could be: "Parse the cadre query statement, extract internal project, rank, and age conditions; 'Changhe Plan' needs to be associated with the company's internal standard project name; rank refers to the definition of 'deputy department head'; age is converted to a range format: [query statement]".
[0055] When a user needs to query data and performs a natural language query on a data terminal, it is necessary to first collect relevant information about the user's input data, such as the content and manner of input. Based on the input method (e.g., voice input, text input, and handwriting input), the input content is uniformly converted into plain text format, providing a data foundation for subsequent processing of the natural language query.
[0056] It should be noted that this application embodiment does not limit the way users input natural language query statements. In actual applications, it can be set according to needs to cover different interaction modes and device scenarios. For example, as one implementation method, the way to input natural language query statements may include at least one of the following: voice input, conversational interaction, template filling input, gesture or graphical input, text box input, and handwriting input.
[0057] In the data query process, to obtain accurate and comprehensive data query results, natural language queries often contain multiple words that describe the user's query intent, target data object, filtering conditions, and operation type. These words that describe the user's query intent, target data object, filtering conditions, or operation type are called query keywords. These query keywords are often independent phrases with independent structures within the natural language query statement; therefore, they can be extracted from the natural language query statement through semantic parsing.
[0058] In one possible implementation, determining multiple query keywords based on natural language query statements can be achieved by: after obtaining the natural language query statement input by the user, performing semantic parsing on the natural language query statement through named entity recognition, regular expression matching, or rule matching to determine the query keywords.
[0059] For example, as one implementation, the data query device may have a built-in bidirectional encoder representations from transformers (BERT) model or SpaCy natural language processing library. After obtaining the natural query statement, the natural query statement is input into the BERT model or SpaCy natural language processing library to determine the query keywords.
[0060] Determined query keywords from user-input natural queries may present the following problems: inconsistencies between query keywords and the company's internal terminology in the database; vague expressions of filtering conditions; unclear logical relationships between multiple conditions; and the inability of independent query keywords to guide the large model in accurate data retrieval. Therefore, it is necessary to construct query suggestion information based on query keywords to guide the large model in parsing based on the business constraint rules associated with each query keyword.
[0061] Business constraint rules are logical conditions or data and terminology specifications that are strongly related to specific business scenarios and are used to restrict or guide query behavior. Business constraint rules can be defined by business experts or generated by relevant devices or models. This application does not restrict the generation method of business constraint rules.
[0062] In one possible implementation, constructing query suggestion information based on query keywords can be achieved as follows: Based on the query keywords, determine the business constraint rules that match the query keywords; based on the matching business constraint rules and the query keywords, generate query suggestion information that represents the business constraint rules associated with each query keyword. For example, as one implementation, the data query device can have a built-in rule repository, which includes multiple business constraint rule-query keyword pairs consisting of query keywords and their corresponding business constraint rules. The business constraint rules corresponding to the query keywords are determined by matching the query keywords, and query suggestion information is generated based on these business constraint rules.
[0063] For example, as another implementation, the rule repository may also include query suggestion-query keyword pairs consisting of query suggestion information and query keywords, and the query suggestion information can be directly determined by matching the query keywords.
[0064] S302. Input the query prompt information and natural language query statement into the large model. Based on the business constraint rules associated with each query keyword, the large model determines the target query conditions that match the query intent of the natural language query statement.
[0065] The target query criteria are query statements or phrases determined by the large model based on query suggestions, conforming to business constraints and matching the user's query intent. The terms in these target query criteria are consistent with the company's internal terminology in the database.
[0066] While large-scale models possess powerful deep semantic understanding and multimodal information fusion capabilities, the accuracy of their interactive intent parsing depends on standardized input prompts. Therefore, when generating target query conditions using large-scale models, natural language queries and relatively accurate query prompts must be input. These prompts clarify the analysis tasks and output requirements of the large-scale model, guiding it to determine target query conditions that match the query intent of the natural language query based on the business constraint rules associated with each query keyword.
[0067] Guided by query prompts, the large model can effectively avoid ambiguous interpretations and clarify the user's query intent through related business constraint rules, thereby generating target query conditions that match the query intent of the natural language query statement.
[0068] The large-scale model possesses superior language understanding and analysis capabilities, enabling it to more accurately understand user query intent through query prompts and output query conditions that conform to business constraints and match the query intent. Guided by query prompts, the target query conditions output by the large-scale model conform to business constraints and match the user's query intent, thereby ensuring the accuracy of the target query conditions.
[0069] In one possible implementation, the query keywords that are not internal enterprise terms in the natural language query statement are converted into internal enterprise terms by the guidance of query prompt information, the query keywords with a relatively vague description range are converted into accurate range words or intervals, and multiple parallel conditions are connected by logical connectors.
[0070] For example, guided by the query prompt message "'Changhe Plan' needs to be associated with the company's internal standard project name," the query keyword "Changhe Plan" is converted into the corresponding internal term, such as "Changhe Project." Guided by the query prompt message "Job level refers to the definition of 'Deputy Department Head,'" "Deputy Department Head" is determined as the precise job level range, i.e., "Decision-making level head and deputy head." Guided by the query prompt message "parse the cadre query statement and extract internal project, job level, and age conditions," "around X years old" is converted into an age range [X-2, X+2], and "joined the company after X years" is converted into an entry time range [X-01-01, current date]. Furthermore, natural language logical operators can be converted into machine-recognizable logic; for example, "and" and "again" are converted to "AND," "or" and "either" are converted to "OR," and "except..." is converted to "NOT."
[0071] S303. Perform data query in the database based on the target query conditions.
[0072] After determining the target query conditions based on the business constraint rules associated with each query keyword using a large model, the target query conditions need to be validated. Specifically, this can be done by checking whether the words in the target query conditions match internal enterprise terminology to determine the word accuracy of the target query conditions, and by inputting the target query conditions and natural language query statements into an intent classification model to determine the intent coverage of the target query conditions.
[0073] If the intent coverage is less than the preset value, it means that the target query conditions are missing some query conditions in the natural language query statement, and the target query conditions cannot cover the user's query intent. Therefore, it is necessary to redetermine the query keywords and regenerate the target query conditions. If the word accuracy is less than the preset value, it means that the words in the target query conditions are not correctly mapped to the company's internal terms. As a result, it may be impossible to query data corresponding to the company's internal terms or to query incorrect data. Therefore, it is necessary to redetermine the query keywords and regenerate the target query conditions.
[0074] In one possible implementation, S303 can be implemented as follows: Input the target query conditions and natural language query statement into an intent classification model, and determine the word accuracy and intent coverage of the target query conditions through the intent classification model. If the word accuracy is less than a first threshold or the intent coverage is less than a second threshold, redetermine the query keywords. If the word accuracy is greater than the first threshold and the intent coverage is greater than the second threshold, perform a data query in the database based on the target query conditions.
[0075] For example, intent coverage can be determined by checking the completeness of the target query conditions, such as detecting whether the "internal project" dimension is missing. Terminology accuracy can also be determined by checking the accuracy of the terms in the target query conditions, such as whether "Changhe Project" is correctly mapped to internal company terminology, or whether the project code is valid. If intent coverage or terminology accuracy does not meet the preset conditions, the target query conditions are regenerated until both intent coverage and terminology accuracy fail to meet the preset conditions.
[0076] By validating the target query conditions, the problem of inaccurate target query conditions caused by large model parsing errors is effectively avoided, thus improving the accuracy of target query conditions and consequently improving the accuracy of data query results.
[0077] As can be seen from the technical solutions S301-S303 above, this application processes natural language query statements to construct query prompt information that describes business constraint rules for query keywords. This information instructs the large model to generate query conditions that conform to the business constraint rules for query keywords and match the query intent of the natural language query statement. The large model possesses powerful language understanding and result generation capabilities, enabling it to more accurately understand the user's query intent through the query prompt information and output query conditions that conform to the business constraint rules and match the query intent. The query conditions generated by the large model can retrieve relatively accurate data query results in the database. Therefore, this application can improve the accuracy of data query results.
[0078] In some embodiments, such as Figure 4As shown, in the above method S302, the application server layer determines multiple query keywords based on the natural language query statement, specifically including: S401-S403.
[0079] S401. Extract multiple initial keywords from a natural language query statement.
[0080] Natural language queries may include vague keywords that describe the user's query intent, target data object, filter conditions, and operation type. Generating query suggestions using vague keywords may lead to the large model generating incorrect target query conditions. These vague keywords are called initial keywords.
[0081] Initial keywords in natural language queries are characterized by semantic ambiguity. These initial keywords typically contain referential terms, such as "this position" or "that project," or broad terms, such as "a better colleague" or "recent plans," or other semantically ambiguous words. Specifically, initial keywords can be extracted by detecting referential or broad terms.
[0082] One possible implementation is to extract multiple initial keywords from a natural language query statement by matching the initial keywords in the natural language query statement based on a fuzzy thesaurus.
[0083] A fuzzy thesaurus is a collection of semantically ambiguous words and predefined matching rules. These predefined matching rules can identify semantically ambiguous words. For example, a natural language query like "help me find cadres who have performed well recently" would, through fuzzy thesaurus matching, extract the initial keywords "recently" and "performing well recently."
[0084] It should be noted that fuzzy word matching is a method chosen to extract multiple initial keywords from natural language query statements. It is not limited to fuzzy word matching. Other methods can also be used, such as extracting initial keywords through pre-trained models or through manual annotation. This application does not limit these methods.
[0085] S402. For each initial keyword, determine the similarity between the initial keyword and each preset keyword in the preset keyword set.
[0086] The preset keywords are the standard expressions of the initial keywords in a natural language query statement.
[0087] After extracting multiple initial keywords from a natural language query, it is necessary to convert these semantically ambiguous initial keywords into semantically clear words. This can be done by selecting predefined keywords that can replace the initial keywords from a pool of semantically clear predefined keywords. Specifically, the predefined keywords that can replace the initial keywords can be determined by calculating the similarity between predefined keywords. These predefined keywords are words or phrases with a single, unambiguous meaning, directly corresponding to database query conditions, and suitable for the application scenario.
[0088] In one possible implementation, for each initial keyword, determining the similarity between the initial keyword and each preset keyword in the preset keyword set can be achieved by: calculating the cosine similarity between the initial keyword and the preset keywords, and determining the cosine similarity as the similarity between the initial keyword and each preset keyword in the preset keyword set. Alternatively, the semantic vectors of the initial keyword and the preset keywords can be extracted, and the similarity can be determined by calculating the distance between the semantic vectors. Alternatively, the initial keyword and the preset keywords can be input into a pre-trained similarity calculation model to determine the similarity. This application does not limit this approach. In one possible implementation, the cosine similarity between the preset keywords and the initial keyword satisfies the following formula 1: Formula 1.
[0089] in, The semantic vector of the preset keywords. This is the semantic vector of the initial keyword.
[0090] S403. Select preset keywords that meet the preset similarity conditions as the query keywords corresponding to the initial keywords.
[0091] For each initial keyword, after determining the similarity between the initial keyword and each preset keyword in the preset keyword set, it is necessary to determine the query keyword that can replace the initial keyword based on the similarity. Specifically, this can be done by setting preset conditions (e.g., the similarity must be greater than a preset value). If the similarity between the preset keyword and the initial keyword meets the preset conditions, the preset keyword is determined as the query keyword corresponding to the initial keyword.
[0092] For example, the preset condition could be whether the similarity between the initial keyword and each preset keyword is greater than a preset value. Alternatively, it could be whether the weighted average of multi-dimensional similarities is greater than a preset value. This application does not limit this approach.
[0093] For each initial keyword, there may be multiple preset keywords that meet the preset conditions. Therefore, when multiple preset keywords meet the preset conditions, it is necessary to select one preset keyword as the query keyword corresponding to the initial keyword. Specifically, this can be achieved by sorting the similarity between the multiple preset keywords that meet the preset conditions and their corresponding initial keywords, and determining the preset keyword with the highest similarity as the query keyword corresponding to the initial keyword. In one possible implementation, when multiple preset keywords meet the preset conditions, selecting one preset keyword can be achieved by determining the preset keyword with the highest similarity as the query keyword corresponding to the initial keyword.
[0094] For example, for the initial keyword "large-screen device", two preset keywords that meet the preset similarity criteria are selected: "large-screen mobile phone" and "large-screen TV". The similarity between "large-screen mobile phone" and "large-screen device" is 0.8, and the similarity between "large-screen TV" and "large-screen device" is 0.7. Among the two preset keywords that meet the preset similarity criteria of 0.7 or greater, the preset keyword with the highest similarity is "large-screen mobile phone". Therefore, "large-screen mobile phone" is determined as the query keyword for "large-screen device".
[0095] For each initial keyword, there may be no preset keyword that meets the preset conditions. This means that among the multiple preset keywords, there is no preset keyword that can semantically express the initial keyword. If a preset keyword that does not meet the preset conditions is selected as the query keyword for the initial keyword, it may lead to incomplete or incorrect semantic expression. To ensure coverage of the user's query intent, the initial keyword can be determined as the query keyword. In one possible implementation, if the similarity between the initial keyword and each preset keyword does not meet the preset conditions, the application server layer determines the initial keyword as the query keyword.
[0096] As can be seen from the above technical solutions S401-S403, this application, by pre-setting multiple preset keywords and converting some initial keywords into preset keywords, transforms ambiguous words in natural language query statements into accurate and clear words, thereby providing an accurate data foundation for subsequently determining the query intent of natural language query statements and improving the accuracy of data query results.
[0097] In some embodiments, such as Figure 5 As shown, in the above method S303, the target query conditions include: keywords, business constraint rules, and logical connectors. Multiple query keywords are determined based on natural language query statements, specifically including: S501-S502.
[0098] S501, the data platform layer inputs keywords, business constraint rules, and logical connectors into a standardized format conversion program to obtain a query statement in a preset format.
[0099] The target query condition is a query statement or phrase determined by the large model based on query prompts, conforming to business constraints and matching the user's query intent. Guided by query prompts and semantic analysis by the large model, the target query condition ensures compliance with business constraints and matches the user's query intent, enabling relatively accurate data retrieval. For databases that do not support direct data retrieval using target query conditions, these conditions can be converted into query statements written in program code that the database can recognize.
[0100] The standardized format conversion program is a Java-based program that can receive input keywords, business constraint rules, and logical connectors, and convert them into query code written in Java, i.e., query statements in a preset format.
[0101] For example, as one implementation method, a standardized format conversion program can be installed in the data query device. After the target query conditions are generated, the data query device can input the target query conditions into the standardized format conversion program to generate a query statement in a preset format.
[0102] After obtaining the query statement in the preset format, in order to perform data query in the database, the query statement in the preset format needs to be sent to the data storage device. To prevent data leakage, the query statement in the preset format needs to be encrypted and encapsulated. Specifically, the query statement in the preset format can be encrypted and encapsulated using a data encryption algorithm to generate an encrypted request body, and then the encrypted request body is sent to the data storage device for data query.
[0103] In one possible implementation, the data platform layer uses a standardized format conversion program to convert keywords, business constraint rules, and logical connectors into standardized query statements, and then encrypts and encapsulates the standardized query statements to generate an application programming interface (API) request body.
[0104] For example, a query statement in a preset format can be encrypted using a symmetric encryption algorithm, a field-level encryption algorithm, or an asymmetric encryption algorithm. This application does not limit the encryption method for the preset format query statement.
[0105] S502. Perform data retrieval in the database based on a query statement with a preset format.
[0106] The preset query statement is a query code written in Java. It meets the syntax requirements of the database deployed on the data storage device, and can be used directly to query data in the database with a fast query speed.
[0107] As a feasible implementation method, data querying in the database based on a pre-formatted query statement can be achieved as follows: in response to receiving a pre-formatted query statement, inputting the pre-formatted query statement into the database and obtaining the data query results.
[0108] For example, as one implementation, a database query program can be installed in the data storage device. The database query program provides a query statement input window, which includes a "Start Query" control. In response to receiving a query statement in a preset format, the data storage device inputs the query statement in the preset format into the query statement input window and triggers "Start Query", thus starting the data storage device to perform a data query.
[0109] Data is queried in the database based on a pre-formatted query statement. After the query results are determined, they need to be given to the user. Specifically, the query results can be visualized to obtain a front-end renderable format, and then sent to the data terminal for display.
[0110] In one possible implementation, a data query is performed in the database based on a query statement in a preset format to determine the query results. Then, the query results can be converted into data query results in a format that can be rendered by the front end, and the data query results in this format can be sent to the front end presentation layer to display information such as cadre name, unit, rank, and project experience in the form of a list or card.
[0111] As can be seen from the above technical solutions S501-S502, this application generates a preset format query statement by inputting keywords, business constraint rules, and logical connectors into a standardized format conversion program. This preset format query statement can perform data queries in the database quickly, shortening the data query time and thus improving the efficiency of data query.
[0112] In some embodiments, such as Figure 6 As shown, the process of training a large model includes S601-S603.
[0113] S601. Obtain industry corpus and enterprise corpus.
[0114] Industry-specific corpora are collections of text data that are collected, organized, and labeled based on a specific industry sector. They are used to train large-scale models, enabling the models to recognize specialized terminology within that industry. Enterprise corpora are collections of internal text data that are collected, organized, and labeled for a specific enterprise or organization. They are also used to train large-scale models, enabling the models to recognize specialized terminology within that enterprise or organization.
[0115] For example, the industry corpus includes standards for cadre ranks (e.g., "above the deputy level of management" includes both the head and deputy head of the decision-making level), age range definitions (e.g., "the age range of around 40 years old is 38-42 years old"), and general rules such as industry terminology mapping. The enterprise corpus includes internal unit aliases, internal project names, department abbreviations, and other enterprise-specific information. The industry corpus and the enterprise corpus can be collectively referred to as a two-level knowledge base.
[0116] Before determining the target query conditions that match the query intent of the natural language query statement based on the business constraint rules associated with each query keyword through the large model, if the initial large model without industry and enterprise corpora is used directly, it may fail to identify specific industry fields and internal professional terms in the query keywords and natural language query statements. Therefore, the initial large model needs to be trained.
[0117] Before training a large model using industry and enterprise corpora, it is necessary to obtain industry and enterprise corpora. Specifically, industry corpora can be obtained from publicly available data in multiple specific industry sectors, and enterprise corpora can be obtained by integrating internal enterprise data.
[0118] In one possible implementation, obtaining industry corpora and enterprise corpora can be achieved by: determining industry corpora based on publicly available data from multiple industry sectors, and determining enterprise corpora based on internal enterprise data.
[0119] For example, as one possible implementation, the data query device can be equipped with a data acquisition program connected to the internet. This program can acquire publicly available data from multiple specific industry sectors from the internet at preset time intervals and construct an industry corpus based on this publicly available data. Alternatively, the data acquisition program can be connected to a data storage device and can acquire internal enterprise data from a database deployed on the data storage device at preset time intervals and construct an enterprise corpus based on this internal data.
[0120] S602. Based on industry corpora and enterprise corpora, fine-tune the initial large model.
[0121] The two-level knowledge base, consisting of an industry corpus and an enterprise corpus, includes professional terminology from both the industry and enterprise sectors. By using this two-level knowledge base to train the initial large-scale model, it gains the ability to recognize these professional terms. Specifically, the data in the two-level knowledge base can be divided into training and testing sets. The initial large-scale model is then trained using the training set data and a pre-defined loss function.
[0122] In one possible implementation, based on industry corpora and enterprise corpora, fine-tuning the initial large model can be achieved by minimizing the prediction error of the initial large model based on the cross-entropy loss function, and optimizing the model parameters of the initial large model through gradient descent.
[0123] For example, as one implementation, the data query device is equipped with an initial large model fine-tuning optimization program. In response to receiving an industry corpus and an enterprise corpus, the initial large model is trained based on the industry corpus, the enterprise corpus, the cross-entropy loss function, and the gradient descent method.
[0124] S603. If the initial large model after fine-tuning meets the preset conditions, the initial large model is determined as the large model.
[0125] In one possible implementation, the preset condition for determining the completion of the initial large model training is: input the query prompts containing test set data into the initial large model; if the initial large model can accurately identify the test set data in the query prompts, then the initial large model training is complete.
[0126] For example, the query suggestion containing test set data is "identify 'Blue Ocean Project' and associate it with the company's internal standard project name". If the large model can accurately output the target query condition containing "Blue Ocean Project" and associate it with the company's internal standard project name, it can be said that the initial large model training is complete.
[0127] As can be seen from the above technical solution for training the large model, by training the initial large model with an industry corpus containing general rules and industry terms and an enterprise corpus containing internal enterprise terms, the large model can recognize industry terms and internal enterprise terms in natural language query statements, and can recognize fuzzy expressions based on general rules, thereby improving the accuracy of user intent recognition.
[0128] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. It is understood that, in order to achieve the above functions, the data query device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the data query method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0129] This application also provides a data query device. This data query device can be a server, the CPU of the server, a data query module within the server, or a client for data querying within the server.
[0130] This application embodiment can divide the data query device into functional modules or functional units according to the above method examples. For example, each function can be divided into its own functional modules or functional units, or two or more functions can be integrated into one processing unit. The integrated modules can be implemented in hardware or in software functional modules or functional units. The module or unit division in this application embodiment is illustrative and represents only one logical functional division; other division methods may be used in actual implementation.
[0131] When dividing each function into modules according to its corresponding function. Figure 7 A structural diagram of a data query device provided in this application is shown below. Figure 7 As shown, the data query device can be used to perform... Figure 3 , Figure 4 , Figure 5The data query method shown is illustrated in the diagram. The data query device 70 includes: a determining module 701, a receiving module 702, an input module 703, and a query module 704. The receiving module 702 is used in response to receiving a natural language query statement. The determining module 701 is used to determine multiple query keywords based on the natural language query statement, and is also used to construct query hint information based on the query keywords. The query hint information represents the business constraint rules associated with each query keyword. The input module 703 is used to input the query hint information and the natural language query statement into a large model. The determining module 701 is also used to determine target query conditions matching the query intent of the natural language query statement based on the business constraint rules associated with each query keyword through the large model. The query module 704 is used to perform data queries in the database based on the target query conditions.
[0132] This application also provides a data query device; the data query device can be used to execute the data query method provided in any of the above embodiments. Figure 8 This is a schematic diagram of the structure of a data query device 80 provided in an embodiment of this application. Figure 8 As shown, the data query device 80 may include a processor 801, a bus 802, a communication interface 803, and a memory 804.
[0133] Furthermore, the data query device 80 may also include a communication interface 803 and a memory 804. The processor 801, the memory 804, and the communication interface 803 can be connected via a bus 802.
[0134] The processor 801 can be a CPU, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 801 can also be other devices with processing capabilities, such as circuits, devices, or software modules, without limitation.
[0135] Bus 802 is used to transmit information between the components included in the data query device 80.
[0136] Communication interface 803 is used to communicate with other devices or other communication networks. These other communication networks can be Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc. Communication interface 803 can be a module, circuit, communication interface, or any device capable of enabling communication.
[0137] The memory 804 is used to store instructions. These instructions can be computer programs.
[0138] The memory 804 can be a read-only memory (ROM) or other type of static storage device that can store static information and / or instructions; it can also be a random access memory (RAM) or other type of dynamic storage device that can store information and / or instructions; it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, etc., without limitation.
[0139] It should be noted that the memory 804 can exist independently of the processor 801 or can be integrated with the processor 801. The memory 804 can be used to store instructions, program code, or some data, etc. The memory 804 can be located inside or outside the data query device 80, without limitation. The processor 801 is used to execute the instructions stored in the memory 804 to implement the data query method provided in the following embodiments of this application.
[0140] In one example, processor 801 may include one or more CPUs.
[0141] As an optional implementation, the data query device 80 includes multiple processors.
[0142] As an optional implementation, the data query device 80 also includes an output device and an input device, which are not shown in the figure.
[0143] In this embodiment of the application, the chip system may be composed of chips or may include chips and other discrete devices.
[0144] This disclosure also provides a computer-readable storage medium storing instructions that, when executed by a processor of an electronic device, enable the electronic device to perform the data query method provided in the embodiments of this disclosure described above.
[0145] This disclosure also provides a computer program product containing instructions that, when run on an electronic device, cause the electronic device to execute the data query method provided in the above-described embodiments of this disclosure.
[0146] The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires; a portable computer disk drive; a hard disk drive; a random access memory (RAM); a read-only memory (ROM); an erasable programmable read-only memory (EPROM); a register; a hard disk drive; an optical fiber; a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination thereof; or any other form of computer-readable storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). In the embodiments of this application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0147] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0148] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0149] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the classified units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0150] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0151] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, essentially, or the part that contributes to the prior art, or a complete or partial classification of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0152] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data query method, characterized in that, The method includes: In response to receiving a natural language query statement, multiple query keywords are determined based on the natural language query statement, and query suggestion information is constructed based on the query keywords; the query suggestion information is used to represent the business constraint rules associated with each query keyword; The query prompt information and the natural language query statement are input into the large model. Based on the business constraint rules associated with each query keyword, the large model determines the target query conditions that match the query intent of the natural language query statement. Data is queried in the database based on the target query conditions.
2. The method according to claim 1, characterized in that, The determination of multiple query keywords based on the natural language query statement includes: Extract multiple initial keywords from the natural language query statement; For each initial keyword, determine the similarity between the initial keyword and each preset keyword in the preset keyword set; Select preset keywords that meet preset similarity criteria as the query keywords corresponding to the initial keywords.
3. The method according to claim 2, characterized in that, The method further includes: If the similarity between the initial keyword and each of the preset keywords does not meet the preset conditions, the initial keyword is determined as the query keyword.
4. The method according to claim 1, characterized in that, The process of querying data in the database based on the target query conditions includes: The target query conditions and the natural language query statement are input into the intent classification model, and the word accuracy and intent coverage of the target query conditions are determined by the intent classification model. If the word accuracy is less than a first threshold or the intent coverage is less than a second threshold, the query keywords will be re-determined. If the word accuracy is greater than the first threshold and the intent coverage is greater than the second threshold, a data query is performed in the database based on the target query conditions.
5. The method according to claim 1, characterized in that, The target query conditions include: keywords, business constraint rules, and logical connectors; The process of querying data in the database based on the target query conditions includes: Input the keywords, business constraint rules, and logical connectors into a standardized format conversion program to obtain a query statement in a preset format; Data is retrieved from the database based on the preset query statement.
6. The method according to claim 1, characterized in that, The method further includes: Acquire industry corpora and enterprise corpora; Based on the industry corpus and the enterprise corpus, the initial large model was fine-tuned; If the initial large model after fine-tuning meets the preset conditions, the initial large model is determined as the large model.
7. The method according to claim 6, characterized in that, The industry corpus includes at least one of the following: business constraint rules for at least one industry, industry terminology mapping rules for the at least one industry, and multimodal data for the at least one industry. The enterprise corpus includes at least one of the following: internal terms of at least one enterprise, abbreviations of the internal terms, and alternative names of the internal terms.
8. A data query device, characterized in that, The device includes: a determining module, a receiving module, an input module, and a query module; The receiving module is configured to respond to receiving a natural language query statement; the determining module is configured to determine multiple query keywords based on the natural language query statement; and the determining module is further configured to construct query prompt information based on the query keywords; the query prompt information is used to represent the business constraint rules associated with each query keyword. The input module is used to input the query prompt information and the natural language query statement into the large model. The determination module is also used to determine the target query conditions that match the query intent of the natural language query statement based on the business constraint rules associated with each query keyword through the large model. The query module is used to perform data queries in the database based on the target query conditions.
9. An electronic device, characterized in that, include: A processor and a communication interface; the communication interface is coupled to the processor, the processor being configured to run computer programs or instructions to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium storing instructions, characterized in that, When the computer executes the instruction, the computer performs the method described in any one of claims 1-7.
Citation Information
Cited By
Artificial intelligence interactive combined query method, system and electronic device
CN122220483A