Natural language conversion method and device and computer equipment

By receiving, filtering and converting natural language query statements, the problem of difficult to remove irrelevant content in user input in the prior art is solved, and the accuracy and user experience of database query are improved.

CN120277088APending Publication Date: 2025-07-08SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311863870.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately filter out content that is not related to database queries in the natural language input by users, resulting in poor user experience and reduced trust.

Method used

By receiving query statements input from natural language, using semantic analysis and similarity matching to filter out obviously unrelated content, generate query statements that can be identified by the database, and convert them into SQL statements through Text2SQL technology, and then filter the difficult-to-reply parts through entropy calculation to obtain the target query statement.

Benefits of technology

It accurately filters out content that is not related to database queries in natural language, improving the accuracy and user experience of query results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277088A_ABST
    Figure CN120277088A_ABST
Patent Text Reader

Abstract

The embodiment of the invention is applicable to the technical field of semantic recognition, and provides a natural language conversion method and device and computer equipment, and the method comprises the following steps: receiving a first query statement which is input in a natural language and is used for querying a target database; filtering statements irrelevant to the target database in the first query statement to obtain a second query statement; generating a third query statement based on the second query statement, wherein the third query statement is a statement which is expressed in a database language and can be recognized by the target database; and filtering the third query statement to obtain a target query statement. By adopting the method, the content irrelevant to the query of the target database in the query statement can be accurately filtered out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application belong to the technical field of semantic recognition, and particularly relate to a natural language conversion method, apparatus, and computer device. Background Art

[0002] Text2SQL is a technology that converts a user's natural language into an executable SQL statement. By applying Text2SQL, it is convenient for the operating object to better understand the true idea contained in the user's natural language and execute it accordingly. Exemplarily, for a database, the user can use natural language to express the idea of querying certain information in the database. By converting the user's natural language into an SQL statement that the database can recognize, the database can find the corresponding information and feedback it to the user.

[0003] When the user uses natural language to express the query requirement for the database, the natural language actually used by the user may include a large amount of content, and there may be inputs that are completely irrelevant to the database query in this content. For example, the natural language input by the user is "How are you?" or "Where to go on weekends?" and other content that is obviously irrelevant to the database query. It is difficult for the model to convert it into an appropriate SQL statement. Or, when facing a natural language that is difficult to recognize, if the model converts it into an inappropriate SQL statement, it will not only bring inconvenience to the actual use experience of the user, but also affect the user's trust in the model, and may lead to the user reducing the use of such models.

[0004] Therefore, how to quickly filter out the content in the natural language input by the user that is irrelevant to the database query is an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a natural language conversion method, apparatus, and computer device to solve the problems existing in the prior art, so as to accurately filter out the content in the natural language that is irrelevant to the database query.

[0006] The first aspect of the embodiments of the present application provides a natural language conversion method, including:

[0007] Receiving a first query statement for querying a target database input in natural language;

[0008] Filtering out the statements in the first query statement that are irrelevant to the target database to obtain a second query statement;

[0009] Generating a third query statement based on the second query statement, where the third query statement is a statement that can be recognized by the target database represented in database language;

[0010] Filter the third query statement to obtain a target query statement.

[0011] The second aspect of the embodiments of the present application provides a natural language conversion device, including:

[0012] A statement receiving module, configured to receive a first query statement for querying a target database input in natural language;

[0013] A first filtering module, configured to filter out statements irrelevant to the target database in the first query statement to obtain a second query statement;

[0014] A statement generating module, configured to generate a third query statement based on the second query statement, where the third query statement is a statement recognizable by the target database represented in database language;

[0015] A second filtering module, configured to filter the third query statement to obtain a target query statement.

[0016] The third aspect of the embodiments of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the natural language conversion method described in the first aspect above is implemented.

[0017] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the natural language conversion method described in the first aspect above is implemented.

[0018] The fifth aspect of the embodiments of the present application provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the natural language conversion method described in the first aspect above.

[0019] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0020] Applying the conversion method provided by the embodiments of the present application, when a computer device receives a first query statement for querying a target database input in natural language, the computer device can perform a preliminary screening on the first query statement, filter out the content in the first query statement that is obviously irrelevant to the query of the target database, and obtain a second query statement also represented in natural language. Then, the computer device can generate a third query statement recognizable by the target database based on the second query statement. For the third query statement, the computer device can reprocess it, filter out the part with a relatively high answering difficulty and difficult to obtain an accurate answer result, and obtain a target query statement. The computer device can use the target query statement to query in the target database and obtain an accurate query result. Through two filtrations, the embodiments of the present application can accurately filter out the content in natural language that is irrelevant to database query, so that the final query result meets the actual needs of users and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic diagram of a natural language conversion method provided by the embodiments of the present application;

[0023] Figure 2 It is a schematic diagram of another natural language conversion method provided by the embodiments of the present application;

[0024] Figure 3 It is a schematic diagram of a natural language conversion process provided by the embodiments of the present application;

[0025] Figure 4 It is a schematic diagram of a natural language conversion device provided by the embodiments of the present application;

[0026] Figure 5 It is a schematic diagram of a computer device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] In the following description, specific details such as specific system structures and technologies are proposed for illustration rather than limitation in order to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0028] The technical solution of the present application will be described below through specific embodiments.

[0029] Referring to Figure 1 , a schematic diagram of a natural language conversion method provided by an embodiment of the present application is shown, which may specifically include the following steps:

[0030] S101. Receive a first query statement for querying a target database input in natural language.

[0031] This method can be applied to a computer device, that is, the execution subject of the embodiment of the present application is a computer device. By executing each step in the filtering method provided by the embodiment of the present application, the computer device can filter out the content in the user's natural language that is irrelevant to database query. Then, the computer device can convert the filtered statement into a language recognizable by the database. For example, the remaining part of the natural language obtained after filtering is converted into SQL language. SQL language is a structured query language applied to databases. Using SQL language, data can be quickly queried in the database to obtain corresponding data records.

[0032] In a possible implementation manner of the embodiment of the present application, the computer device may be an electronic device such as a mobile phone, a tablet computer, or a desktop computer. The computer device can be communicatively connected to the database. The user can obtain the data to be queried from the database by interacting with the computer device. That is, the computer device can provide an interaction interface for interacting with the user. For example, the interaction operation may be an operation of inputting a query statement in the form of natural language. The interaction operation of the user on the interaction interface can be received and processed by the computer device, and the processed interaction operation can be transmitted to the database, and the database executes the operation corresponding to the interaction operation. For example, the database queries a certain type of data and returns the corresponding query result. The query result can be displayed to the user through the interaction interface of the computer device.

[0033] In the embodiment of the present application, the target database may refer to the database corresponding to the interaction operation of the user on the computer device. Exemplarily, if the user hopes to query a certain type of data, the database storing this type of data is the target database. The target database may be a single database or may include multiple databases, and the number thereof may be determined according to the actual query requirements. The embodiment of the present application does not limit this.

[0034] In the embodiments of the present application, the first query statement may refer to the original query statement input by a user on a computer device, and this query statement may be input into the computer device in the form of natural language. In one possible implementation, the user may directly input the first query statement in text form on the computer device; alternatively, the user may also speak the first query statement to the computer device in voice, and after receiving the voice information input by the user, the computer device may convert the voice information into text form to obtain the first query statement in text form. The embodiments of the present application do not limit the specific manner in which the user inputs the first query statement.

[0035] Exemplarily, the first query statement may be "Hello, how many boys are there in the whole grade?"

[0036] S102. Filter out the statements in the first query statement that are irrelevant to the target database to obtain a second query statement.

[0037] In the embodiments of the present application, filtering out the statements in the first query statement that are irrelevant to the target database may refer to filtering out the content in the first query statement that is obviously irrelevant to database query, and this process can be regarded as a primary screening of the first query statement.

[0038] In one possible implementation of the embodiments of the present application, a preprocessing module may be included in the computer device, and the computer device may use this preprocessing module to perform a primary screening on the first query statement and filter out the obviously irrelevant content therein.

[0039] In the embodiments of the present application, the preprocessing module may use semantic analysis to analyze the specific meaning of each statement segment in the first query statement to determine whether it belongs to the content related to database query, so as to determine whether to filter the corresponding statement segment. Exemplarily, in the first query statement in the foregoing example, through semantic analysis, it can be known that the statement segment "Hello" does not belong to the content related to database query, and the preprocessing module may delete this statement segment.

[0040] In another possible implementation of the embodiments of the present application, the preprocessing module may also use the method of similarity matching to perform a primary screening on the first query statement. Exemplarily, the preprocessing module may calculate the similarity between each statement segment in the first query statement and the information describing the data stored in the target database respectively, and determine whether each statement segment is related to database query by analyzing the similarity between the two. If a certain statement segment is related to database query, then this statement segment may be retained, otherwise, the preprocessing module may delete this statement segment.

[0041] After filtering out the content that is obviously irrelevant to the database query in the first query statement, the remaining query statement is the second query statement, and the second query statement is still a query statement in natural language form. For example, in the first query statement of the foregoing example, after filtering out the irrelevant information "Hello", the second query statement is "How many boys are there in the whole grade?"

[0042] S103. Generate a third query statement based on the second query statement, where the third query statement is a statement recognizable by the target database represented in database language.

[0043] In the embodiment of the present application, the third query statement can be a statement that can be accurately recognized by the database. In the case of database query, a statement that can be accurately recognized by the database is also a statement that can be normally understood by the computer device, and the computer device can perform a query operation in the database according to its requirements to obtain a corresponding data query result.

[0044] In a possible implementation manner of the embodiment of the present application, the third query statement can be an SQL statement. The computer device can apply the Text2SQL technology or model to convert the second query statement represented in natural language into the third query statement.

[0045] S104. Filter the third query statement to obtain a target query statement.

[0046] In the embodiment of the present application, filtering the third query statement may refer to reprocessing the SQL statement that has been converted into a statement recognizable by the database, and the purpose is to further filter out some content that is difficult to distinguish and for which the computer device is difficult to obtain an accurate answer.

[0047] In a possible implementation manner of the embodiment of the present application, the computer device may include a post-processing module, and the computer device can use this post-processing module to process the third query statement to filter out the content that is difficult to answer.

[0048] By re-filtering the third query statement, a target query statement that can accurately match the actual requirements of the original first query statement can be obtained. The target query statement can be a statement of the same type as the third query statement. For example, the target query statement can also be an SQL statement.

[0049] In an embodiment of the present application, when a computer device receives a first query statement for querying a target database input in natural language, the computer device can perform a preliminary screening on the first query statement, filter out the content in the first query statement that is significantly irrelevant to the query of the target database, and obtain a second query statement also represented in natural language. Then, the computer device can generate a third query statement recognizable by the target database based on the second query statement. For the third query statement, the computer device can perform reprocessing on it, filter out the parts with greater answering difficulty and difficult to obtain accurate answering results, and obtain a target query statement. The computer device can use the target query statement to query in the target database and obtain an accurate query result. Through two filters in the embodiment of the present application, the content irrelevant to database query in natural language can be accurately filtered out, so that the final query result meets the actual needs of users and improves the user experience.

[0050] Referring to Figure 2 , a schematic diagram of another natural language conversion method provided by an embodiment of the present application is shown, which may specifically include the following steps:

[0051] S201. Receive a first query statement for querying a target database input in natural language.

[0052] In an embodiment of the present application, the target database may refer to the database corresponding to the interactive operations of the user on the computer device. The target database may be a single database or may include multiple databases. The first query statement may refer to the original query statement input by the user on the computer device, and the first query statement may be input into the computer device in the form of text or voice information. For the first query statement in voice form, the computer device can use speech-to-text technology to convert it into a query statement in text form.

[0053] S202. Obtain formatting information of the target database, where the formatting information is information representing the data stored in the target database in a preset standard format.

[0054] In an embodiment of the present application, the formatting information of the target database may be information representing the data stored in the target database in a preset standard format. Through the formatting information, the computer device can quickly understand what types of data are stored in the target database, what fields are included in each data table in the database, the meanings represented by different fields, and the corresponding value-taking situations, etc.

[0055] In a possible implementation manner of an embodiment of the present application, when obtaining the formatting information of the target database, first, the column names of each data table stored in the target database and the data values corresponding to the column names can be identified, and then, based on the column names and their corresponding data values, the formatting information of the target database can be generated.

[0056] Exemplarily, multiple databases are stored in the target database. For example, data table 1, data table 2, and so on. Among them, data table 1 includes multiple fields, and each field corresponds to a column in data table 1. Therefore, the field names of each field in data table 1 can be regarded as the column names of each column. For example, data table 1 includes fields such as "age" and "gender". Correspondingly, the specific data stored under each field is also the data value corresponding to the column name. For example, the "gender" field represents gender, and under this field, code items can be used to represent male and female respectively. That is, the number 1 is used to represent male, and the number 0 is used to represent female. Therefore, under the "gender" field, its values only include two cases, that is, only the cases where the value is 1 and the value is 0.

[0057] After identifying the above information, the formatted information of the target database can be generated.

[0058] As an example of the embodiment of the present application, by identifying the data stored in the target database, the computer device can obtain the following formatted information of the target database:

[0059] Table_name: Explanation of the table name of data table 1

[0060] age: Age,

[0061] gender: Gender ({1: Male, 0: Female}),

[0062] column3: Grade (['First grade', 'Second grade', 'Third grade', 'Fourth grade', 'Fifth grade', 'Sixth grade'])

[0063] Table2_name: Explanation of the table name of data table 2 ...

[0065] In the formatted information of the above example, the general information of each data table in the target database and the data stored under each field are listed respectively.

[0066] S203. Calculate the similarity between each statement in the first query statement and the formatted information respectively.

[0067] In the embodiment of the present application, the first query statement may include multiple statements. For example, in the first query statement "Hello, how many boys are there in the whole grade?" in the above example, at least the first query statement can be divided into two different statements, that is, statement 1 is "Hello", and statement 2 is "How many boys are there in the whole grade?". The division of the first query statement can be implemented based on word segmentation and / or semantic recognition methods, and the embodiment of the present application does not limit this.

[0068] After dividing the first query statement into multiple statements, the similarity between each statement and the formatting information of the target database can be calculated separately.

[0069] In a possible implementation manner of the embodiment of the present application, calculating the similarity between each statement and the formatting information of the target database may refer to calculating the cosine similarity therebetween. When calculating the cosine similarity, each statement in the first query statement can be converted into a first vector, and then the column names in the target database and their corresponding data values can be converted into a second vector. By calculating the cosine similarity between the first vector and the second vector, the similarity between each statement and the formatting information of the target database can be obtained.

[0070] In a specific implementation, the target database includes multiple fields (multiple columns), and the data stored in each field may not be the same. When calculating the similarity between each statement and the formatting information of the target database, the formatting information of each data table can be split into multiple pieces. For example, the formatting information of each data table can be split according to different fields or column names to obtain multiple pieces of information.

[0071] Exemplarily, in the database formatting information of the foregoing example, for the formatting information of Data Table 1, it at least includes fields such as "Table_name", "age", "gender", and "column3". For each field and its stored data value, it can be considered separately. Therefore, after splitting the formatting information of Data Table 1, at least 4 pieces of information can be obtained, namely:

[0072] Information 1: Table_name: Explanation of the name of Data Table 1

[0073] Information 2: age: Age,

[0074] Information 3: gender: Gender ({1: male, 0: female}),

[0075] Information 4: column3: Grade (['First grade', 'Second grade', 'Third grade', 'Fourth grade', 'Fifth grade', 'Sixth grade'])

[0076] For the above 4 pieces of information, the computer device can extract features from each piece of information separately, and the extracted features can be represented by vectors. Therefore, after the above 4 pieces of information are subjected to feature extraction, they can be converted into at least 4 vectors, that is, 4 second vectors. The corresponding relationship between each piece of the above information and the second vector can be expressed as:

[0077] Information 1: Table_name: Explanation of the name of Data Table 1 → Second vector 1

[0078] Information 2: age: age, → Second vector 2

[0079] Information 3: gender: gender ({1: male, 0: female}), → Second vector 3

[0080] Information 4: column3: grade (['first grade','second grade', 'third grade', 'fourth grade', 'fifth grade','sixth grade']) → Second vector 4

[0081] On the other hand, the features of each statement in the first query statement can be extracted in the same way and the extracted features can be represented using vectors. Each statement in the first query statement can be represented by a vector. Exemplarily, in the foregoing example, statement 1 “Hello” can be converted into a vector after feature extraction, and statement 2 “How many boys are there in the whole grade?” can also be converted into a vector after feature extraction. In this way, the above first query statement can be converted into 2 vectors, that is, 2 first vectors. The corresponding relationship between each of the above statements and the first vector can be expressed as:

[0082] Statement 1: “Hello” → First vector 1

[0083] Statement 2: “How many boys are there in the whole grade?” → First vector 2

[0084] When calculating the similarity between the first query statement and the formatted information of the target database, the cosine similarity between the first vector corresponding to each statement in the first query statement and the second vector corresponding to the formatted information can be directly calculated. The first vector corresponding to each statement can calculate the similarity with N second vectors corresponding to the formatted information.

[0085] Exemplarily, for statement 1, the cosine similarity between first vector 1 and second vector 1, second vector 2, second vector 3, and second vector 4 can be calculated respectively. In this way, 4 cosine similarities can be calculated for statement 1. For statement 2, the cosine similarity between first vector 2 and second vector 1, second vector 2, second vector 3, and second vector 4 can be calculated respectively. Therefore, 4 cosine similarities can also be calculated for statement 2.

[0086] The process of converting each statement and each piece of information into corresponding vectors in the foregoing example is text vectorization. Through text vectorization, the above text information such as each statement or each piece of information can be represented as a vector that can express its text semantics, which is to represent the semantics of the text with a numerical vector. In practical applications, for each split statement and each piece of formatted information in the database, the same feature extraction network can be used for feature extraction and converted into vectors based on the extracted features.

[0087] In a possible implementation manner of the embodiments of the present application, a Chinese general text representation model GTE or a Chinese-English semantic vector model BGE (BAAI General Embedding) can be used to convert each statement and each formatted information of the database into corresponding vectors.

[0088] In another possible implementation manner of the embodiments of the present application, when applying the GTE model or the BGE model, relevant data of the database query scenario involved in the embodiments of the present application can also be combined to fine-tune the model, further improving the accuracy of vector conversion and the accuracy of the cosine similarity calculated based on the vectors subsequently, so that the relevance between each statement determined based on the cosine similarity and the database query purpose is more accurate. S204. Delete the statements corresponding to the target similarities in the similarities that are less than the first preset threshold to obtain a second query statement.

[0089] In the embodiments of the present application, it is possible to determine whether each statement is relevant to the query of the target database according to the magnitude relationship between the similarity and the preset threshold. Specifically, the statements corresponding to the target similarities in the similarities that are less than the first preset threshold can be deleted to obtain a second query statement.

[0090] In a possible implementation manner of the embodiments of the present application, since N similarities can be calculated for each statement, it can be considered that the statement is irrelevant to the query of the target database when all N similarities are less than the first preset threshold.

[0091] Specifically, for any statement in the first query statement, if the cosine similarity between the first vector corresponding to the statement and any second vector is less than the first preset threshold, it can be considered that the statement is an obviously irrelevant statement, and the computer device can delete the statement to obtain a second query statement.

[0092] Exemplarily, for statement 1 "Hello" in the above example, the cosine similarity between the corresponding first vector 1 and 4 second vectors is calculated, and the obtained results are all less than the set first preset threshold. Therefore, it can be considered that the statement is an obviously irrelevant statement to the target database, and the computer device can delete it.

[0093] For statement 2 "How many boys are there in the whole grade?" in the above example, the cosine similarity between the corresponding first vector 2 and 4 second vectors is calculated. Among the obtained results, the cosine similarities between the first vector 2 and the second vector 1 and the second vector 2 are less than the set first preset threshold, but the cosine similarities between the first vector 2 and the second vector 3 and the second vector 4 are greater than the first preset threshold. Therefore, it can be considered that the statement is not obviously irrelevant to the target database, and the computer device can retain the statement.

[0094] Therefore, the finally retained second query statement is "How many boys are there in the whole grade?"

[0095] S205. Generate a third query statement based on the second query statement, where the third query statement is a statement recognizable by the target database represented in database language.

[0096] In the embodiment of the present application, the third query statement may be a statement sequence, such as an SQL statement sequence. After the computer device performs a preliminary screening on the first query statement to obtain the second query statement, it may generate a third query statement based on the second query statement. Specifically, the computer device may splice the second query statement with the field names of each data table in the database formatting information, and the spliced information may be predicted using a model to obtain a third query statement represented by an SQL statement sequence.

[0097] Exemplarily, after the second query statement "How many boys are there in the whole grade?" in the foregoing example is spliced with the field names of each data table in the target database formatting information, the spliced information may be expressed as:

[0098] How many boys are there in the whole grade?|Table_name|age: age, gender: gender, column3: grade table2_name|......

[0099] The above spliced information may be input into the Text2SQL model for model prediction. The output of the Text2SQL model is the SQL statement sequence corresponding to the above spliced information, as well as the probability distribution of each token in the SQL statement sequence. A token is the basic building block of text processing, which represents a discrete element in the text, and this element may be a word, character, sub-word, or character, etc. Of course, using the above Text2SQL model to convert the spliced information into the corresponding SQL statement sequence is only a possible implementation manner of the embodiment of the present application. In practical applications, the spliced information may also be input into other models that can perform statement conversion. For example, the pre-trained language model MT5 may also be used to convert the spliced information, and the corresponding SQL statement can also be obtained. The embodiment of the present application does not limit the statement conversion model used.

[0100] S206. Calculate the entropy of each token in the third query statement.

[0101] In the embodiment of the present application, according to the entropy calculation formula, the entropy of each token in the third query statement is calculated as:

[0102] -Σp*log(p)

[0103] Among them, p is the probability distribution of the corresponding token.

[0104] S207. Filter the third query statement according to the entropy of each of the said tokens to obtain a target query statement.

[0105] After calculating the entropy of each token in the third query statement, the third query statement can be further filtered to obtain a target query statement.

[0106] In a possible implementation manner of the embodiment of the present application, for any query statement included in the third query statement, the average value of the entropy of each token in the query statement can be determined; if the average value of the entropy of each token in the query statement is greater than a second preset threshold, the query statement can be deleted.

[0107] Specifically, after calculating the entropy of each token in the third query statement, the computer device can take the average of the entropy of all tokens to obtain an average value, which represents the uncertainty of the entire SQL statement sequence. The computer device can set a threshold S. When the average value of the above entropy is greater than the threshold S, it can be considered that the uncertainty of the model output is relatively large. Therefore, it can be determined that the input is an unanswerable input, and no SQL sequence is output, further filtering out relatively difficult-to-judge irrelevant inputs. The above threshold S can be set artificially according to the actually required filtering effect, or set according to empirical values.

[0108] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0109] For the convenience of understanding, a specific example is given below to introduce the natural language conversion method provided by the embodiment of the present application.

[0110] As Figure 3 shown, it is a schematic diagram of a natural language conversion process provided by the embodiment of the present application. According to the Figure 3 shown conversion process, the computer device can interact with the user and receive a query statement input by the user in natural language. For example, the query statement is "Hello, how many boys are there in the whole grade?" The above query statement can be input by the user in text form in the interaction interface of the computer device and transmitted to the computer device, or can be transmitted to the computer device in the form of voice information by the user through the audio acquisition module provided by the computer device. If the query statement is transmitted in the form of voice information, the computer device can convert the received voice information into text form.

[0111] Figure 3The irrelevant input filtering module in [the above] can be the preprocessing module in the foregoing embodiments. The irrelevant input filtering module in the computer device can be used to handle the received query statement. It should be noted that naming the module for processing the query statement in the computer device as the irrelevant input filtering module or the preprocessing module is merely an example, and the name of the module can be determined according to actual needs, and the embodiments of the present application do not limit this.

[0112] As Figure 3 shown, on the one hand, the irrelevant input filtering module can interact with the database to obtain the formatted information of the database. On the other hand, the irrelevant input filtering module can receive the query statement input by the user and process it.

[0113] Among them, the irrelevant input filtering module can sort out the formatted information of the database according to the specific data stored in the database. Specifically, the irrelevant input filtering module can first identify the column names of each data table stored in the database and the data values corresponding to the column names, and then generate the formatted information of the database according to the column names and their corresponding data values.

[0114] By sorting out the database information, a possible representation of the formatted information of the database can be obtained as:

[0115] Table_name: Explanation of the table name of Data Table 1

[0116] age: Age,

[0117] gender: Gender ({1: male, 0: female}),

[0118] column3: Grade ([‘First grade’, ’Second grade’, ’Third grade’, ’Fourth grade’, ’Fifth grade’, ’Sixth grade’])

[0119] Table2_name: Explanation of the table name of Data Table 2 ...

[0121] The processing of the query statement input by the user received by the irrelevant input filtering module focuses on how to determine the content in the statement that is significantly irrelevant to the database query.

[0122] In the embodiments of the present application, the first query statement can be divided into multiple statements, and the similarity between each statement and the formatted information of the database can be calculated respectively.

[0123] Specifically, the irrelevant input filtering module can convert each statement in the query statement into a first vector, and then convert the column names in the database and their corresponding data values into a second vector. In this way, multiple first vectors and multiple second vectors can be obtained. By calculating the cosine similarity between each first vector and each second vector, the similarity between each statement and the formatted information of the database can be obtained.

[0124] Since N similarities can be calculated for each statement, when all N similarities are less than a preset threshold, it can be considered that the statement has nothing to do with the database query, and then it can be deleted.

[0125] Exemplarily, for statement 1 "Hello" in the above example, the corresponding first vector 1 calculates the cosine similarity with 4 second vectors, and the obtained results are all less than the set first preset threshold. Therefore, it can be considered that this statement is significantly irrelevant to the target database, and the computer device can delete it. For statement 2 "How many boys are there in the whole grade?" in the above example, the corresponding first vector 2 calculates the cosine similarity with 4 second vectors. Among the obtained results, the cosine similarities between the first vector 2 and the second vector 1 and the second vector 2 are less than the set first preset threshold, but the cosine similarities between the first vector 2 and the second vector 3 and the second vector 4 are greater than the first preset threshold. Therefore, it can be considered that this statement is not significantly irrelevant to the target database, and the computer device can retain this statement. Therefore, the finally retained second query statement is "How many boys are there in the whole grade?".

[0126] As Figure 3 shown, the second query statement processed by the irrelevant input filtering module can be input into the Text2SQL model for model prediction.

[0127] The input information for model prediction can be obtained by concatenating the second query statement with the field names of each data table in the database formatted information. For example, the concatenated input information can be expressed as:

[0128] How many boys are there in the whole grade?|Table_name|age: age, gender: gender, column3: grade table2_name|......

[0129] The above concatenated information can be input into the Text2SQL model for model prediction, and an SQL statement sequence and the probability distribution of each token in this statement sequence are output.

[0130] As Figure 3 shown, the output of the Text2SQL model can be used as the input information for the entropy filtering module. Figure 3The entropy filtering module in it may be the post-processing module in the foregoing embodiments.

[0131] The entropy filtering module may calculate the entropy of each token according to the probability distribution of each token in the SQL statement sequence. After calculating the entropy of each token in the third query statement, the computer device may take the average of the entropies of all tokens to obtain an average value, and use this value to represent the uncertainty of the entire SQL statement sequence. The computer device may set a threshold S. When the average value of the above entropy is greater than the threshold S, it can be considered that the uncertainty of the model output is relatively large. Therefore, it can be determined that the input is an unanswerable input, and the output of the SQL sequence is not performed, further filtering out irrelevant inputs that are difficult to judge.

[0132] Referring to Figure 4 , a schematic diagram of a natural language conversion device provided by an embodiment of the present application is shown, which may specifically include a statement receiving module 401, a first filtering module 402, a statement generating module 403, and a second filtering module 404, where:

[0133] The statement receiving module 401 is configured to receive a first query statement for querying a target database input in natural language;

[0134] The first filtering module 402 is configured to filter out statements irrelevant to the target database in the first query statement to obtain a second query statement;

[0135] The statement generating module 403 is configured to generate a third query statement based on the second query statement, and the third query statement is a statement recognizable by the target database represented in database language;

[0136] The second filtering module 404 is configured to filter the third query statement to obtain a target query statement.

[0137] In a possible implementation manner of the embodiment of the present application, the first filtering module 402 may specifically be configured to:

[0138] Obtain formatting information of the target database, where the formatting information is information representing data stored in the target database in a preset standard format;

[0139] Calculate the similarity between each statement in the first query statement and the formatting information respectively;

[0140] Delete the statements corresponding to the target similarity less than the first preset threshold in the similarity to obtain a second query statement.

[0141] In the embodiment of the present application, the first filtering module 402 may further be used for:

[0142] Identify the column names of each data table stored in the target database and the data values corresponding to the column names;

[0143] Generate formatted information of the target database according to the column names and the corresponding data values.

[0144] In the embodiment of the present application, the first filtering module 402 can also be used for:

[0145] Convert each statement in the first query statement into a first vector;

[0146] Convert the column names and the corresponding data values into a second vector;

[0147] Calculate the cosine similarity between the first vector and the second vector.

[0148] In the embodiment of the present application, the first filtering module 402 can also be used for:

[0149] For any statement in the first query statement, if the cosine similarity between the first vector corresponding to the statement and any second vector is less than the first preset threshold, delete the statement to obtain a second query statement.

[0150] In a possible implementation manner of the embodiment of the present application, the second filtering module 404 can specifically be used for:

[0151] Calculate the entropy of each token in the third query statement;

[0152] Filter the third query statement according to the entropy of each token to obtain a target query statement.

[0153] In the embodiment of the present application, the second filtering module 404 can also be used for:

[0154] For any query statement included in the third query statement, determine the average value of the entropy of each token in the query statement;

[0155] When the average value of the entropy of each token in the query statement is greater than the second preset threshold, delete the query statement.

[0156] A natural language conversion device provided by an embodiment of the present application can implement each step in the foregoing method embodiments when applying this device.

[0157] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the description in the method embodiment part.

[0158] Refer to Figure 5, showing a schematic diagram of a computer device provided by an embodiment of the present application. As Figure 5 shown, the computer device 500 in the embodiment of the present application includes: a processor 510, a memory 520, and a computer program 521 stored in the memory 520 and executable on the processor 510. When the processor 510 executes the computer program 521, it implements the steps in each of the above embodiments of the natural language conversion method, such as Figure 1 the steps S101 to S104 shown. Alternatively, when the processor 510 executes the computer program 521, it implements the functions of each module / unit in each of the above device embodiments, such as Figure 4 the functions of the modules 401 to 404 shown.

[0159] Exemplarily, the computer program 521 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 520 and executed by the processor 510 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments can be used to describe the execution process of the computer program 521 in the computer device 500. For example, the computer program 521 can be divided into a statement receiving module, a first filtering module, a statement conversion module, and a second filtering module. The specific functions of each module are as follows:

[0160] The statement receiving module is used to receive a first query statement for querying a target database input in natural language;

[0161] The first filtering module is used to filter out the statements in the first query statement that are irrelevant to the target database to obtain a second query statement;

[0162] The statement conversion module is used to convert the second query statement into a third query statement, and the third query statement is a statement recognizable by the target database represented in database language;

[0163] The second filtering module is used to filter the third query statement to obtain a target query statement.

[0164] The computer device 500 can be an electronic device capable of implementing each step in the foregoing method embodiments. The computer device 500 can be a desktop computer, a cloud server, and other devices. The computer device 500 may include, but is not limited to, a processor 510 and a memory 520. Those skilled in the art can understand, Figure 5This is merely an example of the computer device 500, and does not constitute a limitation on the computer device 500. It may include more or fewer components than those shown in the figure, or combine some components, or have different components. For example, the computer device 500 may also include input / output devices, network access devices, buses, etc.

[0165] The processor 510 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0166] The memory 520 may be an internal storage unit of the computer device 500, such as the hard disk or memory of the computer device 500. The memory 520 may also be an external storage device of the computer device 500, such as a plug-in hard disk equipped on the computer device 500, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 520 may also include both the internal storage unit and the external storage device of the computer device 500. The memory 520 is used to store the computer program 521 and other programs and data required by the computer device 500. The memory 520 may also be used to temporarily store data that has been output or is to be output.

[0167] The embodiments of the present application also disclose a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the natural language conversion method described in the foregoing various embodiments is implemented.

[0168] The embodiments of the present application also disclose a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the natural language conversion method described in the foregoing various embodiments is implemented.

[0169] The embodiments of the present application also disclose a computer program product. When the computer program product runs on a computer, the computer is caused to execute the natural language conversion method described in each of the foregoing embodiments.

[0170] The above-described embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit the same. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A natural language conversion method, characterized in that, Including: Receiving a first query statement for querying a target database, which is input in natural language; Filtering statements in the first query statement that are irrelevant to the target database to obtain a second query statement; Generating a third query statement based on the second query statement, where the third query statement is a statement recognizable by the target database represented in database language; Filtering the third query statement to obtain a target query statement.

2. The method according to claim 1, wherein The filtering the statements in the first query statement that are irrelevant to the target database to obtain a second query statement includes: Obtaining formatting information of the target database, where the formatting information is information representing the data stored in the target database in a preset standard format; Calculating the similarity between each statement in the first query statement and the formatting information respectively; Deleting the statements corresponding to the target similarity with a similarity less than a first preset threshold to obtain a second query statement.

3. The method according to claim 2, wherein The obtaining the formatting information of the target database includes: Identifying the column names of each data table stored in the target database and the data values corresponding to the column names; Generating the formatting information of the target database according to the column names and the corresponding data values.

4. The method according to claim 2, characterized in that The calculating the similarity between each statement in the first query statement and the formatting information respectively includes: Converting each statement in the first query statement into a first vector; Converting the column names of each data table and the corresponding data values into a second vector; Calculating the cosine similarity between the first vector and the second vector.

5. The method according to claim 4, wherein The deleting the statements corresponding to the target similarity with a similarity less than a first preset threshold to obtain a second query statement includes: For any statement in the first query statement, if the cosine similarity between the first vector corresponding to the statement and any second vector is less than the first preset threshold, then deleting the statement to obtain a second query statement.

6. The method according to any one of claims 1-5, characterized in that, The filtering the third query statement to obtain a target query statement includes: Calculating the entropy of each token in the third query statement; Filtering the third query statement according to the entropy of each token to obtain a target query statement.

7. The method according to claim 6, characterized in that, The filtering the third query statement according to the entropy of each token to obtain a target query statement includes: Determining the average value of the entropy of each token in the query statement for any query statement included in the third query statement; When the average value of the entropy of each token in the query statement is greater than a second preset threshold, deleting the query statement.

8. A natural language conversion device, characterized in that, Including: A statement receiving module, configured to receive a first query statement for querying a target database, which is input in natural language; A first filtering module, configured to filter statements in the first query statement that are irrelevant to the target database to obtain a second query statement; A statement generating module, configured to generate a third query statement based on the second query statement, where the third query statement is a statement recognizable by the target database represented in database language; A second filtering module, configured to filter the third query statement to obtain a target query statement.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the natural language conversion method according to any one of claims 1-7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the natural language conversion method according to any one of claims 1-7 is implemented.