A natural language query retrieval system based on deep learning
Patent Information
- Application Number
- CN202610487549.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-14
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-04-14
AI Technical Summary
其中,基于规则或模板的方法通常依赖人工预先定义自然语言到数据库字段及查询语句之间的映射规则,该类方法在规则覆盖范围内能够实现较为稳定的查询效果,但在面对复杂、多样化的自然语言表达时,往往需要频繁维护规则集合,扩展性和通用性较差
在本发明所提出的基于深度学习的自然语言查询检索系统中,通过对目标结构化数据库的元数据信息进行系统化获取和建模,并结合字段名称的命名模式分析以及字段取值样本的特征信息,构建字段级语义表示,使系统能够从结构和数据两个层面对数据库字段语义进行综合理解。相较于现有技术中仅依赖字段名称进行语义匹配的方式,本发明能够有效缓解字段命名不规范、缩写、业务化命名等情况带来的语义歧义问题,从而提升自然语言查询与数据库字段之间的匹配准确性,为后续查询解析和字段选择奠定稳定的语义基础。
Smart Images

Figure CN122332526B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language interaction, and more particularly to a natural language query and retrieval system based on deep learning. Background Technology
[0002] With the widespread application of information systems and data management systems, enterprises, institutions, and various data service platforms commonly use structured databases to store and manage business data. Structured databases typically organize data in the form of tables, fields, and relationships, enabling efficient support for the storage, retrieval, and computation of large-scale data. However, in practical applications, ordinary users often lack the professional knowledge of structured query languages, making it difficult to directly write structured query statements for flexible and efficient data retrieval and analysis. Therefore, how to query structured databases using natural language has become an important research direction in the field of data interaction.
[0003] Existing technologies for natural language query retrieval mainly include rule-based methods, template-based methods, and deep learning-based methods that have been gradually developing in recent years. Rule-based or template-based methods typically rely on manually pre-defined mapping rules between natural language and database fields and query statements. While these methods can achieve relatively stable query results within the scope of the rules, they often require frequent maintenance of rule sets when dealing with complex and diverse natural language expressions, resulting in poor scalability and generality. Furthermore, these methods are difficult to adapt to changes in database structure; once field names or table structures are adjusted, the original rules often become invalid, leading to high maintenance costs. Summary of the Invention
[0004] One objective of this invention is to propose a natural language query and retrieval system based on deep learning. This invention fully utilizes deep learning technology, natural language processing technology, and structured data modeling technology to jointly model metadata information, field value characteristics, and historical query mapping information of structured databases, constructing field-level semantic representations, and achieving accurate mapping between natural language queries and structured query statements. By semantically encoding natural language queries and combining query context information to perform joint disambiguation processing on candidate fields, this invention can accurately determine the target field set in complex query scenarios with multiple fields and constraints, and generate structured query statements based on constraint graphs. This invention possesses advantages such as strong field semantic understanding capabilities, high accuracy in query ambiguity resolution, strong adaptability to complex query conditions, and good system scalability.
[0005] A natural language query and retrieval system based on deep learning according to an embodiment of the present invention includes: The metadata acquisition module is used to acquire metadata information from the target structured database. The field semantic generation module is used to perform naming pattern analysis on field names based on metadata information and generate field-level semantic representations by combining field value samples. The semantic card generation module is used to update the field-level semantic representation and encapsulate it into semantic cards based on the mapping records of historical natural language queries and structured query statements during system operation. The query semantics generation module is used to receive natural language queries input by users and generate query-level semantic representations; The field selection module is used to calculate the similarity value between the query-level semantic representation and the corresponding field-level semantic representation of the semantic card to generate a candidate field set, and to perform joint disambiguation processing on the fields in the candidate field set in combination with the context information in the natural language query to generate the target field set. The constraint construction module is used to extract time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions from natural language queries and map them to the target field set to construct a query constraint graph. The statement generation and validation module is used to generate structured query statements based on query constraint graphs and perform semantic consistency validation. The execution and clarification module is used to execute structured query statements to obtain retrieval results when the semantic consistency check passes, and to trigger an interactive clarification process and update the query constraint graph based on the clarification results when the semantic consistency check fails, thereby generating updated structured query statements.
[0006] Optionally, modules can be integrated using the following methods: Obtain metadata information from the target structured database; The naming pattern of field names in metadata information is analyzed, and combined with field value samples, a field-level semantic representation is generated. Based on the mapping records of historical natural language queries and structured query statements during system operation, update the field-level semantic representation and encapsulate it into semantic cards; Receive natural language queries input by users, preprocess the natural language queries, perform context modeling and semantic encoding, and generate query-level semantic representations; The similarity value is calculated based on the query-level semantic representation and the corresponding field-level semantic representation of the semantic card to generate a candidate field set. Then, combined with the context information in the natural language query, the fields in the candidate field set are subjected to joint disambiguation processing to generate the target field set. Extract time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions from natural language queries, and map them to the target field set to construct a query constraint graph; A structured query statement is generated based on the query constraint graph, and semantic consistency verification is performed on the structured query statement. When the semantic consistency check passes, the structured query statement is executed to obtain the retrieval results. When the semantic consistency check fails, the interactive clarification process is triggered, and the query constraint graph is updated based on the clarification results to generate the updated structured query statement.
[0007] Optionally, the system directory table of the database is accessed through a read-only database connection to read the table structure information, field names, field data types, primary key relationships, and foreign key relationships of the target structured database. Field value samples are obtained through a restricted query method. Under the read-only database connection condition, the restricted query method performs controlled data reading operations on the fields in the target structured database. The controlled data reading operations include sampling and reading field values, and unique value statistics. The table structure information, field names, field data types, primary key relationships, foreign key relationships, and field value samples are uniformly organized and encapsulated to form metadata information.
[0008] Optionally, the generation of the field-level semantic representation specifically includes: The field names in the metadata information are standardized, and the standardized field names are segmented into sub-words to obtain a sequence of field name sub-words; Naming pattern analysis is performed based on the field name sub-word sequence to generate a field name pattern feature vector; Generate field type feature vectors based on the field data types in the metadata information; When the field value sample is numerical, statistical processing is performed on the field value sample to generate a distribution feature vector; When the field value sample is an enumeration type, perform statistical processing on the set of enumeration values in the field value sample to generate an enumeration feature vector; The field name sub-word sequence is matched with the word set in the external knowledge base to generate an external knowledge base matching feature vector. The external knowledge base refers to the word set that provides general vocabulary semantic association information. The feature fusion process is performed on the field name pattern feature vector, field type feature vector, distribution feature vector, enumeration feature vector, and external knowledge base matching feature vector to generate the field-level semantic representation of the corresponding field.
[0009] Optionally, the formation of the semantic card specifically includes: Obtain the mapping records between historical natural language queries and corresponding structured query statements generated during system operation; Parse the structured query statement in the mapping record to obtain the fields referenced in the structured query statement; Generate corresponding query-level semantic representations based on historical natural language queries; Based on the referenced fields and their corresponding query-level semantic representations in structured query statements, the association features between fields and query semantics are generated; Based on the association features, the field-level semantic representation of the corresponding field is updated to generate the updated field-level semantic representation; The updated field-level semantic representation is associated and encapsulated with the corresponding field to form a semantic card.
[0010] Optionally, the generation of the query-level semantic representation specifically includes: Receive natural language query text input by the user, perform normalization processing on the natural language query text, and obtain natural language query text in a unified representation form; Generate a term sequence based on the normalized natural language query text; Based on the term sequence, a corresponding vector representation is generated for each term, forming a term vector sequence; Context modeling is performed based on the term vector sequence to obtain a context representation sequence of the contextual relationships between terms; The context representation sequence is summarized to generate the corresponding query-level semantic representation.
[0011] Optionally, the generation of the target field set specifically includes: Similarity values are calculated based on query-level semantic representation and the corresponding field-level semantic representation of semantic cards; Based on similarity values, the corresponding fields in the semantic cards are filtered to generate a set of candidate fields; For each field in the candidate field set, calculate the context consistency value based on the context information in the natural language query; For each field in the candidate field set, a joint disambiguation score is calculated based on the similarity value and the contextual consistency value; The candidate field set is filtered based on the joint disambiguation score of each field to generate the target field set.
[0012] Optionally, the construction of the query constraint graph specifically includes: Identify time-related expressions from natural language queries and generate time constraints; Identify expressions related to numerical conditions from natural language queries and generate numerical threshold constraints; Identify the content of join conditions from natural language queries and generate logical relationship constraints; Identify sorting-related expressions from natural language queries and generate sorting criteria; The time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions are mapped to the corresponding fields in the target field set to construct a query constraint graph.
[0013] Optionally, a structured query statement is generated based on the target field set, time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions in the query constraint diagram. Semantic consistency checks are performed on the structured query statement. These checks include field usage consistency checks and logical constraint conflict checks. Field usage consistency checks involve sequentially traversing the fields in the structured query statement and comparing each traversed field with the target field set, marking fields not belonging to the target field set as inconsistent fields. Logical constraint conflict checks involve extracting the time constraints and numerical threshold constraints associated with each field in the structured query statement and comparing the numerical threshold constraints associated with the same field. When the value ranges of the numerical threshold constraints associated with the same field do not intersect, the same field is marked as a conflicting field. Logical constraint conflict checks involve combining and checking the numerical threshold constraints connected to the same field based on the logical relationship constraints in the structured query statement. When the value ranges obtained from the combined checks do not intersect, the same field is marked as a conflicting field.
[0014] Optionally, the update of the structured query statement specifically includes: When the semantic consistency check passes, a structured query is executed to obtain the retrieval results; When the semantic consistency check fails, retrieve the fields that were marked as inconsistent or conflicting during the semantic consistency check process; Based on inconsistent or conflicting fields, a clarification question is generated, and an interactive clarification process is triggered. The clarification question is output to the user and user feedback is received, resulting in a clarification result. Based on the clarification results, the time constraints, numerical threshold constraints, logical relationship constraints, or sorting conditions associated with the inconsistent or conflicting fields in the query constraint graph are updated to obtain the updated query constraint graph, which forms the updated structured query statement.
[0015] The beneficial effects of this invention are: In the deep learning-based natural language query and retrieval system proposed in this invention, the metadata information of the target structured database is systematically acquired and modeled. Combined with naming pattern analysis of field names and feature information from field value samples, a field-level semantic representation is constructed, enabling the system to comprehensively understand the semantics of database fields from both structural and data perspectives. Compared to existing technologies that rely solely on field names for semantic matching, this invention effectively alleviates semantic ambiguity caused by non-standard field naming, abbreviations, and business-specific naming, thereby improving the matching accuracy between natural language queries and database fields and laying a stable semantic foundation for subsequent query parsing and field selection.
[0016] This invention utilizes the mapping records between historical natural language queries and structured query statements accumulated during system operation to continuously update field-level semantic representations. These updated representations are then encapsulated as semantic cards, enabling field semantics to dynamically evolve with actual usage. Compared to existing technologies that use statically unchanged field semantics, this invention continuously corrects and enhances field semantic expression capabilities based on actual user query behavior. This improves the system's adaptability to changes in business semantics and user query habits, thereby maintaining high query accuracy and stability over long-term operation.
[0017] In the natural language query parsing stage, this invention semantically encodes the user-input natural language query to generate a query-level semantic representation. It then calculates a candidate field set based on the similarity between the query-level semantic representation and the field-level semantic representations in the semantic cards. Simultaneously, it incorporates contextual information to perform joint disambiguation processing on the candidate fields. Compared to existing methods that select fields based on a single similarity index, this invention comprehensively considers the relationship between the overall query semantics and contextual constraints. This effectively reduces the probability of incorrect field selection when multiple fields are semantically similar, thereby more accurately determining the target field set that matches the user's query intent.
[0018] Furthermore, based on the determined target field set, this invention further extracts time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions from natural language queries, and maps these constraints uniformly to the target field set to construct a query constraint graph. By structuring the relationship between fields and multiple types of constraints through the query constraint graph, this invention can maintain the integrity and consistency of query semantics in complex query scenarios, effectively avoiding the problems of lost, mismatched, or conflicting constraints, and providing a clear and verifiable intermediate representation for the generation of structured query statements.
[0019] By performing semantic consistency checks on the structured query statements generated from the query constraint graph, and introducing an interactive clarification mechanism when the check fails, this invention guides users to confirm and correct key fields or constraints when the system automatically parses ambiguities or conflicts. This dynamically updates the query constraint graph and regenerates the structured query statements. This mechanism ensures the accuracy of query results while avoiding the direct return of results that do not match the user's intent due to a single parsing error, significantly improving the reliability and user experience of natural language query retrieval systems in practical applications.
[0020] In summary, this invention, through field-level semantic modeling, dynamic updating of semantic cards, context-based joint disambiguation, and a query generation mechanism driven by query constraint graphs, effectively overcomes the shortcomings of existing technologies in complex natural language query scenarios, such as insufficient understanding of field semantics, limited ambiguity handling capabilities, and unstable parsing of multiple constraints. It can significantly improve the accuracy, stability, and applicability of natural language query retrieval, and has good engineering application value and promotion prospects. Attached Figure Description
[0021] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of a natural language query and retrieval system based on deep learning proposed in this invention; Figure 2 This is a schematic diagram illustrating the construction of semantic cards in a natural language query and retrieval system based on deep learning proposed in this invention. Figure 3 This is a schematic diagram illustrating the construction of a query constraint graph for a deep learning-based natural language query retrieval system proposed in this invention. Detailed Implementation
[0022] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0023] refer to Figures 1-3 A natural language query and retrieval system based on deep learning, comprising: The metadata acquisition module is used to acquire metadata information from the target structured database. The field semantic generation module is used to perform naming pattern analysis on field names based on metadata information and generate field-level semantic representations by combining field value samples. The semantic card generation module is used to update the field-level semantic representation and encapsulate it into semantic cards based on the mapping records of historical natural language queries and structured query statements during system operation. The query semantics generation module is used to receive natural language queries input by users and generate query-level semantic representations; The field selection module is used to calculate the similarity value between the query-level semantic representation and the corresponding field-level semantic representation of the semantic card to generate a candidate field set, and to perform joint disambiguation processing on the fields in the candidate field set in combination with the context information in the natural language query to generate the target field set. The constraint construction module is used to extract time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions from natural language queries and map them to the target field set to construct a query constraint graph. The statement generation and validation module is used to generate structured query statements based on query constraint graphs and perform semantic consistency validation. The execution and clarification module is used to execute structured query statements to obtain retrieval results when the semantic consistency check passes, and to trigger an interactive clarification process and update the query constraint graph based on the clarification results when the semantic consistency check fails, thereby generating updated structured query statements.
[0024] In this embodiment, the modules are interconnected using the following method: Obtain metadata information from the target structured database; The naming pattern of field names in metadata information is analyzed, and combined with field value samples, a field-level semantic representation is generated. Based on the mapping records of historical natural language queries and structured query statements during system operation, update the field-level semantic representation and encapsulate it into semantic cards; It receives natural language queries input by users, performs preprocessing, context modeling, and semantic encoding on the natural language queries, and generates query-level semantic representations; The similarity value is calculated based on the query-level semantic representation and the corresponding field-level semantic representation of the semantic card. A candidate field set is generated, and the fields in the candidate field set are jointly disambiguated by combining the context information in the natural language query to generate the target field set. Extract time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions from natural language queries, and map them to the target field set to construct a query constraint graph; A structured query statement is generated based on the query constraint graph, and semantic consistency is checked on the structured query statement. When the semantic consistency check passes, the structured query statement is executed to obtain the retrieval results. When the semantic consistency check fails, the interactive clarification process is triggered, and the query constraint graph is updated based on the clarification results to generate the updated structured query statement.
[0025] In this embodiment, the system directory table of the database is accessed through a read-only database connection to read the table structure information, field names, field data types, primary key relationships, and foreign key relationships of the target structured database. Field value samples are obtained through a restricted query method. Under the read-only database connection condition, the restricted query method performs controlled data reading operations on the fields in the target structured database. The controlled data reading operations include sampling and reading field values and unique value statistics. The table structure information, field names, field data types, primary key relationships, foreign key relationships, and field value samples are uniformly organized and encapsulated to form metadata information.
[0026] In this embodiment, the generation of field-level semantic representation specifically includes: The field names in the metadata information are standardized, and the standardized field names are segmented into sub-words to obtain a sequence of field name sub-words; The standardization process specifically involves: unifying the character format in field names, standardizing the use of uppercase and lowercase letters in field names, identifying symbols that separate semantic units in field names, standardizing the field name parts separated by underscores, hyphens, or delimiters, identifying abbreviations in field names and expanding or standardizing them in accordance with common field naming rules, and processing redundant characters in field names that do not have semantic distinguishing function, to obtain standardized field names. Naming pattern analysis is performed based on the field name sub-word sequence to generate a field name pattern feature vector; The naming pattern analysis process is as follows: Analyze the positional relationship of each word in the field name sub-word sequence within the field name, identifying the order and combination of words in the field name; statistically analyze the recurring word forms in the field name sub-word sequence to obtain the naming rules of words in the field name; combine the morphological features of words in the field name sub-word sequence to distinguish words representing time, quantity, identifier, or object meanings; after completing the above analysis, summarize the information of the field naming structure in the field name sub-word sequence and organize it into a field name pattern feature vector according to a unified dimension. Generate field type feature vectors based on the field data types in the metadata information; The generation process is as follows: classify and identify the data type of the field, distinguishing whether the field is numeric, character, time, or Boolean; after identifying the data type of the field, determine the usage characteristics of the field in numerical operations, character matching, or time comparison based on the type attributes corresponding to the data type of the field; encode and organize the type attributes corresponding to the data type of the field according to a unified dimension, and use the encoded and organized result as the field type feature vector; When the field value sample is numerical, statistical processing is performed on the field value sample to generate a distribution feature vector; The statistical processing steps are as follows: First, perform statistical processing on the field value sample set, calculating the central tendency and dispersion of the field value samples in the numerical distribution to obtain the mean and standard deviation of the field value samples. Second, perform sorting analysis on the field value sample set to determine multiple quantiles of the field value samples in the numerical distribution. Third, statistically analyze the unique values in the field value sample set to calculate the proportion of unique values in the field value sample set, obtaining the proportion of unique values. Finally, organize the mean, standard deviation, quantiles, and proportion of unique values of the field value samples in a unified order to form a distribution feature vector for segment numerical distribution characteristics. When the field value sample is an enumeration type, perform statistical processing on the set of enumeration values in the field value sample to generate an enumeration feature vector; The statistical processing is as follows: The values in the field value sample are traversed, and different enumeration values appearing in the field value sample are identified and summarized to form an enumeration value set; the occurrence of each enumeration value in the enumeration value set in the field value sample set is statistically analyzed, and the occurrence frequency of each enumeration value is recorded; after completing the statistical analysis of the occurrence of enumeration values, the number of enumeration values in the enumeration value set is statistically analyzed to obtain the number of enumeration values contained in the field value sample; the number of enumeration values and the distribution information of each enumeration value in the field value sample are organized in a unified order to form an enumeration feature vector. The field name sub-word sequence is matched with the term set in the external knowledge base to generate an external knowledge base matching feature vector. The external knowledge base refers to the term set that provides general vocabulary semantic association information. The matching process is as follows: A set of terms related to the semantics of the field is read from an external knowledge base. The terms in this set are then processed in a unified manner. Each term in the field name sub-word sequence is compared with each term in the term set to identify the correspondence between the sub-words in the field name sub-word sequence and the terms in the term set in terms of literal form or semantic association. After comparing the field name sub-words with the terms, successfully matched sub-words and terms are recorded, and the matching status between the field name sub-word sequence and the term set is statistically analyzed. After the matching statistics are completed, the matching results between the field name sub-word sequence and the external knowledge base term set are organized in a unified order to form an external knowledge base matching feature vector. The field name pattern feature vector, field type feature vector, distribution feature vector, enumeration feature vector, and external knowledge base matching feature vector are fused to generate a field-level semantic representation of the corresponding field. The feature fusion process is as follows: Dimensional alignment is performed on the field name pattern feature vector, field type feature vector, distribution feature vector, enumeration feature vector, and external knowledge base matching feature vector. After dimensional alignment, these feature vectors are combined in a unified order to obtain a fused feature vector. The fused feature vector is then normalized and weighted. The weighting coefficients are obtained from the field data type, statistical information of the field value samples, and the matching results of the external knowledge base matching feature vector. The weighted fused feature vector is then used as the field-level semantic representation of the corresponding field.
[0027] In this embodiment, the formation of semantic cards specifically includes: Obtain the mapping records between historical natural language queries and corresponding structured query statements generated during system operation; The specific implementation method is as follows: After the natural language query is processed and the corresponding structured query statement is successfully generated, the system associates and saves the original input content of the natural language query and the corresponding generated structured query statement to form a mapping record between the natural language query and the structured query statement. Parse the structured query statement in the mapping record to obtain the fields referenced in the structured query statement; The specific parsing process is as follows: The structured query text corresponding to the historical natural language query is read from the mapping record; the overall structure of the structured query is traversed to determine the positional relationship of each segment within the structured query; the field references in the structured query are identified, including field names appearing in query conditions, filter conditions, sorting conditions, and aggregate expressions; after identifying the field references, the field names appearing in the structured query are extracted, and the extracted field names are matched with fields in the target structured database to obtain the referenced fields in the structured query; Generate corresponding query-level semantic representations based on historical natural language queries; The generation process is as follows: The text content of historical natural language queries is read from the mapping records and formatted. Semantic unit identification is then performed on the formatted text content. Semantic unit identification refers to determining the semantic elements and their arrangement relationships contained in the historical natural language query. After semantic unit identification, the relationships between the semantic elements in the historical natural language query are analyzed to form a representation of the overall semantic structure of the historical natural language query. After the relationship analysis is completed, the representation results are organized into a vector-based data representation, serving as the query-level semantic representation corresponding to the historical natural language query. Based on the referenced fields and their corresponding query-level semantic representations in structured query statements, the association features between fields and query semantics are generated; The generation process is as follows: For each structured query statement in the mapping record, based on the statement structure, determine the position of each referenced field in the structured query statement, and distinguish the usage of the field in the filtering conditions, sorting conditions, aggregation expressions, or field selection parts; based on the usage of the field in the structured query statement, associate the query-level semantic representation with the field to form a data representation of the correspondence between the field and the query-level semantic representation; after completing the correspondence, measure the data representation of the correspondence to obtain a numerical result of the degree of association between the field and the query-level semantic representation; and organize the numerical result and the usage of the field in the structured query statement together as the association feature between the field and the query semantics. Based on the association features, the field-level semantic representation of the corresponding field is updated to generate the updated field-level semantic representation; The update process is as follows: read the field-level semantic representation corresponding to the referenced field in the structured query statement, and read the association features corresponding to the field; adjust the field-level semantic representation according to the degree of association between the field and the query semantics reflected by the association features. The adjustment process refers to introducing semantic association information corresponding to the association features during the update process of the field-level semantic representation; after the adjustment process of the field-level semantic representation is completed, output the adjusted field-level semantic representation as the updated field-level semantic representation. The updated field-level semantic representation is associated and encapsulated with the corresponding field to form a semantic card.
[0028] In this embodiment, the generation of query-level semantic representation specifically includes: Receive natural language query text input by the user, perform normalization processing on the natural language query text, and obtain natural language query text in a unified representation form; The standardization process is as follows: First, the character formats in the natural language query text are standardized; second, the uppercase and lowercase characters appearing in the natural language query text are consistent; third, the symbols separating semantic components in the natural language query text are identified, and the text content separated by the symbols is standardized; fourth, after standardizing the symbols, consecutive whitespace characters appearing in the natural language query text are standardized; and finally, the standardized natural language query text is used as a natural language query text with a unified representation. Generate a term sequence based on the normalized natural language query text; The generation process is as follows: the normalized natural language query text is sequentially traversed, and the basic language units composed of adjacent characters are identified based on the arrangement relationship of characters in the text; the basic language units are divided according to the spacing relationship between characters in the natural language query text, and the natural language query text is split into multiple terms with independent semantic meanings; after the term division is completed, the terms are arranged according to the order in which they appear in the natural language query text, forming a term sequence composed of multiple terms in sequence; Based on the term sequence, a corresponding vector representation is generated for each term, forming a term vector sequence; The specific implementation method is as follows: each term in the term sequence is processed sequentially according to its order in the term sequence; for each term, a vector representation of the semantic features of the term is generated based on the character composition information of the term and its usage characteristics in natural language; after the vector representation of each term in the term sequence is generated, the vector representations corresponding to each term are arranged according to the original order of the term in the term sequence to form a term vector sequence composed of multiple term vectors in sequence; Context modeling is performed based on the term vector sequence to obtain a context representation sequence of the contextual relationships between terms; The specific process of context modeling is as follows: The word vector sequence is sequentially traversed according to the order of the words in the word sequence; during the traversal, the word vector corresponding to the current word is combined with the word vectors corresponding to the preceding and following words to comprehensively analyze the relationship between the current word and its adjacent words, thus determining the contextual relationship of the current word in the word sequence; after analyzing the contextual relationship of each word, the information of the contextual relationship of each word in the word sequence is organized into corresponding vector representations; after processing all words in the word sequence, the vector representations corresponding to each word are arranged according to their original order in the word sequence to form a contextual representation sequence. The context representation sequence is summarized to generate the corresponding query-level semantic representation.
[0029] In this embodiment, the generation of the target field set specifically includes: Similarity values are calculated based on query-level semantic representation and the corresponding field-level semantic representation of semantic cards; The calculation process is as follows: obtain the query-level semantic representation, and sequentially read the field-level semantic representations corresponding to the fields in each semantic card; for the current semantic card, perform a dimension-by-dimensional comparison between the query-level semantic representation and the field-level semantic representation corresponding to the semantic card, and comprehensively calculate the corresponding values of the query-level semantic representation and the field-level semantic representation in each dimension to obtain the numerical calculation result; after completing the dimension-by-dimensional comparison process, perform scale consistency processing on the numerical calculation result to form the similarity value of the field corresponding to the current semantic card; Based on similarity values, the corresponding fields in the semantic cards are filtered to generate a set of candidate fields; The filtering process involves obtaining similarity values for fields in each semantic card, establishing a one-to-one correspondence between similarity values and corresponding fields, comparing the similarity values for each field, and sorting the fields in descending order of similarity value. After sorting, fields are filtered based on similarity values, with fields having similarity values greater than or equal to a preset similarity threshold being selected into the candidate field set. The similarity threshold refers to the minimum degree of similarity between a specified field and the query-level semantic representation. Fields with similarity values less than the preset similarity threshold are removed, forming the candidate field set. For each field in the candidate field set, calculate the context consistency value based on the context information in the natural language query; The calculation process involves retrieving contextual information from the natural language query. This contextual information includes constraint fragments, modifier fragments, and pointer fragments related to field selection within the natural language query. For each field in the candidate field set, the field name and corresponding field value samples in the target structured database are read, and the field name is compared with the contextual information. After completing the comparison between the field name and contextual information, the consistency of the constraint fragments in the natural language query and the field value samples is verified to determine the degree of matching between the value format, value range, or value category appearing in the natural language query and the field value samples. After completing the above comparison and consistency verification, the results of the correspondence between the field name and contextual information, as well as the results of the consistency verification between the field value samples and contextual information, are summarized to obtain the contextual consistency value corresponding to the field. For each field in the candidate field set, a joint disambiguation score is calculated based on the similarity value and the contextual consistency value; The calculation process is as follows: For each field in the candidate field set, obtain the similarity value and the context consistency value corresponding to the field, and establish a one-to-one correspondence between the field and the similarity value, and between the field and the context consistency value. Perform numerical range consistency processing on the similarity value and the context consistency value. After completing the numerical range consistency processing, perform joint calculation on the similarity value and the context consistency value. The joint calculation process includes summing the similarity value and the context consistency value and averaging the summing results to obtain the joint disambiguation score corresponding to the field. The candidate field set is filtered based on the joint disambiguation score corresponding to each field to generate the target field set; The filtering process is as follows: the fields in the candidate field set are sorted in descending order of their joint disambiguation scores; after sorting, filtering is performed based on the joint disambiguation scores, which includes adding fields with joint disambiguation scores greater than or equal to the target filtering threshold to the target field set and removing fields with joint disambiguation scores less than the target filtering threshold from the target field set; after completing the above filtering process, the target field set is output.
[0030] In this embodiment, the construction of the query constraint graph specifically includes: Identify time-related expressions from natural language queries and generate time constraints; The identification process is as follows: sequentially scan the natural language query to locate the expression content describing the time range, time point, or time period in the natural language query; parse the time-related expression content to determine the start and end positions of the time indicated by the expression content; after determining the start and end positions of the time, combine the start and end positions of the time to form a time constraint that limits the time range. Identify expressions related to numerical conditions from natural language queries and generate numerical threshold constraints; The identification process is as follows: the natural language query is scanned segment by segment to locate the expression content in the natural language query that describes the quantity, range limitation or numerical comparison; the expression content related to numerical conditions is parsed to obtain the numerical information contained therein and the comparison relationship between numerical values; the numerical information and the comparison relationship are combined to form the numerical threshold constraint that limits the numerical conditions. Identify the content of join conditions from natural language queries and generate logical relationship constraints; The identification process is as follows: perform sequential analysis on the natural language query to locate the expression content used to connect multiple conditions in the natural language query; parse the expression content of the connection conditions to distinguish the occurrence position of different connection methods in the natural language query and their corresponding connection order; after completing the parsing of the expression content of the connection conditions, generate logical relationship constraints describing the connection methods between multiple conditions based on the arrangement relationship of each connection condition in the natural language query. Identify sorting-related expressions from natural language queries and generate sorting criteria; The identification process is as follows: perform sequential analysis on the natural language query to locate the expression content in the natural language query that describes the order of the results; parse and process the expression content related to sorting to determine the sorting basis and sorting order indicated in the natural language query; after determining the sorting basis and sorting order, combine the sorting basis and sorting order to form sorting conditions that describe the sorting method of the results. Map time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions to the corresponding fields in the target field set to construct a query constraint graph; The construction process is as follows: For time constraints, based on the semantic correspondence between time-related expressions in the natural language query and fields in the target field set, the fields corresponding to the time constraints are determined, and a correspondence between time constraints and fields is established; for numerical threshold constraints, based on the semantic correspondence between numerical condition expressions in the natural language query and fields in the target field set, the fields corresponding to the numerical threshold constraints are determined, and a correspondence between numerical threshold constraints and fields is established; for sorting conditions, based on the semantic correspondence between sorting-related expressions in the natural language query and fields in the target field set, the fields corresponding to the sorting conditions are determined, and a correspondence between sorting conditions and fields is established; after establishing the correspondence between time constraints, numerical threshold constraints, and sorting conditions and fields, the connection order between time constraints and numerical threshold constraints is organized according to logical relationship constraints, and logical relationship constraints are associated with the correspondences; after completing the above mapping and organization, the target field set, time constraints, numerical threshold constraints, logical relationship constraints, sorting conditions, and their respective correspondences with fields are uniformly encapsulated to form a query constraint graph.
[0031] In this embodiment, a structured query statement is generated based on the target field set, time constraints, numerical threshold constraints, logical relation constraints, and sorting conditions in the query constraint diagram. Semantic consistency checks are performed on the structured query statement, including field usage consistency checks and logical constraint conflict checks. Field usage consistency checks involve sequentially traversing the fields in the structured query statement and comparing each traversed field with the target field set, marking fields not belonging to the target field set as inconsistent fields. Logical constraint conflict checks involve extracting the time constraints and numerical threshold constraints associated with each field in the structured query statement and comparing the numerical threshold constraints associated with the same field. When the value ranges of the numerical threshold constraints associated with the same field do not overlap, the same field is marked as a conflicting field. Logical constraint conflict checks involve combining and checking the numerical threshold constraints connected to the same field based on the logical relation constraints in the structured query statement. When the value ranges obtained from the combined checks do not overlap, the same field is marked as a conflicting field.
[0032] In this embodiment, updating the structured query statement specifically includes: When the semantic consistency check passes, a structured query is executed to obtain the retrieval results; The acquisition process is as follows: a structured query statement that has passed semantic consistency verification is sent to the target structured database. The target structured database parses the structured query statement and executes a query operation based on the fields, time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions contained in the structured query statement. After completing the query operation, the target structured database returns query result data that matches the structured query statement. After receiving the query result data, the query result data is processed to form the retrieval results. When the semantic consistency check fails, retrieve the fields that were marked as inconsistent or conflicting during the semantic consistency check process; Based on inconsistent or conflicting fields, a clarification question is generated, and an interactive clarification process is triggered to output the clarification question to the user and receive user feedback information, thus forming a clarification result; The clarification results are generated as follows: For each inconsistent or conflicting field, read the associated time constraints, numerical threshold constraints, logical relationship constraints, or sorting conditions; after reading the associated constraints, analyze each of them to determine the constraint content that caused the field to be marked as inconsistent or conflicting; after determining the constraint content, construct clarification questions based on the field and its associated constraint content to confirm the constraint values, constraint relationships, or sorting methods of the field; when there are multiple inconsistent or conflicting fields, generate corresponding clarification questions for each inconsistent or conflicting field, and collect the generated clarification questions to form a clarification question set; after the clarification question set is generated, trigger the interactive clarification process, output the clarification question set to the user and receive user feedback information, and use the user feedback information as the clarification result; Based on the clarification results, the time constraints, numerical threshold constraints, logical relationship constraints, or sorting conditions associated with inconsistent or conflicting fields in the query constraint graph are updated to obtain the updated query constraint graph, which forms the updated structured query statement. The update process is as follows: When the clarification result indicates that at least one of the time constraint, numerical threshold constraint, logical relation constraint, or sorting condition needs to be corrected, the time constraint, numerical threshold constraint, logical relation constraint, or sorting condition associated with the inconsistent or conflicting field in the query constraint graph is replaced with the corresponding constraint content in the clarification result; when the clarification result indicates that the time constraint, numerical threshold constraint, logical relation constraint, or sorting condition associated with the inconsistent or conflicting field is invalid, the corresponding time constraint, numerical threshold constraint, logical relation constraint, or sorting condition associated with the inconsistent or conflicting field in the query constraint graph is deleted; when the clarification result indicates that a new constraint is added, the corresponding time constraint, numerical threshold constraint, logical relation constraint, or sorting condition is added to the query constraint graph for the inconsistent or conflicting field; after completing the above update process, the association between the inconsistent or conflicting field and its associated time constraint, numerical threshold constraint, logical relation constraint, and sorting condition in the query constraint graph is reorganized to form an updated query constraint graph.
[0033] Example 1: In a certain enterprise-level data analysis platform, the underlying layer uses a relational structured database to store business data. The database contains multiple business tables, such as an order table, a customer table, and a sales record table. These tables have a large number of fields, and the field names clearly reflect business characteristics, such as "cust_id", "order_amt", "sale_region", and "created_dt". The main users of this platform are business personnel and managers, who lack the ability to write structured query languages and typically rely on graphical interfaces or fixed reports to obtain data. In actual use, business personnel frequently need to submit ad-hoc, combined data query requests using natural language, such as "Query customers in East China whose order amounts exceeded 100,000 last month, and sort them by amount from highest to lowest". In the existing system, such queries often require technical personnel to manually translate the natural language requests into structured query statements, which is inefficient and prone to inaccurate results due to misunderstandings.
[0034] In this application scenario, the deep learning-based natural language query and retrieval system proposed in this invention is introduced. The system first acquires metadata information from the target structured database in read-only mode, analyzing the table structure, field names, field data types, and field value samples. Based on this, the system performs naming pattern analysis on the field names and, combined with the distribution characteristics of the field value samples, generates field-level semantic representations. For example, for the "order_amt" field, the system can accurately model its semantics as "order amount" by combining the monetary meaning in its field name and the continuous numerical distribution characteristics of the field value samples; for the "created_dt" field, the system can accurately identify it as a time field by combining the time indication characteristics in its field name and the time format characteristics of the field value samples.
[0035] During system operation, the platform gradually accumulates a large number of mapping records between historical natural language queries and structured query statements. For example, "order amount greater than 50,000" corresponds to "order_amt>50000", and "last week" corresponds to "created_dt BETWEEN …". Based on these mapping records, the system dynamically updates the field-level semantic representation and encapsulates the updated field-level semantic representation into semantic cards, enabling the field semantics to continuously adapt to actual business usage scenarios. This process effectively solves the problem of static field semantics in existing technologies, which are difficult to adapt to business evolution.
[0036] When a business user inputs a natural language query, "Query customers in East China whose order amount exceeded 100,000 last month, and sort them by amount from highest to lowest," the system performs semantic encoding on the natural language query, generating a query-level semantic representation. Based on the similarity between this query-level semantic representation and the field-level semantic representations in the semantic cards, it calculates and generates a set of candidate fields. During this process, the system can simultaneously recognize contextual information such as "order amount," "region," "time," and "sorting," and through a joint disambiguation mechanism, accurately distinguishes whether "customer" refers to the customer identifier field rather than the contact person field, thus generating an accurate set of target fields.
[0037] After determining the target field set, the system further extracts time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions from the natural language query, and maps these constraints to the target field set to construct a query constraint graph. Through the structured representation of the relationship between fields and constraints using the query constraint graph, the system can ensure the integrity of the query semantics. After generating the structured query statement, the system performs semantic consistency checks on the structured query statement. When potential conflicts are found in field usage or constraint combinations, an interactive clarification process is triggered, guiding the user to confirm key conditions, thereby avoiding the return of query results that do not match the user's true intent.
[0038] In this embodiment, the method of the present invention was compared with the platform's existing natural language query module based on rules and simple semantic matching. Common natural language queries in real business environments were selected as test samples, covering complex query scenarios with multiple fields and multiple constraints. The system performance was evaluated by statistically analyzing field matching accuracy, structured query statement generation success rate, and user first-time query success rate.
[0039] Table 1. Comparison of Natural Language Query and Retrieval Results
[0040] As shown in Table 1, under the same natural language query test sample conditions, the proposed solution significantly outperforms existing technologies in both field matching accuracy and multi-constraint query parsing success rate. This demonstrates that through field-level semantic representation, dynamic updating of semantic cards, and a context-based joint disambiguation mechanism, the system can more accurately understand the intent of natural language queries. The proposed solution also exhibits significant advantages in structured query statement generation success rate and first-time query success rate, indicating that the query constraint graph and semantic consistency verification mechanism effectively reduce parsing errors and ambiguities in complex query scenarios.
[0041] Meanwhile, it can be seen that the average query response time of the present invention is basically the same as that of existing technologies, indicating that while improving query accuracy, it does not significantly increase system computational overhead, demonstrating good engineering usability. The proportion of queries requiring manual intervention is significantly reduced, reflecting that the present invention can effectively reduce reliance on technical personnel in real-world business environments, improving data query efficiency and user experience.
[0042] As can be seen from the above embodiments, the natural language query and retrieval system based on deep learning proposed in this invention can effectively solve the problems of insufficient field semantic understanding, limited ability to resolve contextual ambiguity, and unstable parsing of multiple constraints in existing technologies under real business data environments and complex natural language query scenarios. It significantly improves the accuracy and practicality of natural language query and retrieval and has good application value.
[0043] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A deep learning-based natural language query retrieval system, characterized by, include: The metadata acquisition module is used to acquire metadata information from the target structured database. The field semantic generation module is used to perform naming pattern analysis on field names based on metadata information and generate field-level semantic representations by combining field value samples. The semantic card generation module is used to update the field-level semantic representation and encapsulate it into semantic cards based on the mapping records of historical natural language queries and structured query statements during system operation. The query semantics generation module is used to receive natural language queries input by users and generate query-level semantic representations; The field selection module is used to calculate the similarity value between the query-level semantic representation and the corresponding field-level semantic representation of the semantic card to generate a candidate field set, and to perform joint disambiguation processing on the fields in the candidate field set in combination with the context information in the natural language query to generate the target field set. The constraint construction module is used to extract time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions from natural language queries and map them to the target field set to construct a query constraint graph. The statement generation and validation module is used to generate structured query statements based on query constraint graphs and perform semantic consistency validation. The execution and clarification module is used to execute structured query statements to obtain retrieval results when the semantic consistency check passes, and to trigger an interactive clarification process and update the query constraint graph based on the clarification results when the semantic consistency check fails, thereby generating updated structured query statements. 2.The deep learning based natural language query retrieval system of claim 1, wherein, The modules are connected in the following way: Obtain metadata information from the target structured database; The naming pattern of field names in metadata information is analyzed, and combined with field value samples, a field-level semantic representation is generated. Based on the mapping records of historical natural language queries and structured query statements during system operation, update the field-level semantic representation and encapsulate it into semantic cards; Receive natural language queries input by users, preprocess the natural language queries, perform context modeling and semantic encoding, and generate query-level semantic representations; The similarity value is calculated based on the query-level semantic representation and the corresponding field-level semantic representation of the semantic card to generate a candidate field set. Then, combined with the context information in the natural language query, the fields in the candidate field set are subjected to joint disambiguation processing to generate the target field set. Extract time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions from natural language queries, and map them to the target field set to construct a query constraint graph; A structured query statement is generated based on the query constraint graph, and semantic consistency verification is performed on the structured query statement. When the semantic consistency check passes, the structured query statement is executed to obtain the retrieval results. When the semantic consistency check fails, the interactive clarification process is triggered, and the query constraint graph is updated based on the clarification results to generate the updated structured query statement.
3. The natural language query and retrieval system based on deep learning according to claim 2, characterized in that, By accessing the system directory tables of the database through a read-only database connection, the system reads the table structure information, field names, field data types, primary key relationships, and foreign key relationships of the target structured database. It then obtains field value samples through a restricted query method. This restricted query method, under read-only database connection conditions, performs controlled data read operations on the fields in the target structured database. These controlled data read operations include sampling field values and unique value statistics. The system then uniformly organizes and encapsulates the table structure information, field names, field data types, primary key relationships, foreign key relationships, and field value samples to form metadata information.
4. The natural language query and retrieval system based on deep learning according to claim 2, characterized in that, The generation of the field-level semantic representation specifically includes: The field names in the metadata information are standardized, and the standardized field names are segmented into sub-words to obtain a sequence of field name sub-words; Naming pattern analysis is performed based on the field name sub-word sequence to generate a field name pattern feature vector; Generate field type feature vectors based on the field data types in the metadata information; When the field value sample is numerical, statistical processing is performed on the field value sample to generate a distribution feature vector; When the field value sample is an enumeration type, perform statistical processing on the set of enumeration values in the field value sample to generate an enumeration feature vector; The field name sub-word sequence is matched with the word set in the external knowledge base to generate an external knowledge base matching feature vector. The external knowledge base refers to the word set that provides general vocabulary semantic association information. The feature fusion process is performed on the field name pattern feature vector, field type feature vector, distribution feature vector, enumeration feature vector, and external knowledge base matching feature vector to generate the field-level semantic representation of the corresponding field.
5. A natural language query and retrieval system based on deep learning according to claim 2, characterized in that, The formation of the semantic card specifically includes: Obtain the mapping records between historical natural language queries and corresponding structured query statements generated during system operation; Parse the structured query statement in the mapping record to obtain the fields referenced in the structured query statement; Generate corresponding query-level semantic representations based on historical natural language queries; Based on the referenced fields and their corresponding query-level semantic representations in structured query statements, the association features between fields and query semantics are generated; Based on the association features, the field-level semantic representation of the corresponding field is updated to generate the updated field-level semantic representation; The updated field-level semantic representation is associated and encapsulated with the corresponding field to form a semantic card.
6. The natural language query and retrieval system based on deep learning according to claim 2, characterized in that, The generation of the query-level semantic representation specifically includes: Receive natural language query text input by the user, perform normalization processing on the natural language query text, and obtain natural language query text in a unified representation form; Generate a term sequence based on the normalized natural language query text; Based on the term sequence, a corresponding vector representation is generated for each term, forming a term vector sequence; Context modeling is performed based on the term vector sequence to obtain a context representation sequence of the contextual relationships between terms; The context representation sequence is summarized to generate the corresponding query-level semantic representation.
7. A natural language query and retrieval system based on deep learning according to claim 2, characterized in that, The generation of the target field set specifically includes: Similarity values are calculated based on query-level semantic representation and the corresponding field-level semantic representation of semantic cards; Based on similarity values, the corresponding fields in the semantic cards are filtered to generate a set of candidate fields; For each field in the candidate field set, calculate the context consistency value based on the context information in the natural language query; For each field in the candidate field set, a joint disambiguation score is calculated based on the similarity value and the contextual consistency value; The candidate field set is filtered based on the joint disambiguation score of each field to generate the target field set.
8. A natural language query and retrieval system based on deep learning according to claim 2, characterized in that, The construction of the query constraint graph specifically includes: Identify time-related expressions from natural language queries and generate time constraints; Identify expressions related to numerical conditions from natural language queries and generate numerical threshold constraints; Identify the content of join conditions from natural language queries and generate logical relationship constraints; Identify sorting-related expressions from natural language queries and generate sorting criteria; The time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions are mapped to the corresponding fields in the target field set to construct a query constraint graph.
9. A natural language query and retrieval system based on deep learning according to claim 2, characterized in that, A structured query statement is generated based on the target field set, time constraints, numerical threshold constraints, logical relationship constraints, and sorting conditions in the query constraint diagram. Semantic consistency checks are performed on the structured query statement, including field usage consistency checks and logical constraint conflict checks. Field usage consistency checks involve sequentially traversing the fields in the structured query statement and comparing each traversed field with the target field set, marking fields not belonging to the target field set as inconsistent fields. Logical constraint conflict checks involve extracting the time constraints and numerical threshold constraints associated with each field in the structured query statement and comparing the numerical threshold constraints associated with the same field. When the value ranges of the numerical threshold constraints associated with the same field do not overlap, the same field is marked as a conflicting field. Logical constraint conflict checks involve combining and checking the numerical threshold constraints connected to the same field based on the logical relationship constraints in the structured query statement. When the value ranges obtained from the combined checks do not overlap, the same field is marked as a conflicting field.
10. A natural language query and retrieval system based on deep learning according to claim 2, characterized in that, The update of the structured query statement specifically includes: When the semantic consistency check passes, a structured query is executed to obtain the retrieval results; When the semantic consistency check fails, retrieve the fields that were marked as inconsistent or conflicting during the semantic consistency check process; Based on inconsistent or conflicting fields, a clarification question is generated, and an interactive clarification process is triggered. The clarification question is output to the user and user feedback is received, resulting in a clarification result. Based on the clarification results, the time constraints, numerical threshold constraints, logical relationship constraints, or sorting conditions associated with the inconsistent or conflicting fields in the query constraint graph are updated to obtain the updated query constraint graph, which forms the updated structured query statement.
Citation Information
Patent Citations
Data query method and system for converting natural language into database query language
CN121210495A
Intelligent data query system and method based on natural language processing
CN121255832A