An intelligent search method, system, device and medium for government affairs data
The intelligent search method for government data uses a database cluster and domain dictionaries to enhance search efficiency and accuracy by identifying entity words and indicators, addressing inefficiencies in traditional search methods.
Patent Information
- Application Number
- CN202310235509.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-03-13
AI Technical Summary
The existing government data search methods are inefficient and low in accuracy, making it difficult to meet complex business needs.
By building a government database cluster and a domain dictionary, entity recognition, target indicator matching and target entity matching are carried out, query statements are constructed, and consistency is verified using the knowledge graph to achieve efficient and accurate search.
It improves the speed and accuracy of government data search, reduces the search burden, ensures that the search results meet the user's true intentions, and reduces errors.
Smart Images

Figure CN116595027B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer information processing, and particularly relates to an intelligent search method, system, device and medium for government affairs data. Background Art
[0002] At present, in the field of government affairs, the amount of data is huge, the data involves many fields and data formats, and the data is updated frequently, making it difficult to achieve standardization, scale and intelligence. Coupled with the large number of professional terms in government affairs data, it is difficult to understand the business and data redundancy. The business required by a department often needs to screen and summarize data in multiple fields. The traditional search method is inefficient and not easy to use, and can no longer meet the growing business needs. Summary of the Invention
[0003] The technical problem to be solved by the present invention is that for search texts for government affairs, the existing search methods have low search efficiency and low search accuracy for search texts. To solve this technical problem, the present invention provides an intelligent search method, system, device and medium for government affairs data.
[0004] The technical solution of the present invention to solve the above technical problem is as follows:
[0005] An intelligent search method for government affairs data includes:
[0006] Step S1, obtaining a search text for government affairs;
[0007] Step S2, according to the search text, a pre-established government affairs database cluster and a pre-constructed government affairs domain dictionary, determining an entity word set, a target index and a target entity in the search text. The government affairs database cluster stores database tables of government affairs data in multiple fields. Each database table of government affairs data in each field corresponds to one database. Each database table includes multiple columns of government affairs data. Each column of government affairs data corresponds to a field. For each field, the field is an identifier of the corresponding column. The field includes an index field and a non-index field. The target index is the index field, and the target entity is the non-index field. The government affairs domain dictionary includes multiple domain dictionaries, and each database corresponds to one domain dictionary;
[0008] Step S3, constructing a query statement corresponding to the search text according to the entity word set, the target index and the target entity;
[0009] Step S4, determining a search result corresponding to the search text according to the query statement.
[0010] The beneficial effects of the present invention are as follows: Starting from the characteristics of government affairs data, by performing entity recognition, target index matching, and target entity matching on the search text, the entity word set, target index, and target entity in the search text are efficiently and accurately determined. According to each entity word, target index, and target entity in the entity word set and their corresponding fields in the government affairs database cluster, a query statement corresponding to the search text is constructed. According to the query statement, the search result corresponding to the search text is quickly determined from the corresponding database tables in the government affairs database cluster, without searching in the database tables corresponding to each field included in the huge government affairs database cluster, greatly reducing the search burden, improving the search speed and search accuracy of the search text, avoiding the problem that the search result accuracy is reduced due to the inability to understand the true intention of the user and thus unable to search for the search result that meets the true intention of the user.
[0011] Based on the above technical solutions, the present invention can also be improved as follows.
[0012] Further, step S3 includes:
[0013] Each entity word, the target index, and the target entity in the entity word set are respectively used as a target field value;
[0014] For each of the target field values, according to the government affairs database cluster, the target field corresponding to the target field value is determined;
[0015] According to each of the target field values and the target field corresponding to each target field value among the target field values, a query statement corresponding to the search text is constructed.
[0016] The beneficial effect of adopting the above further solution is: The corresponding target field is determined according to the target field value, a query statement corresponding to the search text is established, and the search is performed according to the query statement, ensuring the efficiency of obtaining the search result corresponding to the search text.
[0017] Further, the method further includes:
[0018] For each of the target field values, it is verified whether there is consistency between the target field value and the target field corresponding to the target field value through a pre-constructed government affairs knowledge graph;
[0019] For each of the target field values, if there is consistency between the target field value and the target field corresponding to the target field value, the target field is used as the correct field corresponding to the target field value; if there is no consistency between the target field value and the target field corresponding to the target field value, the correct field corresponding to the target field value is determined according to the government affairs knowledge graph and the government affairs domain dictionary;
[0020] Constructing the query statement corresponding to the search text according to each of the target field values and the target field corresponding to each target field value in each of the target field values includes:
[0021] Constructing the query statement corresponding to the search text according to each of the target field values and the correct field corresponding to each target field value in each of the target field values.
[0022] The beneficial effect of adopting the above further solution is that by using the knowledge graph to perform consistency detection on the target field value and its corresponding target field, the problem that one type of data can correspond to multiple fields included in the same database table is avoided, achieving the effect of eliminating ambiguity and ensuring the accuracy of the search results corresponding to the search text determined by this method.
[0023] Further, the search text includes entity words and non-entity words;
[0024] The step S2 includes:
[0025] Step S2.1, performing entity recognition on the search text to obtain the entity word set and non-entity word set corresponding to the search text;
[0026] Step S2.2, determining the target index corresponding to the search text according to the non-entity word set and the government affairs database cluster;
[0027] Step S2.3, determining the remaining text set in the search text according to the entity word set and the target index, where the remaining text set is the text set corresponding to the search text except the entity word set and the target index;
[0028] Step S2.4, determining the target database table and target database corresponding to the search text according to the target index, where the target database table is a database table corresponding to the target index, and the target database is the database corresponding to the target database table;
[0029] Step S2.5, determining the target entity corresponding to the search text according to the remaining text set, the target database table and the target database.
[0030] The beneficial effect of adopting the above further solution is that in the government affairs database cluster, the dimensions and corresponding business logics of each database table are not completely the same, and there is a situation where the same index field corresponds to multiple database tables. By determining the target database table and target database corresponding to the search text according to the target index, it is convenient to determine the target entity corresponding to the search text and improve the recognition accuracy of the search text.
[0031] Further, step S2.2 includes:
[0032] Step A1: Determine the text matrix corresponding to the non-entity word set and the two-dimensional index matrix corresponding to the preset index data set through a pre-constructed similar text automatic generation model. The index data set is a set of index fields included in each database table in the government affairs database cluster. Each element in the two-dimensional index matrix corresponds to a similar text generated by the similar text automatic generation model for one of the index fields. Each row of elements in the two-dimensional index matrix corresponds to a set of similar texts formed by all the similar texts corresponding to one of the index fields. For each of the index fields, convert the set of similar texts corresponding to the index field into a one-dimensional index matrix corresponding to the index field.
[0033] Step A2: For each of the index fields in the index data set, determine the first similarity between the one-dimensional index matrix corresponding to the index field and the text matrix. According to each of the first similarities, determine whether there is a pending index in the index data set whose first similarity is greater than a preset first similarity threshold.
[0034] Step A3: If there is a pending index, determine the target index according to the first similarity corresponding to each of the pending indexes. If there is no pending index, execute step A4.
[0035] Step A4: Segment the non-entity word set to obtain a non-entity word segmentation result, which includes multiple phrases. Perform splicing processing according to the multiple phrases to generate short sentences corresponding to the search text. Each short sentence is a text obtained by splicing at least one of the phrases.
[0036] Step A5: For each of the short sentences, determine the target index according to the second similarity between the short sentence and each of the index fields.
[0037] The beneficial effect of adopting the above further solution is as follows: According to the calculated first similarity and the first similarity threshold, each index field is divided into a pending index and a non-pending index (a non-pending index is an index field that does not belong to a pending index). Then, according to the divided pending index or non-pending index, the determination of the target index is completed. The present invention completes the matching of the target index in a threshold shunting manner, solving the problem of low recognition accuracy of correctly identifying the target index due to the existence of a large number of highly similar index fields in government affairs data.
[0038] Further, for each of the fields, the field type of the field is divided into an enumeration type and a non-enumeration type. The field value corresponding to the field with the enumeration type is an enumeration value.
[0039] In step S2.4, the determining of the target database table and the target database corresponding to the search text according to the target index includes:
[0040] If at least two of the multiple database tables have the target index, the database tables with the target index are used as the pending database tables. For each pending database table, the target database table and the target database are determined according to the similarity between the remaining text set and each enumerated value included in the pending database table;
[0041] If only one of the multiple database tables has the target index, this database table is used as the target database table, and the database corresponding to the target database table is used as the target database;
[0042] Among them, for each pending database table, the determining of the target database table and the target database according to the similarity between the remaining text set and each enumerated value included in the pending database table includes:
[0043] For each pending database table, according to the number of matches corresponding to each enumerated value in the pending database table, the comparison value corresponding to the pending database table is determined. The comparison value is the sum of the number of matches corresponding to each enumerated value in the pending database table. For each match number, the match number is the number of third similarities that meet the preset comparison conditions among all the third similarities corresponding to the enumerated value. The third similarity corresponding to the enumerated value is the similarity between the enumerated value and the remaining text in the remaining text set;
[0044] According to the comparison value corresponding to each database table, the target database and the target database table are determined.
[0045] The beneficial effect of adopting the above further solution is that in the government affairs database cluster, the dimensions and corresponding business logics of each database table are not completely the same, and there is a situation where the same index field corresponds to multiple database tables. By determining the target database and the target database table according to the third similarity corresponding to each enumerated value in the pending database table, the recognition accuracy of the search text is improved.
[0046] Further, step S2.5 includes:
[0047] Step B1, determining the target domain dictionary according to the target database. The target domain dictionary is the domain dictionary corresponding to the target database. According to the target domain dictionary, the remaining text set is segmented to obtain the remaining fine-grained entity segmentation result, and the remaining fine-grained entity segmentation result includes fine-grained entities;
[0048] Step B2. For each of the fine-grained entities, determine whether there is a corresponding relationship between the fine-grained entity and each field in the target database table according to the target domain dictionary. If there is a corresponding relationship between the fine-grained entity and the field in the target database table, then use the fine-grained entity as the target entity; if there is no corresponding relationship between each fine-grained entity and each field in the target database table, then execute Step B3;
[0049] Step B3. For each of the fine-grained entities, if the fine-grained entity meets the preset entity expansion condition, then expand the fine-grained entity to obtain the expanded entity corresponding to the fine-grained entity;
[0050] Step B4. For each of the expanded entities, use the pre-constructed similar text automatic generation model to determine the expansion matrix corresponding to the expanded entity and the two-dimensional field matrix corresponding to the target database table. Each element in the two-dimensional field matrix respectively corresponds to the similar text generated by the similar text automatic generation model for a target field. Each row of elements in the two-dimensional field matrix respectively corresponds to a set of similar texts formed by all the similar texts corresponding to a target field. The target fields are the fields included in the target database table; for each target field, convert the set of similar texts corresponding to the target field into a one-dimensional field matrix corresponding to the target field;
[0051] Step B5. For each of the expanded entities, determine the fourth similarity corresponding to the expanded entity. The fourth similarity is the similarity between the expansion matrix corresponding to the expanded entity and the one-dimensional field matrix;
[0052] Step B6. Determine the target entity according to the fourth similarity corresponding to each expanded entity.
[0053] The beneficial effects of adopting the above further solution are as follows: A large number of rare words and specialized words are contained in government affairs data. When it is determined that there is no corresponding relationship between all fine-grained entities and each field in the target database table according to the target domain dictionary, entity expansion is performed to obtain expanded entities, and entity linking is performed according to the expanded entities and the fields in the target database table, so as to infer the target entity, realizing the automatic inference of the target entity in the search text and the efficient processing of the search text.
[0054] To solve the above technical problems, the present invention also provides an intelligent search system for government affairs data, including:
[0055] A search text acquisition module, configured to acquire a search text for government affairs;
[0056] A target text acquisition module, configured to determine a set of entity words, target metrics, and target entities in the search text according to the search text, a pre-established government affairs database cluster, and a pre-constructed government affairs domain dictionary. The government affairs database cluster stores database tables of government affairs data in multiple domains. Each database table of government affairs data in each domain corresponds to one of the databases. Each database table includes multiple columns of government affairs data, and each column of government affairs data corresponds to a field. For each field, the field is an identifier of the corresponding column. The field includes a metric field and a non-metric field. The target metric is the metric field, and the target entity is the non-metric field. The government affairs domain dictionary includes multiple domain dictionaries, and each database corresponds to one of the domain dictionaries;
[0057] A query statement construction module, configured to construct a query statement according to the set of entity words, the target metrics, and the target entities;
[0058] A search text query module, configured to determine a search result corresponding to the search text according to the query statement.
[0059] To solve the above technical problems, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the intelligent search method for government affairs data as described above is implemented.
[0060] To solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the intelligent search method for government affairs data as described above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a schematic flowchart of the intelligent search method for government affairs data in the present invention;
[0062] Figure 2 It is a schematic diagram of an associated information table corresponding to the metric field in the present invention;
[0063] Figure 3 It is a schematic diagram of the structure of the intelligent search system for government affairs data in the present invention. DETAILED DESCRIPTION
[0064] The principles and features of the present invention are described below. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0065] Government affairs data is aggregated from multiple fields, covering more than 30 fields such as transportation, civil affairs, energy, and justice. There are nearly 3,000 major statistical categories involved in government affairs data. The definitions of statistical data categories in each field are complex, and the statistical requirements are diverse. The data formats of different departments are inconsistent. Some data content is single. It is difficult to meet complex business needs by using machine learning methods or methods based on fixed rules to extract index fields from search texts.
[0066] For nearly 10% of the index fields in government affairs data, the difference in the number of characters is within 3. Nearly 45% of the index fields do not contain dimension information and have a character count less than 6. Nearly 30% of the index fields contain dimension information and have a character count greater than 10. The lengths of the index fields are severely polarized, with difficult-to-unify forms and high similarity. Search in the government affairs field mainly involves short text matching. At the same time, search flexibility needs to be considered, and there are a large number of highly similar interference items, which brings great difficulties to search problems and index field matching.
[0067] When identifying government affairs data, in addition to identifying conventional time, location, and organization names, it is also necessary to identify thousands of dimensions such as enterprise nature, marital status, emergency event classification, employment type, and industry classification. In addition, the identification of dimensions involves problems such as large differences in industries and numerous professional names and abbreviations. If a named entity recognition model is used to identify dimension data in search texts, there are problems such as insufficient training corpus, large workload, and difficult-to-precise identification.
[0068] To solve the problems of low search efficiency and low search accuracy of existing search methods for search texts of government affairs data, the present invention provides an intelligent search method, system, device, and medium for government affairs data. By combining data indexing (i.e., determining the attributes of the fields corresponding to the government affairs data in each column of the database table, where the attributes of the fields are index fields or non-index fields, and the index fields include time, location, person name, organization name, computable fields, enumerated values, age), index system construction (i.e., determining the index fields included in each database table in the government affairs database cluster, as well as the query conditions and calculation methods corresponding to each index field, and establishing a corresponding relationship between the database table and the index field), semantic analysis (i.e., determining the target index corresponding to the search text according to the preset index data set), and government affairs domain dictionary construction, precise and efficient search of the search text is achieved.
[0069] Embodiment 1
[0070] As Figure 1 shown, this embodiment provides an intelligent search method for government affairs data, including:
[0071] Step S1, obtaining a search text for government affairs;
[0072] Step S2: Based on the search text, the pre-established government affairs database cluster, and the pre-constructed government affairs domain dictionary, determine the set of entity words, target indicators, and target entities in the search text. The government affairs database cluster stores database tables of government affairs data in multiple domains. Each database table of government affairs data in each domain corresponds to one database. Each database table includes multiple columns of government affairs data, and each column of government affairs data corresponds to a field. For each field, the field is the identifier of the corresponding column. The field includes an indicator field and a non-indicator field. The target indicator is the indicator field, and the target entity is the non-indicator field. The government affairs domain dictionary includes multiple domain dictionaries, and each database corresponds to one domain dictionary;
[0073] Step S3: Based on the set of entity words, the target indicators, and the target entities, construct a query statement corresponding to the search text;
[0074] Step S4: Based on the query statement, determine the search result corresponding to the search text.
[0075] In the present invention, one database table may have multiple indicator fields; the target entity corresponds to the field included in the database table.
[0076] Among them, for each database, the domain dictionary corresponding to the database includes a thesaurus and a relation word table constructed according to the database. The thesaurus includes fields of the enumeration type and indicator fields in the database. The relation word table includes the mapping relationships between each field in the database and its equivalent words. There are a large number of proper nouns and idiomatic expressions in the government affairs domain, and there are many rare words in the database table. By constructing the relation word table, it is convenient to perform semantic analysis subsequently, making the search results accurate, flexible, and easy to use. For example, in the database table, there are summary fields such as "Current value of liquor production", "Cumulative value of liquor production", and "Year-on-year growth of liquor production". When constructing the domain dictionary, words such as "Current value" and "Quarterly value" can be equated with the word "Current value of liquor production", words such as "Cumulative" and "Total" can be equated with the word "Cumulative value of liquor production", and words such as "Growth rate" and "Growth" can be equated with the word "Year-on-year growth of liquor production". When the user searches for "How much did the liquor production increase in 2018", it can be matched with "Year-on-year growth of liquor production".
[0077] Among them, in Step S1, the search text is determined through the following steps:
[0078] Obtain the original text;
[0079] Clean the original text to delete the colloquial words and redundant special symbols contained in the original text, and obtain the cleaned text;
[0080] If the length of the cleaned text is greater than a preset text length threshold, then intercept the cleaned text to obtain a short text, and perform a normalization process on the short text to obtain the search text;
[0081] If the length of the cleaned text is less than or equal to the text length threshold, then perform a normalization process on the cleaned text to obtain the search text.
[0082] In this embodiment, performing a normalization process on the short text to obtain the search text includes:
[0083] If the short text contains time text, for each piece of time text, use the time text as the text to be confirmed, determine whether the text to be confirmed is in a standard time format. If not, then convert the text to be confirmed into a standard time format to obtain the normalized time text, and replace the text to be confirmed in the short text with the normalized time text to obtain the search text.
[0084] In this embodiment, that the short text contains time text means that the short text contains time information, and the standard time format refers to the data format of the data type of time type stored in the government affairs database cluster; the method of performing a normalization process on the cleaned text to obtain the search text is the same as the method of performing a normalization process on the short text to obtain the search text, and the same parts will not be elaborated. This method clears redundant symbols and vocabulary in the original text to obtain the cleaned text corresponding to the original text, and then performs a normalization process according to the cleaned text to complete the problem rewriting of the original text, thereby obtaining the search text corresponding to the original text, which is convenient for subsequent entity recognition of the search text.
[0085] Among them, the search text includes entity words and non-entity words, and the step S2 includes:
[0086] Step S2.1, perform entity recognition on the search text to obtain the entity word set and non-entity word set corresponding to the search text;
[0087] Step S2.2, determine the target index corresponding to the search text according to the non-entity word set and the government affairs database cluster;
[0088] Step S2.3, determine the remaining text set in the search text according to the entity word set and the target index, and the remaining text set is the text set corresponding to the search text except for the entity word set and the target index;
[0089] Step S2.4. Determine the target database table and target database corresponding to the search text according to the target metric. The target database table is a database table corresponding to the target metric, and the target database is the database corresponding to the target database table.
[0090] Step S2.5. Determine the target entity corresponding to the search text according to the remaining text set, the target database table, and the target database.
[0091] Among them, in step S2.1, the entity word set includes multiple entity words, and the entity word set is determined through the following steps:
[0092] Identify the first entity words included in the search text through a pre-constructed time and number recognition rule to obtain a first entity word set. For each of the first entity words, the entity represented by the first entity word has an attribute including at least one of time or number. In this embodiment, the time and number recognition rule adopts an existing automatic recognition rule for recognizing time and numbers.
[0093] Identify the second entity words included in the search text through a pre-constructed entity detection model to obtain a second entity word set. For each of the second entity words, the entity represented by the second entity word has an attribute including at least one of location, organization name, or person name. In this embodiment, the entity detection model adopts an existing BERT+CRF model.
[0094] The entity word set includes the first entity word set and the second entity word set. In this embodiment, the entity words corresponding to the first entity words are stored in the entity word set in the order of the positions of the first entity words in the search text.
[0095] In the statistical requirements for government affairs data, statistics including time, numbers, locations, organization names, and person names are required. This method performs entity recognition on the search text through a time and number recognition rule and an entity detection model, which is efficient and accurate.
[0096] Among them, in step S2.1, the non-entity word set is determined according to the search text and the entity word set. In this embodiment, the non-entity word set is determined through the following steps:
[0097] Convert the search text into a string to obtain a target string.
[0098] For each entity word in the entity word set, convert the entity word into a string to obtain an entity character, and replace the entity character in the target string with a specified delimiter (for example: "&&", space).
[0099] According to the specified delimiter, split the target string to obtain a plurality of first substrings, and store each of the first substrings as a non-entity word in the non-entity word set in the order of the positions of the first substrings in the target string. The non-entity word set is represented as a = {a1, a2, …, a i}, where i represents the total number of non-entity words, and a m (1 ≤ m ≤ i) represents the m-th non-entity word in the non-entity word set.
[0100] Among them, the step S2.2 includes:
[0101] Step A1, through a pre-constructed similar text automatic generation model, determine the text matrix corresponding to the non-entity word set and the two-dimensional index matrix corresponding to the preset index data set. The index data set is the set of index fields included in each database table in the government affairs database cluster. Each element in the two-dimensional index matrix corresponds to a similar text generated by the similar text automatic generation model for a corresponding index field. Each row of elements in the two-dimensional index matrix (i.e., each dimension of data in the two-dimensional index matrix) corresponds to a set of similar texts formed by all similar texts corresponding to a corresponding index field; for each index field, convert the set of similar texts corresponding to the index field into a one-dimensional index matrix, and each element in the one-dimensional index matrix corresponds to a similar text generated by the similar text automatic generation model for a corresponding index field.
[0102] Step A2, for each index field in the index data set, determine the first similarity between the one-dimensional index matrix corresponding to the index field and the text matrix. According to each of the first similarities, determine whether there is a pending index in the index data set whose first similarity is greater than a preset first similarity threshold; in this embodiment, the cosine similarity algorithm is used to calculate the cosine similarity between the one-dimensional index matrix and the text matrix, and the calculated cosine similarity is used as the first similarity between the one-dimensional index matrix and the text matrix, that is, the semantic matching between each index field and the non-entity word set is completed through the cosine similarity algorithm, and the value of the first similarity threshold is 0.8.
[0103] Step A3, if there is a pending index, determine the target index according to the first similarity corresponding to each pending index; if there is no pending index, execute step A4; in this embodiment, if there is a pending index, the index field corresponding to the largest first similarity is used as the target index.
[0104] Step A4: Segment the non-entity word set to obtain a non-entity word segmentation result, where the non-entity word segmentation result includes multiple phrases; perform splicing processing based on the multiple phrases to generate short sentences corresponding to the search text, and each short sentence is a text obtained by splicing at least one of the phrases. In this embodiment, the jieba word segmentation tool is used to complete the word segmentation operation on the non-entity word set, and the splicing principle for generating the short sentences based on the multiple phrases is: splice the phrases according to the position order of each phrase in the search text.
[0105] Step A5: For each short sentence, determine the target index according to the second similarity between the short sentence and each of the index fields. In this embodiment, the edit distance algorithm is used to calculate the edit distance between the short sentence and the index field, and the calculated edit distance is used as the second similarity between the short sentence and the index field. If there is a second similarity greater than the first similarity threshold, the index field corresponding to the maximum second similarity is used as the target index; if there is no second similarity greater than the first similarity threshold, it indicates that there is no data corresponding to the search text stored in the government affairs database.
[0106] Among them, there are two methods to determine the index field. One is to set it manually. For example, manually extract the index fields "number of enterprises", "enterprise tax payment amount", "total registered capital of enterprises", and "number of large-scale enterprises" from the enterprise database table, and then define corresponding query conditions and calculation methods for each of the extracted index fields. For example, the "total registered capital of enterprises" is to sum the "registered capital" column. The other is to determine it through existing index indexing tools. For example, in the gold production database table, there are statistical items "cumulative value of gold production" and "year-on-year growth of gold production". Both of them belong to the same category, but if they are used as separate search index fields, it will cause interference in search due to excessive similarity. Therefore, the index fields corresponding to these two statistical items are defined as "gold production", that is, the part before the "_" is extracted. The advantage of this is that it makes the search more accurate and flexible. When the user searches for "what is the gold production", it will query both "cumulative value of gold production" and "year-on-year growth of gold production" at the same time. When the user needs to perform a precise search for "how much is the year-on-year growth of gold production in the second quarter of 2020", it will match the statistical item "year-on-year growth of gold production" in combination with the corresponding domain dictionary, thus realizing fuzzy query and precise query. When the data formats of the government affairs data in each column of the database have a unified rule, it is preferred to use the index indexing tool to complete the construction of the index system. Otherwise, the manual setting method is used to complete the construction of the index system.
[0107] Figure 2 is an associated information table for the index field obtained by manual sorting. Figure 2Among them, IndicatorName represents the name of the indicator field, SourceID is the ID of the database table to which the indicator field belongs, TableName represents the table name of the database table to which the indicator field belongs, IndicatorCondition represents the calculation method corresponding to the indicator field, and OutputAgg represents the calculation rule corresponding to the indicator field.
[0108] In this embodiment, the similar text automatic generation model adopts the SimBERT model, that is, the non-entity word set is used as the input, and the text matrix is output through the SimBERT model. The indicator data set is used as the input, and the two-dimensional indicator matrix is output through the SimBERT model; the indicator fields included in each database table in the government affairs database cluster are set manually, and the indicator data set is expressed as b = {b1, b2,..., b j}, j represents the total number of the indicator fields, and b n (1 ≤ n ≤ j) represents the nth indicator field in the indicator data set;
[0109] Each element in the text matrix respectively corresponds to a similar text generated by the similar text automatic generation model for a non-entity word. Each row of elements in the text matrix respectively corresponds to a set of similar texts formed by all the similar texts corresponding to a non-entity word; for each non-entity word, the set of similar texts corresponding to the non-entity word is transformed into a one-dimensional text matrix corresponding to the non-entity word. The text matrix can be expressed as:
[0110]
[0111] The one-dimensional text matrix obtained by transforming the mth (1 ≤ m ≤ i) non-entity word in the non-entity word set through the similar text automatic generation model is expressed as:
[0112] a m,new = [a m1 a m2 a m3 … a mk
[0113] The two-dimensional indicator matrix can be expressed as:
[0114]
[0115] The one-dimensional indicator matrix obtained by transforming the nth (1 ≤ n ≤ j) indicator field in the indicator data set through the similar text automatic generation model is expressed as:
[0116] b n,new = [b n1 bn2 b n3 … b nk ].
[0117] Among them, in the step S2.3, the determining the remaining text set in the search text according to the entity word set and the target index includes:
[0118] Convert the search text into a string to obtain a target string;
[0119] For each entity word in the entity word set, convert the entity word into a string to obtain an entity character, and replace the entity character in the target string with a specified delimiter;
[0120] Convert the target index into a string to obtain an index character, and replace the index character in the target string with a specified delimiter;
[0121] According to the specified delimiter, split the target string to obtain a plurality of second substrings, and use each second substring as the remaining text and store it in the remaining text set in the order of the positions of the second substrings in the target string.
[0122] Among them, for each field, the field type of the field is divided into an enumeration type and a non-enumeration type, and the field value corresponding to the field of the enumeration type is an enumeration value;
[0123] In the step S2.4, the determining the target database table and the target database corresponding to the search text according to the target index includes:
[0124] If at least two database tables among multiple database tables have the target index, then use the database tables with the target index as the pending database tables. For each pending database table, determine the target database table and the target database according to the similarity between the remaining text set and each enumeration value included in the pending database table;
[0125] If only one database table among multiple database tables has the target index, then use this database table as the target database table, and use the database corresponding to the target database table as the target database;
[0126] For each pending database table, the determining the target database table and the target database according to the similarity between the remaining text set and each enumeration value included in the pending database table includes:
[0127] For each of the to-be-determined database tables, according to the matching numbers corresponding to each enumerated value in the to-be-determined database table, determine the comparison value corresponding to the to-be-determined database table. The comparison value is the sum of the matching numbers corresponding to each enumerated value in the to-be-determined database table. For each of the matching numbers, the matching number is the number of third similarities that meet the preset comparison condition among all the third similarities corresponding to the enumerated value. The third similarity corresponding to the enumerated value is the similarity between the enumerated value and the remaining text in the remaining text set. In this embodiment, the enumerated values in each of the database tables can be set manually, and the cosine similarity algorithm is used to calculate the similarity between the enumerated value and the remaining text in the remaining text set. The comparison condition is that the third similarity is greater than the preset second similarity threshold.
[0128] According to the comparison value corresponding to each of the database tables, determine the target database and the target database table. In this embodiment, the to-be-determined database table corresponding to the largest comparison value is used as the target database table, and the database to which the target database table belongs is used as the target database.
[0129] Among them, the step S2.5 includes:
[0130] Step B1, according to the target database, determine the target domain dictionary. The target domain dictionary is the domain dictionary corresponding to the target database. According to the target domain dictionary, segment the remaining text set to obtain the remaining fine-grained entity segmentation result. The remaining fine-grained entity segmentation result includes fine-grained entities. In this embodiment, the loading of the target domain dictionary and the segmentation operation of the remaining text set are completed through the segmentation tool - jieba.
[0131] Step B2, for each of the fine-grained entities, according to the target domain dictionary, determine whether there is a corresponding relationship between the fine-grained entity and each field in the target database table. If there is a corresponding relationship between the fine-grained entity and the field in the target database table, then use the fine-grained entity as the target entity. If there is no corresponding relationship between each of the fine-grained entities and each of the fields in the target database table, then execute step B3. In this embodiment, for each of the fine-grained entities, when the target domain dictionary contains the association information between the fine-grained entity and a field included in the target database table, it is determined that there is a corresponding relationship between the fine-grained entity and the field in the target database table.
[0132] Step B3, for each of the fine-grained entities, if the fine-grained entity meets the preset entity expansion condition, then expand the fine-grained entity to obtain the expanded entity corresponding to the fine-grained entity.
[0133] Step B4: For each of the extended entities, determine the extended matrix corresponding to the extended entity and the two-dimensional field matrix corresponding to the target database table through a pre-constructed similar text automatic generation model. Each element in the two-dimensional field matrix corresponds to a similar text generated by the similar text automatic generation model for a target field. Each row of elements in the two-dimensional field matrix corresponds to a set of similar texts formed by all the similar texts corresponding to a target field. The target fields are the fields included in the target database table. For each target field, convert the set of similar texts corresponding to the target field into a one-dimensional field matrix corresponding to the target field.
[0134] Step B5: For each of the extended entities, determine the fourth similarity corresponding to the extended entity. The fourth similarity is the similarity between the extended matrix corresponding to the extended entity and the one-dimensional field matrix.
[0135] Step B6: Determine the target entity according to the fourth similarity corresponding to each extended entity. In this embodiment, the extended entity corresponding to the maximum fourth similarity is used as the target entity.
[0136] In this embodiment, the entity extension condition is that the length L of the fine-grained entity satisfies 2 ≤ L ≤ 4. The entity extension of the fine-grained entity to obtain the extended entity corresponding to the fine-grained entity includes: concatenating the fine-grained entity with the fine-grained entity adjacent to and after the fine-grained entity in the search text to obtain the extended entity corresponding to the fine-grained entity. For each extended entity, the method for determining the fourth similarity corresponding to the extended entity is the same as the method for determining the first similarity between the one-dimensional index matrix corresponding to the index field and the text matrix, and the same parts will not be elaborated again.
[0137] Among them, step S3 includes:
[0138] Take each entity word in the entity word set, the target index, and the target entity as a target field value respectively.
[0139] For each target field value, determine the target field corresponding to the target field value according to the target database.
[0140] Construct a query statement corresponding to the search text according to each target field value and the target field corresponding to each target field value among each target field value.
[0141] In this embodiment, constructing the query statement corresponding to the search text according to each of the target field values and the target field corresponding to each target field value in each of the target field values includes:
[0142] Taking the target field corresponding to each target field value in each of the target field values as a filling field;
[0143] Determining the query condition and calculation method corresponding to each of the filling fields;
[0144] Establishing an SQL query statement according to each of the target field values, each of the filling fields, the query condition and calculation method corresponding to each of the filling fields, that is, constructing the query statement corresponding to the search text.
[0145] Optionally, the method further includes:
[0146] For each of the target field values, verifying whether there is consistency between the target field value and the target field corresponding to the target field value through a pre-constructed government knowledge graph;
[0147] For each of the target field values, if there is consistency between the target field value and the target field corresponding to the target field value, then taking the target field as the correct field corresponding to the target field value, and if there is no consistency between the target field value and the target field corresponding to the target field value, then determining the correct field corresponding to the target field value according to the government knowledge graph and the target domain dictionary;
[0148] The constructing the query statement corresponding to the search text according to each of the target field values and the target field corresponding to each target field value in each of the target field values includes:
[0149] Constructing the query statement corresponding to the search text according to each of the target field values and the correct field corresponding to each target field value in each of the target field values.
[0150] Most government affairs data is structured data. There are often multiple types of data such as time, location, and enumeration in a database table. When the search text contains multiple times and locations, it is necessary to determine whether it is linked to the correct time and location. This method constructs a government affairs knowledge graph to solve the problem that one type of data can correspond to multiple fields included in the same database table, achieving the effect of eliminating ambiguity. In this embodiment, the government affairs knowledge graph includes the knowledge graph corresponding to each database. For each database, according to the relationship between the fields in the database, the triples corresponding to the database are manually constructed to obtain the knowledge graph corresponding to the database. For each target field value, the knowledge graph corresponding to the target database is used to verify whether there is consistency between the target field value and the target field corresponding to the target field value, that is, entity disambiguation is performed using the knowledge graph, thereby improving the search accuracy of this method for the search text.
[0151] After the user inputs the original text (i.e., the search question entered by the user), this method first cleans the original text to obtain the cleaned text, obtains the corresponding search text according to the cleaned text to complete the question rewriting, then performs entity recognition on the search text. After removing the entity words in the search text, index matching (corresponding to step S2.2 above, index matching includes semantic matching, similarity matching, and enumeration value hit count matching) is performed to obtain the target index. Then, according to the entity words and the target index, the remaining text set is determined. According to the remaining text set, the target database table and the target database are determined, and the domain dictionary corresponding to the target database is loaded. Then, the remaining text set is segmented to obtain the corresponding remaining fine-grained entity segmentation result. Entity matching (corresponding to step B2 above) / and entity expansion (corresponding to step B3 above) / and entity linking (corresponding to steps B4 and B5 above) are performed according to the remaining fine-grained entity segmentation result to obtain the target entity. According to the entity words, target entities, target indexes, and their corresponding query conditions and calculation methods included in the search text, an SQL query statement is established, and the search result corresponding to the search text can be quickly queried in the corresponding target database.
[0152] Embodiment 2
[0153] Based on the same principle as the intelligent search method for government affairs data described in the above Embodiment 1, this embodiment provides an intelligent search system for government affairs data, as Figure 3 shown, including:
[0154] A search text acquisition module, configured to acquire a search text for government affairs, where the search text includes entity words and non-entity words;
[0155] A target text acquisition module, configured to determine an entity word set, a target indicator, and a target entity in the search text according to the search text, a pre-established government affairs database cluster, and a pre-constructed government affairs domain dictionary; the government affairs database cluster stores database tables of government affairs data in multiple domains, each database table of government affairs data in each domain corresponds to one of the databases, each database table includes multiple columns of government affairs data, each column of government affairs data corresponds to a field, for each field, the field is an identifier of the corresponding column, the field includes an indicator field and a non-indicator field, the target indicator is the indicator field, and the target entity is the non-indicator field; for each field, the field type is divided into an enumeration type and a non-enumeration type, and the field value corresponding to the field of the enumeration type is an enumeration value; the government affairs domain dictionary includes multiple domain dictionaries, and each database corresponds to one of the domain dictionaries;
[0156] A query statement construction module, configured to construct a query statement according to the entity word set, the target indicator, and the target entity;
[0157] A search text query module, configured to determine a search result corresponding to the search text according to the query statement.
[0158] Wherein, the target text acquisition module includes:
[0159] An entity recognition unit, configured to perform entity recognition on the search text to obtain an entity word set and a non-entity word set corresponding to the search text;
[0160] A target indicator determination unit, configured to determine a target indicator corresponding to the search text according to the non-entity word set and the government affairs database cluster;
[0161] A remaining text set determination unit, configured to determine a remaining text set in the search text according to the entity word set and the target indicator, and the remaining text set is a text set corresponding to the search text except for the entity word set and the target indicator;
[0162] A target database table and database determination unit, configured to determine a target database table and a target database corresponding to the search text according to the target indicator, the target database table is a database table corresponding to the target indicator, and the target database is a database corresponding to the target database table;
[0163] A target entity determination unit, configured to determine a target entity corresponding to the search text according to the remaining text set, the target database table, and the target database.
[0164] Wherein, the entity recognition unit includes:
[0165] The first entity word determination subunit is configured to identify the first entity words included in the search text through a pre-constructed time and number recognition rule, so as to obtain a first entity word set. For each of the first entity words, the first entity word represents an entity whose attribute includes at least one of time or number;
[0166] The second entity word determination subunit is configured to identify the second entity words included in the search text through a pre-constructed entity detection model, so as to obtain a second entity word set. For each of the second entity words, the second entity word represents an entity whose attribute includes at least one of location, organization name, or person name;
[0167] The entity word set includes the first entity word set and the second entity word set.
[0168] Among them, the target index determination unit includes:
[0169] The first processing subunit is configured to determine a text matrix corresponding to the non-entity word set and a two-dimensional index matrix corresponding to a preset index data set through a pre-constructed similar text automatic generation model. The index data set is a set of index fields included in all database tables in the government affairs database cluster. Each element in the two-dimensional index matrix corresponds to a similar text generated by the similar text automatic generation model for a corresponding index field. Each row of elements in the two-dimensional index matrix corresponds to a set of similar texts formed by all similar texts corresponding to a corresponding index field; for each index field, convert the set of similar texts corresponding to the index field into a one-dimensional index matrix corresponding to the index field;
[0170] The second processing subunit is configured to, for each index field in the index data set, determine a first similarity between the one-dimensional index matrix corresponding to the index field and the text matrix, and determine whether there is a pending index in the index data set whose first similarity is greater than a preset first similarity threshold according to each of the first similarities;
[0171] The third processing subunit is configured to, if there is a pending index, determine a target index according to the first similarity corresponding to each pending index; if there is no pending index, execute the fourth processing subunit;
[0172] The fourth processing subunit is configured to perform word segmentation on the non-entity word set to obtain a non-entity word segmentation result, where the non-entity word segmentation result includes a plurality of phrases; perform splicing processing according to the plurality of phrases to generate a short sentence corresponding to the search text, and each short sentence is a text obtained by splicing at least one of the phrases;
[0173] A fifth processing subunit, configured to determine a target indicator for each of the short sentences according to a second similarity between the short sentence and each of the indicator fields.
[0174] Wherein, the target database and table determination unit is specifically configured to:
[0175] If at least two of the multiple database tables have the target indicator, then use the database tables having the target indicator as candidate database tables, and for each of the candidate database tables, determine a target database table and a target database according to the similarity between the remaining text set and each enumerated value included in the candidate database table;
[0176] If only one of the multiple database tables has the target indicator, then use this database table as the target database table, and use the database corresponding to the target database table as the target database.
[0177] Wherein, when the target database and table determination unit is configured to determine a target database table and a target database according to the similarity between the remaining text set and each enumerated value included in each of the candidate database tables, it is specifically configured to:
[0178] For each of the candidate database tables, determine a comparison value corresponding to the candidate database table according to the number of matches corresponding to each enumerated value in the candidate database table, where the comparison value is the sum of the number of matches corresponding to each enumerated value in the candidate database table, and for each of the number of matches, the number of matches is the number of third similarities that meet a preset comparison condition among all the third similarities corresponding to the enumerated value, and the third similarity corresponding to the enumerated value is the similarity between the enumerated value and the remaining text in the remaining text set;
[0179] Determine the target database and the target database table according to the comparison value corresponding to each of the database tables.
[0180] Wherein, the target entity determination unit includes:
[0181] A first determination subunit, configured to determine a target domain dictionary according to the target database, where the target domain dictionary is the domain dictionary corresponding to the target database, and perform word segmentation on the remaining text set according to the target domain dictionary to obtain a remaining fine-grained entity word segmentation result, and the remaining fine-grained entity word segmentation result includes fine-grained entities;
[0182] A second determination subunit, configured to, for each of the fine-grained entities, determine whether there is a corresponding relationship between the fine-grained entity and each field in the target database table according to the target domain dictionary. If there is a corresponding relationship between the fine-grained entity and the field in the target database table, the fine-grained entity is used as the target entity; if there is no corresponding relationship between each fine-grained entity and each field in the target database table, the third determination subunit is executed;
[0183] A third determination subunit, configured to, for each of the fine-grained entities, if the fine-grained entity meets a preset entity expansion condition, perform entity expansion on the fine-grained entity to obtain an extended entity corresponding to the fine-grained entity;
[0184] A fourth determination subunit, configured to, for each of the extended entities, determine an extended matrix corresponding to the extended entity and a two-dimensional field matrix corresponding to the target database table through a pre-constructed similar text automatic generation model. Each element in the two-dimensional field matrix respectively corresponds to a similar text generated by the similar text automatic generation model for a target field. Each row of elements in the two-dimensional field matrix respectively corresponds to a set of similar texts formed by all similar texts corresponding to a target field. The target field is a field included in the target database table; for each target field, convert the set of similar texts corresponding to the target field into a one-dimensional field matrix corresponding to the target field;
[0185] A fifth determination subunit, configured to, for each of the extended entities, determine a fourth similarity corresponding to the extended entity, where the fourth similarity is the similarity between the extended matrix corresponding to the extended entity and the one-dimensional field matrix;
[0186] A sixth determination subunit, configured to determine a target entity according to the fourth similarity corresponding to each extended entity.
[0187] Wherein, the query statement construction module includes:
[0188] A target field value determination unit, configured to use each entity word in the entity word set, the target index, and the target entity as a target field value respectively;
[0189] A target field determination unit, configured to, for each of the target field values, determine a target field corresponding to the target field value according to the government affairs database cluster;
[0190] A query statement construction unit, configured to construct a query statement corresponding to the search text according to each of the target field values and the target field corresponding to each of the target field values.
[0191] Optionally, the system further includes a consistency detection module, and the consistency detection module is configured to:
[0192] For each of the target field values, verify whether there is consistency between the target field value and the corresponding target field through a pre-constructed government affairs knowledge graph;
[0193] For each of the target field values, if there is consistency between the target field value and the corresponding target field, then use the target field as the correct field corresponding to the target field value; if there is no consistency between the target field value and the corresponding target field, then determine the correct field corresponding to the target field value according to the government affairs knowledge graph and the government affairs domain dictionary;
[0194] The query statement construction unit is configured to construct a query statement corresponding to the search text according to each of the target field values and the correct field corresponding to each of the target field values.
[0195] Optionally, the system further includes a search text determination module, and the search text determination module includes:
[0196] A text acquisition unit, configured to acquire the original text;
[0197] A text cleaning unit, configured to clean the original text to delete the colloquial words and redundant special symbols included in the original text, and obtain a cleaned text;
[0198] A first standardization unit, configured to, if the length of the cleaned text is greater than a preset text length threshold, intercept the cleaned text to obtain a short text, and perform standardization processing on the short text to obtain the search text;
[0199] A second standardization unit, configured to, if the length of the cleaned text is less than or equal to the text length threshold, perform standardization processing on the cleaned text to obtain the search text.
[0200] Embodiment III
[0201] To solve the above technical problems, this embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the intelligent search method for government affairs data as described in Embodiment I.
[0202] Embodiment IV
[0203] To solve the above technical problems, this embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the intelligent search method for government affairs data as described in Embodiment 1.
[0204] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0205] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0206] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. An intelligent search method for government affairs data, characterized in that, Including: Step S1: Obtain the search text for government affairs. Step S2: According to the search text, the pre-established government affairs database cluster, and the pre-constructed government affairs domain dictionary, determine the entity word set, target indicators, and target entities in the search text. The government affairs database cluster stores database tables of government affairs data in multiple domains. Each database table of government affairs data in each domain corresponds to a database. Each database table includes multiple columns of government affairs data, and each column of government affairs data corresponds to a field. For each field, the field is the identifier of the corresponding column. The field includes an indicator field and a non-indicator field. The target indicator is the indicator field, and the target entity is the non-indicator field. The government affairs domain dictionary includes multiple domain dictionaries, and each database corresponds to one of the domain dictionaries. Step S3: Construct a query statement corresponding to the search text according to the entity word set, the target indicators, and the target entities. Step S4: Determine the search result corresponding to the search text according to the query statement. The search text includes entity words and non-entity words. The step S2 includes: Step S2.1: Perform entity recognition on the search text to obtain the entity word set and non-entity word set corresponding to the search text. Step S2.2: Determine the target indicators corresponding to the search text according to the non-entity word set and the government affairs database cluster. Step S2.3: Determine the remaining text set in the search text according to the entity word set and the target indicators. The remaining text set is the text set corresponding to the search text except for the entity word set and the target indicators. Step S2.4: Determine the target database table and target database corresponding to the search text according to the target indicators. The target database table is a database table corresponding to the target indicator, and the target database is the database corresponding to the target database table. Step S2.5: Determine the target entities corresponding to the search text according to the remaining text set, the target database table, and the target database. The step S2.2 includes: Step A1: Through a pre-constructed similar text automatic generation model, determine the text matrix corresponding to the non-entity word set and the two-dimensional indicator matrix corresponding to the preset indicator data set. The indicator data set is the set of indicator fields included in each database table in the government affairs database cluster. Each element in the two-dimensional indicator matrix corresponds to a similar text generated by the similar text automatic generation model for a corresponding indicator field. Each row of elements in the two-dimensional indicator matrix corresponds to a set of similar texts formed by all similar texts corresponding to a corresponding indicator field. For each indicator field, convert the set of similar texts corresponding to the indicator field into a one-dimensional indicator matrix corresponding to the indicator field. Step A2: For each of the index fields in the index dataset, determine the first similarity between the one-dimensional index matrix corresponding to the index field and the text matrix. According to each of the first similarities, determine whether there is a pending index in the index dataset whose first similarity is greater than a preset first similarity threshold. Step A3: If there is a pending index, determine the target index according to the first similarity corresponding to each of the pending indexes; if there is no pending index, execute Step A4. Step A4: Segment the non-entity word set to obtain a non-entity word segmentation result, where the non-entity word segmentation result includes multiple phrases; perform a splicing process according to the multiple phrases to generate a short sentence corresponding to the search text, and each short sentence is a text obtained by splicing at least one of the phrases. Step A5: For each short sentence, determine the target index according to the second similarity between the short sentence and each index field.
2. The method according to claim 1, wherein Step S3 includes: Take each entity word in the entity word set, the target index, and the target entity as a target field value respectively. For each target field value, determine the target field corresponding to the target field value according to the government affairs database cluster. Construct a query statement corresponding to the search text according to each target field value and the target field corresponding to each target field value in each target field value.
3. The method according to claim 2, wherein The method further includes: For each target field value, verify whether there is consistency between the target field value and the target field corresponding to the target field value through a pre-constructed government affairs knowledge graph. For each target field value, if there is consistency between the target field value and the target field corresponding to the target field value, take the target field as the correct field corresponding to the target field value; if there is no consistency between the target field value and the target field corresponding to the target field value, determine the correct field corresponding to the target field value according to the government affairs knowledge graph and the government affairs domain dictionary. The constructing a query statement corresponding to the search text according to each target field value and the target field corresponding to each target field value in each target field value includes: Construct a query statement corresponding to the search text according to each target field value and the correct field corresponding to each target field value in each target field value.
4. The method according to claim 1, wherein For each field, the field type of the field is divided into an enumeration type and a non-enumeration type, and the field value corresponding to the field of the enumeration type is an enumeration value. In Step S2.4, the determining the target database table and the target database corresponding to the search text according to the target index includes: If at least two database tables among the multiple database tables have the target index, take the database tables with the target index as pending database tables. For each pending database table, determine the target database table and the target database according to the similarity between the remaining text set and each enumeration value included in the pending database table. If only one of the multiple database tables has the target metric, then use this database table as the target database table and use the database corresponding to the target database table as the target database; Among them, for each of the to-be-determined database tables, determining the target database table and the target database according to the similarity between the remaining text set and each enumerated value included in the to-be-determined database table includes: For each of the to-be-determined database tables, determine the comparison value corresponding to the to-be-determined database table according to the matching number corresponding to each enumerated value in the to-be-determined database table. The comparison value is the sum of the matching numbers corresponding to each enumerated value in the to-be-determined database table. For each of the matching numbers, the matching number is the number of third similarities that meet the preset comparison conditions among all the third similarities corresponding to the enumerated value. The third similarity corresponding to the enumerated value is the similarity between the enumerated value and the remaining text in the remaining text set; Determine the target database and the target database table according to the comparison value corresponding to each database table.
5. The method according to claim 1, wherein The step S2.5 includes: Step B1, determine the target domain dictionary according to the target database. The target domain dictionary is the domain dictionary corresponding to the target database. According to the target domain dictionary, segment the remaining text set to obtain the remaining fine-grained entity segmentation result, and the remaining fine-grained entity segmentation result includes fine-grained entities; Step B2, for each of the fine-grained entities, determine whether there is a corresponding relationship between the fine-grained entity and each field in the target database table according to the target domain dictionary. If there is a corresponding relationship between the fine-grained entity and the field in the target database table, then use the fine-grained entity as the target entity; if there is no corresponding relationship between each fine-grained entity and each field in the target database table, then execute Step B3; Step B3, for each of the fine-grained entities, if the fine-grained entity meets the preset entity expansion condition, then expand the fine-grained entity to obtain the expanded entity corresponding to the fine-grained entity; Step B4, for each of the expanded entities, determine the expansion matrix corresponding to the expanded entity and the two-dimensional field matrix corresponding to the target database table through the pre-constructed similar text automatic generation model. Each element in the two-dimensional field matrix corresponds to a similar text generated by a target field through the similar text automatic generation model. Each row of elements in the two-dimensional field matrix corresponds to a set of similar texts formed by all the similar texts corresponding to a target field. The target field is the field included in the target database table; for each target field, convert the set of similar texts corresponding to the target field into a one-dimensional field matrix corresponding to the target field; Step B5, for each of the expanded entities, determine the fourth similarity corresponding to the expanded entity. The fourth similarity is the similarity between the expansion matrix corresponding to the expanded entity and the one-dimensional field matrix; Step B6: Determine the target entity according to the fourth similarity corresponding to each of the extended entities.
6. An intelligent search system for government affairs data, characterized in that, Including: A search text acquisition module, configured to acquire a search text for government affairs; A target text acquisition module, configured to determine an entity word set, a target indicator, and a target entity in the search text according to the search text, a pre-established government affairs database cluster, and a pre-constructed government affairs domain dictionary. The government affairs database cluster stores database tables of government affairs data in multiple domains, each database table of government affairs data in each domain corresponds to a database, each database table includes multiple columns of government affairs data, and each column of government affairs data corresponds to a field. For each field, the field is an identifier of the corresponding column, and the field includes an indicator field and a non-indicator field. The target indicator is the indicator field, and the target entity is the non-indicator field. The government affairs domain dictionary includes multiple domain dictionaries, and each database corresponds to one of the domain dictionaries; A query statement construction module, configured to construct a query statement according to the entity word set, the target indicator, and the target entity; A search text query module, configured to determine a search result corresponding to the search text according to the query statement; The target text acquisition module includes: An entity recognition unit, configured to perform entity recognition on the search text to obtain an entity word set and a non-entity word set corresponding to the search text; A target indicator determination unit, configured to determine a target indicator corresponding to the search text according to the non-entity word set and the government affairs database cluster; A remaining text set determination unit, configured to determine a remaining text set in the search text according to the entity word set and the target indicator. The remaining text set is a text set corresponding to the search text except for the entity word set and the target indicator; A target library and table determination unit, configured to determine a target database table and a target database corresponding to the search text according to the target indicator. The target database table is a database table corresponding to the target indicator, and the target database is a database corresponding to the target database table; A target entity determination unit, configured to determine a target entity corresponding to the search text according to the remaining text set, the target database table, and the target database; The entity recognition unit includes: A first entity word determination subunit, configured to identify a first entity word included in the search text through a pre-constructed time and number recognition rule to obtain a first entity word set. For each of the first entity words, the first entity word represents an entity whose attribute includes at least one of time or number; A second entity word determination subunit, configured to identify a second entity word included in the search text through a pre-constructed entity detection model to obtain a second entity word set. For each of the second entity words, the second entity word represents an entity whose attribute includes at least one of location, organization name, or person name; The entity word set includes the first entity word set and the second entity word set; Among them, the target indicator determination unit includes: The first processing subunit is configured to determine a text matrix corresponding to the non-entity word set and a two-dimensional index matrix corresponding to a preset index data set through a pre-constructed similar text automatic generation model. The index data set is a set of index fields included in all database tables in the government affairs database cluster. Each element in the two-dimensional index matrix corresponds to a similar text generated by the similar text automatic generation model for one of the index fields. Each row of elements in the two-dimensional index matrix corresponds to a set of similar texts formed by all similar texts corresponding to one of the index fields. For each of the index fields, convert the set of similar texts corresponding to the index field into a one-dimensional index matrix corresponding to the index field. The second processing subunit is configured to, for each of the index fields in the index data set, determine a first similarity between the one-dimensional index matrix corresponding to the index field and the text matrix, and determine whether there is a pending index in the index data set whose first similarity is greater than a preset first similarity threshold according to each of the first similarities. The third processing subunit is configured to, if there is a pending index, determine a target index according to the first similarity corresponding to each of the pending indexes; if there is no pending index, execute the fourth processing subunit. The fourth processing subunit is configured to segment the non-entity word set to obtain a non-entity word segmentation result, where the non-entity word segmentation result includes multiple phrases; perform splicing processing according to the multiple phrases to generate a short sentence corresponding to the search text, and each short sentence is a text obtained by splicing at least one of the phrases. The fifth processing subunit is configured to, for each of the short sentences, determine a target index according to a second similarity between the short sentence and each of the index fields.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the intelligent search method for government affairs data as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the intelligent search method for government affairs data as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-data-source NL2SQL system based on semantic rules and multi-dimensional model
CN112559550A