Intelligent query system and method for power transmission and distribution production data based on natural language interaction

By constructing a cross-database semantic correlation and dynamic data map, the accuracy and efficiency of power transmission and distribution production data query are solved, data integration and sharing are realized, and the success rate of user query and system satisfaction are improved.

CN120256451AActive Publication Date: 2025-07-04ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD

Patent Information

Application Number
CN202510750132.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Traditional query methods are difficult to quickly and accurately obtain power transmission and distribution production data, especially when the results are interoperable between different databases, and the accuracy rate is low when the user's query intention is fuzzy.

Method used

By collecting multi-source database data, setting up hard rules and soft rules to establish cross-database semantic associations, building dynamic data maps, monitoring metadata increments for updates, and using natural language processing to extract query intents to generate visual results.

Benefits of technology

It realizes data fusion and sharing, improves query accuracy and efficiency, reduces user operation steps, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256451A_ABST
    Figure CN120256451A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent query system and method for power transmission and distribution production data based on natural language interaction, and relates to the technical field of knowledge maps. The method comprises the steps that a multi-source database in power transmission and distribution production is collected, a basic metadata unit is constructed, statistical characteristics are calculated, and a metadatabase is constructed; the method comprises the following steps: setting a hard rule and a soft rule for a metadatabase, establishing semantic association of cross-database fields by utilizing the two rules, and constructing an initial dynamic data graph; the dynamic data graph edges are endowed with weights through the relation strength; metadata increment is monitored, and an updating strategy is set to update the dynamic data atlas; an entity in the query statement is extracted, ambiguity judgment is carried out, the entity is mapped into a dynamic data graph to formulate an active guiding strategy, and a user query intention is obtained; reasoning in the dynamic data graph according to the query intention of the user to obtain a query result, and constructing a visual chart by using the query result to display the query result to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graphs, and particularly to an intelligent query system and method for power transmission and distribution production data based on natural language interaction. Background Art

[0002] With the continuous expansion of the power grid scale and the improvement of the informatization level of the power system, a large amount of data is generated in the power transmission and distribution production process, including equipment operation data, power grid topology data, maintenance records, fault reports, etc. These data are scattered in different systems and databases with various formats, and it is difficult to quickly and accurately obtain the required information through traditional query methods. The power transmission and distribution business involves multiple professional fields and links. When staff conduct data queries, they often need to comprehensively consider multiple factors and conditions. For example, to query the operation status of specific types of equipment in a certain area, it is necessary to associate multiple data sources such as equipment ledgers, real-time monitoring data, and historical fault records. The writing of traditional query statements is complex and requires high technical requirements for business personnel. With the continuous progress of artificial intelligence technologies such as natural language processing, knowledge graphs, and deep learning, new ideas and methods are provided for solving the problem of power transmission and distribution production data query. NLP technology can understand human natural language and convert the user's query intention into instructions executable by a computer.

[0003] However, when using natural language processing technology to analyze the user's query intention and simplify the user's data query process nowadays, when the user's question is vague and there are no clear keywords, it greatly affects the accuracy of intelligent query. Moreover, it is difficult to interoperate between different databases in power transmission and distribution production data, and the generated results have large errors when the user queries. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent query system and method for power transmission and distribution production data based on natural language interaction to solve the problems raised in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions: An intelligent query method for power transmission and distribution production data based on natural language interaction, the method comprising the following steps: S100. Collect data from multiple-source databases in power transmission and distribution production, standardize all data in the multiple-source databases, extract three types of metadata, namely field names, data types, and constraint conditions of the data in the multiple-source databases, construct a basic metadata unit and calculate statistical features, and construct a metadata database; Further, the specific steps for constructing a basic metadata unit and calculating statistical features are as follows: S101. Collect multi-source databases in power transmission and distribution production. The multi-source databases include relational databases, time series databases, and GIS spatial databases; scan the data in the three databases, and use the SQL query algorithm to extract the field names, data types, and constraint conditions of the data in the three databases to construct the basic metadata unit M i ={name i ,type i ,constraints i}; where M i represents the basic metadata unit of the i-th data, name i represents the field name of the i-th data extracted, type i represents the data type of the i-th data extracted, constraints i represents the constraint condition of the i-th data extracted; S102. Calculate the mean, variance, and skewness of each data. Use the mean, variance, and skewness of the data as the statistical features of each data, and use the statistical features to construct the enhanced metadata as M i * =M i ⋃{u i ,σ 2 i ,skewness i}, where M i * represents the i-th enhanced metadata, u i represents the mean of the i-th data, σ 2 i represents the variance of the i-th data, skewness i represents the skewness of the i-th data; Output the enhanced metadata of each data; Standardize the basic metadata unit and the enhanced metadata, and combine them to construct a metadata database.

[0006] By standardizing all the data in the multi-source database, the data format and specifications can be unified, the data inconsistency and errors can be reduced, and the data accuracy and integrity can be improved, thus providing a high-quality data foundation for subsequent data analysis and applications.

[0007] Extracting metadata such as field names, data types, and constraint conditions and constructing a metadata database makes the structure and attributes of the data clearer and more explicit, facilitating data managers and developers to understand the data, perform data maintenance, update, and management, and also helping new data users to get started quickly.

[0008] Calculating statistical features provides a basis for in-depth data analysis, enabling users to understand features such as data distribution and trends, and providing strong support for decision-making. For example, in power transmission and distribution production, the rules of equipment operation data can be discovered through statistical analysis to perform maintenance and fault prevention in advance.

[0009] S200. Set hard rules and soft rules for the meta-database, establish semantic associations across database fields using the two rules, and construct an initial dynamic data graph using the data and semantic associations in the meta-database; Furthermore, the specific steps for constructing an initial dynamic data graph using the data and semantic associations in the meta-database are as follows: S201. Set hard rules, specifically: Use a professional dictionary in the power field to search for field names in the meta-database, traverse the field names in the meta-database, and when a professional term identical to the field name in the meta-database is found in the professional dictionary in the power field, mark the field name in the meta-database using the data type and attributes of the professional term in the professional dictionary in the power field; Judge the marks of each field name in the meta-database, mark the field names with the same marks as forced associations, and finally output a hard rule association list; S202. Set soft rules, specifically: Segment the field names in the meta-database and convert them into vector format, calculate the cosine similarity of vectors of different field names in the meta-database, and the formula is: ; In the formula, S text represents the cosine similarity between the i-th and j-th field names, v i represents the i-th field name vector, and v j represents the j-th field name vector; Calculate the distribution similarity between different field names using JS divergence, and calculate the comprehensive similarity between different field names by combining cosine similarity and distribution similarity. The formula is: ; In the formula, S ij represents the comprehensive similarity between the i-th and j-th field names, and JS(p i ||p j ) represents the distribution similarity between the i-th and j-th field names; Set a similarity threshold Sy using experience. When S ij ≥Sy, establish an association between the i-th field name and the j-th field name; Finally, output a soft rule association matrix; S203. Combine hard rules and soft rules to establish associations for field names in different multi-source databases, set priorities, with the priority of hard rules being higher than that of soft rules; when the associations established by hard rules and soft rules for the same field name are different, give priority to the associations of hard rules; use the metadata in different multi-source databases as nodes and the associations in the hard rule association list and soft rule association matrix as edges to construct an initial dynamic data graph.

[0010] Set hard rules and soft rules to establish semantic associations for cross-database fields, connect the data originally scattered in different databases, break data islands, achieve data fusion and sharing, and improve the utilization value of data.

[0011] Constructing an initial dynamic data graph can visually display the relationships between data in a graphical way, facilitating users to quickly understand the associations and dependencies between data, discover potential patterns and rules, and provide a more comprehensive perspective for problem analysis and decision-making in power transmission and distribution production.

[0012] S300. Calculate the relationship strength between different nodes in the dynamic data graph and assign weights to the edges of the dynamic data graph using the relationship strength. Further, the specific steps for assigning weights to the edges of the dynamic data graph using the relationship strength are as follows: S301. Real-time collect the work logs in power transmission and distribution production within the past 24 hours, the co-occurrence times of different field names, where the co-occurrence times represent the number of times different field names appear simultaneously in the work logs, calculate the distance between different field names based on the GIS data in the GIS database, and calculate the relationship strength of each edge in the dynamic data graph. The formula is: ; In the formula, R ij represents the relationship strength of the edge between the i-th node and the j-th node, C(N i , N j ) represents the co-occurrence times of the i-th field name and the j-th field name, and L ij represents the distance between the field names of the i-th node and the j-th node. Assign weights to each edge in the dynamic data graph using the calculated relationship strength and output the dynamic data graph with weights.

[0013] Calculating the relationship strength between different nodes in the dynamic data graph and assigning weights to the edges can highlight the important relationships between data, enabling users to pay more attention to key associations when viewing the graph, improving the efficiency and pertinence of data analysis. For example, in fault troubleshooting, it is possible to quickly locate important equipment and data related to the fault.

[0014] S400 monitors the metadata increment, and updates the dynamic data graph according to the update policy set based on the metadata increment; Further, the specific steps for updating the dynamic data graph according to the metadata increment and setting the update policy are as follows: S401 monitors three types of metadata in the basic metadata unit and the statistical features in the enhanced metadata in the metadata database in real time. When a new field is added and the statistical features deviate in the metadata database, the associations and association strengths are recalculated in the dynamic data graph to update the dynamic data graph.

[0015] Monitoring the metadata increment and updating the dynamic data graph according to the increment setting can timely reflect the data changes, ensure the timeliness of the data, keep the graph always consistent with the actual data, and provide accurate information for users. For example, in the power transmission and distribution production, the real-time update of equipment operation data can help the operation and maintenance personnel timely master the equipment status.

[0016] S500 inputs the user's real-time query statement, extracts the entities in the query statement, performs fuzziness determination, maps the entities to the dynamic data graph to formulate an active guidance strategy, and obtains the user's query intention; Further, the specific steps for obtaining the user's query intention are as follows: S501 collects the user's real-time query statement, extracts the query statement entities, matches the query statement entities with the field names in the dynamic data graph to obtain the corresponding graph nodes, and calculates the fuzziness of the user's query statement. The formula is: ; In the formula, Fuzz represents the fuzziness of the user's query statement, T no represents the user's query statement entities not found during the matching in the dynamic data graph, In c represents the information entropy of the user's query statement, and Tz represents the total number of nodes in the dynamic data graph; S502, when it is judged that Fuzz > Fs, where Fs is the fuzziness threshold set according to historical query experience, triggers the system's active guidance; uses natural language processing algorithms to generate active guidance statements to ask the user, determines the user's query target range according to the successfully matched graph nodes, and enumerates all targets within the user's query target range by the enumeration method to determine the user's query intention.

[0017] Extracting the entities in the query statement, performing fuzziness determination, and mapping the entities to the dynamic data graph to formulate an active guidance strategy can more accurately understand the user's query intention, avoid inaccurate query results caused by unclear user expressions or incomplete query statements, and improve the success rate of user queries.

[0018] The active guidance strategy can help users express their needs more accurately, provide query results that better meet user expectations, reduce the number of user query operation steps, improve query efficiency, enhance the user experience, and increase user satisfaction with the system.

[0019] S600. Infer query results in the dynamic data graph according to the user's query intention, and use the query results to construct a visualization chart to display the query results to the user.

[0020] Furthermore, the specific steps for using the query results to construct a visualization chart to display the query results to the user are as follows: S601. Search in the dynamic data graph according to the user's query intention to obtain corresponding nodes, extract all user query intention nodes and associated metadata in the current dynamic data graph to form a data subgraph; extract each event node and timestamp in the data subgraph, and generate a user query result based on the time series. S602. Construct the generated user query result into a visualization chart and display it to the user.

[0021] Infer query results in the dynamic data graph according to the user's query intention, and use the query results to construct a visualization chart to display to the user, which can present data in an intuitive way, making it easier for users to understand the query results and quickly obtain key information, so as to make more effective decisions. For example, in the power transmission and distribution production management, the visualization chart can help managers quickly understand the production operation situation and make reasonable decisions.

[0022] The intelligent query system for power transmission and distribution production data based on natural language interaction, the intelligent query system for power transmission and distribution production data includes a data collection module, a meta-database construction module, a dynamic data graph module, a graph update module, a query statement analysis module, and a result generation module; The data collection module is used to collect data from multi-source databases and standardize all data in the multi-source databases; The meta-database construction module is used to extract three types of metadata, namely field names, data types, and constraint conditions of the data in the multi-source databases, construct basic metadata units and calculate statistical features, and construct a meta-database; The dynamic data graph module is used to set hard rules and soft rules for the meta-database, establish semantic associations across database fields using the two rules, construct an initial dynamic data graph using the data and semantic associations in the meta-database, and assign weights to the graph edges; The graph update module is used to monitor metadata increments and update the dynamic data graph according to the set update strategy based on the metadata increments; The query statement analysis module is used to input the user's real-time query statement, extract the entities in the query statement, determine the ambiguity, map the entities to the dynamic data graph to formulate an active guidance strategy, and obtain the user's query intention. The result generation module is used to infer the query result in the dynamic data graph according to the user's query intention, and construct a visualization chart using the query result to display the query result to the user.

[0023] The dynamic data graph module includes a rule setting unit and a weight calculation unit. The rule setting unit is used to set hard rules and soft rules for the meta database, and establish semantic associations across database fields using the two rules. The weight calculation unit is used to calculate the relationship strength between different nodes in the dynamic data graph, and assign weights to the edges of the dynamic data graph using the relationship strength.

[0024] The query statement analysis module includes a fuzzy judgment unit and an active guidance unit. The fuzzy judgment unit is used to calculate the ambiguity of the user's real-time query statement, and determine whether the ambiguity is greater than the ambiguity threshold to obtain whether the user's real-time query statement is ambiguous. The active guidance unit is used to generate an active guidance statement to ask the user using natural language processing algorithms when it is determined that the user's real-time query statement is ambiguous, and obtain the query intention.

[0025] Compared with the prior art, the beneficial effects of the present invention are: 1. The present invention sets hard rules and soft rules to establish semantic associations across database fields, connects the data originally scattered in different databases, breaks data islands, realizes data fusion and sharing, and improves the utilization value of data.

[0026] 2. The present invention deeply infers the user's query intention through the dynamic data graph, outputs the query result obtained by the user's query, provides a query result that better meets the user's expectations, and can also reduce the user's query operation steps and improve the query efficiency.

[0027] 3. The present invention extracts the entities in the query statement, determines the ambiguity, and maps the entities to the dynamic data graph to formulate an active guidance strategy, which can more accurately understand the user's query intention, avoid inaccurate query results caused by unclear user expressions or incomplete query statements, and improve the success rate of user queries. Description of the Drawings

[0028] Figure 1 It is the module distribution diagram of the intelligent query system for power transmission and distribution production data based on natural language interaction of the present invention. Figure 2Schematic diagram of the steps of the intelligent query method for power transmission and distribution production data based on natural language interaction according to the present invention. Detailed implementation manners

[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution An intelligent query method for power transmission and distribution production data based on natural language interaction, the method includes the following steps: S100. Collect data from multi-source databases in power transmission and distribution production, standardize all data in the multi-source databases, extract three types of metadata including field names, data types, and constraint conditions of the data in the multi-source databases, construct a basic metadata unit, calculate statistical features, and construct a metadata database; The specific steps for constructing the basic metadata unit and calculating statistical features are as follows: S101. Collect multi-source databases in power transmission and distribution production. The multi-source databases include relational databases, time series databases, and GIS spatial databases; scan the data in the three databases, and use the SQL query algorithm to extract the field names, data types, and constraint conditions of the data in the three databases, and construct a basic metadata unit M i ={name i , type i , constraints i}; where M i represents the basic metadata unit of the i-th data, name i represents the field name of the i-th data extracted, type i represents the data type of the i-th data extracted, and constraints i represents the constraint condition of the i-th data extracted; S102. Calculate the average value, variance, and skewness of each data, use the average value, variance, and skewness of the data as the statistical features of each data, and use the statistical features to construct enhanced metadata as M i * =M i ⋃{u i , σ 2 i , skewness i}, where M i *Denote the i-th enhanced metadata, u i Denote the average value of the i-th data, σ 2 i Denote the variance of the i-th data, skewness i Denote the skewness of the i-th data; output the enhanced metadata of each data; standardize the basic metadata unit and the enhanced metadata, and combine them to build a metadata database.

[0031] By standardizing all the data in the multi-source database, it is possible to unify the data format and specification, reduce data inconsistency and errors, improve data accuracy and integrity, and thus provide a high-quality data foundation for subsequent data analysis and applications.

[0032] Extract metadata such as field names, data types, and constraint conditions and build a metadata database, making the structure and attributes of the data clearer and more explicit, facilitating data managers and developers to understand the data, perform data maintenance, update and management, and also helping new data users to get started quickly.

[0033] Calculating statistical features provides a basis for in-depth data analysis, enabling users to understand data distribution, trends and other features, and providing strong support for decision-making. For example, in power transmission and distribution production, the rules of equipment operation data can be discovered through statistical analysis to perform maintenance and fault prevention in advance.

[0034] S200. Set hard rules and soft rules for the metadata database, establish semantic associations across database fields using the two rules, and construct an initial dynamic data graph using the data and semantic associations in the metadata database; The specific steps for constructing an initial dynamic data graph using the data and semantic associations in the metadata database are as follows: S201. Set hard rules, specifically: use a professional dictionary in the power field to search for field names in the metadata database, traverse the field names in the metadata database, and when a professional term identical to the field name in the metadata database is found in the professional dictionary in the power field, mark the field name in the metadata database using the data type and attributes of the professional term in the professional dictionary in the power field; judge the marks of each field name in the metadata database, mark the field names with the same marks as forced associations, and finally output a hard rule association list; S202. Set soft rules, specifically: segment the field names in the metadata database and convert them into vector format, calculate the cosine similarity of different field name vectors in the metadata database, and the formula is: ; In the formula, S text Denote the cosine similarity between the i-th and j-th field names, v i Denote the i-th field name vector, vj represents the j-th field name vector; the distribution similarity between different field names is calculated using JS divergence, and the comprehensive similarity between different field names is calculated by combining cosine similarity and distribution similarity. The formula is: ; In the formula, S ij represents the comprehensive similarity between the i-th and j-th field names, and JS(p i ||p j ) represents the distribution similarity between the i-th and j-th field names; the similarity threshold Sy is set empirically. When S ij ≥Sy, an association is established between the i-th field name and the j-th field name; finally, a soft rule association matrix is output. S203. Combine hard rules and soft rules to establish associations for field names in different multi-source databases, and set priorities. The priority of hard rules is higher than that of soft rules; when the associations constructed by hard rules and soft rules for the same field name are different, the association of hard rules is preferred; use the metadata in different multi-source databases as nodes, and use the associations in the hard rule association list and soft rule association matrix as edges to construct an initial dynamic data graph.

[0035] Set hard rules and soft rules to establish semantic associations for cross-database fields, connect the data originally scattered in different databases, break data islands, realize data fusion and sharing, and improve the utilization value of data.

[0036] Constructing an initial dynamic data graph can visually display the relationships between data in a graphical way, facilitate users to quickly understand the associations and dependencies between data, discover potential patterns and rules, and provide a more comprehensive perspective for problem analysis and decision-making in power transmission and distribution production.

[0037] S300. Calculate the relationship strength between different nodes in the dynamic data graph, and assign weights to the edges of the dynamic data graph using the relationship strength. The specific steps for assigning weights to the edges of the dynamic data graph using the relationship strength are as follows: S301. Real-time collect the work logs in power transmission and distribution production in the past 24 hours, the co-occurrence times of different field names. The co-occurrence times represent the number of times different field names appear simultaneously in the work logs. Calculate the distance between different field names based on the GIS data in the GIS database, and calculate the relationship strength of each edge in the dynamic data graph. The formula is: ; In the formula, R ij represents the relationship strength of the edge between the i-th node and the j-th node, C(N i , N jrepresents the co-occurrence times of the i-th field name and the j-th field name, L ij represents the distance between the i-th node and the j-th node field name; Use the calculated relationship strength to assign weights to each edge in the dynamic data graph, and output a weighted dynamic data graph.

[0038] Calculating the relationship strength between different nodes in the dynamic data graph and assigning edge weights can highlight the important relationships between data, enabling users to pay more attention to key associations when viewing the graph, improving the efficiency and pertinence of data analysis. For example, in fault troubleshooting, it is possible to quickly locate important devices and data related to the fault.

[0039] S400. Monitor the metadata increment, and set an update strategy according to the metadata increment to update the dynamic data graph; The specific steps for updating the dynamic data graph by setting an update strategy according to the metadata increment are as follows: S401. Real-time monitor the three types of metadata in the basic metadata unit in the metadata database and the statistical features in the enhanced metadata. When a new field is added or the statistical features deviate in the metadata database, recalculate the associations and association strengths in the dynamic data graph to update the dynamic data graph.

[0040] Monitoring the metadata increment and setting an update strategy according to the increment to update the dynamic data graph can timely reflect the changes in the data, ensure the timeliness of the data, keep the graph always consistent with the actual data, and provide accurate information for users. For example, in power transmission and distribution production, the real-time update of equipment operation data can help maintenance personnel timely master the equipment status.

[0041] S500. Input the user's real-time query statement, extract the entities in the query statement, perform fuzziness determination, map the entities to the dynamic data graph to formulate an active guidance strategy, and obtain the user's query intention; The specific steps for obtaining the user's query intention are as follows: S501. Collect the user's real-time query statement, extract the query statement entities, match the query statement entities in the dynamic data graph using the field names to obtain the corresponding graph nodes, and calculate the fuzziness of the user's query statement. The formula is: ; In the formula, Fuzz represents the fuzziness of the user's query statement, T no represents the user query statement entities that are not found during the matching in the dynamic data graph, In c represents the information entropy of the user's query statement, and Tz represents the total number of nodes in the dynamic data graph; S502. When it is determined that Fuzz > Fs, where Fs is the ambiguity threshold set according to historical query experience, the system actively guides. The system uses natural language processing algorithms to generate active guidance statements to ask the user, determines the user's query target range based on the successfully matched graph nodes, and enumerates all targets within the user's query target range by the enumeration method to determine the user's query intention.

[0042] Extract the entities in the query statement, perform ambiguity determination, and map the entities to the dynamic data graph to formulate an active guidance strategy, which can more accurately understand the user's query intention, avoid inaccurate query results caused by unclear user expressions or incomplete query statements, and improve the success rate of user queries.

[0043] The active guidance strategy can help users more accurately express their needs, provide query results that better meet user expectations, reduce the number of user query operation steps at the same time, improve query efficiency, enhance the user experience, and increase user satisfaction with the system.

[0044] S600. Infer the query result in the dynamic data graph according to the user's query intention, and use the query result to construct a visualization chart to display the query result to the user.

[0045] The specific steps of using the query result to construct a visualization chart to display the query result to the user are as follows: S601. Search in the dynamic data graph according to the user's query intention to obtain the corresponding nodes, extract all user query intention nodes and associated metadata in the current dynamic data graph to form a data subgraph; extract each event node and timestamp in the data subgraph to generate a user query result based on the time series. S602. Construct the generated user query result into a visualization chart and display it to the user.

[0046] Infer the query result in the dynamic data graph according to the user's query intention, and use the query result to construct a visualization chart to display it to the user, which can present data in an intuitive way, make it easier for users to understand the query result, quickly obtain key information, and thus make more effective decisions. For example, in power transmission and distribution production management, the visualization chart can help managers quickly understand the production operation situation and make reasonable decisions.

[0047] The intelligent query system for power transmission and distribution production data based on natural language interaction, where the intelligent query system for power transmission and distribution production data includes a data collection module, a metadata database construction module, a dynamic data graph module, a graph update module, a query statement analysis module, and a result generation module; The data collection module is used to collect data from multi-source databases and standardize all data in the multi-source databases. The meta-database construction module is used to extract three types of metadata, namely field names, data types, and constraint conditions, from multi-source databases, construct basic metadata units, calculate statistical features, and construct a meta-database; The dynamic data graph module is used to set hard rules and soft rules for the meta-database, establish semantic associations across database fields using the two rules, construct an initial dynamic data graph using the data and semantic associations in the meta-database, and assign weights to the graph edges; The graph update module is used to monitor metadata increments and update the dynamic data graph according to the update strategy set based on the metadata increments; The query statement analysis module is used to input a user's real-time query statement, extract entities in the query statement, determine the ambiguity, map the entities to the dynamic data graph to formulate an active guidance strategy, and obtain the user's query intention; The result generation module is used to infer query results in the dynamic data graph according to the user's query intention, and construct a visualization chart using the query results to display the query results to the user.

[0048] The dynamic data graph module includes a rule setting unit and a weight calculation unit; The rule setting unit is used to set hard rules and soft rules for the meta-database and establish semantic associations across database fields using the two rules; The weight calculation unit is used to calculate the relationship strength between different nodes in the dynamic data graph and assign weights to the edges of the dynamic data graph using the relationship strength.

[0049] The query statement analysis module includes a fuzzy judgment unit and an active guidance unit; The fuzzy judgment unit is used to calculate the ambiguity of the user's real-time query statement and determine whether the ambiguity is greater than the ambiguity threshold to obtain whether the user's real-time query statement is ambiguous; The active guidance unit is used to generate an active guidance statement to ask the user when it is determined that the user's real-time query statement is ambiguous, and obtain the query intention.

[0050] Embodiment 1: Extraction and association of transformer temperature fields; Extract the field temp from the transformer table in InfluxDB, with the type being Float and the constraint being non-null.

[0051] Calculate the statistical features, with the average being 85.3 degrees Celsius, the variance being 4.1, and the skewness being 2; standardize the field name to transformer_temp; The corresponding professional term cannot be found in the professional dictionary in the power field using the hard rule; the soft rule is selected to calculate the similarity, and the comprehensive similarity with the device_temp field name in the relational database is 0.72. Assuming the similarity threshold is 0.6, it is determined that there is an association between the field names transformer_temp and device_temp:

[0052] Example 2: The user's query statement is "Which area has the most faults recently?" The extracted entities are "faults, areas, etc." Keywords such as "fault" and "area" are matched to the graph nodes, and the fuzziness is calculated to be 0.64. Assuming the fuzziness threshold is 0.5, it is determined that the user's query is fuzzy. Active guidance is started. Based on the graph nodes of the user's query, the user's query target range is obtained as "East China, a certain city, one month, one week, etc." The natural language processing algorithm is used to generate an active guidance question as "Please select the area range: A. East China region B. Distribution network of a certain city", and it is asked sequentially using the enumeration method; the user's query intention is obtained as "Count the number of faults in the East China region in the past 7 days."

[0053] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

Claims

1. An intelligent query method for power transmission and distribution production data based on natural language interaction, characterized in that: The method includes the following steps: S100. Collect multi-source databases in power transmission and distribution production, standardize all data in the multi-source databases, extract three types of metadata, namely field names, data types, and constraint conditions of the data in the multi-source databases, construct basic metadata units and calculate statistical features, and construct a metadata database; S200. Set hard rules and soft rules for the metadata database, establish semantic associations across database fields using the two rules, and construct an initial dynamic data graph using the data and semantic associations in the metadata database; S300. Calculate the relationship strength between different nodes in the dynamic data graph, and assign weights to the edges of the dynamic data graph using the relationship strength; S400. Monitor metadata increments, and set update strategies according to the metadata increments to update the dynamic data graph; S500. Input a user's real-time query statement, extract entities in the query statement, perform fuzziness determination, map the entities to the dynamic data graph to formulate an active guidance strategy, and obtain the user's query intention; S600. Infer query results in the dynamic data graph according to the user's query intention, and construct a visualization chart using the query results to display the query results to the user.

2. The intelligent query method for power transmission and distribution production data based on natural language interaction according to claim 1, wherein: The specific steps for constructing basic metadata units and calculating statistical features in S100 are as follows: S101. Collect multi-source databases in power transmission and distribution production. The multi-source databases include relational databases, time-series databases, and GIS spatial databases. Scan the data in the three databases, and use the SQL query algorithm to extract the field names, data types, and constraint conditions of the data in the three databases, and construct the basic metadata unit M i ={name i , type i , constraints i}; where M i represents the basic metadata unit of the i-th data, name i represents the field name of the i-th data extracted, type i represents the data type of the i-th data extracted, and constraints i represents the constraint condition of the i-th data extracted; S102. Calculate the mean, variance, and skewness of each data. Use the mean, variance, and skewness of the data as the statistical features of each data, and construct enhanced metadata M using the statistical features i * =M i ⋃{u i , σ 2 i , skewness i}, where M i * represents the i-th enhanced metadata, u i represents the mean of the i-th data, σ 2 i represents the variance of the i-th data, and skewness i represents the skewness of the i-th data; Output enhanced metadata for each data; Standardize the basic metadata units and enhanced metadata, and combine them to construct a metadata database.

3. The intelligent query method for power transmission and distribution production data based on natural language interaction according to claim 2, wherein: The specific steps for constructing an initial dynamic data graph using the data and semantic associations in the metadata database in S200 are as follows: S201. Set hard rules, specifically: use a professional dictionary in the power field to search for field names in the metadata database, traverse the field names in the metadata database, and when a professional term identical to the field name in the metadata database is found in the professional dictionary in the power field, mark the field name in the metadata database using the data type and attributes of the professional term in the professional dictionary in the power field; Judge the markings of each field name in the metadata database, mark the field names with the same markings as forced associations, and finally output a hard rule association list; S202. Set soft rules, specifically: segment the field names in the metadata database and convert them into vector formats, calculate the cosine similarity of the vectors of different field names in the metadata database, and the formula is: ; In the formula, S text represents the cosine similarity between the i-th and j-th field names, v i represents the i-th field name vector, v j represents the j-th field name vector; the distribution similarity between different field names is calculated using the JS divergence, and the comprehensive similarity between different field names is calculated by combining the cosine similarity and the distribution similarity. The formula is as follows: ; In the formula, S ij represents the comprehensive similarity between the i-th and j-th field names, and JS(p i ||p j ) represents the distribution similarity between the i-th and j-th field names; the similarity threshold Sy is set empirically. When S ij ≥Sy, an association is established between the i-th field name and the j-th field name; finally, the soft rule association matrix is output. S203. Combine the hard rules and soft rules to construct associations of field names in different multi-source databases, set priorities, and the priority of the hard rules is greater than that of the soft rules; When the associations constructed by the hard rules and soft rules for the same type of field name are different, give priority to the associations of the hard rules; Use the metadata in different multi-source databases as nodes, and use the associations in the hard rule association list and the soft rule association matrix as edges to construct an initial dynamic data graph.

4. The intelligent query method for power transmission and distribution production data based on natural language interaction according to claim 3, characterized in that: The specific steps for assigning weights to the edges of the dynamic data graph using the relationship strength in S300 are as follows: S301. Collect the work logs in power transmission and distribution production within the past 24 hours in real time, the co-occurrence times of different field names, where the co-occurrence times represent the number of times different field names appear simultaneously in the work logs. Calculate the distances between different field names based on the GIS data in the GIS database, and calculate the relationship strength of each edge in the dynamic data graph. The formula is: ; In the formula, R ij represents the relationship strength of the edge between the i-th node and the j-th node, C(N i , N j ) represents the co-occurrence times of the i-th field name and the j-th field name, and L ij represents the distance between the field names of the i-th node and the j-th node; Assign weights to each edge in the dynamic data graph using the calculated relationship strength, and output the weighted dynamic data graph.

5. The intelligent query method for power transmission and distribution production data based on natural language interaction according to claim 4, characterized in that: The specific steps for updating the dynamic data graph according to the metadata increment setting update strategy in S400 are as follows: S401. Monitor the three types of metadata in the basic metadata unit and the statistical features in the enhanced metadata in the metadata database in real time. When a new field is added or the statistical features deviate in the metadata database, recalculate the associations and association strengths in the dynamic data graph to update the dynamic data graph.

6. The intelligent query method for power transmission and distribution production data based on natural language interaction according to claim 5, characterized in that: The specific steps for obtaining the user's query intention in S500 are as follows: S501. Collect the user's real-time query statement, extract the query statement entity, match the query statement entity with the field names in the dynamic data graph to obtain the corresponding graph nodes, and calculate the fuzziness of the user's query statement. The formula is: ; In the formula, Fuzz represents the fuzziness of the user query statement, and T no represents the user query statement entity that is not found when matching in the dynamic data graph, and In c represents the information entropy of the user query statement, and Tz represents the total number of nodes in the dynamic data graph; S502. When it is determined that Fuzz > Fs (where Fs is the fuzziness threshold set according to historical query experience), trigger the system's active guidance. Use natural language processing algorithms to generate an active guidance statement to ask the user, determine the user's query target range based on the successfully matched graph nodes, and enumerate all the targets within the user's query target range to determine the user's query intention.

7. The intelligent query method for power transmission and distribution production data based on natural language interaction according to claim 6, wherein: The specific steps for constructing a visualization chart using the query results to display the query results in S600 are as follows: S601. Search in the dynamic data graph according to the user's query intention to obtain the corresponding nodes. Extract all the user query intention nodes and the associated metadata in the current dynamic data graph to form a data subgraph. Extract each event node and timestamp in the data subgraph to generate the user's query results based on the time series. S602. Construct the generated user query results into a visualization chart and display it to the user.

8. An intelligent query system for power transmission and distribution production data based on natural language interaction, characterized in that: The intelligent query system for power transmission and distribution production data includes a data collection module, a metadata database construction module, a dynamic data graph module, a graph update module, a query statement analysis module, and a result generation module. The data collection module is used to collect data from multi-source databases and standardize all the data in the multi-source databases. The metadata database construction module is used to extract three types of metadata, namely the field names, data types, and constraint conditions of the data in the multi-source databases, construct the basic metadata unit and calculate the statistical features, and construct the metadata database. The dynamic data graph module is used to set hard rules and soft rules for the metadata database, establish semantic associations between cross-database fields using the two rules, construct an initial dynamic data graph using the data and semantic associations in the metadata database, and assign weights to the graph edges. The graph update module is used to monitor the metadata increment and update the dynamic data graph according to the metadata increment setting update strategy. The query statement analysis module is used to input the user's real-time query statement, extract the entities in the query statement, determine the ambiguity, map the entities to the dynamic data graph to formulate an active guidance strategy, and obtain the user's query intention; The result generation module is used to infer the query result in the dynamic data graph according to the user's query intention, and use the query result to construct a visualization chart to display the query result to the user.

9. The intelligent query system for power transmission and distribution production data based on natural language interaction according to claim 8, wherein: The dynamic data graph module includes a rule setting unit and a weight calculation unit; The rule setting unit is used to set hard rules and soft rules for the meta-database, and establish semantic associations across database fields using the two rules; The weight calculation unit is used to calculate the relationship strength between different nodes in the dynamic data graph, and assign weights to the edges of the dynamic data graph using the relationship strength.

10. The intelligent query system for power transmission and distribution production data based on natural language interaction according to claim 8, wherein: The query statement analysis module includes a fuzzy judgment unit and an active guidance unit; The fuzzy judgment unit is used to calculate the ambiguity of the user's real-time query statement, and judge whether the ambiguity is greater than the ambiguity threshold to obtain whether the user's real-time query statement is ambiguous; The active guidance unit is used to generate an active guidance statement to ask the user using natural language processing algorithms when it is judged that the user's real-time query statement is ambiguous, and obtain the query intention.

Citation Information

Patent Citations

  • Power data intelligent search method based on resource map

    CN119336831A

  • Webpage content security processing method based on knowledge graph

    CN119670070A

  • Cross-unit data management method

    CN119989418A

  • Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search

    US20240386015A1

Cited By

  • Production data query method and device based on atlas

    CN121166713A

  • A method and apparatus for querying production data based on graphs

    CN121166713B

  • Human-computer interaction dialogue method, system and equipment based on natural language and medium

    CN121257719A