A power supply reliability all-service intelligent retrieval method and system

CN117149815BActive Publication Date: 2025-12-19STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311110834.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-30
Publication Date
2025-12-19
Estimated Expiration
2043-08-30

AI Technical Summary

Technical Problem

[0004]本发明所要解决的技术问题在于现有技术检索方法无法针对专业性比较强的供电可靠性管理方面的数据进行转换和查询,从而容易导致检索失败

Benefits of technology

[0038] (1) The power supply reliability management system data is stored in the graph database, and the natural language is standardized, instead of directly querying the natural language, so that the problem that the natural language cannot accurately query the result due to too strong professional nature is avoided, then the words are segmented and the word types are screened, and the natural language is converted into a cypher query statement based on the matching rule and the decision tree, so that the problems of insufficient generalization ability based on the matching rule, underfitting and overfitting of the decision tree are effectively avoided, the query accuracy is further improved, and the retrieval failure is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117149815B_ABST
    Figure CN117149815B_ABST
Patent Text Reader

Abstract

The application discloses a power supply reliability full-service intelligent retrieval method and system, and the method comprises the following steps: storing power supply reliability management system data into a graph database; preprocessing natural language to obtain a standardized input sentence; performing word segmentation on the standardized input sentence, performing word segmentation screening on the word segmentation list, and generating a corresponding cypher query statement based on a rule matching principle and a decision tree; generating a cypher query statement based on rule matching and a decision tree for the user input sentence, respectively, performing consistency calculation on the coded results of the two, when there is consistency, querying the corresponding data through the cypher query statement, and when the two are inconsistent, converting the cypher query statement generated by the decision tree into corresponding natural language to ask the inputter for confirmation; the application has the advantages that the query accuracy is improved, and retrieval failure is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data retrieval, in particular to a power supply reliability full-service intelligent retrieval method and system. BACKGROUND

[0002] Power supply reliability management plays an increasingly important role. In the current domestic mainstream power supply reliability management system, the retrieval of various services is realized by menu options, different menus are entered to retrieve different types of services, and different parameter requirements of a certain service type are retrieved by selecting or assigning the set parameters in each menu. On the one hand, power supply reliability management involves many types of service retrieval, including retrieval and query of line, line segment, and user equipment account, query of basic power outage information such as power outage event, power outage line segment, and power outage user, query of repeated power outage situations such as repeated power outage line segment and repeated power outage user, and query of special power outages such as power outage of more than 100 households, power outage of more than 100 times, and power outage longer than 12 hours. The parameters to be set during retrieval are numerous and professional, including the setting of multiple parameters such as region characteristics, power outage time, power outage nature, responsibility reason, time and number of households, number of times of power outage, and time length. The concepts and calculation rules of power outage time and number of households, the number of times of power outage, the classification and division method of power outage responsibility reason, and the region characteristics form a system compared with other professions, and it requires a lot of effort for non-power supply reliability professionals to master the existing system service retrieval method, which brings a lot of manpower consumption to reliability professional management.

[0003] Chinese patent publication No. CN114625754A discloses a sentence query method, device, electronic equipment and computer readable storage medium, relating to the technical field of data query. The method includes: obtaining a query sentence expressed in natural language; performing word segmentation on the query sentence to obtain a target segmentation result; the target segmentation result includes at least one word and the category to which the at least one word belongs; the number of words contained in each category is determined respectively, and the target graph style corresponding to the query result of the query sentence is determined according to the number of words contained in each category; and the query result of the target graph style is displayed. This avoids single display of the graph corresponding to the query result, realizes flexible display of the graph, and improves user experience. However, this patent application directly queries the natural language, and in the aspect of power supply reliability management, the data types are relatively professional. A business content may have many different expressions, and direct natural language query will result in incorrect results and retrieval failure. SUMMARY

[0004] The technical problem to be solved by the present application is that the existing retrieval method cannot convert and query data in the aspect of power supply reliability management which is relatively professional, thereby easily leading to retrieval failure.

[0005] The application solves the above technical problems by the following technical means: a power supply reliability full-service intelligent retrieval method, comprising the following steps:

[0006] Step one: store power supply reliability management system data into a graph database;

[0007] Step two: preprocess natural language to obtain standardized input sentences;

[0008] Step three: segment the standardized input sentences, perform part-of-speech filtering on the segmented word list, establish judgment rules for word nodes and corresponding word conversion rules based on rule matching principles, thereby converting words of different parts of speech into cypher query statements, or constructing a business data set, inputting the business data set into a decision tree, training the decision tree, and generating corresponding cypher query statements from the trained decision tree;

[0009] Step four: generate cypher query statements from user input sentences according to rule matching and decision tree, respectively, and perform consistency calculation on the coded results of the two, when there is consistency, query the corresponding data through the cypher query statement, when the two are inconsistent, save the input sentence and the generated cypher query statement as difficult data, convert the cypher query statement generated by the decision tree into corresponding natural language to ask the inputter for confirmation, and obtain positive replies to query by the cypher query statement generated by the decision tree, and negative replies to query by the cypher query statement converted by rule matching.

[0010] Further, the step one comprises:

[0011] Store the power supply reliability management system data into a neo4j graph database, establish a power supply reliability full-service data knowledge graph, including substation, line, line segment, user, these types of equipment, various levels of units, and power outage events, power outage lines, power outage users, these operation data, establish jurisdictional connection between units and jurisdictional equipment, establish mounting connection between various levels of equipment, establish containing relationship connection between various levels of power outage data, and establish power outage connection between equipment and power outage data.

[0012] Further, the step two comprises:

[0013] Step 201: format conversion is performed on date data;

[0014] Step 202: load the part-of-speech corresponding to the symbol with a dot into the graph database;

[0015] Step 203: convert the expression manner with inconsistent format in the input sentence into a unified format;

[0016] Step 204: multi-dimensionally encode the users according to attributes, unit level, and unit name, and multi-dimensionally encode the equipment according to attributes and equipment type, and load the results into the graph database; multi-dimensionally encode the business words, and add the results to the graph database.

[0017] Further, the step 201 includes:

[0018] system current date and time information; convert the Chinese character representation of the year, month, and quarter into a digital + year or month format; extract the year or month or day from the string containing the year, month, and day; convert the date data into a fixed format string; complete the year and month attributes of the date, and complete the month according to the query time for months without months, and complete the year according to the query time for years without years; convert the date representing a time period into two time points.

[0019] Further, the step 204 includes:

[0020] The unit includes four levels of units of provincial company, municipal company, county company, and power supply station, which correspond to the first two dimensions of the word as dw.sh, dw.sgs, dw.xgs, and dw.gds, wherein dw represents the attribute of the unit, and sh, sgs, xgs, and gds represent the unit level of the provincial company, municipal company, county company, and power supply station, respectively, such as a municipal company, whose word encoding is dw.sgs.name, and name represents the unit name. The equipment includes substation, line, line segment, and user, which correspond to the first two dimensions of the word as sb.bdz, sb.xl, sb.xd, and sb.yh, wherein sb represents the attribute of the equipment, and bdz, xl, xd, and yh represent the equipment type of substation, line, line segment, and user, respectively.

[0021] The word of the business word includes the following categories: the first category, the first dimension of the word is yw, such as the word of the public variable, which is yw.tz.yhxz, wherein the first dimension, the data of the type yw is the corresponding data in the graph database, the second dimension tz represents the account data, and the third dimension yhxz indicates the user nature field; the second category is ywt, such as the word of the cancellation date, which is ywt.tz.zxrq, wherein ywt indicates that the word is a field name in the database system related to the date, and the third dimension zxrq represents the responsibility reason field; the third category is ywh, such as the word of the event serial number, which is ywh.yx.sjxh, wherein the first dimension ywh indicates that the word is a field name in the system with the storage type of number and symbol string, yx represents the running data, and sjxh represents the event serial number field.

[0022] Further, the research and judgment rules for establishing word nodes based on the rule matching principle and the corresponding word conversion rules are established, so as to convert words of different parts of speech into cypher query statements, including:

[0023] (1) The device sb class takes out the corresponding part of speech third symbol, and the third symbol is the field corresponding to the sb class in the system. The corresponding cypher query field is obtained through the python string splicing technology; the yw class is similar to the sb class.

[0024] (2) The dw class compares the unit information carried by the string with the unit information of the query account itself. If the carrying unit is within the jurisdiction of the query account unit, the carrying unit is taken as the condition to join the cypher query statement. If it is not within the authority of the query account unit, the unit information of the query account itself is inputted and the carrying unit information is ignored.

[0025] (3) The ywt class processes the user input statement, ywt word list and time word list through the ywt function. When the ywt word list is empty and the time word list has data, and when the user input statement has words such as "power failure" and "heavy power failure" relationship operation information, the time word list is converted into cypher query statement segment as the time data of the power failure time field. When the ywt word list and the time word list are not empty, they are matched in the ratio of 1:2, and each ywt word range is converted into a cypher query statement segment between two time words. For example, the start time of someone is queried. The start time has a time interval, and the word to be queried is limited between two time points, which is also between two time words. The ywt function refers to the process of "when the ywt word list is empty and the time word list has data, and when the user input statement has words such as "power failure" and "heavy power failure" relationship operation information, and each ywt word range is converted into a cypher query statement segment between two time words".

[0026] (4) The ywh class contains time, household, time length, abandoned power and other operation data attribute words, which are converted into cypher query statement segments together with the combination of more than X comparison words and numbers followed by them.

[0027] (5) ywsb, ywdw, as the extension of sb, dw class judgment rules, when ywsb and sb are the same type of equipment, do not add new statements, when the two levels are different, add the relationship statement between ywsb and sb type equipment, ywsb points to sb, directly query sb; when ywdw and dw are the same type of unit, do not add new statements, when the two levels are different, add the relationship statement between ywdw and dw type unit, ywdw points to dw, directly query dw; ywsb, ywdw is a generalization of sb, dw class expression, such as sb can represent the specific type of equipment, and ywsb can only represent that it belongs to the device, but the specific device type is not clear, so ywsb, ywdw is a more general rule.

[0028] (6) ywbj class, including greater than, less than, equal to these comparison classes, need to be adjacent data, date class combination into a whole, with adjacent ywt, ywh type association, together constitute its research and judgment rules;

[0029] The word segmentation sequence obtained by the user input statement is sequentially and sequentially judged according to the above rules, and the class condition node is lit for the condition that meets the condition, and the corresponding code conversion rule is operated to complete the conversion of all words in the word segmentation sequence, and the corresponding cypher statement segment is obtained. According to the cypher syntax rule, the code segments are combined to complete the cypher query statement.

[0030] Further, the business data set is constructed, comprising:

[0031] According to the device type query, the name fuzzy search, the device area feature classification query, the device running state query, the device belongs to the superior device, the query of the device belongs to the multi-level superior device, the device belongs to each level of unit query, the device hangs the next level device or multi-level next level device query, the device in a time period, a power failure property or several power failure properties, a power failure responsibility reason or several responsibility reasons, the power failure situation query, the device in a time period, a power failure property or several power failure properties or a power failure responsibility reason or several responsibility reasons cause the power failure range to exceed the input value, the power failure situation query, the device hanging device in a time period, a power failure property or several power failure properties or a power failure responsibility reason or several responsibility reasons cause the power failure range to exceed the input value, the power failure situation query, the device account information query contained in a power failure, the device account information query contained in a power failure of a certain area feature, the device details query of a certain area feature range in a unit caused by or / and a certain power failure property more than X times of power failure, the device details query of a certain area feature range in a unit caused by or / and a certain power failure property more than X of power failure, the device details query of a certain area feature range in a unit caused by or / and a certain power failure property more than X hours of power failure, the statistical query of a certain responsibility reason or power failure property and a certain area feature range in a unit within a certain time period, the total time or number of households affected by power failure, the device query of a certain responsibility reason or power failure property and a certain area feature range in a unit within a certain time period, the total time or number of households affected by power failure, the maximum or minimum of the total time or number of households affected by power failure.

[0032] Further, the inputting the business data set into the decision tree, training the decision tree, generating the corresponding cypher query statement by the trained decision tree comprises:

[0033] Based on the business data set, the word segmentation sequence obtained by the user input statement is matched in order, and the conversion of all words in the word segmentation sequence is completed by operating according to the corresponding coding conversion rule, to obtain the corresponding cypher statement segment. The cypher query statement is formed by combining the code segments according to the cypher syntax rule to form a test data set. The word segmentation and the cypher query statement segment are coded to generate the decision tree of the word segmentation to the cypher query statement segment and the cypher query statement segment to the word segmentation decision tree. The difference between the cypher statement segment generated by the decision tree and the cypher statement segment corresponding to the test data set is used to build a loss function, the parameters of the decision tree are adjusted, the decision tree is trained, and the training is stopped until the loss function value is minimum or the preset training times are reached. The corresponding cypher query statement is generated by using the trained decision tree.

[0034] Further, the step four further comprises:

[0035] The result feedback option is provided for the returned data, and if the user does not select or selects satisfaction, it indicates that the query is valid, and the difficult data and processing mode are stored in the difficult processing set as a subsequent query reference, if the user selects that the query result does not match the query target, it indicates that the query generation has a problem, and is put into the problem data set, and subsequently, the problem data set is added to the decision tree training data set to retrain and update the decision tree.

[0036] The application further provides a power supply reliability full-service intelligent retrieval system, which comprises a storage medium, and the storage medium stores a computer program, and the computer program can be executed by a processor to complete the above method.

[0037] The application has the following advantages:

[0038] (1) The power supply reliability management system data is stored in the graph database, and the natural language is standardized, instead of directly querying the natural language, so that the problem that the natural language cannot accurately query the result due to too strong professional nature is avoided, then the words are segmented and the word types are screened, and the natural language is converted into a cypher query statement based on the matching rule and the decision tree, so that the problems of insufficient generalization ability based on the matching rule, underfitting and overfitting of the decision tree are effectively avoided, the query accuracy is further improved, and the retrieval failure is avoided.

[0039] (2) The application converts the user input natural language into a cyper query statement based on the knowledge graph and word segmentation, searches the corresponding business data from the graph database, and realizes all business queries in a natural language manner in one query interface, compared with a conventional system, the system page switching is avoided, the professional personnel system function learning cost is saved, the professional personnel employment difficulty is reduced, and the time spent on system operation in daily power supply reliability management analysis is reduced.

[0040] (3) The application creates a multi-dimensional professional word type system according to the power supply reliability business data characteristics, realizes professional and fine marking of the word type of the reliability data, optimizes the word segmentation function, realizes support for the multi-dimensional word type system user dictionary and support for the Roman numeral and Chinese character date type date representation mode, summarizes the power supply reliability daily management query business type, establishes an intelligent decision tree, realizes accurate conversion of the natural language to the power supply reliability business retrieval from multiple dimensions, and realizes the power supply reliability full-service intelligent retrieval function. BRIEF DESCRIPTION OF DRAWINGS

[0041] Figure 1 A flowchart of a power supply reliability full-service intelligent retrieval method disclosed by the embodiment of the application. DETAILED DESCRIPTION

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] Embodiment 1

[0044] As Figure 1 shown, a full-service intelligent retrieval method for power supply reliability includes the following steps:

[0045] S1: Store the data of the power supply reliability management system into the graph database; the specific process is as follows:

[0046] Store the data of the power supply reliability management system into the neo4j graph database, and establish a knowledge graph of full-service data for power supply reliability, including equipment of types such as substations, lines, line segments, and users, various hierarchical units, as well as operation data such as power outage events, power outage line segments, and power outage users. Establish a jurisdiction connection between the unit and the subordinate equipment, establish a mounting connection between equipment at all levels, establish an inclusion relationship connection between power outage data at all levels, and establish a power outage connection between the equipment and the power outage data.

[0047] S2: Preprocess the natural language to obtain a standardized input statement; the specific process is as follows:

[0048] S201: Convert the format of the date data; specifically, the preliminary processing of the date type is encapsulated in a class named now. The functions of the now class include:

[0049] 1) The trans function, implemented through a match case statement, functions to convert Chinese characters representing months such as this year, last year, current year, first half of the year, current month, this month, last month, first quarter, second quarter, third quarter, fourth quarter, second half of the year, and January, February, etc. into the format of number + year (month) of date and time. For example, when the input is the first quarter, the output is January to March.

[0050] 2) The Dif function functions to extract, according to needs, the 2023 year, or June, or 5th day from a string in the format of 2023-06-05 for the input function through a regular expression. The regular expression is "(\d{2,4} year)(\d{1,2} month)(\d{1,2} day)?". The function takes two parameters, one is the date string, and the other is the date flag to be extracted. Taking 2023-06-05 as an example, when the flag parameter is "y", 2023 year is extracted; when it is "m", June is extracted; when it is "d", 5th day is extracted.

[0051] 3) s2t function, the input string is converted into python date data using the striptime() function of the datetime module of python.

[0052] 4) t2s function, the python date data is converted into a fixed format string format "%Y-%m-%d %H:%M:%S" using the strftime() function of the datetime module of python. Through the combination of s2t and t2s functions, the date string in the format of 2023-5-6 is unified into the format of 2023-5-6.

[0053] 5) Get function: use the now function of the datetime module of python to obtain the current date and time information. Take the current time August 21, 2023 as an example, the function has an input parameter tag, when tag is "y", take out the current year 2023, when tag is "m", take out the current month August, when tag is "ym", take out the current year and month 2023 August. In addition, d: day, md: month and day, ymd: year, month and day.

[0054] In the trans function, the get function is called to obtain this year, last year, this month, last month, and other date and time information related to the query time. Through the combination of trans, s2t and t2s functions, this year, January and other date formats are unified into the format of 2023-5-6.

[0055] 6) cmplt function, complete the year and month attributes of the date. If the month is not included, it is completed according to the month of the query time, and if the year is not included, it is completed according to the year of the query time. For example, July, in 2023, it will be completed as 2023-7.

[0056] 7) gstime function, convert a date representing a time period into two time points, such as 2023-7, which is converted into the time period from 2023-7-1 to 2023-8-1. The conversion method is to use the s2t and t2s functions of the nowday module, first convert the date string into python date using the s2t function, and use the timedelt function of the datime library and the relativedelta function of the dateutil.relativedelta library to realize the addition and subtraction of years, days and months.

[0057] S202: load the part-of-speech corresponding to the symbol with dots into the graph database;

[0058] The function is mainly implemented by the add_dic function. In the current regular expression, the part of speech does not contain “.”, so for words containing “.”, it cannot be correctly loaded from the dictionary. In the re-implemented function, the loading method of the word is changed to load the corresponding part of speech of the symbol with a dot. The user dictionary is read line by line, and the function is used to obtain the word body, frequency and part of speech of the word.

[0059] S203: converting the expression format in the input sentence that is not uniform into a uniform format;

[0060] The function can be implemented by the cut function. The cut function: through the cut function, the input of the user search box is segmented. The first step is to extract the date words in the input sentence such as 2023 May and the like through regular expression, temporarily add the jieba user dictionary, and realize the recognition of the date by jieba. The second step is to replace the non-uniform voltage level expression in the user input sentence such as kilovolt, KV, kv and the like into kV which is uniformly used in the system. The third step is to segment the user input sentence, take out the date data with the part of speech t, and uniformly convert the date words expressed by pure Chinese characters such as this year and July into date strings in the format of 2023 July through the trans function of nowday class, and replace the date data in the user input sentence with the converted date string. The fourth step is to dynamically add the jieba user dictionary in the user input sentence converted in the format of 2023 July to the jieba user dictionary by repeating the first step. The fifth step is to segment the processed user input sentence again. The segmentation result is returned in list format, and each word includes two parameters of word and part of speech.

[0061] S204: multi-dimensional part of speech coding of users according to attributes, unit level and unit name, multi-dimensional part of speech coding of equipment according to attributes and equipment type, and loading the results into the graph database; multi-dimensional part of speech coding of business words, and adding the results to the graph database, the specific process is:

[0062] The units include four levels of units of provincial companies, municipal companies, county companies and power supply stations, which correspond to the first two dimensions of the word nature as dw.sh, dw.sgs, dw.xgs and dw.gds. Among them, dw represents the attribute of the unit, and sh, sgs, xgs and gds represent the unit level of the provincial company, the municipal company, the county company and the power supply station, respectively. For example, the word nature code of a municipal company is dw.sgs.name, and name represents the unit name. The equipment includes substations, lines, line segments and users, which correspond to the first two dimensions of the word nature as sb.bdz, sb.xl, sb.xd and sb.yh. Among them, sb represents the attribute of the equipment, and bdz, xl, xd and yh represent the equipment types of the substation, the line, the line segment and the user, respectively.

[0063] The word nature of the business word includes the following types: the first type, the first dimension of the word nature is yw, such as the word nature of the public variable is yw.tz.yhxz. The first dimension, the yw type data has corresponding data in the graph database, the second dimension tz represents the account data, and the third dimension yhxz indicates the user nature field; the second type is ywt, such as the word nature of the cancellation date is ywt.tz.zxrq, wherein ywt indicates that the word is a field name related to date in the database system, and the third dimension zxrq represents the responsibility reason field; the third type is ywh, such as the event serial number, whose word nature is ywh.yx.sjxh, wherein the first dimension ywh indicates that the word is a field name in the system, the storage type of which is a digital symbol string, yx represents the running data, and sjxh represents the event serial number field.

[0064] S3: Tokenizing the standardized input sentence, screening the word list after tokenization, establishing judgment rules of word nodes and corresponding word conversion rules based on rule matching principle, so as to convert words of different word nature into cypher query statement, or build business data set, input the business data set into decision tree, train the decision tree, and generate corresponding cypher query statement from the trained decision tree;

[0065] Among them, the part-of-speech filtering of the word list after segmentation is mainly realized by the ext function. Since the part-of-speech is dimensioned, when the corresponding dimensional words are needed, the rbck function is used to extract them, which is equivalent to extracting words of different dimensions. Specifically, the ext function: this function is mainly used to filter the corresponding word list list according to the part-of-speech of the cut function segmentation result. For each word in list, split the part-of-speech by the split(“.”) function according to the dimension, compare the input part-of-speech symbol with the segmented part-of-speech of each dimension, and if they are equal, add it to the output list. Finally, output all selected word lists. The rbck function: part-of-speech dimension selection. Input word data with part-of-speech and selected dimension. Split the part-of-speech of the word data by the split(“.”) function, and output the corresponding dimension part-of-speech string according to the input selected dimension.

[0066] The judgment rules of the word nodes and the corresponding word conversion rules are established based on the rule matching principle, so that the words of different parts of speech are converted into cypher query statements, which include:

[0067] (1) The device sb class takes out the corresponding part-of-speech third symbol, and the third symbol is the field corresponding to the sb class in the system. The corresponding cypher query field is obtained through the python string splicing technology; the yw class is similar to the sb class processing;

[0068] (2) The dw class compares the unit information carried by the string with the unit information of the query account itself. If the carried unit is within the jurisdiction of the query account unit, the carried unit is taken as the condition to join the cypher query statement. If it is not within the authority of the query account unit, the unit information of the query account itself is input and the carried unit information is ignored;

[0069] (3) The ywt class processes the user input statement, ywt word list and time word list through the ywt function. When the ywt word list is empty and the time word list has data, and when the user input statement has words such as “power failure” and “heavy power failure”, the time word list is converted into cypher query statement segment as the time data of the power failure time field. When the ywt word list and the time word list are not empty, match them in the ratio of 1:2, and convert each ywt word range between two time words into a cypher query statement segment;

[0070] (4) The ywh class contains time, household, time length, and abandoned power supply, which involves running data attribute words, and the combination of these comparison words + numbers followed by greater than X is converted into a cypher query statement segment;

[0071] (5) ywsb, ywdw, as the extension of sb, dw type judgment rules, when ywsb and sb are the same type of equipment, do not add new statements, when the two levels are different, add the relationship statement between ywsb and sb type equipment, ywsb points to sb, direct query sb; when ywdw and dw are the same type of unit, do not add new statements, when the two levels are different, add the relationship statement between ywdw and dw type unit, ywdw points to dw, direct query dw;

[0072] (6) ywbj class, including greater than, less than, equal to these comparison classes, need to be adjacent to the data, date class combination into an integral whole, with the adjacent ywt, ywh type associated, together constitute its research and judgment rules;

[0073] The word sequence obtained by the user input statement is judged according to the above rules in the order of first and second order from back to front. The corresponding cypher statement segment is obtained by lighting the class condition node according to the corresponding coding conversion rule for the word sequence in the word sequence. According to the cypher grammar rule, the cypher query statement is completed by combining each code segment.

[0074] Construct a business dataset, including: queries based on equipment type, fuzzy name search, equipment region feature classification query, equipment operating status query, query of equipment's parent equipment, query of multiple levels of parent equipment, query of equipment's various levels of units, query of equipment connected to subordinate equipment or multiple levels of subordinate equipment, query of power outages of a specific type or several types of power outages, or power outage responsibility or several reasons within a time period for a specific equipment, query of power outages of a specific type or several types of power outages, or power outage responsibility or several reasons within a time period for a specific equipment causing a power outage range to exceed the input value, and query of power outages of a specific equipment connected to a specific equipment causing a power outage range to exceed the input value within a time period. This query includes: equipment ledger information for a power outage; equipment ledger information for a specific area within a power outage; detailed query of a type of equipment within a unit that has experienced more than X power outages due to or / and a specific type of power outage within a specific area; detailed query of a type of equipment within a unit that has experienced more than X power outages due to or / and a specific type of power outage within a specific area; detailed query of a type of equipment within a unit that has experienced more than X hours of power outage due to or / and a specific type of power outage within a specific area; statistical query of a unit within a specific time period for a specific cause or type of power outage and within a specific area, including the total time affected by power outages and the number of affected households or households; and query of the equipment with the largest or smallest total time affected by power outages due to a specific cause or type of power outage and within a specific area within a specific time period.

[0075] Based on the business dataset, the word segmentation sequence obtained from the user input statement is matched according to the order of the words and then processed according to the corresponding encoding conversion rules to convert all words in the word segmentation sequence and obtain their corresponding Cypher statement segments. The code segments are then combined according to Cypher syntax rules to form a Cypher query statement, forming a test dataset. The word segmentation and Cypher query statement segments are encoded to generate decision trees from word segmentation to Cypher query statement segments and from Cypher query statement segments to word segmentation. The difference between the Cypher statement segments generated by the decision trees and the corresponding Cypher statement segments in the test dataset is used to construct a loss function. The parameters of the decision trees are adjusted and the decision trees are trained until the loss function value is minimized or the preset number of training iterations is reached. The trained decision trees are then used to generate the corresponding Cypher query statements.

[0076] S4: according to the rule matching and the decision tree, respectively, the cypher query statement is generated, the consistency of the coded results is calculated, when there is consistency, the corresponding data is queried through the cypher query statement, when the two are inconsistent, the input statement and the generated cypher query statement are saved as difficult data, the cypher query statement generated by the decision tree is converted into the corresponding natural language to ask the inputter for confirmation, and the positive reply is queried according to the cypher query statement generated by the decision tree, and the negative reply is queried according to the cypher query statement converted by the rule matching.

[0077] At the same time, the result feedback option is provided for the returned data, the user does not select or selects satisfaction, which indicates that the query is valid, the difficult data and the processing method are stored in the difficult processing set as a subsequent query reference, the user selects the query result which does not match the query target, which indicates that the query generation has a problem, and is put into the problem data set, and subsequently, the problem data set is added to the decision tree training data set to retrain and update the decision tree.

[0078] Through the above technical scheme, the power supply reliability management system data is stored in the graph database, and the natural language is standardized, instead of directly querying the natural language, so that the problem that the natural language cannot accurately query the result due to too strong professional nature is avoided, then the words are segmented and the word types are screened, and the natural language is converted into the cypher query statement based on the matching rule and the decision tree, which can effectively avoid the problems of insufficient generalization ability based on the matching rule, underfitting and overfitting of the decision tree, and further improve the query accuracy and avoid retrieval failure.

[0079] Embodiment 2

[0080] Based on embodiment 1, embodiment 2 of the present application further provides a power supply reliability full-service intelligent retrieval system, the system comprises a storage medium, the storage medium stores a computer program, the computer program can be executed by a processor to complete the method in embodiment 1.

[0081] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A full-service intelligent retrieval method for power supply reliability, characterized in that, Includes the following steps: Step 1: Store the power supply reliability management system data into the graph database; Step 2: Preprocess the natural language to obtain standardized input statements; Step 3: Segment the standardized input statement into words, filter the word list by part of speech, establish the judgment rules for word nodes and the corresponding word encoding rules based on the rule matching principle, so as to convert words of different parts of speech into Cypher query statements, or construct a business dataset, input the business dataset into a decision tree, train the decision tree, and generate the corresponding Cypher query statements from the trained decision tree; Step 4: For the user input statement, generate Cypher query statements based on rule matching and decision tree respectively. Perform consistency calculation on the encoded results of the two. When there is consistency, retrieve the corresponding data through the Cypher query statement. When the two are inconsistent, save the input statement and the generated Cypher query statement as difficult data. Convert the Cypher query statement generated by the decision tree into the corresponding natural language to ask the user for confirmation. If the answer is positive, query according to the Cypher query statement generated by the decision tree. If the answer is negative, query according to the Cypher query statement converted by rule matching.

2. The intelligent power supply reliability retrieval method according to claim 1, characterized in that, Step one includes: The power supply reliability management system data is stored in the neo4j graph database to establish a full-service data knowledge graph for power supply reliability. This graph includes equipment types such as substations, lines, line segments, and users, as well as operational data such as power outage events, outage line segments, and outage users. It establishes management connections between units and their subordinate equipment, connection links between equipment at all levels, inclusion relationships between outage data at all levels, and outage connections between equipment and outage data.

3. The intelligent power supply reliability retrieval method according to claim 1, characterized in that, Step two includes: Step 201: Convert the format of the date data; Step 202: Load the parts of speech corresponding to the dotted symbols into the graph database; Step 203: Convert inconsistent expressions in the input statement into a uniform format; Step 204: Perform multi-dimensional part-of-speech coding on users according to attributes, unit level, and unit name; perform multi-dimensional part-of-speech coding on devices according to attributes and device type, and load the results into the graph database; perform multi-dimensional part-of-speech coding on business terms, and add the results to the graph database.

4. The intelligent power supply reliability retrieval method according to claim 3, characterized in that, Step 201 includes: The system displays the current date and time information; converts Chinese characters representing the year, month, and quarter into a format of numbers plus year or month; extracts the year, month, or day from strings containing year, month, and day; converts date data into a fixed format string; completes the year and month attributes of dates, using the month specified in the query for dates without a month and the year specified in the query for dates without a year; and converts dates representing a time period into two time points.

5. The intelligent power supply reliability retrieval method according to claim 3, characterized in that, Step 204 includes: The units include four levels: provincial company, municipal company, county company, and power supply station. The first two dimensions of their respective parts of speech are dw.sh, dw.sgs, dw.xgs, and dw.gds. Here, dw indicates the attribute is "unit," and sh, sgs, xgs, and gds represent the unit level as provincial company, municipal company, county company, and power supply station, respectively. For example, a municipal company's part-of-speech code is dw.sgs.name, where name represents the unit name. Equipment includes substations, lines, line segments, and users. The first two dimensions of their respective parts of speech are sb.bdz, sb.xl, sb.xd, and sb.yh. Here, sb indicates the attribute is "equipment," and bdz, xl, xd, and yh represent the equipment type as substation, line, line segment, and user, respectively. The parts-of-speech tags for business terms fall into the following categories: The first category includes terms with the first dimension being 'yw', such as 'public variable' with the part-of-speech tag 'yw.tz.yhxz'. Here, the first dimension, 'yw', indicates that the data exists in the graph database; the second dimension, 'tz', indicates that it's ledger data, not operational data; and the third dimension, 'yhxz', indicates that it's a user-related field. The second category includes 'ywt', such as 'cancellation date' with the part-of-speech tag 'ywt.tz.zxrq'. Here, 'ywt' indicates that this term is a field name in a date-related database system; and the third dimension, 'zxrq', indicates that it's a field representing the reason for responsibility. The third category includes 'ywh', such as 'event sequence number' with the part-of-speech tag 'ywh.yx.sjxh'. Here, the first dimension, 'ywh', indicates that this term is a field name in the system with a storage type of number or symbolic string; 'yx' indicates operational data; and 'sjxh' indicates the event sequence number field.

6. The intelligent power supply reliability retrieval method according to claim 5, characterized in that, The rule-based matching principle is used to establish word node analysis rules and corresponding word encoding rules, thereby converting words of different parts of speech into Cypher query statements, including: (1) For the device sb class, extract the corresponding part-of-speech third category symbol. The third category symbol is the field corresponding to the sb class in the system. Obtain the corresponding Cypher query field through Python string concatenation technology; yw is processed similarly to the sb class. (2) dw class, by comparing the unit information brought in by the string with the unit information of the query account itself, if the brought-in unit is within the jurisdiction of the query account, the brought-in unit is added to the Cypher query statement as a condition; if it is not within the unit permissions of the query account, the query account's own unit information is input and the brought-in unit information is ignored. (3) ywt class, which processes the user input statement, ywt word list and time word list through the ywt function. When the ywt word list is empty but the time word list has data, and when the user input statement contains words such as "power outage" and "restart" which are related to the operation information, the time word list is converted into a Cypher query statement segment as the time data of the power outage time field. When neither the ywt word list nor the time word list is empty, the two are matched in a 1:2 ratio, and the relationship between each ywt word and the two time words is converted into a Cypher query statement segment. (4) The ywh class contains terms related to operational data attributes such as time-household, household number, duration, and abandoned power supply, which are combined with comparison terms such as greater than X followed by numbers and converted into Cypher query statement segments. (5) ywsb and ywdw are extended judgment rules for sb and dw. When ywsb and sb have the same device type, no new statement is added. When the two have different levels, a relationship statement between ywsb and sb is added, with ywsb pointing to sb and sb being queried directly. When ywdw and dw have the same unit type, no new statement is added. When the two have different levels, a relationship statement between ywdw and dw is added, with ywdw pointing to dw and dw being queried directly. (6) The ywbj class includes comparison words such as greater than, less than, and equal to. It needs to be combined with adjacent data and date classes to form a whole, and associated with neighboring types such as ywt and ywh to jointly constitute its judgment rules. The word segmentation sequence obtained from the user input statement is processed according to the above rules in two orders: sequential and backward. For those that meet the conditions, the corresponding condition nodes are highlighted, and the corresponding encoding conversion rules are applied to complete the conversion of all words in the word segmentation sequence, thereby obtaining the corresponding Cypher statement segments. The code segments are then combined according to the Cypher syntax rules to complete the Cypher query statement.

7. The intelligent power supply reliability retrieval method according to claim 6, characterized in that, The construction of the business dataset includes: Search by equipment type, name (fuzzy search), equipment area feature classification, equipment operating status, parent / subsidiary level, affiliated unit level, connected subordinate or multi-level subordinate equipment, power outage type or cause of power outage for a specific equipment within a given time period, power outage exceeding the input range due to a specific type or cause of power outage for a specific equipment within a given time period, and power outage exceeding the input range for a specific type or cause of power outage among connected equipment within a given time period. A single power outage may contain... This includes querying equipment ledger information, querying equipment ledger information for a specific area within a power outage, querying detailed information for a certain type of equipment within a unit's area that has experienced more than X power outages due to or / and a certain type of power outage, querying detailed information for a certain type of equipment within a unit's area that has experienced more than X power outages due to or / and a certain type of power outage, querying detailed information for a certain type of equipment within a unit's area that has experienced more than X hours of power outage due to or / and a certain type of power outage, querying detailed information for a certain type of equipment within a unit's area that has experienced more than X hours of power outage due to or / and a certain type of power outage, querying statistical information for a unit's area within a certain time period and the total number of households affected by power outages due to a certain cause or nature of power outages and within a certain area, and querying the equipment with the largest or smallest total number of households affected by power outages due to a certain cause or nature of power outages and within a certain area within a certain time period.

8. The intelligent retrieval method for full-service power supply reliability according to claim 7, characterized in that, The process of inputting the business dataset into a decision tree, training the decision tree, and generating corresponding Cypher query statements from the trained decision tree includes: Based on the business dataset, the word segmentation sequence obtained from the user input statement is matched according to the order of the words and then processed according to the corresponding encoding conversion rules to convert all words in the word segmentation sequence and obtain their corresponding Cypher statement segments. The code segments are then combined according to Cypher syntax rules to form a Cypher query statement, forming a test dataset. The word segmentation and Cypher query statement segments are encoded to generate decision trees from word segmentation to Cypher query statement segments and from Cypher query statement segments to word segmentation. The difference between the Cypher statement segments generated by the decision trees and the corresponding Cypher statement segments in the test dataset is used to construct a loss function. The parameters of the decision trees are adjusted and the decision trees are trained until the loss function value is minimized or the preset number of training iterations is reached. The trained decision trees are then used to generate the corresponding Cypher query statements.

9. The intelligent retrieval method for full-service power supply reliability according to claim 1, characterized in that, Step four also includes: The system provides feedback options for the returned data. If the user does not select an option or selects "satisfied," it indicates that the query is valid. Difficult data and processing methods are stored in the "difficult processing set" for future query reference. If the user selects "the query result does not match the query target," it indicates that there is a problem with the query generation. The data is then added to the "problem dataset." Subsequently, the decision tree training dataset is added to the problem dataset to retrain and update the decision tree.

10. A full-service intelligent retrieval system for power supply reliability, characterized in that, The system includes a storage medium storing a computer program that can be executed by a processor to perform the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Statement query method and device, electronic equipment and computer readable storage medium

    CN114625754A

  • Data processing query method and device based on OLAP pre-calculation model

    CN112148719A

  • Database operation method and device

    CN112783921A