Generation method, device and equipment of telecontrol configuration description file and readable storage medium
By filtering matching participles in the substation detection information table and configuration description file, the remote configuration description file is automatically generated, which solves the problem of low generation efficiency in the existing technology and achieves fast and efficient file generation.
Patent Information
- Application Number
- CN202510393312.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the generation of remote configuration description files is inefficient and complicated, resulting in too long file generation time.
By determining the word segmentation in the substation detection information table and configuration description file, filter out the matching candidate measurement point description, and automatically generate the remote configuration description file using word segmentation matching to reduce the calculation amount and improve the generation efficiency.
It realizes rapid automatic generation of remote configuration description files, reduces manual intervention and improves generation efficiency.
Smart Images

Figure CN120407781A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product for generating a remote configuration description file. Background Art
[0002] With the rapid development of smart grids, the configuration complexity of substations has been continuously increasing. A remote configuration description (RCD) file is a very important file that records the control parameters of various devices in a substation and can help achieve remote monitoring and control of the substation.
[0003] In related technologies, an RCD file is usually generated manually. However, this method is cumbersome to operate, resulting in an excessively long file generation time and a problem of low efficiency in generating remote configuration description files. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product for generating a remote configuration description file that can improve the efficiency of generating remote configuration description files for the above technical problems.
[0005] In a first aspect, the present application provides a method for generating a remote configuration description file, including:
[0006] Determine at least one first word segment included in each first measurement point description in the detection information table regarding the substation, and determine at least one second word segment included in each second measurement point description in the substation configuration description file;
[0007] For each first measurement point description, screen out at least one candidate measurement point description from multiple second measurement point descriptions according to the number of word segments of the first word segments included in the first measurement point description;
[0008] For each first word segment of the first measurement point description, screen out a second word segment that matches the first word segment from the second word segments involved in at least one candidate measurement point description;
[0009] Based on the second word segments that match each first word segment in each first measurement point description, determine a candidate measurement point description that matches each first measurement point description;
[0010] According to the substation configuration description file, determine the remote configuration description information of the candidate measurement point description that matches each first measurement point description, and generate a remote configuration description file based on each remote configuration description information.
[0011] Second aspect, the present application also provides a device for generating a telecontrol configuration description file, including:
[0012] A word segmentation determination module, configured to determine at least one first word segment included in each first measurement point description in the detection information table regarding the substation, and determine at least one second word segment included in each second measurement point description in the substation configuration description file;
[0013] A measurement point description screening module, configured to, for each first measurement point description, screen out at least one candidate measurement point description from multiple second measurement point descriptions according to the number of word segments included in the first word segments of the first measurement point description;
[0014] A word segment screening module, configured to, for each first word segment of the first measurement point description, screen out second word segments matching the first word segment from the second word segments involved in at least one candidate measurement point description;
[0015] A measurement point description matching module, configured to determine, based on the second word segments matching each first word segment in each first measurement point description, candidate measurement point descriptions matching each first measurement point description;
[0016] A file generation module, configured to determine, according to the substation configuration description file, telecontrol configuration description information of candidate measurement point descriptions matching each first measurement point description, and generate a telecontrol configuration description file based on each telecontrol configuration description information.
[0017] Third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0018] Determine at least one first word segment included in each first measurement point description in the detection information table regarding the substation, and determine at least one second word segment included in each second measurement point description in the substation configuration description file;
[0019] For each first measurement point description, screen out at least one candidate measurement point description from multiple second measurement point descriptions according to the number of word segments included in the first word segments of the first measurement point description;
[0020] For each first word segment of the first measurement point description, screen out second word segments matching the first word segment from the second word segments involved in at least one candidate measurement point description;
[0021] Based on the second word segments matching each first word segment in each first measurement point description, determine candidate measurement point descriptions matching each first measurement point description;
[0022] Determine the telecontrol configuration description information of the candidate measuring point descriptions that match each first measuring point description according to the substation configuration description file, and generate a telecontrol configuration description file based on each piece of telecontrol configuration description information.
[0023] Fourthly, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0024] Determine at least one first participle included in each first measuring point description in the detection information table regarding the substation, and determine at least one second participle included in each second measuring point description in the substation configuration description file;
[0025] For each first measuring point description, screen out at least one candidate measuring point description from multiple second measuring point descriptions according to the number of words of the first participles included in the first measuring point description;
[0026] For each first participle of the first measuring point description, screen out the second participles that match the first participles from the second participles involved in at least one candidate measuring point description;
[0027] Based on the second participles that match each first participle in each first measuring point description, determine the candidate measuring point descriptions that match each first measuring point description;
[0028] Determine the telecontrol configuration description information of the candidate measuring point descriptions that match each first measuring point description according to the substation configuration description file, and generate a telecontrol configuration description file based on each piece of telecontrol configuration description information.
[0029] Fifthly, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0030] Determine at least one first participle included in each first measuring point description in the detection information table regarding the substation, and determine at least one second participle included in each second measuring point description in the substation configuration description file;
[0031] For each first measuring point description, screen out at least one candidate measuring point description from multiple second measuring point descriptions according to the number of words of the first participles included in the first measuring point description;
[0032] For each first participle of the first measuring point description, screen out the second participles that match the first participles from the second participles involved in at least one candidate measuring point description;
[0033] Based on the second participles that match each first participle in each first measuring point description, determine the candidate measuring point descriptions that match each first measuring point description;
[0034] Based on the substation configuration description file, determine the telecontrol configuration description information of the candidate measurement point descriptions that match each first measurement point description, and generate a telecontrol configuration description file based on each piece of telecontrol configuration description information.
[0035] The above-mentioned method, device, computer device, computer-readable storage medium, and computer program product for generating a telecontrol configuration description file determine at least one first participle included in each first measurement point description in the detection information table regarding the substation, and determine at least one second participle included in each second measurement point description in the substation configuration description file; for each first measurement point description, according to the number of words of the first participles included in the first measurement point description, at least one candidate measurement point description is pre-screened from multiple second measurement point descriptions from the dimension of the number of words, avoiding subsequent matching processing for all second participles of all second measurement point descriptions, greatly simplifying the calculation amount. For each first participle of the first measurement point description, second participles that match the first participle are screened out from the second participles involved in at least one candidate measurement point description; based on the second participles that match each first participle in each first measurement point description, second candidate measurement point descriptions that match each first measurement point description are determined. The first measurement point description and the candidate measurement point description are automatically matched through the participle matching dimension between the first participles and the second participles. Then, according to the substation configuration description file, the telecontrol configuration description information of the candidate measurement point descriptions that match each first measurement point description is accurately determined, and a telecontrol configuration description file can be quickly and automatically generated based on each piece of telecontrol configuration description information. In the whole process, preliminary screening is performed through the number of words to reduce the subsequent matching workload, and the automatic generation of the telecontrol configuration description file is achieved through participle matching, without manual generation, improving the generation efficiency of the telecontrol configuration description file. Description of the Drawings
[0036] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 It is an application environment diagram of the method for generating a telecontrol configuration description file in an embodiment;
[0038] Figure 2 It is a flowchart of the method for generating a telecontrol configuration description file in an embodiment;
[0039] Figure 3 It is a flowchart of the first participle determination step in an embodiment;
[0040] Figure 4 The structural block diagram of a generating device for a telecontrol configuration description file in an embodiment;
[0041] Figure 5 The internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0042] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0043] The method for generating a telecontrol configuration description file provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. The method for generating a telecontrol configuration description file provided by the embodiments of the present application can be executed independently by the terminal 102 or the server 104, or can be executed collaboratively by the terminal 102 and the server 104.
[0044] Taking the collaborative execution of the terminal 102 and the server 104 as an example for illustration: The terminal sends a detection information table about a substation to the server 104. The server 104 obtains the substation configuration description file. The server 104 determines at least one first word segment included in each first measurement point description in the detection information table about the substation, and determines at least one second word segment included in each second measurement point description in the substation configuration description file; for each first measurement point description, the server 104 screens out at least one candidate measurement point description from multiple second measurement point descriptions according to the number of first word segments included in the first measurement point description; for each first word segment of the first measurement point description, the server 104 screens out a second word segment that matches the first word segment from the second word segments involved in at least one candidate measurement point description; the server 104 determines a candidate measurement point description that matches each first measurement point description based on the second word segment that matches each first word segment in each first measurement point description; the server 104 determines the telecontrol configuration description information of the candidate measurement point description that matches each first measurement point description according to the substation configuration description file, and generates a telecontrol configuration description file based on each telecontrol configuration description information.
[0045] Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0046] In an exemplary embodiment, as Figure 2 shown, a method for generating a telecontrol configuration description file is provided. Taking the case where this method is applied to a computer device (which can be Figure 1 the terminal 102 or Figure 1 the server 104 in
[0047] Step 202: Determine at least one first word segment included in each first measurement point description in the detection information table regarding the substation, and determine at least one second word segment included in each second measurement point description in the substation configuration description file.
[0048] Among them, the detection information table regarding the substation is issued by the dispatching. The detection information table records the measurement and control devices required for the dispatching to perform telecontrol and detection. It can be understood that the detection information table includes multiple first measurement point descriptions. The first measurement point description refers to the measurement point description in the detection information table. The measurement point description includes the device name of the measurement and control device for telecontrol and detection and the status description of the measurement and control device. The measurement point description is used to describe the measurement points deployed in the corresponding measurement and control device. At least one measurement point is deployed in each measurement and control device. The measurement points can be remote measurement, remote signaling, remote control, and remote pulse measurement points. The status description is used to reflect whether the measurement and control device is abnormal or the type of abnormality that occurs. The first word segment refers to the word segment obtained by segmenting the first measurement point description.
[0049] The substation configuration description file (SCD, Substation Configuration Description) is used to describe the configuration information of the substation. The substation configuration description file covers the overall situation of the substation, the detailed parameters and functions of the measurement and control devices and logical nodes, the communication network settings, the automation system function configuration, etc. In one embodiment, the substation configuration description file includes multiple second measurement point descriptions and the telecontrol configuration description information of each second measurement point description. The telecontrol configuration description information reflects the configuration parameters required for the corresponding measurement and control device to perform telecontrol and detection. For example, the signal type of the signal collected by the measurement and control device when performing telecontrol and detection. The second word segment refers to the word segment obtained by segmenting the second measurement point description.
[0050] Optionally, in response to a current file generation request, the computer device obtains a detection information table of a substation and a substation configuration description file. The computer device generates a first dictionary corresponding to the detection information table based on the detection information table, and tokenizes each first measurement point description based on the first dictionary to obtain at least one first token corresponding to each first measurement point description. Exemplarily, the first dictionary is generated in real time based on the current file generation request. For example, in response to the current file generation request, after obtaining the detection information table of the substation, for each first measurement point description in the detection information table, based on the characters in the first measurement point description, a plurality of candidate words for the first measurement point description are determined, and the left entropy and right entropy of each candidate word are calculated, and the first token of the first measurement point description is screened out from the plurality of candidate words. Alternatively, the computer device preprocesses the measurement point descriptions of the obtained detection information table of the substation and the substation configuration description file respectively to screen out the measurement point descriptions with non-standard naming and formats, and continues to execute the steps of determining the first token and the second token based on the preprocessed detection information table and substation configuration description file.
[0051] For example: For each first measurement point description, identify a plurality of characters in the first measurement point description, and generate a plurality of candidate words based on the plurality of characters and the character order. The length of each candidate word can be the same or different. For example, there are 4 characters, which are "water", "fruit", "very", and "sweet" in order according to the character order. Then the generated candidate words can be "water", "fruit", "very", "sweet", "fruit", "very" and so on. For a candidate word w, the left entropy of the candidate word w is calculated using the following formula :
[0052] (1)
[0053] where represents the set of contexts that may appear on the left side of the candidate word w, that is, the set composed of all words or phrases that may appear on the left side of the candidate word w in the text. is the context in the context set. refers to the relative frequency of the context c1 and the candidate word w appearing, expressed as: . For example, in the text "I drink coffee, he also drinks coffee, and we buy coffee together", for the word "coffee", the context set is {drink, buy}. Then, the number of times "drink" appears on the left side of "coffee" is 2, and the number of times "buy" appears on the left side of "coffee" is 1. Then, P(drink) = 2 / 3. Similarly, the right entropy of the candidate word w is calculated using the following formula :
[0054] (2)
[0055] where Denote the set of contexts that may appear to the right of the candidate word w, that is, the set composed of all words or phrases that may appear to the right of the candidate word w in the text. c is the context in this context set. It refers to the relative frequency of the context c2 and the candidate word w, expressed as: .
[0056] It should be noted that the left entropy is used to measure the uncertainty of the context distribution on the left side of the candidate word w, and the right entropy is used to measure the uncertainty of the context distribution on the right side of the candidate word w.
[0057] After calculating the left entropy and right entropy of each candidate word w in the first measurement point description, fuse the left entropy and right entropy to obtain a fused entropy value. Based on the comparison result between the fused entropy value and the fusion threshold, determine whether the candidate word w can be used as the first word segment, that is, whether the candidate word w is a word.
[0058] Optionally, in response to the current file generation request, the computer device obtains a first historical dictionary of the detection information table. The first historical dictionary is generated based on the historical detection information tables obtained for each of the historical file generation requests. The computer device performs word segmentation on each first measurement point description according to the first historical dictionary to obtain at least one first word segment corresponding to each first measurement point description.
[0059] Optionally, after the computer device obtains the substation configuration description file in response to the current file generation request, the computer device generates a second dictionary corresponding to the substation configuration description file based on the substation configuration description file, and performs word segmentation on each second measurement point description based on the second dictionary to obtain at least one second word segment corresponding to each second measurement point description. Exemplarily, the second dictionary is generated in real time based on the current file generation request. For example, after obtaining the substation configuration description file in response to the current file generation request, for each second measurement point description in the substation configuration description file, based on the characters in the second measurement point description, determine multiple candidate words for the second measurement point description. By calculating the left entropy and right entropy of each candidate word, screen out the second word segment of the second measurement point description from the multiple candidate words. For the calculation of the left entropy and right entropy of each candidate word in the second measurement point description, it can be calculated with reference to the above formulas (1)-(2), which will not be elaborated here. After calculating the left entropy and right entropy of each candidate word w in the second measurement point description, fuse the left entropy and right entropy to obtain a fused entropy value. Based on the comparison result between the fused entropy value and the fusion threshold, determine whether the candidate word w can be used as the second word segment, that is, whether the candidate word w is a word.
[0060] Optionally, in response to a current file generation request, the computer device obtains a second historical dictionary for the substation configuration description file, where the second historical dictionary is generated based on the historical substation configuration description files obtained for respective historical file generation requests. The computer device tokenizes each second measurement point description according to the second historical dictionary to obtain at least one second token corresponding to each second measurement point description.
[0061] Step 204: For each first measurement point description, at least one candidate measurement point description is screened out from multiple second measurement point descriptions according to the number of tokens included in the first measurement point description.
[0062] It should be noted that for any two first measurement point descriptions and second measurement point descriptions, the more matching the first measurement point description and the second measurement point description are, the more the number of tokens of the first measurement point description and the second measurement point description is the same.
[0063] Therefore, in one embodiment, screening out at least one candidate measurement point description from multiple second measurement point descriptions according to the number of tokens included in the first measurement point description includes: obtaining a preset token increment, determining a token number range corresponding to the first measurement point description according to the number of tokens included in the first measurement point description and the token increment; screening out second measurement point descriptions whose token numbers are within the token number range from multiple second measurement point descriptions according to the number of tokens included in each second measurement point description, and taking the screened-out second measurement point descriptions as candidate measurement point descriptions.
[0064] Among them, the preset token increment is set in advance, which can be 0, or 1, or 2. The smaller the value of the token increment, the better the screening effect. Exemplarily, for each first measurement point description, the computer device obtains the token increment, determines the number of tokens of the first tokens included in the first measurement point description, takes the difference between the number of tokens and the token increment as the first boundary, and takes the sum of the number of tokens and the token increment as the second boundary, and determines the token number range corresponding to the first measurement point description based on the first boundary and the second boundary. Determine the number of tokens of the second tokens included in each second measurement point description, and take the second measurement point descriptions whose token numbers are within the token number range as candidate measurement point descriptions, that is, the token numbers of the candidate measurement point descriptions are all less than or equal to the second boundary and greater than or equal to the first boundary.
[0065] In this embodiment, through the token increment and the number of tokens of the first tokens in the first measurement point description, it is possible to screen out candidate measurement point descriptions with a high possibility of matching the first measurement point description from a large number of second measurement point descriptions. The screening completed according to the number of tokens not only narrows the range of measurement point description matching, but also reduces the workload of subsequent measurement point description matching, effectively improving the generation efficiency of the telecontrol configuration description file.
[0066] Step 206: For each first participle described in the first measurement point, screen out the second participles that match the first participle from the second participles involved in at least one candidate measurement point description.
[0067] Among them, the second participle that matches the first participle means that the first participle is the same as the second participle, or the meanings of the first participle and the second participle are the same.
[0068] Optionally, for each first participle in each first measurement point description, the computer device determines the second participles of each candidate measurement point description, and queries the second participles that match the first participle from all the second participles of all candidate measurement point descriptions.
[0069] Exemplarily, the computer device first queries the second participles that are the same as the first participle from all the second participles of all candidate measurement point descriptions. If any, the queried second participles are used as the second participles that match the first participle. If not, it is determined whether there are synonyms of the first participle among all the second participles of all candidate measurement point descriptions. If any, the second participles that are synonymous with the first participle are used as the second participles that match the first participle.
[0070] In one embodiment, screening out the second participles that match the first participle from the second participles involved in at least one candidate measurement point description includes: for each first participle, if there is no second participle that is the same as the first participle among the second participles involved in at least one candidate measurement point description, query the synonyms that are synonymous with the second participle from the synonym library corresponding to the substation; if there are synonyms among the second participles involved in at least one candidate measurement point description, use the synonyms as the second participles that match the first participle.
[0071] In one embodiment, the step of constructing the synonym library includes: after the computer device obtains the first participle, according to the first measurement point description to which each first participle belongs, construct a first word library, and the first word library records at least one first participle of each first measurement point description. Similarly, after the computer device obtains the second participle, according to the second measurement point description to which each second participle belongs, construct a second word library, and the second word library records at least one second participle of each second measurement point description. The computer device queries the first participles and second participles with a synonymous relationship according to the first word library and the second word library, and constructs a synonym library based on the first participles and second participles with a synonymous relationship. It can be understood that for the first participles and second participles with a synonymous relationship, the synonym of the first participle is the second participle, and the synonym of the second participle is the first participle.
[0072] In one embodiment, the steps of specifically constructing a synonym library include: for each first word segment, the computer device calculates the edit distance (Levenshtein distance) between the first word segment and each second word segment in the second word library respectively. When the minimum edit distance is less than the distance threshold, the second word segment corresponding to the minimum edit distance is used as the synonym of the first word segment, that is, the second word segment corresponding to the minimum edit distance has a synonymous relationship with the first word segment. Among them, the edit distance represents the minimum number of operations required to transform one word into another, including inserting, deleting, and replacing characters. If the edit distance between two words is small, it means they are relatively similar in form and may also be relatively close in semantics. Exemplarily, the edit distance can be calculated according to the following formula (3):
[0073] (3)
[0074] Among them, D(A,B) is the Levenshtein distance (i.e., the minimum number of edit operations) between string A and string B. |A| is the length of string A. |B| is the length of string B. In this embodiment, string A is the first word segment, and string B is the second word segment. is the edit distance between two strings, and its value range is [0,1]. 0 means the two strings are exactly the same, and 1 means the two strings are completely different.
[0075] It should be noted that the smaller the edit distance between the first word segment and the second word segment, the more similar the glyphs are, and the greater the probability of having a synonymous relationship. Of course, in some scenarios, a larger edit distance does not necessarily mean there is no synonymous relationship. For example, "beautiful" and "pretty" have a large edit distance but have similar meanings. Therefore, in the case where the synonym library is constructed based on the edit distance, to avoid missing words with similar meanings but dissimilar glyphs, the method further includes: if there is no synonym among the second word segments involved in at least one candidate measurement point description, the large language model is called, and a synonym query prompt text is constructed based on the first word segment and all second word segments, and the query result is output, and the query result indicates whether there is a second word segment that is similar or the same in meaning as the first word segment.
[0076] In this embodiment, for each first word segment, if there is no second word segment identical to the first word segment among the second word segments involved in at least one candidate measurement point description, the synonym library can be used to query the second word segment that has a synonymous relationship with the first word segment, so as to avoid failing to identify the second word segment with the same meaning, and thus improve the accuracy of generating the subsequent telecontrol configuration description file.
[0077] Step 208, based on the second word segments matched by each first word segment in each first measurement point description, determine the candidate measurement point descriptions that match each first measurement point description.
[0078] Optionally, for each first measurement point description, when it is determined that each first word segment in the first measurement point description has a matching second word segment, check whether the candidate measurement point descriptions to which the second word segments matched by each first word segment belong are the same. If they are the same, determine that the candidate measurement point description to which the second word segment matched by each first word segment belongs matches the first measurement point description.
[0079] Step 210: According to the substation configuration description file, determine the telecontrol configuration description information of the candidate measurement point description matched by each first measurement point description, and generate a telecontrol configuration description file based on each telecontrol configuration description information.
[0080] Optionally, for each first measurement point description, the computer device determines the candidate measurement point description matched by the first measurement point description, obtains the telecontrol configuration description information corresponding to the candidate measurement point description from the substation configuration description file, and generates a telecontrol configuration description file based on each telecontrol configuration description information.
[0081] In the above method for generating a telecontrol configuration description file, by determining at least one first word segment included in each first measurement point description in the detection information table of the substation, at least one second word segment included in each second measurement point description in the substation configuration description file is determined; for each first measurement point description, according to the number of first word segments included in the first measurement point description, at least one candidate measurement point description is pre-screened from multiple second measurement point descriptions from the dimension of the number of words, avoiding subsequent matching processing for all second word segments of all second measurement point descriptions, greatly simplifying the calculation amount. For each first word segment of the first measurement point description, a second word segment that matches the first word segment is screened out from the second word segments involved in at least one candidate measurement point description; based on the second word segments that match each first word segment in each first measurement point description, a second candidate measurement point description that matches each first measurement point description is determined. The first measurement point description and the candidate measurement point description are automatically matched through the word segment matching dimension between the first word segment and the second word segment. Then, according to the substation configuration description file, the telecontrol configuration description information of the candidate measurement point description matched by each first measurement point description is accurately determined, and a telecontrol configuration description file can be quickly and automatically generated based on each telecontrol configuration description information. In the whole process, the subsequent matching workload is reduced by preliminary screening through the number of words, and the telecontrol configuration description file is automatically generated through word segment matching without manual generation, improving the generation efficiency of the telecontrol configuration description file.
[0082] In one embodiment, as Figure 3 shown, it is a schematic flowchart of the first word segment determination step in one embodiment. Determining at least one first word segment included in each first measurement point description in the detection information table of the substation includes:
[0083] Step 302: Based on the first historical dictionary regarding the detection information table, perform word segmentation on each first measurement point description in the detection information table to obtain the word segmentation results of each first measurement point description.
[0084] Step 304: For each successful word segmentation result indicating successful word segmentation, determine at least one first word segment included in the corresponding first measurement point description based on the successful word segmentation result.
[0085] Step 306: For each failed word segmentation result indicating failed word segmentation, based on the failed word segmentation result, determine the characters in the field where word segmentation was not successful in the corresponding first measurement point description. Based on the characters in the field where word segmentation was not successful, determine multiple candidate words.
[0086] Optionally, for the characters and their order in the field where word segmentation was not successful in the first measurement point description, generate multiple candidate words. For example, the first measurement point description is "110v kV bus sectional measurement and control equipment failure". The successfully segmented words are "110v", "measurement and control device", and "failure". The field that was not successfully recognized is "kV bus sectional". Then, the candidate words can be "thousand", "volt", "bus", "sectional", "kV", "kV bus", "volt bus", etc.
[0087] Step 308: For each candidate word, determine the confidence level of the candidate word according to the position of the candidate word in the first measurement point description. Based on the confidence level of each candidate word, re-segment the field where word segmentation was not successful to determine at least one first word segment included in the first measurement point description corresponding to the failed word segmentation result.
[0088] Optionally, after determining the confidence level of each candidate word, according to the confidence level of each candidate word, screen out the candidate words that meet the confidence level requirements from the multiple candidate words, and use the screened candidate words with confidence levels as the first word segments included in the first measurement point description corresponding to the failed word segmentation result. For example, use the candidate words with confidence levels greater than the confidence level threshold as the candidate words that meet the confidence level requirements.
[0089] In one embodiment, determining the confidence level of the candidate word according to the position of the candidate word in the first measurement point description includes: determining the average word length in the first historical dictionary, determining the corresponding boundary coefficient according to the position of the candidate word in the first measurement point description. The position includes the boundary position and the non-boundary position. The boundary coefficient corresponding to the boundary position is greater than the boundary coefficient corresponding to the non-boundary position; according to the preset window, calculate the left entropy and right entropy of the candidate word respectively, and determine the confidence level of the candidate word based on the boundary coefficient, left entropy, right entropy, and average word length.
[0090] Among them, the average word length of the first historical dictionary is obtained by calculating the average value of the number of words in each first word segmentation in the first historical dictionary. The boundary position can be the beginning or the end of a sentence in the first measurement point description. The non-boundary position is neither the beginning nor the end of a sentence in the first measurement point description. Optionally, the computer device determines the boundary entropy of the candidate word according to the left entropy, the right entropy, and the boundary coefficient of the candidate word, and based on the boundary entropy and the average word length, the confidence level of the candidate word.
[0091] Exemplarily, for each candidate word w, the boundary entropy of the candidate word is calculated by the following formula (4):
[0092] (4)
[0093] Among them, is the left entropy, is the right entropy. The calculation processes of the left entropy and the right entropy refer to the previous formulas (1)-(2). is the number of times the candidate word w appears as an independent unit in the corpus. In this embodiment, it can be understood as the number of times the candidate word w appears within a preset window, or the number of times the candidate word appears in the telecontrol configuration description file, which can be set according to actual needs and is not specifically limited. is the boundary coefficient. When in the boundary position, the boundary coefficient is 1.2; when in the non-boundary position, the boundary coefficient is 1.0. The preset window refers to the window that slides in the first measurement point description. It can be set to take the candidate word w as the center of the preset window and determine the preset window according to the preset window size.
[0094] Next, the confidence level G of the candidate word can be calculated using formula (5):
[0095] (5)
[0096] Among them, is the appearance frequency of the candidate word w within the preset window. For example, taking the candidate word w as the center of the preset window, the preset window is determined to count . N is the number of words within the preset window. |w| is the character length of the candidate word w, and E(|w|) is the average word length of the first historical dictionary. is the boundary entropy of the candidate word w.
[0097] In this embodiment, the boundary coefficient is adaptively adjusted according to the position of the candidate word in the first measurement point description to obtain a more matching confidence level, effectively avoiding the misjudgment of word segmentation of the candidate word due to too small a confidence level caused by directly fusing the left entropy and the right entropy when the candidate word is in the boundary position. In this way, the effect of word segmentation can be improved.
[0098] In one embodiment, the method further includes: determining new word segmentations obtained after re-segmenting the fields for which word segmentation fails; adding the new word segmentations to the first historical dictionary to obtain a first updated dictionary, and based on the first updated dictionary, updating the average word length.
[0099] Exemplarily, after obtaining the first updated dictionary, recalculate the average word length based on each first word segmentation in the first updated dictionary to obtain the updated average word length. In this way, after obtaining the next file generation request, more accurate word segmentation can be performed based on the first updated dictionary and the updated average word length to ensure the accuracy of word segmentation.
[0100] In this embodiment, through the first historical dictionary, word segmentation is performed on each first measurement point description in the detection information table in advance, avoiding directly generating the first dictionary according to the detection information table and improving efficiency. For the words in the fields for which word segmentation fails, after determining the candidate words, the first word segmentations in the fields for which word segmentation fails are screened out through the confidence levels of the candidate words, ensuring the effectiveness of word segmentation.
[0101] In a specific embodiment, the specific steps are as follows:
[0102] Step 1: After the computer device obtains the current file generation request, obtain the detection information table about the substation and the substation configuration description file.
[0103] Step 2: Based on the first historical dictionary regarding the detection information table, perform word segmentation processing on each first measurement point description in the detection information table to obtain the word segmentation results of each first measurement point description; for each successful word segmentation result indicating successful word segmentation, determine at least one first word segmentation included in the corresponding first measurement point description based on the successful word segmentation result; for each failed word segmentation result indicating failed word segmentation, based on the failed word segmentation result, determine the words in the fields for which word segmentation fails in the corresponding first measurement point description, and based on the words in the fields for which word segmentation fails, determine multiple candidate words; for each candidate word, according to the position of the candidate word in the first measurement point description, determine the confidence level of the candidate word, and based on the confidence level of each candidate word, re-segment the fields for which word segmentation fails to determine at least one first word segmentation included in the first measurement point description corresponding to the failed word segmentation result.
[0104] Specifically, determine the average word length in the first historical dictionary, and according to the position of the candidate word in the first measurement point description, determine the corresponding boundary coefficient. The positions include boundary positions and non-boundary positions, and the boundary coefficient corresponding to the boundary position is greater than the boundary coefficient corresponding to the non-boundary position; according to the preset window, calculate the left entropy and right entropy of the candidate word respectively, and based on the boundary coefficient, left entropy, right entropy, and average word length, determine the confidence level of the candidate word.
[0105] Specifically, determine the new word segmentation obtained after re-segmenting the fields with unsuccessful word segmentation; add the new word segmentation to the first historical dictionary to obtain the first updated dictionary, and based on the first updated dictionary, update the average word length.
[0106] Step 3: The computer device performs word segmentation on each second measurement point description in the substation configuration description file based on the second historical dictionary of the substation configuration description file to obtain the word segmentation results of each second measurement point description; for each successful word segmentation result indicating successful word segmentation, determine at least two second word segmentations included in the corresponding second measurement point description based on the successful word segmentation result; for each failed word segmentation result indicating failed word segmentation, based on the failed word segmentation result, determine the characters in the field with unsuccessful word segmentation in the corresponding second measurement point description, and based on the characters in the field with unsuccessful word segmentation, determine multiple candidate words; for each candidate word, determine the confidence level of the candidate word according to the position of the candidate word in the second measurement point description, and based on the confidence level of each candidate word, re-segment the field with unsuccessful word segmentation to determine at least two second word segmentations included in the second measurement point description corresponding to the failed word segmentation result.
[0107] Specifically, determine the average word length in the second historical dictionary, determine the corresponding boundary coefficient according to the position of the candidate word in the second measurement point description, where the position includes the boundary position and the non-boundary position, and the boundary coefficient corresponding to the boundary position is greater than the boundary coefficient corresponding to the non-boundary position; calculate the left entropy and right entropy of the candidate word respectively according to the preset window, and determine the confidence level of the candidate word based on the boundary coefficient, left entropy, right entropy and average word length.
[0108] Specifically, determine the new word segmentation obtained after re-segmenting the fields with unsuccessful word segmentation; add the new word segmentation to the second historical dictionary to obtain the second updated dictionary, and based on the second updated dictionary, update the average word length. The calculations of left entropy, right entropy, etc. in the above embodiments for determining the second word segmentation are similar to those in the embodiments for determining the first word segmentation and can be referred to the previous text.
[0109] Step 4: For each first measurement point description, obtain the preset word increment, and determine the word quantity range corresponding to the first measurement point description according to the word quantity of the first word segmentations included in the first measurement point description and the word increment; according to the word quantity of the second word segmentations included in each second measurement point description, screen out the second measurement point descriptions whose word quantity is within the word quantity range from multiple second measurement point descriptions, and use the screened second measurement point descriptions as candidate measurement point descriptions.
[0110] Step 5: For each first participle described in the first measurement point, for each first participle, if there is no second participle identical to the first participle among the second participles involved in at least one candidate measurement point description, query for synonyms synonymous with the second participle from the synonym library corresponding to the substation; if there are synonyms among the second participles involved in at least one candidate measurement point description, use the synonyms as the second participles matching the first participle.
[0111] Step 6: Based on the second participles matching each first participle in each first measurement point description, determine the candidate measurement point descriptions matching each first measurement point description; according to the substation configuration description file, determine the telecontrol configuration description information of the candidate measurement point descriptions matching each first measurement point description, and generate a telecontrol configuration description file based on each telecontrol configuration description information.
[0112] In this embodiment, by determining at least one first participle included in each first measurement point description in the detection information table regarding the substation, at least one second participle included in each second measurement point description in the substation configuration description file is determined; for each first measurement point description, according to the number of words of the first participles included in the first measurement point description, at least one candidate measurement point description is pre-screened from multiple second measurement point descriptions from the dimension of the number of words, avoiding subsequent matching processing for all second participles of all second measurement point descriptions, greatly simplifying the calculation amount. For each first participle of the first measurement point description, the second participles matching the first participle are screened out from the second participles involved in at least one candidate measurement point description; based on the second participles matching each first participle in each first measurement point description, the second candidate measurement point descriptions matching each first measurement point description are determined. The first measurement point description and the candidate measurement point description are automatically matched through the participle matching dimension between the first participle and the second participle. Then, according to the substation configuration description file, the telecontrol configuration description information of the candidate measurement point descriptions matching each first measurement point description is accurately determined, and a telecontrol configuration description file can be quickly and automatically generated based on each telecontrol configuration description information. In the whole process, the subsequent matching workload is reduced by preliminary screening through the number of words, and the automatic generation of the telecontrol configuration description file is achieved through participle matching, without manual generation, improving the generation efficiency of the telecontrol configuration description file.
[0113] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0114] Based on the same inventive concept, an embodiment of the present application also provides a device for generating a telecontrol configuration description file for implementing the method for generating a telecontrol configuration description file involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the device for generating a telecontrol configuration description file provided below can refer to the limitations on the method for generating a telecontrol configuration description file in the above text, and will not be repeated here.
[0115] In an exemplary embodiment, as Figure 4 shown, a device 400 for generating a telecontrol configuration description file is provided, including: a word segmentation determination module 402, a measuring point description screening module 404, a word segmentation screening module 406, a measuring point description matching module 408, and a file generation module 410, where:
[0116] The word segmentation determination module 402 is used to determine at least one first word segment included in each first measuring point description in the detection information table regarding the substation, and determine at least one second word segment included in each second measuring point description in the substation configuration description file;
[0117] The measuring point description screening module 404 is used to, for each first measuring point description, screen out at least one candidate measuring point description from multiple second measuring point descriptions according to the number of words in the first word segments included in the first measuring point description;
[0118] The word segmentation screening module 406 is used to, for each first word segment of the first measuring point description, screen out a second word segment that matches the first word segment from the second word segments involved in at least one candidate measuring point description;
[0119] The measuring point description matching module 408 is used to determine a candidate measuring point description that matches each first measuring point description based on the second word segments that match each first word segment in each first measuring point description;
[0120] A file generation module 410 is configured to determine, according to the substation configuration description file, telecontrol configuration description information of candidate measurement point descriptions matching each first measurement point description, and generate a telecontrol configuration description file based on each piece of telecontrol configuration description information.
[0121] In one embodiment, a word segmentation determination module 402 is configured to perform word segmentation processing on each first measurement point description in the detection information table based on a first historical dictionary regarding the detection information table, to obtain a word segmentation result of each first measurement point description; for each successful word segmentation result indicating successful word segmentation, determine at least one first word segment included in the corresponding first measurement point description based on the successful word segmentation result; for each failed word segmentation result indicating failed word segmentation, determine characters in a field where word segmentation fails in the corresponding first measurement point description, and determine multiple candidate words based on the characters in the field where word segmentation fails; for each candidate word, determine a confidence level of the candidate word according to a position of the candidate word in the first measurement point description, and perform re-word segmentation on the field where word segmentation fails based on the confidence level of each candidate word, to determine at least one first word segment included in the first measurement point description corresponding to the failed word segmentation result.
[0122] In one embodiment, the word segmentation determination module 402 is configured to determine an average word length in the first historical dictionary, and determine a corresponding boundary coefficient according to a position of the candidate word in the first measurement point description, where the position includes a boundary position and a non-boundary position, and the boundary coefficient corresponding to the boundary position is greater than the boundary coefficient corresponding to the non-boundary position; calculate a left entropy and a right entropy of the candidate word respectively according to a preset window, and determine the confidence level of the candidate word based on the boundary coefficient, the left entropy, the right entropy, and the average word length.
[0123] In one embodiment, the word segmentation determination module 402 is configured to determine new word segments obtained after performing re-word segmentation on the field where word segmentation fails; add the new word segments to the first historical dictionary to obtain a first updated dictionary, and update the average word length based on the first updated dictionary.
[0124] In one embodiment, a measurement point description screening module 404 is configured to obtain a preset word increment, determine a word quantity range corresponding to the first measurement point description according to a word quantity of first word segments included in the first measurement point description and the word increment; screen out second measurement point descriptions whose word quantity is within the word quantity range from multiple second measurement point descriptions according to a word quantity of second word segments included in each second measurement point description, and use the screened-out second measurement point descriptions as candidate measurement point descriptions.
[0125] In one embodiment, the word segmentation screening module 406 is configured to, for each first word segment, if there is no second word segment identical to the first word segment among the second word segments involved in at least one candidate measurement point description, query for synonyms synonymous with the second word segment from the synonym library corresponding to the substation; if there is such a synonym among the second word segments involved in at least one candidate measurement point description, use the synonym as the second word segment that matches the first word segment.
[0126] Each module in the above device for generating a telecontrol configuration description file can be implemented in whole or in part by software, hardware, or a combination thereof. The above modules can be embedded in the processor in the computer device in hardware form or be independent of the processor, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0127] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for generating a telecontrol configuration description file.
[0128] Those skilled in the art can understand that Figure 5 the structure shown in [the figure] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0129] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0130] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.
[0131] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.
[0132] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0133] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0134] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0135] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for generating a remote control configuration description file, characterized in that, The method includes: Determine at least one first word segment included in each first measurement point description in the detection information table regarding the substation, and determine at least one second word segment included in each second measurement point description in the substation configuration description file; For each first measurement point description, based on the number of word segments of the first word segments included in the first measurement point description, screen out at least one candidate measurement point description from multiple second measurement point descriptions; For each first word segment of the first measurement point description, screen out the second word segments that match the first word segment from the second word segments involved in at least one candidate measurement point description; Based on the second word segments that match each first word segment in each first measurement point description, determine the candidate measurement point descriptions that match each first measurement point description; According to the substation configuration description file, determine the telecontrol configuration description information of the candidate measurement point descriptions that match each first measurement point description, and generate a telecontrol configuration description file based on each telecontrol configuration description information.
2. The method according to claim 1, wherein The determination of at least one first word segment included in each first measurement point description in the detection information table regarding the substation includes: Based on the first historical dictionary regarding the detection information table, perform word segment processing on each first measurement point description in the detection information table to obtain the word segment result of each first measurement point description; For each successful word segment result indicating successful word segmentation, determine at least one first word segment included in the corresponding first measurement point description based on the successful word segment result; For each failed word segment result indicating failed word segmentation, based on the failed word segment result, determine the characters in the field where word segmentation was not successful in the corresponding first measurement point description, and determine multiple candidate words based on the characters in the field where word segmentation was not successful; For each candidate word, determine the confidence level of the candidate word according to the position of the candidate word in the first measurement point description, and based on the confidence level of each candidate word, re-segment the field where word segmentation was not successful to determine at least one first word segment included in the first measurement point description corresponding to the failed word segment result.
3. The method according to claim 2, wherein The determination of the confidence level of the candidate word according to the position of the candidate word in the first measurement point description includes: Determine the average word length in the first historical dictionary, and determine the corresponding boundary coefficient according to the position of the candidate word in the first measurement point description. The position includes boundary positions and non-boundary positions, and the boundary coefficient corresponding to the boundary position is greater than the boundary coefficient corresponding to the non-boundary position; According to a preset window, calculate the left entropy and right entropy of the candidate word respectively, and determine the confidence level of the candidate word based on the boundary coefficient, left entropy, right entropy, and average word length.
4. The method according to claim 2, wherein The method further includes: Determine the new word segments obtained after re-segmenting the field where word segmentation was not successful; Add the new word segments to the first historical dictionary to obtain a first updated dictionary, and update the average word length based on the first updated dictionary.
5. The method according to claim 1, wherein The screening of at least one candidate measurement point description from multiple second measurement point descriptions according to the number of word segments of the first word segments included in the first measurement point description includes: Obtain a preset word increment, and determine the word number range corresponding to the first measurement point description according to the number of word segments of the first word segments included in the first measurement point description and the word increment; According to the number of words of the second participles included in each second measurement point description, screen out the second measurement point descriptions whose number of words is within the word number range from multiple second measurement point descriptions, and use the screened second measurement point descriptions as candidate measurement point descriptions.
6. The method according to claim 1, wherein The screening out of the second participles that match the first participle from the second participles involved in at least one candidate measurement point description includes: For each first participle, if there is no second participle identical to the first participle among the second participles involved in at least one candidate measurement point description, query the synonyms synonymous with the second participle from the synonym library corresponding to the substation; If the synonym exists among the second participles involved in at least one candidate measurement point description, use the synonym as the second participle that matches the first participle.
7. An apparatus for generating a telecontrol configuration description file, characterized in that The device includes: A participle determination module, configured to determine at least one first participle included in each first measurement point description in the detection information table regarding the substation, and determine at least one second participle included in each second measurement point description in the substation configuration description file; A measurement point description screening module, configured to, for each first measurement point description, screen out at least one candidate measurement point description from multiple second measurement point descriptions according to the number of words of the first participles included in the first measurement point description; A participle screening module, configured to, for each first participle of the first measurement point description, screen out the second participles that match the first participle from the second participles involved in at least one candidate measurement point description; A measurement point description matching module, configured to determine the candidate measurement point descriptions that match each first measurement point description based on the second participles that match each first participle in each first measurement point description; A file generation module, configured to determine the telecontrol configuration description information of the candidate measurement point descriptions that match each first measurement point description according to the substation configuration description file, and generate a telecontrol configuration description file based on each telecontrol configuration description information.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.