A method and device for classifying device information

The method and apparatus use a maximum probability path algorithm to segment and classify electric power secondary device data, addressing disorganization and loss issues, enhancing data integration and retrieval efficiency.

CN113918723BActive Publication Date: 2025-07-15GUANGDONG POWER GRID CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111415304.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-25
Publication Date
2025-07-15
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

In the prior art, the operating data of secondary equipment is stored in an unstructured format, resulting in information confusion, missing and lost, and increasing the difficulty of maintenance.

Method used

The maximum probability path algorithm is used to process the word segmentation of the running data, calculate the data similarity value, and realize the classification storage of the data.

Benefits of technology

It realizes the integration and classification of equipment information, reduces the user's data processing burden, and improves data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113918723B_ABST
    Figure CN113918723B_ABST
Patent Text Reader

Abstract

The present invention discloses a classification method and device for device information. The method includes: obtaining operation data of a secondary device to be processed; performing word segmentation processing on the operation data by using a maximum probability path algorithm to obtain word segmentation data; calculating a data similarity value between the secondary device to be processed and the devices stored previously by using the word segmentation data; and storing the operation data of the secondary device to be processed in a classified manner according to the data similarity value. The present invention can obtain the operation data of a device, perform word segmentation on the device based on the operation data of the device, and finally determine whether the operation data of the device is the operation data of a stored device according to the word segmentation data, so as to achieve the effects of data integration and classification, thereby reducing the burden of users in processing data and improving the data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment information arrangement, and particularly to a classification method and device for equipment information. Background Art

[0002] Power secondary equipment (referred to as secondary equipment for short) is auxiliary equipment for monitoring, measuring, controlling, protecting, and regulating primary equipment in a power system. That is, equipment that does not directly contact electrical energy. With the gradual expansion of the scale of the power system, the requirements for the safety and reliability of the power system are becoming increasingly strict. Therefore, the maintenance and management of power secondary equipment have become particularly important.

[0003] Due to the variety and large quantity of secondary equipment, it is necessary to update and maintain the data of secondary equipment in a timely manner. The traditional method of data update and maintenance is to manually perform on-site maintenance on the equipment, fill in the corresponding documents after the maintenance, and finally enter them into the management system (such as a computer), and then the management system processes them uniformly.

[0004] However, the currently commonly used data management method has the following technical problems: Since the operation data such as defects, maintenance, and setting sheets of secondary equipment accumulated in the early stage are usually stored in an unstructured data format, the storage of each piece of information is relatively chaotic, there is no correlation between information and information, it is difficult for users to query information, and the information entered later often easily overwrites the previous information, which easily leads to the lack and loss of information of different secondary equipment, making it impossible for technicians to view the previous working conditions of the equipment, thereby increasing the difficulty of maintenance and repair of the equipment by technicians. Summary of the Invention

[0005] The present invention proposes a classification method and device for equipment information. The method can perform word segmentation recognition on the operation data of the equipment to determine the equipment type corresponding to the operation data, so as to achieve the effect of integrating and classifying data.

[0006] The first aspect of the embodiment of the present invention provides a classification method for equipment information, and the method includes: obtaining the operation data of the secondary equipment to be processed;

[0007] Performing word segmentation processing on the operation data by using the maximum probability path algorithm to obtain segmented data;

[0008] Calculating the data similarity value between the secondary equipment to be processed and the previously stored equipment by using the segmented data;

[0009] Classifying and storing the operation data of the secondary equipment to be processed according to the data similarity value.

[0010] In a possible implementation of the first aspect, the operation data includes historical operation data;

[0011] The tokenizing the operation data by using the maximum probability path algorithm to obtain tokenized data includes:

[0012] Searching for descriptive text data for describing the state of the secondary device to be processed from the historical operation data;

[0013] Extracting a plurality of phrases from the descriptive text data according to a preset vocabulary, and obtaining the phrase order of the plurality of phrases;

[0014] Taking each phrase in the phrase order as a vertex, drawing a directed graph of phrases with the plurality of phrases according to the phrase order, and assigning weights to the paths of every two directly connected vertices in the directed graph of phrases to obtain a plurality of weights;

[0015] Calculating the tokenization probability values corresponding to each tokenization scheme in a preset tokenization group for the plurality of weights by using the maximum probability path algorithm to obtain a plurality of tokenization probability values, where a plurality of tokenization schemes are stored in the preset tokenization group;

[0016] Selecting the target tokenization probability value with the largest value from the plurality of tokenization probability values, and dividing the historical operation data according to the tokenization scheme corresponding to the target tokenization probability value to obtain tokenized data.

[0017] In a possible implementation of the first aspect, the operation data includes real-time operation data;

[0018] The tokenizing the operation data by using the maximum probability path algorithm to obtain tokenized data includes:

[0019] Obtaining the data type corresponding to the real-time operation data;

[0020] Searching for the corresponding tokenization rule according to the data type;

[0021] Dividing the real-time operation data according to the tokenization rule to obtain tokenized data.

[0022] In a possible implementation of the first aspect, the calculating the data similarity value between the secondary device to be processed and the previously stored devices by using the tokenized data includes:

[0023] Calculating the data similarity value between the tokenized data and the previously stored data corresponding to each previously stored device by using a similarity calculation algorithm.

[0024] In a possible implementation of the first aspect, the classifying and storing the operation data of the secondary device to be processed according to the data similarity value includes:

[0025] When the data similarity value is greater than a preset threshold, classify and store the operation data of the secondary device to be processed in the prior device database of the prior storage device corresponding to the preset threshold;

[0026] When the data similarity value is less than a preset threshold, create a storage database to be processed corresponding to the secondary device to be processed, and classify and store the operation data of the secondary device to be processed in the storage database to be processed.

[0027] The second aspect of the embodiments of the present invention provides a classification device for device information. The device includes: an acquisition module, configured to acquire operation data of a secondary device to be processed;

[0028] A word segmentation module, configured to perform word segmentation processing on the operation data by using a maximum probability path algorithm to obtain segmented data;

[0029] A calculation module, configured to calculate a data similarity value between the secondary device to be processed and a previously stored device by using the segmented data;

[0030] A classification module, configured to classify and store the operation data of the secondary device to be processed according to the data similarity value.

[0031] In a possible implementation manner of the second aspect, the operation data includes historical operation data;

[0032] The word segmentation module is further configured to:

[0033] Search for descriptive text data for describing the state of the secondary device to be processed from the historical operation data;

[0034] Extract a plurality of phrases from the descriptive text data according to a preset word library, and obtain the phrase order of the plurality of phrases;

[0035] Use each phrase in the phrase order as a vertex, draw the plurality of phrases into a phrase directed graph according to the phrase order, and assign weights to the paths of every two directly connected vertices in the phrase directed graph to obtain a plurality of weights;

[0036] Calculate the word segmentation probability values corresponding to each word segmentation scheme in a preset word segmentation group by using the maximum probability path algorithm for the plurality of weights, to obtain a plurality of word segmentation probability values, where a plurality of word segmentation schemes are stored in the preset word segmentation group;

[0037] Screen the largest target word segmentation probability value from the plurality of word segmentation probability values, and divide the historical operation data according to the word segmentation scheme corresponding to the target word segmentation probability value to obtain segmented data.

[0038] In a possible implementation of the second aspect, the operation data includes real-time operation data;

[0039] The word segmentation module is further configured to:

[0040] Obtain the data type corresponding to the real-time operation data;

[0041] Find the corresponding word segmentation rule according to the data type;

[0042] Divide the real-time operation data according to the word segmentation rule to obtain word segmentation data.

[0043] In a possible implementation of the second aspect, the calculation module is further configured to:

[0044] Use a similarity calculation algorithm to calculate the data similarity values between the word segmentation data and the prior data corresponding to each previously stored device.

[0045] In a possible implementation of the second aspect, the classification module is further configured to:

[0046] When the data similarity value is greater than a preset threshold, classify and store the operation data of the secondary device to be processed in the prior device database of the previously stored device corresponding to the preset threshold;

[0047] When the data similarity value is less than the preset threshold, create a to-be-processed storage database corresponding to the secondary device to be processed, and classify and store the operation data of the secondary device to be processed in the to-be-processed storage database.

[0048] Compared with the prior art, the classification method and device for device information provided by the embodiments of the present invention have the beneficial effects that: the present invention can obtain the operation data of the device, perform word segmentation on the device based on the operation data of the device, and finally determine whether the operation data of the device is the operation data of the stored device according to the word segmentation data, so as to achieve the effect of data integration and classification, thereby reducing the burden of users in processing data and improving the data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a schematic flowchart of a classification method for device information provided by an embodiment of the present invention;

[0050] Figure 2 is a schematic structural diagram of a classification device for device information provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.

[0052] The currently commonly used data management methods have the following technical problems: Since the operation data such as defects, maintenance, and setting sheets of secondary equipment recorded and accumulated in the early stage are usually stored in an unstructured data format, the storage of each piece of information is relatively chaotic, and there is no any correlation between information and information, making it difficult for users to query information. Moreover, the information entered later often easily overwrites the information entered earlier, easily leading to situations such as missing and lost information of different secondary equipment, making it impossible for technicians to view the previous working conditions of the equipment, and thus increasing the difficulty of maintenance and repair of the equipment by technicians.

[0053] To solve the above problems, the following specific embodiments will be used to introduce and illustrate in detail a classification method for equipment information provided by the embodiments of the present application.

[0054] Referring to Figure 1 , a flowchart showing a classification method for equipment information provided by an embodiment of the present invention is shown.

[0055] Among them, by way of example, the classification method for equipment information may include:

[0056] S11. Obtain the operation data of the secondary equipment to be processed.

[0057] The secondary equipment to be processed may be secondary equipment that has not been classified and stored. The operation data may be various state data of the secondary equipment during operation. For example, it may be maintenance information, alarm information, monitoring information, various setting information, defect information, and so on.

[0058] The operation data of the secondary equipment to be processed can be obtained by real-time detection and back-checking the data recorded in the past.

[0059] S12. Perform word segmentation processing on the operation data using the maximum probability path algorithm to obtain segmented data.

[0060] Since the operation data of the secondary equipment to be processed contains various information contents, if all the information contents are screened and judged one by one, it will increase the processing time and reduce the processing efficiency. The operation data can be segmented correspondingly to obtain the keywords in the operation data, and finally the equipment can be classified based on its keywords to improve the data processing efficiency.

[0061] In an optional embodiment, the operation data includes historical operation data.

[0062] Since the previous historical data may be filled in or processed by different maintenance personnel, and the formats and data structures are different when filled in by different people. In order to perform word segmentation on data with different structures or formats, as an example, step S12 may include the following sub-steps:

[0063] Sub-step S121: Search for descriptive text data in the historical operation data that describes the status of the secondary equipment to be processed.

[0064] Optionally, the descriptive text data may be the evaluation content of the secondary equipment by previous maintenance personnel during the maintenance or repair of the secondary equipment. For example, it may be leakage, good, aging, etc.; or for another example, the equipment has been used for 10 years and there has been a situation of unstable voltage, etc.

[0065] Sub-step S122: Extract a number of phrases from the descriptive text data according to a preset word library, and obtain the phrase order of the number of phrases.

[0066] The preset word library may contain multiple user-preset words, and each word may correspond to a phrase.

[0067] In application, phrase extraction can be performed on the descriptive text data according to various words contained in the preset word library, so that a number of phrases can be obtained.

[0068] After obtaining a number of phrases, the number of phrases can be sorted according to the time order of obtaining the number of phrases, and their arrangement order can be obtained to get the phrase order.

[0069] Sub-step S123: Take each phrase in the phrase order as a vertex, draw the number of phrases into a phrase directed graph according to the phrase order, and assign weights to the paths of every two directly connected vertices in the phrase directed graph to obtain a number of weights.

[0070] For example, for vertex A → vertex B, the path weight between vertices A and B is the weight of B (if B is the end vertex, the weight is 0).

[0071] At this time, the original problem is transformed into a single-source shortest path problem, and the optimal solution can be obtained by dynamic programming.

[0072] Sub-step S124: Calculate the word segmentation probability values corresponding to each word segmentation scheme in the preset word segmentation group for the number of weights by using the maximum probability path algorithm to obtain a plurality of word segmentation probability values, and there are multiple word segmentation schemes stored in the preset word segmentation group.

[0073] In one embodiment, a user can set multiple word segmentation schemes, and each word segmentation scheme can correspond to a word segmentation rule. Since the historical operation data in the past may be filled in by different maintenance personnel, and different maintenance personnel may have different writing habits, a word segmentation scheme can be correspondingly divided according to the writing habit of each maintenance personnel, and finally multiple word segmentation schemes are aggregated to obtain a preset word segmentation group.

[0074] During application, the maximum probability path algorithm can be used to calculate the word segmentation probability values that are close to each other between the several weights and each word segmentation scheme respectively, so that multiple word segmentation probability values can be obtained.

[0075] Sub-step S125: Screen the target word segmentation probability value with the largest value from the multiple word segmentation probability values, and segment the historical operation data according to the word segmentation scheme corresponding to the target word segmentation probability value to obtain segmented data.

[0076] Since each word segmentation probability value represents the proximity distance between the several weights and the word segmentation scheme, screening the target word segmentation probability value with the largest value from multiple word segmentation probability values can obtain the word segmentation scheme that is the closest, so that the historical operation data can be segmented using the word segmentation rules included in this word segmentation scheme, and finally the corresponding segmented data can be obtained.

[0077] Specifically, during the word segmentation process, the several input phrases can be a string T1, T2,..., T n , and the output segmented data can be a word string S = W1, W2,..., W m , where m <= n. For a specific string T, there may be multiple corresponding word segmentation schemes S, and the word segmentation operation is to find the word segmentation scheme with the largest probability among these word segmentation schemes S, that is, to segment the most likely word sequence from the input string.

[0078] Calculate the probability that the word segmentation scheme for the target sentence T is S, where S = {s1, s2,..., s m}:

[0079] Find the word segmentation scheme S that maximizes the probability, and its calculation formula is as follows:

[0080]

[0081] Its algorithm method can be: first, for a sub-string T to be segmented, all candidate words s1, s2,..., s are taken out in order from left to right m , then the probability values P(s i ) of each candidate word are found in the dictionary, and all the left adjacent words of each candidate word are recorded; then according to each calculated cumulative probability, the best left adjacent word of each candidate word is compared at the same time; if the current word s mis the last word of the string T, and the cumulative probability P(s m ) is the largest, then s m is the end word of T; starting from s m , in the order from right to left, the best left neighbor word of each word is output in turn, which is the word segmentation result of T.

[0082] By screening the corresponding word segmentation rules for word segmentation, the historical operation data can be effectively segmented correspondingly to improve the accuracy of word segmentation.

[0083] In one embodiment, the operation data includes real-time operation data.

[0084] The real-time operation data is the data detected by the secondary equipment during real-time operation, and this data can be the data detected by the maintenance personnel on-site in real time at the current time node.

[0085] Since there are many types of secondary equipment, there may be newly added secondary equipment for maintenance, and the newly added secondary equipment may not be classified. It is necessary to perform corresponding word segmentation on the real-time operation data of the secondary equipment to determine the type of the secondary equipment, so that corresponding information storage management can be performed according to its equipment type.

[0086] In an alternative embodiment, step S12 may include the following sub-steps:

[0087] Sub-step S126, obtain the data type corresponding to the real-time operation data.

[0088] Specifically, the data type may be the input data format type.

[0089] Sub-step S127, find the corresponding word segmentation rule according to the data type.

[0090] In one embodiment, when the maintenance personnel perform maintenance on the newly added secondary equipment, the way they fill in the data is preset by the user. Therefore, the newly added real-time operation data can be added and collected in the way preset by the user. Therefore, by obtaining the data type, the corresponding data format type can be obtained.

[0091] After obtaining its data format type, the corresponding word segmentation rule can be found based on the data format type. The word segmentation rule can also be preset by the user. Different word segmentation rules can correspond to different data format types.

[0092] Sub-step S128, divide the real-time operation data according to the word segmentation rule to obtain segmented data.

[0093] After determining the word segmentation rule, the real-time operation data can be divided according to the word segmentation rule, so as to obtain the corresponding segmented data.

[0094] By quickly searching for the data types of the real-time operation data of newly added secondary devices, the word segmentation rules for the real-time operation data can be quickly determined, and finally word segmentation can be performed according to these word segmentation rules to improve the word segmentation efficiency.

[0095] S13. Calculate the data similarity value between the secondary device to be processed and the previously stored devices by using the word segmentation data.

[0096] After obtaining the word segmentation data, the similarity value can be calculated based on the word segmentation data and the data of the previously stored secondary devices, so as to determine whether the current secondary device to be processed is a previously stored secondary device.

[0097] Since the various status data of the same device are similar, through the calculation of the similarity value of the data, the association between the data can be effectively determined, and it can be quickly determined whether the current secondary device to be processed is a previously stored secondary device.

[0098] In one embodiment, step S13 may include the following sub-steps:

[0099] Sub-step S131. Calculate the data similarity value between the word segmentation data and the previous data corresponding to each previously stored device by using a similarity calculation algorithm.

[0100] Specifically, each previously stored secondary device has its corresponding storage space to store its corresponding data, and the data similarity value between the word segmentation data and the data corresponding to each previously stored secondary device can be calculated.

[0101] S14. Classify and store the operation data of the secondary device to be processed according to the data similarity value.

[0102] After determining the data similarity value, it can be determined whether the current secondary device to be processed is a previously stored secondary device according to the size of the value, so that the operation data of the current secondary device to be processed can be stored correspondingly.

[0103] In order to accurately store the operation data of the current secondary device to be processed, in one embodiment, step S14 may include the following sub-steps:

[0104] Step S141. When the data similarity value is greater than a preset threshold, classify and store the operation data of the secondary device to be processed into the previous device database of the previously stored device corresponding to the preset threshold.

[0105] Step S142. When the data similarity value is less than the preset threshold, create a storage database to be processed corresponding to the secondary device to be processed, and classify and store the operation data of the secondary device to be processed into the storage database to be processed.

[0106] In another embodiment, since the segmented data will be used to calculate similarity values with the data of multiple prior secondary devices, multiple data similarity values can be calculated. If there are multiple data similarity values greater than a preset threshold, then the data similarity value with the largest value is selected from the multiple data similarity values, and the storage area is determined based on the data similarity value with the largest value.

[0107] In this embodiment, the embodiment of the present invention provides a method for classifying device information. The beneficial effect is that the present invention can obtain the operation data of the device, segment the device based on the operation data of the device, and finally determine whether the operation data of the device is the operation data of the stored device according to the segmented data, so as to achieve the effect of data integration and classification, thereby reducing the burden of the user to process data and improving the data processing efficiency.

[0108] The embodiment of the present invention also provides a device information classification device. Refer to Figure 2 , which shows a schematic structural diagram of a device information classification device provided by an embodiment of the present invention.

[0109] Among them, by way of example, the device information classification device may include:

[0110] An acquisition module 201, configured to acquire the operation data of the secondary device to be processed;

[0111] A segmentation module 202, configured to perform segmentation processing on the operation data by using the maximum probability path algorithm to obtain segmented data;

[0112] A calculation module 203, configured to calculate the data similarity value between the secondary device to be processed and the previously stored devices by using the segmented data;

[0113] A classification module 204, configured to classify and store the operation data of the secondary device to be processed according to the data similarity value.

[0114] Optionally, the operation data includes historical operation data;

[0115] The segmentation module is further configured to:

[0116] Search for descriptive text data for describing the state of the secondary device to be processed from the historical operation data;

[0117] Extract several phrases from the descriptive text data according to a preset word library, and obtain the phrase order of the several phrases;

[0118] Take each phrase in the phrase order as a vertex, draw the several phrases into a phrase directed graph according to the phrase order, and assign weights to the paths of every two directly connected vertices in the phrase directed graph to obtain several weights;

[0119] Use the maximum probability path algorithm to calculate the word segmentation probability values corresponding to each word segmentation scheme in the preset word segmentation group for the several weights, and obtain multiple word segmentation probability values. Multiple word segmentation schemes are stored in the preset word segmentation group;

[0120] Select the target word segmentation probability value with the largest numerical value from the multiple word segmentation probability values, and divide the historical operation data according to the word segmentation scheme corresponding to the target word segmentation probability value to obtain word segmentation data.

[0121] Optionally, the operation data includes real-time operation data;

[0122] The word segmentation module is further configured to:

[0123] Obtain the data type corresponding to the real-time operation data;

[0124] Find the corresponding word segmentation rule according to the data type;

[0125] Divide the real-time operation data according to the word segmentation rule to obtain word segmentation data.

[0126] Optionally, the calculation module is further configured to:

[0127] Use the similarity calculation algorithm to calculate the data similarity values between the word segmentation data and the prior data corresponding to each previously stored device.

[0128] Optionally, the classification module is further configured to:

[0129] When the data similarity value is greater than the preset threshold, classify and store the operation data of the secondary device to be processed into the prior device database of the previously stored device corresponding to the preset threshold;

[0130] When the data similarity value is less than the preset threshold, create a to-be-processed storage database corresponding to the secondary device to be processed, and classify and store the operation data of the secondary device to be processed into the to-be-processed storage database.

[0131] Furthermore, an embodiment of the present application further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the classification method of device information as described in the above embodiment.

[0132] Further, an embodiment of the present application also provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the method for classifying device information as described in the above embodiment.

[0133] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also considered within the protection scope of the present invention.

Claims

1. A method for classifying device information, characterized in that, The method includes: Obtaining the operation data of the secondary device to be processed; Performing word segmentation processing on the operation data by using the maximum probability path algorithm to obtain word segmentation data; Calculating the data similarity value between the secondary device to be processed and the previously stored devices by using the word segmentation data; Classifying and storing the operation data of the secondary device to be processed according to the data similarity value; The operation data includes historical operation data; The performing word segmentation processing on the operation data by using the maximum probability path algorithm to obtain word segmentation data includes: Searching the historical operation data for descriptive text data used to describe the state of the secondary device to be processed; Extracting a plurality of phrases from the descriptive text data according to a preset thesaurus, and obtaining the phrase order of the plurality of phrases; Taking each phrase in the phrase order as a vertex, drawing a phrase directed graph of the plurality of phrases according to the phrase order, and assigning weights to the paths of every two directly connected vertices in the phrase directed graph to obtain a plurality of weights; Calculating the word segmentation probability values corresponding to each word segmentation scheme in a preset word segmentation group by using the maximum probability path algorithm for the plurality of weights, to obtain a plurality of word segmentation probability values, where a plurality of word segmentation schemes are stored in the preset word segmentation group; Screening the target word segmentation probability value with the largest value from the plurality of word segmentation probability values, and dividing the historical operation data according to the word segmentation scheme corresponding to the target word segmentation probability value to obtain word segmentation data; The operation data includes real-time operation data; The performing word segmentation processing on the operation data by using the maximum probability path algorithm to obtain word segmentation data includes: Obtaining the data type corresponding to the real-time operation data; Searching for the corresponding word segmentation rule according to the data type; Dividing the real-time operation data according to the word segmentation rule to obtain word segmentation data.

2. The classification method of device information according to claim 1, characterized in that, The calculating the data similarity value between the secondary device to be processed and the previously stored devices by using the word segmentation data includes: Calculating the data similarity value between the word segmentation data and the previously stored data corresponding to each previously stored device by using a similarity calculation algorithm.

3. The classification method of device information according to claim 2, characterized in that The classifying and storing the operation data of the secondary device to be processed according to the data similarity value includes: When the data similarity value is greater than a preset threshold, classifying and storing the operation data of the secondary device to be processed into the prior device database of the previously stored device corresponding to the preset threshold; When the data similarity value is less than a preset threshold, creating a to-be-processed storage database corresponding to the secondary device to be processed, and classifying and storing the operation data of the secondary device to be processed into the to-be-processed storage database.

4. A classification device for device information, characterized in that, The device includes An obtaining module, configured to obtain the operation data of the secondary device to be processed; A word segmentation module, configured to perform word segmentation processing on the operation data by using the maximum probability path algorithm to obtain word segmentation data; A calculating module, configured to calculate the data similarity value between the secondary device to be processed and the previously stored devices by using the word segmentation data; A classifying module, configured to classify and store the operation data of the secondary device to be processed according to the data similarity value; The operation data includes historical operation data; The word segmentation module is further configured to: Find the descriptive text data for describing the status of the secondary device to be processed from the historical operation data; Extract several phrases from the descriptive text data according to a preset word library, and obtain the phrase order of the several phrases; Take each phrase in the phrase order as a vertex, draw the several phrases into a phrase directed graph according to the phrase order, and assign weights to the paths of every two directly connected vertices in the phrase directed graph to obtain several weights; Use the maximum probability path algorithm to calculate the word segmentation probability values corresponding to each word segmentation scheme in the preset word segmentation group for the several weights, and obtain multiple word segmentation probability values. Multiple word segmentation schemes are stored in the preset word segmentation group; Screen the target word segmentation probability value with the largest value from the multiple word segmentation probability values, and segment the historical operation data according to the word segmentation scheme corresponding to the target word segmentation probability value to obtain segmented data; The operation data includes real-time operation data; The word segmentation module is further configured to: Obtain the data type corresponding to the real-time operation data; Find the corresponding word segmentation rule according to the data type; Segment the real-time operation data according to the word segmentation rule to obtain segmented data.

5. The classification device for device information according to claim 4, characterized in that The calculation module is further configured to: Use a similarity calculation algorithm to calculate the data similarity values between the segmented data and the prior data corresponding to each prior stored device; 6. The classification device for device information according to claim 5, characterized in that, The classification module is further configured to: When the data similarity value is greater than a preset threshold, classify and store the operation data of the secondary device to be processed in the prior device database of the prior storage device corresponding to the preset threshold; When the data similarity value is less than a preset threshold, create a to-be-processed storage database corresponding to the secondary device to be processed, and classify and store the operation data of the secondary device to be processed in the to-be-processed storage database.

Citation Information

Patent Citations

  • Method and device for matching texts

    CN102411583A

  • Power distribution network secondary equipment type identification method and system

    CN106022950A