Structural information processing method and device, computer readable storage medium and electronic equipment
By building a structural information knowledge base and using prefix trees and tree charts to match and sort, the problem that structural information cannot be uniformly standardized, and efficient standardization and consistency processing of structural information is achieved.
Patent Information
- Application Number
- CN202510500388.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The prior art cannot effectively standardize structural information, which makes it difficult to ensure consistency of structural information.
Build a structural information knowledge base, by building a prefix tree and tree chart, splitting structural information using keywords and regular expressions, matching and sorting, obtaining standardized structural information, and updating the knowledge base.
It realizes efficient standardization processing and consistency of structural information, records the optimization and adjustment process of structural information, and ensures that structural information from different sources is standardized and consistent.
Smart Images

Figure CN120578784A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data modeling and analysis, and in particular to a method and device for processing structural information, a computer-readable storage medium, and an electronic device. Background Art
[0002] In the process of modernization, structural information will be continuously optimized and adjusted to make it more adaptable to the needs of modern scenarios. In order to ensure the integrity and consistency of structural information, a scientific and standardized method is needed to adjust the structural information accordingly.
[0003] However, manual adjustment methods are prone to produce erroneous information and cause a waste of human resources. In addition, changes such as adjustment or cancellation of elements in structural information often occur during the development process, which increases the difficulty of adjusting structural information.
[0004] Therefore, related technologies cannot effectively perform unified and standardized processing on structural information, and cannot guarantee the consistency of structural information. Summary of the Invention
[0005] The main purpose of the present disclosure is to provide a structural information processing method, device, computer-readable storage medium and electronic device to solve the problem in related technologies that structural information cannot be effectively and uniformly standardized and the consistency of structural information cannot be guaranteed.
[0006] In order to achieve the above objectives, the first aspect of the present disclosure provides a method for processing structure information, comprising:
[0007] Constructing a structure information knowledge base, wherein the structure information knowledge base includes structure information and corresponding states, wherein the structure information includes at least one of the following: currently used standard structure information, historical structure information, and to-be-confirmed structure information; the state corresponding to the standard structure information is a standard state, the state corresponding to the historical structure information is a historical state, and the to-be-confirmed structure information is a to-be-confirmed state;
[0008] Constructing a prefix tree and a tree graph corresponding to the structural information according to the structural information and the corresponding state;
[0009] Determine a set of structure information tuples to be matched based on the structure information; one piece of structure information corresponds to one tuple in the set of structure information tuples to be matched;
[0010] For each tuple in the set of structure information tuples to be matched, matching the tuple with the prefix tree and the tree graph respectively to obtain a set of candidate structure information corresponding to the tuple, and calculating the maximum matching degree between the tuple and each candidate structure information in the set of candidate structure information;
[0011] Determining target structure information corresponding to the structure information based on the target candidate structure information corresponding to the maximum matching degree, wherein the target structure information is used to represent normalized structure information;
[0012] The structure information knowledge base is updated according to the target structure information to standardize the structure information in the structure information knowledge base, and the association relationship between the historical structure information and the standard structure information in the structure information knowledge base is recorded.
[0013] Optionally, further, determining a set of structure information tuples to be matched based on the structure information includes:
[0014] Based on the keywords, the data type and data features of the structural information, multiple rounds of recursive regular matching are used to extract the hierarchical information and attribute information from the structural information, and label the hierarchical information and attribute information with corresponding tags. The data types include: string type, value type, and encoding type; the data features include: string length, value precision, and encoding standard; and the attribute information includes at least one of the following: unit code, postal code, area code, longitude, and latitude;
[0015] The extracted information at each level is combined according to the level, and non-compliant combinations are eliminated to obtain a set of to-be-matched structural information tuples corresponding to the structural information.
[0016] Optionally, further, matching the tuple with the prefix tree and the tree graph respectively to obtain a candidate structure information set corresponding to the tuple includes:
[0017] Searching the tuple in the prefix tree to obtain a prefix tree matching result, wherein the prefix tree matching result includes all prefix paths to be matched and matched nodes;
[0018] According to the prefix tree matching result, the node corresponding to the tree diagram is located, and the matching status of the tuple with the node corresponding to the tree diagram and its child nodes is checked to obtain a candidate structure information set corresponding to the tuple.
[0019] Optionally, further, the calculating the maximum matching degree between the tuple and each candidate structure information in the candidate structure information set includes:
[0020] Converting the tuple and each candidate structure information in the candidate structure information set into a structure information vector respectively;
[0021] Obtaining a matching degree between the tuple and each candidate structure information in the candidate structure information set by calculating a matching degree of each component in the structure information vector corresponding to each candidate structure information in the candidate structure information set;
[0022] The matching degrees between the tuple and each candidate structure information in the candidate structure information set are sorted to obtain the maximum matching degree.
[0023] Optionally, further, determining the target structure information corresponding to the structure information according to the target candidate structure information corresponding to the maximum matching degree includes:
[0024] If the maximum matching degree is greater than or equal to a preset threshold, determining the target structure information corresponding to the structure information according to the state of the target candidate structure information recorded in the structure information knowledge base, and performing necessary information filling on the target candidate structure information based on the to-be-matched structure information corresponding to the tuple, so as to expand the structure information knowledge base; wherein the state includes a standard state, a historical state, and a pending confirmation state;
[0025] If the maximum matching degree is equal to 1 and the candidate structure information set contains one candidate structure information, the target candidate structure information is used as the target structure information, and based on the to-be-matched structure information corresponding to the tuple, the target candidate structure information is filled with necessary information to expand the structure information knowledge base;
[0026] If the maximum matching degree is equal to 0, the most complete structural information corresponding to the tuple is obtained, and the most complete structural information is marked as a pending confirmation state to update the structural information knowledge base.
[0027] Optionally, further, constructing a prefix tree and a tree graph corresponding to the structural information according to the structural information and the corresponding state includes:
[0028] Determine a key-value pair based on the structural information and the corresponding state; wherein the key in the key-value pair is the abbreviation of the unit or department, and the value in the key-value pair is a detailed list or set consisting of detailed information of the unit or department with the same abbreviation;
[0029] Based on the key-value pairs, construct the prefix tree, and record the corresponding status in the prefix tree;
[0030] generating a tree diagram according to the structural information;
[0031] The fusion information in the detailed information is a word segmentation array and corresponds to the nodes in the tree diagram.
[0032] Optionally, further, the constructing of a structure information knowledge base includes:
[0033] Obtaining structural information of each unit or part and corresponding structural optimization and adjustment log information; wherein the structural information includes at least one of the following: full name, abbreviation, alias or code name, hierarchy, unit code, parent unit code, postal code, area code, longitude, and latitude; the structural optimization and adjustment log information includes changes between old and new names and changes in recorded affiliation; the structural optimization and adjustment log information is used to query historical structural information and determine the status of the structural information;
[0034] A structural information knowledge base is generated based on the structural information of each unit or part and the corresponding structural optimization and adjustment log information; wherein the historical structural information and the standard structural information in the structural information knowledge base are associated with each other through foreign keys.
[0035] A second aspect of the present disclosure provides a structure information processing device, comprising:
[0036] a first processing unit configured to construct a structure information knowledge base, the structure information knowledge base including structure information and corresponding states, the structure information including at least one of the following: currently used standard structure information, historical structure information, and to-be-confirmed structure information; the state corresponding to the standard structure information being a standard state, the state corresponding to the historical structure information being a historical state, and the to-be-confirmed structure information being a to-be-confirmed state;
[0037] A second processing unit is configured to construct a prefix tree and a tree graph corresponding to the structural information according to the structural information and the corresponding state;
[0038] A third processing unit is configured to determine a set of structure information tuples to be matched based on the structure information; one piece of structure information corresponds to one tuple in the set of structure information tuples to be matched;
[0039] a fourth processing unit, configured to match each tuple in the set of structure information tuples to be matched with the prefix tree and the tree graph, respectively, to obtain a set of candidate structure information corresponding to the tuple, and to calculate a maximum degree of match between the tuple and each candidate structure information in the set of candidate structure information;
[0040] A fifth processing unit is configured to determine target structure information corresponding to the structure information based on the target candidate structure information corresponding to the maximum matching degree, wherein the target structure information is used to represent the normalized structure information;
[0041] The sixth processing unit is configured to update the structure information knowledge base according to the target structure information, so as to normalize the structure information in the structure information knowledge base, and record the association between the historical structure information and the standard structure information in the structure information knowledge base.
[0042] A third aspect of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the structural information processing method provided by any one of the first aspects.
[0043] The fourth aspect of the present disclosure provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor executes the structural information processing method provided in any one of the first aspects.
[0044] A fifth aspect of the present disclosure provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the structural information processing method provided by any one of the first aspects.
[0045] In the structural information processing method provided in the embodiment of the present disclosure, a structural information knowledge base is constructed by collecting structural information, and based on the structural information to be standardized and the corresponding state in the structural information knowledge base, a prefix tree and a tree diagram corresponding to the structural information are constructed to match the structural information to be matched corresponding to the structural information. By calculating the matching degree and sorting, standardized structural information is obtained to update the structural information knowledge base, and then the standardized structural information is output, and the state corresponding to the structural information is adjusted and the association between the historical structural information and the standard structural information in the structural information knowledge base is recorded. This can provide a basis for the adjustment process from historical structural information to standard structural information, achieve unified standardization of all structural information in the structural information knowledge base, and record the optimization adjustment process of the structural information at the same time, thereby achieving efficient standardization processing while tracking and recording the adjustment process of the structural information, so that the structural information from different sources has the technical effect of standardization and consistency, thereby solving the technical problem that the relevant technology cannot effectively perform unified standardization processing on the structural information and cannot ensure the consistency of the structural information. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0047] Figure 1 A flowchart of a method for processing structural information provided by an embodiment of the present disclosure;
[0048] Figure 2 A schematic diagram of a scenario of a structural information processing method provided in an embodiment of the present disclosure;
[0049] Figure 3 A flowchart of a method for processing structural information provided in yet another embodiment of the present disclosure;
[0050] Figure 4 ER diagram of the structural information knowledge base provided by the embodiment of the present disclosure;
[0051] Figure 5 A schematic diagram of a prefix tree data structure and a tree graph data structure provided by an embodiment of the present disclosure;
[0052] Figure 6 A block diagram of a structural information processing device provided in an embodiment of the present disclosure;
[0053] Figure 7 A block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0054] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.
[0055] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate for the embodiments of the present disclosure described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products, or apparatus.
[0056] In this disclosure, terms such as "upper," "lower," "left," "right," "front," "back," "top," "bottom," "inner," "outer," "center," "vertical," "horizontal," "transverse," and "longitudinal" indicate positions or locations based on the positions or locations shown in the accompanying drawings. These terms are primarily intended to better describe this disclosure and its embodiments and are not intended to limit the devices, elements, or components indicated to having a specific orientation, or to being constructed or operated in a specific orientation.
[0057] Furthermore, some of the above terms may be used to express other meanings besides indicating a position or location. For example, the term "on" may also be used to express a dependency or connection in certain circumstances. Those skilled in the art will understand the specific meanings of these terms in this disclosure based on the specific circumstances.
[0058] Furthermore, the terms "installed," "disposed," "provided with," "connected," "connected," and "socketed" should be interpreted broadly. For example, they can refer to fixed connections, removable connections, or integral structures; mechanical connections or electrical connections; direct connections, indirect connections through an intermediary, or internal communication between two devices, elements, or components. Those skilled in the art will understand the specific meanings of these terms in this disclosure based on the specific circumstances.
[0059] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0060] Currently, manual adjustments are prone to generating erroneous information and wasteful of human resources. Furthermore, changes such as the adjustment or removal of structural information elements often occur during development, increasing the difficulty of adjusting structural information. Consequently, related technologies are unable to effectively standardize structural information and guarantee its consistency.
[0061] In order to solve the above problems, the technical concept of the present invention is to use the structural information knowledge base to construct a structural information prefix tree and tree graph, and obtain standardized structural information by extracting a set of structural information tuples to be matched from the structural information, matching and sorting them with the nodes in the prefix tree and tree graph, and then updating the structural information knowledge base, so that the structural information from each system source is standardized and consistent.
[0062] In practical applications, the present disclosure may be implemented by a structure information processing device, which can be deployed in electronic devices such as terminals and servers. Users can use electronic devices equipped with the structure information processing device to process structure information from various sources, normalizing and ensuring consistency. The structure information may be organizational structure information, which is not specifically limited here.
[0063] The following describes the method for processing structural information in detail, taking the structural information as organizational structural information as an example.
[0064] For example, see Figure 1 As shown, Figure 1 A flow chart of a structural information processing method provided for an embodiment of the present disclosure. The structural information processing method here can be used to collect structural information in various business information systems (wherein, the format, data structure, information integrity, etc. of the structural information in different business information systems may be different; structural information refers to information that describes the organizational form and association method of the internal components of the system, object or data and their mutual relationships), and build a structural information knowledge base based on the structural information in each business information system, and perform standardization and consistency processing on the structural information knowledge base. Wherein, the structural information processing here can be applied to structural information related to any scenario, such as: security scenario, company scenario, school scenario, etc. For security scenario: structural information includes structural information of various security units (for example, M1-level security unit, M2-level security unit, M3-level security unit, etc.); for company scenario: structural information includes group, second-level unit, third-level unit, fourth-level unit, etc. For school scenario: school headquarters, branch school, etc., no specific limitation is made to the scenario and structural information here.
[0065] Specifically, the structure information processing method may include the following steps (combined with Figure 2 As shown, Figure 2 Scenario diagram of the structural information processing method provided in an embodiment of the present disclosure):
[0066] S101: Build a structural information knowledge base.
[0067] Among them, the structural information includes multiple levels of structure (taking the security scenario as an example below, such as: M1-level security unit, M2-level security unit, M3-level security unit, M4-level security unit, M5-level security unit, M6-level security unit, M7-level security unit, M8-level security unit, etc.), as well as the full name (name), abbreviation (short_name), alias or code (alias_name), level (level), unit code (code), parent unit code (parent_code), postal code (zip_code), area code (city_code), longitude (longitude), latitude (latitude) and other information of each unit (or unit) or department. The structural information knowledge base includes three types of structural information: currently used standardized structural information (that is, the currently used standard structural information, which will not be repeated below), historical structural information, and to-be-confirmed structural information. Optionally, the structural information knowledge base also includes the credibility corresponding to the structural information. When the credibility corresponding to the to-be-confirmed structural information reaches a threshold, the to-be-confirmed structural information can be adjusted to historical structural information or standard structural information.
[0068] (1.1) The currently used standardized structural information, as the standardized output result of the structural information disclosed herein, has a corresponding data state (or status, which will not be described in detail below) of status: S (here, referring to the standard state);
[0069] (1.2) Historical structural information that has been abandoned due to structural optimization adjustments such as reorganization and diversion is used to track and record the structural optimization and adjustment process. The corresponding data status is status:H (here refers to the historical status);
[0070] (1.3) The data status of the unconfirmed pending structure information is status: U (here refers to the pending confirmation status).
[0071] Standard structure information (S) and historical structure information (H) primarily come from documented sources, while unverified structure information (U) comes from external input within the system (e.g., a structure information processing system or a structure information normalization system, which will not be discussed further below). During system operation, the reliability of the unverified structure information is adjusted. When the reliability meets certain requirements, it is converted into standard structure information (S) or historical structure information (H).
[0072] S102: Constructing a structural information prefix tree and a tree graph.
[0073] The standard structure information (S), historical structure information (H) and to-be-confirmed structure information (U) in the information knowledge base are used to construct the structure information prefix tree data structure and tree graph data structure, and the status (status: S|H|U) is used to distinguish different types of structure data in the prefix tree nodes and tree graph nodes.
[0074] (2.1) In the prefix tree data structure, the root node is an empty node, and the remaining nodes consist of key-value pairs. The key is the abbreviation of the unit or department (SHORT_NAME), and the value is a list of detailed information about the unit or department with the same abbreviation (e.g., INFO_LIST, containing NAME, ALIAS_NAME, LEVEL, CODE, PARENT_CODE, ZIP_CODE, CITY_CODE, LONGITUDE, LATITUDE, INTEGRATE_INFO (integrated information, a word set of detailed information; for example, integrate_info: [BBM1-level security unit, CCM2-level security unit, DDM3-level security unit, EEM4-level security unit, FFN5-level security unit]), STATUS) (or a set).
[0075] (2.2) In the tree graph data structure, the root node is an empty node, and the remaining nodes are composed of hierarchical structures such as M1-level support unit, M2-level support unit, M3-level support unit, M4-level support unit, M5-level support unit, M6-level support unit, M7-level support unit, and M8-level support unit (or N1-level support unit, not specifically limited here), and the leaf nodes include basic units or departments such as N1-level support unit, N2-level support unit, N3-level support unit, N4-level support unit, N5-level support unit, N6-level support unit, N7-level support unit, and N8-level support unit.
[0076] (2.3) Each unit or department in the prefix tree node "value" list corresponds to a node in the tree diagram.
[0077] S103: Split and extract data containing structural information.
[0078] Using keywords, data types and data features of structural information, structural information is obtained from data containing structural information (such as text, which will not be described in detail below, and the storage or presentation method of data containing structural information is not limited here).
[0079] (3.1) Keywords include M1 security unit, M2 security unit, M3 security unit, M4 security unit, M5 security unit, M6 security unit, M7 security unit, M8 security unit, N2 security unit, N3 security unit, N4 security unit, N5 security unit, N6 security unit, N7 security unit, N8 security unit, etc.; data types include string type, numeric type, encoding type, etc.; data features include string length, numeric precision, encoding standard, etc.
[0080] (3.2) Using keywords, data types, and data features, regular expressions are constructed. Multiple rounds of recursive matching are used to perform multi-dimensional segmentation of data containing structural information (e.g., text) to obtain structural information fragments and assign corresponding labels to these fragments. For data where valid information fragments cannot be extracted through regular matching, manual annotation is used to extract structural information and assign corresponding labels.
[0081] (3.3) The structural information fragments are combined according to the levels of M1-level security unit, M2-level security unit, M3-level security unit, M4-level security unit, M5-level security unit, M6-level security unit, M7-level security unit, M8-level security unit, and N2-level security unit, N3-level security unit, N4-level security unit, N5-level security unit, N6-level security unit, N7-level security unit, and N8-level security unit to obtain the set of structural information tuples to be matched S orig Wherein, each or every piece of information containing structural data (or each structural information or each structural information fragment, which will not be described in detail below) corresponds to a set of structural information tuples to be matched.
[0082] S104: Structural information matching and sorting.
[0083] The set of structure information tuples to be matched S orig The tuples in are matched with the structure information prefix tree and tree graph (nodes) to obtain the candidate structure information set S prop , calculate the maximum matching degree between the structure information to be matched and the candidate structure information in the candidate structure information set.
[0084] The matching degree between structural information is calculated by using the structural information vector. That is, the structural information is first converted into a structural information vector, and then the matching degree between the structural information vectors is obtained by calculating the matching degree of each component (here refers to the information component) of the structural information vector (for example, the matching degree of each component in the tuple vector and each component in the corresponding candidate structural information vector). Thus, the matching degree between the structural information is obtained. Formula (1) is used to express it as follows (i.e., the formula for calculating the matching degree between structural information):
[0085]
[0086] in:
[0087] T i ∈S orig is the i-th structure information to be matched;
[0088] P j ∈S prop is the jth candidate structure information;
[0089] VT i is the structural information vector corresponding to the i-th structural information to be matched;
[0090] VP j is the structural information vector corresponding to the j-th candidate structural information;
[0091] VT i (k) is the kth information component of the i-th structure information vector to be matched;
[0092] VP j (k) is the kth information component of the jth candidate structure information vector;
[0093] Match(VT i (k),VP j (k)) is the matching degree between the kth information components. For string type and encoding type components, the matching degree is 1 when the two components are the same, otherwise the matching degree is 0; for numeric type components, the matching degree is 1 when the difference between the two components is less than a given threshold, otherwise the matching degree is 0.
[0094] An example of the components in the structure information vector is as follows:
[0095] (M1 security unit, M2 security unit, M3 security unit, M4 security unit, M5 security unit, M6 security unit, M7 security unit, M8 security unit, unit code, postal code, area code, longitude, latitude)
[0096] S105: Output the normalized structural information or update the structural information database.
[0097] According to the maximum matching degree obtained in S104, the normalized structural information is output or the structural information knowledge base is updated.
[0098] Assume that the maximum matching result obtained in S104 is:
[0099] (T maxi ,P maxj ,m maxi,maxj )=arg max{Match(T i ,P j):T i ∈S orig ,P j ∈S prop}
[0100] (5.1) If the maximum matching degree m maxi,maxj If it is greater than or equal to a given threshold (usually 2), then output P maxj As the normalized structural information, T maxi P maxj The necessary supplement is made to the structural information in the knowledge base, that is, the structural information knowledge base is supplemented with data containing structural information to achieve the expansion of the structural knowledge base information.
[0101] Optionally, when the candidate structure information P maxj The state is the historical state (ie, status: H), that is, the candidate structure information P maxj When it is historical structure information, P is obtained through the association relationship maxj The corresponding standard structure information currently used is the same as P maxj Together they serve as the normalized structural information obtained from the data containing structural information; when the candidate structural information P maxj The status is pending confirmation (status: U), that is, the candidate structure information P maxj If it is the structure information to be confirmed, then add P in the structure information knowledge base maxj credibility.
[0102] (5.2) If the maximum matching degree m maxi,maxj is equal to 1, and the candidate structure information set S prop There is only one candidate structure information P maxj , then output P maxj As the normalized structural information, T maxi P maxj The necessary supplement is made to the structural information in the knowledge base, that is, the structural information knowledge base is supplemented with data containing structural information to achieve the expansion of the structural knowledge base information.
[0103] (5.3) If the maximum matching degree m maxi,maxj Equal to 0 (that is, no T is found in the structure information prefix tree maxi Matching structure information), then take the set of structure information tuples to be matched S orig The structure information tuple to be matched with the most complete information is marked as pending confirmation status (status: U) and added to the structure information knowledge base.
[0104] Illustrative examples:
[0105] 1. Original data:
[0106] BBM1-level security unit CCM2-level security unit DDM3-level security unit EEM4-level security unit FFN5-level security unit
[0107] 2. Through multiple rounds of recursive regular matching, we get:
[0108] {BBM1-level support unit: M1-level support unit, CCM2-level support unit: M2-level support unit, DDM3-level support unit: M3-level support unit, EEM4-level support unit: M4-level support unit, FFN5-level support unit: department}
[0109] 3. Get the set of structure information tuples to be matched:
[0110] S orig
[0111] ={(BB M1-level support unit: M1-level support unit, CC M2-level support unit: M2-level support unit, DD M3-level support unit: M3-level support unit, EE M4-level support unit: M4-level support unit, N5-level support unit: department)}
[0112] 4. Structural information vector to be matched:
[0113] (BB M1-level support unit: M1-level support unit, CC M2-level support unit: M2-level support unit, DD M3-level support unit: M3-level support unit, EE M4-level support unit: M4-level support unit, N5-level support unit: department)
[0114] 5. The structure information vector to be matched is matched with the structure information prefix tree and tree graph nodes to obtain the maximum matching degree and normalized structure information:
[0115] BBM1-level security unit CCM2-level security unit DDM3-level security unit EEM4-level security unit FFN5-level security unit
[0116] This paper uses a structured information knowledge base to construct a prefix tree and dendrogram of structured information. It then splits and extracts data containing structured information using keywords and regular expressions, matches and sorts the nodes in the prefix tree and dendrogram, and thereby obtains standardized structured information or updates the structured information knowledge base. This method achieves matching and standardized processing of structured information, while also recording the optimization and adjustment process of structured information to ensure the standardization and consistency of structured information.
[0117] It should be noted that the collection, storage, use, processing, transmission, provision and disclosure of data and other information involved in the technical solution of the present disclosure are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0118] The following specific embodiments are used to describe the technical solution of the present disclosure in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0119] Structural information refers to the systematic arrangement of departments and positions within an organization, as well as their interrelationships. It defines the organizational hierarchy, division of responsibilities, reporting lines, and collaboration methods, and serves as the fundamental framework for organizational management and operations. The following uses structural information from an assurance scenario as an example to explain how to process structural information.
[0120] The present disclosure provides a method for processing structure information. Figure 3 As shown, the method includes the following steps S301 to S306:
[0121] S301: Construct a structure information knowledge base, which includes structure information and corresponding status. The structure information includes at least one of the following: currently used standard structure information, historical structure information, and to-be-confirmed structure information; the status corresponding to the standard structure information is the standard status, the status corresponding to the historical structure information is the historical status, and the to-be-confirmed structure information is the to-be-confirmed status.
[0122] S302: Construct a prefix tree and a tree graph corresponding to the structural information according to the structural information and the corresponding state.
[0123] S303: Determine a set of structure information tuples to be matched based on the structure information; one piece of structure information corresponds to one tuple in the set of structure information tuples to be matched.
[0124] S304: For each tuple in the set of structure information tuples to be matched, match the tuple with the prefix tree and the tree graph respectively to obtain a candidate structure information set corresponding to the tuple, and calculate the maximum matching degree between the tuple and each candidate structure information in the candidate structure information set.
[0125] S305: Determine target structure information corresponding to the structure information according to the target candidate structure information corresponding to the maximum matching degree, where the target structure information is used to represent normalized structure information.
[0126] S306: updating the structure information knowledge base according to the target structure information to standardize the structure information in the structure information knowledge base, and recording the association relationship between the historical structure information and the standard structure information in the structure information knowledge base.
[0127] In the embodiment of the present disclosure, a structural information knowledge base is constructed by collecting structural information, and based on the structural information to be standardized and the corresponding status in the structural information knowledge base, a prefix tree and a tree diagram corresponding to the structural information are constructed to match the structural information to be matched corresponding to the structural information. By calculating the matching degree and sorting, standardized structural information is obtained to update the structural information knowledge base, and then the standardized structural information is output, and the status corresponding to the structural information is adjusted and the association between the historical structural information and the standard structural information in the structural information knowledge base is recorded. This can provide a basis for the adjustment process from historical structural information to standard structural information, achieve unified standardization of all structural information in the structural information knowledge base, and record the optimization adjustment process of the structural information at the same time, thereby achieving efficient standardization processing while tracking and recording the adjustment process of the structural information, so that the structural information from each system source has standardization and consistency.
[0128] Based on the data containing structural information collected from various business information systems, a structural information knowledge base is constructed by extracting the structural information. Each piece of data containing structural information (or each piece of structural information in the constructed structural information knowledge base, or other structural data in the constructed structural information knowledge base other than standard structural data) can be treated as structural information to be normalized and subjected to normalization. The structural information in the structural information knowledge base can include three types of data: currently used standardized structural information (i.e., currently used standard structural information), historical structural information, and structural information to be confirmed.
[0129] One or a piece of structural information to be normalized corresponds to one piece of structural information to be matched or a tuple in a set of structural information tuples to be matched; matching and sorting operations are performed on each tuple in the set of structural information tuples to be matched, and normalized structural information corresponding to the target candidate structural information corresponding to each tuple or each principle is determined. Specifically, the following operations are performed on each tuple:
[0130] Calculate the matching degree between the current tuple and each candidate structure information in its corresponding candidate structure information set, and sort the matching degrees;
[0131] Determine the result corresponding to the maximum matching degree, where the result corresponding to the maximum matching degree includes the current tuple or the structural information to be normalized corresponding to the current tuple, the maximum matching degree value, and the target candidate structural information in the candidate structural information set corresponding to the current tuple; the target candidate structural information here refers to the candidate structural information that calculates the maximum matching degree.
[0132] Optionally, the constructing of a structure information knowledge base includes:
[0133] Obtaining structural information of each unit or part and corresponding structural optimization and adjustment log information; wherein the structural information includes at least one of the following: full name, abbreviation, alias or code name, hierarchy, unit code, parent unit code, postal code, area code, longitude, and latitude; the structural optimization and adjustment log information includes changes between old and new names and changes in recorded affiliation; the structural optimization and adjustment log information is used to query historical structural information and determine the status of the structural information;
[0134] A structural information knowledge base is generated based on the structural information of each unit or part and the corresponding structural optimization and adjustment log information; wherein the historical structural information and the standard structural information in the structural information knowledge base are associated with each other through foreign keys.
[0135] In the embodiment of the present disclosure, Figure 4 As shown, Figure 4 The entity-relationship (ER) diagram of the structural information knowledge base is shown. The structural information knowledge base contains structural information of the M1-level security unit, M2-level security unit, M3-level security unit, M4-level security unit, M5-level security unit, M6-level security unit, M7-level security unit, and M8-level security unit. In the knowledge base (here refers to the structural information knowledge base, which will not be repeated below), each unit or department can contain structural information of at least one of the following: full name (name), abbreviation (short_name), alias or code (alias_name), level (level), unit code (code), parent unit code (parent_code), postal code (zip_code), area code (city_code), longitude (longitude), and latitude (latitude). The full name (name) includes the complete name of the unit, such as BBM1-level support unit, CCM2-level support unit, DDM3-level support unit, EEM4-level support unit, and FFN5-level support unit. The short name (short_name) is the unit or department name agreed upon within the unit, such as the Political Work Department or the Weather Station. Unit and department information will be gradually updated as the system is used.
[0136] The knowledge base contains three types of structural information: the currently used standard structural information (S), the historical structural information (H) that has been abandoned due to reorganization and diversion, and the unconfirmed structural information (U) whose nature cannot be clearly determined for the time being.
[0137] To enable queries on historical structure information, the knowledge base includes log information on structure optimization and adjustment, including changes in old and new names, changes in directory relationships, etc. To track and record the adjustment process of structure information, historical structure information is linked to standard structure information through foreign keys.
[0138] Optionally, constructing a prefix tree and a tree graph corresponding to the structural information according to the structural information and the corresponding state includes:
[0139] Determine a key-value pair based on the structural information and the corresponding state; wherein the key in the key-value pair is the abbreviation of the unit or department, and the value in the key-value pair is a detailed list or set consisting of detailed information of the unit or department with the same abbreviation;
[0140] Based on the key-value pairs, construct the prefix tree, and record the corresponding status in the prefix tree;
[0141] generating a tree diagram according to the structural information;
[0142] The fusion information in the detailed information is a word segmentation array and corresponds to the nodes in the tree diagram.
[0143] In the embodiment of the present disclosure, at least one of the standard structure information (S), historical structure information (H) and to-be-confirmed structure information (U) in the information knowledge base is used to construct a prefix tree and a tree diagram of the structure information. In the prefix tree nodes and the tree diagram nodes (or the nodes in the tree diagram), status is used to distinguish different types of structure data (or structure information, which will not be repeated below). When the knowledge base is initially constructed, one unit or department corresponds to at least one node, which is a one-to-many relationship. The goal is that one unit or department corresponds to one node to achieve consistency of the structure information. However, one node may correspond to multiple units.
[0144] Specifically, the prefix tree node is composed of a key-value pair. The key is the abbreviation of the unit or department, and the value is a list (or set) of detailed information of the unit or department with the same abbreviation. To facilitate data comparison and verification, the integrated information (integrate_info) in the details (here refers to the detailed information) is a word array, corresponding to the nodes in the tree diagram, see Figure 5 As shown, Figure 5 A schematic diagram showing a prefix tree data structure and a tree graph data structure is shown, wherein Figure 5 The left side is a prefix tree data structure (including a root node and other nodes consisting of key-value pairs; the other nodes include: Operation and Maintenance / key: FFN5-level Support Unit / key: Weather Station; the Operation and Maintenance node includes key: A Operation and Maintenance Research Institute / key: B Operation and Maintenance Support Team); Figure 5The right side of the diagram shows a tree diagram data structure (BBM1 level security unit <- CCM2 level security unit <- DDM3 level security unit <- EEM4 level security unit <- FFN5 level security unit). For example, the prefix tree (or prefix tree node)
[0145]
[0146]
[0147] Corresponding to the nodes in the tree diagram:
[0148] BBM1-level support unit<-CCM2-level support unit<-DDM3-level support unit<-EEM4-level support unit<-FFN5-level support unit
[0149] Optionally, determining a set of structure information tuples to be matched based on the structure information includes:
[0150] Based on the keywords, the data type and data features of the structural information, multiple rounds of recursive regular matching are used to extract the hierarchical information and attribute information from the structural information, and label the hierarchical information and attribute information with corresponding tags. The data types include: string type, value type, and encoding type; the data features include: string length, value precision, and encoding standard; and the attribute information includes at least one of the following: unit code, postal code, area code, longitude, and latitude;
[0151] The extracted information at each level is combined according to the level, and non-compliant combinations are eliminated to obtain a set of to-be-matched structural information tuples corresponding to the structural information.
[0152] In the embodiment of the present disclosure, a regular expression is constructed based on keywords such as M1-level security unit, M2-level security unit, M3-level security unit, M4-level security unit, M5-level security unit, M6-level security unit, M7-level security unit, M8-level security unit, N2-level security unit, N3-level security unit, N4-level security unit, N5-level security unit, N6-level security unit, N7-level security unit, and N8-level security unit, as well as data types and data features such as unit code, postal code, area code, longitude, and latitude. Keywords, data types and data features, as well as regular expressions are stored in a configuration file and expanded and adjusted as needed. The configuration file is stored in a redis cache and persisted on the application system server.
[0153] In order to extract structural information as accurately as possible, multiple rounds of recursive regular matching are used to split the data containing structural information into multiple dimensions, extract the abbreviations of each level of information, as well as information such as unit code, postal code, area code, longitude, latitude, etc. (here, it can be referred to as attribute information), and assign corresponding labels (for example: M1-level security unit, M2-level security unit, M3-level security unit, M4-level security unit, N5-level security unit, etc.). For example, "BBM1-level security unit CCM2-level security unit DDM3-level security unit EEM4-level security unit FFN5-level security unit" is obtained through multiple rounds of recursive regular matching:
[0154] {BBM1-level support unit: M1-level support unit, CCM2-level support unit: M2-level support unit, DDM3-level support unit: M3-level support unit, EEM4-level support unit: M4-level support unit, FFN5-level support unit: department}
[0155] The extracted information is combined according to the levels of "M1-level security unit, M2-level security unit, M3-level security unit, M4-level security unit, M5-level security unit, M6-level security unit, M7-level security unit, M8-level security unit", and after eliminating unreasonable combinations, the set of structural information tuples to be matched S is obtained. orig For example, from "{BBM1-level security unit:M1-level security unit, CCM2-level security unit:M2-level security unit, DDM3-level security unit:M3-level security unit, EEM4-level security unit:M4-level security unit, FFN5-level security unit: department}", we can get a set of tuples (here refers to the set of structural information tuples to be matched):
[0156] {(BBM1-level support unit: M1-level support unit, CCM2-level support unit: M2-level support unit, DDM3-level support unit: M3-level support unit, EEM4-level support unit: M4-level support unit, FFN5-level support unit: department)}
[0157] An example of eliminating unreasonable combinations:
[0158] XXM8 level security unit YYM6 level security unit, XXM2 level security unit system application department
[0159] Among them, for data containing structural information that cannot be effectively extracted through regular matching, manual annotation is used to extract structural information and assign corresponding labels.
[0160] Optionally, matching the tuple with the prefix tree and the tree graph respectively to obtain a candidate structure information set corresponding to the tuple includes:
[0161] Searching the tuple in the prefix tree to obtain a prefix tree matching result, wherein the prefix tree matching result includes all prefix paths to be matched and matched nodes;
[0162] According to the prefix tree matching result, the node corresponding to the tree diagram is located, and the matching status of the tuple with the node corresponding to the tree diagram and its child nodes is checked to obtain a candidate structure information set corresponding to the tuple.
[0163] Here, one piece of structural information to be normalized (e.g., data containing structural information to be normalized) corresponds to one structural information tuple to be matched or a set of structural information tuples to be matched (the set may include one tuple). The structural information to be normalized here may be each structural information in the constructed structural information knowledge base, or may be at least one structural information among standard structural information, historical structural information, and structural information to be confirmed (e.g., structural information other than standard structural information).
[0164] The following takes each structural information in the constructed structural information knowledge base as an example (no further details will be given below). Based on each data containing structural information or each structural information, each corresponding structural information tuple to be matched can be determined, thereby forming a set of structural information tuples to be matched. Then, a set of candidate matching items is generated for each tuple to be matched (that is, the structural information tuple or tuple to be matched, no further details will be given below), for example: one tuple corresponds to at least one candidate structural information.
[0165] In the embodiment of the present disclosure, taking a tuple as an example, each level of information in the structure information tuple to be matched is matched with the structure information prefix tree and the tree diagram node (or tree diagram, which will not be repeated below), and the details of the matching results are saved in the candidate structure information set.
[0166] Specifically, the matching process includes:
[0167] Step a1. Tuple preprocessing: standardize each structural information tuple to be matched. This may include operations such as word segmentation, denoising, and normalization.
[0168] Step a2. Prefix tree matching: Search each tuple in the prefix tree to find all possible matching prefix paths; record the matched nodes and the matching degree.
[0169] Step a3. Tree Node Matching: Locate the corresponding node in the tree based on the prefix tree matching results. Check the matching between the tuple and the node and its children, taking into account relationships such as sibling nodes and parent-child nodes.
[0170] Step a4. Candidate set generation: Generate a set of candidate matching items for each tuple to be matched. Each candidate item contains a matching node and a matching score, for example, retaining the top-K most matching candidates.
[0171] The matching algorithm may include one of the following methods or a combination of them:
[0172] Path matching degree: the degree of consistency between the tuple path and the tree path;
[0173] Node attribute matching: similarity of attributes such as name, type, and level;
[0174] Structural similarity: similarity of positional relationships in the tree structure;
[0175] Fuzzy matching: handles situations where names are not exactly the same but have similar semantics.
[0176] The final output is a "candidate structure information set", in which each tuple to be matched corresponds to a set of possible matching results, which can be used for subsequent precise matching or manual verification.
[0177] Optionally, calculating the maximum matching degree between the tuple and each candidate structure information in the candidate structure information set includes:
[0178] Converting the tuple and each candidate structure information in the candidate structure information set into a structure information vector respectively;
[0179] Obtaining a matching degree between the tuple and each candidate structure information in the candidate structure information set by calculating a matching degree of each component in the structure information vector corresponding to each candidate structure information in the candidate structure information set;
[0180] The matching degrees between the tuple and each candidate structure information in the candidate structure information set are sorted to obtain the maximum matching degree.
[0181] In the embodiment of the present disclosure, the structure information to be matched T is calculated by the following steps: i ∈S orig (or refers to the i-th tuple in the set of structure information tuples to be matched, which will not be repeated below) and the candidate structure information P j ∈S prop The degree of match between:
[0182] Step b1: The structure information to be matched T i ∈S orig and candidate structure information P j ∈S prop Converted into the structure information vector VT to be matched i and candidate structure information vector VP jAn example of each component in the structure information vector is as follows:
[0183] (M1 security unit, M2 security unit, M3 security unit, M4 security unit, M5 security unit, M6 security unit, M7 security unit, M8 security unit, unit code, postal code, area code, longitude, latitude)
[0184] Step b2: Calculate the structure information vector VT to be matched i and candidate structure information vector VP j The matching degree m ij :=Match(VT i ,VP j )(Combined with the above formula (1)), where the matching degree is the sum of the matching degrees of each component in the structural information vector. For string type components such as M1-level security unit, M2-level security unit, M3-level security unit, M4-level security unit, M5-level security unit, M6-level security unit, M7-level security unit, M8-level security unit, unit code, postal code, area code, etc., when the two components are the same, the matching degree is 1, otherwise the matching degree is 0; for numerical type components such as longitude and latitude, when the difference between the two components is less than a given threshold, the matching degree is 1, otherwise the matching degree is 0.
[0185] Step b3: Sort the matching degree between the structure information to be matched and the candidate structure information, and output (T i ,P j ,m ij ,k ij ). Among them, T i ∈S orig is the i-th structure information to be matched, P j ∈S prop is the jth candidate structure information, mi j is the structure information to be matched T i and candidate structure information P j The matching degree between them, k ij Sort the results by matching degree.
[0186] Optionally, determining the target structure information corresponding to the structure information according to the target candidate structure information corresponding to the maximum matching degree includes:
[0187] If the maximum matching degree is greater than or equal to a preset threshold, determining the target structure information corresponding to the structure information according to the state of the target candidate structure information recorded in the structure information knowledge base, and performing necessary information filling on the target candidate structure information based on the to-be-matched structure information corresponding to the tuple, so as to expand the structure information knowledge base; wherein the state includes a standard state, a historical state, and a pending confirmation state;
[0188] If the maximum matching degree is equal to 1 and the candidate structure information set contains one candidate structure information, the target candidate structure information is used as the target structure information, and based on the to-be-matched structure information corresponding to the tuple, the target candidate structure information is filled with necessary information to expand the structure information knowledge base;
[0189] If the maximum matching degree is equal to 0, the most complete structural information corresponding to the tuple is obtained, and the most complete structural information is marked as a pending confirmation state to update the structural information knowledge base.
[0190] Optionally, determining the target structure information corresponding to the structure information according to the state of the target candidate structure information recorded in the structure information knowledge base includes:
[0191] When the target candidate structure information is recorded in the structure information knowledge base in a historical state (i.e., the target candidate structure information is historical structure information), the currently used standard structure information corresponding to the target candidate structure information is obtained through an association relationship, and the currently used standard structure information corresponding to the target candidate structure information and the target candidate structure information are used together as the target structure information to adjust the to-be-matched structure information corresponding to the maximum matching degree to the standard structure information, and at the same time adjust the credibility or status corresponding to the to-be-matched structure information and / or the target candidate structure information;
[0192] When the target candidate structure information is recorded in the structure information knowledge base in a standard state, the target candidate structure information is used as the target structure information;
[0193] When the state of the candidate structure information is to be confirmed, the credibility of the candidate structure information in the structure information knowledge base is increased;
[0194] The structure information knowledge base is updated using the target structure information.
[0195] In the embodiment of the present disclosure, the result with the largest matching degree is obtained (T maxi ,P maxj ,m maxi,maxj), that is, the result corresponding to the maximum matching degree, including the structure information to be matched, the target candidate structure information and the maximum matching degree.
[0196] Specifically, if the maximum matching degree m maxi,maxj Greater than or equal to a given threshold (usually 2, indicating T maxi With P maxj At least two data items are the same), then output P maxj As the normalized structural information obtained from the data containing structural information, at the same time, the structural information to be matched T maxi For candidate structure information P maxj Make necessary information filling to achieve the expansion of the structure knowledge base information. In particular, when the candidate structure information P maxj When the status is historical structure information (ie status: H), P is obtained through the association relationship. maxj The corresponding standard structure information currently used is the same as P maxj Together they serve as the normalized structural information obtained from the data containing structural information; when the candidate structural information P maxj When the status is pending confirmation (status: U), the structure information knowledge base P maxj credibility.
[0197] If the maximum matching degree m maxi,maxj Equal to 1 (ie T maxi With P maxj There is only one data item in common), and the candidate structure information set S prop There is only one candidate structure information P maxj , then output P maxj As the normalized structural information obtained from the data containing structural information, T maxi P maxj The necessary structural information is supplemented in the structure information knowledge base, that is, the structural information knowledge base is supplemented with data containing structural information.
[0198] If the maximum matching degree m maxi,maxj Equal to 0 (that is, no T is found in the structure information prefix tree maxi Matching structure information), then take the set of structure information tuples to be matched S orig The structure information tuple to be matched with the most complete information is marked as pending confirmation (status: U) and added to the structure information knowledge base.
[0199] From the above description, it can be seen that the present disclosure achieves the following technical effects: using structural information to construct prefix tree and tree graph data structures, using prefix tree and tree graph data structures to achieve matching and normalization processing of structural information, and at the same time recording the optimization and adjustment process of structural information to ensure the standardization and consistency of structural information.
[0200] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0201] The disclosed embodiment also provides a structural information processing system for implementing the above method embodiment, which includes three main modules: a data splitting and extraction module, an information matching and sorting module, and a knowledge base updating module.
[0202] The data splitting and extraction module is used to determine a set of structure information tuples to be matched based on the structure information to be normalized; wherein one or a piece of structure information to be normalized corresponds to a tuple in the set of structure information tuples to be matched;
[0203] an information matching and sorting module, configured to match each tuple in the to-be-matched structure information set with the prefix tree and the tree graph, respectively, to obtain a candidate structure information set corresponding to the tuple, and to calculate a matching degree between the tuple and each candidate structure information in the candidate structure information set; to determine a maximum matching degree between the tuple and each candidate structure information in the candidate structure information set by sorting the matching degrees; and to determine a target structure information corresponding to the tuple based on the target candidate structure information corresponding to the maximum matching degree; the target structure information is used to represent normalized structure information;
[0204] The knowledge base updating module is used to update the structure information knowledge base according to the target structure information, so as to standardize the structure information in the structure information knowledge base and record the association relationship between the historical structure information and the standard structure information in the structure information knowledge base.
[0205] Specifically, the data splitting and extraction module is primarily used to split and extract institutional information from data (text) containing structural information. This includes information at various levels, such as M1, M2, M3, M4, M5, M6, M7, and M8, as well as information such as unit codes, postal codes, area codes, longitude, and latitude. The structural information is then combined hierarchically to obtain the structural information to be matched (corresponding to a tuple). This module maintains keywords, the data types and characteristics of the structural information, and a list of regular expressions, which can be expanded and adjusted as needed.
[0206] Information matching and sorting module: mainly used to match the structural information to be matched with the structural information in the address prefix tree and tree diagram, calculate the matching degree and sort it to obtain standardized structural information.
[0207] Knowledge base update module: mainly used to maintain standard structure information, historical structure information and to-be-confirmed structure information, and to build and update address prefix trees and tree graphs based on the structure information knowledge base information.
[0208] The present disclosure utilizes structural information to construct prefix tree and tree graph data structures, utilizes prefix tree and tree graph data structures to achieve matching and normalization of structural information, and simultaneously records the optimization and adjustment process of structural information to ensure the normalization and consistency of structural information.
[0209] The present disclosure also provides a structural information processing device for implementing the above method embodiment. Figure 6 As shown, the structure information processing device 60 includes:
[0210] The first processing unit 601 is configured to construct a structure information knowledge base, wherein the structure information knowledge base includes structure information and corresponding states, wherein the structure information includes at least one of the following: currently used standard structure information, historical structure information, and to-be-confirmed structure information; the state corresponding to the standard structure information is the standard state, the state corresponding to the historical structure information is the historical state, and the to-be-confirmed structure information is the to-be-confirmed state;
[0211] The second processing unit 602 is configured to construct a prefix tree and a tree graph corresponding to the structural information according to the structural information and the corresponding state;
[0212] The third processing unit 603 is configured to determine a set of structure information tuples to be matched based on the structure information; one structure information corresponds to one tuple in the set of structure information tuples to be matched;
[0213] The fourth processing unit 604 is configured to match each tuple in the set of structure information tuples to be matched with the prefix tree and the dendrogram, respectively, to obtain a candidate structure information set corresponding to the tuple, and calculate a maximum matching degree between the tuple and each candidate structure information in the candidate structure information set;
[0214] A fifth processing unit 605 is configured to determine target structure information corresponding to the structure information based on the target candidate structure information corresponding to the maximum matching degree, where the target structure information is used to represent normalized structure information;
[0215] The sixth processing unit 606 is configured to update the structure information knowledge base according to the target structure information, so as to normalize the structure information in the structure information knowledge base, and record the association between the historical structure information and the standard structure information in the structure information knowledge base.
[0216] Optionally, when the third processing unit 603 determines the set of structure information tuples to be matched according to the structure information, the following steps are specifically performed:
[0217] Based on the keywords, the data type and data features of the structural information, multiple rounds of recursive regular matching are used to extract the hierarchical information and attribute information from the structural information, and label the hierarchical information and attribute information with corresponding tags. The data types include: string type, value type, and encoding type; the data features include: string length, value precision, and encoding standard; and the attribute information includes at least one of the following: unit code, postal code, area code, longitude, and latitude;
[0218] The extracted information at each level is combined according to the level, and non-compliant combinations are eliminated to obtain a set of to-be-matched structural information tuples corresponding to the structural information.
[0219] Optionally, when the fourth processing unit 604 matches the tuple with the prefix tree and the tree graph respectively to obtain a candidate structure information set corresponding to the tuple, the fourth processing unit 604 specifically includes:
[0220] Searching the tuple in the prefix tree to obtain a prefix tree matching result, wherein the prefix tree matching result includes all prefix paths to be matched and matched nodes;
[0221] According to the prefix tree matching result, the node corresponding to the tree diagram is located, and the matching status of the tuple with the node corresponding to the tree diagram and its child nodes is checked to obtain a candidate structure information set corresponding to the tuple.
[0222] Optionally, when calculating the maximum matching degree between the tuple and each candidate structure information in the candidate structure information set, the fourth processing unit 604 specifically includes:
[0223] Converting the tuple and each candidate structure information in the candidate structure information set into a structure information vector respectively;
[0224] Obtaining a matching degree between the tuple and each candidate structure information in the candidate structure information set by calculating a matching degree of each component in the structure information vector corresponding to each candidate structure information in the candidate structure information set;
[0225] The matching degrees between the tuple and each candidate structure information in the candidate structure information set are sorted to obtain the maximum matching degree.
[0226] Optionally, when the fifth processing unit 605 determines the target structure information corresponding to the structure information according to the target candidate structure information corresponding to the maximum matching degree, the following steps are specifically performed:
[0227] If the maximum matching degree is greater than or equal to a preset threshold, determining the target structure information corresponding to the structure information according to the state of the target candidate structure information recorded in the structure information knowledge base, and performing necessary information filling on the target candidate structure information based on the to-be-matched structure information corresponding to the tuple, so as to expand the structure information knowledge base; wherein the state includes a standard state, a historical state, and a pending confirmation state;
[0228] If the maximum matching degree is equal to 1 and the candidate structure information set contains one candidate structure information, the target candidate structure information is used as the target structure information, and based on the to-be-matched structure information corresponding to the tuple, the target candidate structure information is filled with necessary information to expand the structure information knowledge base;
[0229] If the maximum matching degree is equal to 0, the most complete structural information corresponding to the tuple is obtained, and the most complete structural information is marked as a pending confirmation state to update the structural information knowledge base.
[0230] Optionally, when the second processing unit 602 constructs a prefix tree and a tree graph corresponding to the structure information according to the structure information and the corresponding state, the following steps are specifically performed:
[0231] Determine a key-value pair based on the structural information and the corresponding state; wherein the key in the key-value pair is the abbreviation of the unit or department, and the value in the key-value pair is a detailed list or set consisting of detailed information of the unit or department with the same abbreviation;
[0232] Based on the key-value pairs, construct the prefix tree, and record the corresponding status in the prefix tree;
[0233] generating a tree diagram according to the structural information;
[0234] The fusion information in the detailed information is a word segmentation array and corresponds to the nodes in the tree diagram.
[0235] Optionally, when executing the construction of the structure information knowledge base, the first processing unit 601 specifically includes:
[0236] Obtaining structural information of each unit or part and corresponding structural optimization and adjustment log information; wherein the structural information includes at least one of the following: full name, abbreviation, alias or code name, hierarchy, unit code, parent unit code, postal code, area code, longitude, and latitude; the structural optimization and adjustment log information includes changes between old and new names and changes in recorded affiliation; the structural optimization and adjustment log information is used to query historical structural information and determine the status of the structural information;
[0237] A structural information knowledge base is generated based on the structural information of each unit or part and the corresponding structural optimization and adjustment log information; wherein the historical structural information and the standard structural information in the structural information knowledge base are associated with each other through foreign keys.
[0238] The specific manner in which each unit in the above device embodiment performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0239] The present disclosure also provides an electronic device, such as Figure 7 As shown, the electronic device includes one or more processors 71 and a memory 72. Figure 7 A processor 71 is taken as an example.
[0240] The controller may further include an input device 73 and an output device 74 .
[0241] The processor 71, the memory 72, the input device 73 and the output device 74 may be connected via a bus or other means. Figure 7 The bus connection is taken as an example.
[0242] The processor 71 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips. The general-purpose processor can be a microprocessor or any conventional processor.
[0243] Memory 72, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs, and modules, such as the program instructions / modules corresponding to the control method in the embodiments of the present disclosure. Processor 71 executes the non-transitory software programs, instructions, and modules stored in memory 72 to execute various server functional applications and data processing, thereby implementing the structural information processing method of the above-mentioned method embodiment.
[0244] The memory 72 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the processing device operated by the server, etc. In addition, the memory 72 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 72 may optionally include a memory remotely located relative to the processor 71, and these remote memories may be connected to a network connection device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0245] The input device 73 can receive input digital or character information and generate key signal input related to user settings and function control of the processing device of the server. The output device 74 can include a display device such as a display screen.
[0246] One or more modules are stored in the memory 72 and, when executed by one or more processors 71 , perform the above-described method.
[0247] An embodiment of the present disclosure further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute and implement the structural information processing method described above.
[0248] An embodiment of the present disclosure further provides a computer program product, including a computer program, which implements the structural information processing method described above when executed by a processor.
[0249] Those skilled in the art will appreciate that all or part of the processes in the above method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FM), a hard disk drive (HDD), or a solid-state drive (SSD). The storage medium can also include a combination of the above types of memory.
[0250] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for processing structural information, characterized in that: include: Constructing a structure information knowledge base, wherein the structure information knowledge base includes structure information and corresponding states, wherein the structure information includes at least one of the following: currently used standard structure information, historical structure information, and to-be-confirmed structure information; the state corresponding to the standard structure information is a standard state, the state corresponding to the historical structure information is a historical state, and the to-be-confirmed structure information is a to-be-confirmed state; Constructing a prefix tree and a tree graph corresponding to the structural information according to the structural information and the corresponding state; Determining a set of structure information tuples to be matched based on the structure information; One piece of structural information corresponds to one tuple in the set of structural information tuples to be matched; For each tuple in the set of structure information tuples to be matched, matching the tuple with the prefix tree and the tree graph respectively to obtain a set of candidate structure information corresponding to the tuple, and calculating the maximum matching degree between the tuple and each candidate structure information in the set of candidate structure information; Determining target structure information corresponding to the structure information based on the target candidate structure information corresponding to the maximum matching degree, wherein the target structure information is used to represent normalized structure information; The structure information knowledge base is updated according to the target structure information to standardize the structure information in the structure information knowledge base, and the association relationship between the historical structure information and the standard structure information in the structure information knowledge base is recorded.
2. The method according to claim 1, characterized in that The step of determining a set of structure information tuples to be matched based on the structure information includes: Based on the keywords, the data type and data features of the structural information, multiple rounds of recursive regular matching are used to extract the hierarchical information and attribute information from the structural information, and label the hierarchical information and attribute information with corresponding tags. The data types include: string type, value type, and encoding type; the data features include: string length, value precision, and encoding standard; and the attribute information includes at least one of the following: unit code, postal code, area code, longitude, and latitude; The extracted information at each level is combined according to the level, and non-compliant combinations are eliminated to obtain a set of to-be-matched structural information tuples corresponding to the structural information.
3. The method according to claim 1, characterized in that The step of matching the tuple with the prefix tree and the tree graph respectively to obtain a candidate structure information set corresponding to the tuple includes: Searching the tuple in the prefix tree to obtain a prefix tree matching result, wherein the prefix tree matching result includes all prefix paths to be matched and matched nodes; According to the prefix tree matching result, the node corresponding to the tree diagram is located, and the matching status of the tuple with the node corresponding to the tree diagram and its child nodes is checked to obtain a candidate structure information set corresponding to the tuple.
4. The method according to claim 1, wherein The calculating the maximum matching degree between the tuple and each candidate structure information in the candidate structure information set includes: Converting the tuple and each candidate structure information in the candidate structure information set into a structure information vector respectively; Obtaining a matching degree between the tuple and each candidate structure information in the candidate structure information set by calculating a matching degree of each component in the structure information vector corresponding to each candidate structure information in the candidate structure information set; The matching degrees between the tuple and each candidate structure information in the candidate structure information set are sorted to obtain the maximum matching degree.
5. The method according to claim 1, wherein The determining, based on the target candidate structure information corresponding to the maximum matching degree, target structure information corresponding to the structure information includes: If the maximum matching degree is greater than or equal to a preset threshold, determining the target structure information corresponding to the structure information according to the state of the target candidate structure information recorded in the structure information knowledge base, and performing necessary information filling on the target candidate structure information based on the to-be-matched structure information corresponding to the tuple, so as to expand the structure information knowledge base; wherein the state includes a standard state, a historical state, and a pending confirmation state; If the maximum matching degree is equal to 1 and the candidate structure information set contains one candidate structure information, the target candidate structure information is used as the target structure information, and based on the to-be-matched structure information corresponding to the tuple, the target candidate structure information is filled with necessary information to expand the structure information knowledge base; If the maximum matching degree is equal to 0, the most complete structural information corresponding to the tuple is obtained, and the most complete structural information is marked as a pending confirmation state to update the structural information knowledge base.
6. The method according to any one of claims 1 to 5, characterized in that The step of constructing a prefix tree and a tree graph corresponding to the structural information according to the structural information and the corresponding state includes: Determine a key-value pair based on the structural information and the corresponding state; wherein the key in the key-value pair is the abbreviation of the unit or department, and the value in the key-value pair is a detailed list or set consisting of detailed information of the unit or department with the same abbreviation; Based on the key-value pairs, construct the prefix tree, and record the corresponding status in the prefix tree; generating a tree diagram according to the structural information; The fusion information in the detailed information is a word segmentation array and corresponds to the nodes in the tree diagram.
7. The method according to any one of claims 1 to 5, characterized in that The constructing of the structure information knowledge base includes: Obtaining structural information of each unit or part and corresponding structural optimization and adjustment log information; wherein the structural information includes at least one of the following: full name, abbreviation, alias or code name, hierarchy, unit code, parent unit code, postal code, area code, longitude, and latitude; the structural optimization and adjustment log information includes changes between old and new names and changes in recorded affiliation; the structural optimization and adjustment log information is used to query historical structural information and determine the status of the structural information; A structural information knowledge base is generated based on the structural information of each unit or part and the corresponding structural optimization and adjustment log information; wherein the historical structural information and the standard structural information in the structural information knowledge base are associated with each other through foreign keys.
8. A structural information processing device, characterized in that: include: a first processing unit configured to construct a structure information knowledge base, the structure information knowledge base including structure information and corresponding states, the structure information including at least one of the following: currently used standard structure information, historical structure information, and to-be-confirmed structure information; the state corresponding to the standard structure information being a standard state, the state corresponding to the historical structure information being a historical state, and the to-be-confirmed structure information being a to-be-confirmed state; A second processing unit is configured to construct a prefix tree and a tree graph corresponding to the structural information according to the structural information and the corresponding state; A third processing unit is configured to determine a set of structure information tuples to be matched based on the structure information; One piece of structural information corresponds to one tuple in the set of structural information tuples to be matched; a fourth processing unit, configured to match each tuple in the set of structure information tuples to be matched with the prefix tree and the tree graph, respectively, to obtain a set of candidate structure information corresponding to the tuple, and to calculate a maximum degree of match between the tuple and each candidate structure information in the set of candidate structure information; A fifth processing unit is configured to determine target structure information corresponding to the structure information based on the target candidate structure information corresponding to the maximum matching degree, wherein the target structure information is used to represent the normalized structure information; The sixth processing unit is configured to update the structure information knowledge base according to the target structure information, so as to normalize the structure information in the structure information knowledge base, and record the association between the historical structure information and the standard structure information in the structure information knowledge base.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the structural information processing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor executes the structural information processing method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Data structure chart generation method and device, data structure chart updating method and device, electronic equipment and storage medium
CN113282795A
Fault-tolerant matching method and device for automobile information
CN114896381A
Information input method and device, equipment, storage medium and program product
CN115113740A
Geographic entity positioning analysis method and device and computer equipment
CN116431625A
Power grid fault risk disposal method and device based on knowledge graph
CN116644810A