Knowledge Graph-based Foreign Trade Data Classification and Management System

By conducting terminology modeling and graph path construction in the foreign trade data classification management system, the problems of insufficient semantic disassembly capability and lack of path expansion mechanism in the existing system are solved, more refined semantic analysis and more accurate term path mapping are achieved, and the accuracy and reliability of classification management are improved.

CN120045712BActive Publication Date: 2025-06-24庞学伟
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510504425.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-06-24
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing foreign trade data classification management system lacks fine word order modeling in the term processing process, resulting in insufficient semantic disassembly between terms, easy to produce structural mismatch during classification, and lacks a path expansion mechanism based on node level and morphemes logical structure, which cannot accurately reflect the deep relationship between semantic transmission paths between terms.

Method used

The term word order modeling module is used to extract morphemes and word order fragments, and the term original word order sequence template is constructed, and the morphemes are established from the top layer to the last layer, and the relationship between the connecting node numbers and adjacent nodes of all edges in the path is collected to build a graph path structure. At the same time, a path cross analysis and endpoint node frequency statistics mechanism are introduced to form a term cross path set, and semantic genera judgment is performed based on this.

Benefits of technology

It improves the word order restoration ability of term analysis, enhances the graph linkage between term structures, strengthens the capture ability of semantic overlap, improves the semantic focus effect of attribute labels, realizes a complete path closed loop from terminology to node classification, and enhances the systematization and traceability of classification management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045712B_ABST
    Figure CN120045712B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data classification management, specifically a foreign trade data classification management system based on a knowledge graph. The system includes a term word order modeling module, a graph path generation module, a path intersection analysis module, a semantic category determination module, and a classification structure output module. In the present invention, by performing morpheme-level word segmentation and word order extraction on foreign trade terms, an original word order template is constructed to enhance the restoration ability of semantic expression order. A directed path is established through morpheme-level label sorting to strengthen the hierarchical clarity of the term structure and the logical relationship between nodes. The semantic intersection area is extracted by using node intersection and end point frequency analysis to improve the accuracy of term association determination. The semantic attribution is selected according to the end point label frequency, and a node classification closed loop is formed by combining the label and the path mapping, connecting all links of term graph modeling, path construction, intersection analysis and attribution output, and fully improving the classification management process of foreign trade data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data classification management, and particularly to a foreign trade data classification management system based on a knowledge graph. Background Art

[0002] The technical field of data classification management includes relevant methods and systems for sorting, annotating, archiving, and managing various structured or unstructured data according to specific standards or rules. The core content lies in uniformly processing, classifying, and identifying a large amount of data with heterogeneous sources, diverse formats, and complex semantics to achieve efficient information organization and utilization. It covers multiple links such as data preprocessing, attribute extraction, label assignment, semantic normalization, and category mapping, and constructs a complete classification management system by combining means such as natural language processing, database management, and knowledge modeling, and is applicable to various information application scenarios such as enterprises, scientific research, and commerce.

[0003] Among them, a foreign trade data classification management system based on a knowledge graph refers to an information processing system that classifies and manages import and export trade data based on constructing and utilizing an entity relationship network. The data items targeted mainly include foreign trade data fields such as commodity attributes, trade flows, trading entities, and customs clearance times. An ontology modeling method is used to construct a special knowledge graph for the foreign trade field, and multi-dimensional semantic classification and data linking are carried out based on entity matching and relationship reasoning technologies. By matching existing knowledge nodes with field values in the data through a rule engine, the category attribution and the determination of the upper and lower position relationships are completed in the knowledge graph structure to achieve structured classification and management of a large amount of heterogeneous foreign trade data.

[0004] In the existing foreign trade data classification management process, in the process of term processing, there is a lack of fine-grained modeling at the morpheme level, and only indexing and annotation are carried out around the surface features of terms, without deeply exploring the word order features of the internal structure of terms, resulting in insufficient semantic disassembling ability between terms and prone to structural mismatch phenomena during classification. In terms of path construction, static label mapping is mostly used, lacking a path expansion mechanism based on node levels and morpheme logical structures, and unable to accurately reflect the deep relationship of the semantic transmission path between terms. In semantic judgment, cross-path analysis and end-point label weight mechanisms are not introduced, ignoring the aggregated semantic value carried by intersection nodes between terms, and the attribution judgment is easily affected by insufficient label coverage, resulting in low stability of classification results. In the term classification stage, there is a lack of a node organization system based on path mapping, and no closed-loop attribution chain from terms to nodes is formed, resulting in problems such as coarse classification granularity and scattered term label mapping. For example, in multiple trade entries with similar term fragments, the existing methods are difficult to distinguish different semantic contexts through structure, resulting in multi-label overlap and classification cross-interference, affecting the reliability and expansion efficiency of the overall classification system. Summary of the Invention

[0005] The object of the present invention is to solve the deficiencies existing in the prior art, and a foreign trade data classification and management system based on a knowledge graph is proposed.

[0006] To achieve the above object, the present invention adopts the following technical solutions: The foreign trade data classification and management system based on a knowledge graph includes:

[0007] The term word order modeling module obtains the term text in the foreign trade data, extracts the morpheme index by word segmentation, judges the first position and frequency of the keyword, maps and generates the morpheme arrangement sequence, and constructs the original term word order sequence template;

[0008] The graph path generation module sorts the morphemes based on the original term word order sequence template, establishes a directed path of morphemes from the top layer to the bottom layer according to the sorting result, collects the connection node numbers and adjacent node relationships of all edges in the path, and constructs the graph path structure;

[0009] The path intersection analysis module extracts the term path node sequence according to the graph path structure, compares the intersection nodes and counts the frequency of the end nodes, screens the cross paths, and obtains the term cross path set;

[0010] The semantic category determination module collects the semantic labels of the end nodes based on the end nodes in the term cross path set, sorts them according to the appearance frequency and matches the path end labels, and obtains the term semantic attribution label group;

[0011] The classification structure output module counts the graph classification nodes to which each label belongs according to the term semantic attribution label group, divides the term path under the corresponding nodes, establishes the node and term path classification attribution relationship structure, and generates the foreign trade data classification structure table.

[0012] As a further solution of the present invention, the original term word order sequence template includes a morpheme index arrangement structure, a keyword word order segment set, and a keyword frequency weight model. The graph path structure includes a morpheme level mapping relationship, a directed path node chain, and a node connection relationship set. The term cross path set includes a path intersection node set, an end node appearance frequency distribution, and a cross node screening result. The term semantic attribution label group includes a semantic label frequency ranking, a path end semantic label mapping, and a belonging category label matching result. The foreign trade data classification structure table includes a classification node identifier, a term path grouping result, and a classification attribution relationship mapping result.

[0013] As a further solution of the present invention, the term word order modeling module includes:

[0014] The morpheme extraction sub-module obtains the technical term text in the foreign trade data, performs morpheme-level word segmentation on the technical term text, extracts the morpheme sets of each term, records the index positions of each morpheme in the term, compares the relationship between the first occurrence position of the keyword in the morpheme arrangement list and the number of morphemes, classifies by term, and obtains the keyword morpheme index distribution result;

[0015] The original order segment construction sub-module extracts the morpheme segments corresponding to the keywords in the original text of the term according to the keyword morpheme index distribution result, intercepts the morphemes based on the index interval of the keyword in the term, constructs a morpheme segment set according to the positions of the intercepted morphemes of each keyword, and recombines it in combination with the term to which the morpheme belongs, and obtains the keyword original order segment set;

[0016] The word order template generation sub-module counts the occurrence frequency values of all keywords according to the keyword original order segment set, performs position rearrangement processing on the morpheme segment set based on the original order arrangement of the keywords in the technical term text, splices the original order segments of multiple keywords in the same term according to the first occurrence position, and uses the formula:

[0017] ;

[0018] Calculate the total weight of the original word order of the term , and perform classification and integration in combination with the word order weight result corresponding to each term to obtain the original word order sequence template of the term, where represents the first index position of the th item of the keyword, represents the th occurrence frequency value of the keyword, represents the length of the morpheme segment corresponding to the keyword, represents the number of keywords in the term, represents the total number of morphemes in the term, represents the total number of all keywords in the term.

[0019] As a further solution of the present invention, the graph path generation module includes:

[0020] The hierarchical sorting sub-module, based on the original word order sequence template of the term, combines the hierarchical labels of each morpheme node of the term, compares and sorts all morpheme nodes according to the priority value of their hierarchical labels, rearranges the positions of the morphemes of the term in the order from the top layer to the bottom layer, establishes a rearrangement sequence index table, and obtains the morpheme sorting index value;

[0021] The path construction sub-module obtains the adjacent node set in the morpheme rearrangement sequence according to the morpheme sorting index value, numbers each pair of adjacent nodes separately and records their connection directions, and uses the formula:

[0022] ;

[0023] Calculate the offset strength value of the morpheme directed path , combine all node connection relationships in the path, integrate the structural edge information, and generate morpheme directed path graph data, where represents the index number of the th morpheme in the sorted list, represents the index number of the th morpheme in the sorted list, represents the connection span value of the th morpheme, represents the connection span value of the th morpheme, represents the absolute value of the hierarchical label difference between the th and the th morphemes, is the number of edges in the path;

[0024] The node structure extraction sub-module collects the node numbers and adjacent node pair relationships in all connected edges according to the morpheme directed path graph data, constructs a node mapping table based on the adjacent structure relationships, stores the upstream and downstream relationship types and connection directions between each morpheme, and obtains the graph path structure.

[0025] As a further solution of the present invention, the path intersection analysis module includes:

[0026] The path extraction sub-module collects the node sequences in any two term paths based on the graph path structure, sequentially extracts the node number information under each path and establishes a term node mapping set, marks the term identification and path length parameters to which each path belongs, and obtains the term path node number value;

[0027] The intersection comparison sub-module calls the node number sequences of any two term paths according to the term path node number value, performs an intersection comparison operation on the node sets of the two paths, extracts all the end node numbers in the intersection, respectively counts the number of times such nodes appear in different paths, and uses the formula:

[0028] ;

[0029] Calculate the end point deviation value of the intersection path , compare it item by item with the path intersection judgment reference value, screen out the path pair combinations with the deviation value less than or equal to the reference value, and establish a set of intersection numbers of paths that meet the conditions, where represents the number of times the th intersection node appears in path A, represents the number of times the same node appears in path B, Indicates the total number of occurrences of the node in all path sets, which is the total number of intersection nodes;

[0030] Based on the set of path intersection numbers that meet the conditions, the path screening sub-module queries the original term path identifiers according to the path combinations corresponding to the numbers, integrates the term identifiers and path intersection node information, and establishes a term path relationship linked list to generate a set of term intersection paths.

[0031] As a further solution of the present invention, the semantic category determination module includes:

[0032] The semantic label collection sub-module collects the set of semantic labels to which each end node belongs based on the end nodes in the set of term intersection paths, performs index mapping between the term paths and their end labels, and generates a set of path end semantic labels;

[0033] The label frequency statistics sub-module performs a repetition count operation on all semantic labels based on the set of path end semantic labels, records the number of times each semantic label appears in the term path set, and performs sorting processing from high to low according to the number of occurrences to obtain a sorted semantic label sequence;

[0034] The category label determination sub-module performs a matching judgment on the set of labels corresponding to the end nodes in the term paths according to the sorted semantic label sequence, screens the label item that is the most forward in the sorted sequence in each path as the semantic category to which the path belongs, and integrates the attribution labels of all term paths to obtain a set of term semantic attribution labels.

[0035] As a further solution of the present invention, the classification structure output module includes:

[0036] The node extraction sub-module collects the graph classification nodes corresponding to each label according to the set of term semantic attribution labels, records the associated path numbers and the number of corresponding term path sets in each classification node, determines the matching index between the semantic label and the graph node, and obtains the label attribution node number value;

[0037] The path classification sub-module divides the corresponding term paths into each graph classification node based on the label attribution node number value with the semantic label as the classification basis, establishes a two-way corresponding structure between the term path number and the node number, extracts the list of path numbers attributed to each node, and obtains the node path attribution quantity value;

[0038] The structure generation sub-module performs structure mapping integration on the graph classification nodes and the subordinate term path numbers according to the node path attribution quantity value, outputs the classification node index, the corresponding semantic label and the total number of paths, determines the attribution of the node and the term path, and generates a foreign trade data classification structure table.

[0039] Compared with the prior art, the advantages and positive effects of the present invention are:

[0040] In the present invention, morpheme-level segmentation and word order fragment extraction are performed on term texts in foreign trade data, and the original word order template is constructed by comparing the position relationship of keywords in the morpheme sequence, so as to realize fine-grained modeling of the semantic structure of the term, improve the word order restoration ability of term parsing, set priorities according to the morpheme node level labels, construct directed paths from the top layer to the bottom layer for the terms, establish multi-level semantic paths through the connection relationship between node numbers and edges, enhance the graph linkage between term structures, introduce node intersection calculation and terminal frequency statistics mechanism, form semantic intersection areas between paths, strengthen the ability to capture semantic overlap, based on The attribution labels of the terminal nodes in the cross-path are sorted and screened, and the label with the highest frequency is used as the basis for path classification to improve the semantic focus effect of the attribution label. The term path mapping and structural aggregation of the graph classification nodes are carried out in combination with semantic labels to achieve a complete path closed loop from terminology to node classification, enhance the systematization and traceability of classification management, and the overall process is based on morpheme structure, semantic path, cross-relationship, and attribution label to build a dynamic semantic recognition and classification mapping system, improve the accuracy of term organization, structural coordination between paths, and the semantic carrying capacity of graph nodes, and greatly improve the classification management process of foreign trade data. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a system flow chart of the present invention;

[0042] Figure 2 A flow chart of the term order modeling module of the present invention;

[0043] Figure 3 This is a flow chart of the graph path generation module of the present invention;

[0044] Figure 4 This is a flow chart of the path intersection analysis module of the present invention;

[0045] Figure 5 This is a flow chart of the semantic category determination module of the present invention;

[0046] Figure 6 This is a flow chart of the classification structure output module of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0048] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as length, width, up, down, front, back, left, right, vertical, horizontal, top, bottom, inside, outside, etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "plurality" is two or more unless otherwise specifically defined.

[0049] Please refer to Figure 1 , the foreign trade data classification and management system based on the knowledge graph includes:

[0050] The term word order modeling module obtains the term text in the foreign trade data, performs morpheme-level word segmentation on the term text and extracts the morpheme index, compares the first occurrence position of the keyword in the morpheme arrangement list with the number of morphemes, extracts the original word order fragment corresponding to the keyword, and combines the occurrence frequency values of all keywords to construct a term original word order sequence template;

[0051] The graph path generation module, based on the term original word order sequence template, combines the hierarchical labels of each morpheme node of the term, sorts the morphemes according to the priority of the hierarchical labels, establishes a directed path of morphemes from the top layer to the bottom layer according to the sorting result, collects the connection node numbers and adjacent node relationships of all edges in the path, and constructs a graph path structure;

[0052] The path intersection analysis module, according to the graph path structure, extracts the node sequences in any two term paths, performs an intersection comparison of the path node sets on the two node sequences, counts the number of times the end node appears in different paths in the intersection path, and compares it with the path intersection judgment reference value to screen the term paths that meet the cross-node conditions to obtain a term cross-path set;

[0053] The semantic category determination module, based on the end nodes appearing in the term cross-path set, collects the semantic label set to which the end nodes belong, performs a sorting operation on the repetition times of the semantic label set, combines the semantic labels of the end nodes corresponding to each term path, and performs a matching judgment with the sorted label set, and selects the semantic label with the highest occurrence frequency value as the belonging category label of the target term path to generate a term semantic belonging label group;

[0054] The classification structure output module, according to the term semantic belonging label group, counts the graph classification nodes to which each label belongs, extracts the total number of the belonging path list and the belonging term path set in the classification nodes, divides the term paths into the corresponding nodes according to the semantic belonging label as the classification basis, establishes the classification belonging relationship structure between the nodes and the term paths, and generates a foreign trade data classification structure table.

[0055] The original word order sequence template for terms includes a morpheme index arrangement structure, a keyword word order fragment set, and a keyword frequency weight model. The atlas path structure includes a morpheme level mapping relationship, a directed path node chain, and a node connection relationship set. The term cross-path set includes a path intersection node set, an end node occurrence frequency distribution, and a cross-node screening result. The term semantic attribution label group includes a semantic label frequency ranking, a path end semantic label mapping, and an attribution generic label matching result. The foreign trade data classification structure table includes a classification node identifier, a term path grouping result, and a classification attribution relationship mapping result.

[0056] Please refer to Figure 2 , the term word order modeling module includes:

[0057] The morpheme extraction sub-module obtains the term text in the foreign trade data, performs morpheme-level word segmentation on the term text, extracts the morpheme set of each term, records the index position of each morpheme in the term, compares the relationship between the first occurrence position of the keyword in the morpheme arrangement list and the number of morphemes, classifies by term, and obtains the keyword morpheme index distribution result;

[0058] To obtain the term text in the foreign trade data, it is necessary to first collect the actual foreign trade corpus from specific industry scenarios. For example, extract term texts such as export plastic packaging film and stainless steel pipe fittings from the customs declaration list. For each term text, apply the separation logic to segment it into morphemes by words. For example, segment export plastic packaging film into export, plastic, packaging, and film, then number each morpheme, and its index order is 1 to 4 in sequence. Then identify the keywords from the morpheme set. For example, set the keywords as packaging and film from the database, judge that their first occurrence positions in the morpheme list are 3 and 4 respectively, and at the same time count the total number of morphemes in this term as 4. Compare the keyword index with the total number of morphemes. If the index position is less than the median of the total number of morphemes, it is recorded as a preposition morpheme, otherwise it is recorded as a postposition morpheme. Use the preposition or postposition identifier to participate in the subsequent word order judgment. In this process, it is necessary to set a keyword index interval judgment threshold for dividing the importance of morphemes. For example, set the preposition position as morpheme index ≤ 2, and the postposition as index ≥ 3. Classify and mark each keyword through this threshold, and then summarize the distribution status of all keywords in each term, and establish a morpheme index mapping array [3, 4] of keywords, bind the result to the term identifier, and construct a term index distribution matrix. If multiple terms are processed in this process, the following example data can be formed. For example, export stainless steel pipe fittings → [1, 2, 3, 4], the keywords steel and pipe have indexes 3 and 4, the total number of morphemes in this term is 4, and the number of keywords is 2. Thus, the keyword morpheme index distribution value is [(steel, 3), (pipe, 4)], and it is classified into the keyword vector corresponding to the term, forming the result keyword morpheme index distribution result.

[0059] The original order segment construction sub-module extracts the morpheme segments corresponding to the keywords in the original text of the terms according to the keyword morpheme index distribution results, intercepts the morphemes based on the index intervals of the keywords in the terms, constructs a set of morpheme segments according to the positions of the intercepted morphemes of each keyword, and reorganizes them in combination with the terms to which the morphemes belong to obtain a set of original order segments of the keywords;

[0060] Based on the keyword morpheme index distribution results, read the terms and their keyword index information one by one, intercept the morphemes according to the keyword positions in the original text of the terms, divide the original order segments by the intercepted intervals. For example, in the term "export plastic packaging film", the position of the keyword "packaging" is the 3rd morpheme, so the morphemes from the 3rd to the 4th are extracted to form the segment "packaging film". Further, all the original order segments of the keywords in the term are extracted, and multiple segments are marked with term attribution and combined into a set of morpheme segments. For example, for the term "export plastic packaging film", the keyword segment set {packaging film} is generated. If a term has multiple keywords, such as "export stainless steel pipe fittings", and its keywords "steel" and "pipe" are located at the 3rd and 4th morphemes respectively, then the intercepted segment is "steel pipe fittings", and the morpheme set is {steel pipe fittings}. For terms with multiple keywords, the morpheme slices are combined by sorting the keyword indexes from small to large. If the slice interval is less than or equal to 1, the segments are merged; if the interval is greater than 1, the respective segments are retained. Finally, a set of original order segments corresponding to multiple keywords in each term is constructed, and this set is organized by term to form a term original order structure library, which is the resulting set of original order segments of the keywords.

[0061] The word order template generation sub-module counts the occurrence frequency values of all keywords according to the set of original order segments of the keywords, performs a position rearrangement process on the set of morpheme segments based on the original order of the keywords in the term text, splices the original order segments of multiple keywords in the same term according to the first occurrence position, and uses the formula:

[0062] ;

[0063] Calculate the total weight value of the original word order of the term , and perform classification and integration in combination with the word order weight results corresponding to each term to obtain the original word order sequence template of the term. Among them, represents the first index position of the th item of the keyword, represents the th item of the keyword's occurrence frequency value, represents the length of the morpheme segment corresponding to this keyword, represents the number of keywords in this term, represents the total number of morphemes in the term, represents the total number of all keywords in the term;

[0064] According to the set of original-order fragments of keywords, read the keyword set and their occurrence frequency values in each term. For the term "export stainless steel pipe fittings", set the occurrence frequencies of the keywords "steel" and "pipe" to 2 and 3 respectively, read their original-order indexes as 3 and 4, and the corresponding morpheme fragment lengths as 1 and 2. The total number of keywords in this term is 2, and the total number of morphemes is 4. Substitute into the formula:

[0065] ;

[0066] Substitute the parameter values into the formula:

[0067] Group 1: , , , , , ;

[0068] Group 2: , , , , , ;

[0069] The calculation process is as follows:

[0070] The value of the first operation: ;

[0071] The value of the second operation: ;

[0072] Sum and average:

[0073] ;

[0074] This result indicates that the total weight of the original word order of the term "export stainless steel pipe fittings" is 2.68. Thus, the word order weights of multiple terms are classified and integrated to form the word order template corresponding to each term. For example, the weight of "export plastic packaging film" is 1.75, and the weight of "stainless steel pipe fittings" is 2.68. After sorting or clustering them, the original word order sequence template of the result terms can be generated.

[0075] Table 1 Summary of term word order weights:

[0076] ;

[0077] Table 1 lists the total weights of different terms in the reconstruction of the original word order, which can be used for word order template classification and structural induction.

[0078] The formula focuses on the reconstruction of the original word order of keywords in the term, reflecting the comprehensive influence of factors such as the position characteristics, frequency weights, and structural complexity of keywords in the word order. Among them, the first occurrence position of keywords Indicates its structural priority in the morpheme sequence. The smaller the value, the more forward it is, directly reflecting its word order dominant position; the frequency of keyword appearance is processed by the square root to be , which is used to reflect the marginal impact of its non-linear growth and prevent high-frequency words from causing asymmetric interference to the word order structure; the length of the keyword morpheme fragment and the total number of keywords The product of represents the keyword structure density in the term, and then divided by the total number of morphemes in the term to form , which is used as a normalized measure of the term structure complexity. The overall weighted expression is , which is used to balance the keyword position advantage, frequency influence and structure density contribution, combines the absolute value operation to unify the word order directionality, and finally sums and averages the influence values of all keywords in the term , to obtain a unified measurement result of the word order weight of the term. Therefore, this formula realizes a word order modeling mechanism of position priority + frequency reconciliation + structure compensation in terms of structure, making the word order template more representative and comparable as a whole.

[0079] Please refer to Figure 3 , the graph path generation module includes:

[0080] Based on the original word order sequence template of the term, the hierarchical sorting sub-module combines the hierarchical labels of each morpheme node of the term, compares and sorts all morpheme nodes according to the priority value of their hierarchical labels, rearranges the positions of the term morphemes in the order from the top layer to the bottom layer, establishes a rearrangement sequence index table, and obtains the morpheme sorting index value;

[0081] Based on the original word order sequence template of terms, first extract all morpheme nodes in the term and their corresponding hierarchical labels. For example, the term "remote sensing device module" can be disassembled into morphemes "remote", "sensing", "device", and "module", and the corresponding hierarchical labels are 3, 2, 1, and 0 respectively, indicating that "module" is the top layer, and so on downwards. Subsequently, sort all morpheme nodes according to the priority of their hierarchical labels. Here, it is necessary to judge the label values. The lower the priority, the closer it is to the top layer. For example, the morpheme "module" with a hierarchical label of 0 is sorted first, followed by "device", "sensing", and "remote". After sorting, the morpheme order obtained is "module - device - sensing - remote". In the specific sorting process, by comparing whether the label value of the adjacent morpheme is less than the previous node, if so, move the current morpheme forward. After sorting, generate a rearrangement sequence index value for each morpheme, numbered sequentially from 0, such as 0 for "module", 1 for "device", etc. This index is used for subsequent path establishment. If the number of terms increases, for example, a new term "intelligent analysis terminal device" is added, where the morpheme "terminal" is at hierarchical level 0, "device" is at 1, "analysis" is at 2, and "intelligent" is at 3, then the sorting result is "terminal - device - analysis - intelligent". This sorting rule ensures the top-down structural order, and the indexes are from 0 to 3, as shown in Table 2.

[0082] Table 2 Morpheme Hierarchical Sorting Index Table:

[0083] ;

[0084] As shown in Table 2, the hierarchical label directly affects the sorting order. The comparison operation in the sorting can be attributed to performing greater than or less than judgments on the label values of the morpheme nodes, that is, if the label value of the previous item is greater than that of the latter item, then swap their orders. Through multiple rounds of judgments, the overall sorting process is completed, and this process finally generates the morpheme sorting index value.

[0085] The path construction sub-module obtains the set of adjacent nodes in the morpheme rearrangement sequence according to the morpheme sorting index value, numbers each pair of adjacent nodes separately and records their connection directions, using the formula:

[0086] ;

[0087] Calculate the offset intensity value of the morpheme directed path , and combine the connection relationships of all nodes in the path, integrate the structure edge information, and generate the morpheme directed path graph data. Among them, represents the index number of the th morpheme in the sorted list, represents the index number of the th morpheme in the sorted list, represents the connection span value of the th morpheme, represents the connection span value of the th morpheme, represents the The absolute value of the hierarchical label difference between the th morpheme, is the number of edges in the path;

[0088] According to the morpheme sorting index value, the sorted morpheme sequence is concatenated bit by bit to establish adjacent node connection pairs. For example, in the previous example, the module-device-sensing-remote corresponding node indexes are 0-1-2-3, and the adjacent connection pairs are (0, 1), (1, 2), (2, 3). Each pair of connection pairs represents a segment of the edge of the morpheme path. In path construction, the node numbers of each connected edge need to be collected and the structural offset is calculated, that is, the latter node number minus the former node number. For example, the offset of (1, 2) is 1. In addition, the hierarchical difference needs to be combined. Assuming that the hierarchy of the node module is 0 and the device is 1, the hierarchical difference is 1. To quantify the structural offset strength of each segment of the edge in the path, a connection span parameter is introduced. Let the span between module and device be 2, and the span between device and sensing be 3, and so on. The following formula is used for structural strength calculation:

[0089] Let , , , then:

[0090] The first segment: ;

[0091] The second segment: ;

[0092] The third segment: ;

[0093] The sum is:

[0094] ;

[0095] This result shows that the morpheme path offset strength value is 6.0. This value represents the directional connection strength of the term under the sorting structure and can be used as a measurement basis for the path organization logic, for subsequent unified structure layout and link direction in the atlas path construction, and finally generate the morpheme directed path atlas data.

[0096] The formula integrates the structural and hierarchical information among multiple parameters to characterize the connection strength characteristics of the path between morpheme nodes. Among them, represents the position difference between adjacent nodes in the sorting, reflecting the sequence span between the connected nodes, is the sum of the morpheme spans connecting two nodes, measuring the information load density in the path. The product of the two and then divided by 2 represents the weighted average contribution of this path edge to the connection strength. Further considering the influence of the hierarchical label difference, It represents the absolute difference between two nodes at the semantic level. Using the square root operation reflects the non-linear weakening trend that as the level difference increases, the influence of the boundary change on the path strength tends to flatten out. Subtracting the above two parts and taking the absolute value can suppress the directional influence of the level fluctuation on the path contribution, thereby forming a unified positive structure strength value. Finally, perform a summation operation on all path edges and integrate them into the overall structure offset strength value, which reflects the organizational rationality and hierarchical distribution characteristics of the path as a whole.

[0097] The node structure extraction sub-module collects the node numbers and adjacent node pair relationships in all connection edges according to the morpheme directed path graph data, constructs a node mapping table based on the adjacent structure relationship, stores the upstream and downstream relationship types and connection directions between each morpheme, and obtains the graph path structure.

[0098] According to the morpheme directed path graph data, collect the starting node number and ending node number of each connection edge, and on this basis, establish a connection mapping table to distinguish the upstream and downstream node position relationships. If the smaller numbered one is the starting point, the direction is defined as positive; if the larger numbered one is the starting point, the direction is negative. For example, the node number pair (1, 2) is positive, and (3, 2) is negative. Complete the node direction classification by judging the number relationship of each connection pair one by one. Taking module-device as an example, module is 0 and device is 1, and the defined direction is 0→1. Similarly, sensing-remote is 2→3, forming an edge direction mapping table. Then classify and store all edges and their corresponding directions. By number normalization, a unified format for node connections can be established among multiple terms. This mapping structure is further used to extract the structural paths between nodes in the morpheme graph and record them as a structure table, and finally obtain the graph path structure.

[0099] Please refer to Figure 4 , the path intersection analysis module includes:

[0100] Based on the graph path structure, the path extraction sub-module collects the node sequences in any two-term paths, sequentially extracts the node number information under each path and establishes a term node mapping set, marks the term identification and path length parameters to which each path belongs, and obtains the term path node number value.

[0101] When extracting the node sequences of any two term paths in the atlas path structure, first, the set of node paths in the term atlas data should be obtained. On this basis, unique number identifiers need to be assigned to the two target terms. For example, the path node numbers of term A are {101, 103, 107, 110}, and the path node numbers of term B are {102, 103, 108, 110}. When extracting the node number sequence, it should be arranged according to the order in which the nodes appear in the path. If there are duplicate nodes in the node sequence, the first occurrence position should be retained according to the path order and the remaining positions should be discarded. Furthermore, an index mapping table of the two-term path nodes is established. The table should record the node number, the position index of the node in the path, the corresponding term identifier, and the path length value. For example, the path length of term A is 4, and the path length of term B is also 4. At this time, the node numbers 103 and 110 exist in both terms, indicating an intersection part, and the record is as follows

[0102] Table 3: Node mapping table:

[0103] ;

[0104] As shown in Table 3, the table records the position index of each node in different paths and the intersection judgment situation. Among them, whether it is an intersection is a boolean field. If the node number exists in both term paths, it is set to yes, otherwise it is set to no. This judgment does not require a fuzzy interval. Finally, the node numbers with an intersection of yes in the table are extracted and converted into an intersection path node set, and the term path node number value can be obtained.

[0105] The intersection comparison sub-module calls the node number sequences of any two term paths according to the term path node number value, performs an intersection comparison operation on the node sets of the two paths, extracts all the end node numbers in the intersection, and respectively counts the number of times such nodes appear in different paths. Using the formula:

[0106] ;

[0107] Calculate the end point deviation value of the intersection path , and compare it item by item with the path intersection judgment reference value. Filter out the path pair combinations with the deviation value less than or equal to the reference value, and establish a set of intersection number sets of paths that meet the conditions. Among them, represents the number of times the th intersection node appears in path A, represents the number of times the same node appears in path B, represents the total number of times the node appears in all path sets, is the total number of intersection nodes;

[0108] When performing an intersection comparison process on a set of nodes according to the term path node number values, the set of intersection nodes should be selected first. In this embodiment, the intersection nodes are {103, 110}. Count the number of occurrences of these two nodes in Path A and Path B respectively. Since the node numbers in each path are unique, each of the two nodes appears once in Path A and Path B. Then calculate the total number of occurrences in all path sets. For example, if there are 5 term paths in the system, where 103 appears in 3 of them and 110 appears in 4 of them, the intersection node statistical parameters are as follows:

[0109] Table 4: Statistical parameter record:

[0110] ;

[0111] Substitute the data in Table 4 into the formula for calculation to get:

[0112] ;

[0113] According to the calculation result, the intersection path end point deviation value is 0, and the path crossing judgment reference value is set to 0.5. The basis is the median value of the discrete tolerance interval obtained after normalizing the distribution range of the intersection node deviation amounts in the whole path set. Specifically, the distribution range of the deviation values of all path pairs in the intersection in the statistical sample set is [0, 1.3], and the deviation values of more than 70% of the path pairs are concentrated in the interval [0, 0.5]. Therefore, 0.5 is set as the conservative boundary point for screening structurally consistent path combinations. This reference value tends to decrease as the path node density of the whole graph increases. The increase in density leads to an increase in the coincidence degree between nodes, an increase in the cross-overlap phenomenon, and a corresponding reduction in the tolerance space for deviations. Therefore, it is statistically reasonable and graphically structure-corresponding to select a deviation value less than 0.5 as the judgment threshold. Set the path crossing judgment reference value to 0.5, and the numerical deviation result is lower than the reference value. Therefore, this path pair meets the cross-node condition and enters the subsequent screening process. This result shows that there is no obvious difference in the occurrence rules of the intersection nodes between the two paths, and they have structural consistency, and a set of path intersection numbers that meet the conditions can be established.

[0114] The formula is based on the comprehensive measurement of the difference in the occurrence of the path node intersection in the two paths. Its core lies in measuring whether the occurrence distribution of the intersection nodes in Path A and Path B is consistent. First, is used to strengthen the influence of the nodes that frequently appear in Path A on the deviation value. Among them, is squared, which can amplify the weight difference in the case of repeated node occurrences, and then improve the sensitivity to the interference of frequent nodes; this difference value is then divided by , the purpose is to smooth the nodes frequently appearing in all paths, reduce the dominance of high-frequency nodes over the overall deviation in the form of taking the square root, and add 1 to prevent the denominator from being zero when a node appears only once; after calculating the results of each node, the absolute value is used to ensure that the deviation value is always non-negative, avoiding the positive and negative cancellation from affecting the overall judgment. Finally, the deviation values of all intersection nodes are summed and then divided by the number of intersection nodes , to complete the normalization process to obtain the average level of the deviation distribution between nodes, so as to measure the structural consistency between two paths. The overall formula takes into account the node frequency difference, the global frequency distribution and the influence amplitude of a single node, enabling the deviation value to reasonably reflect the similarities and differences in the cross-structure between paths.

[0115] The path screening sub-module queries the original term path identifiers according to the set of intersection numbers of the paths that meet the conditions and the corresponding path combinations according to the numbers, integrates the term identifiers and the path intersection node information, and establishes a term path relationship linked list to generate a set of term cross paths;

[0116] After obtaining the set of intersection numbers of the paths that meet the conditions, it is necessary to reverse query the original term identifiers and their path structure information corresponding to the intersection node set. For example, the paths corresponding to the numbers 103 and 110 are term A and term B. It is necessary to retrieve their path structures from the term path index table and generate a set of term pairs. Then, an association linked list between the term pairs and their intersection node sets is constructed. The fields in the linked list include the number of term A, the number of term B, the intersection node set, the path length ratio, the number of intersection nodes, etc. If the path lengths of term A and term B are both 4 and the number of intersection nodes is 2, then the path length ratio is 1 and the proportion of intersection nodes is 0.5. Furthermore, a cross-path tuple {term A, term B, {103, 110}, 1, 0.5} is established, and the combinations with path length ratios deviating from the interval [0.5, 2] are screened out, and the qualified path tuples are stored in the final structure table to obtain the set of term cross paths.

[0117] Please refer to Figure 5 , the semantic category determination module includes:

[0118] The semantic label collection sub-module collects the set of semantic labels to which each end node belongs based on the end nodes appearing in the set of term cross paths, performs index mapping between the term path and its end label, and generates a set of path end semantic labels;

[0119] Based on the end nodes that appear in the set of term cross paths, after collecting each end node of the term path, it is necessary to first extract the semantic tags marked for this node in the semantic database, and construct the corresponding relationship between each node and the semantic tags through the number mapping method. Assume that there are paths P1, P2, and P3, and their end nodes are N7, N12, and N15 respectively. Then, according to the structure of the term graph, it is retrieved that the label corresponding to N7 is the supply chain, N12 is the payment method, and N15 is the transaction process, forming an initial label set; during the further extraction process, the situation where multiple paths have duplicate end nodes needs to be considered. For example, the end nodes of paths P4 and P5 are both N7, and their labels are also unified as the supply chain. The processing of such nodes is achieved by one-to-one matching of the path number and the node number, and the mapping relationship is stored in the form of an array to ensure the efficiency of retrieval and subsequent operations; if the number of term paths is 20 and the end nodes involve 10 semantic tags, the following structure is recorded in a two-dimensional table, as shown in Table 5.

[0120] Table 5 Distribution Table of Semantic Tags of End Nodes:

[0121] ;

[0122] As shown in Table 5, for each term path, the labels corresponding to its end nodes have been clearly listed through number indexing. During the collection process, the paths are sorted by path number, and the labels corresponding to the node numbers are queried one by one to ensure the integrity and consistency of label acquisition, and then a semantic label group for the path end is established.

[0123] Based on the semantic label group of the path end, the label frequency statistics sub-module performs the operation of counting the number of repetitions for all semantic tags, records the number of times each semantic tag appears in the set of term paths, and performs sorting from high to low according to the number of appearances to obtain a sorted semantic tag sequence;

[0124] After obtaining the semantic label group of the path end, it is necessary to perform frequency statistics on all semantic tags. The operation process is achieved by recording the cumulative number of times the label appears in the term path. A frequency array F is used to count the number of times each label number appears. For example, in the path set, the supply chain label appears 5 times, the payment method label appears 3 times, and the transaction process label appears 2 times. Then the initial frequency array is , and the corresponding index structure of the label needs to be clarified before sorting. Assume that the label numbers are t1, t2, and t3, corresponding to the above labels; when sorting, it is arranged in descending order of frequency, that is ; In frequency calculation, if the end nodes of some paths are empty, the path needs to be skipped and not included in the frequency statistics; to improve the statistical accuracy, semantic tags can also be filtered according to the path coverage rate. This filtering criterion is set based on the representativeness of semantic tags. Specifically, the total number of term paths P is set to the total number of valid paths entered during system construction. If the occurrence times of a semantic tag are less than 10% of P, it is considered a low-frequency and unrepresentative tag. This 10% threshold is taken from the lower bound of the optimized interval of semantic classification accuracy. This lower bound setting is derived from the sample stability analysis. When the tag coverage rate is less than one-tenth of the total sample, its semantic stability shows a linear decay trend. This value is adjusted according to the change in the total number of term paths. For example, when the total number of paths is 20, this threshold is set to 2. If the total number of paths increases to 50, the corresponding threshold is 5. This value is not fixed as a constant but dynamically adjusted according to the number of input samples of term paths to ensure the sensitivity and dynamic adaptability of the filtering criterion to the label distribution within the system. If this benchmark is set, tags with a frequency of 1 such as t4: import and export restrictions will be excluded; finally, the sorting result is an ordered pair array of tag numbers and their frequencies. , that is, generate a sorted semantic tag sequence.

[0125] The generic label determination sub-module performs a matching judgment on the label set corresponding to the end nodes in the term path according to the sorted semantic tag sequence, screens the label item that is the most forward in position in the sorted sequence for each path as the semantic category to which the path belongs, and integrates the attribution labels of all term paths to obtain the term semantic attribution label group;

[0126] According to the sorted semantic tag sequence, perform a matching judgment on the label set corresponding to the end nodes in the term path in turn. The matching process uses a label comparison mechanism. For the semantic label of the end node of each path, retrieve its position number in the sorted label sequence. If the label is in a forward position in the sorted sequence (such as the top 3), it is marked as a successful match; for example, the end label of path P1 is the supply chain, ranked 1st in the sorted list, and the match is successful. The end label of path P6 is enterprise credit, ranked 7th, so it is not selected; the successfully matched label is regarded as the path attribution label. If multiple labels are simultaneously matched at the path end, the final generic category is determined according to the label with the most forward ranking; finally, integrate the corresponding relationships between all term paths and their attribution labels to construct term path generic label pairs, such as , and generate the term semantic attribution label group one by one.

[0127] Please refer to Figure 6 , the classification structure output module includes:

[0128] The node extraction sub-module collects the graph classification nodes corresponding to each label according to the term semantic attribution label group, records the associated path numbers in each classification node and the number of their corresponding term path sets, determines the matching index between the semantic label and the graph node, and obtains the label attribution node number value;

[0129] According to the term semantic attribution label group, successively collect the corresponding relationship information of each label in the term path, traverse the term paths attributed to each semantic label, extract the graph classification node numbers bound in the graph structure, and for the paths with the same semantic label, count the distribution of all their graph node numbers. For example, if paths 1, 2, and 3 are bound under label a and belong to nodes 101, 101, and 102 respectively, it can be known that the frequency of node 101 corresponding to this label is 2, and the frequency of node 102 is 1. Therefore, a one-to-many matching table structure between the semantic label and the node number needs to be established for subsequent path attribution judgment. During the counting process, it is necessary to judge whether the node number value is unique. If it is unique, the label belongs to this node. If it is not unique, the frequency needs to be counted and the maximum value is taken as the basis for attribution judgment. In the above process, the actually selected semantic labels are label a, label b, label c, and label d respectively. Their corresponding paths come from the previous path aggregation stage, and the node numbers are collected from the structure values constructed in the path mapping table. After cleaning their mapping information, a binding index dictionary structure from the semantic label to the graph node is constructed, and the node number distribution quantity information under this index structure is extracted, so as to obtain the label attribution node number value.

[0130] The path classification sub-module divides the corresponding term paths into each graph classification node according to the semantic label based on the label attribution node number value, establishes a two-way corresponding structure between the term path number and the node number, extracts the list of path numbers attributed to each node, and obtains the node path attribution quantity value;

[0131] Based on the label attribution node number value, the term paths are corresponding to the corresponding classification node numbers according to their semantic labels. During this process, all term path numbers need to be traversed, and their semantic labels need to be queried, and then mapped to their bound node numbers, and the path numbers are classified into the node path set. During this process, the path numbers attributed to the same node number need to be recorded in the same structure. For example, when the node number is 101, its attributed path set may include paths 1, 4, 8, etc. When counting, the path set quantity is accumulated. After performing such operations on all node numbers, the number of paths attributed to each node can be obtained. For example, there are 15 paths attributed to node 101, 9 paths attributed to node 102, 12 paths attributed to node 103, and 7 paths attributed to node 104. A corresponding structure of number and quantity is established for these statistical results, so as to obtain the node path attribution quantity value.

[0132] Based on the node path attribution quantity value, the structure generation sub-module integrates the classification nodes of the graph and the subordinate term path numbers through structure mapping, outputs the classification node index, corresponding semantic labels and the total number of paths, determines the attribution of nodes and term paths, and generates a foreign trade data classification structure table;

[0133] Based on the node path attribution quantity value, for each node in the graph, collect its node number, bound semantic label and the total number of path numbers it belongs to, combine the three pieces of information into one to construct a structure record. During the construction process, extract the number of its path sets from the previous structure for each node number and mark it in the structure table, then perform a reverse lookup on the semantic label bound to the node and record it together. Use the semantic label, node number, and path quantity as the three field values of the output structure. Finally, establish the structure of the classification structure table. Each row in the structure table records the classification structure information of a node, and the fields are shown in Table 1. The mapping relationship between the semantic label, path and node is clearly marked, thereby generating a foreign trade data classification structure table.

[0134] Table 6 Foreign Trade Semantic Label Classification Statistical Table:

[0135] ;

[0136] As shown in Table 6, there are obvious distribution differences in the semantic labels among the attributed nodes. This table serves as the data basis source for the subsequent classification structure table.

[0137] The above is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. The foreign trade data classification management system based on knowledge graph is characterized by: The system comprises: The terminology word order modeling module obtains terminology texts in foreign trade data, extracts morpheme indexes through word segmentation, determines the first position and frequency of keywords, maps and generates morpheme arrangement sequences, and constructs terminology original word order sequence templates; The graph path generation module sorts the morphemes based on the original word order sequence template of the term, establishes a directed path of morphemes from the top layer to the bottom layer according to the sorting result, collects the connection node numbers and adjacent node relationships of all edges in the path, and constructs a graph path structure; The path intersection analysis module extracts the term path node sequence according to the graph path structure, compares the intersection nodes and counts the frequency of the terminal nodes, screens the intersection paths, and obtains the term intersection path set; The semantic category determination module collects semantic labels of the terminal nodes based on the terminal nodes in the term intersection path set, sorts them by frequency of occurrence, and matches the path terminal labels to obtain a term semantic attribution label group; The classification structure output module counts the graph classification nodes to which each label belongs according to the semantic attribution label group of the term, divides the term path into corresponding nodes, establishes the classification attribution relationship structure between the node and the term path, and generates a foreign trade data classification structure table; The term order modeling module includes: The morpheme extraction submodule obtains the term text in the foreign trade data, performs morpheme-level segmentation on the term text, extracts the morpheme set of each term, records the index position of each morpheme in the term, compares the relationship between the first appearance position of the keyword in the morpheme arrangement list and the number of morphemes, classifies by term, and obtains the keyword morpheme index distribution result; The original sequence fragment construction submodule extracts the morpheme fragments corresponding to the keyword in the original text of the term according to the keyword morpheme index distribution result, intercepts the morphemes based on the index interval of the keyword in the term, constructs a morpheme fragment set according to the position of each keyword intercepted morpheme, and reorganizes the morphemes in combination with the terms to which the morphemes belong, to obtain a keyword original sequence fragment set; The word order template generation submodule counts the occurrence frequency values ​​of all keywords according to the original sequence fragment set of keywords, and performs position rearrangement processing on the morpheme fragment set based on the original order of keywords in the term text, and sequentially splices the original sequence fragments of multiple keywords in the same term according to the first appearance position, using the formula: ; Calculate the sum of the original word order weights of the terms , combined with the word order weight results corresponding to each term, classify and integrate them to obtain the original word order sequence template of the term, where, Representative keywords the first index position of the item, Representative keywords The frequency value of the item, Indicates the length of the morpheme segment corresponding to the keyword. Indicates the number of keywords in the term. Indicates the total number of morphemes in a term, Indicates the number of all keywords in the term.

2. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The term original word order sequence template includes a morpheme index arrangement structure, a keyword word order fragment set, and a keyword frequency weight model; the graph path structure includes a morpheme hierarchical mapping relationship, a directed path node chain, and a node connection relationship set; the term cross path set includes a path intersection node set, a terminal node occurrence frequency distribution, and a cross node screening result; the term semantic attribution label group includes a semantic label frequency ranking, a path terminal semantic label mapping, and an attribution category label matching result; the foreign trade data classification structure table includes a classification node identifier, a term path grouping result, and a classification attribution relationship mapping result.

3. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The graph path generation module includes: The hierarchical sorting submodule compares and sorts all morpheme nodes according to their hierarchical label priority values ​​based on the original word order sequence template of the term and the hierarchical labels of each morpheme node of the term, rearranges the positions of the term morphemes in order from the top layer to the bottom layer, establishes a rearranged sequence index table, and obtains a morpheme sorting index value; The path construction submodule obtains the adjacent node set in the morpheme rearrangement sequence according to the morpheme sorting index value, and numbers each pair of adjacent nodes to record their connection direction, using the formula: ; Calculate the morpheme directional path offset strength value , combining all node connection relationships in the path, integrating structural edge information, and generating morpheme directed path graph data, where Indicates The index number of the morpheme in the sorted list, Indicates The index number of the morpheme in the sorted list, Indicates The connection span value of morphemes, Indicates The connection span value of morphemes, Indicates The first The absolute value of the level label difference between morphemes, is the number of edges in the path; The node structure extraction submodule collects the node numbers and adjacent node pair relationships in all connecting edges according to the morpheme directed path graph data, builds a node mapping table based on the adjacent structural relationship, stores the upstream and downstream relationship types and connection directions between each morpheme, and obtains the graph path structure.

4. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The path intersection analysis module includes: The path extraction submodule collects the node sequences in any two term paths based on the graph path structure, extracts the node number information under each path in turn and establishes a term node mapping set, marks the term identifier and path length parameter of each path, and obtains the term path node number value; The intersection comparison submodule calls the node number sequences of any two term paths according to the node number values ​​of the term paths, performs an intersection comparison operation on the node sets of the two paths, extracts all the terminal node numbers in the intersection, and counts the number of times such nodes appear in different paths, using the formula: ; Calculate the intersection path endpoint deviation value , compare them one by one with the path intersection judgment benchmark value, select the path pair combinations with deviation values ​​less than or equal to the benchmark value, and establish a set of path intersection numbers that meet the conditions, where, Indicates The number of times the intersection node appears in path A, represents the number of times the same node appears in path B, Indicates the total number of occurrences of the node in all path sets. is the total number of intersection nodes; The path screening submodule queries the original term path identifier according to the path combination corresponding to the number according to the set of path intersection numbers that meet the conditions, integrates the term identifier and the path intersection node information, establishes a term path relationship link list, and generates a term intersection path set.

5. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The semantic category determination module comprises: The semantic label collection submodule collects the semantic label set to which each terminal node belongs based on the terminal nodes in the term intersection path set, performs index mapping between the term path and its terminal label, and generates a path terminal semantic label group; The tag frequency statistics submodule performs a repetition count operation on all semantic tags based on the path endpoint semantic tag group, records the number of times each semantic tag appears in the term path set, and sorts them from high to low according to the number of occurrences to obtain a sorted semantic tag sequence; The category label determination submodule performs matching judgment on the label set corresponding to the terminal node in the term path according to the sorted semantic label sequence, selects the label item in each path that is at the front of the sorted sequence as the path corresponding to the semantic category, integrates the attribution labels of all term paths, and obtains the term semantic attribution label group.

6. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The classification structure output module includes: The node extraction submodule collects the graph classification nodes corresponding to each label according to the term semantic attribution label group, records the number of associated path numbers and their corresponding term path sets in each classification node, determines the matching index between the semantic label and the graph node, and obtains the label attribution node number value; The path classification submodule divides the corresponding term path into each graph classification node based on the node number value of the label and the semantic label as the classification basis, establishes a bidirectional correspondence structure between the term path number and the node number, extracts the path number list of each node, and obtains the node path attribution quantity value; The structure generation submodule integrates the graph classification nodes and the subordinate term path numbers according to the node path attribution quantity value, outputs the classification node index, the corresponding semantic label and the total number of paths, determines the attribution of the nodes and term paths, and generates a foreign trade data classification structure table.

Citation Information

Patent Citations

  • Automatic translation method for terms in large texts

    CN103488628A

  • Knowledge base system, inter-word meaning relation determination method in the same system and computer program

    JP2005157823A