Foreign trade data classification management system based on knowledge graph
By modeling term order and building map paths in the foreign trade data classification management system, the problems of insufficient semantic disassembly capability and lack of path expansion mechanism in the existing system are solved, and higher classification management accuracy and system reliability are achieved.
Patent Information
- Application Number
- CN202510504425.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The existing foreign trade data classification management system lacks fine word order modeling in the term processing process, resulting in insufficient semantic disassembly between terms, easy to produce structural mismatch during classification, and lacks a path expansion mechanism based on node level and morphemes logical structure, which cannot accurately reflect the deep relationship between semantic transmission paths between terms.
The term text in foreign trade data is obtained through the term term order modeling module, morphemes and word order fragments are extracted, and the term original word order sequence template is constructed. Then, based on the graph path generation module, a morphemes directed path from the top layer to the last layer is established to build a graph path structure. Through the path cross analysis module, the term path node sequence is extracted, the intersection nodes are compared, and the end node frequency is counted, the intersection paths are filtered to obtain the term cross path set. Finally, based on the semantic class attribute determination module, the semantic labels of the end point node are collected, sorted by occurrence frequency, matched the path end point labels, and obtained the term semantic attribute label group.
Through fine-grained term order modeling and multi-level map path construction, the term order restoration and semantic linkage of term analysis are improved, the graph linkage and semantic overlap capture capabilities between term structures are enhanced, and the classification management accuracy of foreign trade data and the reliability of system are improved.
Smart Images

Figure CN120045712A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data classification management, and particularly to a foreign trade data classification management system based on a knowledge graph. Background Art
[0002] The technical field of data classification management includes relevant methods and systems for sorting, annotating, archiving, and managing various structured or unstructured data according to specific standards or rules. The core content lies in uniformly processing, classifying, and identifying a large amount of data with heterogeneous sources, diverse formats, and complex semantics to achieve efficient information organization and utilization. It covers multiple links such as data preprocessing, attribute extraction, label assignment, semantic normalization, and category mapping, and constructs a complete classification management system by combining means such as natural language processing, database management, and knowledge modeling, which is applicable to various information application scenarios such as enterprises, scientific research, and commerce.
[0003] Among them, a foreign trade data classification management system based on a knowledge graph refers to an information processing system that classifies and manages import and export trade data based on constructing and utilizing an entity relationship network. The data items targeted mainly include foreign trade data fields such as commodity attributes, trade flows, trading entities, and customs clearance times. An ontology modeling method is used to construct a special knowledge graph for the foreign trade field, and multi-dimensional semantic classification and data linking are performed based on entity matching and relationship reasoning technologies. By matching existing knowledge nodes with field values in the data through a rule engine, the category attribution and the determination of the upper and lower position relationships are completed in the knowledge graph structure to achieve structured classification and management of a large amount of heterogeneous foreign trade data.
[0004] In the existing foreign trade data classification management process, in the term processing process, there is a lack of fine-grained modeling at the morpheme level, and only indexing and annotation are carried out around the surface features of the terms, without deeply exploring the word order features of the internal structure of the terms, resulting in insufficient semantic disassembling ability between terms and prone to structural mismatch phenomena during classification. In terms of path construction, static label mapping is mostly used, lacking a path expansion mechanism based on node levels and morpheme logical structures, and unable to accurately reflect the deep relationship of the semantic transmission path between terms. In semantic judgment, cross-path analysis and end-point label weight mechanisms are not introduced, ignoring the aggregated semantic value carried by the intersection nodes between terms, and the attribution judgment is easily affected by insufficient label coverage, resulting in low stability of the classification results. In the term classification stage, there is a lack of a node organization system based on path mapping, and a closed-loop attribution chain from terms to nodes is not formed, resulting in problems such as coarse classification granularity and scattered term label mapping. For example, in multiple trade entries with similar term fragments, the existing methods are difficult to distinguish different semantic contexts through structure, resulting in multi-label overlap and classification cross-interference, affecting the reliability and expansion efficiency of the overall classification system. Summary of the Invention
[0005] The object of the present invention is to solve the disadvantages existing in the prior art, and a foreign trade data classification and management system based on a knowledge graph is proposed.
[0006] To achieve the above object, the present invention adopts the following technical solutions: The foreign trade data classification and management system based on a knowledge graph includes: The term word order modeling module obtains the term text in the foreign trade data, performs word segmentation to extract the morpheme index, judges the first position and frequency of the keyword, maps and generates the morpheme arrangement sequence, and constructs the original term word order sequence template; The graph path generation module sorts the morphemes based on the original term word order sequence template, establishes a directed path of morphemes from the top layer to the bottom layer according to the sorting result, collects the connection node numbers and adjacent node relationships of all edges in the path, and constructs the graph path structure; The path intersection analysis module extracts the term path node sequence according to the graph path structure, compares the intersection nodes and counts the frequency of the end nodes, filters the intersection paths, and obtains the term intersection path set; The semantic category determination module collects the semantic labels of the end nodes based on the end nodes in the term intersection path set, sorts them according to the appearance frequency and matches the path end labels, and obtains the term semantic attribution label group; The classification structure output module counts the graph classification nodes to which each label belongs according to the term semantic attribution label group, divides the term path under the corresponding nodes, establishes the node and term path classification attribution relationship structure, and generates the foreign trade data classification structure table.
[0007] As a further solution of the present invention, the original term word order sequence template includes a morpheme index arrangement structure, a keyword word order segment set, and a keyword frequency weight model. The graph path structure includes a morpheme level mapping relationship, a directed path node chain, and a node connection relationship set. The term intersection path set includes a path intersection node set, an end node appearance frequency distribution, and a cross node screening result. The term semantic attribution label group includes a semantic label frequency ranking, a path end semantic label mapping, and a belonging category label matching result. The foreign trade data classification structure table includes a classification node identifier, a term path grouping result, and a classification attribution relationship mapping result.
[0008] As a further solution of the present invention, the term word order modeling module includes: The morpheme extraction sub-module obtains the term text in the foreign trade data, performs morpheme-level word segmentation on the term text, extracts the morpheme set of each term, records the index position of each morpheme in the term, compares the relationship between the first appearance position of the keyword in the morpheme arrangement list and the number of morphemes, classifies by term, and obtains the keyword morpheme index distribution result; The original sequence fragment construction sub-module extracts the morpheme fragments corresponding to the keywords in the original text of the terms according to the keyword morpheme index distribution result, intercepts the morphemes based on the index interval of the keywords in the terms, constructs a morpheme fragment set according to the positions of the intercepted morphemes of each keyword, and recombines them in combination with the terms to which the morphemes belong, so as to obtain the original sequence fragment set of the keywords; The word order template generation sub-module counts the occurrence frequency values of all keywords according to the original sequence fragment set of the keywords, performs position rearrangement processing on the morpheme fragment set based on the original order of the keywords in the term text, splices the original sequence fragments of multiple keywords in the same term according to the first occurrence position, and uses the formula: ; Calculate the total weight of the original word order of the term , and perform classification and integration in combination with the word order weight results corresponding to each term to obtain the original word order sequence template of the term. Among them, represents the first index position of the th item of the keyword, represents the occurrence frequency value of the th item of the keyword, represents the length of the morpheme fragment corresponding to the keyword, represents the number of keywords in the term, represents the total number of morphemes in the term, represents the total number of all keywords in the term.
[0009] As a further solution of the present invention, the map path generation module includes: The hierarchical sorting sub-module compares and sorts all morpheme nodes according to the priority value of their hierarchical labels based on the original word order sequence template of the term, rearranges the positions of the morphemes of the term in the order from the top layer to the bottom layer, establishes a rearrangement sequence index table, and obtains the morpheme sorting index value; The path construction sub-module obtains the adjacent node set in the morpheme rearrangement sequence according to the morpheme sorting index value, numbers each pair of adjacent nodes and records their connection directions respectively, and uses the formula: ; Calculate the offset intensity value of the directed path of the morpheme , and integrate the structure edge information in combination with the connection relationships of all nodes in the path to generate the directed path map data of the morpheme. Among them, represents the index number of the th morpheme in the sorted list, represents the index number of the th morpheme in the sorted list, represents the The connection span value of a morpheme, indicating the connection span value of the indicating the absolute value of the hierarchical label difference between the th morpheme and the is the number of edges in the path; The node structure extraction sub-module collects the node numbers and adjacent node pair relationships in all connection edges according to the morpheme directed path graph data, constructs a node mapping table based on the adjacent structure relationship, stores the upstream and downstream relationship types and connection directions between each morpheme, and obtains the graph path structure.
[0010] As a further solution of the present invention, the path intersection analysis module includes: The path extraction sub-module collects the node sequences in any two term paths based on the graph path structure, sequentially extracts the node number information under each path and establishes a term node mapping set, marks the term identification and path length parameters to which each path belongs, and obtains the term path node number value; The intersection comparison sub-module calls the node number sequences of any two term paths according to the term path node number value, performs an intersection comparison operation on the node sets of the two paths, extracts all the end node numbers in the intersection, respectively counts the number of times such nodes appear in different paths, and uses the formula: ; Calculate the intersection path end point deviation value and compare it item by item with the path intersection judgment reference value, screen the path pair combinations with the deviation value less than or equal to the reference value, and establish a set of path intersection numbers that meet the conditions. Among them, indicates the number of times the th intersection node appears in path A, indicates the number of times the same node appears in path B, indicates the total number of times the node appears in all path sets, is the total number of intersection nodes;
[0011] As a further solution of the present invention, the semantic category determination module includes: The semantic label collection sub-module collects the semantic label sets to which each end node belongs based on the end nodes in the term intersection path set, performs an index mapping between the term path and its end label, and generates a path end semantic label group; The label frequency statistics sub-module performs a repetition count operation on all semantic labels based on the semantic label group at the path end point, records the number of times each semantic label appears in the term path set, and performs a sorting process from high to low according to the number of appearances to obtain a sorted semantic label sequence; The generic label determination sub-module performs a matching judgment on the label set corresponding to the end node in the term path according to the sorted semantic label sequence, screens the label item that is the most forward in the sorted sequence in each path as the attributed semantic category corresponding to the path, and integrates the attributed labels of all term paths to obtain a term semantic attribution label group.
[0012] As a further solution of the present invention, the classification structure output module includes: The node extraction sub-module collects the graph classification nodes corresponding to each label according to the term semantic attribution label group, records the associated path numbers and the number of corresponding term path sets in each classification node, determines the matching index between the semantic label and the graph node, and obtains the label attribution node number value; The path classification sub-module divides the corresponding term paths into each graph classification node based on the label attribution node number value, with the semantic label as the classification basis, establishes a two-way corresponding structure between the term path number and the node number, extracts the path number list attributed to each node, and obtains the node path attribution quantity value; The structure generation sub-module performs a structural mapping integration of the graph classification node and the subordinate term path numbers according to the node path attribution quantity value, outputs the classification node index, the corresponding semantic label and the total number of paths, determines the attribution of the node and the term path, and generates a foreign trade data classification structure table.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, morpheme-level word segmentation and word order fragment extraction are performed on the technical text in foreign trade data. By comparing the positional relationships of keywords in the morpheme sequence, an original word order template is constructed to achieve fine-grained modeling of the semantic structure of terms, enhance the word order restoration ability of term parsing, set priorities according to the hierarchical labels of morpheme nodes, construct a directed path from the top layer to the bottom layer for terms, establish a multi-level semantic path through the connection relationship between node numbers and edges, enhance the graph linkage between term structures, introduce a node intersection calculation and end point frequency statistics mechanism to form a semantic intersection area between paths, strengthen the ability to capture semantic overlap, based on the sorting and screening of the end point node attribution labels in the cross paths, use the label with the highest frequency as the basis for path categorization, enhance the semantic focusing effect of the attribution label, combine semantic labels to perform term path mapping and structural aggregation on the graph classification nodes, achieve a complete path closed-loop from terms to node classification, enhance the systematization and traceability of classification management, and the overall process constructs a dynamic semantic recognition and classification mapping system relying on morpheme structure, semantic path, cross relationship, and attribution label, improving the accuracy of term organization, the structural coordination between paths, and the semantic carrying capacity of graph nodes, and greatly improving the classification management process of foreign trade data. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is the system flow chart of the present invention; Figure 2 is the flow chart of the term word order modeling module of the present invention; Figure 3 is the flow chart of the graph path generation module of the present invention; Figure 4 is the flow chart of the path cross analysis module of the present invention; Figure 5 is the flow chart of the semantic categorization determination module of the present invention; Figure 6 is the flow chart of the classification structure output module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0015] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0016] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms length, width, up, down, front, back, left, right, vertical, horizontal, top, bottom, inner, outer, etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0017] Please refer to Figure 1 , the foreign trade data classification management system based on the knowledge graph includes: The term word order modeling module obtains the term text in the foreign trade data, performs morpheme-level word segmentation on the term text and extracts morpheme indexes, compares the first occurrence position of the keyword in the morpheme arrangement list with the number of morphemes, extracts the original word order segment corresponding to the keyword, and combines the occurrence frequency values of all keywords to construct a term original word order sequence template; The graph path generation module, based on the term original word order sequence template, combines the hierarchical labels of each morpheme node of the term, sorts the morphemes according to the priority of the hierarchical labels, establishes a directed path of morphemes from the top layer to the bottom layer according to the sorting result, collects the connection node numbers and adjacent node relationships of all edges in the path, and constructs a graph path structure; The path intersection analysis module, according to the graph path structure, extracts the node sequences in any two term paths, performs an intersection comparison of the path node sets on the two node sequences, counts the number of times the end node appears in different paths in the intersection path, and compares it with the path intersection judgment reference value to screen the term paths that meet the cross-node conditions to obtain a term cross-path set; The semantic category determination module, based on the end nodes appearing in the term cross-path set, collects the semantic label set to which the end nodes belong, performs a sorting operation on the repetition times of the semantic label set, combines the semantic labels of the end nodes corresponding to each term path, and performs a matching judgment with the sorted label set, selects the semantic label with the highest occurrence frequency value as the attribution category label of the target term path, and generates a term semantic attribution label group; The classification structure output module, according to the term semantic attribution label group, counts the graph classification nodes to which each label belongs, extracts the total number of the belonging path list and the belonging term path set in the classification nodes, divides the term paths into the corresponding nodes according to the semantic attribution label as the classification basis, establishes the classification attribution relationship structure between the nodes and the term paths, and generates a foreign trade data classification structure table.
[0018] The original word order sequence template of terms includes a morpheme index arrangement structure, a keyword word order segment set, and a keyword frequency weight model. The atlas path structure includes a morpheme level mapping relationship, a directed path node chain, and a node connection relationship set. The term cross-path set includes a path intersection node set, an end node occurrence frequency distribution, and a cross-node screening result. The term semantic attribution label group includes a semantic label frequency ranking, a path end semantic label mapping, and an attribution generic label matching result. The foreign trade data classification structure table includes a classification node identifier, a term path grouping result, and a classification attribution relationship mapping result.
[0019] Please refer to Figure 2 , the term word order modeling module includes: The morpheme extraction sub-module obtains the term text in the foreign trade data, performs morpheme-level word segmentation on the term text, extracts the morpheme set of each term, records the index position of each morpheme in the term, compares the relationship between the first occurrence position of the keyword in the morpheme arrangement list and the number of morphemes, classifies by term, and obtains the keyword morpheme index distribution result; To obtain the term text in the foreign trade data, it is necessary to first collect the actual foreign trade corpus from specific industry scenarios. For example, extract term texts such as export plastic packaging film and stainless steel pipe fittings from the customs declaration list. For each term text, apply the separation logic to split it into morphemes by words. For example, split export plastic packaging film into export, plastic, packaging, and film, and then number each morpheme. Their index order is 1 to 4 in sequence. Then identify the keywords from the morpheme set. For example, set the keywords as packaging and film from the database. Judge their first occurrence positions in the morpheme list as 3 and 4 respectively. At the same time, count the total number of morphemes in this term as 4. Compare the keyword index with the total number of morphemes. If the index position is less than the median of the total number of morphemes, it is recorded as a preposition morpheme, otherwise it is recorded as a postposition morpheme. Use the preposition or postposition identifier to participate in the subsequent word order judgment. In this process, a keyword index interval judgment threshold needs to be set to divide the importance of morphemes. For example, set the preposition position as morpheme index ≤ 2 and the postposition as index ≥ 3. Classify and mark each keyword through this threshold, and then summarize the distribution status of all keywords in each term, and establish a morpheme index mapping array [3, 4] of keywords. Bind the result to the term identifier to construct a term index distribution matrix. If multiple terms are processed in this process, the following example data can be formed. For example, export stainless steel pipe fittings → [1, 2, 3, 4], the keywords steel and pipe have indexes 3 and 4. The total number of morphemes in this term is 4, and the number of keywords is 2. Thus, the keyword morpheme index distribution value is [(steel, 3), (pipe, 4)], and it is classified into the keyword vector corresponding to the term to form the result keyword morpheme index distribution result.
[0020] The original sequence fragment construction sub-module extracts the morpheme fragments corresponding to the keywords in the original text of the terms according to the keyword morpheme index distribution result, intercepts the morphemes based on the index range of the keywords in the terms, constructs a set of morpheme fragments according to the positions of the intercepted morphemes of each keyword, and reorganizes them in combination with the terms to which the morphemes belong to obtain a set of original sequence fragments of the keywords; Based on the keyword morpheme index distribution result, read the terms and their keyword index information one by one, intercept the morphemes according to the keyword positions in the original text of the terms, divide the original sequence fragments by the intercept interval. For example, in the term "export plastic packaging film", the keyword "packaging" is at the 3rd morpheme, so the morphemes from the 3rd to the 4th are extracted to form the fragment "packaging film". Further, extract all the original sequence fragments of the keywords in the term, and mark the attribution of multiple fragments to form a set of morpheme fragments. For example, for the term "export plastic packaging film", the keyword fragment set {packaging film} is generated. If a term has multiple keywords, such as "export stainless steel pipe fittings", and its keywords "steel" and "pipe" are at the 3rd and 4th morphemes respectively, then the intercepted fragment is "steel pipe fittings", and the morpheme set is {steel pipe fittings}. For terms with multiple keywords, perform morpheme slicing and combination in ascending order of keyword index sorting. If the slice interval is less than or equal to 1, merge the fragments; if the interval is greater than 1, retain the respective fragments. Finally, construct a set of original sequence fragments corresponding to multiple keywords in each term, and organize this set in terms of terms to form a term original sequence structure library, which is the result keyword original sequence fragment set.
[0021] The word order template generation sub-module counts the occurrence frequency values of all keywords according to the set of original sequence fragments of the keywords, performs position rearrangement processing on the set of morpheme fragments based on the original order of the keywords in the term text, splices the original sequence fragments of multiple keywords in the same term according to the first occurrence position, and uses the formula: ; Calculate the total weight of the original word order of the term , and perform classification and integration in combination with the word order weight results corresponding to each term to obtain the original word order sequence template of the term. Among them, represents the first index position of the th item of the keyword, represents the occurrence frequency value of the th item of the keyword, represents the length of the morpheme fragment corresponding to this keyword, represents the number of keywords in this term, represents the total number of morphemes in the term, represents the total number of all keywords in the term; According to the set of original sequence fragments of keywords, read the set of keywords and their occurrence frequency values in each term. For the term "export stainless steel pipe fittings", set the occurrence frequencies of the keywords "steel" and "pipe" to 2 and 3 respectively, read their original sequence indexes as 3 and 4 respectively, and the corresponding morpheme fragment lengths as 1 and 2 respectively. The total number of keywords in this term is 2, and the total number of morphemes is 4. Substitute into the formula: ; Substitute the parameter values into the formula: Group 1: , , , , , ; Group 2: , , , , , ; The calculation process is as follows: The value of the first operation: ; The value of the second operation: ; Sum and average: ; This result indicates that the total weight of the original word order of the term "export stainless steel pipe fittings" is 2.68. Thus, the word order weights of multiple terms are classified and integrated to form the word order template corresponding to each term. For example, the weight of "export plastic packaging film" is 1.75, and the weight of "stainless steel pipe fittings" is 2.68. After sorting or clustering them, the original word order sequence template of the result terms can be generated.
[0022] Table 1 Summary of term word order weights: ; Table 1 lists the total weights of different terms in the reconstruction of the original word order, which can be used for word order template classification and structural induction.
[0023] The formula focuses on the reconstruction of the original word order of keywords in the term, reflecting the comprehensive influence of factors such as the position characteristics, frequency weight, and structural complexity of keywords in the word order. Among them, the position where the keyword first appears represents its structural priority in the morpheme sequence. The smaller the value, the more forward it is, directly reflecting its dominant position in the word order; the keyword occurrence frequency is processed by taking the square root to be , which is used to reflect its marginal influence of non-linear growth and prevent high-frequency words from causing asymmetric interference to the word order structure; the keyword morpheme fragment length and the total number of keywords The product, representing the keyword structure density in the term, is then divided by the total number of morphemes in the term to form , which serves as a normalized measure of the term structure complexity. The overall weighted expression is , used to balance the keyword position advantage, frequency influence, and structural density contribution. By combining the absolute value operation to unify the word order directionality, finally, the influence values of all keywords in the term are summed and averaged , obtaining a unified measurement result of the term word order weight. Therefore, this formula realizes a word order modeling mechanism of position priority + frequency reconciliation + structural compensation in terms of structure, making the word order template more representative and comparable as a whole
[0024] Please refer to Figure 3 , the graph path generation module includes: The hierarchical sorting sub-module, based on the original word order sequence template of the term and in combination with the hierarchical labels of each morpheme node of the term, compares and sorts all morpheme nodes according to their hierarchical label priority values, rearranges the positions of the term morphemes in the order from the top layer to the bottom layer, establishes a rearranged sequence index table, and obtains the morpheme sorting index value Based on the original word order sequence template of the term, first extract all morpheme nodes in the term and their corresponding hierarchical labels. For example, the term "remote sensing device module" can be disassembled into morphemes "remote", "sensing", "device", "module", and the corresponding hierarchical labels are 3, 2, 1, 0 respectively, indicating that "module" is the top layer, and so on downwards. Subsequently, all morpheme nodes are sorted according to their hierarchical label priorities. Here, it is necessary to judge the label values. The lower the priority, the closer it is to the top layer. For example, the morpheme "module" with a hierarchical label of 0 is sorted first, followed by "device", "sensing", "remote". After the sorting is completed, the morpheme order is "module - device - sensing - remote". In the specific sorting process, by comparing whether the adjacent morpheme label value is less than the previous node, if so, the current morpheme is moved forward. After the sorting is completed, a rearranged sequence index value is generated for each morpheme, numbered sequentially from 0, such as "module" is 0, "device" is 1, etc. This index is used for subsequent path establishment. If the number of terms increases, for example, a new term "intelligent analysis terminal device" is added, where the morpheme "terminal" is at hierarchical level 0, "device" is at 1, "analysis" is at 2, and "intelligent" is at 3, then the sorting result is "terminal - device - analysis - intelligent". This sorting rule ensures the top-down structural order, and the indexes are from 0 to 3, as shown in Table 2
[0025] Table 2 Morpheme Hierarchical Sorting Index Table: ; As shown in Table 2, the hierarchical labels directly affect the sorting order. The comparison actions in the sorting operation can be attributed to performing greater than or less than judgments on the label values of the morpheme nodes. That is, if the label value of the previous item is greater than that of the latter item, their order is swapped. The overall sorting process is completed through multiple rounds of judgments, and this process finally generates a morpheme sorting index value.
[0026] The path construction sub-module obtains the set of adjacent nodes in the morpheme rearrangement sequence according to the morpheme sorting index value, numbers each pair of adjacent nodes respectively and records their connection directions, and uses the formula: ; Calculate the offset intensity value of the morpheme directed path , combined with the connection relationships of all nodes in the path, integrate the structural edge information to generate morpheme directed path map data, where, represents the index number of the th morpheme in the sorted list, represents the index number of the th morpheme in the sorted list, represents the connection span value of the th morpheme, represents the connection span value of the th morpheme, represents the absolute value of the hierarchical label difference between the th and the th morphemes, is the number of edges in the path; According to the morpheme sorting index value, connect the sorted morpheme sequences bit by bit to establish adjacent node connection pairs. For example, in the previous example, module-device-sensing-remote corresponds to node indexes 0-1-2-3, and the adjacent connection pairs are (0, 1), (1, 2), (2, 3). Each pair of connection pairs represents an edge of the morpheme path. In path construction, it is necessary to collect the node numbers of each connection edge and calculate the structural offset, that is, the latter node number minus the former node number. For example, the offset of (1, 2) is 1. In addition, the hierarchical difference needs to be combined. Assuming that the hierarchy of the node module is 0 and the device is 1, the hierarchical difference is 1. To quantify the structural offset intensity of each edge in the path, a connection span parameter is introduced. Let the span between module-device be 2, device-sensing be 3, and so on. The following formula is used for structural strength calculation: Let , , , then: The first segment: ; The second segment: ; The third segment: ; The sum is: ; The result shows that the morpheme path deviation intensity value is 6.0. This value represents the directional connection strength of the term under the sorting structure, and can be used as a measurement basis for the path organization logic, to unify the structure layout and link direction in the subsequent construction of the atlas path, and finally generate the morpheme directed path atlas data.
[0027] The formula fuses the structural and hierarchical information among multiple parameters to characterize the connection strength characteristics of the paths between morpheme nodes. Among them, represents the position difference of adjacent nodes in the sorting, reflecting the sequence span between the connected nodes, is the sum of the morpheme spans connecting two nodes, measuring the information load density in the path. The product of the two is divided by 2, indicating the weighted average contribution of this path edge to the connection strength. Further considering the influence of the hierarchical label difference, represents the absolute difference between two nodes in the semantic hierarchy. Using the square root operation reflects the non-linear weakening trend that as the hierarchical difference increases, the influence of the boundary change on the path strength tends to be gentle. Subtracting the above two parts and taking the absolute value can suppress the directional influence of the hierarchical fluctuation on the path contribution, so as to form a unified positive structure strength value. Finally, the summation operation is performed on all path edges to integrate into the overall structure deviation intensity value, reflecting the organizational rationality and hierarchical distribution characteristics of the path as a whole.
[0028] The node structure extraction sub-module collects the node numbers and adjacent node pair relationships in all connection edges according to the morpheme directed path atlas data, constructs a node mapping table based on the adjacent structure relationship, stores the upstream and downstream relationship types and connection directions between each morpheme, and obtains the atlas path structure; According to the morpheme directed path atlas data, the starting node number and the ending node number of each connection edge are collected. On this basis, a connection mapping table is established to distinguish the upstream and downstream node position relationships. If the smaller numbered one is the starting point, the direction is defined as positive. If the larger numbered one is the starting point, the direction is reverse. For example, the node number pair (1, 2) is positive, and (3, 2) is reverse. By judging the number relationship of the connection pairs one by one, the node direction classification is completed. Taking module-device as an example, module is 0, device is 1, and the defined direction is 0→1. Similarly, sensing-remote is 2→3, forming an edge direction mapping table. Then all edges and the corresponding directions are classified and stored. By number normalization, a unified format of node connection can be established among multiple terms. This mapping structure is further used to extract the structural paths between nodes in the morpheme atlas and record them as a structure table, and finally the atlas path structure is obtained.
[0029] Please refer to Figure 4 , the path intersection analysis module includes: Based on the graph path structure, the path extraction sub-module collects the node sequences in any two term paths, sequentially extracts the node number information under each path, establishes a term node mapping set, marks the term identifiers and path length parameters to which each path belongs, and obtains the term path node number values. When extracting the node sequences of any two term paths in the graph path structure, first, the node path set in the term graph data should be obtained. On this basis, unique number identifiers should be assigned to the two target terms. For example, the path node numbers of term A are {101, 103, 107, 110}, and the path node numbers of term B are {102, 103, 108, 110}. When extracting the node number sequence, it should be arranged according to the order in which the nodes appear in the path. If there are duplicate nodes in the node sequence, the first occurrence position should be retained according to the path order, and the remaining positions should be discarded. Then, a two-term path node index mapping table is established. The table should record the node numbers, the position indexes of the nodes in the path, the corresponding term identifiers, and the path length values. For example, the path length of term A is 4, and the path length of term B is also 4. At this time, the node numbers 103 and 110 exist in both terms, indicating that there is an intersection part, and the record is as follows Table 3: Node mapping table: ; As shown in Table 3, the table records the position indexes of each node in different paths and the intersection determination situation. Among them, whether there is an intersection is a Boolean field. If the node number exists in both term paths, it is set to yes, otherwise it is set to no. This determination does not require a fuzzy interval. Finally, the node numbers with yes in whether there is an intersection in the table are extracted and converted into an intersection path node set, and the term path node number values can be obtained.
[0030] Based on the term path node number values, the intersection comparison sub-module calls the node number sequences of any two term paths, performs an intersection comparison operation on the node sets of the two paths, extracts all the end node numbers in the intersection, and respectively counts the number of times such nodes appear in different paths. The formula is used: ; Calculate the end point deviation value of the intersection path , and compare it item by item with the path intersection judgment reference value. Filter out the path pair combinations with deviation values less than or equal to the reference value, and establish a set of intersection numbers of paths that meet the conditions. Among them, represents the number of times the th intersection node appears in path A, represents the number of times the same node appears in path B, represents the total number of times the node appears in all path sets, is the total number of intersection nodes; When performing an intersection comparison process on a set of nodes according to the term path node number values, the set of intersection nodes should be selected first. In this embodiment, the intersection nodes are {103, 110}. The occurrences of these two nodes in Path A and Path B are counted respectively. Since the node numbers in each path are unique, each of the two nodes appears 1 time in Path A and Path B. Then, calculate the total number of times they appear in all path sets. For example, if there are 5 term paths in the system, where 103 appears in 3 of them and 110 appears in 4 of them, the intersection node statistical parameters are as follows: Table 4: Statistical parameter record: ; Substitute the data in Table 4 into the formula and calculate to get: ; According to the calculation result, the end point deviation value of the intersection path is 0, and the path crossing judgment reference value is set to 0.5. The basis is the median value of the discrete tolerance interval obtained after normalizing the distribution range of the intersection node deviation amounts in the entire path set. Specifically, the distribution range of the deviation values of the intersection nodes in all path pairs in the statistical sample set is [0, 1.3], and the deviation values of more than 70% of the path pairs are concentrated in the interval [0, 0.5]. Therefore, 0.5 is set as the conservative boundary point for screening structurally consistent path combinations. This reference value tends to decrease as the path node density of the entire graph increases. The increase in density leads to an increase in the coincidence degree between nodes and an increase in the cross-overlap phenomenon, and the tolerance space for deviations decreases accordingly. Therefore, it is statistically reasonable and graphically structure-corresponding to select a deviation value less than 0.5 as the judgment threshold. Set the path crossing judgment reference value to 0.5, and the numerical deviation result is lower than the reference value. Therefore, this path pair meets the cross node condition and enters the subsequent screening process. This result shows that there is no obvious difference in the occurrence rules of the intersection nodes between the two paths, and they have structural consistency, and a set of path intersection numbers that meet the conditions can be established.
[0031] The formula is based on a comprehensive measure of the differences in the occurrences of path node intersections in the two paths. Its core lies in measuring whether the occurrence distributions of the intersection nodes in Path A and Path B are consistent. First, is used to strengthen the influence of the nodes that frequently appear in Path A on the deviation value, where is squared, which can amplify the weight difference in the case of repeated node occurrences, and then improve the sensitivity to the interference of frequent nodes; this difference value is then divided by , the purpose is to smooth the nodes frequently appearing in all paths, reduce the dominance of high-frequency nodes over the overall deviation in the form of taking the square root, and add 1 to prevent the problem of the denominator being zero when a node appears only once; after the calculation results of each node, the absolute value is used to ensure that the deviation value is always non-negative, avoiding the positive and negative cancellation from affecting the overall judgment. Finally, the deviation values of all intersection nodes are summed up and then divided by the number of intersection nodes. , the normalization process is completed to obtain the average level of the deviation distribution between nodes, so as to measure the structural consistency degree between two paths. The overall formula takes into account the node frequency difference, the global frequency distribution, and the influence amplitude of a single node, enabling the deviation value to reasonably reflect the similarities and differences in the cross-structure between paths.
[0032] The path screening sub-module queries the original term path identifiers according to the set of intersection numbers of the paths that meet the conditions and the corresponding path combinations according to the numbers, integrates the term identifiers and the path intersection node information, and establishes a term path relationship linked list to generate a set of term cross paths. After obtaining the set of intersection numbers of the paths that meet the conditions, it is necessary to reverse query the original term identifiers and their path structure information corresponding to the intersection node set. For example, if the paths corresponding to the numbers 103 and 110 are Term A and Term B, it is necessary to retrieve their path structures from the term path index table and generate a set of term pairs. Then, an association linked list between the term pairs and their cross-node sets is constructed. The fields in the linked list include the number of Term A, the number of Term B, the cross-node set, the path length ratio, the number of cross nodes, etc. If the path lengths of Term A and Term B are both 4 and the number of cross nodes is 2, then the path length ratio is 1 and the cross-node ratio is 0.5. Furthermore, a cross-path tuple {Term A, Term B, {103, 110}, 1, 0.5} is established, and the combinations with path length ratios deviating from the interval [0.5, 2] are screened out, and the qualified path tuples are stored in the final structure table to obtain the set of term cross paths.
[0033] Please refer to Figure 5 , the semantic category determination module includes: The semantic label collection sub-module collects the set of semantic labels to which each end node belongs based on the end nodes appearing in the set of term cross paths, performs index mapping between the term path and its end label, and generates a path end semantic label group. Based on the end nodes that appear in the set of term cross paths, after collecting each end node of the term path, it is necessary to first extract the semantic tags marked for this node in the semantic database, and construct the corresponding relationship between each node and the semantic tag through the number mapping method. Suppose there are paths P1, P2, and P3, and their end nodes are N7, N12, and N15 respectively. Then, according to the term graph structure, it is retrieved that the tag corresponding to N7 is the supply chain, N12 is the payment method, and N15 is the transaction process, forming an initial tag set; during the further extraction process, the situation where multiple paths have duplicate end nodes needs to be considered. For example, the end nodes of paths P4 and P5 are both N7, and their tags are also unified as the supply chain. The processing of such nodes is achieved by matching the path numbers and node numbers one by one, and the mapping relationship is stored in the form of an array to ensure the efficiency of retrieval and subsequent operations; if the number of term paths is 20 and the end nodes involve 10 semantic tags, the following structure is recorded in a two-dimensional table, as shown in Table 5.
[0034] Table 5 Distribution Table of Semantic Tags of End Nodes: ; As shown in Table 5, for each term path, the tags corresponding to its end nodes have been clearly listed through number indexing. During the collection process, the paths are sorted by path number, and the tags corresponding to the node numbers are queried one by one to ensure the integrity and consistency of tag acquisition, and then a semantic tag group for the path end is established.
[0035] The tag frequency statistics sub-module performs a repetition count operation on all semantic tags based on the semantic tag group for the path end, records the number of times each semantic tag appears in the term path set, and performs a sorting process from high to low according to the number of occurrences to obtain a sorted semantic tag sequence; After obtaining the semantic tag group for the path end, it is necessary to perform frequency statistics on all semantic tags. The operation process is achieved by recording the cumulative number of times the tag appears in the term path. The frequency array F is used to count the number of occurrences of each tag number. For example, in the path set, the supply chain tag appears 5 times, the payment method tag appears 3 times, and the transaction process tag appears 2 times. Then the initial frequency array is , and the corresponding index structure of the tag needs to be clarified before sorting. Suppose the tag numbers are t1, t2, and t3 respectively, corresponding to the above tags; when sorting, it is arranged in descending order of frequency, that is ; In frequency calculation, if the end nodes of some paths are empty, the path needs to be skipped and not included in the frequency statistics; to improve the statistical accuracy, semantic tags can also be filtered according to the path coverage rate. This filtering criterion is set based on the representativeness of the semantic tags. Specifically, the total number of term paths P is set to the total number of valid paths entered during system construction. If the occurrence times of a semantic tag are less than 10% of P, it is considered a low-frequency and non-representative tag. This 10% threshold is taken from the lower bound of the optimized interval of semantic classification accuracy. This lower bound setting is derived from the sample stability analysis. When the tag coverage rate is less than one-tenth of the total sample, its semantic stability shows a linear decay trend. This value is adjusted according to the change in the total number of term paths. For example, when the total number of paths is 20, the threshold is set to 2. If the total number of paths increases to 50, the corresponding threshold is 5. This value is not fixed as a constant but dynamically adjusted according to the number of input samples of term paths to ensure the sensitivity and dynamic adaptability of the filtering criterion to the label distribution within the system. If this benchmark is set, tags with a frequency of 1 such as t4: import and export restrictions will be excluded; finally, the sorting result is an ordered pair array of tag numbers and their frequencies. , that is, generate a sorted semantic tag sequence.
[0036] The generic tag determination sub-module performs a matching judgment on the label set corresponding to the end nodes in the term path according to the sorted semantic tag sequence, screens the label item that is the most forward in position in the sorted sequence for each path as the semantic category to which the path belongs, and integrates the attribution labels of all term paths to obtain the term semantic attribution label group; According to the sorted semantic tag sequence, perform a matching judgment on the label set corresponding to the end nodes in the term path in turn. The matching process uses a label comparison mechanism. For the semantic label of the end node of each path, retrieve its position number in the sorted label sequence. If the label is in the front position in the sorted sequence (for example, the top 3), it is marked as a successful match; for example, the end label of path P1 is "supply chain", ranked 1st in the sorted list, and the match is successful. The end label of path P6 is "enterprise credit", ranked 7th, so it is not selected; the successfully matched label is regarded as the path attribution label. If multiple labels are matched at the end of the path at the same time, the final generic category is determined according to the label with the most forward ranking; finally, integrate the corresponding relationships between all term paths and their attribution labels to construct term path generic label pairs, such as , and generate the term semantic attribution label group one by one.
[0037] Please refer to Figure 6 , the classification structure output module includes: The node extraction sub-module collects the graph classification nodes corresponding to each label according to the term semantic attribution label group, records the associated path numbers and the quantity of the corresponding term path sets in each classification node, determines the matching index between the semantic label and the graph node, and obtains the label attribution node number value; According to the semantic attribution tag group of terms, successively collect the corresponding relationship information of each tag in the term path, traverse the term path to which each semantic tag belongs, extract the graph classification node numbers bound in the graph structure, and for the paths with the same semantic tag, count the distribution of all its graph node numbers. For example, if paths 1, 2, and 3 are bound under tag a and belong to nodes 101, 101, and 102 respectively, it can be known that the frequency of node 101 corresponding to this tag is 2, and the frequency of node 102 is 1. Therefore, a many-to-one matching table structure between semantic tags and node numbers needs to be established for subsequent path attribution judgment. During the statistical process, it is necessary to judge whether the node number value is unique. If it is unique, the tag belongs to this node; if it is not unique, the frequency needs to be counted and the maximum value is used as the basis for attribution judgment. In the above process, the actually selected semantic tags are tag a, tag b, tag c, and tag d respectively. Their corresponding paths come from the previous path aggregation stage, and the node numbers are collected from the structure values constructed in the path mapping table. After cleaning its mapping information, a binding index dictionary structure from semantic tags to graph nodes is constructed, and the node number distribution quantity information under this index structure is extracted to obtain the tag attribution node number value.
[0038] Based on the tag attribution node number value, the path classification sub-module divides the corresponding term paths into each graph classification node according to the semantic tag as the classification basis, establishes a two-way corresponding structure between the term path number and the node number, extracts the list of path numbers belonging to each node, and obtains the node path attribution quantity value; Based on the tag attribution node number value, the term paths are corresponded to the corresponding classification node numbers according to their semantic tags. In this process, all term path numbers need to be traversed, and their semantic tags need to be queried, and then mapped to their bound node numbers, and the path numbers are classified into the node path set. During this process, the path numbers belonging to the same node number need to be recorded in the same structure. For example, when the node number is 101, its attribution path set may include paths 1, 4, 8, etc. When counting, the path set quantity is accumulated. After performing such operations on all node numbers, the number of paths belonging to each node can be obtained. For example, there are 15 paths belonging to node 101, 9 paths belonging to node 102, 12 paths belonging to node 103, and 7 paths belonging to node 104. A corresponding structure of number and quantity is established for these statistical results to obtain the node path attribution quantity value.
[0039] The structure generation sub-module maps and integrates the graph classification nodes and the subordinate term path numbers according to the node path attribution quantity value, outputs the classification node index, the corresponding semantic tag and the total number of paths, determines the attribution of nodes and term paths, and generates a foreign trade data classification structure table; According to the node path attribution quantity value, for each node in the graph, collect its node number, bound semantic label, and the total number of path numbers to which it belongs. Combine these three pieces of information into a single structure record. During the construction process, extract the quantity of its path set from the pre-order structure for each node number and mark it in the structure table. Then, perform a reverse lookup on the semantic label bound to this node and record it together. Use the semantic label, node number, and path quantity as the three field values of the output structure. Finally, establish the classification structure table structure. Each row in the structure table records the classification structure information of a node, and the fields are shown in Table 1. The mapping relationship between the semantic label, path, and node is clearly marked, thereby generating the foreign trade data classification structure table.
[0040] Table 6 Foreign Trade Semantic Label Classification Statistical Table: ; As shown in Table 6, there are obvious distribution differences in the semantic labels among the attributed nodes. This table serves as the data basis source for the subsequent classification structure table.
[0041] The above is only a preferred embodiment of the present invention and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. The foreign trade data classification management system based on knowledge graph is characterized by: The system comprises: The terminology word order modeling module obtains terminology texts in foreign trade data, extracts morpheme indexes through word segmentation, determines the first position and frequency of keywords, maps and generates morpheme arrangement sequences, and constructs terminology original word order sequence templates; The graph path generation module sorts the morphemes based on the original word order sequence template of the term, establishes a directed path of morphemes from the top layer to the bottom layer according to the sorting result, collects the connection node numbers and adjacent node relationships of all edges in the path, and constructs a graph path structure; The path intersection analysis module extracts the term path node sequence according to the graph path structure, compares the intersection nodes and counts the frequency of the terminal nodes, screens the intersection paths, and obtains the term intersection path set; The semantic category determination module collects semantic labels of the terminal nodes based on the terminal nodes in the term intersection path set, sorts them by frequency of occurrence, and matches the path terminal labels to obtain a term semantic attribution label group; The classification structure output module counts the graph classification nodes to which each label belongs according to the term semantic attribution label group, divides the term path into corresponding nodes, establishes the node and term path classification attribution relationship structure, and generates a foreign trade data classification structure table.
2. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The term original word order sequence template includes a morpheme index arrangement structure, a keyword word order fragment set, and a keyword frequency weight model; the graph path structure includes a morpheme hierarchical mapping relationship, a directed path node chain, and a node connection relationship set; the term cross path set includes a path intersection node set, a terminal node occurrence frequency distribution, and a cross node screening result; the term semantic attribution label group includes a semantic label frequency ranking, a path terminal semantic label mapping, and an attribution category label matching result; the foreign trade data classification structure table includes a classification node identifier, a term path grouping result, and a classification attribution relationship mapping result.
3. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The term order modeling module includes: The morpheme extraction submodule obtains the term text in the foreign trade data, performs morpheme-level segmentation on the term text, extracts the morpheme set of each term, records the index position of each morpheme in the term, compares the relationship between the first appearance position of the keyword in the morpheme arrangement list and the number of morphemes, classifies by term, and obtains the keyword morpheme index distribution result; The original sequence fragment construction submodule extracts the morpheme fragments corresponding to the keyword in the original text of the term according to the keyword morpheme index distribution result, intercepts the morphemes based on the index interval of the keyword in the term, constructs a morpheme fragment set according to the position of each keyword intercepted morpheme, and reorganizes the morphemes in combination with the terms to which the morphemes belong, to obtain a keyword original sequence fragment set; The word order template generation submodule counts the occurrence frequency values of all keywords according to the original sequence fragment set of keywords, and performs position rearrangement processing on the morpheme fragment set based on the original order of keywords in the term text, and sequentially splices the original sequence fragments of multiple keywords in the same term according to the first appearance position, using the formula: ; Calculate the sum of the original word order weights of the terms , combined with the word order weight results corresponding to each term, classify and integrate them to obtain the original word order sequence template of the term, where, Representative keywords the first index position of the item, Representative keywords The frequency value of the item, Indicates the length of the morpheme segment corresponding to the keyword. Indicates the number of keywords in the term. Indicates the total number of morphemes in a term, Indicates the number of all keywords in the term.
4. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The graph path generation module includes: The hierarchical sorting submodule compares and sorts all morpheme nodes according to their hierarchical label priority values based on the original word order sequence template of the term and the hierarchical labels of each morpheme node of the term, rearranges the positions of the term morphemes in order from the top layer to the bottom layer, establishes a rearranged sequence index table, and obtains a morpheme sorting index value; The path construction submodule obtains the adjacent node set in the morpheme rearrangement sequence according to the morpheme sorting index value, and numbers each pair of adjacent nodes to record their connection direction, using the formula: ; Calculate the morpheme directional path offset strength value , combining all node connection relationships in the path, integrating structural edge information, and generating morpheme directed path graph data, where Indicates The index number of the morpheme in the sorted list, Indicates The index number of the morpheme in the sorted list, Indicates The connection span value of morphemes, Indicates The connection span value of morphemes, Indicates The first The absolute value of the level label difference between morphemes, is the number of edges in the path; The node structure extraction submodule collects the node numbers and adjacent node pair relationships in all connecting edges according to the morpheme directed path graph data, builds a node mapping table based on the adjacent structural relationship, stores the upstream and downstream relationship types and connection directions between each morpheme, and obtains the graph path structure.
5. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The path intersection analysis module includes: The path extraction submodule collects the node sequences in any two term paths based on the graph path structure, extracts the node number information under each path in turn and establishes a term node mapping set, marks the term identifier and path length parameter of each path, and obtains the term path node number value; The intersection comparison submodule calls the node number sequences of any two term paths according to the node number values of the term paths, performs an intersection comparison operation on the node sets of the two paths, extracts all the terminal node numbers in the intersection, and counts the number of times such nodes appear in different paths, using the formula: ; Calculate the intersection path endpoint deviation value , compare them one by one with the path intersection judgment benchmark value, select the path pair combinations with deviation values less than or equal to the benchmark value, and establish a set of path intersection numbers that meet the conditions, where, Indicates The number of times the intersection node appears in path A, represents the number of times the same node appears in path B, Indicates the total number of occurrences of the node in all path sets. is the total number of intersection nodes; The path screening submodule queries the original term path identifier according to the path combination corresponding to the number according to the set of path intersection numbers that meet the conditions, integrates the term identifier and the path intersection node information, establishes a term path relationship link list, and generates a term intersection path set.
6. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The semantic category determination module comprises: The semantic label collection submodule collects the semantic label set to which each terminal node belongs based on the terminal nodes in the term intersection path set, performs index mapping between the term path and its terminal label, and generates a path terminal semantic label group; The tag frequency statistics submodule performs a repetition count operation on all semantic tags based on the path endpoint semantic tag group, records the number of times each semantic tag appears in the term path set, and sorts them from high to low according to the number of occurrences to obtain a sorted semantic tag sequence; The category label determination submodule performs matching judgment on the label set corresponding to the terminal node in the term path according to the sorted semantic label sequence, selects the label item in each path that is at the front of the sorted sequence as the path corresponding to the semantic category, integrates the attribution labels of all term paths, and obtains the term semantic attribution label group.
7. The foreign trade data classification management system based on knowledge graph according to claim 1 is characterized in that: The classification structure output module includes: The node extraction submodule collects the graph classification nodes corresponding to each label according to the term semantic attribution label group, records the number of associated path numbers and their corresponding term path sets in each classification node, determines the matching index between the semantic label and the graph node, and obtains the label attribution node number value; The path classification submodule divides the corresponding term path into each graph classification node based on the node number value of the label and the semantic label as the classification basis, establishes a bidirectional correspondence structure between the term path number and the node number, extracts the path number list of each node, and obtains the node path attribution quantity value; The structure generation submodule integrates the graph classification nodes and the subordinate term path numbers according to the node path attribution quantity value, outputs the classification node index, the corresponding semantic label and the total number of paths, determines the attribution of the nodes and term paths, and generates a foreign trade data classification structure table.
Citation Information
Patent Citations
Automatic translation method for terms in large texts
CN103488628A
Knowledge base system, inter-word meaning relation determination method in the same system and computer program
JP2005157823A
Frequency information equipped word set generation method, program, program storage medium, frequency information equipped word set generation device, text index word production device, full text retrieval device and text classification device
JP2006243976A
Cited By
Knowledge graph construction method based on BNCT, terminal and medium
CN120236716A
Learner argumentation process effect evaluation method in critical thinking training mode
CN120470133A
Methods for evaluating the effectiveness of learners' argumentation process under critical thinking training models
CN120470133B
Neurosurgery patient follow-up visit management system
CN120496891A
Data element high-quality construction and application integrated operation platform
CN120804081A