Custom hierarchical processing method and system for multi-source data
Through custom hierarchical processing methods, the multi-source data is encoded and associated with tree structures and multi-forktree models, solving the problems of insufficient flexibility and functionality of traditional data fusion methods, and achieving efficient and flexible data fusion and management.
Patent Information
- Application Number
- CN202510668877.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Traditional multi-source data fusion methods have problems such as insufficient flexibility, low data structure abstraction level and insufficient functionality, and it is difficult to adapt to dynamically changing data environments and business needs.
By using custom hierarchical processing methods, by establishing a pre-database and standard database, using a tree structure and a multi-forktree model to encode and associate the data, hierarchical organization, fine-grained development and normalized management of the data.
It improves the flexibility and abstraction level of data fusion, reduces the cost of data fusion design and development, realizes dynamic adjustment, cropping, deformation and other functions of data sets, and improves overall implementation efficiency.
Smart Images

Figure CN120180377A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method and system for custom hierarchical processing of multi-source data. Background Art
[0002] The statements in this part only relate to the background art related to the present invention and do not necessarily constitute the prior art.
[0003] Multi-source data fusion refers to the organic combination and analysis of data information from multiple different sources to improve the accuracy, comprehensiveness, and reliability of information processing. Multi-source data fusion can make the obtained information more complete, give full play to the potential of various data, improve the data processing results, and promote the scientific and automated decision-making.
[0004] Multi-source data fusion methods involve multiple disciplinary fields, such as information processing, pattern recognition, artificial intelligence, database technology, etc. Common multi-source data fusion methods are as follows: First, the method based on weighted average: This method mainly comprehensively processes data from multiple sources with different weights. Common comprehensive methods include arithmetic average, weighted average, geometric average, etc. By reasonably setting the weights, the information from different sources can obtain corresponding weights to reflect its importance; Second, the method based on feature extraction: This method mainly converts the data from multiple sources into data in the same feature space through feature extraction and then fuses them. Usually, the methods of feature extraction include principal component analysis, wavelet transform, independent component analysis, etc. This method can reduce the heterogeneity of data sources, extract key feature information, and thus improve the effect of data fusion; Third, the method based on model: This method mainly models the data from different data sources based on a specific model and then fuses them. Among them, common models include neural network models, decision tree models, grey system models, etc. By using a specific model to describe the relationship between data sources, various data can be better fused, and the accuracy of the fusion result can be improved; Fourth, the method based on decision rules: This method comprehensively processes the results of multiple data sources based on decision rules. Common decision rules include simple majority voting, weighted majority voting, logistic regression, etc. By reasonably using decision rules to comprehensively process the results of multiple data sources, the effect of data fusion can be effectively improved.
[0005] However, traditional data fusion methods have the following disadvantages: 1. The fusion mode is not flexible. Traditional data fusion methods often require predefined data patterns and fusion rules, and are less adaptable to dynamically changing data environments. When the data structure or data type changes, the entire data fusion process needs to be redesigned and developed, making it difficult to quickly respond to current business requirements; 2. The business attribute abstraction level of the data structure is relatively low. The traditional tree-shaped data fusion method does not abstract the content between upper and lower level data. Information is only stored in leaf nodes, and intermediate-level child nodes do not store information, resulting in deficiencies in business logic, storage space, and usage patterns; 3. Insufficient functionality. Along with business changes, when the data set needs to be adjusted as a whole, only the existing data pattern can be invalidated, and then the data structure design and warehousing work need to be carried out again. There is a lack of functions such as overall or partial cropping and deformation of the data set, affecting the overall implementation efficiency. Summary of the Invention
[0006] Therefore, the object of the present invention is to provide a method and system for custom hierarchical processing of multi-source data, realizing an automatic data processing method for hierarchical organization, fine-grained expansion, and normalized management of multi-source data.
[0007] To achieve the above object, a method for custom hierarchical processing of multi-source data provided by the present invention includes the following steps: S1. Establish a pre-database to store the original data obtained from various channels; S2. Establish a standard database, set up multiple standard data sets in the standard database, create the node positions of the data in a tree structure in each standard data set, and encode the data according to the node positions; S3. According to the node positions, associate the data according to a preset encoding rule, and update the encoding of the associated data according to the associated node positions; the preset encoding rule includes: the data structure encoding is composed of the data bit and the level bit. Define the encoding length M, and take the maximum data level number N in the standard database, then the level bit is [Lg(N)] + 1, and the data bit is M - ([Lg(N)] + 1); The data bit is divided into data bits of each level from left to right according to the level number N, and the length of each level data bit is [(M - ([Lg(N)] + 1)) / N]; S4. Obtain the user's data processing request, and perform hierarchical processing on the node positions of the relevant data according to the user's processing request.
[0008] Further preferably, when the pre-database stores the original data, the data structure between the original data is synchronously stored.
[0009] Further preferably, in S2, when creating the node positions of data in each standard dataset using a tree structure, the tree structure is a multi-way tree, including a root node, intermediate nodes, and leaf nodes; among them, the root node and intermediate nodes store the common content of the part below this node, and the leaf nodes store the individual content of this node; each intermediate node includes 1 parent node and any number of child nodes; the root node is the unique identifier related to the described entity; the root node, intermediate nodes, and leaf nodes are all independent physical data tables; under the same dataset, the number of data entries of the root node, intermediate nodes, and leaf nodes is the same, and the data relationships are in one-to-one correspondence.
[0010] Further preferably, in S2, according to the node positions, data structure encoding is performed, and the encoding follows the following principles: In the same standard database, the data structure encoding lengths are the same, and the encoding bits use decimal numbers; The nodes under the same parent node are sorted in order from left to right, and naturally increase on this layer of data bits, with a step size of 1.
[0011] The data bits of each node are concatenated in order from top to bottom for all nodes from the root node to this node, and the data bits below this node level are all set to 0.
[0012] The root node, intermediate nodes, and leaf nodes of the multi-way tree of each standard dataset form independent physical data tables according to the association relationship; under the same dataset, the number of data entries of the root node, intermediate nodes, and leaf nodes is the same, and the data relationships are in one-to-one correspondence; All nodes are represented using vectors, including: using the Tf-idf method to extract the vector features of the text attributes of each node and perform vector representation; For all node vectors, the following formula is used for calculation to complete the normalization of all node vectors: ; Among them, X2 represents the vector after normalization of the current node, and x1 to xn are the original vectors of the child nodes under the current node.
[0013] Further preferably, in S4, when the user's data processing request is data splicing, N - 1 data buffer areas are prepared, a data pre-reading operation is performed, the data of the root node and each intermediate node are placed into the corresponding-level data buffer areas respectively, and data splicing is performed to generate each complete piece of data, which specifically includes the following steps: S411. Generate a request code according to the data processing request, and perform a search on the physical data table of the standard dataset using the pre-order breadth-first traversal method; S412. Obtain the root node data content into the data buffer area; S413. Read all the direct child nodes under the current node N1 in the order from left to right. When the read node matches the request code, record the node data content into the data buffer. S414. Determine whether there is a data bit +1 of the current node. If so, perform the next operation; if not, return to the upper level. S415. Repeatedly perform the operations of S413 - S414. Each time, based on the latest node recorded in the data buffer, read its direct child nodes below. S416. When all N - 1 positions in the data buffer are filled, generate the spliced data, and use the hashing method to generate 256 - bit data fingerprints for each piece of data as the unique identifier of the data.
[0014] Further preferably, in S4, when the user's data processing request is data cropping or splicing, Connect the P node, the lower - level node of the dataset A to be cropped, under the Q node of the dataset B, including: 1. Identify the nodes to be cropped: Identify the nodes to be cropped and all their child nodes except the root node. If the node to be cropped contains multiple child nodes, aggregate the vectors of all child nodes into the vector of the node to be cropped by averaging. 2. Select alternative placement points Traverse the entire tree structure to find all nodes at the same level or higher levels as the node to be cropped as candidate nodes. Let the node to be cropped be A and the candidate node be B. Calculate the similarity score between the node to be cropped and the candidate location node. Define a similarity threshold, take the candidate nodes exceeding the threshold as alternative placement points, and sort them according to the scores. All intermediate nodes and leaf nodes under the P node are transferred to be under the Q node following the P node. The mutual relationships of the nodes under the P node remain unchanged; the access rights, security levels, encryption levels, and de - sensitization levels of the nodes under the P node meet the definitions of the Q node.
[0015] Further preferably, in S4, when splicing and cropping data, it also includes: Verify the legality of the splicing target position. The legality verification of the splicing target position includes the following process: Obtain the hierarchical information of the target position and compare it with the maximum level. When the target position level is greater than the maximum level, it is judged as illegal. Determine whether the target position is a parent node. When the target position is a non-root node, query the database according to the parent node identifier of the target position to check if there is a corresponding parent node record; if there is no parent node record, it is determined to be illegal. Judge the ID repeatability of the nodes at the target position to be spliced. Use a deep learning model to map a low-dimensional embedding vector to each node in the tree structure. The embedding vector includes the node structure and the attributes of the node data; calculate the vector similarity between the nodes at the target position and the original nodes to determine whether the nodes are repeated and the compatibility of the nodes at the target position.
[0016] Further preferably, in S4, it also includes judging whether the spliced or cut tree structure meets the integrity constraints, including judging whether the tree structure meets the following after splicing or cutting: a. Node uniqueness constraint; globally search the node identifier field in the multi-way tree to ensure that the identifier of each node in the tree structure is unique within the entire tree structure; b. Parent-child relationship integrity constraint; for each non-root node, judge that the standard database points to its parent node record through a foreign key to ensure that each non-root node has and only has one parent node; c. Hierarchical continuity constraint: Let the set of levels be L = {1, 2, 3,..., k}. When inserting a new node, assume that the level of the node to be inserted is i belonging to the set L. Query the database to confirm whether its upper-level node exists. If it does not exist, the insertion operation is blocked.
[0017] The present invention also provides a custom hierarchical processing system for multi-source data for implementing the above-mentioned custom hierarchical processing method for multi-source data, including: A data storage module for establishing a pre-database to store the original data obtained from various channels; A database establishment module for establishing a standard database, setting up multiple standard data sets in the standard database, creating the node positions of the data in a tree structure in each standard data set, and encoding the data according to the node positions; A data encoding module for associating the data according to the node positions according to the preset encoding rules, and updating the encoding of the associated data according to the associated node positions; the preset encoding rules include the following content: the data structure encoding is composed of the data bit and the level bit spliced together. Define the encoding length M. According to the multi-way tree hierarchical structure, take the maximum data level number N in the standard database, then the level bit is [Lg(N)] + 1, and the data bit is M - ([Lg(N)] + 1); The data bit is divided into the data bits of each level from left to right according to the level number N, and the length of each level data bit is [(M - ([Lg(N)] + 1)) / N]; The hierarchical processing module obtains the user's data processing request and hierarchically processes the node positions of relevant data according to the user's processing request.
[0018] The present invention also provides an electronic device, including: a memory storing computer program instructions; a processor, when the computer program instructions are executed by the processor, implementing the steps of the above-mentioned multi-source data custom hierarchical processing method.
[0019] The present invention also provides a computer-readable storage medium, which is used to store instructions. When the stored instructions run on a computer, the computer is caused to execute the steps of the above-mentioned multi-source data custom hierarchical processing method.
[0020] A multi-source data custom hierarchical processing method and system disclosed in the present application, compared with the prior art, has at least the following advantages: 1. Improve the flexibility of data fusion. By means of customization, multi-level, and dynamic adjustment, reduce the design and development costs of data fusion to adapt to the dynamically changing data environment, including: the data model can be custom-designed; the data model can be set to any multi-level, and each part level can be different; the data model can be dynamically adjusted during use; a data model coding rule is designed for various data operations.
[0021] 2. Improve the abstraction level of the data structure. The present invention further abstracts the data meaning. Each level of node saves the general attributes of all its downward nodes, and they are spliced when needed, improving the flexibility of the data processing system in terms of business logic, storage space, usage mode, etc.
[0022] 3. Expand the functions of the data processing system. The custom multi-level data structure can be freely transformed, adjusted, etc., and at the same time, according to the designed data structure coding system, the whole or part of the data can be conveniently cut, deformed, etc. Description of the Drawings
[0023] Figure 1 It is a schematic structural diagram of the multi-source data custom hierarchical processing method provided by the present invention.
[0024] Figure 2 It is a schematic diagram of the data splicing process provided by the present application.
[0025] Figure 3 It is a schematic structural diagram of the multi-source data custom hierarchical processing system provided by the present invention. Detailed Embodiments
[0026] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0027] As shown Figure 1 in the figure, a method for custom hierarchical processing of multi-source data provided by an embodiment of the present invention on the one hand includes the following steps: S1. Establish a pre-database to store the original data obtained from various channels; mark the data attributes according to the data categories; when storing the original data in the pre-database, synchronously store the data structure between the original data. Through this step, the original data with a wide range of sources and diverse formats are uniformly aggregated, providing a data basis for subsequent systematic processing. Whether the data comes from imported data sources, interface data sources, database extraction, or file import, etc., it can be properly stored in the pre-database. At the same time, when storing the original data in the pre-database, the data structure between the original data is synchronously stored, which helps to understand the internal relationship of the original data in subsequent processing and provides a reference for data conversion and integration.
[0028] S2. Establish a standard database, set up multiple standard data sets in the standard database, create the node positions of the data in a tree structure in each standard data set, and encode the data according to the node positions; here, the tree structure is a multi-way tree, including a root node, intermediate nodes, and leaf nodes; among them, the root node and intermediate nodes store the common content of the lower part of the node, and the leaf nodes store the individual content of the node; each intermediate node includes 1 parent node and any number of child nodes; the root node is the unique identifier related to the described entity; the root node, intermediate nodes, and leaf nodes are all independent physical data tables; the number of data entries of the root node, intermediate nodes, and leaf nodes under the same data set is the same, and the data relationships correspond one by one. Each node includes several entity attributes, and the entity attributes are represented in text form including entity type and entity description.
[0029] The tree structure can clearly show the hierarchical relationship between the data, providing support for the orderly organization and efficient retrieval of the data. Each node corresponds to specific data, and its position determines the level and order in the entire data system, while the encoding based on the node position becomes the key identifier for subsequent data association and processing.
[0030] Encode the data structure according to the node positions, and the encoding follows the following principles: 1. In the same standard database, the lengths of the data structure encodings are the same, and the encoding bits use decimal numbers; 2. The nodes under the same parent node are sorted in sequence from left to right, and naturally increase on this layer of data bits, with a step size of 1; 3. The data bits of each node are concatenated in sequence from the root node to this node in a top-down order, and all data bits lower than the level of this node are set to 0.
[0031] In addition, the root nodes, intermediate nodes, and leaf nodes of the multi-way tree of each standard data set form independent physical data tables according to the association relationship; the number of data entries of the root nodes, intermediate nodes, and leaf nodes under the same data set is the same, and the data relationships correspond one by one. This design facilitates data storage, management, and query.
[0032] All nodes are represented by vectors, including: using the Tf-idf method to extract the vector features of the text attributes of each node and perform vector representation; Calculate all node vectors according to the following formula to complete the normalization of all node vectors: ; where X2 represents the vector of the current node after normalization, and x1 to xn are the original vectors of the child nodes under the current node.
[0033] S3. According to the node positions, associate the data according to the preset coding rules, and update the coding of the associated data according to the associated node positions; perform data structure coding according to the node positions, and the coding follows the following principles: In the same standard database, the data structure coding lengths are the same, and the coding bits use decimal; The data structure coding is composed of the data bits and the hierarchy bits spliced together. Define the coding length M, and take the maximum data hierarchy number N in the standard database. Then the hierarchy bit is [Lg(N)] + 1, and the data bit is M - ([Lg(N)] + 1); This design makes the coding take into account both the data hierarchy and the specific identification information. For example, if the maximum data hierarchy number N is 10, the hierarchy bit is [Lg(10)] + 1 = 2. If the coding length M is set to 12, then the data bit is 12 - 2 = 10.
[0034] The data bits are divided into data bits of each hierarchy from left to right according to the hierarchy number N, and the length of each layer of data bits is [(M - ([Lg(N)] + 1)) / N]; The nodes under the same parent node are sorted in order from left to right, and naturally increase on this layer of data bits with a step size of 1. Such a division ensures that there is a corresponding coding space for each layer of data and the distribution is relatively reasonable. Taking the above example, the length of each layer of data bits is (12 - 2) / 10 = 1, and in actual applications, it will be adapted by rounding and other methods.
[0035] Further preferably, the data bits of each node are spliced in order from top to bottom for all nodes from the root node to this node, and the data bits below the hierarchy of this node are all set to 0.
[0036] S4. Obtain the user's data processing request, and perform hierarchical processing on the node positions of the relevant data according to the user's processing request.
[0037] As Figure 3 shown, in S4, when the user's data processing request is data splicing, prepare N - 1 data buffer areas, perform data pre - reading operations, place the data of the root node and each intermediate node into the corresponding - level data buffer areas respectively, according to the encoding of each layer, adopt the pre - order traversal method to read the left child node of each node starting from the root node until there is no left child node, and splice each piece of data in the data table from top to bottom in the order of the passing nodes; use the hash method to generate a 256 - bit data fingerprint for each piece of data as the unique identifier of the data.
[0038] As Figure 2 shown, it specifically includes the following steps: S411. Generate a request code according to the data processing request, and perform a search on the standard data set structure by adopting the pre - order breadth - first traversal method; S412. Obtain the data content of the root node and place it into the data buffer area; S413. Read all the direct child nodes under the current node N1 in the order from left to right. When the read node matches the request code, record the node data content into the data buffer area; S414. Determine whether there is a data bit + 1 of the current node. If it exists, perform the next operation. If not, return to the previous level; S415. Repeatedly perform the operations of S413 - S414, and each time, based on the latest node recorded in the data buffer area, read its direct child nodes; S416. When all N - 1 positions in the data buffer area are filled, generate spliced data, and use the hash method to generate a 256 - bit data fingerprint for each piece of data as the unique identifier of the data.
[0039] Further preferably, in S4, when the user's data processing request is data cropping, In the same standard database, dock the data set at the lower - level node P and the data set at the upper - level node Q of the data to be cropped, and connect the lower - level node P of the data set A to be cropped under the Q node of the data set B; Connecting the lower - level node P of the cropped data set A under the Q node of the data set B includes; 1. Identify the nodes to be cropped: Identify the nodes to be cropped except the root node and all their child nodes; If the node to be cropped contains multiple child nodes, aggregate the vectors of all child nodes into the vector of the node to be cropped by averaging; 2. Select alternative placement points Traverse the entire tree structure to find all nodes at the same level or higher levels as the node to be cut as candidate nodes; Let the node to be cut be A and the candidate node be B. According to Calculate the similarity score between the node to be cut and the candidate position node; Define a similarity threshold, take the candidate nodes exceeding the threshold as alternative placement points, and sort them according to the scores; All intermediate nodes and leaf nodes under the P node are transferred to under the Q node following the P node; The mutual relationships of the nodes under the P node remain unchanged; the access rights, security levels, encryption levels, and desensitization levels of the nodes under the P node all meet the definitions of the Q node.
[0040] After data is cut or spliced, the data association relationship is re-established. At this time, in the standard database, the data tables of the root node, intermediate nodes, and child nodes are automatically updated; Obtain the data set and encoding at the upper-level node Q of the cut data, and re-encode the encodings of the lower-level node P and each node under the P node according to the relationships of other nodes at the upper-level node Q and the preset encoding rules.
[0041] Further preferably, when data is spliced and cut, to prevent possible damage to the integrity of the tree structure, the following measures are taken: 1. Establish integrity constraints for the tree structure: 1.1 Node uniqueness constraint: Let the node set be S. For the node identification field, set a unique index for it to ensure that for any two nodes n1, n2 ∈ S (n1 ≠ n2), at the database design level, set a unique index for the node identification field to ensure that the identification of each node in the tree structure is unique within the entire structure, avoid the situation where the same data appears at multiple different node positions, and ensure data consistency. For example, in the database table of the enterprise organizational structure tree structure, set a unique index for the department number field to prevent the situation where two department numbers are the same.
[0042] 1.2 Parent-child relationship integrity: Use the foreign key constraint of the database to maintain the parent-child relationship. For each non-root node, its record points to the record of its parent node through a foreign key. Let the node set be S. For any non-root node element n ∈ S, ensure that each non-root node has and only has one parent node, and the root node has no parent node. Taking the file system tree structure database as an example, in the table storing folder information, the foreign key field of the non-root directory folder record points to the record of its parent folder, and the foreign key field of the root directory record is empty, so as to maintain the hierarchical logic of the tree structure and prevent problems such as node hanging or circular references.
[0043] 1.3 Hierarchical Continuity: In the program logic of data insertion or update operations, add hierarchical continuity checks. Let the set of levels be L = {1, 2, 3,..., k}. When inserting a new node, assume the level of the node to be inserted is i which belongs to the set L. Confirm whether its upper-level node exists through database query. If it does not exist, prevent the insertion operation. For example, in the product classification tree structure, when inserting a third-level specific product classification node, first check whether the second-level major category nodes and the first-level general category nodes exist to ensure the continuity between levels in the tree structure and no gaps.
[0044] 2. Verify the legality of the splicing target position: 2.1 Hierarchical Range Verification: Set a global variable in the program code to record the maximum level of the tree structure. When performing data splicing operations involving the target position, obtain the level information of the target position and compare it with the maximum level. If the maximum level of the tree structure is 5 and the target position level is 6, it is determined to be illegal, terminate the splicing operation and return an error message to the user.
[0045] 2.2 Parent Node Existence Verification: When performing data splicing, for non-root target positions, according to the parent node identifier of the target position, query in the database whether there is a corresponding parent node record. For example, in the family tree structure, if a new member node is to be added to a certain position, first query whether the specified parent node exists in the genealogy database. If it does not exist, return an error message to prevent illegal node connections.
[0046] 2.3 Node Type Compatibility Verification: When designing the database, set a type field for each node to record the node type (such as category node, product node, etc.). During data splicing operations, according to the predefined connection rules for different types of nodes, perform legality verification based on the type fields of the source node and the target node. For example, in the e-commerce product classification tree structure, a category node can be connected to a sub-category node or a product node, but a product node cannot be connected to a category node anymore. When the splicing target position involves connecting different types of nodes, verify the type fields to ensure the operation is legal.
[0047] As Figure 3 shown, the present invention also provides a multi-source data custom hierarchical processing system for implementing the above multi-source data custom hierarchical processing method, including: A data storage module for establishing a pre-database to store the original data obtained from various channels; A database establishment module for establishing a standard database, setting up multiple standard data sets in the standard database, creating node positions for the data in a tree structure in each standard data set, and encoding the data according to the node positions; A data encoding module, according to the node positions, associates the data according to a preset encoding rule, and the encoded data is updated according to the node positions after association. A hierarchical processing module, obtains a user's data processing request, and hierarchically processes the node positions of relevant data according to the user's processing request.
[0048] The present invention also provides an electronic device, including: a memory storing computer program instructions; a processor, when the computer program instructions are executed by the processor, implementing the steps of the above-mentioned method for custom hierarchical processing of multi-source data.
[0049] The present invention also provides a computer-readable storage medium, which is used to store instructions. When the stored instructions run on a computer, the computer is caused to execute the steps of the above-mentioned method for custom hierarchical processing of multi-source data.
[0050] Obviously, the above embodiments are only examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A custom hierarchical processing method for multi-source data, characterized in that, It includes the following steps: S1. Establish a pre-database to store the original data obtained from various channels; S2. Establish a standard database, set up multiple standard data sets in the standard database, create the node positions of the data in a tree structure in each standard data set, and encode the data according to the node positions; S3. According to the node positions, associate the data according to the preset encoding rules, and update the encoding of the associated data according to the associated node positions; The preset encoding rules include: The data structure encoding is composed of the data bit and the level bit spliced together. Define the encoding length M. According to the multi-way tree level structure, take the maximum data level number N in the standard database, then the level bit is [Lg(N)] + 1, and the data bit is M - ([Lg(N)] + 1); The data bit is divided into data bits of each level from left to right according to the level number N, and the length of each level of data bit is defined as [(M - ([Lg(N)] + 1)) / N]; S4. Obtain the user's data processing request, and hierarchically process the node positions of the relevant data according to the user's processing request.
2. The custom hierarchical processing method for multi-source data according to claim 1, characterized in that, When the pre-database stores the original data, the data structure between the original data is stored synchronously.
3. The custom hierarchical processing method for multi-source data according to claim 1, characterized in that, In S2, when creating the node positions of the data in a tree structure in each standard data set, the tree structure is a multi-way tree, including a root node, intermediate nodes, and leaf nodes; among them, the root node and intermediate nodes store the common content of the part below this node, and the leaf nodes store the individual content of this node; each intermediate node includes 1 parent node and any number of child nodes; the root node is the unique identifier related to the described entity; each node includes several entity attributes, and the entity attributes are represented in text form including entity type and entity description.
4. The custom hierarchical processing method for multi-source data according to claim 1, characterized in that, In S2, according to the node positions, perform data structure encoding, and the encoding follows the following principles: In the same standard database, the data structure encoding lengths are the same, and the encoding bits use decimal; The nodes under the same parent node are sorted in order from left to right, and naturally increase on this layer of data bits, with a step size of 1; The data bit of each node is spliced in sequence from the root node to all the nodes of this node in a top-down order, and the data bits below the level of this node are all set to 0.
5. The custom hierarchical processing method for multi-source data according to claim 4, characterized in that, The root nodes, intermediate nodes, and leaf nodes of the multi-way tree of each standard data set form independent physical data tables according to the association relationship; Under the same data set, the number of data of the root node, intermediate nodes, and leaf nodes is the same, and the data relationships correspond one by one; Use vector representation for all nodes, including: using the Tf-idf method to extract the vector features of the text attributes of each node and perform vector representation; Calculate all node vectors according to the following formula to complete the normalization of all node vectors: ; Among them, X2 represents the vector after normalization of the current node, and x1 to xn are the original vectors of the child nodes under the current node.
6. The custom hierarchical processing method for multi-source data according to claim 5, characterized in that, In S4, when the user's data processing request is data splicing, prepare N - 1 data buffer areas, perform data pre-reading operations, place the data of the root node and each intermediate node into the corresponding level data buffer areas respectively, and perform data splicing to generate each complete data, which specifically includes the following steps: S411. Generate a request code according to the data processing request, and perform a pre-order breadth-first traversal search on the physical data table of the standard data set; S412. Obtain the root node data content into the data buffer; S413. Read all direct child nodes under the current node N1 in the order from left to right. When the read node matches the request code, record the node data content into the data buffer; S414. Determine whether there is a data bit +1 of the current node. If so, perform the next operation. If not, return to the previous level; S415. Repeatedly perform the operations of S413 - S414. Each time, based on the latest node recorded in the data buffer, read its direct child nodes; S416. When all N - 1 positions in the data buffer are filled, generate spliced data, and use the hashing method to generate 256-bit data fingerprints for each piece of data as the unique identifier of the data.
7. The custom hierarchical processing method for multi-source data according to claim 1, characterized in that, In S4, when the user's data processing request is data cropping or splicing, In the same standard database, dock the data set at the lower-level node P and the data set at the upper-level node Q of the data to be cropped, and connect the lower-level node P of the cropped A data set under the Q node of the B data set, including; 1. Identify the node to be cropped: Identify the node to be cropped other than the root node and all its child nodes; If the node to be cropped contains multiple child nodes, aggregate the vectors of all child nodes into the vector of the node to be cropped by averaging; 2. Select alternative placement points Traverse the entire tree structure to find all nodes at the same level or higher levels as the node to be cropped as candidate nodes; Let the node to be cut be A and the candidate node be B. According to Calculate the similarity score between the node to be cut and the candidate position node; Define a similarity threshold, and use the candidate nodes exceeding the threshold as alternative placement points and sort them according to the scores; All intermediate nodes and leaf nodes under the P node are transferred to under the Q node following the P node; The mutual relationships of the nodes under the P node remain unchanged; the access rights, security levels, encryption levels, and de-sensitization levels of the nodes under the P node satisfy the definitions of the Q node.
8. The custom hierarchical processing method for multi-source data according to claim 1, wherein, In S4, when data is spliced and cropped, it also includes: Verify the legality of the splicing target position. The legality verification of the splicing target position includes the following process: Obtain the hierarchical information of the target position and compare it with the maximum level. When the target position level is greater than the maximum level, it is judged as illegal; Judge whether the target position is a parent node: If the target position is a non-root node, query whether there is a corresponding parent node record in the database according to the parent node identifier of the target position; if there is no parent node record, it is judged as illegal.
9. The custom hierarchical processing method for multi-source data according to claim 1, wherein, In S4, it also includes judging whether the spliced or cropped tree structure satisfies the integrity constraints, including judging whether the tree structure satisfies after splicing or cropping: a. Node uniqueness constraint; Search the node identifier field in the multi-way tree universe to ensure that the identifier of each node in the tree structure is unique within the entire tree structure; b. Parent-child relationship integrity constraint; For each non-root node, judge the record pointing to its parent node through the foreign key in the standard database to ensure that each non-root node has and only has one parent node; c. Hierarchical continuity constraint: Let the set of levels be L = {1, 2, 3,..., k}. When inserting a new node, assume that the level of the node to be inserted is i, which belongs to the set L. Check whether its upper-level node exists through database query. If it does not exist, the insertion operation is blocked.
10. A custom hierarchical processing system for multi-source data, wherein, A method for custom hierarchical processing of multi-source data according to any one of claims 1-9 above, comprising: A data storage module for establishing a pre-database to store the original data obtained from various channels; A database establishment module for establishing a standard database, setting up multiple standard data sets in the standard database, creating the node positions of the data in a tree structure in each standard data set, and encoding the data according to the node positions; A data encoding module for associating the data according to the node positions according to a preset encoding rule, and updating the encoding of the associated data according to the associated node positions; the preset encoding rule includes the following content: the data structure encoding is composed of the data bit and the level bit. Define the encoding length M. According to the multi-way tree hierarchical structure, take the maximum data level number N in the standard database, then the level bit is [Lg(N)] + 1, and the data bit is M - ([Lg(N)] + 1); The data bit is divided into data bits of each level from left to right according to the level number N, and the length of each level data bit is [(M - ([Lg(N)] + 1)) / N]; A hierarchical processing module for obtaining the user's data processing request and performing hierarchical processing on the node positions of the relevant data according to the user's processing request.
Citation Information
Patent Citations
Method and device for establishing reverse index of internet of things intelligent equipment
CN106126646A
Picture-based poem making method, device and equipment and storage medium
CN115080786A
Method, device and equipment for merging and tracing multi-source tree data and chain data
CN115827725A
Tree coding method and system
CN115994141A
Tree structure data dynamic management method based on node hierarchical attributes and recursive updating
CN119719048A