Method for dynamically appending and managing JSON data based on target keywords
By using syntactic dependency analysis and semantic tagging mapping, the problems of semantic inconsistency and insufficient adaptability in JSON data management are solved. This enables accurate semantic positioning and dynamic appending of JSON data, improving the accuracy and efficiency of data operations and enhancing adaptability and intelligence.
Patent Information
- Application Number
- CN202511241072.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing technologies lack a deep understanding of the semantic hierarchy of keywords in JSON data management, and cannot effectively capture the semantic relationships between keywords. This leads to semantic inconsistencies when appending data, and makes it difficult to adaptively adjust to achieve accurate positioning and intelligent organization. Furthermore, it cannot accurately identify explicit and implicit reference relationships between JSON data nodes, affecting data consistency maintenance and efficient access.
By converting target keywords into feature vectors through syntactic dependency analysis, generating a keyword affinity matrix, performing adaptive classification vector clustering to obtain keyword cluster families, assigning semantic labels to cluster families, establishing a mapping relationship between semantic labels and JSON data nodes, constructing a semantic addressing table, locating target nodes and extracting explicit and implicit reference relationships, constructing a dependency topology graph, calculating the association strength between nodes and allocating data blocks, and achieving accurate semantic positioning and dynamic appending of JSON data.
It improves the accuracy and efficiency of JSON data manipulation, reduces the error rate, achieves modular management, enhances data maintainability and scalability, and strengthens the ability to understand user intent, adaptability, and intelligence.
Smart Images

Figure CN120744117B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a JSON data dynamic appending and management method based on target keywords. BACKGROUND
[0002] With the rapid development of Internet technology, JSON data format has become the mainstream format of current data exchange and storage due to its lightweight, easy to read and write, and cross-language characteristics. In large-scale application systems, JSON data structure often needs to be dynamically appended and managed according to business needs, especially in complex data relationship processing, big data analysis, and knowledge graph construction scenarios. Effective semantic recognition and dynamic management of JSON data are particularly important.
[0003] Traditional JSON data management usually adopts simple key-value pair matching and path navigation methods. This method lacks semantic understanding ability when dealing with complex JSON structures, and it is difficult to identify implicit data association relationships. Although there have been some studies on JSON data management based on keywords, there are still significant deficiencies.
[0004] The prior art still has some defects in the dynamic appending and management of JSON data based on keywords. The prior art lacks a deep understanding of the semantic hierarchy of keywords, cannot effectively capture the semantic association between keywords, and leads to semantic inconsistency problems when appending data. Traditional JSON data management methods usually rely on predefined data structure patterns, and lack adaptive adjustment ability when facing dynamic business needs, making it difficult to achieve precise positioning and intelligent organization of data. The prior art fails to establish an effective data dependency analysis mechanism, which cannot accurately identify explicit and implicit reference relationships between JSON data nodes, thereby affecting data consistency maintenance and efficient access, especially in complex nested structure JSON data processing. SUMMARY
[0005] The embodiments of the present application provide a JSON data dynamic appending and management method based on target keywords, which can solve the problems in the prior art.
[0006] In a first aspect, the embodiments of the present application provide a JSON data dynamic appending and management method based on target keywords, comprising:
[0007] receiving target keywords and JSON data;
[0008] converting the target keywords into feature vectors using syntax dependency analysis, generating a keyword affinity matrix through cosine similarity calculation, and performing adaptive classification vector clustering operation based on the keyword affinity matrix to obtain a keyword clustering family;
[0009] assigning a semantic identifier to the keyword cluster family, establishing a mapping relationship between the semantic identifier and a node in the JSON data, and generating a semantic addressing table;
[0010] locating a target node in the JSON data according to the semantic addressing table, extracting an explicit reference relationship and an implicit reference relationship to construct a dependency topology graph, calculating a correlation strength between nodes, constructing a data block according to the correlation strength, and assigning a globally unique block identifier to the data block;
[0011] retrieving a candidate semantic identifier based on the target keyword in the semantic addressing table, determining a target data block and a write hierarchical position, writing the JSON data to be appended to the target data block, updating a node correlation relationship in the dependency topology graph, and recording block change information;
[0012] updating the semantic addressing table and the dependency topology graph according to the block change information.
[0013] In an optional embodiment, the target keyword is converted into a feature vector using syntax dependency analysis, a keyword affinity matrix is generated by cosine similarity calculation, and adaptive classification vector clustering operation is performed based on the keyword affinity matrix to obtain the keyword cluster family, including:
[0014] word segmentation is performed on the target keyword to obtain a word sequence, part-of-speech tagging is used to obtain part-of-speech features of each word in the word sequence, a syntax dependency tree is constructed based on the part-of-speech features and position information, syntax relationship features are extracted from the syntax dependency tree, and the part-of-speech features and the syntax relationship features are combined to generate a semantic feature matrix;
[0015] dimensionality reduction is performed on the semantic feature matrix based on orthogonal transformation, feature dimensions are sequentially selected according to variance contribution rates, and corresponding feature values are combined to form a target feature vector;
[0016] a cosine distance between any two target feature vectors is calculated to obtain a similarity value, an initial similarity matrix is constructed according to the similarity value, and element values in the initial similarity matrix are normalized to determine an affinity matrix;
[0017] the mean and variance of each row element in the affinity matrix are counted, a dynamic threshold is set based on the mean and variance, element values less than the dynamic threshold are set to zero to obtain an optimized affinity matrix;
[0018] a Laplacian matrix of the optimized affinity matrix is calculated, a feature vector of the Laplacian matrix is extracted to construct a feature space, and adaptive classification vector clustering operation is performed in the feature space to obtain the keyword cluster family.
[0019] In an alternative embodiment, the adaptive clustering operation of the classification vectors in the feature space to obtain the keyword cluster family comprises:
[0020] Mapping the keywords in the feature space as sample points;
[0021] Establishing a search radius for each sample point, counting the number of samples within the search radius to obtain a local density value; calculating the Euclidean distance between any two sample points, for each target sample point, selecting the smallest Euclidean distance among all sample points with a local density value greater than the target sample point to determine the contrast sample point, and recording the Euclidean distance between the target sample point and the contrast sample point as the distance value;
[0022] Calculating the product of the local density value and the distance value to obtain a density index, sorting the density index in descending order, setting a density index threshold value according to the numerical distribution of the density index, and taking the sample points greater than the density index threshold value as initial clustering points;
[0023] Building a classifier with the initial clustering points as the classification vectors of the classifier; inputting sample points one by one, calculating the similarity between the sample points and each classification vector, and selecting the classification vector with the highest similarity as the matching vector;
[0024] Based on the position of the current input sample point and the matching vector, adjusting the position of the matching vector by a preset distance ratio, and adjusting other classification vectors adjacent to the matching vector;
[0025] Repeating the operation until the classification vectors are stable and unchanged, grouping the sample points associated with the same classification vector into a class to obtain the keyword cluster family.
[0026] In an alternative embodiment, the keyword cluster family is assigned a semantic identifier, and a mapping relationship between the semantic identifier and the nodes in the JSON data is established to generate a semantic addressing table, which comprises:
[0027] Counting the word frequency of the keywords in the keyword cluster family, extracting the part-of-speech features and syntactic dependency relations of the keywords to construct semantic vectors, calculating the semantic topic category of the keyword cluster family according to the semantic vectors, and assigning a semantic identifier to the keyword cluster family according to the preset hierarchical semantic coding rules;
[0028] Parsing the hierarchical structure of the JSON data to obtain JSON nodes, extracting the key names and data types of the JSON nodes, determining the parent-child relationship and sibling relationship between the JSON nodes, constructing a tree structure adjacency matrix of the JSON nodes, and calculating the relationship distance between the JSON nodes according to the tree structure adjacency matrix;
[0029] Calculate a text similarity of the semantic identifier and key name information of the JSON node, construct a semantic environment vector based on a relationship distance between the JSON nodes, and obtain a semantic correlation matrix by weighting and fusing the text similarity and the semantic environment vector;
[0030] Filter the semantic correlation matrix by setting a mapping threshold, eliminate redundant mapping relationships based on a transitive closure, process conflicts of one-to-many mapping and many-to-one mapping, and determine a mapping relationship between the semantic identifier and the JSON node;
[0031] Construct a semantic identifier index and a node path index to form a bidirectional mapping structure based on the mapping relationship, and generate a semantic addressing table according to the bidirectional mapping structure.
[0032] In an optional embodiment, locate a target node in the JSON data according to the semantic addressing table, extract an explicit reference relationship and an implicit reference relationship to construct a dependency topology graph, calculate an association strength between nodes, and construct data blocks according to the association strength, and assign a globally unique block identifier to the data blocks, including:
[0033] Locate a target node according to the semantic addressing table, and extract a reference relationship of the target node, including an explicit reference relationship and an implicit reference relationship, wherein the key-value correspondence relationship in the target node is extracted to determine foreign key references and ID references as the explicit reference relationship, and the co-occurrence frequency of node attribute values and the similarity of naming rules are calculated to determine the implicit reference relationship;
[0034] Construct a dependency topology graph by taking the target node as a graph node and the reference relationship as a connection edge, calculate a transitive closure matrix of the dependency topology graph to identify a node subset of circular references, and optimize the dependency topology graph by using a minimum spanning tree algorithm;
[0035] Calculate a distance association factor by calculating a shortest path length between nodes in the dependency topology graph, obtain a reference association factor by counting a citation frequency, calculate a feature association factor by calculating a data type matching degree and a numerical interval intersection, and calculate an update association factor by calculating a correlation coefficient of a node data update time sequence, and obtain an association strength between nodes by weighted summation of the distance association factor, the reference association factor, the feature association factor, and the update association factor;
[0036] Divide a node set into candidate blocks according to a preset association strength threshold, calculate a clustering coefficient by calculating a ratio of an average degree to a maximum degree of nodes in the candidate block, merge and optimize the candidate blocks according to the clustering coefficient, calculate a number of shared nodes between the candidate blocks, and assign the shared nodes to a block with the largest clustering coefficient to obtain data blocks;
[0037] Generate a globally unique block identifier according to the association strength, the clustering coefficient, and the reference relationship of the data blocks.
[0038] In an alternative embodiment, based on the target keyword, the candidate semantic identifiers are retrieved in the semantic addressing table, the target data block and the write hierarchical position are determined, the JSON data to be added is written into the target data block, the node association relationship in the dependency topology graph is updated, and the block change information is recorded, including:
[0039] According to the target keyword, a set of candidate semantic identifiers is retrieved in the semantic addressing table, the time sequence activity and the reference heat of each semantic identifier in the set of candidate semantic identifiers are calculated, the priority score of the semantic identifier is calculated by weighting, and the semantic identifier with the highest priority score is selected to determine the target data block;
[0040] The JSON data to be added is analyzed and structured to build an attribute association tree, the inheritance depth and the association width between nodes are calculated based on the attribute association tree, the hierarchical position of data writing is determined in combination with the inheritance depth and the association width, and the JSON data to be added is written into the target data block;
[0041] A bidirectional reference index is built for the new node in the dependency topology graph, the circular reference path between nodes is identified based on the bidirectional reference index, the association relationship with the weakest reference strength in the circular reference path is marked as a soft connection, the soft connection only retains the one-way reference between nodes and is automatically disassociated after the expiration of the invalidation timestamp, the closed loop path in the circular reference path is blocked through the soft connection, the node association relationship in the dependency topology graph is updated, and the block change information is recorded.
[0042] In a second aspect of the embodiment of the application, a JSON data dynamic addition and management system based on a target keyword is provided, including:
[0043] The first unit is configured to receive a target keyword and JSON data;
[0044] The second unit is configured to convert the target keyword into a feature vector using syntax dependency analysis, generate a keyword affinity matrix through cosine similarity calculation, and perform adaptive classification vector clustering operation based on the keyword affinity matrix to obtain a keyword clustering family;
[0045] The third unit is configured to assign a semantic identifier to the keyword clustering family, establish a mapping relationship between the semantic identifier and a node in the JSON data, and generate a semantic addressing table;
[0046] The fourth unit is configured to locate a target node in the JSON data according to the semantic addressing table, construct a dependency topology graph by extracting explicit reference relationships and implicit reference relationships, calculate the association strength between nodes, construct data blocks according to the association strength, and assign a globally unique block identifier to the data blocks;
[0047] The fifth unit is configured to search for a candidate semantic identifier in the semantic addressing table based on a target keyword, determine a target data block and a write level position, write the JSON data to be added into the target data block, update a node association relationship in the dependency topology graph, and record block change information.
[0048] The sixth unit is configured to update the semantic addressing table and the dependency topology graph according to the block change information.
[0049] In a third aspect, an electronic device is provided, including:
[0050] a processor;
[0051] a memory for storing processor-executable instructions;
[0052] The processor is configured to invoke the instructions stored in the memory to execute the method described above.
[0053] In a fourth aspect, a computer-readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.
[0054] In the embodiment of the present application, by introducing syntax dependency analysis and semantic identifier mapping, accurate semantic positioning and dynamic addition of JSON data are realized, the accuracy and efficiency of data operation are greatly improved, the operation error rate caused by complex data structure is significantly reduced, the association between data can be intelligently identified and reasonably organized; the dependency topology graph and the data block are constructed, the modular management of JSON data is realized, the problems of complex large-scale JSON data structure and difficult maintenance in the traditional method are solved, the maintainability and expansibility of data are improved, and the traceability of data operation is realized through the block change record mechanism; adaptive classification vector clustering and keyword affinity matrix realize semantic management of JSON data, effectively solve the problem of inaccurate keyword matching in the traditional method, improve the understanding ability of user intent, make the data addition more in line with the actual business demand, and greatly enhance the adaptability and intelligence. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 A flowchart of a JSON data dynamic addition and management method based on a target keyword according to an embodiment of the present application is shown in the figure.
[0056] Figure 2 A keyword clustering logic flowchart is shown in the figure. DETAILED DESCRIPTION
[0057] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0058] The technical solutions of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and some embodiments can not be described again for the same or similar concepts or processes.
[0059] Figure 1 The flowchart of the method for dynamically appending and managing JSON data based on target keywords in the embodiments of the present application is shown in FIG. 1, which comprises the following steps. Figure 1 The method comprises the following steps.
[0060] Receiving a target keyword and JSON data.
[0061] Converting the target keyword into a feature vector by syntax dependency analysis, generating a keyword affinity matrix by cosine similarity calculation, and performing adaptive classification vector clustering operation based on the keyword affinity matrix to obtain a keyword clustering family.
[0062] Assigning a semantic identifier to the keyword clustering family, establishing a mapping relationship between the semantic identifier and a node in the JSON data, and generating a semantic addressing table.
[0063] Locating a target node in the JSON data according to the semantic addressing table, extracting explicit reference relationships and implicit reference relationships to construct a dependency topology graph, calculating the correlation strength between nodes, constructing data blocks according to the correlation strength, and assigning a globally unique block identifier to the data blocks.
[0064] Retrieving a candidate semantic identifier based on the target keyword in the semantic addressing table, determining a target data block and a write hierarchical position, writing the JSON data to be appended into the target data block, updating the node correlation relationship in the dependency topology graph, and recording block change information.
[0065] Updating the semantic addressing table and the dependency topology graph according to the block change information.
[0066] In an optional embodiment, converting the target keyword into a feature vector by syntax dependency analysis, generating a keyword affinity matrix by cosine similarity calculation, and performing adaptive classification vector clustering operation based on the keyword affinity matrix to obtain a keyword clustering family comprises the following steps.
[0067] The target keyword is segmented to obtain a word sequence, the part-of-speech of each word in the word sequence is obtained by part-of-speech tagging, a syntax dependency tree is constructed based on the part-of-speech and position information, a syntactic relationship feature is extracted from the syntax dependency tree, and the part-of-speech feature and the syntactic relationship feature are combined to generate a semantic feature matrix;
[0068] The semantic feature matrix is reduced in dimension based on an orthogonal transformation, and the feature dimensions are sequentially selected according to the variance contribution rate, and the corresponding feature values are combined to form a target feature vector;
[0069] The cosine distance between any two target feature vectors is calculated to obtain a similarity value, and an initial similarity matrix is constructed according to the similarity value, and the element values in the initial similarity matrix are normalized to determine an affinity matrix;
[0070] The mean and variance of each row element in the affinity matrix are counted, a dynamic threshold is set based on the mean and variance, the element values less than the dynamic threshold are set to zero to obtain an optimized affinity matrix;
[0071] The Laplacian matrix of the optimized affinity matrix is calculated, the feature vector of the Laplacian matrix is extracted to construct a feature space, and adaptive classification vector clustering operations are performed in the feature space to obtain a keyword clustering family.
[0072] In a specific embodiment, a set of target keywords is received as input, such as “smartphone screen”, “mobile display technology”, “high-definition display screen”, and other keywords related to the technical field. For each target keyword, use a segmentation tool (such as a conditional random field-based segmentation algorithm) to split it into a word sequence. Taking “smartphone screen” as an example, the segmentation result is [“smart”, “phone”, “screen”].
[0073] The word sequence after segmentation is processed by part-of-speech tagging, and a part-of-speech tagging method based on a hidden Markov model is used to assign a part-of-speech label to each word. In the above example, the tagging result is [“smart / adj”, “phone / n”, “screen / n”], indicating that “smart” is an adjective, and “phone” and “screen” are nouns.
[0074] Based on the part-of-speech tagging result and the position information of the word in the sequence, a syntax dependency tree is constructed. A transition-based dependency analysis algorithm is used to identify the dependency relationship between words. In the example, “smart” modifies “phone”, forming an attribute relationship (ATT), and “phone” and “screen” form an attribute relationship (ATT), with “screen” as the core word (HED).
[0075] From the syntactic dependency tree, syntactic relation features are extracted, including the dependency relation type and its direction between words. In the example, the extracted syntactic relation features include "smartphone: ATT" and "screen: ATT". The part-of-speech features and syntactic relation features are combined to generate a semantic feature matrix. For the example keyword, the formed feature matrix contains three rows (corresponding to three words), and each row contains part-of-speech features (one-hot encoding) and syntactic relation features (one-hot encoding of dependency relation type).
[0076] To reduce the feature dimension and extract key information, the semantic feature matrix is subjected to orthogonal transformation dimension reduction processing. Specifically, the principal component analysis (PCA) method is used. The covariance matrix of the feature matrix is calculated, and its eigenvalues and eigenvectors are solved. For example, for a 10x10 feature matrix, 10 eigenvalues and their corresponding eigenvectors can be obtained.
[0077] According to the size of the eigenvalues, the variance contribution rate of each eigenvalue is calculated. Assuming that the variance contribution rates of the first three eigenvalues are 45%, 30%, and 15%, respectively, and the cumulative contribution rate reaches 90%, the first three feature dimensions are selected. The corresponding three eigenvectors are combined to form the target eigenvector. For "smartphone screen" in the example, a three-dimensional eigenvector [0.75, 0.42, -0.18] can be obtained.
[0078] For all target keywords, their eigenvectors are generated respectively, and the cosine similarity between any two eigenvectors is calculated. The cosine similarity is obtained by calculating the cosine value of the angle between two vectors, with a value range of [-1, 1], and the larger the value, the more similar the two vectors. For example, the cosine similarity between the eigenvectors of "smartphone screen" and "mobile display technology" is 0.82, indicating that the two keywords have high semantic relevance.
[0079] An initial similarity matrix S is constructed according to the calculated similarity values. For N keywords, S is an N x N matrix, where S[i,j] represents the similarity between the i-th and j-th keywords. The elements in matrix S are normalized by dividing each element value by the maximum value in the matrix to obtain the normalized affinity matrix A.
[0080] To optimize the affinity matrix, the mean μ and variance σ of each row element in the affinity matrix A are calculated. For example, the mean μ of the first row element is 0.45, and the variance σ is 0.08. Based on the mean and variance, a dynamic threshold value is set, and the threshold value is calculated as τ = μ - α x σ, where α is an adjustment parameter with a value of 1.5. For the first row in the example, the dynamic threshold value τ = 0.45 - 1.5 x 0.08 = 0.33.
[0081] Set the element values in the affinity matrix A that are less than the corresponding dynamic threshold to 0 to obtain an optimized affinity matrix A'. This step can effectively filter out weak correlations and retain strong correlations, improving clustering accuracy.
[0082] Calculate the degree matrix D of the optimized affinity matrix A', where D is a diagonal matrix with diagonal elements being the sum of the corresponding row elements of A'. According to the degree matrix D and the affinity matrix A', calculate the Laplacian matrix L = D - A'. Extract the eigenvectors of L, sort them from small to large according to the eigenvalues, and select the top K eigenvectors to construct a feature space, where K is the expected number of clusters.
[0083] In the constructed feature space, perform the K-means clustering algorithm to adaptively cluster the keywords. In specific implementation, randomly initialize K cluster centers, and iteratively perform the assignment and update steps until convergence. The assignment step assigns each keyword to the nearest cluster center, and the update step recalculates the center point of each cluster.
[0084] After clustering, the keywords in the example may be divided into three cluster families: {“smartphone screen”, “mobile display technology”, “high-definition display screen”}, {“artificial intelligence algorithm”, “machine learning model”}, and {“wireless charging technology”, “fast charging solution”}. These cluster families represent different technical fields or topics.
[0085] In an alternative embodiment, performing adaptive classification vector clustering operations in the feature space to obtain keyword cluster families includes:
[0086] Mapping the keywords in the feature space to sample points;
[0087] Establishing a search radius for each sample point, counting the number of samples within the search radius to obtain a local density value; calculating the Euclidean distance between any two sample points, for each target sample point, selecting the sample point with the smallest Euclidean distance from all sample points with a local density value greater than the target sample point, determining the comparison sample point, and recording the Euclidean distance between the target sample point and the comparison sample point as the distance value;
[0088] Calculating the product of the local density value and the distance value to obtain a density indicator, sorting the density indicators in descending order, setting a density indicator threshold value according to the numerical distribution of the density indicators, and taking the sample points greater than the density indicator threshold value as initial cluster points;
[0089] Constructing a classifier with the initial cluster points as the classification vectors of the classifier; inputting the sample points one by one, calculating the similarity between the sample points and each classification vector, and selecting the classification vector with the highest similarity as the matching vector;
[0090] Based on the current input sample point and the position of the matching vector, the position of the matching vector is adjusted by a preset distance ratio, and other classification vectors adjacent to the matching vector are also adjusted;
[0091] The repeated execution is performed until the classification vectors are stable and unchanged, the sample points associated with the same classification vector are classified into a category, and a keyword clustering family is obtained.
[0092] In a specific embodiment, each keyword in the feature space can be represented as a vector with a specific feature dimension. In order to effectively cluster keywords, it is necessary to map these keywords to sample points in the feature space. Specifically, a word vector model can be used to convert keywords into a fixed-dimensional vector representation, such as a 300-dimensional vector space, and each keyword corresponds to a sample point in the feature space.
[0093] After the sample points are mapped, a search radius is established for each sample point to determine its local density value. The search radius can be set according to the data distribution, for example, set to 20% of the average distance between all sample points in the feature space. For each sample point, the number of other sample points falling within its search radius range is counted, which is the local density value of the sample point. For example, if the search radius of sample point P1 is 0.5, there are 7 other sample points within the radius range, then the local density value of P1 is 7.
[0094] Subsequently, the Euclidean distance between any two sample points is calculated. The Euclidean distance calculation is based on the square root of the sum of the squares of the coordinate differences of the sample points in each dimension of the feature space. For each target sample point in the feature space, among all sample points with a local density value greater than the target sample point, the sample point with the smallest Euclidean distance from the target sample point is found as the comparison sample point. The Euclidean distance between the target sample point and the comparison sample point is recorded as the distance value. For example, the local density value of sample point P2 is 5, among all sample points with a local density value greater than 5, the Euclidean distance between sample points P3 and P2 is the smallest, which is 0.8, then the distance value of P2 is 0.8.
[0095] Based on the local density value and the distance value, the density index of each sample point is calculated, which is the product of the local density value and the distance value. The density index reflects the potential of the sample point as a clustering center. All sample points are sorted in descending order according to the density index, and the numerical distribution of the density index is analyzed. Generally, the density index of the real clustering center point will be significantly higher than that of other points. By observing the distribution curve of the density index, a suitable density index threshold can be set. For example, if the density index of the first 5 sample points after sorting is 25, 23, 22, 12, and 11, and the density index of the subsequent sample points is less than 10 and uniformly distributed, the density index threshold can be set to 20, and the first 3 sample points are selected as the initial clustering points.
[0096] When constructing the classifier, these initial clustering points are taken as the classification vectors of the classifier. The classification vector represents a clustering center. The remaining sample points are input one by one, and the similarity between each sample point and each classification vector is calculated. The similarity can be calculated based on the reciprocal of the Euclidean distance, and the smaller the distance, the higher the similarity. For the input sample point, the classification vector with the highest similarity is selected as its matching vector. For example, the similarity of sample point P4 with classification vectors V1, V2 and V3 is 0.7, 0.5 and 0.3 respectively, and V1 is the matching vector of P4.
[0097] Based on the positional relationship between the current input sample point and the matching vector, the position of the matching vector is adjusted. The adjustment uses a preset distance ratio, for example 0.1, which means that the matching vector moves 10% of the distance in the direction of the current sample point. If the distance between the current sample point P5 and its matching vector V2 is 2.0, then the new position of V2 will move 0.2 units of distance in the direction of P5. At the same time, the other classification vectors adjacent to the matching vector also need to be adjusted to maintain the topological relationship between the classification vectors. The adjustment amplitude of the adjacent vectors can be set to 50% of the adjustment amplitude of the matching vector, and the direction is consistent.
[0098] The classification and adjustment process is repeated until the change in the position of all classification vectors in the last two rounds of adjustment is less than a preset threshold (such as 0.01), and it is considered that the classification vectors are stable and unchanged. At this time, all sample points associated with the same classification vector (i.e. taking the classification vector as the matching vector) are classified into a class, forming a keyword clustering family.
[0099] To illustrate with a specific case, suppose there is a keyword set {“computer”, “laptop”, “tablet”, “mobile phone”, “camera”, “printer”, “earphone”, “sound box”}, and the coordinates of the keywords after being mapped to a two-dimensional feature space by a word vector model are (1.2, 2.3), (1.5, 2.1), (1.8, 2.4), (4.2, 1.1), (4.5, 1.3), (4.7, 1.0), (3.1, 4.2), (3.4, 4.5) respectively. The search radius is set to 0.5, and the local density values of each point are calculated to be 3, 3, 3, 3, 3, 3, 2, 2 respectively. After distance value calculation and density index sorting, (“computer”, 1.2, 2.3), (“mobile phone”, 4.2, 1.1) and (“earphone”, 3.1, 4.2) are selected as the initial clustering points. After multiple rounds of classification vector position adjustment, three keyword clustering families are finally formed: {“computer”, “laptop”, “tablet”}, {“mobile phone”, “camera”, “printer”} and {“earphone”, “sound box”}, which correspond to different types of electronic devices.
[0100] As shown in FIG. 1, Figure 2 A keyword clustering logic flowchart is shown.
[0101] In an optional implementation, a semantic identifier is assigned to the keyword cluster family, a mapping relationship between the semantic identifier and the nodes in the JSON data is established, and generating a semantic addressing table includes:
[0102] Word frequency statistics are performed on the keywords in the keyword cluster family, a part-of-speech feature and a syntactic dependency relationship of the keywords are extracted to construct a semantic vector, a semantic topic category of the keyword cluster family is calculated according to the semantic vector, and a semantic identifier is assigned to the keyword cluster family in combination with a preset hierarchical semantic coding rule;
[0103] A hierarchical structure of the JSON data is parsed to obtain JSON nodes, a key name and a data type of the JSON nodes are extracted, a parent-child relationship and a sibling relationship between the JSON nodes are determined, a tree structure adjacency matrix of the JSON nodes is constructed, and a relationship distance between the JSON nodes is calculated according to the tree structure adjacency matrix;
[0104] A text similarity between the semantic identifier and the key name information of the JSON nodes is calculated, a semantic environment vector is constructed based on the relationship distance between the JSON nodes, and a semantic correlation degree matrix is obtained by weighted fusion of the text similarity and the semantic environment vector;
[0105] A mapping threshold is set to filter the semantic correlation degree matrix, redundant mapping relationships are eliminated based on a transitive closure, conflicts of one-to-many mapping and many-to-one mapping are processed, and a mapping relationship between the semantic identifier and the JSON nodes is determined;
[0106] A semantic identifier index and a node path index are constructed based on the mapping relationship to form a bidirectional mapping structure, and a semantic addressing table is generated according to the bidirectional mapping structure.
[0107] In a specific implementation, when performing word frequency statistics on the keywords in the keyword cluster family, a sliding window method is used to count the word frequency. For example, for a cluster family including keywords such as “user information”, “personal data”, and “account data”, the frequency of “user” is calculated to be 0.35, the frequency of “information” is calculated to be 0.25, the frequency of “data” is calculated to be 0.2, the frequency of “account” is calculated to be 0.15, and the frequency of “data” is calculated to be 0.05. When extracting the part-of-speech feature, a word segmenter is used to split the keywords into word units and label the part-of-speech, such as “user / noun” and “information / noun”. The syntactic dependency relationship is obtained through dependency syntax analysis, for example, “personal / adjective data / noun” indicates that “personal” modifies “data”. The word frequency, part-of-speech feature, and syntactic dependency relationship are integrated into a 300-dimensional semantic vector.
[0108] When calculating the semantic topic category, the semantic vector is matched with the pre-trained topic model. Assuming that the preset topic categories include "user management", "order processing", "payment system", etc., by calculating the cosine similarity, it is determined that the semantic topic category of this keyword cluster family is "user management", and the similarity score is 0.78. The hierarchical semantic coding rule is designed based on a three-level structure of "big category-middle category-small category", such as the "user management" big category is coded as "U", the "personal information" middle category under it is coded as "UI", and the "basic information" small category is coded as "UIB". Accordingly, the semantic identifier "UIB001" is assigned to this keyword cluster family.
[0109] When parsing the hierarchy of JSON data, a depth-first traversal algorithm is used. Taking the JSON data related to user information as an example, it contains the following structure: {"user":{"basicInfo":{"name":"","age":0,"gender":""},"contactInfo":{"phone":"","email":""}}}. The key names of the nodes are extracted, such as "user", "basicInfo", "name", etc., and the data types are determined, such as "object", "string", "number", etc. The parent-child relationship is determined by the nesting level, such as "user" is the parent node of "basicInfo"; the sibling relationship is determined by the sibling nodes, such as "name", "age", "gender" are sibling nodes.
[0110] When constructing the tree structure adjacency matrix, the elements in the matrix represent the connection relationship between nodes, and the value 1 represents direct connection and the value 0 represents no direct connection. For example, in an 8x8 adjacency matrix, the connection relationship between the user node and the basicInfo node is represented as matrix[0][1]=1, indicating that there is a direct connection from user to basicInfo. Through the adjacency matrix, the relationship distance between nodes can be calculated, such as the distance from user to name is 2, because it needs to pass through the basicInfo node to reach the name node.
[0111] When calculating the text similarity of semantic identifiers and JSON node key names, the text is preprocessed, including word segmentation, stop word removal, etc. For example, the text description "user basic information" associated with the semantic identifier "UIB001" and the key name of the JSON node "basicInfo" are processed to calculate the cosine similarity of the word embedding vector as 0.82. The semantic environment vector takes into account the context information of the JSON node, such as the parent node of the "basicInfo" node is "user" and the sibling node is "contactInfo". The relationship distance is converted into a weight, such as the distance from the parent node is 1 corresponding to the weight 0.8, and the distance from the sibling node is 2 corresponding to the weight 0.5. The calculation method of the weighted fusion of text similarity and semantic environment vector is: similarity x 0.7 + environment vector score x 0.3, to get the final semantic correlation matrix.
[0112] When setting the mapping threshold to filter the semantic correlation matrix, the threshold is set to 0.65, and the mapping relationship below the threshold is filtered out. The specific method of eliminating redundant mapping relationships based on transitive closure is: if semantic identifier 1 is mapped to node 1, node 1 is the parent node of node 2, and semantic identifier 1 is also mapped to node 2, then the mapping of semantic identifier 1 to node 1 is retained and the mapping of semantic identifier 1 to node 2 is deleted. When processing one-to-many mapping, compare the semantic correlation of multiple target nodes and select the highest correlation to retain. For example, the semantic identifier "UIB001" is mapped to "basicInfo" (correlation 0.82) and "contactInfo" (correlation 0.68) at the same time, and the mapping relationship with "basicInfo" is retained. When processing many-to-one mapping, analyze the hierarchical relationship of semantic identifiers, and preferentially retain the mapping relationship of more specific semantic identifiers. For example, "UIB001" (user basic information) and "UI002" (user information) are both mapped to the "basicInfo" node, and the mapping relationship of "UIB001" is retained because it represents a more specific semantic category.
[0113] When constructing a bidirectional mapping structure, two index tables are created. The semantic identifier index table stores the mapping of semantic identifiers to JSON node paths, such as {"UIB001":" / user / basicInfo"}; the node path index table stores the mapping of JSON node paths to semantic identifiers, such as {" / user / basicInfo":"UIB001"}. The final semantic addressing table contains these two index tables, and additional metadata such as creation time, version number, etc. is added to quickly locate specific nodes in JSON data through semantic identifiers, or to find corresponding semantic identifiers through node paths, to achieve semantic data access and retrieval.
[0114] Through the above steps, the mapping relationship between the semantic identification of the keyword cluster family and the JSON data node is established, and a semantic addressing table convenient for query and access is generated, thereby providing basic support for subsequent semantic-based data operation.
[0115] In an optional implementation, a target node in the JSON data is located according to the semantic addressing table, an explicit reference relationship and an implicit reference relationship are extracted to construct a dependency topology graph, an association strength between nodes is calculated, and a data block is constructed according to the association strength, and a globally unique block identifier is assigned to the data block, including:
[0116] The target node is located according to the semantic addressing table, and a reference relationship of the target node is extracted, including an explicit reference relationship and an implicit reference relationship, wherein the key-value pair relationship in the target node is extracted to determine the foreign key reference and the ID reference as the explicit reference relationship, and the co-occurrence frequency of the node attribute value and the similarity of the naming rule are calculated to determine the implicit reference relationship;
[0117] The target node is taken as a graph node, and the reference relationship is taken as a connection edge to construct a dependency topology graph, a transitive closure matrix of the dependency topology graph is calculated to identify a node subset of a circular reference, and a minimum spanning tree algorithm is used to optimize the dependency topology graph;
[0118] A distance association factor is calculated by calculating the shortest path length between nodes in the dependency topology graph, a reference association factor is calculated by counting the citation frequency, a feature association factor is calculated by calculating the data type matching degree and the numerical interval intersection, and an update association factor is calculated by calculating the correlation coefficient of the node data update time sequence, and the distance association factor, the reference association factor, the feature association factor and the update association factor are weighted and summed to obtain the association strength between nodes;
[0119] According to a preset association strength threshold, a node set is divided into a candidate block, an aggregation coefficient is calculated by calculating the ratio of the average degree to the maximum degree of the nodes in the candidate block, the candidate block is merged and optimized according to the aggregation coefficient, the number of shared nodes between the candidate blocks is calculated, and the shared nodes are assigned to the block with the largest aggregation coefficient, thereby obtaining a data block;
[0120] A globally unique block identifier is generated according to the association strength, the aggregation coefficient and the reference relationship of the data block.
[0121] In one specific implementation, a target node in JSON data is located according to a semantic addressing table. The semantic addressing table contains a mapping relationship between a field path and semantic information, for example, "customer.address.city" is mapped to "the city where the customer is located". By traversing the JSON document structure, the field path in the semantic addressing table is matched to determine the location of the target node. For the JSON data of an e-commerce system, the key business nodes such as "orders", "customers", and "products" can be located.
[0122] When extracting the reference relationship of the target node, it is divided into two categories: explicit reference and implicit reference. Explicit reference includes foreign key reference and ID reference. For example, "customer_id: 1001" in the "order" node points to the record with "id: 1001" in the "customers" collection, forming a foreign key reference relationship; "ref: # / definitions / product" in the "product_info" node points to the "product" definition in the document, forming an ID reference. Implicit reference is determined by calculating the co-occurrence frequency of node attribute values and the similarity of naming rules. For example, the "shipping_address" and "billing_address" nodes have similar field structures (street, city, zip), the similarity of naming rules is 0.85, and they co-occur in multiple transaction records, the co-occurrence frequency is 0.92, so they are determined to be implicitly referenced.
[0123] The target node is taken as a graph node, and the reference relationship is taken as a connection edge to construct a dependency topology graph. For example, in an order management system, the "Order" node references the "Customer" node, the "Customer" node references the "Address" node, and the "Order" node references the "Product" node, forming an initial dependency topology graph. The transitive closure matrix of the graph is calculated to identify the subset of nodes with circular references, such as "Department" referencing "Employee" and "Employee" referencing "Department" forming a circular reference. The Kruskal minimum spanning tree algorithm is used to optimize the dependency topology graph, retaining the edges with the highest association strength and removing redundant connections. In the optimization process, the edge weight of the circular reference is set to a lower value to ensure that it is not preferentially selected into the minimum spanning tree.
[0124] The calculation of the correlation strength between nodes in the dependency topology involves multiple factors. The distance correlation factor is obtained by calculating the shortest path length between nodes, such as the shortest path length from "Order" to "Customer" is 1, and the distance correlation factor is 1.0; the shortest path length from "Order" to "Address" is 2, and the distance correlation factor is 0.5. The citation correlation factor is obtained by counting the frequency of citation, such as the frequency of "Order" citing "Product" is 3.5 products per order on average, and the citation correlation factor is 0.85. The feature correlation factor calculates the data type matching degree and the intersection of value intervals, such as "Product.price" and "OrderItem.unit_price" are both decimal, and the value range has 90% overlap, and the feature correlation factor is 0.9. The update correlation factor calculates the correlation coefficient of the update time sequence of node data, such as the correlation coefficient of the update time sequence of "Inventory" and "OrderItem" nodes is 0.82, indicating that they are highly correlated. The four factors are weighted and summed according to weights of 0.3, 0.25, 0.25, and 0.2 to obtain the comprehensive correlation strength between nodes.
[0125] According to the preset correlation strength threshold 0.65, the node set is divided into candidate blocks. For example, nodes "Order", "OrderItem", and "Shipment" with a correlation strength greater than 0.65 are divided into the same candidate block. The ratio of the average degree to the maximum degree of the nodes in the candidate block is calculated to obtain the clustering coefficient, such as the average degree of block A is 3.2, the maximum degree is 5, and the clustering coefficient is 0.64; the clustering coefficient of block B is 0.58. If the clustering coefficients of two candidate blocks differ by less than 0.1 and there is a shared node, then the merging optimization is performed. For the shared node, its connection degree in each candidate block is calculated, and it is assigned to the block with the highest connection degree. For example, the "PaymentInfo" node is cited by both the "Order" block and the "Customer" block, with a connection degree of 4 in the "Order" block and a connection degree of 2 in the "Customer" block, so "PaymentInfo" is assigned to the "Order" block.
[0126] A globally unique block identifier is generated according to the correlation strength, clustering coefficient, and citation relationship of the data block. The identifier format is "prefix-content hash-timestamp", where the prefix represents the business type, such as "ORD" representing order-related blocks; the content hash is calculated by the node ID and correlation strength within the block; and the timestamp is the block creation time. For example, the identifier of the order-related data block may be "ORD-8F4D2A1B-20230615142538". This identifier uniquely identifies the data block in the distributed system, facilitating subsequent data access, update, and consistency maintenance.
[0127] By the above method, the accurate positioning of the nodes in the JSON data, the extraction of the association relationship, the construction of the data block and the allocation of the globally unique identifier are realized, thereby providing technical support for efficient management and access of complex data structures.
[0128] In an optional embodiment, candidate semantic identifiers are retrieved based on the target keyword in the semantic addressing table, the target data block and the write hierarchical position are determined, the JSON data to be appended is written into the target data block, the node association relationship in the dependency topology graph is updated, and the block change information is recorded, including:
[0129] A set of candidate semantic identifiers is retrieved according to the target keyword in the semantic addressing table, the time sequence activity and the reference heat of each semantic identifier in the set of candidate semantic identifiers are calculated, the priority score of the semantic identifier is obtained by weighted calculation, and the semantic identifier with the highest priority score is selected to determine the target data block;
[0130] The attribute association tree is constructed by performing structural analysis on the JSON data to be appended, the inheritance depth and the association width between nodes are calculated based on the attribute association tree, the hierarchical position of data writing is determined in combination with the inheritance depth and the association width, and the JSON data to be appended is written into the target data block;
[0131] The bidirectional reference index is constructed for the new node in the dependency topology graph, the circular reference path between nodes is identified based on the bidirectional reference index, the association relationship with the weakest reference strength in the circular reference path is marked as a soft connection, the soft connection only retains the one-way reference between nodes and is automatically disassociated after the expiration of the invalidation timestamp, the closed loop path in the circular reference path is blocked through the soft connection, the node association relationship in the dependency topology graph is updated, and the block change information is recorded.
[0132] In a specific embodiment, each semantic identifier is associated with a specific data block and records the time sequence activity and the reference heat. The time sequence activity represents the frequency of the semantic identifier being accessed in the near future, and the reference heat represents the number of times the semantic identifier is referenced by other nodes.
[0133] When JSON data needs to be appended, a target keyword such as “user preference” is received. A set of candidate semantic identifiers related to “user preference” is retrieved in the semantic addressing table, for example, three candidate semantic identifiers [“user_pref_01”, “preference_data_02”, “user_settings_03”] are obtained.
[0134] For each candidate semantic identifier, its temporal activity and reference heat are obtained. Suppose the temporal activity of "user_pref_01" is 0.8, and the reference heat is 0.5; the temporal activity of "preference_data_02" is 0.6, and the reference heat is 0.9; the temporal activity of "user_settings_03" is 0.4, and the reference heat is 0.3. The temporal activity and reference heat are weighted and calculated, the weight of the temporal activity is 0.6, and the weight of the reference heat is 0.4. The priority score of "user_pref_01" is 0.68, the priority score of "preference_data_02" is 0.72, and the priority score of "user_settings_03" is 0.36. Therefore, "preference_data_02" with the highest priority score is selected as the target semantic identifier, and the data block associated with it is determined as the target data block.
[0135] After obtaining the target data block, the JSON data to be appended is analyzed in structure, and an attribute association tree is constructed. Suppose the JSON data to be appended is: {"display": {"theme": "dark", "fontSize": 14}, "notification": {"email": true, "push": {"enabled": true, "quiet_hours": {"start": "22:00", "end": "06:00"}}}}.
[0136] The JSON structure is analyzed, and the attribute association tree is constructed. The root node is the entire JSON object, the first layer of child nodes includes "display" and "notification", and the second layer of child nodes includes "theme", "fontSize", "email" and "push" and the like. The inheritance depth and association width between nodes are calculated, the inheritance depth of "display" and "notification" is 1, and the association width is 2; the inheritance depth of "theme" and "fontSize" is 2, and the association width is 2; the inheritance depth of "push.quiet_hours.start" is 4, and the association width is 1.
[0137] Based on the calculation result, the hierarchical position of data writing is determined. For nodes with an inheritance depth greater than 3 and an association width less than 2, such as the "push.quiet_hours" part, they are processed as independent sub-documents; for nodes with an inheritance depth less than or equal to 3 or an association width greater than or equal to 2, they are directly inlined into the parent document. In this way, the JSON data to be appended is written into the target data block according to the determined hierarchical position.
[0138] After writing data, a bidirectional reference index is constructed for the new node in the dependency topology graph. The dependency topology graph is a directed graph, and the nodes represent data entities, and the edges represent reference relationships. Assuming that there are nodes 1 (user basic information) and node 2 (user permission settings) in the original dependency topology graph, and a new node 3 (user preference settings) is added. A reference relationship pointing to node 1 is created for node 3, and node 2 also references node 3, forming a reference chain of node 1←node 2←node 3.
[0139] Based on the bidirectional reference index, potential circular reference paths are identified. If node 2 is also referenced by node 1, a circular reference node 1←node 3←node 2←node 1 is formed. The reference strength of each reference relationship is calculated, and the reference strength is based on factors such as reference frequency and data dependency. Assuming that the reference strength of node 1←node 2 is 0.8, the reference strength of node 3←node 2 is 0.7, and the reference strength of node 2←node 1 is 0.5, then node 2←node 1 is the weakest association relationship in terms of reference strength.
[0140] Node 2←node 1 is marked as a soft connection, which only retains a one-way reference and sets a 24-hour expiration timestamp. By setting node 2←node 1 as a soft connection, the closed loop in the circular reference path is successfully blocked, avoiding the problem of a dead loop in data dependency.
[0141] The node association relationship in the dependency topology graph is updated, and the block change information is recorded, including the change timestamp, operation type (append), target data block identifier, and written data content digest. The change information is saved to a special block log for subsequent version tracking and data recovery.
[0142] Through the above implementation, the target data block can be efficiently identified and retrieved based on semantic identification, the hierarchical storage of JSON data can be reasonably arranged, the dependency relationship between data nodes can be maintained, the circular reference problem can be avoided, and efficient data management and access can be achieved.
[0143] The JSON data dynamic appending and management system based on a target keyword according to the embodiment of the application comprises:
[0144] A first unit is configured to receive a target keyword and JSON data.
[0145] A second unit is configured to convert the target keyword into a feature vector using syntax dependency analysis, generate a keyword affinity matrix through cosine similarity calculation, and perform adaptive classification vector clustering operation based on the keyword affinity matrix to obtain a keyword clustering family.
[0146] A third unit is configured to assign a semantic identifier to the keyword clustering family, establish a mapping relationship between the semantic identifier and nodes in the JSON data, and generate a semantic addressing table.
[0147] The fourth unit is configured to locate a target node in the JSON data according to the semantic addressing table, extract an explicit reference relationship and an implicit reference relationship to construct a dependency topology graph, calculate a correlation strength between nodes, construct a data block according to the correlation strength, and assign a globally unique block identifier to the data block;
[0148] The fifth unit is configured to search for a candidate semantic identifier in the semantic addressing table based on a target keyword, determine a target data block and a write hierarchical position, write the JSON data to be appended to the target data block, update a node correlation relationship in the dependency topology graph, and record block change information.
[0149] The sixth unit is configured to update the semantic addressing table and the dependency topology graph according to the block change information.
[0150] In a third aspect, an electronic device is provided, including:
[0151] a processor;
[0152] a memory for storing processor-executable instructions;
[0153] The processor is configured to invoke the instructions stored in the memory to execute the method described above.
[0154] In a fourth aspect, a computer-readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the method described above.
[0155] The present application can be a method, device, system and / or computer program product. The computer program product can include a computer-readable storage medium having stored thereon computer-readable program instructions that, when executed by a computer, cause the computer to carry out various aspects of the present application.
[0156] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for dynamic appending and managing JSON data based on target keywords, characterized in that, The method comprises the following steps: receiving a target keyword and JSON data; converting the target keyword into a feature vector by using syntax dependency analysis, generating a keyword affinity matrix by calculating cosine similarity, and performing adaptive classification vector clustering operation based on the keyword affinity matrix to obtain a keyword clustering family; assigning a semantic identifier to the keyword clustering family, establishing a mapping relationship between the semantic identifier and a node in the JSON data, and generating a semantic addressing table; locating a target node in the JSON data according to the semantic addressing table, extracting an explicit reference relationship and an implicit reference relationship to construct a dependency topology graph, calculating the correlation strength between nodes, constructing a data block according to the correlation strength, and assigning a globally unique block identifier to the data block; retrieving a candidate semantic identifier based on the target keyword in the semantic addressing table, determining a target data block and a write hierarchical position, writing the JSON data to be appended to the target data block, updating the node correlation relationship in the dependency topology graph, and recording block change information; updating the semantic addressing table and the dependency topology graph according to the block change information.
2. The method of claim 1, wherein, Converting the target keyword into a feature vector by using syntax dependency analysis, generating a keyword affinity matrix by calculating cosine similarity, and performing adaptive classification vector clustering operation based on the keyword affinity matrix to obtain a keyword clustering family comprises: performing word segmentation on the target keyword to obtain a word sequence, obtaining part-of-speech features of each word in the word sequence by using part-of-speech tagging, constructing a syntax dependency tree based on the part-of-speech features and position information, extracting syntax relationship features from the syntax dependency tree, and combining the part-of-speech features and the syntax relationship features to generate a semantic feature matrix; performing dimension reduction on the semantic feature matrix based on orthogonal transformation, selecting feature dimensions in turn according to variance contribution rates, and combining corresponding feature values to form a target feature vector; calculating the cosine distance between any two target feature vectors to obtain a similarity value, constructing an initial similarity matrix according to the similarity value, and determining an affinity matrix by normalizing the element values in the initial similarity matrix; statistically calculating the mean and variance of each row element in the affinity matrix, setting a dynamic threshold based on the mean and variance, setting the element values less than the dynamic threshold to zero to obtain an optimized affinity matrix; calculating the Laplacian matrix of the optimized affinity matrix, extracting the feature vector of the Laplacian matrix to construct a feature space, and performing adaptive classification vector clustering operation in the feature space to obtain a keyword clustering family.
3. The method of claim 2, wherein, Performing adaptive classification vector clustering operation in the feature space to obtain a keyword clustering family comprises: mapping the keywords in the feature space into sample points; establishing a search radius for each sample point, calculating the local density value by counting the number of samples within the search radius, calculating the Euclidean distance between any two sample points, and for each target sample point, selecting the sample point with the smallest Euclidean distance from all sample points with a local density value greater than the target sample point to determine a comparison sample point, and recording the Euclidean distance between the target sample point and the comparison sample point as a distance value. The product of the local density value and the distance value is calculated to obtain a density index, the density index is sorted in descending order, a density index threshold is set according to the numerical distribution of the density index, and sample points greater than the density index threshold are taken as initial clustering points; A classifier is constructed, and the initial clustering points are taken as classification vectors of the classifier; sample points are input one by one, the similarity between the sample points and each classification vector is calculated, and the classification vector with the highest similarity is selected as a matching vector; Based on the position of the current input sample point and the matching vector, the position of the matching vector is adjusted by a preset distance ratio, and other classification vectors adjacent to the matching vector are also adjusted; The above steps are repeated until the classification vectors are stable, sample points associated with the same classification vector are classified into the same category, and a keyword clustering family is obtained.
4. The method of claim 1, wherein, The semantic addressing table is generated by assigning a semantic identifier to the keyword clustering family and establishing a mapping relationship between the semantic identifier and the nodes in the JSON data, including: The keyword clustering family is assigned a semantic identifier, and a mapping relationship between the semantic identifier and the nodes in the JSON data is established to generate a semantic addressing table, including: The hierarchical structure of the JSON data is parsed to obtain JSON nodes, the key names and data types of the JSON nodes are extracted, the parent-child relationship and sibling relationship between the JSON nodes are determined, the tree structure adjacency matrix of the JSON nodes is constructed, and the relationship distance between the JSON nodes is calculated according to the tree structure adjacency matrix; The text similarity between the semantic identifier and the key name information of the JSON node is calculated, the semantic environment vector is constructed based on the relationship distance between the JSON nodes, and the semantic relevance matrix is obtained by weighting and fusing the text similarity and the semantic environment vector; The mapping threshold is set to filter the semantic relevance matrix, the redundant mapping relationship is eliminated based on the transitive closure, the conflicts of one-to-many mapping and many-to-one mapping are processed, and the mapping relationship between the semantic identifier and the JSON node is determined; The semantic identifier index and the node path index are constructed based on the mapping relationship to form a bidirectional mapping structure, and the semantic addressing table is generated according to the bidirectional mapping structure.
5. The method of claim 1, wherein, According to the semantic addressing table, the target node in the JSON data is located, the explicit reference relationship and the implicit reference relationship are extracted to construct a dependency topology graph, the correlation strength between nodes is calculated, and the data blocks are constructed according to the correlation strength, and a globally unique block identifier is assigned to the data blocks, including: According to the semantic addressing table, the target node is located, and the reference relationship of the target node is extracted, including the explicit reference relationship and the implicit reference relationship, wherein the foreign key reference and the ID reference are determined as the explicit reference relationship by extracting the key-value correspondence relationship in the target node, and the co-occurrence frequency of node attribute values and the similarity of naming rules are determined as the implicit reference relationship; The target node is taken as a graph node, and the reference relationship is taken as a connection edge to construct a dependency topology graph, the transitive closure matrix of the dependency topology graph is calculated to identify a node subset of circular reference, and the dependency topology graph is optimized by using a minimum spanning tree algorithm. A distance correlation factor is calculated according to a shortest path length between nodes in the dependency topology graph, a citation correlation factor is counted according to a citation frequency, a feature correlation factor is calculated according to a data type matching degree and a numerical interval intersection, and an update correlation factor is calculated according to a correlation coefficient of a node data update time sequence, and the distance correlation factor, the citation correlation factor, the feature correlation factor and the update correlation factor are summed to obtain a correlation strength between nodes; According to a preset correlation strength threshold, the node set is divided into candidate blocks, a ratio of an average degree to a maximum degree of nodes in the candidate blocks is calculated to obtain an aggregation coefficient, the candidate blocks are merged and optimized according to the aggregation coefficient, and the number of shared nodes between the candidate blocks is calculated and the shared nodes are distributed to a block with the largest aggregation coefficient to obtain data blocks; A globally unique block identifier is generated according to the correlation strength, the aggregation coefficient and the citation relationship of the data blocks.
6. The method of claim 1, wherein, Based on the target keyword, candidate semantic identifiers are retrieved in the semantic addressing table, a target data block and a write hierarchical position are determined, the JSON data to be appended is written into the target data block, the node correlation relationship in the dependency topology graph is updated, and block change information is recorded, including: According to the target keyword, a set of candidate semantic identifiers is retrieved in the semantic addressing table, the time sequence activity and the citation heat of each semantic identifier in the set of candidate semantic identifiers are calculated, and a priority score of the semantic identifier is obtained by weighted calculation, and the semantic identifier with the highest priority score is selected to determine the target data block; The JSON data to be appended is analyzed to construct an attribute association tree, the inheritance depth and the association width between nodes are calculated based on the attribute association tree, the hierarchical position of data writing is determined in combination with the inheritance depth and the association width, and the JSON data to be appended is written into the target data block; In the dependency topology graph, a bidirectional reference index is constructed for the new node, a circular reference path between nodes is identified based on the bidirectional reference index, the correlation relationship with the weakest reference strength in the circular reference path is marked as a soft connection, the soft connection only retains a one-way reference between nodes and is automatically disassociated after the expiration of the invalidation timestamp, the closed loop path in the circular reference path is blocked through the soft connection, the node correlation relationship in the dependency topology graph is updated, and the block change information is recorded.
7. A JSON data dynamic appending and management system based on target keywords, for implementing the method of any one of the preceding claims 1-6, characterized in that, Comprising: A first unit for receiving a target keyword and JSON data; A second unit for converting the target keyword into a feature vector by using syntax dependency analysis, generating a keyword affinity matrix through cosine similarity calculation, and performing adaptive classification vector clustering operation based on the keyword affinity matrix to obtain a keyword clustering family; A third unit for assigning a semantic identifier to the keyword clustering family, establishing a mapping relationship between the semantic identifier and nodes in the JSON data, and generating a semantic addressing table; A fourth unit for positioning a target node in the JSON data according to the semantic addressing table, constructing a dependency topology graph by extracting an explicit reference relationship and an implicit reference relationship, calculating a correlation strength between nodes, constructing data blocks according to the correlation strength, and assigning a globally unique block identifier to the data blocks; A fifth unit is configured to retrieve a candidate semantic identifier in a semantic addressing table based on a target keyword, determine a target data block and a write level position, write the JSON data to be added into the target data block, update a node association relationship in a dependency topology graph, and record block change information. A sixth unit is configured to update the semantic addressing table and the dependency topology graph according to the block change information.
8. An electronic device, comprising: The method comprises: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the method of any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by the processor, implement the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Data grading and classifying method and device
CN119538118A
Traditional Chinese medicinal material data management system and method for data retrieval
CN120045744A