Large model knowledge persistent storage and retrieval method

Through the persistent storage and retrieval method of large-model knowledge, the problem of insufficient version management and retrieval performance in traditional storage methods is solved, efficient knowledge storage and retrieval is achieved, and storage resource utilization and retrieval accuracy are improved.

CN120353778AInactive Publication Date: 2025-07-22HANGZHOU HONGQIANG TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510429850.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional knowledge storage methods are difficult to effectively support the version management, rapid retrieval and efficient storage optimization of knowledge units in large-scale pre-trained models. In addition, the existing index structure has high query complexity when facing the evolutionary relationship between high-dimensional knowledge data and complex topology, resulting in increased search delay and low storage resource utilization.

Method used

Through knowledge unit evolution trajectory modeling, manifold destruction compression, dynamic redirection indexing and nonlinear storage management, knowledge compression is used to use adaptive fractal networks for knowledge compression, combined with octree and Z-order spatial mapping to optimize storage layout, and dynamically recover storage resources using a dual-channel heat decay model.

Benefits of technology

It realizes efficient knowledge storage and retrieval, reduces redundant data usage, improves storage resource utilization and retrieval accuracy, and optimizes the load balancing and retrieval performance of storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353778A_ABST
    Figure CN120353778A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of storage and retrieval, in particular to a large-model knowledge persistent storage and retrieval method, which comprises the following steps of: capturing version evolution paths of large-model knowledge units in real time, and constructing a four-dimensional knowledge manifold comprising an ontology feature vector and a version association weight for each knowledge unit; extracting a core knowledge skeleton through a topology preserving layer, separating version difference characteristics through an evolution path layer, and generating compressed skeleton-path double codes; constructing a multi-level index structure based on the skeleton-path double coding, wherein the multi-level index structure comprises a static knowledge ontology layer and a dynamic path redirection layer; writing the knowledge unit into a storage device by adopting a hypercube mapping strategy, and establishing a path attenuation model to dynamically recycle a waste version space; according to the method, the long-term storage availability is improved, and the retrieval accuracy and the storage resource utilization rate are improved while efficient evolution management of the knowledge units is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of storage and retrieval, and particularly relates to a method for persistent storage and retrieval of large model knowledge. Background Art

[0002] With the development of large-scale pre-trained models, the complexity of knowledge storage and retrieval has been continuously increasing. Traditional knowledge storage methods usually adopt static databases or key-value storage. When facing a dynamically evolving knowledge system, it is difficult to effectively support version management, fast retrieval, and efficient storage optimization. In addition, the structural complexity of knowledge units in large models is relatively high, involving multi-layer semantic associations, version evolution paths, and cross-modal feature fusion. Traditional linear storage methods are difficult to meet the requirements of efficient management of multi-dimensional dynamic knowledge.

[0003] Traditional knowledge storage methods mainly rely on regular snapshots or incremental storage, and it is difficult to accurately capture the evolution characteristics of knowledge units, resulting in a large amount of redundant data, occupying storage resources and affecting retrieval performance. In addition, existing compression methods are mostly based on simple vector quantization or hash indexing, and do not fully utilize the structural information of knowledge units, resulting in a low storage compression rate and affecting storage scalability.

[0004] Traditional knowledge storage systems mainly rely on B+ trees, hash indexing, or inverted indexing for retrieval. When facing high-dimensional knowledge data and complex topological evolution relationships, the query complexity of the index structure increases significantly with the growth of data volume, resulting in an increase in retrieval latency. In addition, most existing index optimization strategies are based on fixed rules and cannot dynamically adapt to changes in access popularity, knowledge evolution paths, and query frequencies, restricting the retrieval performance of large-scale knowledge storage systems.

[0005] Due to the obvious time-series decay characteristics of the access frequency of knowledge units, traditional storage schemes often have difficulty distinguishing between long-term stable knowledge and short-term active knowledge, resulting in low utilization rate of storage resources. Summary of the Invention

[0006] The present invention provides a method for persistent storage and retrieval of large model knowledge. Through modeling the evolution trajectory of knowledge units, manifold deconstruction and compression, dynamic redirection indexing, and non-linear storage management, an efficient knowledge storage and retrieval mechanism is realized, the utilization rate of storage resources is improved, and the occupation of redundant data is reduced.

[0007] A method for persistent storage and retrieval of large model knowledge includes the following steps:

[0008] S1. Evolution trajectory modeling: Real-time capture the version evolution path of large model knowledge units, and construct a four-dimensional knowledge manifold including an ontology feature vector and a version association weight for each knowledge unit, where the version association weight is dynamically calculated and generated from the length and branch density of the version evolution path;

[0009] S2. Manifold Deconstruction and Compression: Input the four-dimensional knowledge manifold into the adaptive fractal network, extract the core knowledge skeleton through the topology-preserving layer, separate the version difference features through the evolutionary path layer, and generate the compressed skeleton-path dual coding.

[0010] S3. Dynamic Redirection Indexing: Based on the skeleton-path dual coding, construct a multi-level index structure, including a static knowledge ontology layer and a dynamic path redirection layer. Among them, the dynamic path redirection layer automatically generates associated jump paths according to the version access pattern.

[0011] S4. Nonlinear Persistent Storage: According to the spatial distribution of the skeleton coding and the popularity of the path redirection, adopt the hypercube mapping strategy to write knowledge units into the storage device, and establish a path decay model to dynamically recycle the abandoned version space.

[0012] Optionally, the S1 specifically includes:

[0013] S11, Version Change Monitoring: Continuously monitor the change amount of the feature vector of the knowledge unit through the sliding window mechanism. When the change rate of the vector cosine similarity exceeds the change rate threshold δ, trigger a version snapshot, and record the current timestamp and the changed content.

[0014] S12, Version Tree Construction: Aggregate historical version snapshots based on the knowledge unit ID, and construct a version tree structure with timestamps. The tree nodes in the version tree structure include a <V, t, p> triple, where V represents the feature vector, t is the timestamp, and p is the parent version pointer.

[0015] S13, Manifold Generation: Use the differential manifold algorithm to map the version tree into a four-dimensional knowledge manifold.

[0016] Optionally, in the four-dimensional knowledge manifold:

[0017] Dimensions 1-3: Extract the topological features of the version tree through the Graph Attention Network (GAT) to generate the ontology feature vector. The ontology feature vector is represented by the three-dimensional space coordinates (X, Y, Z): (X, Y, Z) = GAT(A), where A is the adjacency matrix of the version tree.

[0018] Dimension 4: Calculate the version association weight W v : W v = αL p + βD b where L p represents the path length, which represents the time span of the knowledge unit in the version evolution process, and D b represents the branch density factor, which represents the branch evolution degree of a certain knowledge unit at different time points.

[0019] Optionally, the S2 specifically includes:

[0020] S21, Manifold Feature Decoupling: Input the ontology feature vector in the four-dimensional knowledge manifold into the topology-preserving layer to extract its core knowledge skeleton. Use the multi-head self-attention mechanism to model the ontology feature vector to capture long-range dependence information. At the same time, prevent gradient disappearance through residual connections, and finally obtain the skeleton encoding. Meanwhile, take the version association weight as the evolution path feature and input it into the evolution path layer, and use the time-sensitive graph convolutional network to process it to separate the evolution features between different versions and generate the path encoding;

[0021] S22, Dynamic Feature Fusion: After extracting the skeleton encoding and the path encoding, adopt a cross-modal interaction method for fusion. Project the skeleton encoding and the path encoding into the same low-dimensional space, and use a gated weighted fusion function to calculate the fusion feature;

[0022] S23, Fractal Compression Encoding: Take the fusion feature as the input, and after being processed by the fractal compressor, generate the compressed skeleton-path dual encoding.

[0023] Optionally, the gated weighted fusion function adaptively adjusts the weight according to the dynamic characteristics of the input features, so that the fused fusion feature comprehensively retains the topological information and the evolution information.

[0024] Optionally, the processing of the fractal compressor in S23 includes fractal folding and fractal unfolding;

[0025] The fractal folding recursively folds the fusion feature, retains its low-frequency components (i.e., stable and core semantic information), and filters out high-frequency noise, including using a combination of max pooling and one-dimensional convolution to gradually compress the feature dimension and extract the skeleton features reflecting the essential structure of the knowledge unit;

[0026] The fractal unfolding sparsely unfolds the fusion feature, focusing on capturing its high-frequency components (i.e., dynamic version evolution information), including using a sparse mask to filter out the weight features and only retaining the path features strongly related to the version evolution.

[0027] Optionally, S3 specifically includes:

[0028] S31, Static Knowledge Ontology Layer Construction: Based on the skeleton encoding in the skeleton-path dual encoding, use spectral clustering to perform category division, generate the ontology category set, determine the category center vector, and calculate the coverage radius of the category to measure the range of feature distribution within the category. Meanwhile, use the semantic distillation model to automatically generate the semantic labels of the category for subsequent retrieval and organization of knowledge units;

[0029] S32. Construction of dynamic path redirection layer: Based on the path encoding in the skeleton-path dual encoding, construct a dynamic association graph to describe the evolution relationship between different versions, providing support for retrieval path optimization;

[0030] S33. Generation of associated jump paths: After receiving a retrieval request, first locate the target category in the static knowledge ontology layer to filter out the category range that meets the query requirements. Then, perform path search in the dynamic association graph. After the path search is completed, filter the top several optimal paths whose scores meet the filtering threshold and output them as the final retrieval results.

[0031] Optionally, the dynamic association graph includes establishing a node set of the graph and constructing an edge set. Each node in the node set corresponds to a corresponding version of a knowledge unit. The edge set represents the evolutionary connection between versions, and edge weights are calculated. The edge weights are jointly calculated by the path encoding similarity and access popularity between versions to construct a complete dynamic association graph.

[0032] Optionally, performing path search in the dynamic association graph includes: initializing a path weight accumulator and using the Monte Carlo tree search algorithm to find the optimal jump path. The optimal jump path is filtered based on path scoring, and the path scoring is based on the accumulation of edge weights. At the same time, a path decay factor is introduced to control the influence of depth on the path score.

[0033] Optionally, the specific steps of S4 are as follows:

[0034] S41. Spatial coordinate transformation: Perform spatial mapping on the skeleton encoding. Use the octree space segmentation algorithm to convert the encoding of the knowledge unit into three-dimensional spatial coordinates and organize them in a high-dimensional storage structure. The three-dimensional spatial coordinates are used to divide the hypercube storage space. At the same time, according to the path popularity information, determine the priority of the storage quadrants, so that knowledge units with high popularity can be stored in areas with higher access efficiency, optimizing the retrieval performance;

[0035] S42. Heat-aware storage: After obtaining the three-dimensional spatial coordinates, use the Z-order curve mapping method to convert the coordinates into physical storage addresses and encode the coordinate values to ensure that knowledge units adjacent in space are continuously distributed in the storage layout;

[0036] S43. Decay model construction: For the stored knowledge units, establish a decay model of path popularity to reflect the access trend of knowledge units over time. The decay model adopts a two-channel decay mechanism to calculate the long-term heat decay and short-term access fluctuation effects respectively, and adjust the storage priority of the data. The long-term decay channel is used to simulate the heat decline of knowledge units in the case of long-term non-access, and the short-term fluctuation channel is used to reflect the influence of recent access frequencies, balancing the historical access records and the current activity level;

[0037] S44, dynamic space recycling: Based on the path heat decay results, regularly detect low-heat knowledge units in storage and update the hypercube storage structure.

[0038] Beneficial effects of the present invention:

[0039] The present invention realizes efficient storage, retrieval and evolution management of knowledge through the collaborative design of knowledge unit evolution trajectory modeling, manifold deconstruction compression, dynamic redirection indexing and nonlinear storage management. Compared with traditional knowledge storage methods, an adaptive fractal network is used for knowledge compression to improve storage efficiency, and an octree and Z-order space mapping are used to optimize the storage layout and reduce retrieval latency. A dual-channel heat decay model and a dynamic storage recovery mechanism are combined to improve long-term storage availability. While ensuring efficient evolution management of knowledge units, the retrieval accuracy and storage resource utilization are improved.

[0040] The present invention adopts an adaptive fractal network to perform manifold feature deconstruction on knowledge units before storage, decouples skeleton information from evolutionary paths, and generates a dual-coding representation for efficient storage through dynamic feature fusion. The fused features undergo fractal folding and unfolding operations to achieve dynamic compression of storage occupancy while ensuring the integrity of knowledge evolution. Compared with traditional knowledge storage methods, this can reduce storage space occupancy and support efficient version evolution backtracking.

[0041] The present invention adopts a dual-channel path heat decay model, comprehensively considers long-term storage trends and short-term access fluctuations, and accurately calculates the life cycle of knowledge units through an exponential decay mechanism. When a knowledge unit is in a low-heat state for a long time, the system marks it as a recycling candidate, and automatically recycles and migrates it through a storage defragmentation engine to maximize storage resource utilization. Combined with a mechanism for dynamically adjusting the recycling threshold, it ensures that storage load optimization and retrieval performance are balanced. Under high load conditions, the storage strategy can be actively adjusted to prevent high-priority data from having access performance degradation due to storage bottlenecks. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 A schematic diagram of a method flow chart of an embodiment of the present invention;

[0044] Figure 2 Schematic diagram of dynamic redirection index according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the accompanying drawings are only for more specifically describing the embodiments and are not intended to specifically limit the present invention.

[0046] It should be pointed out that in the specification, when referring to "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc., it indicates that the described embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge scope of those skilled in the relevant art.

[0047] Generally, terms can be understood at least in part from their use in the context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but instead, at least in part depending on the context, allowing for the existence of other factors that may not be explicitly described.

[0048] As Figure 1 - Figure 2 shown, a method for persistent storage and retrieval of large model knowledge includes the following steps:

[0049] S1. Evolutionary trajectory modeling: Real-time capture the version evolution path of large model knowledge units, and construct a four-dimensional knowledge manifold for each knowledge unit, including an ontology feature vector and a version association weight, where the version association weight is dynamically calculated and generated by the version evolution path length and branch density;

[0050] S2. Manifold deconstruction and compression: Input the four-dimensional knowledge manifold into an adaptive fractal network, extract the core knowledge skeleton through a topology-preserving layer, and separate the version difference features through an evolution path layer to generate a compressed skeleton-path dual encoding;

[0051] S3. Dynamic redirection indexing: Based on the skeleton-path dual encoding, construct a multi-level index structure, including a static knowledge ontology layer and a dynamic path redirection layer, where the dynamic path redirection layer automatically generates an associated jump path according to the version access pattern;

[0052] S4. Nonlinear Persistent Storage: According to the spatial distribution of the skeleton encoding and the popularity of path redirection, adopt the hypercube mapping strategy to write knowledge units into the storage device, and establish a path attenuation model to dynamically recycle the abandoned version space.

[0053] The knowledge units of the large model refer to the independent knowledge fragments inside the large model, which can be one of the following forms:

[0054] Parameter level: The local patterns or important feature representations in the specific weight matrix learned by the large model during training.

[0055] Concept level: The specific facts, logical relationships, and reasoning abilities about a certain field stored in the model.

[0056] Semantic level: The language expression ways learned by the model in the training corpus, including the semantic associations between specific words, phrases, or text fragments.

[0057] Embedding level: The vectorized knowledge representation stored in the vector database.

[0058] Knowledge units are essentially a kind of structured or semi-structured knowledge fragments stored inside the large model. They are represented by feature vectors (V) and have the ability of version evolution.

[0059] S1 specifically includes:

[0060] S11, Version Change Monitoring: Continuously monitor the change amount of the feature vectors of knowledge units through the sliding window mechanism. When the change rate of the vector cosine similarity exceeds the change rate threshold δ, trigger a version snapshot to record the current timestamp and the changed content;

[0061] The sliding window is used to continuously monitor the change of the feature vectors of knowledge units, as follows:

[0062] 1. Initialize the window: Set the window size Define the number of historical versions for monitoring;

[0063] 2. Vector storage: Maintain a queue to store the feature vectors V of the most recent knowledge units t ;

[0064] 3. When a new vector arrives, add the latest feature vector V t to the window. If the window size exceeds then remove the earliest vector (FIFO);

[0065] 4. Calculate the vector change rate: Calculate the cosine similarity between the current vector V t and the previous version V t-1 in the window:

[0066]

[0067] Calculate the change rate of cosine similarity: ΔS t = |S(V t , V t-1 ) - S(V t-1 , V t-2 )|;

[0068] 5. Trigger version snapshot: If the change rate ΔS t > δ, trigger a version snapshot, record this version, and update the version tree, where δ is set to 0.15.

[0069] S12, Version tree construction: Aggregate historical version snapshots based on knowledge unit IDs to construct a timestamped version tree structure. The tree nodes in the version tree structure include a <V, t, p> triple, where V represents the feature vector, t is the timestamp, and p is the parent version pointer;

[0070] S13, Manifold generation: Use the differential manifold algorithm to map the version tree into a four-dimensional knowledge manifold.

[0071] In the four-dimensional knowledge manifold:

[0072] Dimensions 1 - 3: Extract the topological features of the version tree through a graph attention network (GAT) to generate an ontology feature vector. The ontology feature vector is represented by three-dimensional space coordinates (X, Y, Z): (X, Y, Z) = GAT(A), where A is the version tree adjacency matrix;

[0073] Dimension 4: Calculate the version association weight W v : W v = αL p + βD b , where L p represents the path length, which is the time span of the knowledge unit during the version evolution process, and D b represents the branch density factor, which indicates the degree of branch evolution of a certain knowledge unit at different time points;

[0074] The path length L p is calculated by accumulating the distance with time decay:

[0075] The branch density factor D b is calculated as: where α, β are adjustment coefficients, with initial values α = 0.6, β = 0.4, ω is the time decay factor, with a value of 0.05, ε is a very small constant to prevent the denominator from being zero, taking 10 -7 , t now is the current timestamp, t iis the timestamp of the i-th version, n is the number of versions, and t n -t1 represents the evolutionary time span of this knowledge unit. t1 is the timestamp of the earliest version of this knowledge unit (i.e., the creation time), and t n is the timestamp of the latest version of this knowledge unit (i.e., the time of the current latest snapshot).

[0076] GAT(A) represents that the graph attention network acts on the adjacency matrix A of the version tree, and is used to extract the version evolution topological features of the knowledge unit, specifically as follows:

[0077] 1. Input: The adjacency matrix A of the version tree. A represents the structure of the version tree and indicates the connection relationship between knowledge units.

[0078] 2. Calculate the attention weight A ij : GAT uses the self-attention mechanism to calculate the importance of version node i to neighbor node j: Among them, W is a trainable weight matrix (linear transformation parameter), and h i 、h j are the feature vectors of nodes i and j, a is the attention parameter vector, ∥ represents the vector concatenation operation, is the neighbor set of node i.

[0079] 3. Aggregate neighbor features: Calculate the new feature representation:

[0080] σ is the non-linear activation function ReLU.

[0081] 4. After being calculated by GAT, the version relationship of the knowledge unit is mapped to a new feature space, and the topologically enhanced three-dimensional coordinates (X, Y, Z) are obtained.

[0082] S2 specifically includes:

[0083] S21, manifold feature decoupling:

[0084] Skeleton encoding extraction: Input the ontological feature vector of the four-dimensional knowledge manifold into the topology-preserving layer, and use the multi-head self-attention network with residual connection to extract the core knowledge skeleton and generate the dimension-compressed skeleton encoding: Among them, the multi-head self-attention uses 8 heads to capture long-range dependencies, and the residual connection prevents gradient disappearance. The output dimension d1 of the skeleton encoding is 256;

[0085] Path encoding extraction: Input the version association weight into the evolution path layer, and separate the version difference features through the time-sensitive graph convolutional network (TS-GCN) to generate the path encoding:

[0086] Among them: TS-GCN is expressed as: Among them, H (l) is the node feature matrix of the l-th layer, W (l) is the learnable weight matrix, T is the timestamp vector, W t is the time projection matrix, σ is the non-linear activation function, ⊙ is the element-wise multiplication operation, which realizes the time attention weight, making the versions closer in time have a greater impact, Performs graph convolution to capture the topological information between versions, TW t Time projection is used to capture the time impact, ensuring that the updated versions have a greater impact than the old versions. Finally, the output of the L-th layer is the path encoding E p = H (L) , is the node adjacency matrix, represents the normalized connection weight between version i and version j, I represents the indices of all versions directly connected to version i, W v (i, j) represents the version association weight between version i and version j, ∑ I Wv ( The sum of the version association weights between version i and all adjacent versions I;

[0087] Simply put:

[0088]

[0089] 1 + exp(TW t ) → Adjust the influence according to time (newer versions have a greater impact);

[0090] ⊙ (Hadamard multiplication) → Fuse the topological information with the time impact;

[0091] σ(·) non-linear transformation → Enhance the expressive ability and avoid gradient vanishing.

[0092] Finally, TS-GCN extracts a time-sensitive version evolution feature, that is, the path encoding E p , which is used to describe how knowledge units evolve over time and ensure more accurate retrieval between versions.

[0093] S22, Dynamic Feature Fusion: Perform cross-modal interaction on the skeleton encoding E s and the path encoding E p to obtain the fusion feature E fusion :

[0094] Among them, W s 、W p are learnable projection matrices, Denotes matrix multiplication. GatedSum(·) is a gated weighted fusion function, defined as follows:

[0095] GatedSum(E s ,E p ) = αE s +(1 - α)E p ;

[0096] α = σ(W g [E s ;E p );

[0097] Among them, W g is the gated weight matrix, σ is the Sigmoid activation function, and [E s ;E p represents the concatenation operation.

[0098] S23, fractal compression encoding: Input the fused feature E fusion into the fractal compressor, and generate the compressed skeleton-path dual coding C through iterative feature recombination: C = [Fold(E fusion ;r),Unfold(E fusion ;θ)], where:

[0099] Fold is the fractal folding operation, extracting the skeleton feature: By recursively folding the high-dimensional features of E fusion , recursively folding them into a low-dimensional fractal structure (iterated 3 times), retaining the low-frequency components (skeleton features) strongly related to the core semantics. r = 0.4 is the folding rate, controlling the feature compression degree. A combination of max pooling + one-dimensional convolution is used to gradually compress the feature dimension, expressed as:

[0100] Fold(E fusion ;r) = MaxPool(Conv1D(E fusion ,r)), MaxPool is max pooling, and Conv1D is one-dimensional convolution;

[0101] Unfold is the fractal unfolding operation, separating the path feature: By sparsely unfolding E fusion , capturing the high-frequency change components (path features) related to version evolution, retaining the sparse expression of the key path features. θ = 0.35 is the sparse threshold, filtering the low-weight paths.

[0102] The fused feature, as the "intermediate representation", not only retains the key information of the original feature but also generates new associated features, enabling the compressed dual coding to simultaneously possess semantic integrity and version evolution characteristics. In this way, both the global information of the fused feature is retained, and independent compression of the skeleton and path is achieved.

[0103] S3 specifically includes:

[0104] S31, Static Knowledge Ontology Layer Construction: Encode the skeleton E s Input it into the knowledge clustering network, and generate the ontology category set through spectral clustering method n1 is the total number of categories, and each category C q is associated with the following tuples: C q = <V q , R q , T q >, where V q is the category center vector, R q is the coverage radius, and T q is the semantic label;

[0105] The specific process of generating the ontology category set through spectral clustering method is as follows:

[0106] Construct the similarity matrix: Calculate the similarity between the skeleton encodings of knowledge units to measure the association degree between knowledge units;

[0107] Calculate the Laplacian matrix based on the similarity matrix and perform normalization to capture the clustering structure of the data;

[0108] Calculate the eigenvectors: Solve the eigenvalues of the normalized Laplacian matrix, select the eigenvectors corresponding to the smallest multiple eigenvalues, and form a new low-dimensional representation space.

[0109] Perform k-means clustering in the eigenvector space to divide the knowledge units into different categories.

[0110] Generate ontology categories: For each category, calculate the category center vector, coverage radius, and automatically generate the semantic label of the category through the semantic distillation model, and finally form the ontology category set.

[0111] V q is generated by the mean of the skeleton encodings within the category: represents the index of the knowledge units within category C q , and traverse all the knowledge units in this category;

[0112] R q is defined as the standard deviation of the skeleton encodings within the category:

[0113] T q is automatically labeled through the semantic distillation model; Adopt the knowledge distillation technique to distill the large semantic model into a lightweight classification model, calculate the vector similarity between the center vector and the vectors in the semantic label dictionary, screen the label candidate set, and select the semantic label T with the highest matching degree through the association degree q as the final category annotation.

[0114] S32, Construction of Dynamic Path Redirection Layer: Based on path encoding E p Construct a dynamic association graph G = (N, E, J), where the node set N represents the knowledge unit version, the edge set E represents the version evolution relationship, J is the edge weight, and the edge weight J is jointly calculated by the path encoding similarity and the access popularity:

[0115] Among them, sim(E p (i), E p (j)) represents the similarity of path encoding, calculated using cosine similarity, H i is the popularity value of version i, H j is the popularity value of version j, both calculated by statistically counting the access frequencies through a sliding window;

[0116] S33, Generation of Associated Jump Paths: When a retrieval request is received, perform the following operations:

[0117] 1. Locate the target category C in the static knowledge ontology layer k ;

[0118] 2. Initiate path exploration in the dynamic association graph:

[0119] Initialize the path weight accumulator U = {}, U = {} is an empty set, used to store the cumulative path weight values calculated during the path search process, for accumulating the scores of each path during the Monte Carlo tree search process, facilitating the subsequent selection of the optimal jump path;

[0120] Use Monte Carlo tree search (MCTS) to select the optimal jump path:

[0121] Among them, λ is the path attenuation factor, with a value of 0.2, is the jump depth, represents the path depth attenuation factor, controlling the influence weight of longer paths, PathScore is the comprehensive score of the current path, ∑ e∈Path represents the cumulative calculation for all edges e in the path, e, J e is the edge weight of edge e;

[0122] 3. Output the top K associated paths that satisfy as the retrieval results, is the screening threshold for associated jump paths, Among them, is the mean of the historical path scores, used to represent the average path quality, is the standard deviation of the historical path scores, used to measure the dispersion degree of path scores.

[0123] S4 specifically includes:

[0124] S41, Spatial coordinate conversion: Convert the knowledge skeleton encoding E s into three-dimensional space storage coordinates (x, y, z) through the octree space segmentation method, where (x, y, z) = OctreeEncode(E s ), and each coordinate corresponds to a storage quadrant in the hypercube. The storage quadrant priority Q is dynamically determined by the path heat H: where H max is the maximum heat value of the current system, represents rounding to the nearest integer, and OctreeEncode represents the octree space segmentation operation;

[0125] S42, Heat-aware storage: Map the three-dimensional coordinates to physical storage nodes using the Z-order curve:

[0126] where bit f (·) represents the f-th bit binary value of the coordinate value, bit f (x), bit f (y), bit f (z) respectively represent the f-th bit of the three-dimensional coordinates x, y, z in binary representation, and n3 is the number of binary digits;

[0127] S43, Decay model construction: Establish a two-way path heat decay function:

[0128] where λ1 is the long-term decay factor, λ1 = 0.05, λ2 is the short-term fluctuation factor, λ2 = 0.2, is the th access increment, H0 is the initial heat value, represents the th access timestamp, and m is the number of times the path has been accessed;

[0129] S44, Dynamic space recycling: When it is detected that the path heat H(t) < ρ continuously exceeds the time window T w = 24h:

[0130] 1. Mark the corresponding storage quadrant as a recyclable area;

[0131] 2. Start the storage fragmentation reorganization engine to migrate the valid data to the adjacent quadrant;

[0132] 3. Release the abandoned version space to the free pool and update the hypercube space distribution map;

[0133] where the heat threshold ρ = 0.15H max, when the storage load is higher than 80%, adjust the heat threshold ρ = 0.2H max .

[0134] The present invention covers any alternatives, modifications, equivalent methods, and solutions made to the essence and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without these detailed descriptions. Additionally, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.

[0135] The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for persistent storage and retrieval of large model knowledge, characterized in that, It includes the following steps: S1. Evolutionary trajectory modeling: Real-time capture the version evolution path of the large model knowledge units, and construct a four-dimensional knowledge manifold including ontology feature vectors and version association weights for each knowledge unit, where the version association weights are dynamically calculated and generated by the version evolution path length and branch density; S2. Manifold deconstruction and compression: Input the four-dimensional knowledge manifold into an adaptive fractal network, extract the core knowledge skeleton through the topology-preserving layer, separate the version difference features through the evolution path layer, and generate a compressed skeleton-path dual encoding; S3. Dynamic redirection indexing: Based on the skeleton-path dual encoding, construct a multi-level index structure, including a static knowledge ontology layer and a dynamic path redirection layer. Among them, the dynamic path redirection layer automatically generates associated jump paths according to the version access pattern; S4. Nonlinear persistent storage: According to the spatial distribution of the skeleton encoding and the popularity of the path redirection, adopt the hypercube mapping strategy to write the knowledge units into the storage device, and establish a path decay model to dynamically recycle the abandoned version space.

2. A method for persistent storage and retrieval of large model knowledge according to claim 1, characterized in that The specific content of S1 includes: S11. Version change monitoring: Continuously monitor the change amount of the feature vectors of the knowledge units through a sliding window mechanism. When the change rate of the vector cosine similarity exceeds the change rate threshold δ, trigger a version snapshot, and record the current timestamp and the changed content; S12. Version tree construction: Aggregate historical version snapshots based on the knowledge unit ID to construct a version tree structure with timestamps. The tree nodes in the version tree structure include <V, t, p> triples, where V represents the feature vector, t is the timestamp, and p is the parent version pointer; S13. Manifold generation: Use the differential manifold algorithm to map the version tree into a four-dimensional knowledge manifold.

3. A method for persistent storage and retrieval of large model knowledge according to claim 2, characterized in that, In the four-dimensional knowledge manifold: Dimensions 1-3: Extract the topology features of the version tree through a graph attention network to generate ontology feature vectors, and the ontology feature vectors are represented by three-dimensional space coordinates (X, Y, Z); Dimension 4: Calculate the version association weight W v : W v =αL p +βD b , where L p represents the path length, which is the time span of the knowledge unit during the version evolution process, and D b represents the branch density factor, which indicates the degree of branch evolution of a certain knowledge unit at different time points.

4. A method for persistent storage and retrieval of large model knowledge according to claim 1, characterized in that, The specific content of S2 includes: S21. Manifold feature decoupling: Input the ontology feature vectors in the four-dimensional knowledge manifold into the topology-preserving layer to extract its core knowledge skeleton, use the multi-head self-attention mechanism to model the ontology feature vectors to capture long-range dependence information, and at the same time prevent gradient disappearance through residual connections. Finally, obtain the skeleton encoding. At the same time, input the version association weights as the evolution path features into the evolution path layer, and use a time-sensitive graph convolutional network to process them to separate the evolution features between different versions and generate path encoding; S22. Dynamic feature fusion: After extracting the skeleton encoding and the path encoding, use a cross-modal interaction method for fusion, project the skeleton encoding and the path encoding into the same low-dimensional space, and use a gated weighted fusion function to calculate the fusion features; S23. Fractal compression encoding: Use the fusion features as input, and after being processed by a fractal compressor, generate a compressed skeleton-path dual encoding.

5. A method for persistent storage and retrieval of large model knowledge according to claim 4, characterized in that, The gated weighted fusion function adaptively adjusts the weights according to the dynamic characteristics of the input features, so that the fused fusion features comprehensively retain the topological information and the evolution information.

6. A method for persistent storage and retrieval of large model knowledge according to claim 4, characterized in that, The fractal compressor processing in S23 includes fractal folding and fractal unfolding; The fractal folding fuses features through recursive folding, retains its low-frequency components, and filters high-frequency noise, including using a combination of max pooling and one-dimensional convolution to gradually compress the feature dimension and extract the skeleton features reflecting the essential structure of the knowledge unit; The fractal unfolding fuses features through sparse unfolding, focusing on capturing its high-frequency components, including using a sparse mask to filter out weight features and only retaining the path features strongly related to version evolution.

7. A method for persistent storage and retrieval of large model knowledge according to claim 1, characterized in that The S3 specifically includes: S31, construction of the static knowledge ontology layer: Based on the skeleton encoding in the skeleton-path dual encoding, use spectral clustering to perform category division, generate an ontology category set, determine the category center vector, and calculate the coverage radius of the category to measure the range of feature distribution within the category. At the same time, use a semantic distillation model to automatically generate semantic labels for the category; S32, construction of the dynamic path redirection layer: Based on the path encoding in the skeleton-path dual encoding, construct a dynamic association graph to describe the evolution relationship between different versions; S33, generation of associated jump paths: After receiving a retrieval request, first locate the target category in the static knowledge ontology layer to filter out the category range that meets the query requirements, and then perform path search in the dynamic association graph. After the path search is completed, filter the top several optimal paths whose scores meet the filtering threshold and output them as the final retrieval result.

8. A method for persistent storage and retrieval of large model knowledge according to claim 7, characterized in that, The dynamic association graph includes establishing a node set of the graph and constructing an edge set. Each node in the node set corresponds to a corresponding version of a knowledge unit. The edge set represents the evolution connection between versions, and the edge weight is calculated. The edge weight is jointly calculated by the path encoding similarity between versions and the access popularity, constructing a complete dynamic association graph.

9. A method for persistent storage and retrieval of large model knowledge according to claim 7, characterized in that, Performing path search in the dynamic association graph includes: initializing a path weight accumulator and using the Monte Carlo tree search algorithm to find the optimal jump path. The optimal jump path is filtered based on path scoring, and the path scoring is based on the accumulation of edge weights. At the same time, a path decay factor is introduced to control the influence of depth on the path score.

10. The method for persistent storage and retrieval of large model knowledge according to claim 1, characterized in that, The S4 specifically includes: S41, spatial coordinate transformation: Perform spatial mapping on the skeleton encoding, adopt an octree space segmentation algorithm to convert the encoding of the knowledge unit into three-dimensional spatial coordinates, and organize them in a high-dimensional storage structure. The three-dimensional spatial coordinates are used to divide the hypercube storage space. At the same time, according to the path heat information, determine the priority of the storage quadrant; S42, heat-aware storage: After obtaining the three-dimensional spatial coordinates, use the Z-order curve mapping method to convert the coordinates into physical storage addresses and encode the coordinate values to ensure that knowledge units adjacent in space are continuously distributed in the storage layout; S43, construction of the decay model: For the stored knowledge units, establish a decay model of path heat to reflect the access trend of knowledge units over time. The decay model adopts a two-channel decay mechanism to calculate the long-term heat decay and the influence of short-term access fluctuations respectively, and adjusts the storage priority of the data. The long-term decay channel is used to simulate the heat decline of knowledge units in the case of long-term non-access, and the short-term fluctuation channel is used to reflect the influence of recent access frequencies, balancing historical access records and current activity; S44, Dynamic Space Recycling: Based on the path heat decay result, regularly detect the knowledge units with low heat in storage and update the hypercube storage structure.

Citation Information

Cited By

  • Atomic knowledge base-based Internet of Things knowledge fusion method and system

    CN121412198A

  • An internet of things knowledge fusion method and system based on an atomic knowledge base

    CN121412198B

  • Intelligent decision support method based on dynamic knowledge graph and multi-modal fusion

    CN121882221A