Knowledge security evolution updating method and system for newly added information

By constructing a multi-granularity temporal index structure and calculating the temporal reversal congestion factor, the problems of key information being overwhelmed and retrieval queue congestion in existing technologies are solved, enabling stable and secure evolution and updates of the knowledge base in scenarios such as college admissions consultation.

CN121880346APending Publication Date: 2026-04-17NANJING SANLIUJIE NETWORK INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING SANLIUJIE NETWORK INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-01-05
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing knowledge base management systems suffer from problems such as the overload of key information, congestion in retrieval queues, and unstable updates when dealing with complex business scenarios with strict phase boundaries, such as college admissions consultations. They cannot simultaneously ensure the freshness of the data entry and the effectiveness of the business.

Method used

Construct a multi-granularity time-series index structure that includes key nodes in the business cycle. By parsing the entry and effective timestamps, calculate the time-series reversal congestion factor and the security evolution sorting weight, and achieve accurate sorting and stable updates of knowledge data.

Benefits of technology

This ensures high visibility of key information within the time window, prevents core outline documents from being pushed out of the first search results, improves system stability and semantic relevance, and reduces the occurrence of time illusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880346A_ABST
    Figure CN121880346A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a knowledge security evolution updating method and system oriented to newly-added information, and the method comprises the steps: constructing a multi-granularity time sequence index structure containing business cycle key nodes, marking system input timestamps for collected data, and analyzing business effective timestamps; identifying and inputting an effective time difference distribution zone based on the distribution pattern of the difference value of the two, and calculating a time sequence reversal congestion factor representing the priority reversal risk of the data near the key node; combining the factor, the input aging attenuation degree and the service aging fitting degree to calculate a security evolution sorting weight, and generating a structured suppression parameter by using the factor; and finally, performing incremental data fusion on the target node of the index structure according to the weight, and updating the knowledge base. Through two-dimensional time modeling and an anti-congestion mechanism, the problem that core knowledge is submerged due to high-concurrency window period fragmentation information is effectively solved, and orderly and safe evolution of the knowledge base is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a knowledge security evolution and update method and system for new information. Background Technology

[0002] With the widespread application of Retrieval Augmentation (RAG) technology in vertical fields such as educational consulting and admissions Q&A, the system's requirements for the timeliness of knowledge are becoming increasingly stringent. In these scenarios, the knowledge base not only needs to store static documents, but also needs to handle policies, notices, and status information that change dynamically over time.

[0003] Existing knowledge base management systems typically use a single time dimension for data sorting and retrieval, relying primarily on the system entry time or document publication time to determine information freshness. In general information flow scenarios, this "newest is best" strategy can meet most needs, allowing users to prioritize recently generated information. While some advanced systems have introduced the concept of effective time, their processing mechanisms often simply treat effective time as a simple filtering interval or apply a simple linear weighting to it and the entry time.

[0004] However, in complex business scenarios with strict phased boundaries, such as college admissions consultation and application submission, the existing processing methods mentioned above have significant technical limitations. In these scenarios, the generation and effectiveness of data exhibit a unique two-dimensional time distribution pattern: on the one hand, core policy documents and admission regulations are often released well before the business window opens (belonging to advance-released data); on the other hand, near the end of the business window or at key nodes (such as supplementary applications or batch switching), a large number of patch information, surplus / shortage notices, or error correction announcements are generated intensively (belonging to delayed-patch data).

[0005] Using existing single-time sorting or simple linear weighting strategies will lead to a severe time priority inversion problem in the above scenario. Specifically, this manifests as follows: 1. Key information is buried: Near the boundary of business phases, due to the very recent entry time of the delayed patch data, the existing system will prioritize it, which will cause the early-released data (such as general charter and core rules) that are in the core effective period to be squeezed to the back in the search results, making it difficult for users to obtain the truly decisive general information.

[0006] 2. Retrieval queue congestion: During time-sensitive windows, a large amount of fragmented correction information floods in, causing the head of the retrieval index to be occupied by homogeneous, frequently updated data, resulting in a significant decrease in the semantic diversity of the system.

[0007] 3. Instability of evolution and updates: The existing system lacks in-depth modeling of the distribution pattern of the difference between the entry time and the effective time, and cannot identify whether the data belongs to the preview or the remediation. As a result, the knowledge base cannot automatically adjust the weight according to the business stage when it is updated incrementally, which can easily introduce short-term noise and undermine the long-term stability of the knowledge base.

[0008] In summary, how to balance the freshness of data entry with the effectiveness of business operations within a system, especially by preventing fragmented patch information from crowding out core knowledge near critical business nodes, and achieving the secure and orderly evolution of the knowledge base, is a pressing technical challenge that needs to be addressed. Summary of the Invention

[0009] This invention provides a knowledge security evolution and update method and system for new information, which solves the technical problems mentioned in the background art.

[0010] Firstly, a knowledge security evolution and update method for newly added information includes: Construct a multi-granularity time-series index structure that includes key nodes in the business cycle, and mark the collected knowledge data with system-entered timestamps and parse the business effective timestamps; Based on the difference distribution characteristics between the business activation timestamp and the system entry timestamp, the entry activation time difference distribution band is identified, and the time sequence reversal congestion factor, which characterizes the risk of time sequence priority reversal of knowledge data near the key nodes of the business cycle, is calculated. The security evolution ranking weight is calculated by combining the time-inversion congestion factor, the data entry timeliness decay degree, and the business timeliness fit degree, wherein the time-inversion congestion factor is configured to generate a structured suppression parameter for the security evolution ranking weight. Based on the security evolution sorting weight, incremental data fusion is performed on the target nodes of the multi-granularity time-series index structure, and the knowledge base inverted index containing two-dimensional time metadata is updated.

[0011] Secondly, a knowledge security evolution and update system for newly added information, applied in any of the knowledge security evolution and update methods for newly added information described above, characterized in that it includes: The time series parsing and construction module is used to build a multi-granularity time series index structure containing key nodes of the business cycle, mark the collected knowledge data with system-entered timestamps, and parse the business effective timestamps; The risk identification and factor calculation module is used to identify the time difference distribution band of entry and entry based on the difference distribution characteristics between the business entry time stamp and the system entry time stamp, and to calculate the time-series reversal congestion factor that represents the risk of time-series priority reversal of knowledge data near the key nodes of the business cycle. The weight calculation module is used to calculate the security evolution ranking weight by combining the time-inversion congestion factor, the input timeliness decay degree and the business timeliness fit degree, wherein the time-inversion congestion factor is configured to generate a structured inhibition parameter for the security evolution ranking weight. The update and fusion module is used to perform incremental data fusion on the target nodes of the multi-granularity time-series index structure according to the security evolution sorting weight, and update the knowledge base inverted index containing two-dimensional time metadata.

[0012] The beneficial effects of this invention are as follows: 1. This invention introduces a computational mechanism for the distribution band of data entry effective time difference and the time-series reversal congestion factor. By quantifying the distribution pattern of data near key nodes in the business cycle, the system can accurately identify data that is very recent in entry time but is only a lagging patch, and generate structured suppression parameters for it. This effectively solves the problem that in time-sensitive scenarios such as student recruitment consultation, the influx of a large amount of fragmented patch information causes core general outline documents to be squeezed out of the first screen of search results, ensuring the high visibility of key information within the time window.

[0013] 2. This invention abandons the traditional hard threshold-based judgment logic and instead uses continuous mathematical models such as S-shaped functions and exponential decay functions to calculate the structured soft gate weights and security evolution ranking weights. This processing mode ensures that the weight changes are gradual and continuous when the system switches business cycles (such as from the registration period to the admission period), avoiding data jitter or service interruption caused by sudden rule changes, and greatly improving the stability of the system.

[0014] 3. This invention constructs a multi-granularity time-series index structure containing key nodes of the business cycle and explicitly embeds two-dimensional time metadata into the inverted index and semantic vector aggregation process. By introducing an anti-congestion correction kernel during the index construction stage, the system enables the generated node summary vector to naturally possess denoising capabilities. During retrieval, this mechanism ensures that the output results are not only semantically relevant but also safe and within the effective window in terms of business timeliness, effectively reducing the occurrence of time illusion. Attached Figure Description

[0015] Figure 1 This is a flowchart of a knowledge security evolution and update method for new information according to the present invention; Figure 2 This is an architecture diagram of a knowledge security evolution and update system for new information according to the present invention. Detailed Implementation

[0016] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.

[0017] Example 1: As Figure 1 As shown, a knowledge security evolution and update method for new information includes: Construct a multi-granularity time-series index structure that includes key nodes in the business cycle, and mark the collected knowledge data with system-entered timestamps and parse the business effective timestamps; Based on the difference distribution characteristics between the business activation timestamp and the system entry timestamp, the entry activation time difference distribution band is identified, and the time sequence reversal congestion factor, which characterizes the risk of time sequence priority reversal of knowledge data near the key nodes of the business cycle, is calculated. The security evolution ranking weight is calculated by combining the time-inversion congestion factor, the data entry timeliness decay degree, and the business timeliness fit degree, wherein the time-inversion congestion factor is configured to generate a structured suppression parameter for the security evolution ranking weight. Based on the security evolution sorting weight, incremental data fusion is performed on the target nodes of the multi-granularity time-series index structure, and the knowledge base inverted index containing two-dimensional time metadata is updated.

[0018] In a preferred embodiment, a multi-granularity time-series index structure containing key nodes of the business cycle is constructed, including: Establish a hierarchical index tree that includes year, month, and day time level nodes and key business cycle nodes mounted at the bottom layer; Calculate the absolute time difference between the business activation timestamp of the collected knowledge data and the most recent key node in the business cycle, and define this absolute time difference as the business activation interval. The calculation formula is: ; in, For the first The business activation timestamp of the knowledge data. For the first Key time points in each business cycle; Based on the effective time interval of the aforementioned business, a node proximity coefficient, representing the strength of the association between knowledge data and key nodes in the business cycle, is calculated using an exponential decay function. The node proximity coefficient is stored in the index metadata, and the calculation formula is as follows: ; in, The impact scale parameter for key nodes in the business cycle.

[0019] Preferably, the process of constructing a multi-granularity time-series index structure containing key nodes of the business cycle is executed according to the following steps and logic.

[0020] First, the system initializes a hierarchical index tree in the form of a multi-way tree in the in-memory database. The topology of this tree is limited to four levels, from top to bottom: year-time level nodes, month-time level nodes, day-time level nodes, and business cycle key nodes attached to the bottom layer.

[0021] The system reads the pre-configured business configuration table and parses out the set of time points for all key nodes in the business cycle. ,in Indexed by natural numbers, each All are 64-bit Unix timestamps accurate to the second, representing hard time anchors such as the start of the application process or the end of the admission process. These nodes are then sorted in chronological order and attached to the corresponding day-time level nodes.

[0022] When the system collects the first When processing knowledge data, extract the parsed business activation timestamp from its metadata. This timestamp is also in 64-bit Unix timestamp format. The system then executes a binary search algorithm in the ordered set. Mid-positioning The predecessor and successor nodes are calculated separately. The absolute value of the time difference between these two nodes is taken as the minimum value for the service activation time. The calculation logic of this process follows the formula. ,in The unit of measurement is seconds, and the range of values ​​is... .like If the value is empty, the system throws an exception and blocks the index building process because the relative distance cannot be calculated due to the lack of a business anchor point.

[0023] Subsequently, the system based on the calculated service activation time interval The node proximity coefficient is calculated using the exponential decay function. This method quantifies the correlation strength between a piece of knowledge data and the most recent key nodes in the business cycle. Based on the time decay prior of information relevance—that is, the closer knowledge data is to key nodes, the higher its value for user decision-making, and this value decreases sharply and non-linearly with increasing distance—the calculation formula is defined as follows: ,in For the natural constant An exponential function with base 0, output value for Dimensionless floating-point numbers within an interval. (In the formula) Defined as the impact scale parameter of key nodes in the business cycle, its unit is seconds. The value is derived based on the statistical distribution of historical query logs. Specifically, it is obtained by performing Gaussian fitting on the time distribution of user query frequency within the same historical business period, and taking the time length corresponding to the half-peak full width (FWHM) of the distribution as the metric. The optimal value, for example, in the scenario of recruiting volunteers, is set as follows: (i.e., 24 hours) means that when the effective time of knowledge data is 24 hours away from the key node, its correlation strength decays to... For the calculated The system serializes it into a 32-bit single-precision floating-point number and writes it into the corresponding index metadata field in the inverted index. Specifically, if... Exceed The system will Forced truncation to machine minimum To avoid underflow while maintaining non-zero attributes, ensure that all indexed data is associated with business nodes, and complete the feature space mapping from raw time parameters to business semantic relevance.

[0024] It should be noted that the system initializes the set of key nodes in the business cycle by reading a pre-set JSON-formatted business configuration file. This configuration file defines a list of all events with hard time boundaries within the current enrollment year, specifically covering the start and end times of each batch of application submissions, the time of admission results announcement, and the opening time of the supplementary application window, among other core time anchors. When the system starts, it parses this file and standardizes all time points into UTC timestamps. Then, it performs deduplication and sorts the data in ascending order to generate a globally ordered list of key node timelines. This list is resident in memory and serves as a reference for subsequent calculations of the effective time of business operations and the node proximity coefficient. Furthermore, the system supports dynamically updating this list without stopping the service through a hot-loading interface to adapt to sudden policy time adjustments, ensuring that distance calculations are always based on the latest business rules.

[0025] In a preferred embodiment, the process of recording timestamps on the collected knowledge data and parsing the business activation timestamps includes: The moment when the system collects the current knowledge data is recorded is identified and marked as the system's data entry timestamp. ; Identify candidate effective timestamps from the content of knowledge data using information extraction models. And simultaneously output the corresponding extraction confidence score. ,in The value ranges from 0 to 1; Based on the extracted confidence level, a weighted sum is calculated between the candidate effective timestamps and the system-entered timestamps to obtain the final business effective timestamp. The calculation formula is: ; Wherein, when the extraction confidence level When the value approaches zero, the service activation timestamp is numerically and smoothly rolled back to the system entry timestamp to maintain the continuity of the parsing process.

[0026] Preferably, the process of recording timestamps in the system and parsing the business activation timestamps for the collected knowledge data follows the following continuous execution logic and numerical processing flow.

[0027] The system first receives the first data at the data acquisition interface. The instant a knowledge data packet is received, the operating system's kernel-level time function is immediately invoked to obtain the current high-precision Coordinated Universal Time (UTC) and convert it into a 64-bit double-precision floating-point Unix timestamp. This value is uniquely and immutably marked as the system's recorded timestamp. Its unit is seconds, and its range is 100%. .

[0028] Next, the system inputs the unstructured text content of the knowledge data into a pre-trained information extraction model based on the BERT-BiLSTM-CRF architecture. This model performs sequence labeling to identify implicit temporal entities in the text. The temporal entity fragment with the highest log probability in the model's output sequence is parsed and normalized to the Unix timestamp format, defined as a candidate effective timestamp. Meanwhile, the normalized probability value corresponding to the entity fragment in the output probability distribution of the model's Softmax layer is extracted as the extraction confidence score. The range of this value is strictly constrained to a closed interval. If the model fails to identify any valid time entities, the system will perform default boundary processing. Set as And Forced to Subsequently, the system enters the core soft fusion computing phase, based on an uncertainty-aware temporal smoothing mechanism. This mechanism aims to eliminate the risks of gradient truncation and process interruption caused by traditional hard threshold judgment logic (such as if-else branches). The system follows the formula... Calculate the final business activation timestamp In this formula, As the merged output time, it inherits the units and data format of the input timestamp; As a dynamic weighting coefficient, it quantifies the credibility of semantic information; This refers to the complement weight of the system observation time.

[0029] When extracting confidence level Approaching When the model considers the extracted results to be highly unreliable or nonexistent, the product term... tending to And weighted terms tending to This makes the output result Numerically, it smoothly regresses and gradually converges to the system's recorded timestamp. This ensures that even in extreme cases where semantic parsing completely fails, the system can still generate valid time (i.e., the time when data enters the system).

[0030] It should be noted that when constructing the information extraction model for parsing business effective timestamps, the system first selects a large-scale historical enrollment policy document, consultation Q&A, and official announcements as the raw corpus. It then uses the first-character-middle-character-non-entity sequence labeling method to fine-grainedly label the time entities contained within. Specifically, the character sequence corresponding to the start time of the business effective time is labeled as the start first character and start middle character, and the character sequence corresponding to the end time of the business effective time is labeled as the end first character and end middle character, thus constructing a training sample set containing the input text sequence and the corresponding gold standard label sequence. The model architecture integrates a pre-trained language representation layer to capture semantic features and a bidirectional long short-term memory network to model contextual dependencies. At the output end, a conditional random field layer is used to constrain the transition probabilities between labels. During training, the system defines the global loss function as the negative log-likelihood loss of the true label path relative to the scores of all possible paths. This loss function is minimized through backpropagation to iteratively update the model parameters until the model can accurately identify ambiguous time expressions in complex contexts and output extraction confidence scores for subsequent soft fusion calculations.

[0031] In a preferred embodiment, identifying the data entry effective time difference distribution band based on the difference distribution characteristics between the business effective timestamp and the system entry timestamp includes: Business activation timestamps for computing knowledge data Subtract the system entry timestamp The value is defined as the time difference between data entry and effective date. The calculation formula is: ; Based on the sign of the time difference before the data becomes effective, the knowledge data is divided into a set representing positive time differences indicating early release. and the set of negative time differences representing delayed patches ; Introducing the node proximity coefficient As statistical weights, the distribution zone centers of the positive time difference set and the negative time difference set are calculated respectively. and The calculation formula is: ; ; in, This represents the weighted median function; Introducing the node proximity coefficient As statistical weights, the distribution band widths of the positive time difference set and the negative time difference set are calculated respectively. and The calculation formula is: ; ; in, This represents the weighted absolute median difference function. It is the normal uniformity constant; Depend on Comprising an advance release band, by These constitute a delayed patch band, and together they form the time difference distribution band for the entry taking effect.

[0032] Preferably, the process of identifying the time difference distribution band of entry effectiveness based on the difference distribution characteristics between the business effective timestamp and the system entry timestamp follows the following continuous data processing logic and statistical derivation process.

[0033] The system completes the first Business activation timestamp of knowledge data After parsing, immediately read the system entry timestamp that has been permanently stored in memory. Both are 64-bit floating-point numbers accurate to the second.

[0034] The system first performs difference calculations using the formula. Calculate the time difference between data entry and effective date. Its physical dimension is second, and its range of values ​​is 100%. .

[0035] Immediately afterwards, the system based on The sign bit is used for binary logic splitting, and the condition is satisfied. The data will be included in the set representing the positive time difference released in advance, which will meet the conditions. The data is assigned to the negative time difference set representing the lagging patch.

[0036] To accurately capture the distribution center in noisy dynamic data streams, the system abandons mean statistics, which are sensitive to outliers, and instead uses a weighted median function (wMedian). This function introduces a node proximity coefficient. As a statistical weight, specifically, data that is closer to the key nodes in the business cycle ( The larger the value, the greater its temporal characteristics contribute to the congestion risk of the current window; therefore, it should occupy a higher weight in distribution estimation. The system separately calculates the values ​​in the positive and negative time difference sets. Sort the data and accumulate the corresponding weights. The value at which the accumulated weights reach 50% of the total weights is selected as the center of the distribution band, i.e., the center of the positive distribution band. and negative distribution zone center The calculation formula is defined as follows: and .

[0037] Subsequently, to quantify the dispersion of the distribution, i.e., the bandwidth, the system uses the weighted absolute median difference (wMAD) as a robust scaling estimate. The system calculates the scaling factor for each element within the set. With the center of the corresponding distribution zone The absolute residual is then used again with weights. Calculate the weighted median of the residual sequence and introduce the normality constant. Scaling is applied. The derivation of this constant is based on the Gaussian distribution prior in statistics, that is, the inverse function of the cumulative distribution function of the standard normal distribution. The reciprocal of the value at ( Its function is to ensure that the calculation results are unbiased and consistent with the standard deviation when the data is approximately normally distributed.

[0038] The system is based on the formula and The distribution band width of the positive time difference set was calculated separately. Distribution band width of negative time difference set .

[0039] Finally, the system will calculate the parameter tuple. Instantiate as an early release tape, Instantiated as a delayed patch band, both are marked together in memory as the time difference distribution band for entry effectiveness; if either set is empty, the system initializes the corresponding center and width parameters to 0.

[0040] In a preferred embodiment, calculating the time-series inversion congestion factor, which represents the risk of time-series priority reversal arising near key nodes in the business cycle, includes: Calculate the time difference for the entry and activation of the current knowledge data. The normalized distance relative to the early release band and the late patch band is selected, and the minimum value is defined as the intra-band normalized residual. The calculation formula is: ; in, These are the distribution center zones for the early release zone and the delayed patch zone, respectively. These are the distribution band widths for the early release band and the delayed patch band, respectively. To prevent the removal of minute quantities; Calculate the node proximity coefficient The product of the ... The calculation formula is: ; Preferably, the process of calculating the temporal reversal congestion factor, which represents the risk of temporal priority reversal, near the key nodes of the business cycle, follows the following continuous numerical calculation logic.

[0041] The system completes the construction of the time difference distribution band for entry and effect and obtains the parameters of the pre-release band. With delayed patch parameters Then, immediately trigger the action targeting the current [number]. Risk assessment calculation for each piece of knowledge data. The system reads the cached current knowledge data from memory and the time lag before the data entry takes effect. First, calculate the geometric fit of the time difference with respect to the two distribution zones.

[0042] The system calculates separately With advance release of the center absolute difference and the center of the delayed patch band absolute difference These two differences represent the Euclidean distance of the data deviating from the group trend over time. To eliminate the impact of differences in time span (bandwidth) across different business cycles on the distance metric, the system introduces a denominator for normalization, where the denominator represents the corresponding distribution band width. and Plus prevention of zero micro-quantity Here The preferred value is set as This value is set according to the machine precision limitations of double-precision floating-point numbers under the IEEE 754 standard, aiming to prevent extreme convergence of the distribution band from causing... This prevents division by zero anomalies while ensuring the stability of numerical calculations.

[0043] The system calculates two normalized distances using a formula and utilizes a minimum value function. Select the smaller one as the in-band normalized residual. The calculation formula is defined as follows: Based on the maximum membership prior, that is, knowledge data should statistically belong to the distribution pattern closest to it (whether it is released in advance or patched later). The closer the value is to 0, the more precisely the data falls on a high-density temporal distribution ridge, and the higher the structural risk of congestion. Subsequently, the system performs multi-source feature coupling calculations, introducing a node proximity coefficient. (range of values) ), using formula Calculate the final time-reversal congestion factor. In this formula, This constitutes a Laplacian kernel, which transforms the non-negative distance space... Nonlinear mapping to similarity space ,when The kernel function value is 1, and it decays exponentially with increasing distance. This serves as a spatiotemporal gating term, used to constrain the spatial scope of risk. Specifically, only when the knowledge data simultaneously satisfies the condition of being extremely close to key nodes in the business cycle (…) And located in the center of a high-density distribution zone ( Under these two conditions, This will approach the maximum value of 1, thus accurately identifying high-risk data that floods in during critical time windows and highly conforms to congestion distribution patterns. The final calculated... It is a dimensionless normalized floating-point number with a range of values ​​of 1. This value is written to the priority queue management module of the index building engine in real time. If during the calculation process... If missing, the system defaults. To lift risk control restrictions.

[0044] In a preferred embodiment, the security evolution ranking weight is calculated by combining the time-series reversal congestion factor, the data entry timeliness decay, and the business timeliness fit, including: Calculate the current time Subtract the system entry timestamp The difference is used to calculate the negative exponential decay rate, which yields the data entry time-related decay rate. The calculation formula is: ; in, To input the time-depreciation scale; Calculate the current time With business activation timestamp The absolute difference is used to calculate the business timeliness fit by applying negative exponential decay. The calculation formula is: ; in, To ensure business timeliness aligns with standards; Authority coefficient based on source domain And using the S-shaped function The time-reversed congestion factor Mapped to structured soft gate weights The calculation formula is: ; in, For control parameters, This is the softening temperature; Generate vector representations of the time-constrained text and calculate their average semantic similarity within the local neighborhood. And in conjunction with the time-reversed congestion factor Calculate semantic filtering weights The calculation formula is: ; in, For control parameters, This is the softening temperature; Calculate the time-reversed congestion factor. The negative value of the exponential function is the value of the natural exponential function of the exponent, which is defined as the anti-congestion correction kernel. The calculation formula is: ; in, For folding suppression strength; The final security evolution ranking weight is obtained by multiplying the structured soft gate weight, the semantic filtering weight, the input timeliness decay, the business timeliness fit, and the anti-congestion correction kernel. The calculation formula is: ; Preferably, the process of calculating the security evolution ranking weight by combining the time-series reversal congestion factor, the input timeliness decay degree, and the business timeliness fit degree follows the following steps.

[0045] The system first obtains the current calculation time. (Unix timestamp format), for the first The knowledge data is retrieved by reading its fixed system entry timestamp. Using the formula Calculate the time decay rate of data entry ,in Defined as the data entry time-lapse metric, its value is preferably set to [value to be filled in]. The second (i.e., 7 days) parameter is set based on the information freshness half-life theory, aiming to give higher weight to data newly entered within the most recent week; at the same time, it reads the business activation timestamp. Using the formula Calculate the timeliness of business operations ,in Defined as a business timeliness fit metric, its preferred value is set as follows: The second (i.e., 1 day) parameter is set based on the urgency of the business window, ensuring that weight is tilted towards events that are currently in effect or about to take effect. Subsequently, the system calls the domain reputation database to query the authority coefficient of the source domain. (Normalization to) The interval (e.g., the official education examination authority's domain name is set to 1.0), combined with the time-series inversion congestion factor. Using the S-shaped function Calculate the weight of the structured soft gate The calculation formula is: ,in Based on the threshold (preferred value) ), Congestion penalty gain (preferred value) ), Softening temperature (preferred value) ), to utilize the nonlinear characteristics of the logistic regression function, so that when Exceeding the critical point (Right now When ), weight This results in an avalanche-like decline, thus achieving soft suppression of high-congestion-risk data. Next, the system extracts the time-constrained text (e.g., [valid=...]) from the knowledge data, concatenates it with the main text, and encodes it into a high-dimensional vector using the BERT model. The nearest neighbor is then retrieved from the vector index. vectors ( (Preferred value: 20), calculate the semantic density by its mean cosine similarity. and using the formula Calculate semantic filtering weights ,in (Preferred value) )and (Preferred value) ) used to build congestion The joint suppression plane of density, when the data is in a high congestion zone ( (High) and there is a lot of semantic repetition ( (High) The item will be significantly larger than ,lead to The value approaches 0, thus eliminating the accumulation of homogeneous information at the semantic level. Furthermore, the system directly calculates the anti-congestion correction kernel for congestion risk. The formula is ,in For the folding suppression strength, the preferred value is set as follows: This parameter, set based on risk preference, is used to apply a final exponential penalty to high-risk data. Finally, the system multiplies the above components together, i.e. We obtain dimensionless and normalized (theoretically close to but less than 1) safe evolutionary ranking weights. This ensures the system's stability during the knowledge base's evolution, prioritizing authority, timeliness, and preventing congestion.

[0046] It should be noted that the system constructs a high-dimensional vector index based on the hierarchical navigation small-world graph algorithm. The local neighborhood is defined as the set of K nearest neighbor nodes in the vector space that are closest to the semantic vector of the current knowledge data in terms of Euclidean distance. The hyperparameter K is preferably set to 20 to balance the computational cost and the representativeness of local features. When calculating the semantic density, the system uses the semantic vector of the current data as the query vector and performs an approximate nearest neighbor search in the index to obtain the above set of neighbor nodes. Then, it calculates the cosine similarity between the query vector and each neighbor vector in the set and takes the arithmetic mean. This average value is the semantic density, which represents the degree of congestion in the semantic space where the current data is located. If the density is too high, it means that there is a lot of homogeneous information in that semantic direction. The system will trigger an anti-congestion mechanism to suppress the weight of the current data accordingly.

[0047] It should be noted that the system uses a key-value pair-based structured template generation mechanism to construct time-constrained text. First, the parsed business start and end timestamps are converted into strings in standard ISO8601 date format. Then, according to a fixed template, these are filled into specific text fragments with valid start and end times equal to the date string. A special delimiter is used to prepend these fragments to the beginning of the original knowledge data text. When this combined text is fed into the tokenizer of the pre-trained language model for word segmentation, the specific time key-value pairs are mapped to independent position and paragraph codes. This forces the semantic vector to explicitly encode the time interval constraint information in the embedding space, ensuring that subsequent vector retrieval can perceive the semantic boundaries of the time dimension, thereby distinguishing knowledge items with similar content but different start times at the retrieval level.

[0048] It should be noted that during the initialization phase, the system constructs and maintains a hierarchical domain reputation mapping table. This table divides internet domains into three levels of authority. Official domains of government and educational institutions ending in .gov.cn or .edu.cn are marked as Level 1 authoritative sources and assigned a coefficient of 1.0. Whitelisted domains certified by mainstream news media organizations are marked as Level 2 trusted sources and assigned a coefficient of 0.8. Other general commercial domains not on the whitelist are marked as Level 3 ordinary sources and assigned a default coefficient of 0.5. During the calculation process, the system extracts the main domain portion of the knowledge data source URL as the query key and retrieves the corresponding coefficient value in the mapping table. If a match fails, a low-confidence downgrade strategy is triggered, automatically assigning the authority coefficient of the data to 0.1. This quantifies the differences in the contribution of different information sources to the credibility of the knowledge base evolution at the source, ensuring that high-authority information occupies a dominant position in the evolution process.

[0049] In a preferred embodiment, incremental data fusion is performed on the target nodes of the multi-granularity temporal index structure based on the security evolution sorting weight, and the knowledge base inverted index containing two-dimensional temporal metadata is updated, including: Based on the business activation timestamp of the knowledge data Locate its corresponding leaf node in the multi-granularity temporal index structure. ; Using the security evolution ranking weight As a weighting coefficient, for those falling into the same leaf node semantic vectors of all knowledge data within Perform a weighted average calculation to generate the aggregated summary vector for that leaf node. The calculation formula is: ; The aggregated summary vector is passed up and the summary representation of the ancestor node in the multi-granularity temporal index structure is updated to achieve incremental data fusion; Construct a two-layer index structure comprising a time index layer and a semantic index layer, wherein the time index layer locates the time neighborhood based on the multi-granularity temporal index structure, and the semantic index layer stores the semantic vectors of the knowledge data. ; The security evolution ranking weight The time-reversed congestion factor The two-dimensional time metadata is associated and stored in the inverted list of the semantic index layer, and the inverted list is configured for retrieval and querying. The results are then rearranged by multiplying semantic similarity by the security evolution ranking weight, and the scoring formula is as follows: ; Preferably, the process of incrementally fusing data on the target nodes of the multi-granularity temporal index structure and updating the knowledge base inverted index containing two-dimensional temporal metadata according to the security evolution sorting weight follows the following continuous vector space operation and data structure update logic.

[0050] The system first reads the current number Business activation timestamp of knowledge data A tree traversal is performed in the memory-resident multi-granularity time-series index structure, matching year, month, and day nodes sequentially from top to bottom, ultimately pinpointing the unique leaf node to which the timestamp belongs. Subsequently, the system performs incremental vector aggregation calculation based on security weights to update the semantic representation of the leaf nodes to reflect the mainstream high-confidence information features within that time window.

[0051] The system reads the leaf node. The currently stored cumulative weights With cumulative vector sum Combined with the semantic vector of current knowledge data (A 768-dimensional Float32 floating-point vector generated by the BERT model) and the safe evolutionary ordering weights calculated in the preceding steps. Update the accumulated value using atomic operations: and .

[0052] Next, the system uses the formula Calculate the updated aggregated summary vector This formula is equivalent to the one in the image. Based on the weighted centroid shift mechanism, by introducing... As a quality coefficient, it forces the leaf node summary vector In the vector space, data clusters are moved towards centers that are timely, authoritative, and have low congestion risk, thereby automatically suppressing the interference of low-weight patch data on node semantics.

[0053] After the leaf node update is completed, the system triggers a bottom-up recursive propagation mechanism to update the new leaf node. This serves as input to update the summary representation of its parents up to the root node, enabling incremental data fusion across the entire tree.

[0054] Simultaneously, the system maintains a two-level index structure, where the time index layer directly reuses the aforementioned multi-granularity structure to achieve... The system employs time-neighborhood pruning to reduce complexity, and the semantic indexing layer uses an inverted file structure (IVF). When writing to the inverted list, force the safe evolutionary sort weights. Time-reversed congestion factor and two-dimensional time metadata The payload is stored together with the document ID as a compact tuple.

[0055] During the retrieval phase, when the system receives a query And encoded as a vector Then, for the first candidate in the candidate set For each data point, the system first calculates the cosine similarity. Then immediately perform the multiplication rearrangement operation, according to the formula. The final score is obtained.

[0056] This scoring formula constructs a semantic-security joint metric space, where Characterize content relevance, while As a scalar gating mechanism, this is equivalent to amplitude modulation of semantic similarity. If the data... Due to high congestion risk ( Even with a high semantic match, a score close to 0 (either high or expired) will result in a low final score. It will also be forcibly lowered, thereby ensuring that the search results follow the system's safety evolution strategy based on relevance.

[0057] Specifically, all floating-point calculations throughout the process employ the IEEE 754 standard double-precision format to reduce accumulated errors; the vector dimension is fixed at 768 dimensions; and the cosine similarity calculation undergoes L2 normalization preprocessing to ensure that the values ​​are within acceptable limits. Within the range.

[0058] Example 2: Figure 2 As shown, a knowledge security evolution and update system for newly added information, applied in any of the knowledge security evolution and update methods for newly added information, is characterized by comprising: The time series parsing and construction module is used to build a multi-granularity time series index structure containing key nodes of the business cycle, mark the collected knowledge data with system-entered timestamps, and parse the business effective timestamps; The risk identification and factor calculation module is used to identify the time difference distribution band of entry and entry based on the difference distribution characteristics between the business entry time stamp and the system entry time stamp, and to calculate the time-series reversal congestion factor that represents the risk of time-series priority reversal of knowledge data near the key nodes of the business cycle. The weight calculation module is used to calculate the security evolution ranking weight by combining the time-inversion congestion factor, the input timeliness decay degree and the business timeliness fit degree, wherein the time-inversion congestion factor is configured to generate a structured inhibition parameter for the security evolution ranking weight. The update and fusion module is used to perform incremental data fusion on the target nodes of the multi-granularity time-series index structure according to the security evolution sorting weight, and update the knowledge base inverted index containing two-dimensional time metadata.

[0059] It is important to note that all input data described in this solution is acquired in real-time through legal and compliant hardware interfaces with the user's full knowledge, explicit consent, and active cooperation. The preset parameters, prior constants, and statistical means are all derived from publicly available scientific literature data, de-identified general research datasets, or calibration data from laboratory environments, and do not contain any unauthorized sensitive third-party information. The system's data processing is limited to local or volatile memory computation transmitted via encrypted channels. There is no illegal collection, theft, or retention of user biometric data or infringement of user privacy without the user's knowledge. All parameter calls and generation comply with the principles of data minimization, legality, legitimacy, and necessity.

[0060] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.

Claims

1. A knowledge security evolution updating method for new information, characterized in that, include: Construct a multi-granularity time-series index structure that includes key nodes in the business cycle, and mark the collected knowledge data with system-entered timestamps and parse the business effective timestamps; Based on the difference distribution characteristics between the business activation timestamp and the system entry timestamp, the entry activation time difference distribution band is identified, and the time sequence reversal congestion factor, which characterizes the risk of time sequence priority reversal of knowledge data near the key nodes of the business cycle, is calculated. The security evolution ranking weight is calculated by combining the time-inversion congestion factor, the data entry timeliness decay, and the business timeliness fit. The time-inversion congestion factor is configured to generate a structured inhibition parameter for the security evolution ranking weight. Based on the security evolution sorting weight, incremental data fusion is performed on the target nodes of the multi-granularity time-series index structure, and the knowledge base inverted index containing two-dimensional time metadata is updated.

2. The knowledge security evolution and update method for newly added information according to claim 1, characterized in that, Construct a multi-granularity time-series index structure that includes key nodes in the business cycle, including: Establish a hierarchical index tree that includes year, month, and day time level nodes and key business cycle nodes mounted at the bottom layer; Calculate the absolute time difference between the business activation timestamp of the collected knowledge data and the most recent key node of the business cycle, and use this absolute time difference as the business activation interval; Based on the effective time interval of the business, the node proximity coefficient, which represents the correlation strength between knowledge data and key nodes in the business cycle, is calculated using an exponential decay function, and the node proximity coefficient is stored in the index metadata.

3. The knowledge security evolution and update method for newly added information according to claim 1, characterized in that, The system records and parses the timestamps of the collected knowledge data, including: The moment when the system collects the current knowledge data is recorded is identified and marked as the system's input timestamp. The information extraction model is used to identify candidate effective timestamps from the content of knowledge data and the corresponding extraction confidence scores are output simultaneously. Based on the extracted confidence level, a weighted summation is performed on the candidate effective timestamp and the system input timestamp to obtain the final business effective timestamp. When the extraction confidence level approaches zero, the business activation timestamp is numerically and smoothly rolled back to the system entry timestamp to maintain the continuity of the parsing process.

4. The knowledge security evolution and update method for newly added information according to claim 2, characterized in that, Based on the difference distribution characteristics between the business activation timestamp and the system entry timestamp, the data entry activation time difference distribution band is identified, including: The business effective timestamp of the knowledge data is subtracted from the system entry timestamp, and this value is used as the entry effective time difference; Based on the sign of the effective time difference, the knowledge data is divided into a set of positive time differences representing early releases and a set of negative time differences representing delayed patches; Using the node proximity coefficient as a statistical weight, the weighted median of the effective entry time difference in the positive time difference set and the negative time difference set is calculated respectively, and this median is used as the center of the distribution zone. The node proximity coefficient is introduced as a statistical weight. The weighted absolute median difference based on the weighted median is calculated for the entry effective time difference in the positive time difference set and the negative time difference set, respectively. This difference is then multiplied by the normal consistency constant and defined as the distribution band width. The distribution band corresponding to the positive time difference set is used to form the advance release band, and the distribution band corresponding to the negative time difference set is used to form the lag patch band. Together, they form the input effective time difference distribution band.

5. A knowledge security evolution and update method for newly added information according to claim 4, characterized in that, The calculation of the temporal inversion congestion factor, which represents the risk of temporal priority reversal near the key nodes of the business cycle, includes: Calculate the absolute difference between the time difference of the current knowledge data entry and the distribution center of the pre-release zone, divide it by the corresponding distribution zone width, and obtain the positive normalized distance; Calculate the absolute difference between the time difference of the current knowledge data entry and the distribution center of the lag patch band, divide it by the corresponding distribution band width, and obtain the negative normalized distance; The minimum value between the positive normalized distance and the negative normalized distance is selected as the in-band normalized residual; The product of the node proximity coefficient and the value of the natural exponential function with the negative of the in-band normalized residual as the exponent is calculated and defined as the time-inverted congestion factor.

6. The knowledge security evolution and update method for newly added information according to claim 1, characterized in that, Combining the aforementioned time-series reversal congestion factor, data entry timeliness decay, and business timeliness fit, the security evolution ranking weight is calculated, including: Calculate the difference between the current time and the system entry timestamp, apply negative exponential decay calculation to it, and obtain the entry timeliness decay degree; Calculate the absolute difference between the current time and the service activation timestamp, and apply negative exponential decay calculation to it to obtain the service timeliness fit. Based on the authority coefficient of the source domain name, and using the S-shaped function to map the time-inversion congestion factor to a structured soft gate weight with a value between zero and one, wherein the larger the time-inversion congestion factor, the smaller the structured soft gate weight. Generate a vector representation of the time-constrained text, calculate its average semantic similarity in the local neighborhood as the semantic density, and combine the time-inversion congestion factor with the S-shaped function to calculate the semantic filtering weight. Calculate the value of the natural exponential function with the negative of the time-reversal congestion factor as the exponent, and use it as the anti-congestion correction kernel; The security evolution ranking weight is obtained by multiplying the structured soft gate weight, the semantic filtering weight, the input timeliness decay, the business timeliness fit, and the anti-congestion correction kernel.

7. The knowledge security evolution and update method for newly added information according to claim 1, characterized in that, Based on the security evolution ranking weight, incremental data fusion is performed on the target nodes of the multi-granularity temporal index structure, and the knowledge base inverted index containing two-dimensional temporal metadata is updated, including: Based on the business activation timestamp of the knowledge data, locate the corresponding leaf node in the multi-granularity time-series index structure; Using the security evolution ranking weight as a weighting coefficient, the semantic vectors of all knowledge data falling into the same leaf node are weighted and averaged to generate the aggregated summary vector of that leaf node. The aggregated summary vector is passed up and the summary representation of the ancestor node in the multi-granularity temporal index structure is updated to achieve incremental data fusion; A two-layer index structure is constructed, comprising a time index layer and a semantic index layer, wherein the time index layer locates the time neighborhood based on the multi-granularity temporal index structure, and the semantic index layer stores the semantic vectors of knowledge data. The secure evolutionary ranking weight, the temporal reversal congestion factor, and the two-dimensional temporal metadata are associated and stored in the inverted list of the semantic index layer. The inverted list is configured to rearrange the results during retrieval by multiplying the semantic similarity with the secure evolutionary ranking weight.

8. A knowledge security evolution and update system for newly added information, applied in the knowledge security evolution and update method for newly added information as described in any one of claims 1-7, characterized in that, include: The time series parsing and construction module is used to build a multi-granularity time series index structure containing key nodes of the business cycle, mark the collected knowledge data with system-entered timestamps, and parse the business effective timestamps; The risk identification and factor calculation module is used to identify the time difference distribution band of entry and entry based on the difference distribution characteristics between the business entry time stamp and the system entry time stamp, and to calculate the time-series reversal congestion factor that represents the risk of time-series priority reversal of knowledge data near the key nodes of the business cycle. The weight calculation module is used to calculate the security evolution ranking weight by combining the time-inversion congestion factor, the input timeliness decay degree and the business timeliness fit degree, wherein the time-inversion congestion factor is configured to generate a structured suppression parameter for the security evolution ranking weight. The update and fusion module is used to perform incremental data fusion on the target nodes of the multi-granularity time-series index structure according to the security evolution sorting weight, and update the knowledge base inverted index containing two-dimensional time metadata.