Big data text relation extraction management system based on knowledge graph
By introducing hotspot interaction index and event propagation speed index, combined with machine learning evaluation, and dynamically adjusting the data update frequency, the lag problem of timed batch update strategy in highly dynamic data processing is solved, realizing real-time information and efficient resource utilization, and improving the accuracy and efficiency of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-27
AI Technical Summary
When processing highly dynamic data such as financial data or trending news, existing technologies struggle to capture real-time changes through scheduled batch updates, leading to information lag and impacting decision-making accuracy and system efficiency.
We introduce hotspot interaction index and event propagation speed index for dynamic feature analysis, combine machine learning evaluation, divide high-frequency and low-frequency data, dynamically adjust the update frequency, and adopt non-linear weighted difference and exponential adjustment optimization update strategy.
It improves the system's responsiveness to highly dynamic data, ensures the timeliness and accuracy of information, optimizes resource utilization, reduces unnecessary resource waste, and enhances the accuracy and efficiency of the question-and-answer system and search engine.
Smart Images

Figure CN121743469A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of text extraction technology, specifically to a big data text relationship extraction and management system based on knowledge graphs. Background Technology
[0002] Knowledge graph-based big data text relationship extraction and management refers to the automatic identification and extraction of relationships between entities from massive amounts of text data in a big data environment using natural language processing technology, and the structured storage of this information in a knowledge graph for management and application. This process includes data preprocessing, text analysis, entity and relationship extraction, and graph construction and optimization. Through knowledge graph management, the information contained in text can be effectively organized and linked, thereby improving information retrieval, recommendation, and reasoning capabilities. Text relationship extraction and management in a big data environment can not only handle massive amounts of information but also dynamically update and optimize the graph structure with new data input, providing powerful knowledge support for applications such as decision support and intelligent question answering.
[0003] Knowledge graph-based big data text relationship extraction and management can be widely applied in intelligent question answering and search engines. Its core function lies in extracting entities and their relationships from massive amounts of text, transforming unstructured data into a structured knowledge graph, and providing question answering systems and search engines with deep semantic understanding and accurate knowledge matching capabilities. This technology helps intelligent question answering systems accurately interpret user intent and generate more relevant answers; simultaneously, search engines use knowledge graphs to support contextualized searches and multi-level information reasoning, thereby improving the accuracy and richness of retrieval. This capability is particularly valuable in complex queries, cross-domain question answering, and real-time dynamic information needs.
[0004] The existing technology has the following shortcomings:
[0005] When transforming unstructured data into structured knowledge graphs, a timed batch update strategy is typically employed, which involves updating the captured data at fixed time intervals. This strategy is effective for data with low update frequency and long change cycles, efficiently meeting the needs. However, for rapidly changing data such as financial data and trending news, timed batch updates have significant limitations: the fixed update cycle struggles to capture real-time changes, and when users require the latest information, the system generates answers based on outdated data, leading to a significant decrease in the accuracy and relevance of the information. Especially in highly sensitive areas such as financial markets or public opinion monitoring, delayed information can mislead user decisions, causing serious economic losses or other adverse consequences.
[0006] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0007] The purpose of this invention is to provide a knowledge graph-based big data text relationship extraction and management system. By introducing dynamic feature analysis such as hotspot interaction index and event propagation speed index, and intelligent evaluation based on machine learning, the system accurately identifies the real-time demand for highly dynamic data (such as financial data or hot news), classifies it into high-frequency update data, and dynamically adjusts the update frequency. Compared with traditional timed batch update strategies, this mechanism significantly improves the system's real-time response capability and information accuracy, avoiding misleading decisions caused by data lag. Simultaneously, the system classifies data using a heat coefficient and reference threshold, optimizes resource allocation, adopts a timed update strategy for low-frequency update data, and dynamically increases the update frequency of high-frequency data. This significantly improves system efficiency and resource utilization while achieving real-time performance and accuracy, thus solving the problems mentioned in the background technology.
[0008] To achieve the above objectives, the present invention provides the following technical solution: a knowledge graph-based big data text relationship extraction and management system, comprising a data acquisition module, a data processing and feature extraction module, a feature analysis and machine learning evaluation module, a data classification module, a low-frequency data timed update module, and a high-frequency data dynamic update module:
[0009] The data acquisition module first obtains data information from various massive unstructured text data sources;
[0010] The data processing and feature extraction module processes the acquired data into a structured analysis set and extracts key features that reflect the dynamic changes of the data from the analysis set.
[0011] The feature analysis and machine learning evaluation module analyzes and processes the extracted features under the monitoring window, and inputs the analyzed key feature information into the pre-learned machine learning model, which then performs intelligent evaluation of the acquired data information.
[0012] The data classification module, based on the evaluation results of the machine learning model, divides the acquired relevant data information into high-frequency update data and low-frequency update data;
[0013] The low-frequency data timed update module continues to use the preset timed batch update strategy for data that is classified as low-frequency update.
[0014] The high-frequency data dynamic update module dynamically increases the update frequency of data classified as high-frequency updates based on the evaluation results of the machine learning model under the monitoring window and a preset timed batch update strategy.
[0015] Preferably, key features reflecting dynamic changes in data are extracted from the analysis set, including the frequency of user interaction with data content and the speed of event-related information propagation on the network. After acquisition, under the detection window, the frequency of user interaction with data content and the speed of event-related information propagation on the network are analyzed to generate a hot topic interaction index and an event propagation speed index. The hot topic interaction index quantifies the intensity and frequency of user interaction with a certain data content, reflecting the popularity of the content and its trend of becoming a hot topic event. The event propagation speed index quantifies the diffusion rate and coverage of event-related information on the network platform, reflecting the event's propagation capability and real-time influence.
[0016] Preferably, after obtaining the hotspot interaction index and event propagation speed index generated by analyzing the extracted key features, the hotspot interaction index and event propagation speed index are input into a pre-learned machine learning model. The machine learning model generates a data popularity coefficient, and the data popularity coefficient is used to intelligently evaluate the obtained data information.
[0017] Preferably, the data popularity coefficient generated after analyzing the extracted key features under the monitoring window is compared with a pre-set data popularity coefficient reference threshold, and the acquired relevant data information is divided. The division steps are as follows:
[0018] If the data popularity coefficient is greater than or equal to the preset data popularity coefficient reference threshold, the relevant data information obtained will be classified as high-frequency update data.
[0019] If the data popularity coefficient is less than the preset data popularity coefficient reference threshold, the relevant data information will be classified as low-frequency update data.
[0020] Preferably, for data classified as high-frequency updates, based on the evaluation results of the machine learning model under the monitoring window, and based on a preset timed batch update strategy, the update frequency is dynamically increased. The specific steps are as follows:
[0021] By comparing the data popularity coefficient DPHC with a reference threshold for the data popularity coefficient using a non-linear weighted difference, we can assess whether the current data popularity exceeds the reference level. The calculation expression is as follows:
[0022] ,
[0023] In the formula, Δ DPHC DPHC is a weighted data popularity bias, used to measure the overall deviation between the current popularity of data and the benchmark. ref This is the reference threshold for the data popularity coefficient; w1 and w2 are both weighting coefficients.
[0024] By introducing nonlinear transformation and normalization, the update priority of the data is calculated, enabling the system to more accurately measure the impact of data popularity deviation. The calculation expression is as follows:
[0025] ,
[0026] In the formula, UP is the update priority, and tanh(k·|Δ DPHC |) represents the effect of smoothing extreme values using the hyperbolic tangent function, where b controls the sensitivity of the smoothing, and max(|Δ) DPHC |,1) is the normalization factor, sign(Δ DPHC () is the symbol for thermal deviation;
[0027] Based on the preset timed batch update frequency and update priority UP, the actual update frequency is dynamically calculated through exponential adjustment and nonlinear functions. The calculation expression is as follows:
[0028] ,
[0029] In the formula, F dynamic It is the dynamic update frequency, F base τ is the preset timed batch update frequency, τ is the exponential adjustment coefficient, which controls the overall influence range of update priority on update frequency, σ is the fluctuation adjustment coefficient, which controls the influence of dynamic changes on fine frequency adjustments, and e is the natural base.
[0030] The actual performance after monitoring data updates is assessed, and the deviation between the actual update frequency and the theoretical optimal frequency is calculated using the following expression:
[0031] ,
[0032] In the formula, ∈ is the update frequency error, F actual It is the actual update frequency, F optimal This is the theoretically optimal update frequency, calculated as follows:
[0033] ,
[0034] In the formula, D is the dynamic adjustment coefficient, which determines the magnitude of the impact of data dynamics on the optimal frequency;
[0035] Based on the current update frequency error ∈, the update strategy is further improved by adaptively optimizing the exponential adjustment coefficient τ and the fluctuation adjustment coefficient σ of the dynamic update frequency model. The calculation expression is as follows:
[0036] τ new =τ current ·(1-H·∈), σ new =σ current ·(1-H·∈),
[0037] In the formula, τ new It is the optimized exponential adjustment coefficient, σ new It is the optimized fluctuation adjustment coefficient, τ current This is an adjustment factor, currently used to control the overall impact of the Data Heat Factor (DPHC) on the update frequency. σ current It is a sensitivity coefficient, currently used to control the bias Δ of the weighted data heat. DPHC For the coefficient of sensitivity to update frequency, H is the step size coefficient, which is used to control the sensitivity of adjustment;
[0038] The optimized exponential adjustment coefficient τ new and the optimized fluctuation adjustment coefficient σ new Re-enter the dynamic update frequency calculation formula to improve the update strategy. The calculation expression is as follows:
[0039] F dynamic,new =F base ·(1+τ new ·tanh(σ new ·|Δ DPHC |)),
[0040] In the formula, F dynamic,new It refers to the updated dynamic update frequency.
[0041] Preferably, under the monitoring window, the specific steps for analyzing the frequency of user interaction with data content and generating a hotspot interaction index are as follows:
[0042] In the monitoring window, extract interaction behavior data related to the data content from the interaction log to form a behavior set B, where B = {b}. i} = {b1, b2, b3, ..., n}, where b i This refers to the i-th interaction, including the number of clicks, likes, comments, shares, and the time point of the interaction. n is the total number of interactions, and each interaction is represented by b. i The representation is as follows: b i =(C i L i R i S i T i In the formula, C i It refers to the number of times users click on data content, L i It is the number of times users like data content, R i S is the number of times users comment on the data content. i It is the number of times a user shares data content, T i It is the point in time when the interaction occurs;
[0043] To reflect the real-time nature of user interaction frequency, a time decay weight is introduced. The weight value is defined by an exponential function, and the calculation expression is as follows: In the formula, W(T) i ) is the time decay weight, e is the natural base, λ is the time decay factor that controls the decay rate, and t2 is the end time of the monitoring window.
[0044] b for each interaction behavior i The weights are associated with the numerical values of the behaviors. The weighted behavior value is calculated using the following expression:
[0045] B′ i =W(T) i )·(C i +αL i +βR i +γS i )
[0046] In the formula, B i ' is the weighted value of the i-th interaction behavior, where α, β, and γ are the behavior importance weights, used to adjust the number of times the user likes the data content, L. i R, the number of times users comment on the data content i And the number of times users share data content (S) i Impact on the hot topic interaction index;
[0047] To capture the complex characteristics of data content in user interactions, an interaction behavior complexity factor is defined to quantify the correlation between various behaviors. The calculation expression is as follows:
[0048] ,
[0049] In the formula, H(B') is the interaction complexity factor. This is the Shannon information entropy formula, p(b i The weighted value of the i-th interaction behavior is the proportion of the total weighted value, and the calculation expression is as follows:
[0050] ,
[0051] In the formula, B j ′ is the weighted value of the j-th interaction behavior. It is the weighted sum of all interactive behaviors;
[0052] The weighted value B of the i-th interaction behavior i The interaction complexity factor H(B') and the time decay weight W(T) i The hot topic interaction index is generated, and the calculation expression is as follows:
[0053] ,
[0054] In the formula, WDHII is the hotspot interaction index, and κ is the index normalization factor.
[0055] Preferably, under the monitoring window, the specific steps for analyzing the propagation speed of event-related information in the network and generating an event propagation speed index are as follows:
[0056] Under the monitoring window, raw data related to the event is collected from multiple network data sources, and propagation nodes are identified. The calculation expression is as follows:
[0057] ,
[0058] In the formula, N p It is the number of propagation nodes, ω m f is the weight of each data source m. m (t) is the time function of the event-related propagation point in the m-th data source, which records the number of propagation behaviors of the data source in time t, and M is the total number of data sources;
[0059] Analyze the propagation path of event-related information and quantify the propagation depth, i.e., the hierarchical structure of information spreading from the initial propagation node to other nodes. The calculation expression is as follows:
[0060] ,
[0061] In the formula, D p It is a propagation depth index, where L is the total number of propagation levels, and ω is the propagation depth index. l V is the influence factor of the l-th layer of propagation, used to measure the influence weight of each layer of propagation. l It is the number of propagation nodes in the l-th layer, that is, the number of independent nodes that receive information in each layer;
[0062] The propagation speed, i.e., the speed at which information spreads per unit time, is calculated using the increment of propagation nodes and the time interval as the basis. The calculation expression is as follows:
[0063] ,
[0064] In the formula, V p It is the propagation speed index, and Q is the total number of time segments. It is the propagation rate weight of the q-th time segment, ΔN q The number of newly added propagation nodes in the q-th time segment, Δt q It is the duration of the q-th time segment;
[0065] Combined with the number of propagation nodes N p ,D spread depth index p and the propagation speed index V pGenerate the event propagation speed index, calculated using the following expression:
[0066] ,
[0067] In the formula, NEPVI is the event propagation speed index, and S, Y, and Z are all weighting coefficients used to balance the number of propagation nodes N. p ,D spread depth index p and the propagation speed index V p The overall impact on the speed of event propagation.
[0068] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0069] This invention, by introducing the analysis of dynamic data characteristics (such as hot topic interaction index and event propagation speed index) and intelligent evaluation based on machine learning, enables the system to accurately identify the real-time needs of highly dynamic data (such as financial data or trending news), classify it into high-frequency update data, and dynamically adjust the update frequency. Compared to traditional timed batch update strategies, this dynamic update mechanism significantly improves the system's responsiveness to rapidly changing data. For example, for financial market data, the system can capture high-profile events such as price fluctuations or changes in trading volume, update relevant content in a timely manner, and ensure that the answers generated by the question-and-answer system and search engine are based on the latest information, avoiding misleading decisions due to data lag. In highly sensitive areas, such as financial investment or public opinion monitoring, improved real-time performance directly affects users' economic interests and public opinion response strategies, significantly enhancing the system's accuracy and credibility.
[0070] This invention effectively categorizes data into high-frequency and low-frequency update types through a data popularity coefficient and reference threshold mechanism. For low-frequency update data, a timed batch update strategy continues to be used to reduce unnecessary resource waste; for high-frequency update data, the update frequency is dynamically increased, thereby achieving optimal resource utilization under limited computing resources and bandwidth. In this way, the system can meet the real-time requirements of highly dynamic data while conserving resources on low-dynamic data, ensuring overall operational efficiency. For example, dynamic updates of trending news focus only on truly high-profile events, avoiding repeated fetching of less popular information, thus reducing system load. While achieving real-time information delivery, the system optimizes computing resource allocation, ensuring efficient operation while flexibly adapting to dynamic needs. Attached Figure Description
[0071] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0072] Figure 1 This is a schematic diagram of the modules of the knowledge graph-based big data text relationship extraction and management system of the present invention. Detailed Implementation
[0073] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0074] This invention provides, for example Figure 1 The knowledge graph-based big data text relationship extraction and management system shown includes a data acquisition module, a data processing and feature extraction module, a feature analysis and machine learning evaluation module, a data classification module, a low-frequency data timed update module, and a high-frequency data dynamic update module.
[0075] The data acquisition module first obtains data information from various massive unstructured text data sources;
[0076] Acquiring data from a vast amount of unstructured text data sources refers to using technical means to collect relevant information from scattered and unorganized data sources, providing raw data input for subsequent processing and analysis. These data sources typically include news websites, social media platforms, forums, blogs, electronic documents, and corporate reports. The content exists in unstructured form, such as text paragraphs, articles, and conversation logs. Compared to structured data (such as tabular information in a database), unstructured data requires more processing steps to extract useful information. The process of acquiring data involves using methods such as web crawlers, API interfaces, and data stream subscriptions to collect data, ensuring the breadth and representativeness of the data to cover all information that may affect the construction of the knowledge graph.
[0077] The data processing and feature extraction module processes the acquired data into a structured analysis set and extracts key features that reflect the dynamic changes of the data from the analysis set.
[0078] The feature analysis and machine learning evaluation module analyzes and processes the extracted features under the monitoring window, and inputs the analyzed key feature information into the pre-learned machine learning model, which then performs intelligent evaluation of the acquired data information.
[0079] Key features reflecting dynamic changes in data are extracted from the analysis set, including the frequency of user interaction with data content and the speed of event-related information propagation on the network. After acquisition, under the detection window, the frequency of user interaction with data content and the speed of event-related information propagation on the network are analyzed to generate a hot topic interaction index and an event propagation speed index. The hot topic interaction index quantifies the intensity and frequency of user interaction with a certain data content, reflecting the popularity of the content and its trend of becoming a hot topic. The event propagation speed index quantifies the diffusion rate and coverage of event-related information on the network platform, reflecting the event's propagation ability and real-time impact.
[0080] After obtaining the hotspot interaction index and event propagation speed index generated by analyzing the extracted key features, the hotspot interaction index and event propagation speed index are input into a pre-learned machine learning model. The machine learning model generates a data popularity coefficient, and the data popularity coefficient is used to intelligently evaluate the obtained data information.
[0081] A pre-trained machine learning model refers to a model that has already been trained on historical data and possesses predictive or classification capabilities. It can generate meaningful outputs (such as a data popularity coefficient) based on input feature data (such as hotspot interaction indices and event propagation speed indices). The training process typically involves collecting and organizing large amounts of historical data as a training set, labeling or classifying the data, and then using algorithms such as supervised learning, unsupervised learning, or reinforcement learning to optimize the model, enabling it to recognize patterns and regularities in the data. For example, a model for predicting event popularity might use data from historical hot events, inputting past interaction frequency and propagation speed features, and then training the model to output the event's popularity coefficient. After training, the model possesses generalization ability and can make reasonable predictions or classifications for unseen input data.
[0082] The core of this model lies in learning the mapping relationship between features and outcomes. For example, by analyzing surges in interaction frequency and rapid spread of information, the model can determine the potential popularity of an event and the frequency at which it needs to be updated. Model training relies on high-quality training data and carefully designed feature engineering to ensure the model can capture the dynamic changes in key features. Furthermore, to improve prediction accuracy, hyperparameter tuning may be performed, and performance evaluation may be conducted using validation and test sets to ensure its robustness in real-world applications. Once deployed in a real-world system, the model can receive input features in real time and efficiently generate data popularity coefficients, assisting the system in intelligently assessing whether data requires frequent updates and providing a basis for the dynamic update strategies of question-answering systems and search engines.
[0083] A surge in user interaction frequency typically indicates that the extracted data is rapidly changing, as a significant increase in interaction frequency is usually highly correlated with the real-time and dynamic nature of the data content. When users focus on a particular type of data and interact with it frequently (e.g., clicks, comments, shares), it often means that the data is changing rapidly or contains new key information that attracts users' immediate needs. For example, trending news, financial market dynamics, and social media topics are typical examples of rapidly changing data categories, and the timeliness and update frequency of this data directly affect users' decisions and experience. A surge in interaction frequency may also reflect significant changes in the data content within a short period, or a rapid expansion of its relevance and influence. If the system fails to update this data in a timely manner, it will lead to delayed or inaccurate information for users, seriously affecting users' trust in the system. Therefore, a surge in user interaction frequency can serve as an important signal of data dynamism, indicating that this type of data needs frequent updates to meet users' real-time needs and ensure the accuracy and relevance of the information.
[0084] The specific steps for analyzing the frequency of user interaction with data content and generating a hotspot interaction index within the monitoring window are as follows:
[0085] In the monitoring window, extract interaction behavior data related to the data content from the interaction log to form a behavior set B, where B = {b}. i} = {b1, b2, b3, ..., n}, where b i This refers to the i-th interaction, including the number of clicks, likes, comments, shares, and the time point of the interaction. n is the total number of interactions, and each interaction is represented by b. i The representation is as follows: b i =(C i L i R i S i T i In the formula, Ci It refers to the number of times users click on data content, L i It is the number of times users like data content, R i S is the number of times users comment on the data content. i It is the number of times a user shares data content, T i It is the point in time when the interaction occurs;
[0086] Constructing a comprehensive dataset of interactive behaviors is a crucial step in analyzing user interaction patterns with data content and a vital foundation for subsequent feature extraction of dynamic data changes and calculation of hotspot interaction indices. This dataset should include records of all user interactions with the target data, such as clicks, likes, comments, and shares, recording detailed information such as the frequency and timestamps of each behavior to capture the temporal and diverse nature of these interactions. By constructing this behavior set, scattered user interaction data can be systematically organized and structured, providing complete input for subsequent feature analysis and calculations while ensuring the comprehensiveness and operability of the behavioral data. This not only helps identify user focus points on the data but also provides a solid basis for formulating intelligent update strategies, enabling the system to more accurately adjust the data update frequency, thereby improving the efficiency and accuracy of dynamic data processing.
[0087] To reflect the real-time nature of user interaction frequency, a time decay weight is introduced. The weight value is defined by an exponential function, and the calculation expression is as follows: In the formula, W(T) i ) is the time decay weight, e is the natural base, λ is the time decay factor that controls the decay rate, and t2 is the end time of the monitoring window.
[0088] Time decay weighting is a weighting mechanism that dynamically adjusts the importance of an interaction's contribution to the hot topic interaction index based on the distance between the time the interaction occurred and the current time. Using the exponential decay formula, the closer the time is to the current moment, the higher the weight W(T). i The higher the weighting, the lower the weighting. This mechanism highlights the impact of recent interactions while de-emphasizing past actions, thus more accurately reflecting the dynamic changes of data content within the current monitoring window. Here, the time decay weighting helps the hotspot interaction index capture the time sensitivity of user interactions, ensuring that the calculation results are highly correlated with current real-time changes and avoiding interference from historical interactions in dynamic data processing.
[0089] b for each interaction behavior i The weights are associated with the numerical values of the behaviors. The weighted behavior value is calculated using the following expression:
[0090] B′ i =W(T) i )·(C i +αLi +βR i +γS i ),
[0091] In the formula, B i ′ is the weighted value of the i-th interaction behavior, and α, β, and γ are all behavior importance weights, used to adjust the number of times the user likes the data content, respectively. i R, the number of times users comment on the data content i And the number of times users share data content (S) i Impact on the hot topic interaction index;
[0092] Parameters α, β, and γ represent the importance coefficients of likes, comments, and shares in calculating the weighted behavior value, respectively, and are used to adjust the influence weights of different behaviors on the hot topic interaction index. The setting of these coefficients reflects the system's value assessment of different user interaction behaviors. For example, in some scenarios, the sharing behavior γ may be considered to reflect the dissemination and influence of data content more effectively than likes α and comments β, so γ can be set > α and β. Conversely, in scenarios that emphasize the depth of user feedback, the weight of the comment behavior β can be increased. By flexibly adjusting α, β, and γ, the contribution ratio of interaction behaviors can be optimized according to specific application scenarios, making the hot topic interaction index more targeted and applicable, thereby accurately reflecting the dynamic changes in user interaction characteristics and data.
[0093] To capture the complex characteristics of data content in user interactions, an interaction behavior complexity factor is defined to quantify the correlation between various behaviors. The calculation expression is as follows:
[0094] ,
[0095] In the formula, H(B') is the interaction complexity factor. This is the Shannon information entropy formula, p(b i The weighted value of the i-th interaction behavior is the proportion of the total weighted value, and the calculation expression is as follows:
[0096] ,
[0097] In the formula, B j ′ is the weighted value of the j-th interaction behavior. It is the weighted sum of all interactive behaviors;
[0098] The interaction complexity factor assesses the diversity and uniformity of user interactions by quantifying the distribution characteristics of different user behavior types (such as clicks, likes, comments, and shares). Its calculation is based on the information entropy formula, where p(b iThe complexity factor represents the proportion of the first (i) interaction behavior to the total number of interactions. A higher complexity factor value indicates a more even distribution of user interactions across different behavior types, and greater diversity of behaviors; a lower value indicates that interactions are concentrated on a specific behavior type. Here, the interaction complexity factor aims to capture the depth and breadth of user interaction with data content, preventing a single behavior from dominating the index result, thus more comprehensively reflecting the dynamic characteristics of the data. By combining the total number of interactions and the complexity factor, the hotspot interaction index can more accurately measure the overall level of user attention to content, improving the accuracy of dynamic update decisions.
[0099] Shannon's information entropy formula is a mathematical expression used to measure the uncertainty or information diversity in a system. It is widely used in information theory, data analysis, and statistics. The more uniform the probability distribution of all events, the higher the information entropy value, indicating greater uncertainty or information diversity in the system. Conversely, if the probability distribution is concentrated on a few events, the information entropy value will be lower, indicating higher determinism in the system. In interactive behavior analysis, Shannon's information entropy is used to quantify the distribution characteristics of users' different behavior types (such as clicks, likes, comments, and shares). It can effectively measure the diversity and uniformity of behavior, providing an important reference indicator for the analysis of dynamically changing data characteristics.
[0100] The weighted value B of the i-th interaction behavior i The interaction complexity factor H(B′) and the time decay weight W(Ti′) are also mentioned. ) The hotspot interaction index is generated, and the calculation expression is as follows:
[0101] ,
[0102] In the formula, WDHII is the hotspot interaction index, and κ is the index normalization factor.
[0103] The index normalization factor is a parameter used to adjust the calculation results of the hotspot interaction index (or similar indicator). Its purpose is to normalize the index value to a specific range, making it comparable and applicable across different scenarios. Represented by κ, its specific value is flexibly set according to application requirements and data range. For example, in datasets of different sizes, due to differences in the total amount or complexity of behaviors, the calculated hotspot interaction index may have excessively large or small deviations. By introducing a normalization factor, the index can be linearly scaled or normalized, ensuring that the results are both meaningful in practice and easy to compare with results from other scenarios or time periods. Here, the role of the index normalization factor is to balance the dimensional differences of various calculation results, making the output of the thermoelectric interaction index both stable and interpretable, thereby supporting efficient dynamic data update decisions.
[0104] By comprehensively considering the weighted value B of the i-th interaction behavior iThe interaction complexity factor H(B') and the time decay weight W(T) are also considered. i This system generates a hotspot interaction index that accurately reflects the dynamic changes in data. Total behavior reflects the overall frequency of user interactions within a specific time window, demonstrating the popularity of the data content and the level of user attention. Behavioral complexity measures the diversity of distribution among different behavior types (such as clicks, likes, comments, and shares), ensuring that the hotspot index not only considers the total amount of interaction but also captures users' deep engagement with the content and multi-dimensional interactive characteristics. Time weighting introduces a time decay mechanism to highlight the impact of recent interactions, ensuring the index is highly consistent with current data dynamics. By organically combining these three factors, the hotspot interaction index comprehensively and accurately reflects the dynamic characteristics of data, providing a quantitative basis for high-frequency updates and dynamic content management. This enables the system to respond promptly to user needs, improving data timeliness, relevance, and user experience.
[0105] In the monitoring window, a higher Hotspot Interaction Index value, generated after analyzing the frequency of user interaction with data content, indicates a significant increase in user attention and interaction with the data within a short period. This is typically closely related to the real-time and dynamic nature of the data content. Such data may involve trending news, financial market dynamics, breaking events, etc., which change rapidly and have a wide impact, thus requiring frequent updates to meet users' real-time needs and provide a precise experience. Conversely, a lower Hotspot Interaction Index value indicates stable user interaction with the data content, without significant dynamic changes. This type of data usually belongs to a relatively stable category, such as historical records or background knowledge, where a lower update frequency is sufficient.
[0106] The accelerated spread of event-related information online indicates that the extracted data is dynamically changing. The speed of dissemination reflects the trend of information being widely noticed, discussed, and shared by users in a short period, which is usually related to the suddenness, timeliness, and impact of the event. For example, the release of major policies, sudden accidents, or significant fluctuations in financial markets often trigger widespread and rapid dissemination. This rapid dissemination necessitates an extremely high frequency of data updates, as the dynamic content of the event is constantly changing and supplemented in a short period, such as new details, latest developments, and expert analyses. If the system cannot capture these changes in a timely manner, it may provide outdated or incomplete answers, failing to meet users' real-time needs. Therefore, the accelerated dissemination speed not only directly reflects rapid dynamic changes but also indicates that this data type requires a higher frequency of updates to ensure that question-and-answer systems and search engines can generate accurate and relevant answers and respond promptly to users' needs for the latest information.
[0107] In the monitoring window, a higher Hotspot Interaction Index value, generated after analyzing the frequency of user interaction with data content, indicates a significant increase in user attention and interaction with the data within a short period. This is typically closely related to the real-time and dynamic nature of the data content. Such data may involve trending news, financial market dynamics, breaking events, etc., which change rapidly and have a wide impact, thus requiring frequent updates to meet users' real-time needs and provide a precise experience. Conversely, a lower Hotspot Interaction Index value indicates stable user interaction with the data content, without significant dynamic changes. This type of data usually belongs to a relatively stable category, such as historical records or background knowledge, where a lower update frequency is sufficient.
[0108] The accelerated spread of event-related information online indicates that the extracted data is dynamically changing. The speed of dissemination reflects the trend of information being widely noticed, discussed, and shared by users in a short period, which is usually related to the suddenness, timeliness, and impact of the event. For example, the release of major policies, sudden accidents, or significant fluctuations in financial markets often trigger widespread and rapid dissemination. This rapid dissemination necessitates an extremely high frequency of data updates, as the dynamic content of the event is constantly changing and supplemented in a short period, such as new details, latest developments, and expert analyses. If the system cannot capture these changes in a timely manner, it may provide outdated or incomplete answers, failing to meet users' real-time needs. Therefore, the accelerated dissemination speed not only directly reflects rapid dynamic changes but also indicates that this data type requires a higher frequency of updates to ensure that question-and-answer systems and search engines can generate accurate and relevant answers and respond promptly to users' needs for the latest information.
[0109] The specific steps for analyzing the propagation speed of event-related information on the network and generating an event propagation speed index within the monitoring window are as follows:
[0110] Under the monitoring window, raw data related to the event is collected from multiple network data sources, and propagation nodes are identified. The calculation expression is as follows:
[0111] ,
[0112] In the formula, N p It represents the number of propagation nodes, i.e., the total number of propagation behaviors triggered within the monitoring window, ω. m f is the weight of each data source m. m (t) is the time function of the event-related propagation point in the m-th data source, which records the number of propagation behaviors of the data source in time t, and M is the total number of data sources;
[0113] In the process of collecting raw event-related data from multiple network data sources and identifying propagation nodes, the definition and scope of network data sources and propagation nodes are crucial. The following is a detailed explanation, along with classifications and examples of some common and major network data sources and propagation nodes in China.
[0114] 1. Network data source
[0115] Web data sources are the original sources used to collect information related to events, and mainly include the following categories:
[0116] (1) News website
[0117] Description: A network platform that primarily disseminates authoritative content such as breaking news, policy interpretations, and social hot topics.
[0118] Example:
[0119] People's Daily Online and Xinhua News Agency: Release authoritative policy interpretations and major domestic and international events.
[0120] The Paper and Jiemian News: Focus on social hot topics and in-depth reporting.
[0121] Tencent News and NetEase News: Provide comprehensive news content, covering breaking news and trending topics.
[0122] (2) Social media platforms
[0123] Description: The primary source of user-generated content, featuring a large volume of real-time event discussions and sharing.
[0124] Example:
[0125] Weibo: A real-time discussion platform for trending public opinion and breaking news.
[0126] WeChat Official Accounts Platform: Disseminating in-depth content through subscription accounts and articles.
[0127] Douyin and Kuaishou: They rapidly disseminate breaking news and trending topics through short videos.
[0128] (3) Online forums and communities
[0129] Description: A gathering place for users to discuss events and express their opinions, reflecting diverse public opinion.
[0130] Example:
[0131] Zhihu: In-depth Q&A and discussion among users on specific events.
[0132] Hupu: Focuses on sports hot topics, but also involves social discussions.
[0133] Douban: Especially in film, television and cultural events, it reflects public opinion.
[0134] (4) Search engines and content aggregation platforms
[0135] Description: Provides a wealth of articles, images, and videos related to the event through keywords.
[0136] Example:
[0137] Baidu: Provides highly relevant content through Baidu News or trending search terms.
[0138] Toutiao: Pushes the latest events and personalized content based on algorithms.
[0139] (5) Government and authoritative agency websites
[0140] Description: Provides official announcements, policy interpretations, and authoritative data.
[0141] Example:
[0142] The National Bureau of Statistics website publishes official statistical data.
[0143] The State Council Information Office provides interpretations of events and policies at the national level.
[0144] Ministry of Public Security website: disseminating information on sudden security incidents.
[0145] 2. Propagation Nodes
[0146] Propagation nodes are the triggering or carrying points of information propagation, representing the key links in the diffusion and transmission of information in the network. They mainly include the following categories:
[0147] (1) Information release node
[0148] Description: The node that publishes the original event information is usually the starting point for event propagation.
[0149] Example:
[0150] Official accounts of media organizations (such as the People's Daily Weibo).
[0151] Announcements published by government agencies (such as the official website of the Ministry of Emergency Management).
[0152] (2) Information forwarding node
[0153] Description: A node that expands the reach of information through forwarding.
[0154] Example:
[0155] Social media users who share posts (such as celebrities on Weibo sharing trending topics).
[0156] Articles about events shared on WeChat Moments.
[0157] (3) Information comments and interactive nodes
[0158] Description: Nodes that participate in information dissemination through comments, likes, and discussions.
[0159] Example:
[0160] User discussions in the Weibo comment section.
[0161] In-depth discussions in Zhihu answers and comments.
[0162] (4) Information aggregation and reprocessing nodes
[0163] Description: A node that forms new information by organizing, summarizing, or reprocessing original information.
[0164] Example:
[0165] News aggregation platforms (such as Toutiao).
[0166] Secondary analysis and interpretation of the event by self-media accounts.
[0167] (5) Hotspot diffusion nodes
[0168] Description: Rapidly promote the spread of information through high-influence nodes (such as celebrities and media organizations).
[0169] Example:
[0170] The incident was reposted by Weibo influencers and short video creators.
[0171] Secondary dissemination by institutional media.
[0172] Online data sources include news websites, social media, forums, content aggregation platforms, and government agencies, covering multiple sources of event information.
[0173] Propagation nodes are divided into information release nodes, forwarding nodes, comment and interaction nodes, aggregation and processing nodes, and hotspot diffusion nodes, reflecting the propagation path and method of information on the network.
[0174] These network data sources and propagation nodes together constitute a complete ecosystem for the generation, propagation, and diffusion of information within the network, providing a diverse data foundation for calculating the event propagation speed index.
[0175] Analyze the propagation path of event-related information and quantify the propagation depth, i.e., the hierarchical structure of information spreading from the initial propagation node to other nodes. The calculation expression is as follows:
[0176] ,
[0177] In the formula, D p It is a propagation depth index, where L is the total number of propagation levels, and ω is the propagation depth index.l V is the influence factor of the l-th layer of propagation, used to measure the influence weight of each layer of propagation. l It is the number of propagation nodes in the l-th layer, that is, the number of independent nodes that receive information in each layer;
[0178] The propagation depth index is a quantitative indicator of the depth of the hierarchical structure formed by information spreading from the initial propagation node to other nodes in a network. It reflects the diffusion capacity and scope of influence of information across different propagation levels. For example, the greater the propagation depth of a piece of information after multiple layers of propagation, such as primary propagation (direct forwarding) and secondary propagation (forwarding of forwards), the more layers of diffusion and the wider the coverage of the information in the network. The propagation depth index is used to assess the breadth and penetration of information diffusion, revealing whether the information has the potential for sustained propagation. In scenarios such as public opinion monitoring and hotspot analysis, the propagation depth index can help systems identify events with significant impact, determine their dynamism, and thus decide on data update frequency and resource allocation strategies.
[0179] The propagation speed, i.e., the speed at which information spreads per unit time, is calculated using the increment of propagation nodes and the time interval as the basis. The calculation expression is as follows:
[0180] ,
[0181] In the formula, V p It is the propagation speed index, and Q is the total number of time segments. It is the propagation rate weight of the q-th time segment, ΔN q The number of newly added propagation nodes in the q-th time segment, Δt q It is the duration of the q-th time segment;
[0182] The speed of propagation index refers to how quickly information spreads within a network. It reflects the dynamics and real-time nature of information propagation by quantifying the growth rate of information propagation nodes per unit of time. It focuses on the efficiency of information expansion within a specific time window and is typically determined by both the number of new propagation nodes and the time interval. A higher speed of propagation index indicates faster information spread and higher levels of attention and participation from network users.
[0183] Here, the propagation speed index measures the efficiency of information dissemination within a short period, determining whether the information is sufficiently dynamic, thus providing a basis for high-frequency data updates. A high propagation speed index indicates high time sensitivity and popularity of the information, requiring more frequent data updates to ensure that the question-and-answer system and search engine can respond to user needs in real time. Conversely, a low propagation speed index indicates slower information dissemination and weaker dynamics, allowing for a reduction in update frequency and optimization of system resource allocation.
[0184] Combined with the number of propagation nodes N p ,D spread depth index p and the propagation speed index V p Generate the event propagation speed index, calculated using the following expression:
[0185] ,
[0186] In the formula, NEPVI is the event propagation speed index, and S, Y, and Z are all weighting coefficients used to balance the number of propagation nodes N. p ,D spread depth index p and the propagation speed index V p The overall impact on the event propagation speed index;
[0187] S, Y, and Z are the weighting parameters in the event propagation speed exponential formula; they are used to measure the number of propagation nodes N. p ,D spread depth index p and the propagation speed index V p The weights contribute to the final NEPVI (Network Effort Speed Index). Specifically, S reflects the weight of the number of dissemination nodes on the impact of dissemination; more nodes indicate a wider scope of attention for the event. Y represents the weight of dissemination depth; greater depth indicates more layers of information dissemination and wider coverage. Z measures the importance of dissemination speed; faster speed indicates higher efficiency in information diffusion within a short period. These three weights are used to adjust the relative importance of each indicator according to the dynamic needs of dissemination in different scenarios. For example, S can be increased in scenarios emphasizing broad attention, while Z can be increased in sudden events where speed is more important, thus making the NEPVI more adaptable to specific practical needs.
[0188] In the monitoring window, the higher the event propagation speed index value generated after analyzing the speed of event-related information propagation on the network, the more widely the event information is disseminated and shared in a short period of time, and the rapidly expanding scope or intensity of the propagation. This rapid propagation reflects the high timeliness and dynamism of the event, indicating that the data type is dynamically changing and requires frequent updates to ensure that the system can capture the latest changes in a timely manner and generate accurate answers. Conversely, if the event propagation speed index value is low, it indicates that the event information propagates relatively slowly, with limited changes and weak dynamism. The data type does not fall into the category of rapidly changing data and usually does not require frequent updates.
[0189] The machine learning model is not limited here. Any machine learning model that can analyze the hotspot interaction index WDHII and the event propagation speed index NEPVI to generate the data popularity coefficient DPHC is acceptable. In order to realize the technical solution of the present invention, the present invention provides a specific implementation method.
[0190] The formula for generating the Data Popularity Coefficient (DPHC) is as follows:
[0191] ,
[0192] In the formula, f1 and f2 are the preset proportional coefficients of the hotspot interaction index WDHII and the event propagation speed index NEPVI, respectively, and both f1 and f2 are greater than 0.
[0193] As shown in the data popularity coefficient calculation formula, under the monitoring window, the higher the hotspot interaction index value generated after analyzing the frequency of user interaction with data content, and the higher the event propagation speed index value generated after analyzing the propagation speed of event-related information on the network, the higher the data popularity coefficient value generated after analyzing the extracted key features under the monitoring window. This indicates that the data has high user attention and propagation speed, strong information dynamism, and belongs to hotspot or sudden event data. It usually requires high-frequency updates to ensure that the system can capture changes and respond to user needs in a timely manner. Conversely, the lower the data popularity coefficient value, the lower the user interaction frequency, the slower the propagation speed, the weaker the dynamism, and the more stable the data type, which does not require frequent updates.
[0194] The preset proportional coefficients f1 and f2 here are weighting parameters used to balance the influence of the Hotspot Interaction Index (WDHII) and the Event Propagation Speed Index (NEPVI) on the final Data Popularity Index (DPHC) during the calculation of the data popularity index. Because the importance of these two indicators may vary in different scenarios—for example, in some applications, the WDHII may be more important, while in others, the NEPVI may be a more critical indicator of dynamism—by setting the proportions of f1 and f2, the model can adjust the relative contributions of the two indicators according to actual needs, ensuring that the calculated popularity index better reflects the specific scenario and requirements. These preset proportional coefficients are typically based on domain experience, historical data analysis, or the training results of machine learning models.
[0195] The data classification module, based on the evaluation results of the machine learning model, divides the acquired relevant data information into high-frequency update data and low-frequency update data;
[0196] The data popularity coefficient generated after analyzing the extracted key features under the monitoring window will be compared with a pre-set data popularity coefficient reference threshold. The relevant data information will then be divided, and the division steps are as follows:
[0197] If the data popularity coefficient is greater than or equal to the preset data popularity coefficient reference threshold, the relevant data information obtained will be classified as high-frequency update data.
[0198] If the data popularity coefficient is less than the preset data popularity coefficient reference threshold, the relevant data information will be classified as low-frequency update data.
[0199] High-frequency update data refers to data types that change frequently within a short period of time, are highly dynamic, and require real-time updates. This type of data is usually driven by the suddenness of events or high public attention, and its content may be updated or iterated rapidly over time. Low-frequency update data, on the other hand, refers to data types that have a longer change cycle, lower dynamics, and even remain relatively stable over long periods. This type of data typically includes background knowledge, historical events, long-term statistical data, or natural science knowledge.
[0200] The low-frequency data timed update module continues to use the preset timed batch update strategy for data that is classified as low-frequency update.
[0201] For data classified as low-frequency updates, the purpose of continuing to use the pre-defined scheduled batch update strategy is to meet the system's requirements for the integrity and accuracy of this data in an efficient and low-cost manner. This data has a long change cycle and low dynamism; a fixed update interval is sufficient to meet its update needs. For example, historical events or basic knowledge data have relatively stable content and low real-time requirements, therefore frequent updates are unnecessary. Through scheduled batch updates, the system can save computing resources and storage costs while ensuring data consistency, avoiding resource waste caused by unnecessary frequent updates, thereby optimizing the overall system operating efficiency.
[0202] The high-frequency data dynamic update module dynamically increases the update frequency based on the evaluation results of the machine learning model under the monitoring window and a preset timed batch update strategy for data classified as high-frequency updates.
[0203] For data classified as high-frequency updates, based on the evaluation results of the machine learning model under the monitoring window, and using a preset timed batch update strategy, the update frequency is dynamically increased. The specific steps are as follows:
[0204] By comparing the data popularity coefficient DPHC with a reference threshold for the data popularity coefficient using a non-linear weighted difference, we can assess whether the current data popularity exceeds the reference level. The calculation expression is as follows:
[0205] ,
[0206] In the formula, Δ DPHC DPHC is a weighted data popularity bias, used to measure the overall deviation between the current popularity of data and the benchmark. ref This is the reference threshold for the data popularity coefficient; w1 and w2 are both weighting coefficients.
[0207] In the nonlinear weighted difference formula, w1 and w2 are both weighting coefficients, used to control the direct ratio term and the logarithmic adjustment term in the calculation of the heat deviation Δ of the data. DPHC The extent of its influence. Specifically:
[0208] Weighting coefficient w1: Controls the data popularity coefficient DPHC and the reference threshold DPHC. ref The weighting of the direct ratio emphasizes the relative change in overall popularity. If w1 is large, the formula focuses more on the ratio of the current popularity of the data to the reference threshold, reflecting the overall intensity of dynamic changes.
[0209] Weighting coefficient w2: Controls the weight of the logarithmic adjustment term, mainly capturing the subtlety of changes in heat intensity and the non-linear characteristics of marginal changes. If w2 is large, the formula is more sensitive to small deviations, making it suitable for monitoring initial dynamic changes or subtle fluctuations.
[0210] By combining w1 and w2, the formula's adaptability to different types of data can be flexibly adjusted. For example, in highly dynamic scenarios, increasing w1 emphasizes the significance of relative changes; in low-dynamic scenarios, increasing w2 focuses more on capturing subtle changes. This weighting mechanism improves the accuracy and adaptability of data heat assessment, enabling the system to make more reasonable decisions on adjusting the update frequency based on specific needs.
[0211] By combining ratios and logarithms, the nonlinear characteristics of data heat changes can be assessed more accurately, ensuring that the system can respond sensitively even under extreme dynamic conditions or close to the reference threshold.
[0212] By introducing nonlinear transformation and normalization, the update priority of the data is calculated, enabling the system to more accurately measure the impact of data popularity deviation. The calculation expression is as follows:
[0213] ,
[0214] In the formula, UP is the update priority, and tanh(k·|Δ DPHC |) represents the effect of smoothing extreme values using the hyperbolic tangent function, where b controls the sensitivity of the smoothing, and max(|Δ) DPHC |,1) is the normalization factor, sign(Δ DPHC () is the symbol for thermal deviation;
[0215] The significance of update priority
[0216] Update priority (UP) is a core metric for measuring whether data needs dynamic adjustment of update frequency, used to determine the order of resource allocation. Higher priority data indicates greater dynamism and higher real-time requirements for the system, thus necessitating more frequent updates. By calculating update priority, the system can intelligently classify the importance and urgency of data, thereby providing higher processing priority for highly dynamic data while preventing low-dynamic data from consuming excessive resources, ensuring the efficiency and rationality of overall resource allocation.
[0217] The hyperbolic tangent function tanh(x) is a non-linear function often used to smooth extreme values to a finite range. In priority update calculations, using the hyperbolic tangent function can avoid over-amplification of priority calculations when the heat deviation is too large. For example, when the dynamic changes of a certain data are very drastic, directly using linear calculations may result in a much higher priority than other data, causing uneven resource allocation. Through the tanh function, extreme values are compressed to the range [-1, 1], making the priority more stable, suitable for highly dynamic scenarios, while retaining sensitivity to moderate and subtle dynamic changes.
[0218] The normalization factor is a key parameter used to standardize data popularity bias values, ensuring consistency in update priority calculations across different data ranges. The purpose of normalization is to map data with different dimensions (such as high-volatility data from financial markets versus low-dynamic data from social media) to a unified evaluation standard. For example, by using max(|Δ DPHC |,1) As a normalization factor, it can prevent excessively small or large data deviation values from having an extreme impact on priority calculation, thus enhancing the applicability of the formula. The introduction of the normalization factor makes the calculation results of updated priorities both relatively comparable and retains flexible adaptation to various data characteristics.
[0219] The sign of the thermal deviation is sign(Δ) DPHC This is used to indicate the directionality of dynamic data changes and determine the direction of update priority adjustments. If the heat deviation is positive (Δ... DPHC A value greater than 0 indicates that the data's popularity is significantly higher than the reference threshold, requiring an increase in update frequency. If the popularity deviation is negative or close to 0, it indicates insufficient data dynamism, allowing the current update frequency to be maintained or the update intensity to be reduced. By combining symbolic information, the system can make reasonable update decisions based on dynamic changes in different directions, avoiding ineffective update frequency adjustments and improving resource utilization.
[0220] Through nonlinear smoothing and normalization, priority calculation can cope with drastic changes and remain sensitive to small dynamics, thus avoiding waste of system resources.
[0221] Based on the preset timed batch update frequency and update priority UP, the actual update frequency is dynamically calculated through exponential adjustment and nonlinear functions. The calculation expression is as follows:
[0222] ,
[0223] In the formula, F dunamic It is the dynamic update frequency, F base τ is the preset timed batch update frequency, τ is the exponential adjustment coefficient, which controls the overall influence range of update priority on update frequency, σ is the fluctuation adjustment coefficient, which controls the influence of dynamic changes on fine frequency adjustments, and e is the natural base.
[0224] By combining the adjustment mechanisms of exponential and sine functions, the update frequency can be accurately adapted to dynamic changes, while avoiding over-updating when changes approach the threshold.
[0225] Dynamic update frequency refers to a strategy that flexibly adjusts the update cycle or frequency based on the real-time popularity, rate of change, and priority of data. Unlike traditional fixed update frequencies, dynamic update frequencies optimize resource utilization efficiency by assessing the degree of dynamic change in data and allocating more frequent updates to highly dynamic data while maintaining a lower update frequency for inactive data. Its goal is to achieve a balance between the real-time nature and relevance of information and the resource efficiency of the system.
[0226] The actual performance after monitoring data updates is assessed, and the deviation between the actual update frequency and the theoretical optimal frequency is calculated using the following expression:
[0227] ,
[0228] In the formula, ε is the update frequency error, which measures the relative deviation between the actual update frequency and the theoretical optimal frequency, and F actual This refers to the actual update frequency, the frequency of updates that have been performed, F. optimal This is the theoretically optimal update frequency, calculated as follows:
[0229] ,
[0230] In the formula, D is the dynamic adjustment coefficient, which determines the magnitude of the impact of data dynamics on the optimal frequency;
[0231] Based on the current update frequency error ε, the exponential adjustment coefficient τ and fluctuation adjustment coefficient σ of the dynamic update frequency model are adaptively optimized to further improve the update strategy. The calculation expressions are as follows:
[0232] τ new =τ current ·(1-H·∈), σ new =σ current·(1-H·∈),
[0233] In the formula, τ new It is the optimized exponential adjustment coefficient, σ new It is the optimized fluctuation adjustment coefficient, τ current This is an adjustment factor, currently used to control the overall impact of the Data Heat Factor (DPHC) on the update frequency. σ current It is a sensitivity coefficient, currently used to control the bias Δ of the weighted data heat. DPHC For the coefficient of sensitivity to update frequency, H is the step size coefficient, which is used to control the sensitivity of adjustment;
[0234] The optimization step size coefficient H is a key factor controlling the adjustment magnitude during parameter optimization. It determines the optimization algorithm's response speed and adjustment magnitude to error feedback. It is used to uniformly adjust the optimization intensity of the two core coefficients in the update frequency model. The main function of the optimization step size coefficient is to balance the system's adjustment speed: when the step size is large, parameter optimization responds more quickly to errors ∈ , suitable for handling dynamically changing data; when the step size is small, the system adjusts more smoothly, suitable for gradually optimizing stable data, thus avoiding over-adjustment that could lead to frequency fluctuations or system oscillations. By setting H, the system can flexibly adjust its optimization sensitivity according to scenario requirements, ensuring that frequency adjustment is both efficient and stable.
[0235] The optimized exponential adjustment coefficient τ new and the optimized fluctuation adjustment coefficient σ new Re-enter the dynamic update frequency calculation formula to improve the update strategy. The calculation expression is as follows:
[0236] F dynamic,new =F base ·(1+τ new ·tanh(σ new ·|Δ DPHC |)),
[0237] In the formula, F dynamic,new It refers to the updated dynamic update frequency;
[0238] By monitoring errors in real time and optimizing adjustment coefficients, the frequency update strategy can be adaptively adjusted to gradually approach the optimal state, ensuring efficient use of resources while meeting the needs of dynamic data changes.
[0239] For data classified as high-frequency updates, the update frequency is dynamically increased based on the evaluation results of the machine learning model under the monitoring window, according to a preset timed batch update strategy. Its core function is to ensure that the system can respond promptly to rapidly changing data to meet real-time requirements. This type of data typically has high dynamism and timeliness, such as financial data (e.g., stock prices, trading volumes), trending news (e.g., the latest developments in breaking events), or social media discussions (e.g., rapid changes in topic popularity). If the system still adopts a fixed low-frequency update strategy, it may lead to data lag, failing to capture the latest changes, thus affecting the accuracy and relevance of the question-answering system or search engine.
[0240] By dynamically increasing the update frequency, the system can flexibly adjust the update rhythm based on real-time evaluation results, building upon the existing scheduled batch updates. For example, when the machine learning model detects a significant increase in certain key indicators (such as event propagation speed, user interaction frequency, etc.), the system will shorten the data update cycle, quickly capture and process the latest content, ensuring that the information users receive is always up-to-date and accurate. This dynamic adjustment mechanism not only improves the system's adaptability to frequently changing data but also enhances user experience and system reliability. More importantly, this strategy avoids the waste of resources from frequently updating all data, achieving priority processing of high-frequency data and rational allocation of resources, thereby optimizing resource utilization efficiency while ensuring performance.
[0241] This invention, by introducing the analysis of dynamic data characteristics (such as hot topic interaction index and event propagation speed index) and intelligent evaluation based on machine learning, enables the system to accurately identify the real-time needs of highly dynamic data (such as financial data or trending news), classify it into high-frequency update data, and dynamically adjust the update frequency. Compared to traditional timed batch update strategies, this dynamic update mechanism significantly improves the system's responsiveness to rapidly changing data. For example, for financial market data, the system can capture high-profile events such as price fluctuations or changes in trading volume, update relevant content in a timely manner, and ensure that the answers generated by the question-and-answer system and search engine are based on the latest information, avoiding misleading decisions due to data lag. In highly sensitive areas, such as financial investment or public opinion monitoring, improved real-time performance directly affects users' economic interests and public opinion response strategies, significantly enhancing the system's accuracy and credibility.
[0242] This invention effectively categorizes data into high-frequency and low-frequency update types through a data popularity coefficient and reference threshold mechanism. For low-frequency update data, a timed batch update strategy continues to be used to reduce unnecessary resource waste; for high-frequency update data, the update frequency is dynamically increased, thereby achieving optimal resource utilization under limited computing resources and bandwidth. In this way, the system can meet the real-time requirements of highly dynamic data while conserving resources on low-dynamic data, ensuring overall operational efficiency. For example, dynamic updates of trending news focus only on truly high-profile events, avoiding repeated fetching of less popular information, thus reducing system load. While achieving real-time information delivery, the system optimizes computing resource allocation, ensuring efficient operation while flexibly adapting to dynamic needs.
[0243] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A knowledge graph-based big data text relationship extraction and management system, characterized in that, It includes a data acquisition module, a data processing and feature extraction module, a feature analysis and machine learning evaluation module, a data classification module, a low-frequency data timed update module, and a high-frequency data dynamic update module. The data acquisition module first obtains data information from various massive unstructured text data sources; The data processing and feature extraction module processes the acquired data into a structured analysis set and extracts key features that reflect the dynamic changes of the data from the analysis set. The feature analysis and machine learning evaluation module analyzes and processes the extracted features under the monitoring window, and inputs the analyzed key feature information into the pre-learned machine learning model, which then performs intelligent evaluation of the acquired data information. The data classification module, based on the evaluation results of the machine learning model, divides the acquired relevant data information into high-frequency update data and low-frequency update data; The low-frequency data timed update module continues to use the preset timed batch update strategy for data that is classified as low-frequency update. The high-frequency data dynamic update module dynamically increases the update frequency of data classified as high-frequency updates based on the evaluation results of the machine learning model under the monitoring window and a preset timed batch update strategy.
2. The knowledge graph-based big data text relationship extraction and management system according to claim 1, characterized in that: Key features reflecting dynamic changes in data are extracted from the analysis set, including the frequency of user interaction with data content and the speed of event-related information propagation in the network. After acquisition, under the detection window, the frequency of user interaction with data content and the speed of event-related information propagation in the network are analyzed to generate hotspot interaction index and event propagation speed index respectively. The hotspot interaction index quantifies the intensity and frequency of user interaction with a certain data content, reflecting the popularity of the content and its trend of becoming a hot event. The event propagation speed index quantifies the rate and coverage of the spread of information related to a certain event on a network platform, reflecting the event's propagation capability and real-time impact.
3. The knowledge graph-based big data text relationship extraction and management system according to claim 2, characterized in that: After obtaining the hotspot interaction index and event propagation speed index generated by analyzing the extracted key features, the hotspot interaction index and event propagation speed index are input into a pre-learned machine learning model. The machine learning model generates a data popularity coefficient, and the data popularity coefficient is used to intelligently evaluate the obtained data information.
4. The knowledge graph-based big data text relationship extraction and management system according to claim 3, characterized in that: The data popularity coefficient generated after analyzing the extracted key features under the monitoring window will be compared with a pre-set data popularity coefficient reference threshold. The relevant data information will then be divided, and the division steps are as follows: If the data popularity coefficient is greater than or equal to the preset data popularity coefficient reference threshold, the relevant data information obtained will be classified as high-frequency update data. If the data popularity coefficient is less than the preset data popularity coefficient reference threshold, the relevant data information will be classified as low-frequency update data.
5. The knowledge graph-based big data text relationship extraction and management system according to claim 4, characterized in that: For data classified as high-frequency updates, based on the evaluation results of the machine learning model under the monitoring window, and using a preset timed batch update strategy, the update frequency is dynamically increased. The specific steps are as follows: By comparing the data popularity coefficient DPHC with a reference threshold for the data popularity coefficient using a non-linear weighted difference, we can assess whether the current data popularity exceeds the reference level. The calculation expression is as follows: , In the formula, Δ DPHC DPHC is a weighted data popularity bias, used to measure the overall deviation between the current popularity of data and the benchmark. ref This is the reference threshold for the data popularity coefficient; w1 and w2 are both weighting coefficients. By introducing nonlinear transformation and normalization, the update priority of the data is calculated, enabling the system to more accurately measure the impact of data popularity deviation. The calculation expression is as follows: , In the formula, UP is the update priority, and tanh(k·|Δ DPHC |) represents the effect of smoothing extreme values using the hyperbolic tangent function, where b controls the sensitivity of the smoothing, and max(|Δ) DPHC |,1) is the normalization factor, sign(Δ DPHC () is the symbol for thermal deviation; Based on the preset timed batch update frequency and update priority UP, the actual update frequency is dynamically calculated through exponential adjustment and nonlinear functions. The calculation expression is as follows: , In the formula, F dynamic It is the dynamic update frequency, F base τ is the preset timed batch update frequency, τ is the exponential adjustment coefficient, which controls the overall influence range of update priority on update frequency, σ is the fluctuation adjustment coefficient, which controls the influence of dynamic changes on fine frequency adjustments, and e is the natural base. The actual performance after monitoring data updates is assessed, and the deviation between the actual update frequency and the theoretical optimal frequency is calculated using the following expression: , In the formula, ∈ is the update frequency error, F actual It is the actual update frequency, F optimal This is the theoretically optimal update frequency, calculated as follows: , In the formula, D is the dynamic adjustment coefficient, which determines the magnitude of the impact of data dynamics on the optimal frequency; Based on the current update frequency error ∈, the update strategy is further improved by adaptively optimizing the exponential adjustment coefficient τ and the fluctuation adjustment coefficient σ of the dynamic update frequency model. The calculation expression is as follows: t new =t current ·(1-H·∈),σ new =s current ·(1-H·∈), In the formula, τ new It is the optimized exponential adjustment coefficient, σ new It is the optimized fluctuation adjustment coefficient, τ current This is an adjustment factor, currently used to control the overall impact of the Data Heat Factor (DPHC) on the update frequency. σ current It is a sensitivity coefficient, currently used to control the bias Δ of the weighted data heat. DPHC For the coefficient of sensitivity to update frequency, H is the step size coefficient, which is used to control the sensitivity of adjustment; The optimized exponential adjustment coefficient τ new and the optimized fluctuation adjustment coefficient σ new Re-enter the dynamic update frequency calculation formula to improve the update strategy. The calculation expression is as follows: F dynamic,new =F base ·(1+τ new ·tanh(σ new ·|Δ DPHC |)), In the formula, F dynamic,new It refers to the updated dynamic update frequency.
6. The knowledge graph-based big data text relationship extraction and management system according to claim 2, characterized in that, The specific steps for analyzing the frequency of user interaction with data content and generating a hotspot interaction index within the monitoring window are as follows: The specific steps for analyzing the frequency of user interaction with data content and generating a hotspot interaction index within the monitoring window are as follows: In the monitoring window, extract interaction behavior data related to the data content from the interaction log to form a behavior set B, where B = {b}. i } = {b1, b2, b3, ..., n}, where b i This refers to the i-th interaction, including the number of clicks, likes, comments, shares, and the time point of the interaction. n is the total number of interactions, and each interaction is represented by b. i The representation is as follows: b i =(C i L i R i S i T i In the formula, C i It refers to the number of times users click on data content, L i It is the number of times users like data content, R i S is the number of times users comment on the data content. i It is the number of times a user shares data content, T i It is the point in time when the interaction occurs; To reflect the real-time nature of user interaction frequency, a time decay weight is introduced. The weight value is defined by an exponential function, and the calculation expression is as follows: In the formula, W(T) i ) is the time decay weight, e is the natural base, λ is the time decay factor that controls the decay rate, and t2 is the end time of the monitoring window. b for each interaction behavior i The weights are associated with the numerical values of the behaviors. The weighted behavior value is calculated using the following expression: B i ′=W(T i )·(C i +αL i +βR i +γS i ), In the formula, B i ′ is the weighted value of the i-th interaction behavior, and α, β, and γ are all behavior importance weights, used to adjust the number of times the user likes the data content, respectively. i R, the number of times users comment on the data content i And the number of times users share data content (S) i Impact on the hot topic interaction index; To capture the complex characteristics of data content in user interactions, an interaction behavior complexity factor is defined to quantify the correlation between various behaviors. The calculation expression is as follows: , In the formula, H(B′) is the interaction behavior complexity factor. This is the Shannon information entropy formula, p(b i The weighted value of the i-th interaction behavior is the proportion of the total weighted value, and the calculation expression is as follows: , In the formula, B j ′ is the weighted value of the j-th interaction behavior. It is the weighted sum of all interactive behaviors; The weighted value B of the i-th interaction behavior i The interaction complexity factor H(B') and the time decay weight W(T) are also considered. i The hot topic interaction index is generated, and the calculation expression is as follows: , In the formula, WDHII is the hotspot interaction index, and W is the index normalization factor.
7. The knowledge graph-based big data text relationship extraction and management system according to claim 2, characterized in that, The specific steps for analyzing the propagation speed of event-related information on the network and generating an event propagation speed index within the monitoring window are as follows: Under the monitoring window, raw data related to the event is collected from multiple network data sources, and propagation nodes are identified. The calculation expression is as follows: , In the formula, N p It is the number of propagation nodes, ω m f is the weight of each data source m. m (t) is the time function of the event-related propagation point in the m-th data source, which records the number of propagation behaviors of the data source in time t, and M is the total number of data sources; Analyze the propagation path of event-related information and quantify the propagation depth, i.e., the hierarchical structure of information spreading from the initial propagation node to other nodes. The calculation expression is as follows: , In the formula, D p It is a propagation depth index, where L is the total number of propagation levels, and ω is the propagation depth index. l V is the influence factor of the l-th layer of propagation, used to measure the influence weight of each layer of propagation. l It is the number of propagation nodes in the l-th layer, that is, the number of independent nodes that receive information in each layer; The propagation speed, i.e., the speed at which information spreads per unit time, is calculated using the increment of propagation nodes and the time interval as the basis. The calculation expression is as follows: , In the formula, V p It is the propagation speed index, and Q is the total number of time segments. It is the propagation rate weight of the q-th time segment, ΔN q The number of newly added propagation nodes in the q-th time segment, Δt q It is the duration of the q-th time segment; Combined with the number of propagation nodes N p ,D spread depth index p and the propagation speed index V p Generate the event propagation speed index, calculated using the following expression: , In the formula, NEPVI is the event propagation speed index, and S, Y, and Z are all weighting coefficients used to balance the number of propagation nodes N. p ,D spread depth index p and the propagation speed index V p The overall impact on the speed of event propagation.