Edge computing and light-weight AI-based hotspot analysis terminal for basic research institutions and implementation method

By using edge computing and lightweight AI to create a hotspot analysis terminal for grassroots research institutions, the network consumption and latency issues caused by centralized cloud processing have been resolved. This enables localized real-time analysis of scientific literature data and provides personalized hotspot insight support.

CN121167356BActive Publication Date: 2026-05-29GUANGZHOU NANFANG WANFANG DATA CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU NANFANG WANFANG DATA CO LTD
Filing Date
2025-09-08
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for analyzing scientific research hotspots rely on centralized cloud processing, resulting in high network bandwidth consumption, high latency response, and high hardware costs. The lack of local context integration also leads to insufficient personalization of analysis results.

Method used

A hotspot analysis terminal for grassroots research institutions based on edge computing and lightweight AI is adopted. The feature vector set is fused with locally generated context weights through a data packet generation system to generate preliminary hotspot labels with intensity indicators. Core hotspot data packets are then filtered out through adaptive thresholds for topology analysis and differentiated display.

Benefits of technology

It achieves full-process localization and real-time processing of scientific research literature data, significantly reduces network bandwidth consumption, alleviates terminal computing pressure, and provides grassroots researchers with efficient, accurate and low-cost decision support tools. The analysis results deeply integrate the institution's local research preferences and historical background.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167356B_ABST
    Figure CN121167356B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of AI chips, and provides a terminal and an implementation method for hotspot analysis of basic scientific research institutions based on edge computing and light-weight AI. The terminal comprises a vector set acquisition system, a data packet generation system and a differential display system. The method comprises the following steps: a layered feature stripping of a light-weight AI model architecture is utilized to decompose scientific research literature content into multi-dimensional feature tuples such as semantic cores, methodology labels and data reference networks; a set of standardized and light-weight feature vectors are obtained; a preliminary hotspot label with an intensity identifier exceeding an adaptive threshold and core hotspot data associated with the preliminary hotspot label are marked out to generate a core hotspot data packet; a macro skeleton of a core hotspot atlas is presented on a terminal interface; differential highlighting is performed according to the intensity identifier in the core hotspot data packet; and deep detail information corresponding to a hotspot node is dynamically loaded and rendered. The application presents a macro hotspot trend and supports user interaction exploration to obtain deep information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AI chip technology, and in particular to a hotspot analysis terminal and implementation method for grassroots scientific research institutions based on edge computing and lightweight AI. Background Technology

[0002] The development of AI chips has evolved from general-purpose processors to dedicated accelerators. Initially, AI computing primarily relied on general-purpose CPUs, but their parallel processing capabilities were limited. Subsequently, GPUs became the mainstream choice for AI training due to their powerful parallel computing capabilities, but their high power consumption made them unsuitable for edge deployment. Next, FPGAs and dedicated ASICs gradually emerged, offering a better balance between power consumption and performance while maintaining high computing power, making them more suitable for edge computing scenarios. In recent years, to meet the demands of edge AI computing, an increasing number of chip companies have focused on developing low-power, high-performance edge AI chips. These chips typically employ reduced instruction sets, optimized memory architectures, and dedicated computing units to provide sufficient AI computing power within a limited power budget.

[0003] Existing technology 1, application number: CN202411739697.4, discloses a method and system for tracking and visualizing research hotspots in institutions. The method includes: acquiring research literature data from multiple dimensions; performing cluster analysis and heat mining on the research literature data to obtain target research literature data with potential hotspots; tracking and analyzing the heat persistence and heat stability of the target research literature data, and filtering to obtain hot research literature data; and performing visual analysis on the hot research literature data to obtain hotspot visualization analysis results. Although acquiring research literature data from multiple dimensions, tracking and analyzing the heat persistence and heat stability of target research literature data with potential hotspots, filtering to obtain hot research literature data, and finally performing visualization analysis can improve the accuracy of research hotspot acquisition and visualization effect, it involves multiple links such as data uploading, cloud queuing, calculation, and result return. The analysis results are periodic and lagging, and cannot provide researchers with real-time or near-real-time hotspot feedback.

[0004] Existing technology two, application number: CN202010092905.1, discloses a method for exploring the research status of institutions based on topic visualization, including the following steps: acquisition and preprocessing of research data, specifically, identifying the institutions to be studied and acquiring SCI academic literature data of the institutions; extracting the required research fields and preprocessing the acquired research corpus; processing the selected corpus using TF-IDF feature extraction and LDA topic model text mining technology to extract research hot topics and their keywords, and performing academic literature topic clustering; presenting the clustered topics and other dimensional information in the academic literature data in a visual manner, and analyzing the results from multiple dimensions. While this method is beneficial for better understanding and tracking the current research status and development of institutions, allowing researchers to better capture the forefront and hot topics of discipline development and avoid redundant research, it usually requires large amounts of memory and computing resources, and may need to run on server-level hardware, which constitutes a hardware cost barrier for grassroots laboratories or research institutes with limited computing power.

[0005] Current technologies 1 and 2 suffer from drawbacks in existing research hotspot analysis methods. These methods rely on centralized cloud processing, leading to high network bandwidth consumption, high latency, and high hardware requirements. Furthermore, the lack of local context integration results in insufficient personalization of analysis results. Therefore, this invention provides a hotspot analysis terminal and implementation method for grassroots research institutions based on edge computing and lightweight AI. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a hotspot analysis terminal for grassroots research institutions based on edge computing and lightweight AI, comprising:

[0007] The data packet generation system is used to fuse the feature vector set with the locally generated context weights to generate a set of preliminary hotspot labels with intensity indicators; at the same time, the preliminary hotspot labels with intensity indicators exceeding the adaptive threshold and their associated core hotspot data are marked to generate core hotspot data packets.

[0008] Optional, a packet generation system includes:

[0009] The weighting and acquisition subsystem is used to input the feature vector set into the local interest contour generator. By continuously recording and analyzing the implicit feedback data of researchers, it dynamically adjusts the weight values ​​in a lightweight matrix to form a local context weight set that represents the latest research concerns of the research institution.

[0010] The dimension matching subsystem is used to map and align the abstract semantic space of feature vectors with the specific interest space of context weights; each set of feature vectors is assigned an initial strength value that represents the strength of its relevance to the local scientific research background; the set of feature vectors is processed with the local context weight set in the dual-track fusion computing unit to generate a set of undecided labels with initial strength values.

[0011] The intensity threshold comparison subsystem is used to analyze the statistical distribution of all initial intensity values ​​in the set of tags to be judged with initial intensity values, and calculate an adaptive threshold based on the distribution to filter out relatively significant hot spots. The initial intensity value of each tag in the set of tags to be judged is compared with this adaptive threshold. A few preliminary hot spot tags whose initial intensity values ​​exceed the adaptive threshold and their corresponding core hot spot data are captured and extracted from the set of tags to be judged and packaged into a core hot spot data packet.

[0012] Optional, intensity threshold comparison subsystem, including:

[0013] An adaptive threshold generation component is used to analyze the overall distribution of the initial intensity values ​​of all the labels to be judged and obtain a statistical distribution. If the statistical distribution is positively skewed, a threshold baseline is adaptively determined based on the data point density in the high intensity value region to capture true outliers. If the statistical distribution is uniform or negatively skewed, the threshold is determined based on the clustering of intensity values. The statistical distribution of the initial intensity values ​​of the labels to be judged is processed to generate an adaptive threshold.

[0014] The decision signal judgment component is used to synchronously judge the intensity of each tag in the tag set to be judged, compare the initial intensity value of each tag to be judged with the adaptive threshold, and generate a binary capture decision signal. If the initial intensity value exceeds the adaptive threshold, the capture decision signal is true, otherwise it is false.

[0015] The compression coding component is used to transmit the actual capture decision signal to a sparsely coded data packetizer, which compresses and encodes the core hotspot data associated with each captured tag to be judged, and encapsulates it together with the tag to be judged and the strength identifier to form a core hotspot data packet.

[0016] Optional, adaptive threshold generation component, including:

[0017] The diffusion index calculation sub-component is used to transform the initial intensity value set of the tag set to be judged into a distribution pattern identifier and its corresponding density distribution spectrum after multi-moment joint analysis and local density diffusion index calculation.

[0018] The candidate value output sub-component is used to import the distribution pattern identifier and density distribution spectrum into a high-density domain threshold extractor if the distribution pattern identifier indicates that the statistical distribution is positively skewed. The density distribution spectrum of the high-intensity value region is analyzed to find the intensity value corresponding to the critical point. After the distribution pattern identifier and density distribution spectrum are processed by the high-density domain threshold extractor, a threshold baseline candidate value is generated.

[0019] If the distribution pattern identifier indicates that the statistical distribution is uniform or negatively skewed, it is directed to a clustering perceptron, which scans the entire density distribution spectrum, identifies natural clusters of intensity values ​​and numerical gaps between clusters, and outputs a threshold baseline candidate value.

[0020] The smoothing correction subcomponent is used to smooth the threshold baseline candidate values ​​to generate an adaptive threshold.

[0021] Optional, the diffusion index calculation subcomponent includes:

[0022] The density variation spectrum generation module is used to simultaneously solve the third and fourth moments of the initial intensity value set of the label set to be judged; it obtains the dominant factor of the distribution pattern of the third moment and the auxiliary factor of the distribution pattern of the fourth moment; at the same time, the initial intensity value set is fed into a sliding density evaluation window in parallel. For each sliding position, the number of data points in the window is counted and recorded to form a preliminary density profile; the difference ratio of density values ​​between adjacent windows is calculated to capture the areas of drastic changes in the aggregation and dispersion of data points in the value range, and a density variation spectrum is obtained that describes the degree of concentration and the rate of change of data points in each intensity interval.

[0023] The distribution pattern classification module is used to classify distribution patterns according to the rules for determining the interaction between the dominant and auxiliary factors of distribution patterns, and output the distribution pattern classification code.

[0024] The association and encapsulation module is used to classify and encapsulate the distribution pattern as a global pattern identifier, associate and encapsulate it with the detailed local density information contained in the density change spectrum, and generate a distribution pattern identifier containing global pattern and local density features and its matching density distribution spectrum.

[0025] Optional, the distribution pattern classification module includes:

[0026] The spatial mapping processing submodule is used to map the two-dimensional coordinates composed of the dominant distribution morphology factor and the auxiliary distribution morphology factor to the morphology discrimination state space. The morphology discrimination state space is divided into multiple different regions, each region corresponding to a distribution morphology. After the morphology discrimination state space mapping processing, the dominant distribution morphology factor and the auxiliary distribution morphology factor are converted into morphology state codes.

[0027] The morphology comparison execution submodule is used to compare the currently obtained morphology state code with the recent historical morphologies in the local historical distribution prior knowledge base; if the current morphology is consistent with the recent mainstream morphology, the confidence of the morphology state code is increased and it is output directly; if the current morphology deviates significantly, a review process is initiated.

[0028] The morphology coding selection submodule is used to select between two possible neighboring morphology codes based on changes in the dominant and auxiliary factors of the distribution morphology, generating a distribution morphology classification code with confidence weights.

[0029] Optional, the morphology encoding selection submodule includes:

[0030] The descriptor transformation unit is used to extract the original values ​​of the dominant and auxiliary factors of the distribution pattern, along with their implied trend information, from the local historical distribution prior knowledge base. It extracts the numerical sequences of the dominant and auxiliary factors of the distribution pattern over the most recent calculation periods and plots their respective trajectories over time. The current value is compared with the trajectory to analyze the magnitude and direction of the change in the current value relative to its historical trajectory. The dominant and auxiliary factors of the distribution pattern are then transformed into dynamic descriptors.

[0031] The preliminary decision output unit is used to determine, based on the dynamic change descriptor, which of the two adjacent morphological codes the current change behavior tends to lead to, and to obtain a preliminary tendency decision.

[0032] The discrimination result output unit is used to input the preliminary tendency decision and the dynamic change descriptor into the decision confidence fusion unit. It introduces a decision stability constraint, which tends to select the morphological code that minimizes the overall fluctuation of a series of recent morphological discrimination results.

[0033] Optionally, it also includes a vector set acquisition system for non-uniformly segmenting the raw data stream of scientific literature, dynamically adjusting the processing granularity based on the data entropy value, and prioritizing the parsing of segments with high information density; using the hierarchical feature stripping of a lightweight AI model architecture, the content of scientific literature is decomposed into multi-dimensional feature tuples such as semantic core, methodological tags, and data citation networks; resulting in a set of standardized and lightweight feature vectors.

[0034] Optionally, it also includes a differentiated display system for performing topological analysis on core hotspot data packets to construct a relational graph between core hotspots; presenting the macroscopic skeleton of the core hotspot graph on the terminal interface and highlighting it differentiatedly according to the intensity indicators in the core hotspot data packets; when researchers have an interactive intention on a hotspot node of a certain macroscopic skeleton, the system dynamically loads and renders the deep detail information corresponding to the hotspot node according to the requirements.

[0035] This invention provides a method for implementing a hotspot analysis terminal for grassroots research institutions based on edge computing and lightweight AI, comprising the following steps:

[0036] The raw data stream of scientific literature is segmented non-uniformly, and the processing granularity is dynamically adjusted according to the data entropy value, prioritizing the parsing of segments with high information density; using the hierarchical feature stripping of a lightweight AI model architecture, the content of scientific literature is decomposed into multi-dimensional feature tuples such as semantic core, methodological tags, and data citation network; resulting in a set of standardized and lightweight feature vectors.

[0037] The feature vector set is fused with the locally generated context weights to generate a set of preliminary hotspot labels with intensity indicators; at the same time, the preliminary hotspot labels with intensity indicators exceeding the adaptive threshold and their associated core hotspot data are marked to generate core hotspot data packages.

[0038] The system performs topological analysis on core hotspot data packets to construct a relational graph between core hotspots. The macroscopic skeleton of the core hotspot graph is presented on the terminal interface, and differentiated highlighting is performed based on the intensity indicators in the core hotspot data packets. When researchers have an interactive intention on a hotspot node of a certain macroscopic skeleton, the system dynamically loads and renders the deep detailed information corresponding to the hotspot node according to the requirements.

[0039] This invention achieves end-to-end localized and real-time processing of scientific literature data, from its raw form to final insights. It begins with intelligent parsing and feature extraction of the data stream, proceeds through local context-based popularity assessment and data condensation, and culminates in interactive visualization. All calculations and analyses are completed on the terminal, without relying on high-performance cloud servers. This significantly optimizes data efficiency and resource utilization. Through front-end feature extraction and mid-end data filtering, massive amounts of raw literature data are gradually condensed into a very small amount of core hot topics. It significantly reduces network bandwidth consumption, as only the final condensed core data package needs to be uploaded to the cloud for collaboration, while also alleviating the computational pressure on terminal devices continuously processing large-scale data. It provides grassroots researchers with an efficient, accurate, and low-cost decision support tool. It not only quickly presents macro-level hot topics and supports user interactive exploration to obtain in-depth information, but more importantly, its analysis results deeply integrate the institution's local research preferences and historical background, making hot topic insights more personalized and practical, effectively supporting efficient collaboration in a distributed research environment.

[0040] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0041] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0042] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0043] Figure 1 This is a block diagram of a hotspot analysis terminal for grassroots scientific research institutions based on edge computing and lightweight AI in Embodiment 1 of the present invention;

[0044] Figure 2 This is a schematic diagram of the hotspot analysis terminal for grassroots scientific research institutions based on edge computing and lightweight AI in Embodiment 1 of the present invention.

[0045] Figure 3 This is a block diagram of the vector set acquisition system in Embodiment 2 of the present invention;

[0046] Figure 4 This is a block diagram of the data packet generation system in Embodiment 4 of the present invention;

[0047] Figure 5 This is a block diagram of the differentiated display system in Embodiment 10 of the present invention;

[0048] Figure 6 This is a flowchart illustrating the implementation method of the hotspot analysis terminal for grassroots scientific research institutions based on edge computing and lightweight AI in Embodiment 11 of the present invention. Detailed Implementation

[0049] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0050] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0051] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0052] Example 1: As Figure 1 As shown, this embodiment of the invention provides a hotspot analysis terminal for grassroots research institutions that combines edge computing and lightweight AI, comprising:

[0053] The vector set acquisition system is used to perform non-uniform segmentation of the raw data stream of scientific literature, dynamically adjust the processing granularity based on the data entropy value, and prioritize the parsing of segments with high information density; it uses the hierarchical feature stripping of a lightweight AI model architecture to decompose the content of scientific literature into multi-dimensional feature tuples such as semantic core, methodological tags and data citation network; and obtains a set of standardized and lightweight feature vectors.

[0054] The data packet generation system is used to fuse the feature vector set with the locally generated context weights to generate a set of preliminary hotspot labels with intensity indicators; at the same time, the preliminary hotspot labels with intensity indicators exceeding the adaptive threshold and their associated core hotspot data are marked to generate core hotspot data packets.

[0055] The differentiated display system is used to perform topological structure analysis on core hotspot data packets and construct a correlation map between core hotspots. The macroscopic skeleton of the core hotspot map is presented on the terminal interface, and differentiated highlighting is performed according to the intensity indicators in the core hotspot data packets. When researchers have an interactive intention on a hotspot node of a certain macroscopic skeleton, the system dynamically loads and renders the deep detailed information corresponding to the hotspot node according to the needs.

[0056] The working principle and beneficial effects of the above technical solution are as follows: The vector set acquisition system in this embodiment is used to perform non-uniform segmentation of the original data stream of scientific research literature, dynamically adjust the processing granularity according to the data entropy value, and prioritize the parsing of segments with high information density; using the hierarchical feature stripping of the lightweight AI model architecture, the content of scientific research literature is decomposed into multi-dimensional feature tuples such as semantic core, methodological tags, and data citation network; a set of standardized and lightweight feature vector sets is obtained; the data packet generation system is used to fuse the feature vector set with the locally generated context weights to generate a set of preliminary hot spot tags with intensity indicators; at the same time, preliminary hot spot tags with intensity indicators exceeding the adaptive threshold and their associated core hot spot data are marked to generate core hot spot data packets; the differentiated display system is used to perform topological structure analysis on the core hot spot data packets to construct the association map between core hot spots; the macro skeleton of the core hot spot map is presented on the terminal interface, and differentiated highlighting is performed according to the intensity indicators in the core hot spot data packets; when researchers have an interactive intention on a hot spot node of a certain macro skeleton, the deep detail information corresponding to the corresponding hot spot node is dynamically loaded and rendered according to the needs (the specific principle is as follows). Figure 2 (As shown). The above solution achieves localized and real-time processing of scientific literature data from its raw form to final insights. It begins with intelligent parsing and feature extraction of the data stream, proceeds through local context-based popularity assessment and data condensation, and concludes with interactive visualization. All calculations and analyses are completed on the terminal, without relying on high-performance cloud servers. This greatly optimizes data efficiency and resource utilization. Through front-end feature extraction and mid-end data filtering, massive amounts of raw literature data are gradually condensed into a very small amount of core hot information. It significantly reduces network bandwidth consumption, as only the final condensed core data package needs to be uploaded to the cloud for collaboration, while also alleviating the computational pressure on terminal devices to continuously process large-scale data. It provides grassroots researchers with an efficient, accurate, and low-cost decision support tool. It can not only quickly present macro-level hot topics and support users to interactively explore and obtain in-depth information, but more importantly, its analysis results deeply integrate the institution's local research preferences and historical background, making hot topic insights more personalized and practical, effectively supporting efficient collaboration in a distributed research environment.

[0057] Example 2: Figure 3 As shown, based on Embodiment 1, the vector set acquisition system provided in this embodiment of the invention includes:

[0058] The window mechanism processing subsystem is used to trigger an entropy-aware dynamic window mechanism from the raw data stream of scientific literature. This mechanism continuously monitors the information entropy changes in the raw data stream, and the information entropy is calculated based on the frequency and distribution rate of key terms in a specific field within the text. When the entropy-aware dynamic window mechanism detects that the information entropy exceeds a dynamic baseline, it determines that the currently flowing raw data segment of the scientific literature is a high-information-density segment and anchors the starting point of a processing window. The ending point of the processing window is determined by the point where the information entropy falls back below the dynamic baseline, resulting in a set of data segments of varying lengths, each composed of high-information-density segments.

[0059] The stripping strategy execution subsystem is used to progressively strip the data paragraph set according to the hierarchical feature stripping network. The first layer of the hierarchical feature stripping network scans the input data paragraph set to identify and extract semantic cores that summarize the core claims and conclusions of the scientific literature. After the first layer of extraction is completed, the remaining text content and the initially extracted core semantic units are sent to the second layer. The second layer identifies and extracts elements such as research methods, technical routes, and experimental procedures that support the core semantics, forming methodological labels. The remaining text content after the first two layers of stripping, together with the obtained semantic cores and methodological labels, is input into the third layer. The third layer is responsible for parsing the citation information, data sources, and relationships in the literature to construct a structured data citation network.

[0060] The feature normalization processing subsystem is used to feed the multi-dimensional feature tuples obtained from the decomposition of semantic core, methodological labels, and data reference networks into a feature normalization and vector encoder in parallel; it assigns a vectorization sub-model to each type of feature tuple, mapping the unstructured text content and data reference network information to a unified vector space; each vectorization sub-model outputs a corresponding feature vector fragment, which is then concatenated and normalized according to predetermined rules; generating a set of normalized and lightweight feature vectors.

[0061] The working principle and beneficial effects of the above technical solution are as follows: The window mechanism processing subsystem of this embodiment is used to trigger an entropy-aware dynamic window mechanism by the original data stream of scientific literature. The entropy-aware dynamic window mechanism is used to continuously monitor the information entropy changes of the original data stream of scientific literature. The calculation of information entropy is based on the frequency of occurrence and distribution change rate of key terms in a specific field in the text. When the entropy-aware dynamic window mechanism detects that the information entropy value exceeds a dynamic baseline, it determines that the currently flowing original data segment of scientific literature is a high information density segment and anchors the starting point of a processing window. The ending point of the processing window is determined by the point where the information entropy value falls back below the dynamic baseline. A series of data segment sets of varying lengths, composed of high information density segments, are obtained. The stripping strategy execution subsystem is used to perform a progressive stripping strategy on the data segment set according to the hierarchical feature stripping network. The first layer of the hierarchical feature stripping network scans the input data segment set, identifies and extracts the semantic core that summarizes the core claims and conclusions of the scientific literature. After the first layer of extraction, the remaining text content, along with the initially extracted core semantic units, is fed into the second layer. The second layer identifies and extracts elements such as research methods, technical routes, and experimental procedures to support the core semantics, forming methodological labels. The remaining text content after the first two layers of extraction, along with the obtained semantic core and methodological labels, is fed into the third layer. The third layer is responsible for parsing citation information, data sources, and relationships in the literature, constructing a structured data citation network. The feature normalization processing subsystem is used to feed the multi-dimensional feature tuples of the decomposed semantic core, methodological labels, and data citation network into a feature normalization and vector encoder in parallel. A vectorization submodel is assigned to each type of feature tuple, mapping the unstructured text content and data citation network information into a unified vector space. Each vectorization submodel outputs a corresponding feature vector fragment, which is then concatenated and normalized according to predetermined rules, generating a set of standardized and lightweight feature vectors. The aforementioned scheme, based on a dynamic monitoring mechanism of information entropy, can autonomously identify high-value information fragments in the original document stream. It achieves adaptive window segmentation of non-fixed length through dynamic baseline threshold judgment, ensuring targeted and efficient data collection and avoiding ineffective processing of low-information-density content. The hierarchical processing mode mimics the cognitive hierarchy of human reading: first grasping the core arguments, then understanding the methodological support, and finally sorting out citation relationships, forming a complete knowledge deconstruction chain. The intermediate results output at each level serve as input for the next level, maintaining the correlation between knowledge elements. Specialized vectorization models are used for different types of knowledge elements, preserving the professional attributes of various features while achieving unified alignment of the vector space through standardized splicing. This processing method transforms unstructured document content into machine-computable vector representations, while maintaining the topological relationships between semantic cores, methodological support, and citation networks, providing a structured data foundation for subsequent knowledge mining and analysis.

[0062] In summary, this embodiment transforms the original document stream into a standardized vector set carrying multi-dimensional knowledge features, realizing the transformation process from unstructured text to structured knowledge representation, and providing basic data support for the intelligent processing and analysis of scientific research documents.

[0063] Example 3: Based on Example 2, the window mechanism processing subsystem provided in this embodiment of the invention includes:

[0064] The dynamic scanning component is used to dynamically scan the raw data stream of scientific research literature and extract high-frequency core words and phrases from the local historical processing literature of grassroots scientific research institutions, forming a lightweight dynamic domain terminology set with distinctive institutional research characteristics.

[0065] The continuous sliding component is used to input a dynamic domain terminology set and the original data stream of scientific literature into a sliding calculation window, which slides continuously over the original data stream of scientific literature with a fixed initial length. At each position of the sliding calculation window, the frequency of each term in the dynamic domain terminology set appearing in the current sliding calculation window is counted to obtain the original frequency vector. At the same time, the distribution change rate of each term is calculated to obtain the distribution change rate vector, which is the absolute value of the change between the frequency of the term in the current sliding calculation window and the frequency in the previous adjacent sliding calculation window, capturing the instantaneous fluctuation of the term's attention. After the original frequency vector and the distribution change rate vector are subjected to two-factor frequency domain contribution analysis, they are weighted and fused to generate a comprehensive contribution index representing the information activity level in the current window.

[0066] The dynamic baseline confirmation component is used to feed the comprehensive contribution index into a dynamic baseline calibrator, continuously record the historical distribution of the comprehensive contribution index of all sliding calculation windows over a recent period, and determine the current calculation baseline, i.e., the dynamic baseline, in real time by multiplying the median of the historical distribution by a configurable preset scaling factor.

[0067] The working principle and beneficial effects of the above technical solution are as follows: The dynamic scanning component in this embodiment is used to dynamically scan the original data stream of scientific research literature and extract high-frequency core words and phrases from the local historical processing literature of grassroots scientific research institutions, forming a dynamic domain terminology set that is highly characteristic of institutional research and lightweight; the continuous sliding component is used to input the dynamic domain terminology set and the original data stream of scientific research literature into a sliding calculation window, and continuously slide it on the original data stream of scientific research literature with a fixed initial length; at each position of the sliding calculation window, the frequency of each term in the dynamic domain terminology set appearing in the current sliding calculation window is counted to obtain the original frequency vector; at the same time, the frequency of each term is calculated. The distribution change rate is used to obtain a distribution change rate vector, which is the absolute value of the change in frequency of a term in the current sliding calculation window compared to its frequency in the previous adjacent sliding calculation window, capturing the instantaneous fluctuations in term attention. The original frequency vector and the distribution change rate vector are then weighted and fused after two-factor frequency domain contribution analysis to generate a comprehensive contribution index representing the information activity level within the current window. A dynamic baseline confirmation component is used to input the comprehensive contribution index into a dynamic baseline calibrator, continuously recording the historical distribution of the comprehensive contribution index of all sliding calculation windows over a recent period. The current calculation baseline, i.e., the dynamic baseline, is determined in real time by multiplying the median of the historical distribution by a configurable preset scaling factor. The dynamic scanning component of the above scheme enables adaptive term extraction from the scientific literature data stream; it ensures that the term set has domain specificity and timeliness, avoiding the generalization problem of generalized thesaurus, and enabling subsequent calculations to accurately reflect the focus of specific research institutions. The continuous sliding component enables dynamic capture of multi-dimensional information in the literature stream; using a fixed window sliding mechanism, it can not only identify the absolute frequency of occurrence of terms but also perceive short-term fluctuations in their attention, thus more sensitively reflecting the information change trends in the scientific literature stream. The dynamic baseline confirmation component enables adaptive adjustment of the window activation threshold. Based on historical data, the baseline is dynamically calculated, allowing the system to adapt to information density fluctuations across different time periods and research hotspots, avoiding misjudgments or omissions caused by fixed thresholds. The calibration mechanism ensures the stability and flexibility of window segmentation, making the identification of high-information-density segments more accurate.

[0068] In summary, this embodiment can accurately and in real time filter out segments with high information activity from the scientific literature stream, providing a stable input source for subsequent feature extraction and data normalization.

[0069] Example 4: Figure 4 As shown, based on Embodiment 1, the data packet generation system provided in this embodiment of the invention includes:

[0070] The weighting and acquisition subsystem is used to input the feature vector set into the local interest profile generator. By continuously recording and analyzing implicit feedback data such as researchers' query history, literature browsing time and tagging and collection behavior, it dynamically adjusts the weight values ​​in a lightweight matrix to form a local context weight set that represents the latest research concerns of the research institution.

[0071] The dimension matching subsystem is used to map and align the abstract semantic space of feature vectors with the specific interest space of context weights. If a dimension in a feature vector matches a high-weight dimension in the context weights, the dimension value is non-linearly enhanced; if it does not match, it is suppressed. Each set of feature vectors is assigned an initial strength value that characterizes the strength of its relevance to the local research background. The feature vector set is processed with the local context weight set in a dual-track fusion computing unit to generate a set of undecided labels with initial strength values.

[0072] The intensity threshold comparison subsystem is used to analyze the statistical distribution of all initial intensity values ​​in the set of tags to be judged with initial intensity values, and calculate an adaptive threshold based on the distribution to filter out relatively significant hot spots. The initial intensity value of each tag in the set of tags to be judged is compared with this adaptive threshold. A few preliminary hot spot tags whose initial intensity values ​​exceed the adaptive threshold and their corresponding core hot spot data are captured and extracted from the set of tags to be judged and packaged into a core hot spot data packet.

[0073] The working principle and beneficial effects of the above technical solution are as follows: The weight and acquisition subsystem of this embodiment is used to input the feature vector set into the local interest contour generator. By continuously recording and analyzing implicit feedback data such as researchers' query history, literature browsing time, and tagging and collecting behavior, the weight values ​​in a lightweight matrix are dynamically adjusted to form a local context weight set representing the latest research concerns of the research institution; the dimension matching subsystem is used to map and align the abstract semantic space of the feature vectors with the specific interest space of the context weights; if the dimension in the feature vector matches the high-weight dimension in the context weight, the dimension value is non-linearly enhanced; if it does not match, it is suppressed; each set of feature vectors is... An initial intensity value representing the relevance of a label to the local research background is assigned. The feature vector set is processed in a dual-track fusion computing unit with the local context weight set to generate a set of labels to be judged with initial intensity values. An intensity threshold comparison subsystem analyzes the statistical distribution of all initial intensity values ​​in the label set and calculates an adaptive threshold based on this distribution to filter out relatively significant hotspots. The initial intensity value of each label in the label set is compared with this adaptive threshold. A few preliminary hotspot labels with initial intensity values ​​exceeding the adaptive threshold, along with their corresponding core hotspot data, are captured and extracted from the label set and packaged into a core hotspot data package. This scheme captures the multi-dimensional behavioral trajectories of researchers (querying, browsing, and saving) to quantify implicit research tendencies into a computable weight matrix. The real-time updated local context weight set essentially constructs a digital mirror of the institution's research dynamics. The dimension matching subsystem realizes intelligent translation from the semantic space to the interest space. Through non-linear interaction (enhancement / suppression) between weight dimensions and feature vectors, the raw data stream is transformed into a label set bearing the institutional research fingerprint. This mapping process preserves the original semantics of the academic data while incorporating the institution's unique research preferences. The intensity threshold comparison subsystem employs dynamic statistical methods to identify significant hotspots from a massive pool of undecided labels. An adaptive threshold mechanism ensures that hotspot selection aligns with the overall data distribution characteristics while capturing prominent outliers. The resulting core hotspot data package is essentially a condensed version of research value, filtered through behavioral analysis, dimension matching, and intensity screening.

[0074] In summary, this embodiment transforms scattered research behavior data into structured knowledge packages with intensity markers, providing multi-stage refined data raw materials for research trend analysis. The entire process realizes the transformation from unstructured behavior to structured knowledge products.

[0075] Example 5: Based on Example 4, the intensity threshold comparison subsystem provided in this embodiment of the invention includes:

[0076] An adaptive threshold generation component is used to analyze the overall distribution of the initial intensity values ​​of all the labels to be judged and obtain a statistical distribution. If the statistical distribution is positively skewed, a threshold baseline is adaptively determined based on the data point density in the high intensity value region to capture true outliers. If the statistical distribution is uniform or negatively skewed, the threshold is determined based on the clustering of intensity values. The statistical distribution of the initial intensity values ​​of the labels to be judged is processed to generate an adaptive threshold.

[0077] The decision signal judgment component is used to synchronously judge the intensity of each tag in the tag set to be judged, compare the initial intensity value of each tag to be judged with the adaptive threshold, and generate a binary capture decision signal. If the initial intensity value exceeds the adaptive threshold, the capture decision signal is true, otherwise it is false.

[0078] The compression coding component is used to transmit the actual capture decision signal to a sparsely coded data packetizer, which compresses and encodes the core hotspot data associated with each captured tag to be judged, and encapsulates it together with the tag to be judged and the strength identifier to form a core hotspot data packet.

[0079] The working principle and beneficial effects of the above technical solution are as follows: The adaptive threshold generation component of this embodiment is used to analyze the overall distribution of the initial intensity values ​​of all the tags to be judged and obtain a statistical distribution; if the statistical distribution is positively skewed, a threshold baseline is adaptively determined based on the data point density of the high intensity value area to capture the real outliers; if the statistical distribution is uniform or negatively skewed, the threshold is determined based on the clustering of intensity values; the statistical distribution of the initial intensity values ​​of the tags to be judged is processed to generate an adaptive threshold; the decision signal judgment component is used to synchronously judge the intensity of each tag to be judged in the tag set to be judged, and compare the initial intensity value of each tag to be judged with the adaptive threshold; a binary capture decision signal is generated. If the initial intensity value exceeds the adaptive threshold, the capture decision signal is true, otherwise it is false; the compression encoding component is used to transmit the true capture decision signal to a sparsely encoded guided data packetizer, compress and encode the core hotspot data associated with each captured tag to be judged, and encapsulate it together with the tag to be judged and the intensity identifier to form a core hotspot data packet. The above scheme achieves efficient identification and optimized processing of abnormal data. The adaptive threshold generation component establishes a threshold determination mechanism that balances statistical characteristics and data density by dynamically analyzing data distribution characteristics. For positively skewed distributions, it focuses on identifying tail outliers; for uniform or negatively skewed distributions, it classifies them based on clustering characteristics, ensuring that the threshold setting both conforms to data characteristics and accurately distinguishes anomalies. The capture decision component adopts a real-time comparison mechanism to synchronously match the intensity value of each data point with the dynamic threshold. The binary judgment mode forms a clear anomaly judgment standard, providing a clear decision basis for subsequent processing and avoiding processing uncertainties caused by ambiguity. The compression coding component achieves the selection and packaging of abnormal data through sparse coding technology; only data confirmed as abnormal is compressed and encapsulated with hotspot information, which not only preserves key data characteristics but also significantly reduces storage and transmission load, forming a complete data value extraction chain.

[0080] In summary, this embodiment optimizes data processing efficiency and resource utilization while ensuring the accuracy of anomaly identification through a three-level processing flow of dynamic threshold generation, accurate anomaly detection, and efficient data condensation, providing a purified, high-quality anomaly dataset for subsequent data analysis.

[0081] Example 6: Based on Example 5, the adaptive threshold generation component provided in this embodiment of the invention includes:

[0082] The diffusion index calculation sub-component is used to transform the initial intensity value set of the tag set to be judged into a distribution pattern identifier and its corresponding density distribution spectrum after multi-moment joint analysis and local density diffusion index calculation.

[0083] The candidate value output sub-component is used to import the distribution pattern identifier and density distribution spectrum into a high-density domain threshold extractor if the distribution pattern identifier indicates that the statistical distribution is positively skewed. This extractor analyzes the density distribution spectrum of the high-intensity value region to find the intensity value corresponding to the critical point, which is the threshold baseline candidate value that separates the very few abnormal high-value points from the main data. The distribution pattern identifier and density distribution spectrum are processed by the high-density domain threshold extractor to generate a threshold baseline candidate value.

[0084] If the distribution pattern identifier indicates that the statistical distribution is uniform or negatively skewed, it is directed to a clustered perceptron to scan the entire density distribution spectrum, identify natural clusters formed by intensity values ​​and numerically relative blank areas between clusters; the lower boundary of the largest numerically relative blank area or the upper edge of the most significant cluster is determined as the best boundary to distinguish different significance levels, and a threshold baseline candidate value is output.

[0085] The smoothing correction subcomponent is used to smooth the threshold baseline candidate values ​​to generate an adaptive threshold.

[0086] The working principle and beneficial effects of the above technical solution are as follows: The diffusion index calculation sub-component of this embodiment is used to convert the initial intensity value set of the label set to be judged into a distribution pattern identifier and its matching density distribution spectrum after multi-moment joint analysis and local density diffusion index calculation; the candidate value output sub-component is used to import the distribution pattern identifier and density distribution spectrum into a high-density domain threshold extractor if the distribution pattern identifier indicates that the statistical distribution is positively skewed, analyze the density distribution spectrum of the high intensity value region, and find the intensity value corresponding to the critical point, which is the threshold baseline candidate value for separating a very small number of abnormal high value points from the main data. The distribution morphology identifier and density distribution spectrum are processed by a high-density domain threshold extractor to generate a candidate threshold baseline value. If the distribution morphology identifier indicates a uniform or negatively skewed statistical distribution, it is directed to a clustering perceptron to scan the entire density distribution spectrum, identify natural clusters of intensity values ​​and numerically relative blank areas between clusters. The lower boundary of the largest numerically relative blank area or the upper edge of the most significant cluster is determined as the optimal boundary for distinguishing different saliency levels, and a candidate threshold baseline value is output. A smoothing correction subcomponent is used to smooth the candidate threshold baseline value to generate an adaptive threshold. The above scheme, through a diffusion index calculation subcomponent, transforms the original set of intensity values ​​into quantifiable distribution morphology features, providing a data foundation for subsequent threshold determination. The distribution morphology identifier can clearly distinguish three typical data distribution patterns: positively skewed, uniform, and negatively skewed. The candidate value output subcomponent employs differentiated processing strategies for different distribution patterns: for positively skewed data, it focuses on capturing outlier separation points in high-intensity regions; for uniformly or negatively skewed data, it identifies the boundaries of natural clusters to determine the significance level; ensuring that the threshold extraction method always matches the statistical characteristics of the data itself. The smoothing correction subcomponent optimizes the initially extracted threshold, eliminating possible local fluctuations and enhancing the robustness and practicality of the threshold.

[0087] In summary, this embodiment automatically generates reasonable threshold dividing lines based on the actual distribution characteristics of the input data. The entire process does not rely on preset fixed parameters, but dynamically determines the optimal dividing point by analyzing the statistical characteristics of the data itself, demonstrating the adaptive nature of data-driven processing.

[0088] Example 7: Based on Example 6, the diffusion index calculation sub-component provided in this embodiment of the invention includes:

[0089] The density variation spectrum generation module is used to simultaneously solve the third and fourth moments of the initial intensity value set of the label set to be judged; it obtains the dominant factor of the distribution pattern of the third moment and the auxiliary factor of the distribution pattern of the fourth moment; at the same time, the initial intensity value set is fed into a sliding density evaluation window in parallel. For each sliding position, the number of data points in the window is counted and recorded to form a preliminary density profile; the difference ratio of density values ​​between adjacent windows is calculated to capture the areas of drastic changes in the aggregation and dispersion of data points in the value range, and a density variation spectrum is obtained that describes the degree of concentration and the rate of change of data points in each intensity interval.

[0090] The calculation results of the third moment are used to quantify the asymmetry of the data distribution, that is, its skewness; the calculation results of the fourth moment are used to quantify the kurtosis of the data distribution, that is, its sharpness or flatness.

[0091] The distribution pattern classification module is used to classify distribution patterns according to the rules for determining the interaction between the dominant and auxiliary factors of distribution patterns, and output the distribution pattern classification code.

[0092] The association and encapsulation module is used to classify and encapsulate the distribution pattern as a global pattern identifier, associate and encapsulate it with the detailed local density information contained in the density change spectrum, and generate a distribution pattern identifier containing global pattern and local density features and its matching density distribution spectrum.

[0093] The working principle and beneficial effects of the above technical solution are as follows: The density variation spectrum forming module of this embodiment is used to simultaneously solve the third and fourth moments of the initial intensity value set of the label set to be judged; the dominant factor of the distribution pattern of the third moment and the auxiliary factor of the distribution pattern of the fourth moment are obtained; at the same time, the initial intensity value set is fed into a sliding density evaluation window in parallel. For each sliding position, the number of data points in the window is counted and recorded to form a preliminary density profile; the difference ratio of density values ​​between adjacent windows is calculated to capture the drastic changes in the aggregation and dispersion of data points in the value range, and a density variation spectrum is obtained that describes the degree of concentration and the rate of change of data points in each intensity interval. The calculation results of the third moment of the multi-moment joint analyzer are used to quantify the asymmetry of the data distribution, i.e., its skewness; the calculation results of the fourth moment of the multi-moment joint analyzer are used to quantify the kurtosis of the data distribution, i.e., its sharpness or flatness; the distribution pattern classification module is used to classify the distribution pattern according to the judgment rules of the interaction relationship between the dominant and auxiliary factors of the distribution pattern, and output the distribution pattern classification code; the association and encapsulation module is used to associate and encapsulate the distribution pattern classification code as a global pattern identifier with the detailed local density information contained in the density change spectrum, and generate a distribution pattern identifier containing global pattern and local density features and its matching density distribution spectrum. The density change spectrum generation module of the above scheme comprehensively captures data features through a dual-path analysis mechanism: the moment calculation path extracts two core statistics, the third moment (skewness) and the fourth moment (kurtosis), to quantify the overall morphological characteristics of the distribution; the sliding window path generates continuous local density assessment results, recording the changes in the aggregation state of data points in each intensity interval; the distribution morphology classification module establishes decision rules based on statistics: through the combination relationship between the dominant factor (third moment) and the auxiliary factor (fourth moment), the statistical characteristics of the continuous distribution are transformed into discrete classification codes; the association encapsulation module realizes feature integration: establishing a correspondence between the global classification code and the local density change spectrum to form a composite descriptor that simultaneously contains macroscopic morphological judgment and microscopic density change information.

[0094] In summary, this embodiment completes the transformation process from the original intensity value to the standardized distribution descriptor, providing a data feature representation that combines statistical generalization and local detail for subsequent threshold selection; thus, the determination of the distribution pattern considers both the overall statistical characteristics and retains the interpretability of local density changes.

[0095] Example 8: Based on Example 7, the distribution pattern classification module provided in this embodiment of the invention includes:

[0096] The spatial mapping processing submodule is used to map the two-dimensional coordinates composed of the dominant distribution morphology factor and the auxiliary distribution morphology factor to the morphology discrimination state space. The morphology discrimination state space is divided into multiple different regions, each region corresponding to a distribution morphology. After the morphology discrimination state space mapping processing, the dominant distribution morphology factor and the auxiliary distribution morphology factor are converted into morphology state codes.

[0097] The morphology comparison execution submodule is used to compare the currently obtained morphology state code with the recent historical morphologies in the local historical distribution prior knowledge base; if the current morphology is consistent with the recent mainstream morphology, the confidence of the morphology state code is increased and it is output directly; if the current morphology deviates significantly, a review process is initiated.

[0098] The morphology coding selection submodule is used to select between two possible neighboring morphology codes based on changes in the dominant and auxiliary factors of the distribution morphology, and generate a distribution morphology classification code with confidence weights.

[0099] Among them, the two possible proximity morphology codes refer to the codes corresponding to the two different, predefined distribution morphology categories that are closest to the coordinate point position determined by the dominant distribution morphology factor and the auxiliary distribution morphology factor in the morphology discrimination state space.

[0100] The working principle and beneficial effects of the above technical solution are as follows: The spatial mapping processing submodule of this embodiment is used to map the two-dimensional coordinates composed of the dominant distribution morphology factor and the auxiliary distribution morphology factor to the morphology discrimination state space. The morphology discrimination state space is divided into multiple different regions, each region corresponding to a distribution morphology. The dominant distribution morphology factor and the auxiliary distribution morphology factor are converted into morphology state codes after morphology discrimination state space mapping processing. The morphology comparison execution submodule is used to compare the currently obtained morphology state code with the recent historical morphology in the local historical distribution prior knowledge base. If the current morphology is consistent with the recent mainstream morphology, the confidence of the morphology state code is enhanced and it is directly output. If the current morphology deviates significantly, a review process is initiated. The morphology code selection submodule is used to select between two possible adjacent morphology codes based on the changes of the dominant distribution morphology factor and the auxiliary distribution morphology factor, and generate a distribution morphology classification code with confidence weight. Among them, the two possible adjacent morphology codes refer to the codes corresponding to the two different, predefined distribution morphology categories that are closest to the coordinate point position determined by the dominant distribution morphology factor and the auxiliary distribution morphology factor in the morphology discrimination state space. The above scheme maps the dominant and auxiliary factors to a predefined morphological discrimination space, and realizes the transformation from continuous statistics to discrete morphological codes through spatial region division; it compares with the historical prior knowledge base to verify the typicality of the current morphology, and initiates an additional review process for abnormal morphologies to ensure the reliability of the results; when the coordinate point is located in the boundary area of ​​morphological categories, it makes a weighted selection between adjacent codes according to the trend of statistical changes.

[0101] In summary, this embodiment establishes a standardized classification framework through state-space mapping, utilizes historical data to achieve dynamic credibility assessment, and provides an arbitration mechanism based on statistical trends for critical situations. Finally, it outputs classification codes with confidence weights, allowing for probabilistic descriptions of boundary morphologies while maintaining classification consistency. This ensures rapid determination of mainstream morphologies while handling classification uncertainty caused by minor fluctuations in statistical values.

[0102] Example 9: Based on Example 8, the morphology encoding selection submodule provided in this embodiment of the invention includes:

[0103] The descriptor transformation unit is used to extract the original values ​​of the dominant and auxiliary factors of the distribution pattern, along with their implied trend information, from the local historical distribution prior knowledge base. It extracts the numerical sequences of the dominant and auxiliary factors of the distribution pattern over the most recent calculation periods and plots their respective trajectories over time. The current value is compared with the trajectory to analyze the magnitude and direction of the change in the current value relative to its historical trajectory. The dominant and auxiliary factors of the distribution pattern are then transformed into dynamic descriptors.

[0104] The preliminary decision output unit is used to determine, based on the dynamic change descriptor, which of the two adjacent morphological codes the current change behavior tends to lead to, and to obtain a preliminary tendency decision.

[0105] The discrimination result output unit is used to input the preliminary tendency decision and the dynamic change descriptor into the decision confidence fusion unit. It introduces a decision stability constraint, which tends to select the morphological code that minimizes the overall fluctuation of a series of recent morphological discrimination results.

[0106] The specific content of the decision stability constraint is as follows: During the decision confidence fusion stage, the encoding option that minimizes the difference between the most recent consecutive morphological discrimination results is preferentially selected. This constraint ensures that the system output does not exhibit abrupt changes by evaluating the impact of candidate encodings on the smoothness of the historical discrimination sequence, thereby maintaining the temporal consistency of morphological discrimination.

[0107] The working principle and beneficial effects of the above technical solution are as follows: The descriptor conversion unit in this embodiment extracts the original values ​​of the dominant and auxiliary distribution morphology factors and their implied trend information from the local historical distribution prior knowledge base. It then extracts the numerical sequences of the dominant and auxiliary distribution morphology factors over the most recent calculation periods and plots their respective trajectories over time. The current value is compared with the trajectory to analyze the magnitude and direction of the change in the current value relative to its historical trajectory. The dominant and auxiliary distribution morphology factors are converted into dynamic change descriptors. The preliminary decision output unit determines, based on the dynamic change descriptors, which of the two adjacent morphological codes the current change behavior tends to lead to, obtaining a preliminary tendency decision. The discrimination result output unit inputs the preliminary tendency decision and the dynamic change descriptor into the decision confidence fusion unit, introducing a decision stability constraint that tends to select the morphological code that minimizes the overall fluctuation of a series of recent morphological discrimination results. Through a multi-level linked dynamic analysis mechanism, the above solution systematically compares the real-time data of the distribution morphology factors with their historical evolution patterns, ultimately outputting a morphological code discrimination result with temporal stability. The three core units form a progressive processing chain: first, historical trajectories are extracted to establish a dynamic benchmark; second, preliminary morphological prediction is made based on the current deviation features; and finally, the final code is output through stability optimization.

[0108] Example 10: As Figure 5 As shown, based on Embodiment 1, the differentiated display system provided in this embodiment of the invention includes:

[0109] The correlation degree acquisition subsystem is used to feed core hot data packets into a multi-dimensional relation strength calculator. The feature vector corresponding to each hotspot label is regarded as a mass body in an abstract semantic space. By analyzing the cosine similarity and Euclidean distance between the vectors of mass bodies, and incorporating their respective strength labels as mass weights, a set of multi-dimensional relation strength values ​​that quantify the closeness of the correlation between them are calculated. A relation strength matrix describing the pairwise relationship strength between all hotspots is obtained.

[0110] The topology skeleton generation subsystem is used to select the strongest relationship from the relationship strength matrix as the starting point, merge two related hotspots into an initial community; with the initial community as the core, iteratively absorbs external hotspots that have a relationship strength exceeding the dynamic threshold with any hotspot in the community, and gradually expands the community; until all hotspots are classified into a community or exist as isolated points; each community formed in the end is regarded as a macro-level hotspot research cluster; and a primary topology skeleton composed of several hotspot research clusters is generated.

[0111] The association graph construction subsystem is used to pass the primary topological skeleton to the cross-cluster hub connector to discover and establish important connections between different research clusters, connecting discrete clusters into a complete graph; it scans hotspots in different hotspot research clusters to identify connection pairs with a global threshold for potential cross-domain connections, which are identified as cross-cluster hub links; the primary topological skeleton and the identified cross-cluster hub links are integrated by the cross-cluster hub connector to construct a core hotspot association graph.

[0112] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the correlation degree acquisition subsystem is used to send the core hotspot data packets into a multi-dimensional relation strength calculator, treating the feature vector corresponding to each hotspot label as a mass body in an abstract semantic space; by analyzing the cosine similarity and Euclidean distance between the vectors of the mass bodies, and incorporating their respective strength labels as mass weights, a set of multi-dimensional relation strength values ​​quantifying the closeness of the correlation between them is calculated; a relation strength matrix describing the pairwise relationship strength between all hotspots is obtained; the topology skeleton generation subsystem is used to select the relationship with the highest strength from the relation strength matrix as the starting point, and merge two related hotspots into an initial community; with the initial community as the core, iteratively absorbs any hotspot within the community. External hotspots with relationship strength exceeding a dynamic threshold are progressively expanded into communities until all hotspots are categorized into a community or exist as isolated points. Each ultimately formed community is considered a macro-level hotspot research cluster. A primary topological skeleton composed of several hotspot research clusters is generated. The association graph construction subsystem is used to pass the primary topological skeleton to the cross-cluster hub connector to discover and establish important connections between different research clusters, connecting discrete clusters into a complete graph. Hotspots in different hotspot research clusters are scanned to identify connection pairs with a global threshold for potential cross-domain connections, which are identified as cross-cluster hub links. The primary topological skeleton and the identified cross-cluster hub links are integrated by the cross-cluster hub connector to construct a core hotspot association graph. The above scheme transforms discrete core hotspot data into a visual graph with a clear hierarchical structure through a hierarchical association mining and structure optimization process. The system first quantifies the semantic association strength between hotspots, then constructs hotspot community clusters based on the association strength, and finally integrates discrete clusters through cross-cluster key connections to form a topological network reflecting global association characteristics.

[0113] Example 11: As Figure 6 As shown, based on Examples 1-10, the implementation method of the edge computing and lightweight AI hotspot analysis terminal for grassroots scientific research institutions provided by this invention includes the following steps:

[0114] S100: The original data stream of scientific literature is segmented non-uniformly, and the processing granularity is dynamically adjusted according to the data entropy value, prioritizing the parsing of segments with high information density; using the hierarchical feature stripping of a lightweight AI model architecture, the content of scientific literature is decomposed into multi-dimensional feature tuples such as semantic core, methodological tags and data citation network; resulting in a set of standardized and lightweight feature vectors.

[0115] S200: The feature vector set is fused with the locally generated context weights to generate a set of preliminary hotspot labels with intensity indicators; at the same time, the preliminary hotspot labels with intensity indicators exceeding the adaptive threshold and their associated core hotspot data are marked to generate core hotspot data packages;

[0116] S300: Performs topological analysis on core hotspot data packets to construct a relational graph between core hotspots; presents the macroscopic skeleton of the core hotspot graph on the terminal interface, highlighting it differently based on the intensity markers in the core hotspot data packets; when researchers have an interactive intention on a hotspot node of a certain macroscopic skeleton, dynamically loads and renders the deep detail information corresponding to the hotspot node according to the requirements.

[0117] The working principle and beneficial effects of the above technical solution are as follows: This embodiment first performs non-uniform segmentation on the original data stream of scientific research literature, dynamically adjusts the processing granularity based on the data entropy value, and prioritizes the parsing of segments with high information density; using the hierarchical feature stripping of a lightweight AI model architecture, the content of scientific research literature is decomposed into multi-dimensional feature tuples such as semantic core, methodological tags, and data citation networks; a set of standardized and lightweight feature vectors is obtained; secondly, the feature vector set is fused with locally generated context weights to generate a set of preliminary hotspot tags with intensity indicators; at the same time, preliminary hotspot tags with intensity indicators exceeding the adaptive threshold and their associated core hotspot data are marked to generate core hotspot data packages; finally, the core hotspot data packages are subjected to topological structure analysis to construct a correlation graph between core hotspots; the macroscopic skeleton of the core hotspot graph is presented on the terminal interface, and differentiated highlighting is performed according to the intensity indicators in the core hotspot data packages; when researchers have an interactive intention on a hotspot node of a certain macroscopic skeleton, the deep detail information corresponding to the hotspot node is dynamically loaded and rendered according to the needs. The above solution achieves efficient structured processing of massive amounts of scientific research literature; combined with context-aware intelligent weighting, it accurately captures hot topics in the field and forms core data packages; and finally, through topological visualization and interactive design, it presents both the macro-research context and supports the exploration of micro-details, thus constructing an intelligent analysis terminal that can adapt to the characteristics of scientific research data and balance computational efficiency and cognitive depth.

[0118] Embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0119] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of equivalents of this invention, this invention is also intended to include these modifications and variations.

Claims

1. A hotspot analysis terminal for grassroots research institutions based on edge computing and lightweight AI, characterized in that, Include: The vector set acquisition system is used to perform non-uniform segmentation of the raw data stream of scientific literature, dynamically adjust the processing granularity based on the data entropy value, and prioritize the parsing of segments with high information density; it uses the hierarchical feature stripping of a lightweight AI model architecture to decompose the content of scientific literature into multi-dimensional feature tuples of semantic core, methodological tags and data citation network; and obtains a set of standardized and lightweight feature vectors. The data packet generation system is used to fuse the feature vector set with the locally generated context weights to generate a set of preliminary hotspot labels with intensity indicators; at the same time, the preliminary hotspot labels with intensity indicators exceeding the adaptive threshold and their associated core hotspot data are marked to generate core hotspot data packets. The data packet generation system includes: The weighting and acquisition subsystem is used to input the feature vector set into the local interest contour generator. By continuously recording and analyzing the implicit feedback data of researchers, it dynamically adjusts the weight values ​​in a lightweight matrix to form a local context weight set that represents the latest research concerns of the research institution. The dimension matching subsystem is used to map and align the abstract semantic space of feature vectors with the specific interest space of context weights; each set of feature vectors is assigned an initial strength value that characterizes the strength of its relevance to the local research background. The feature vector set is processed with the local context weight set in the dual-track fusion computing unit to generate a set of undecided labels with initial intensity values; The intensity threshold comparison subsystem is used to analyze the statistical distribution of all current initial intensity values ​​in the set of labels to be judged with initial intensity values, and calculate an adaptive threshold based on the distribution to filter out relatively significant hot spots; The initial strength value of each tag in the tag set to be judged is compared with this adaptive threshold; a few preliminary hotspot tags whose initial strength value exceeds the adaptive threshold and their corresponding core hotspot data are captured and extracted from the tag set to be judged and packaged into a core hotspot data packet; The intensity threshold comparison subsystem includes: An adaptive threshold generation component is used to analyze the overall distribution of the initial intensity values ​​of all the labels to be judged and obtain a statistical distribution. If the statistical distribution is positively skewed, a threshold baseline is adaptively determined based on the data point density in the high intensity value region to capture true outliers. If the statistical distribution is uniform or negatively skewed, the threshold is determined based on the clustering of intensity values. The statistical distribution of the initial intensity values ​​of the labels to be judged is processed to generate an adaptive threshold. The decision signal judgment component is used to synchronously judge the intensity of each tag in the tag set to be judged, and compare the initial intensity value of each tag to be judged with the adaptive threshold. Generate a binary capture decision signal. If the initial intensity value exceeds the adaptive threshold, the capture decision signal is true; otherwise, it is false. The compression coding component is used to transmit the real capture decision signal to a sparsely coded data packetizer, compress and encode the core hotspot data associated with each captured tag to be judged, and encapsulate it together with the tag to be judged and the strength identifier to form a core hotspot data packet; An adaptive threshold generation component, comprising: The diffusion index calculation sub-component is used to transform the initial intensity value set of the tag set to be judged into a distribution pattern identifier and its corresponding density distribution spectrum after multi-moment joint analysis and local density diffusion index calculation. The candidate value output sub-component is used to import the distribution pattern identifier and density distribution spectrum into a high-density domain threshold extractor if the distribution pattern identifier indicates that the statistical distribution is positively skewed. The density distribution spectrum of the high-intensity value region is analyzed to find the intensity value corresponding to the critical point. After the distribution pattern identifier and density distribution spectrum are processed by the high-density domain threshold extractor, a threshold baseline candidate value is generated. If the distribution pattern identifier indicates that the statistical distribution is uniform or negatively skewed, it is directed to a clustering perceptron to scan the entire density distribution spectrum and identify natural clusters of intensity values ​​as well as numerical gaps between clusters. Output a threshold baseline candidate value; The smoothing correction subcomponent is used to smooth the threshold baseline candidate values ​​to generate an adaptive threshold.

2. The grassroots research institution hotspot analysis terminal based on edge computing and lightweight AI as described in claim 1, characterized in that, The diffusion index calculation subcomponent includes: The density variation spectrum generation module is used to simultaneously solve the third and fourth moments of the initial intensity value set of the label set to be judged; it obtains the dominant factor of the distribution pattern of the third moment and the auxiliary factor of the distribution pattern of the fourth moment; at the same time, the initial intensity value set is fed into a sliding density evaluation window in parallel. For each sliding position, the number of data points in the window is counted and recorded to form a preliminary density profile; the difference ratio of density values ​​between adjacent windows is calculated to capture the areas of drastic changes in the aggregation and dispersion of data points in the value range, and a density variation spectrum is obtained that describes the degree of concentration and the rate of change of data points in each intensity interval. The distribution pattern classification module is used to classify distribution patterns according to the rules for determining the interaction between the dominant and auxiliary factors of distribution patterns, and output the distribution pattern classification code. The association and encapsulation module is used to classify and encapsulate the distribution pattern as a global pattern identifier, associate and encapsulate it with the detailed local density information contained in the density change spectrum, and generate a distribution pattern identifier containing global pattern and local density features and its matching density distribution spectrum.

3. The grassroots research institution hotspot analysis terminal based on edge computing and lightweight AI as described in claim 2, characterized in that, The distribution pattern classification module includes: The spatial mapping processing submodule is used to map the two-dimensional coordinates composed of the dominant distribution morphology factor and the auxiliary distribution morphology factor to the morphology discrimination state space. The morphology discrimination state space is divided into multiple different regions, each region corresponding to a distribution morphology. After the morphology discrimination state space mapping processing, the dominant distribution morphology factor and the auxiliary distribution morphology factor are converted into morphology state codes. The morphology comparison execution submodule is used to compare the currently obtained morphology state code with the recent historical morphologies in the local historical distribution prior knowledge base; if the current morphology is consistent with the recent mainstream morphology, the confidence of the morphology state code is increased and it is output directly; if the current morphology deviates significantly, a review process is initiated. The morphology coding selection submodule is used to select between two possible neighboring morphology codes based on changes in the dominant and auxiliary factors of the distribution morphology, generating a distribution morphology classification code with confidence weights.

4. The grassroots research institution hotspot analysis terminal based on edge computing and lightweight AI as described in claim 3, characterized in that, The morphology encoding selection submodule includes: The descriptor transformation unit is used to extract the original values ​​of the dominant and auxiliary factors of the distribution pattern, along with their implied trend information, from the local historical distribution prior knowledge base. It extracts the numerical sequences of the dominant and auxiliary factors of the distribution pattern over the most recent calculation periods and plots their respective trajectories over time. The current value is compared with the trajectory to analyze the magnitude and direction of the change in the current value relative to its historical trajectory. The dominant and auxiliary factors of the distribution pattern are then transformed into dynamic descriptors. The preliminary decision output unit is used to determine, based on the dynamic change descriptor, which of the two adjacent morphological codes the current change behavior tends to lead to, and to obtain a preliminary tendency decision. The discrimination result output unit is used to input the preliminary tendency decision and the dynamic change descriptor into the decision confidence fusion unit. It introduces a decision stability constraint, which tends to select the morphological code that minimizes the overall fluctuation of a series of recent morphological discrimination results.

5. The grassroots research institution hotspot analysis terminal based on edge computing and lightweight AI as described in claim 1, characterized in that, It also includes a differentiated display system for performing topological analysis on core hotspot data packets and constructing a relational graph between core hotspots; presenting the macroscopic skeleton of the core hotspot graph on the terminal interface and highlighting it differentiatedly based on the intensity markers in the core hotspot data packets; and dynamically loading and rendering the deep detail information corresponding to the hotspot node when researchers interact with a hotspot node of a macroscopic skeleton, according to their needs.

6. A method for implementing a hotspot analysis terminal for grassroots research institutions based on edge computing and lightweight AI, used to carry the hotspot analysis terminal for grassroots research institutions based on edge computing and lightweight AI as described in any one of claims 1-5, characterized in that, Includes the following steps: The raw data stream of scientific literature is segmented non-uniformly, and the processing granularity is dynamically adjusted according to the data entropy value, prioritizing the parsing of segments with high information density; using the hierarchical feature stripping of a lightweight AI model architecture, the content of scientific literature is decomposed into multi-dimensional feature tuples of semantic core, methodological tags and data citation network; resulting in a set of standardized and lightweight feature vectors. The feature vector set is fused with the locally generated context weights to generate a set of preliminary hotspot labels with intensity indicators; at the same time, the preliminary hotspot labels with intensity indicators exceeding the adaptive threshold and their associated core hotspot data are marked to generate core hotspot data packages. The system performs topological analysis on core hotspot data packets to construct a relational graph between core hotspots. The macroscopic skeleton of the core hotspot graph is presented on the terminal interface, and differentiated highlighting is performed based on the intensity indicators in the core hotspot data packets. When researchers have an interactive intention on a hotspot node of a certain macroscopic skeleton, the system dynamically loads and renders the deep detailed information corresponding to the hotspot node according to the requirements.