Method and system for automatically and accurately tracking academic activities of scholars based on large language model
Patent Information
- Application Number
- CN202511775404.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-11-28
AI Technical Summary
[0004]本申请提供了基于大语言模型的学者学术活动自动精准追踪方法及系统,解决传统学者学术活动追踪方法依赖静态数据,导致学者科研信息自动跟踪的全面性、准确性与实时性不足的技术问题
本申请提供的基于大语言模型的学者学术活动自动精准追踪方法及系统,涉及人工智能技术领域,解决了传统学者学术活动追踪方法依赖静态数据,导致学者科研信息自动跟踪的全面性、准确性与实时性不足的技术问题,达到了通过多源异构数据动态采集、大语言模型驱动的事件抽取与多维度交叉验证机制,实现学者学术活动的动态更新与精准追踪,提升信息获取效率和准确性的技术效果。
Smart Images

Figure CN121616285B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a method and system for automatically and accurately tracking the academic activities of scholars based on a large language model. Background Technology
[0002] With the accelerated globalization and digitalization of scientific research, scholars' research achievements, academic exchanges, and career dynamics have generated a large amount of traceable data in cyberspace. However, current research on tracking scholars' research dynamics still has significant shortcomings: on the one hand, existing research mainly focuses on tracking the output of scholars' papers or analyzing the career mobility of scholars, lacking systematic and dynamic monitoring of scholars' overall academic activities; on the other hand, traditional methods are mostly based on static resume data, which cannot achieve real-time updates of research activities and accurate semantic understanding, making it difficult to comprehensively depict the research ecosystem of scholars.
[0003] Meanwhile, the explosive growth of internet information and the rapid development of artificial intelligence technology have led to an increasingly richer collection of dynamic data on high-level scholars' online research activities, academic appointments, awards, conference reports, and collaborative projects. This multi-source, heterogeneous dynamic information provides a new data foundation for depicting scholars' research behavior, the evolution of their research interests, and their academic influence. However, traditional methods cannot effectively address the problems of data dispersion, delayed updates, and information noise, resulting in ambiguity in academic activity tracking results, delayed updates to relationship networks, and difficulty in extracting effective research events from unstructured text. Summary of the Invention
[0004] This application provides a method and system for automatic and accurate tracking of scholars' academic activities based on a large language model, which solves the technical problem that traditional methods for tracking scholars' academic activities rely on static data, resulting in insufficient comprehensiveness, accuracy and real-time performance in the automatic tracking of scholars' research information.
[0005] Firstly, this application provides a method for automatically and accurately tracking scholars' academic activities based on a large language model, the method comprising: The process involves: constructing a calibrated crawling frequency for users, which is the initial crawling plan; performing feature extraction on target scholars to establish a scholar feature set, which includes activity features, timeliness features, academic contribution features, event-driven features, and relationship network features, and the scholar feature set is equipped with a source reliability identifier; inputting the scholar feature set with the source reliability identifier into a frequency dynamic compensation network to output a frequency correction factor, which is set with a crawling bias; after compensating the calibrated crawling frequency according to the frequency correction factor, performing data collection on target scholars; performing multi-dimensional scholar identity authentication on the target scholar data collection results, which includes research field semantic feature verification, academic relationship network verification, and spatiotemporal consistency verification; using the multi-dimensional scholar identity authentication results to perform event extraction verification of the academic activity extraction model, and performing academic activity tracking and management based on the event extraction verification results.
[0006] Secondly, this application provides an automatic and accurate tracking system for scholars' academic activities based on a large language model, the system comprising: The system comprises the following modules: a crawling frequency construction module (constructing a calibrated crawling frequency for the user, which is the initial crawling plan); a feature extraction module (performing feature extraction of the target scholar, establishing a scholar feature set including activity features, timeliness features, academic contribution features, event-driven features, and relationship network features, with a source reliability identifier); a data acquisition module (inputting the scholar feature set with the source reliability identifier into a frequency dynamic compensation network, outputting a frequency correction factor, which has a crawling bias; after compensating the calibrated crawling frequency according to the frequency correction factor, data acquisition of the target scholar is performed); an identity authentication module (performing multi-dimensional scholar identity authentication on the target scholar data acquisition results, including research field semantic feature verification, academic relationship network verification, and spatiotemporal consistency verification); and an academic activity tracking module (using the multi-dimensional scholar identity authentication results to verify the event extraction of the academic activity extraction model, and managing academic activities based on the event extraction verification results).
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: The method and system for automatic and accurate tracking of scholars' academic activities based on a large language model provided in this application relate to the field of artificial intelligence technology. It solves the technical problem that traditional methods for tracking scholars' academic activities rely on static data, resulting in insufficient comprehensiveness, accuracy, and real-time performance in the automatic tracking of scholars' research information. It achieves the technical effect of dynamic updating and accurate tracking of scholars' academic activities through dynamic acquisition of multi-source heterogeneous data, event extraction driven by a large language model, and a multi-dimensional cross-validation mechanism, thereby improving the efficiency and accuracy of information acquisition. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a flowchart illustrating the method for automatically and accurately tracking scholars' academic activities based on a large language model, as described in this application.
[0010] Figure 2 This is a schematic diagram of the structure of the automatic and accurate tracking system for scholars' academic activities based on a large language model, as proposed in this application.
[0011] Figure labeling: Crawling frequency construction module 11, Feature extraction module 12, Data acquisition module 13, Identity authentication module 14, Academic activity tracking module 15. Detailed Implementation
[0012] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0013] Example 1, as Figure 1 As shown, this application provides a method for automatically and accurately tracking scholars' academic activities based on a large language model. This method includes: Establish the user's calibrated crawling frequency, which is the initial crawling plan.
[0014] Specifically, before initiating the tracking of scholars' academic activities, a configurable initial crawling plan is established based on the needs of the target users for data collection scheduling. This initial crawling plan has a set calibration crawling frequency. For newly added scholars, this calibration crawling frequency defaults to a medium frequency, such as once a week. For existing scholars, the frequency is set according to their activity level. Generally, highly active scholars correspond to a high frequency, such as once daily or every two days; moderately active scholars correspond to a medium frequency, such as once a week; inactive scholars correspond to a low frequency, such as once every two weeks or once a month; and dormant scholars correspond to a very low frequency, such as once a quarter or only passively triggered by collaboration events with other scholars. By setting the crawling frequency based on changes in scholar behavior and event-driven factors, both real-time collection and system stability can be balanced, achieving an efficient academic activity data acquisition process.
[0015] The target scholars are subjected to feature extraction, and a scholar feature set is established. The scholar feature set includes activity features, timeliness features, academic contribution features, event-driven features, and relationship network features. The scholar feature set is equipped with a source reliability identifier.
[0016] Specifically, the study first automatically identifies information fragments related to the target scholar from data sources such as university websites, research institution homepages, academic conference platforms, media reports, and academic databases. Then, utilizing the semantic understanding capabilities of natural language processing and large language models, the text content undergoes topic identification, event classification, and entity matching to extract key attributes reflecting the scholar's research dynamics. This extracts a scholar feature set, which includes activity features, timeliness features, academic contribution features, event-driven features, and relationship network features. Activity features quantify the frequency of a scholar's research activities within a specific time period; timeliness features describe the freshness and update frequency of information to reflect the real-time nature of the scholar's research behavior; academic contribution features measure the importance of events; event-driven features identify key event types that trigger data collection or frequency adjustments; and relationship network features describe the academic collaborations and social connections between the scholar and other scholars and institutions. Furthermore, to ensure the credibility of data sources and the reliability of feature results, a source reliability identifier is attached to each data point during the construction of the scholar feature set. This identifier is generated by weighting factors such as the authority, timeliness, information density, and historical accuracy of the data source, and is used to distinguish between high-credibility and low-credibility sources. Through these steps, a structured, computable scholar feature set with source reliability identifiers can be formed, providing a high-quality input foundation for subsequent frequency dynamic compensation and scholar identity authentication.
[0017] The scholar feature set with source reliability identifier is input into the frequency dynamic compensation network, and the frequency correction factor is output. The frequency correction factor is set with crawling bias. After the crawling frequency compensation is calibrated according to the frequency correction factor, the data collection of the target scholar is performed.
[0018] Specifically, after obtaining the scholar feature set with source reliability indicators, this feature set is input into a pre-constructed frequency dynamic compensation network. This network calculates a frequency correction factor in the computational layer based on the standardized scholar feature set, dynamically determining the current frequency correction factor. This factor also includes a crawling bias parameter, set according to current task requirements or the type of recent events involving the scholar. This bias parameter characterizes the priority of specific scholars or event types in data collection. For example, higher crawling frequency weights are assigned to highly active scholars or significant research events, while the collection frequency is automatically reduced for targets with long-term inactivity or low source reliability. After obtaining the frequency correction factor, the calibrated crawling frequency is compensated and corrected by multiplying the frequency correction factor by the calibrated crawling frequency to generate a real-time updated crawling frequency. Finally, based on the compensated crawling frequency and the set collection constraints, data is collected from the target scholars to construct the target scholar data collection results, ensuring high-frequency updates and dynamic tracking of important scholar information.
[0019] Furthermore, the scholar feature set with source reliability identifiers is input into the frequency dynamic compensation network, and the output frequency correction factor includes: The frequency dynamic compensation network is used to normalize the feature values of the scholar feature set; based on the normalization result, the frequency correction factor is calculated through the computation layer, as follows: ;in, Characterizing the frequency correction factor, The total number of features in the scholar's feature set. Represent any feature in the set of characteristics of scholars. For the first The weight of each feature, Characterizing the first Normalized eigenvalues of a feature.
[0020] Specifically, in the calculation of the frequency correction factor, the input scholar feature set is first normalized to eliminate differences in units and numerical ranges between different features. During this process, each feature value, such as activity level, timeliness score, academic contribution value, event frequency, and number of collaborations among scholars, is standardized according to its overall sample distribution. Minimum-maximum normalization or Z-score is commonly used to map all feature values to [0,1] or a standard normal distribution interval with a mean of 0, thus ensuring comparability of features across dimensions in network computation. Subsequently, the normalized scholar feature set is sent to the input layer of the frequency dynamic compensation network. The input layer passes the normalized feature values from the scholar feature set to the computation layer, which then calculates the frequency correction factor based on the importance weight of each feature and its corresponding normalized feature value. The specific calculation method is as follows: ;in, Characterizing the frequency correction factor, The total number of scholar features in the scholar feature set. The first in the characterization of scholars Item features, For the first The weights of each feature, Characterizing the first The normalized eigenvalues of the features. The frequency correction factor calculated by this formula can compensate for the calibrated crawling frequency, enabling dynamic and personalized data acquisition scheduling. This allows the acquisition process to focus on high-value data, reduce redundant access, and improve overall real-time performance and resource utilization efficiency.
[0021] Furthermore, after calibrating and compensating the crawling frequency according to the frequency correction factor, data collection from the target scholar is performed, including: An event-driven dynamic tracking strategy is established based on the crawling bias. Data collection from the target scholar is performed using the dynamic tracking strategy and a compensated crawling frequency to establish a first data collection result. Data source characteristics are obtained, and data source priorities are configured based on these characteristics. A first collection constraint is established based on the data source priority. A relationship triggering strategy is established based on the target scholar's academic relationship network, and a second collection constraint is established using the relationship triggering strategy. Data collection from the target scholar is performed based on the first collection constraint, the second collection constraint, and the compensated crawling frequency to establish a second data collection result. The target scholar data collection result is constructed based on the first and second data collection results.
[0022] Specifically, the system first establishes an event-driven dynamic tracking strategy based on the crawling bias set on the frequency correction factor. This strategy uses a large language model to identify the types of recent research events of scholars (such as published papers, academic appointments, conference reports, awards, etc.) and predicts the types of new events and related information sources that may occur in the next stage. For example, when it is detected that a scholar has just presented a paper at an international conference, the crawling priority of their personal homepage, laboratory website, and academic social media platforms will be automatically increased; when it is detected that a scholar has just been appointed to a new position, the crawling priority of their institution's official website honors section, personal resume page, and relevant academic social media platforms will be automatically increased. The system then combines the dynamic tracking strategy with the compensated crawling frequency to perform the first data collection and generate the first data collection result. Subsequently, the characteristics of the data sources collected in this round are obtained, such as source category, update frequency, information density, and authority. Based on the normalized weighted results of these data source characteristics, priorities are assigned to different data sources to construct the first collection constraint. This first collection constraint restricts the access to high-value and high-credibility data sources to be prioritized in the next round of collection, thereby optimizing resource utilization. Subsequently, a relationship-triggered strategy is established based on the target scholar's academic network, including co-authorship, mentorship, and institutional colleagues. When a research event involving a relevant scholar is detected, a crawling task for that scholar is immediately triggered. This ensures real-time updates even before the scheduled collection period. For example, when scholar D publishes a new paper, the academic network identifies scholar E as a co-author or corresponding author. In this case, the relationship-triggered strategy automatically triggers an immediate crawl of scholar E's relevant page. Based on this strategy, a second collection constraint is defined to achieve event-driven data collection. Then, under the combined effect of the first and second collection constraints and the compensated crawling frequency, the system executes a second round of data collection, obtaining the second data collection results. Finally, by deduplicating the first and second data collection results, and then integrating the deduplicated results, a unified and complete set of target scholar data is constructed, ensuring the comprehensiveness, timeliness, and consistency of the collected content.
[0023] The data collection results of the target scholars are subjected to multi-dimensional scholar identity authentication, which includes research field semantic feature verification, academic relationship network verification, and spatiotemporal consistency verification.
[0024] Specifically, after data collection, the system performs multi-dimensional scholar identity verification on the target scholars' academic information to ensure the accuracy of the collected data and its consistency with the scholars' actual research activities. This multi-dimensional scholar identity verification process includes research field semantic feature verification, academic relationship network verification, and spatiotemporal consistency verification. Research field semantic feature verification involves extracting keywords and academic terms from the scholar's data collection results to construct a set of features related to the scholar's research field. These features are then compared with the scholar's past research interests and academic background to calculate the semantic similarity, thus proving whether the scholar's identity characteristics are consistent with their research field. Academic relationship network verification is based on the scholar's collaboration records and academic relationship network to determine whether the academic events or collaboration information appearing in the data collection results are consistent with known academic relationship networks, helping to identify whether the scholar participated in genuine research collaborations. Spatiotemporal consistency verification checks the consistency of the collected results in terms of time and space. The system compares the timeline of the scholar's historical data and events to ensure that the time of the event is consistent with the scholar's other activities, avoiding errors caused by time errors or information duplication during data collection. Authentication through these three dimensions can effectively verify the identity of the target scholar, ensure the authenticity and reliability of the collected data, and provide high-quality basic data for subsequent academic activity tracking and research.
[0025] Furthermore, the data collection results of the target scholars will be used for multi-dimensional scholar identity verification, including: The semantic feature verification of the research domain includes: extracting keywords and entities from the data collection results of the target scholars, establishing a domain-related feature set, which includes core research topics, keywords, execution methods, and technical terms; obtaining the research interest description of the target scholars, converting the research interest description into a high-dimensional vector, and calculating the semantic similarity between the high-dimensional vector and the domain-related feature set; and completing the semantic feature verification of the research domain based on the semantic similarity.
[0026] Specifically, when validating the semantic features of the target scholar's data collection results, the system first uses keyword and entity extraction techniques, such as TF-IDF, Conditional Random Field (CRF), and BERT, to identify key information related to the scholar's research from the data collection results. This key information typically includes core vocabulary and terms that reflect the scholar's research content and direction, such as core research themes, keywords, implementation methods, and technical terms. By summarizing this key information, a domain-related feature set for the scholar is constructed, providing foundational data for subsequent verification steps. Next, the system obtains the scholar's research interest description, which usually comes from the scholar's personal homepage, research summary, or publicly available academic archives. This description includes the scholar's research direction, long-term research areas of interest, and key projects participated in. The system then uses embedding models such as Sentence-BERT to transform the research interest description into a high-dimensional vector, allowing the scholar's research interests to be mathematically compared with the domain-related feature set. Next, cosine similarity is used to calculate the semantic similarity between the high-dimensional vector and the domain-related feature set. This measures the similarity between the scholar's research interests and the content in their domain-related feature set. A higher similarity indicates that the scholar's research interests are highly consistent with the themes, methods, and terminology of their academic activities, verifying the correctness of their research field. Furthermore, if the scholar's published papers already exist in the knowledge base, LDA topic modeling is used to determine whether the topic distribution of the new text is consistent with the scholar's mainstream topic distribution. This consistency is then averaged with the semantic similarity to obtain the final semantic similarity. Finally, if this semantic similarity meets a predetermined threshold, the scholar's research field is considered valid and consistent, and the verification passes; otherwise, an anomaly is triggered, and the verification fails. This verification process ensures that the scholar's data collection results are highly consistent with their actual research direction, thereby improving the credibility and validity of the data.
[0027] Furthermore, multi-dimensional scholar identity verification of the target scholar data collection results also includes: Performing academic relationship network verification includes: establishing a relationship extraction feature library; using the relationship extraction feature library to extract relationships from the data collection results of the target scholars; establishing relationship extraction results; using the academic relationship network to compare the relationship extraction results; and completing the academic relationship network verification based on the relationship comparison results.
[0028] Specifically, when validating the academic relationship network of the data collected from target scholars, a relationship extraction feature library is first established. This library contains various extraction rules designed based on common collaboration patterns in academic fields. These rules can help identify collaboration relationships, shared research areas, and joint publications among scholars. For example, rules could be set as follows: "If two scholars co-author the same article, extract the article as a collaborative paper and mark the two scholars as collaborators." "If two scholars have presented or participated in organizing the same academic conference, they have an academic activity collaboration relationship." "If two scholars have participated in the execution or collaboration of the same research project, they have a project collaboration relationship." "If two scholars hold the same or related positions in the same academic institution, they can be considered to have an academic position collaboration relationship." Subsequently, using the established relationship extraction feature library, relationship extraction is performed on the data collected from the target scholars. During this process, by comparing the information extracted from the data collection results with the patterns in the extraction rules, collaboration relationships or other academic connections between scholars are automatically identified, and corresponding relationship extraction results are generated based on the identification results. These relationship extraction results represent the association data between the scholar and other scholars and institutions. Next, the extracted relationships are compared using the academic relationship network of scholars. This network is a multi-dimensional graph neural network model, where nodes represent scholars and edges represent relationships between them, such as collaboration or working together. By comparing the extracted relationships with the target scholar's academic relationship network, it's possible to check if the extracted relationships match known relationships within that network, and whether there are duplicate, conflicting, or new, unidentified relationships. If the extracted relationships match existing relationships in the academic relationship network, the scholar's academic network verification is successful. If inconsistencies or mismatches are found, a warning is issued, and the academic relationship network verification fails. This verification process ensures that the data collection results from academic activities accurately reflect scholars' academic collaborations and social networks, preventing the spread of misinformation and enhancing data reliability.
[0029] Furthermore, multi-dimensional scholar identity verification of the target scholar data collection results also includes: Performing spatiotemporal consistency verification includes: acquiring historical data collected by the target scholar; extracting data from the historical data that meets a preset confidence threshold to construct a timeline for the target scholar; using the timeline to monitor conflicts in the data collection results of the target scholar, including joint detection of institutional conflicts and time conflicts; and completing spatiotemporal consistency verification based on the conflict detection results.
[0030] Specifically, when verifying the spatiotemporal consistency of data collected from the target scholar, the historical data of the target scholar is first acquired. This historical data typically includes records of the scholar's past research activities, academic events, published papers, conference participation, and changes in positions. Then, the confidence levels of these historical data are compared with a pre-set confidence threshold. Data greater than or equal to the pre-set confidence threshold are extracted and arranged chronologically to construct the target scholar's timeline. This timeline is a sequence of events arranged chronologically, showcasing all the scholar's important research activities, job changes, academic publications, etc., and is timestamped. Next, this timeline is used to monitor for conflicts in the data collection results. This conflict monitoring includes joint detection of institutional and temporal conflicts, aiming to check for inconsistencies in the target scholar's academic data. In institutional conflict detection, the system checks the timeline to see if the target scholar has overlapping associations with multiple academic institutions or research teams at different points in time. For example, if data collection results show that the scholar is employed at University A, but the target scholar's timeline shows that the scholar left University A last year, it indicates that the data collection results contain inaccurate information, and an institutional conflict is considered to exist. In time conflict detection, the system analyzes the event timestamps of the target scholar based on the timeline to ensure that the scholar's activities are temporally reasonable. For example, if the scholar's public activities (such as conference reports, changes in academic positions, etc.) overlap in time, or if the time of certain academic events does not match the time of other known activities of the scholar, it will be marked as a time conflict. Finally, based on the conflict detection results, the system performs spatiotemporal consistency verification. If a time conflict or institutional conflict is detected, it will indicate that the data has potential errors or inconsistencies, and the spatiotemporal consistency verification will fail. Spatiotemporal consistency verification will only succeed if there are no conflicts, ensuring the accuracy and reliability of the data.
[0031] The event extraction and verification of the academic activity extraction model is carried out using the results of multi-dimensional scholar identity authentication, and the academic activity tracking and management is carried out based on the results of the event extraction and verification.
[0032] Specifically, after completing multi-dimensional scholar identity verification, the verification result is used as input. Combined with a pre-defined general extraction mode, an academic activity extraction model is used to extract events, yielding the event extraction results. Subsequently, event deduplication and verification are performed based on these results, resulting in event extraction verification results. Afterward, the system will use these verification results for academic activity tracking and management. The goal of this tracking and management is to update scholars' research activities in real-time and accurately. This means that various academic activities, such as new research projects, academic reports, and participation in academic organizations, will be continuously tracked and recorded chronologically. Furthermore, the system will dynamically adjust the data collection frequency and priority based on the priority and importance of different academic events to ensure that the most important academic activities are captured promptly, guaranteeing the reliability and timeliness of research data.
[0033] Furthermore, the event extraction verification of the academic activity extraction model utilizes multi-dimensional scholar identity authentication results, including: A general extraction pattern is constructed, which includes key attributes of event types, including report title, organizer, inviter, and collaborators. The large language model in the academic activity extraction model is used to perform event extraction based on the general extraction pattern and multi-dimensional scholar identity authentication results, and the event extraction results are established.
[0034] Specifically, when using the academic activity extraction model for event extraction, a general extraction pattern is first constructed. This general extraction pattern forms the foundation of the academic activity extraction model, aiming to standardize the extraction methods for different types of academic activities. This allows the system to uniformly identify and extract key information from academic events. The general extraction pattern is the union of all possible outputs, mainly including key attributes related to the event type, such as report title, organizer, inviter, and collaborators. These key attributes help describe the basic elements of academic events. Subsequently, precise prompt words are set for the academic activity extraction model, which is then used to execute specific event extraction tasks. This academic activity extraction model can be built based on GPT or BERT, and based on the constructed general extraction pattern, with the support of multi-dimensional scholar identity authentication results, it can automatically identify and extract key attribute information related to academic activities. This information is then structured to generate event extraction results. These results serve as the foundational data for academic activity tracking, used for subsequent academic activity analysis, tracking, and management to ensure information consistency and reliability.
[0035] Furthermore, establishing event extraction results also includes: The event extraction results are subjected to event deduplication and verification processing, including: performing rule-based deduplication processing on the event extraction results using unique fingerprints and key attributes; using the large language model to vectorize the academic report titles, academic names, and event times in the event extraction results, calculating the similarity of the vectorized results, and performing semantic-based deduplication processing; using a machine learning model to perform model analysis on the event extraction results, and performing model-based deduplication processing; and completing event extraction verification based on rule-based deduplication, semantic-based deduplication, and model-based deduplication.
[0036] Specifically, after obtaining the event extraction results, to ensure the accuracy, uniqueness, and non-duplication of the event data, the results undergo deduplication and verification. This process begins with initial deduplication using the event's unique fingerprint and key attributes. The unique fingerprint is generated by hashing core event information such as scholar ID, event type, standardized event time, and report title, creating a unique identifier. This fingerprint identifies the event's uniqueness; even slight differences in event descriptions can be identified as identical events. This unique fingerprint helps determine if an event is duplicated from existing events in the database, thus avoiding duplicate records. Key attributes use a unique identifier to mark the event; identical identifiers indicate duplicate records. Furthermore, for data crawled from web pages, URL comparison can be used to avoid duplicate data. After rule-based deduplication, the system uses a large language model to vectorize the event titles, academic names, and event times in the extracted results, and uses word embedding technology to convert the text data into vector representations. Subsequently, cosine similarity is used to calculate the similarity between the vectorized result and all event vectors of the same scholar, event type, and adjacent time periods in the knowledge base. If the similarity between two events is higher than a threshold, it indicates that they may be semantically duplicated events. In this case, semantic deduplication can be performed to remove those semantically similar or identical events. After semantic deduplication, the system also uses machine learning models, such as XGBoost and LightGBM, for complex event deduplication analysis. This machine learning model is pre-trained using manually labeled duplicate and non-duplicate event pairs through iterative training steps such as forward propagation, loss calculation (e.g., cross-entropy loss), backpropagation, and parameter optimization. It can intelligently identify duplicate events in complex situations based on samples and rules in historical data, thereby ensuring that the final event data is unique and accurate. Finally, the system comprehensively verifies the results after rule-level deduplication, semantic-level deduplication, and model-level deduplication to complete the event extraction verification. These deduplication layers complement each other, effectively reducing the occurrence of duplicate events and ensuring that each academic activity appears only once in the database, guaranteeing the accuracy, uniqueness, and consistency of the data.
[0037] Furthermore, the event extraction and validation process, based on rule-level deduplication, semantic-level deduplication, and model-level deduplication, also includes: Perform cross-validation on the deduplicated event extraction results, including internal consistency verification, external consistency verification, and source credibility verification; filter the event extraction results based on the cross-validation results, and establish event extraction verification results.
[0038] Specifically, after the event extraction deduplication process is completed, the deduplicated event extraction results undergo cross-validation to further verify their accuracy and reliability. This cross-validation includes three aspects: internal consistency verification, external consistency verification, and source credibility verification. In internal consistency verification, the system checks whether the deduplicated event data maintains consistency in internal attributes and content; that is, it checks the event timestamps to ensure the logical chronological order of events. If multiple records of the same event show inconsistencies in time—for example, if the event occurred after the webpage publication time or after the scholar's birth date—it will be marked as potentially conflicting data. The system also verifies whether key attributes in the event match the scholar's identity characteristics to ensure that the information does not contain internal contradictions. For example, if the scholar's report topic does not match their research field, or if the scholar is invited to give a cutting-edge academic report, there is an internal inconsistency. Through internal consistency verification, it can be ensured that the inherent information of each event is free of errors or conflicts and accurately reflects the actual situation of academic activities. In external consistency verification, the system compares the extracted event results with external data sources to ensure that the extracted academic activity information is consistent with public domain or other authoritative information. For example, if the title of an academic report is already recorded on the official page of an academic conference, the system verifies whether the event actually occurred by comparing the report title, time, and organizing institution. If the time of the event conflicts with the scholar's schedule displayed on an official page, it is verified as an external inconsistency. The system also performs the same relationship network verification on collaborators, inviters, and other individuals involved in academic activities to check whether these individuals have any relationship with the target scholar. Through external consistency verification, it ensures that the extracted events conform to the actual situation and avoids erroneous data extraction and false information. In source credibility verification, the system evaluates the data sources in the event extraction results, that is, it weights the source authority score, information originality score, and publication timeliness score. Among them, the source authority score is usually highest for official institution websites and well-known journal websites, and lower for personal blogs, Wikipedia, and commercial news sites; the information originality score is usually highest for first publication, and lower for reprinted or compiled reports. By assessing the credibility of each data source, the system ensures that the final retained data comes from reliable channels, avoiding erroneous data due to unreliable information sources. After completing the above cross-validation, the system filters the event extraction results based on the validation results, treating the validated events as valid data and retaining them. This establishes an event extraction validation result, ensuring the high quality and reliability of academic activity data and providing authentic and effective foundational data for subsequent data analysis, academic reports, and research management.
[0039] In summary, the automatic and accurate tracking method for scholars' academic activities based on a large language model provided in this application has at least the following technical effects: by dynamically collecting multi-source heterogeneous data, extracting events driven by a large language model, and using a multi-dimensional cross-validation mechanism, the method achieves dynamic updates and accurate tracking of scholars' academic activities, thereby improving the efficiency and accuracy of information acquisition.
[0040] Example 2 is based on the same inventive concept as the method for automatically and accurately tracking scholars' academic activities based on a large language model in the previous examples, such as... Figure 2 As shown, this application provides an automatic and accurate tracking system for scholars' academic activities based on a large language model. The system includes: The crawling frequency construction module 11 constructs the user's calibrated crawling frequency, which is the initially set crawling plan; the feature extraction module 12 performs feature extraction on the target scholar, establishes a scholar feature set, which includes activity features, timeliness features, academic contribution features, event-driven features, and relationship network features, and the scholar feature set is set with source reliability identifiers; the data acquisition module 13 inputs the scholar feature set with source reliability identifiers into the frequency dynamic compensation network, outputs a frequency correction factor, which is set with crawling bias, and performs calibrated crawling frequency compensation according to the frequency correction factor before performing data acquisition on the target scholar; the identity authentication module 14 performs multi-dimensional scholar identity authentication on the target scholar data acquisition results, which includes research field semantic feature verification, academic relationship network verification, and spatiotemporal consistency verification; the academic activity tracking module 15 uses the multi-dimensional scholar identity authentication results to perform event extraction verification of the academic activity extraction model, and performs academic activity tracking management based on the event extraction verification results.
[0041] Furthermore, the data acquisition module 13 includes: An event-driven dynamic tracking strategy is established based on the crawling bias. Data collection from the target scholar is performed using the dynamic tracking strategy and a compensated crawling frequency to establish a first data collection result. Data source characteristics are obtained, and data source priorities are configured based on these characteristics. A first collection constraint is established based on the data source priority. A relationship triggering strategy is established based on the target scholar's academic relationship network, and a second collection constraint is established using the relationship triggering strategy. Data collection from the target scholar is performed based on the first collection constraint, the second collection constraint, and the compensated crawling frequency to establish a second data collection result. The target scholar data collection result is constructed based on the first and second data collection results.
[0042] Furthermore, the data acquisition module 13 includes: The frequency dynamic compensation network is used to normalize the feature values of the scholar feature set; based on the normalization result, the frequency correction factor is calculated through the computation layer, as follows: ;in, Characterizing the frequency correction factor, The total number of features in the scholar's feature set. Represent any feature in the set of characteristics of scholars. For the first The weight of each feature, Characterizing the first Normalized eigenvalues of a feature.
[0043] Furthermore, the identity authentication module 14 includes: The semantic feature verification of the research domain includes: extracting keywords and entities from the data collection results of the target scholars, establishing a domain-related feature set, which includes core research topics, keywords, execution methods, and technical terms; obtaining the research interest description of the target scholars, converting the research interest description into a high-dimensional vector, and calculating the semantic similarity between the high-dimensional vector and the domain-related feature set; and completing the semantic feature verification of the research domain based on the semantic similarity.
[0044] Furthermore, the identity authentication module 14 includes: Performing academic relationship network verification includes: establishing a relationship extraction feature library; using the relationship extraction feature library to extract relationships from the data collection results of the target scholars; establishing relationship extraction results; using the academic relationship network to compare the relationship extraction results; and completing the academic relationship network verification based on the relationship comparison results.
[0045] Furthermore, the identity authentication module 14 includes: Performing spatiotemporal consistency verification includes: acquiring historical data collected by the target scholar; extracting data from the historical data that meets a preset confidence threshold to construct a timeline for the target scholar; using the timeline to monitor conflicts in the data collection results of the target scholar, including joint detection of institutional conflicts and time conflicts; and completing spatiotemporal consistency verification based on the conflict detection results.
[0046] Furthermore, the academic activity tracking module 15 includes: A general extraction pattern is constructed, which includes key attributes of event types, including report title, organizer, inviter, and collaborators. The large language model in the academic activity extraction model is used to perform event extraction based on the general extraction pattern and multi-dimensional scholar identity authentication results, and the event extraction results are established.
[0047] Furthermore, the academic activity tracking module 15 includes: The event extraction results are subjected to event deduplication and verification processing, including: performing rule-based deduplication processing on the event extraction results using unique fingerprints and key attributes; using the large language model to vectorize the academic report titles, academic names, and event times in the event extraction results, calculating the similarity of the vectorized results, and performing semantic-based deduplication processing; using a machine learning model to perform model analysis on the event extraction results, and performing model-based deduplication processing; and completing event extraction verification based on rule-based deduplication, semantic-based deduplication, and model-based deduplication.
[0048] Furthermore, the academic activity tracking module 15 includes: Perform cross-validation on the deduplicated event extraction results, including internal consistency verification, external consistency verification, and source credibility verification; filter the event extraction results based on the cross-validation results, and establish event extraction verification results.
[0049] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0050] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for automatically and accurately tracking scholars' academic activities based on a large language model, characterized in that: The method includes: Establish the user's calibrated crawling frequency, which is the initially set crawling plan; The target scholars are subjected to feature extraction, and a scholar feature set is established. The scholar feature set includes activity features, timeliness features, academic contribution features, event-driven features, and relationship network features. The scholar feature set is equipped with a source reliability identifier. The scholar feature set with source reliability identifier is input into the frequency dynamic compensation network, and the frequency correction factor is output. The frequency correction factor is set with crawling bias. After the crawling frequency compensation is calibrated according to the frequency correction factor, the data collection of the target scholar is performed. The data collection results of target scholars are subjected to multi-dimensional scholar identity authentication, which includes research field semantic feature verification, academic relationship network verification, and spatiotemporal consistency verification. The event extraction and verification of the academic activity extraction model is carried out using the results of multi-dimensional scholar identity authentication, and the academic activity tracking and management is carried out based on the results of the event extraction and verification. The data collection results of the target scholars will be used for multi-dimensional scholar identity verification, including: Perform semantic feature verification in the research domain, including: Keyword and entity extraction is performed on the data collection results of the target scholars to establish a domain-related feature set, which includes core research topics, keywords, execution methods, and technical terms. Obtain the research interest description of the target scholar, convert the research interest description into a high-dimensional vector, and then calculate the semantic similarity between the high-dimensional vector and the domain-related feature set. Based on the semantic similarity, semantic feature verification of the research field is completed; The multi-dimensional scholar identity verification of the target scholar data collection results also includes: Performing academic relationship network verification includes: Establish a relation extraction feature library, and use the relation extraction feature library to extract relations from the data collection results of the target scholars to establish relation extraction results; The academic relationship network is used to compare the extracted relationships, and the academic relationship network is verified based on the comparison results. The multi-dimensional scholar identity verification of the target scholar data collection results also includes: Perform spatiotemporal consistency verification, including: Obtain historical data of the target scholar, extract data that meets a preset threshold from the historical data, and construct a timeline of the target scholar; The timeline of the target scholar is used to monitor conflicts in the data collection results of the target scholar. The conflict monitoring includes the joint detection of institutional conflicts and time conflicts. Complete the spatiotemporal consistency verification based on the conflict detection results; After calibrating and compensating the crawling frequency according to the frequency correction factor, data collection from the target scholars is performed, including: Based on the crawling bias, an event-driven dynamic tracking strategy is established. The target scholar's data is collected using the dynamic tracking strategy and the compensated crawling frequency, and a first data collection result is established. Obtain the data source characteristics for data collection, configure the data source priority according to the data source characteristics, and establish a first collection constraint according to the data source priority; A relationship triggering strategy is established based on the academic relationship network of the target scholar, and a second collection constraint is established using the relationship triggering strategy; Based on the first collection constraint, the second collection constraint, and the compensated crawling frequency, data collection of the target scholars is performed, and a second data collection result is established. The target scholar data collection results are constructed based on the first data collection results and the second data collection results.
2. The method of claim 1, wherein the method comprises: The event extraction validation of the academic activity extraction model utilizes multi-dimensional scholar identity authentication results, including: Construct a general extraction pattern, which includes key attributes of the event type, including report title, organizer, inviter, and collaborators; Using the large language model in the academic activity extraction model, we perform event extraction based on a general extraction pattern to extract multi-dimensional scholar identity authentication results and establish event extraction results.
3. The method of claim 2, wherein the method further comprises: determining a publication date of the publication; and determining a publication date of the publication in the publication database. The event extraction results also include: The event extraction results are subjected to event deduplication and validation processing, including: The rule layer of event extraction results is deduplicated using unique fingerprints and key attributes. The large language model is used to vectorize the academic report titles, academic names, and event times in the event extraction results. The similarity of the vectorized results is calculated, and semantic deduplication is performed. The event extraction results are analyzed using a machine learning model, and deduplication is performed at the model layer. Event extraction and validation are completed through rule-based deduplication, semantic-based deduplication, and model-based deduplication.
4. The method of claim 3, wherein the method further comprises: Event extraction and validation are completed based on rule-level deduplication, semantic-level deduplication, and model-level deduplication. This also includes: Perform cross-validation on the event extraction results after deduplication, including internal consistency verification, external consistency verification, and source credibility verification; Based on the cross-validation results, the event extraction results are filtered to establish event extraction validation results.
5. The scholar academic activity automatic and precise tracking method based on a large language model according to claim 1, wherein, The scholar feature set with source reliability identifiers is input into the frequency dynamic compensation network, which outputs frequency correction factors, including: The frequency dynamic compensation network is used to perform feature value normalization transformation of the scholar feature set; Based on the normalization transformation results, the frequency correction factor is calculated through the computation layer as follows: ; wherein, a frequency correction factor, a total number of features of the scholar feature set, an arbitrary feature of the scholar feature set, a weight of the th feature, a normalized feature value of the th feature.
6. The automatic and accurate tracking system for scholars' academic activities based on large language models, characterized in that, For implementing the method for automatically and accurately tracking the academic activities of scholars based on a large language model as described in any one of claims 1-5, the system comprises: Crawling frequency construction module: Constructs the user's calibrated crawling frequency, which is the initially set crawling plan; Feature extraction module: Performs feature extraction on the target scholar and establishes a scholar feature set, which includes activity features, timeliness features, academic contribution features, event-driven features, and relationship network features. The scholar feature set is set with a source reliability identifier. Data acquisition module: Inputs the scholar feature set with source reliability identifier into the frequency dynamic compensation network, outputs frequency correction factor, the frequency correction factor is set with crawling bias, and after calibrating the crawling frequency compensation according to the frequency correction factor, performs data acquisition of the target scholar; Identity authentication module: Performs multi-dimensional scholar identity authentication on the target scholar data collection results. The multi-dimensional scholar identity authentication includes research field semantic feature verification, academic relationship network verification, and spatiotemporal consistency verification. Academic activity tracking module: use multi-dimensional scholar identity authentication results to perform event extraction verification on the academic activity extraction model, and perform academic activity tracking management according to the event extraction verification results.
Citation Information
Patent Citations
Name disambiguation method orienting Chinese authors in English literature
CN106294677A
Adaptive perception event element extraction method
CN120337901A