Agent-based university homepage information transaction monitoring statistical method and system
By adopting an agent-based method for monitoring and statistically analyzing changes in university homepage information, the problem of incomplete handling of changes in university homepage information was solved. This method achieves automated and accurate identification of changes and timely data feedback, provides structured change reports, and improves data accuracy and management efficiency.
Patent Information
- Application Number
- CN202511521422.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies lack comprehensive and automated systems to process and analyze changes in information on university homepages, making it difficult to guarantee data accuracy and timeliness.
An agent-based method for monitoring and statistically analyzing changes in university homepage information is adopted. This method involves crawling university homepage information, comparing text content, performing LLM analysis and classification, generating change reports, and storing them in a database, thereby achieving automated processing and accurate identification of page anomalies.
It enables comprehensive and automated processing of information changes on university homepages, improving the accuracy and timeliness of identification and classification, and providing structured change information to facilitate decision-making by university administrators.
Smart Images

Figure CN120976910A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to an Agent-based university homepage information change monitoring and statistics method and system. BACKGROUND
[0002] With the continuous development of higher education and the annual update of degree authorization points, the organizational structure, discipline setting, teacher position and society leadership information of universities are also changing. The changes of these information are usually published in real time through the news or announcements of university homepage. In order to obtain and monitor these changes in time, especially in all universities, manual searching and recording of changes not only take time and effort, but also are easy to miss.
[0003] At present, most of the university information update monitoring systems only rely on manual or simple timing crawler to capture the content of the webpage, and there is no comprehensive and automatic system to process and analyze the updated content on the university homepage. In addition, the existing system is also difficult to accurately identify and classify the change content in the page, which leads to the difficulty in guaranteeing the accuracy and timeliness of the data. SUMMARY
[0004] In view of the above deficiencies of the prior art, the purpose of the embodiments of the present application is to provide an Agent-based university homepage information change monitoring and statistics method, which can solve the technical problems that there is no comprehensive and automatic system to process and analyze the updated content on the university homepage, and the existing system is also difficult to accurately identify and classify the change content in the page, leading to the difficulty in guaranteeing the accuracy and timeliness of the data.
[0005] The first aspect of the embodiments of the present application proposes an Agent-based university homepage information change monitoring and statistics method, comprising: S1: capturing university homepage information; S2: comparing the text content of the university homepage information with the last captured university homepage information to determine the change text content; S3: performing preliminary analysis on the change text content through LLM to identify the category of the change text content; S4: selecting a change processing mode corresponding to the category to process the change text content to form change data; S5: aggregating the change data and generating a change report; S6: storing the change data to a database.
[0006] The second aspect of the embodiments of the present application proposes an Agent-based university homepage information change monitoring and statistics system, comprising: a processor and a memory. The memory stores programs or instructions which can be run on the processor, and the programs or instructions are executed by the processor to implement the steps of the Agent-based university homepage information change monitoring and statistics method according to the first aspect.
[0007] According to a third aspect of the embodiments of the present application, a readable storage medium is provided, and the readable storage medium stores programs or instructions, and the programs or instructions are executed by a processor to implement the steps of the Agent-based university homepage information change monitoring and statistics method according to the first aspect.
[0008] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects: In the embodiments of the present application, the updated content on the university homepage can be comprehensively and automatically processed and analyzed, the changed content in the page can be accurately identified and classified, and the accuracy and timeliness of the university homepage information change identification can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0009] The accompanying drawings are only for the purpose of illustrating the specific embodiments, and are not considered as limiting the present application, and in the whole drawings, the same reference signs represent the same components. Obviously, the accompanying drawings in the following description are only some embodiments described in the embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the drawings.
[0010] Figure 1 is a flow diagram of an Agent-based university homepage information change monitoring and statistics method provided by the embodiments of the present application.
[0011] Figure 2 is a structural diagram of an Agent-based university homepage information change monitoring and statistics system provided by the embodiments of the present application. DETAILED DESCRIPTION
[0012] In order to make the person skilled in the art better understand the technical scheme in the embodiments of the present application, the technical scheme of the present application will be described clearly and completely in combination with the drawings, and obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. It should be understood that these descriptions are only exemplary, and are not used to limit the scope of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.
[0013] The Agent-based university homepage information change monitoring and statistics method provided by the embodiments of the present application will be described in detail in combination with the drawings and specific embodiments and application scenarios.
[0014] Reference is made to the drawingsFigure 1 The diagram illustrates a flowchart of an agent-based method for monitoring and statistically analyzing changes in university homepage information, provided by an embodiment of the present invention.
[0015] This invention provides an agent-based method for monitoring and statistically analyzing changes in university homepage information, which may include the following steps: S1: Retrieve information from the homepages of higher education institutions.
[0016] In one possible implementation, S1 specifically involves periodically using web crawling technology to retrieve news announcements and page content that publishes information about universities from the official websites of universities across the country.
[0017] Furthermore, for each university homepage, the content of each crawled webpage is compared with the data from the previous crawl to identify any additions, deletions, or modifications to the information.
[0018] S2: Compare the text content of the information on the homepage of higher education institutions with the information on the homepage of higher education institutions that was crawled last time, and identify the text content that has changed.
[0019] Specifically, for each university homepage, the content of each crawled webpage is compared with the data from the previous crawl to identify any additions, deletions, or modifications to the information.
[0020] In one possible implementation, S2 specifically includes sub-steps S201 to S205: S201: Extract text data from the homepage information of higher education institutions.
[0021] S202: Remove HTML tags, scripts, and advertisements from the text data, retaining only the main text content. Preprocessing ensures consistent data formatting and eliminates interference.
[0022] S203: Segment the main text content and generate a unique hash identifier for each segment.
[0023] S204: Perform fast matching based on hash identifiers to locate candidate difference paragraphs.
[0024] Furthermore, subsequent comparisons are performed only in candidate paragraphs, avoiding the need to run edit distances across the entire webpage, thus improving efficiency.
[0025] S205: Using a text difference analysis algorithm based on Levenshtein edit distance, the text content of candidate difference paragraphs is compared to identify the abnormal text content, which includes added, deleted or modified content.
[0026] Furthermore, during the calculation process, the type of each operation (insertion, deletion, replacement) is recorded, and the context position is marked, so that "new content", "deleted content" and "modified content" can be directly output.
[0027] Optionally, a weighting mechanism can be introduced when calculating edit distance. Keywords (such as "professor" or "new subject") are given higher weights, making changes to key information such as subject settings and teacher positions easier to detect; while the impact of changes to ordinary stop words is reduced.
[0028] In this embodiment of the invention, by using step-by-step processing and rapid matching of hash identifiers, it is possible to efficiently locate and compare potentially changed paragraphs on a webpage, avoiding tedious edit distance calculations for the entire webpage, thereby significantly improving processing speed. Simultaneously, removing irrelevant content ensures data format consistency and accuracy, while Levenshtein-based edit distance difference analysis accurately identifies newly added, deleted, or modified content, further improving detection accuracy and efficiency, and optimizing the ability to process large-scale data.
[0029] In one possible implementation, after S205, S2 further includes: S206: Calculate the semantic similarity of text content using a BERT-based deep semantic model.
[0030] S207: Determine whether the semantic similarity is greater than the similarity threshold. If yes, determine whether the corresponding text content is specifically modified. Otherwise, determine whether the corresponding text content is specifically added or deleted.
[0031] Furthermore, during the calculation process, the type of each operation (insertion, deletion, replacement) is recorded, and the context position is marked, so that "new content", "deleted content", and "modified content" can be directly output.
[0032] Optionally, the similarity threshold can be dynamically adjusted based on text length and paragraph distribution, with a higher similarity threshold for short texts than for long texts. This ensures detection accuracy for texts of different sizes.
[0033] In this embodiment of the invention, by calculating the semantic similarity of text content using a BERT-based deep semantic model, semantic changes between texts can be determined more accurately, surpassing traditional literal comparison. When the semantic similarity is high, the system can accurately identify modified content, avoiding misjudgments as additions or deletions, thus improving detection accuracy. When the similarity is low, the system can accurately determine whether content has been added or deleted, thereby ensuring the comprehensiveness and accuracy of anomaly detection. Especially when processing texts containing similar but different expressions, it greatly improves the ability to recognize subtle semantic changes.
[0034] S3: Using LLM, perform preliminary analysis of the abnormal text content to identify the category of the abnormal text content.
[0035] Optionally, the categories of the change text content include: subject change category, teacher position change category, organizational structure change category, and academic society leader change category.
[0036] In this embodiment of the invention, LLM (Limited Language Management) is used to perform preliminary analysis of the anomaly text content and identify its category, enabling automated and intelligent classification processing. This method can accurately categorize anomaly content into specific categories, such as subject, teacher position, organizational structure, or academic society leader, thereby ensuring that each type of information receives targeted processing and analysis, improving the system's efficiency and accuracy. Furthermore, automatic classification reduces the need for manual intervention, ensuring rapid response and efficient management during large-scale data processing.
[0037] It's important to note that this invention does not allow the LLM to directly extract text. Instead, it first determines whether the differing text contains specified content. This is because executing the extraction all at once would place significant resource demands on both the token and the LLM's execution time. Furthermore, the uncertainty surrounding the input text's presence of the required content would lead to decreased extraction accuracy. In particular, we found that when using a Reasoning model for extraction, if the input does not contain the specified content, the model may fall into a self-doubting loop, unable to stop the "thinking" process. Step-by-step processing not only filters out inputs that do not contain the specified text, allowing for early stopping, but also allows for targeted prompt design, improving output quality.
[0038] S4: Select the anomaly handling method corresponding to the category, process the anomaly text content, and form anomaly data.
[0039] In one possible implementation, S4 specifically includes sub-steps S401 to S404: S401: For text changes related to subject changes, check if any subjects have been added or removed from the college introduction. If so, standardize the changed subject by comparing it with the national first-level subject directory and store it in a JSON structure. Otherwise, stop the process directly.
[0040] S402: For the text content related to teacher position changes, determine if there are names in the non-common substrings. If so, extract the names and positions of the added or removed individuals from the change content using the large model and store them in a JSON structure. Otherwise, stop the process directly.
[0041] S403: For text content related to organizational structure changes, determine if there is an organization in the non-common substring. If so, extract the newly added or removed organizations from the change content using the large model and store them in a JSON structure. Otherwise, stop the process directly.
[0042] S404: For the text content related to changes in the association's leadership position, determine if a name is present in the non-common substring. If so, extract the name and position from the change content using a large model and store them in a JSON structure. Otherwise, stop the process immediately.
[0043] In this embodiment of the invention, by setting specific processing methods for different anomaly categories (such as disciplines, teacher positions, organizational structures, and academic society leaders), the system can perform precise judgments and processing based on the specific anomaly content. This classification processing method ensures that each type of anomaly can be specifically identified and stored in a standardized manner, thereby improving processing efficiency and accuracy. Using a large model to extract specific information about additions or deletions and storing it in a JSON structure helps ensure the structured nature, traceability, and convenience of subsequent analysis of the data. In addition, this method effectively avoids interference from irrelevant content, improving the system's flexibility and intelligence level when processing large-scale data.
[0044] S5: Summarize the abnormal data and generate an abnormal report.
[0045] Optionally, the change report includes specific details of the change, such as newly added or deleted disciplines, personnel position changes, and changes in the person in charge of the academic society.
[0046] In this embodiment of the invention, by summarizing change data and generating detailed change reports, the system can provide users with clear and structured change information, including newly added or deleted disciplines, personnel position changes, and changes in academic society leaders. This method not only improves the readability and operability of the data but also ensures the comprehensiveness and accuracy of the reports, providing real-time and intuitive data support for university administrators and decision-makers, facilitating timely decision-making and adjustments. Simultaneously, the structured reports facilitate subsequent archiving, retrieval, and analysis, enhancing the long-term value of the data.
[0047] S6: Stores abnormal data in the database. This ensures long-term information preservation and subsequent retrieval.
[0048] Furthermore, the system of this invention supports a scheduled automatic execution function, periodically crawling and comparing the content of university homepages to generate updated anomaly reports. The crawling frequency and task execution time can be configured according to actual needs.
[0049] In this embodiment of the invention, by storing anomaly data in a database, the system ensures long-term preservation and convenient retrieval of information, guaranteeing data traceability and the integrity of historical records. The scheduled automatic execution function enables the system to periodically crawl and compare content from university homepages, generating updated anomaly reports to ensure real-time information updates and timely feedback. Configurable crawling frequency and task execution time allow the system to flexibly adapt to different needs, optimize resource utilization, improve the system's automation level and ability to respond to changes, and further enhance the efficiency and accuracy of monitoring.
[0050] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, the updated content on university homepages can be processed and analyzed comprehensively and automatically, and the abnormal content on the pages can be accurately identified and classified, thereby improving the accuracy and timeliness of identifying abnormal information on university homepages.
[0051] Reference manual attached Figure 2 The diagram shows a schematic of the structure of a university homepage information anomaly monitoring and statistics system based on Agent provided in an embodiment of the present invention.
[0052] This invention provides an Agent-based monitoring and statistical system 20 for monitoring changes in university homepage information, comprising: a processor 201 and a memory 202; The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the steps of the above-described agent-based method for monitoring and statistically analyzing changes in university homepage information, and achieve the same technical effect. To avoid repetition, this invention will not elaborate further.
[0053] It should be understood that the processor 201 in this embodiment of the invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0054] It should also be understood that the memory 202 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0055] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0056] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0057] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0058] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0059] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0060] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0061] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0062] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0063] This invention provides a readable storage medium comprising: storing a program or instructions on the readable storage medium, wherein when the program or instructions are executed by a processor, the program or instructions implement the steps of the above-described agent-based method for monitoring and statistically analyzing changes in university homepage information, and achieve the same technical effect. To avoid repetition, this invention will not elaborate further.
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the protection scope of the present invention.
Claims
1. A method for monitoring and statistically analyzing changes in university homepage information based on agent, characterized in that, include: S1: Retrieves information from the homepages of higher education institutions; S2: Compare the text content of the information on the homepage of the higher education institution with the information on the homepage of the higher education institution that was previously crawled, and determine the abnormal text content; S3: Using LLM, perform a preliminary analysis of the abnormal text content to identify the category of the abnormal text content; S4: Select the anomaly processing method corresponding to the category, process the anomaly text content, and form anomaly data; S5: Summarize the abnormal data and generate an abnormality report; S6: Store the abnormal data in the database.
2. The agent-based method for monitoring and statistically analyzing changes in university homepage information according to claim 1, characterized in that, Specifically, S1 is: We regularly use web scraping technology to extract news announcements and page content that publishes information about universities from the official websites of universities across the country.
3. The agent-based method for monitoring and statistically analyzing changes in university homepage information according to claim 1, characterized in that, S2 specifically includes: S201: Extract text data from the homepage information of the higher education institutions; S202: Remove HTML tags, scripts, and advertisements from the text data, retaining the main text content; S203: The main text content is segmented, and a unique hash identifier is generated for each segment; S204: Perform fast matching based on the hash identifier to locate candidate difference segments; S205: Using a text difference analysis algorithm based on Levenshtein edit distance, the candidate difference paragraphs are compared to identify the abnormal text content, which includes added, deleted, or modified content.
4. The agent-based method for monitoring and statistically analyzing changes in university homepage information according to claim 3, characterized in that, Following S205, S2 further includes: S206: Calculate the semantic similarity of text content using a BERT-based deep semantic model; S207: Determine whether the semantic similarity is greater than the similarity threshold; if yes, determine that the corresponding text content is specifically modified content; otherwise, determine that the corresponding text content is specifically added content or deleted content.
5. The agent-based method for monitoring and statistically analyzing changes in university homepage information according to claim 4, characterized in that, The similarity threshold can be dynamically adjusted according to the text length and paragraph distribution, and the similarity threshold for short texts is greater than that for long texts.
6. The agent-based method for monitoring and statistically analyzing changes in university homepage information according to claim 1, characterized in that, The categories of the changed text content include: subject-specific changes, teacher position changes, organizational structure changes, and changes of academic society leaders.
7. The agent-based method for monitoring and statistically analyzing changes in university homepage information according to claim 6, characterized in that, S4 specifically includes: S401: For the text content of the subject change category, determine whether there are any newly added or removed subjects in the college introduction; if so, compare the changed subject with the national first-level subject directory for standardization and store it in JSON structure; otherwise, stop the process directly. S402: For the change text content of the teacher position change category, determine whether there are names in the non-common substring; if so, extract the added or removed names and positions in the change content through the large model and store them in JSON structure; otherwise, stop the process directly. S403: For the change text content of the organizational structure change category, determine whether there is an organization in the non-common substring; if so, extract the newly added or removed organizations in the change content through the large model and store them in JSON structure; otherwise, stop the process directly. S404: For the change text content of the aforementioned association leader change category, determine whether there is a person's name in the non-common substring; if so, extract the person's name and position from the change content through the large model and store them in JSON structure; otherwise, stop the process directly.
8. The agent-based method for monitoring and statistically analyzing changes in university homepage information according to claim 1, characterized in that, The change report includes specific details of the change, such as the addition or deletion of disciplines, personnel positions, and changes in the heads of academic societies.
9. A user-based monitoring and statistical system for changes in information on university homepages, characterized in that: include: Processor and memory; The memory stores programs or instructions that can run on the processor, and when the program or instructions are executed by the processor, they implement the steps of the agent-based method for monitoring and statistically analyzing changes in university homepage information as described in any one of claims 1 to 8.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions, which, when executed by a processor, implement the steps of the agent-based method for monitoring and statistically analyzing changes in university homepage information as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method and system for detecting hostile attack on Internet information system
CN103281177A
Program ABI interface compatibility calculation method based on Linux system
CN114510267A
Judgment method for intelligently detecting substantive change of webpage content based on text embedding
CN117786270A
Text processing method and device, equipment, storage medium and program product
CN118447521A
Digital information traceability management system
CN120145334A