An interdisciplinary information interaction method, device, medium and program product
By acquiring academic activities and registration data from target databases within universities, determining subject attribute sets, and making precise recommendations, the problem of low efficiency in cross-disciplinary information exchange has been solved, and efficient dissemination of cross-disciplinary information has been achieved.
Patent Information
- Application Number
- CN202410921735.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-10
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-07-10
AI Technical Summary
The fragmented release of academic activity information across different disciplines within universities forces scholars and students to spend a significant amount of time browsing multiple platforms, resulting in incomplete information and low efficiency in interdisciplinary information exchange.
By acquiring academic activity data and registration datasets from different disciplines through the target database, the subject attribute sets of users and projects are determined, and the precise push of cross-disciplinary activities is achieved using preset push methods.
It has improved the efficiency of users obtaining information about activities of interest, promoted information exchange between different disciplines, and enhanced the efficiency of interdisciplinary information exchange.
Smart Images

Figure CN118964722B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information interaction technology, and in particular to an interdisciplinary information interaction method, device, medium and program product. Background Technology
[0002] With the rapid development of information technology and the deepening of educational informatization, the interaction between different disciplines within universities has become increasingly important. Interdisciplinary exchange and cooperation not only help improve students' comprehensive qualities and innovative abilities but also promote the overall improvement of the university's teaching level and research strength.
[0003] In related technologies, various colleges within universities typically employ multiple information dissemination channels to disseminate information about academic activities across their disciplines. These channels include, but are not limited to, the university's official website, academic seminars, internal college websites, college bulletin boards, and academic journals. Specifically, each college selects appropriate channels to promote relevant academic activities based on its own professional characteristics and research directions. For example, an engineering college might publish information about lectures on the latest technological developments on the university's official website and also promote them through the college bulletin board; a humanities college might publish details of literature seminars through its internal college website and academic journals, thereby disseminating information about various academic activities.
[0004] However, by adopting the above approach, different colleges within a university disseminate information about academic activities of various disciplines through different information dissemination channels. This decentralized dissemination not only requires scholars and students to spend a lot of time browsing multiple information dissemination platforms, but may also result in the collection of incomplete or outdated information about academic activities of various disciplines. Moreover, this fragmented collection method may not be able to quickly and comprehensively understand the academic activities of various disciplines, thus leading to low efficiency in interdisciplinary information exchange in related technologies. Summary of the Invention
[0005] This application provides a method, device, medium, and program product for interdisciplinary information exchange, which can improve the efficiency of interdisciplinary information exchange.
[0006] In a first aspect, this application provides a cross-disciplinary information interaction method applied to the aforementioned electronic device. The method includes: obtaining N first activity datasets from a target database, and obtaining M registration datasets from the target database, wherein the N first activity datasets include key information extracted from academic activities of different disciplines, and the M registration datasets are sets of registration records generated when M target users successfully register for Q target publishing projects, where N, M, and Q are all positive integers greater than or equal to 1; determining a target subject attribute set based on the M registration datasets, and determining a target invited subject attribute set based on the Q target publishing projects, wherein the target subject attribute set is a set of subject areas pre-set by the M target users, and the target invited subject attribute set is a set of invited subject areas set by P project publishers when publishing Q target publishing projects, where P is a positive integer greater than or equal to 1; determining N target activity projects based on the N first activity datasets, and performing target push operations on the N target activity projects and the Q target publishing projects based on the target subject attribute set, the target invited subject attribute set, and a preset push method.
[0007] By adopting the above technical solution, N first activity datasets corresponding to academic activities in different disciplines are obtained through the target database, and M registration datasets corresponding to registration records are obtained through the target database. This allows for comprehensive acquisition of cross-disciplinary activity information and users' subject preferences. Furthermore, the target subject attribute set for users is determined through the M registration datasets, and the target invited subject attribute set is determined through Q target published projects, thus matching user preferences with the invited subjects of project publishers. Subsequently, N target activity projects are determined based on the first activity datasets, and a target push operation is performed on the N target activity projects and Q target published projects by combining the target subject attribute set, the invited subject attribute set, and a preset push method. This achieves precise push notifications of cross-disciplinary academic activities, improves the efficiency of users obtaining information about activities of interest, promotes information interaction between different disciplines, solves the technical problem of low efficiency in cross-disciplinary information interaction in related technologies, and achieves the technical effect of improving the efficiency of cross-disciplinary information interaction.
[0008] In a second aspect, embodiments of this application provide an electronic device comprising: one or more processors and a memory; the memory is coupled to one or more processors and is used to store computer program code, the computer program code including computer instructions, wherein one or more processors invoke the computer instructions to cause the electronic device to perform the method described in the first aspect and any possible implementation thereof.
[0009] Thirdly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on an electronic device, cause the electronic device to perform the method described in the first aspect and any possible implementation thereof.
[0010] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect and any possible implementation thereof. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating an interdisciplinary information interaction method in an embodiment of this application;
[0012] Figure 2 This is a schematic diagram of a two-way selection process in an embodiment of this application;
[0013] Figure 3 This is a schematic diagram of a message push process in an embodiment of this application;
[0014] Figure 4 This is an entity relationship diagram of users, disciplines, and projects in the embodiments of this application;
[0015] Figure 5 This is a schematic diagram of the architecture of an interdisciplinary collaborative system as described in this application.
[0016] Figure 6 This is a schematic diagram of the physical device structure of an electronic device in an embodiment of this application. Detailed Implementation
[0017] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.
[0018] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0019] This application provides a cross-disciplinary information exchange method, see [link / reference] Figure 1 , Figure 1 This is a flowchart illustrating an interdisciplinary information interaction method in an embodiment of this application, including the following steps:
[0020] Step S101: Obtain N first activity datasets from the target database, and obtain M registration datasets from the target database. The N first activity datasets include key information extracted from academic activities in different disciplines, and the M registration datasets are sets of registration records generated when M target users successfully register for Q target published projects. N, M, and Q are all positive integers greater than or equal to 1.
[0021] The target database is used to store information such as activity datasets and registration datasets. The first activity dataset refers to a collection of key information extracted from academic activities across various disciplines, including but not limited to activity titles, release dates, event dates, speakers, and locations. The registration dataset is a summary of registration records generated after target users successfully register for target-published projects. Target users refer to users who register and use the system (i.e., the interdisciplinary collaboration system). Target-published projects refer to projects published within the system by the project publisher for target users to register for. N, M, and Q are all positive integers, representing the number of the first activity dataset, the number of registration datasets, and the number of target-published projects, respectively.
[0022] Specifically, during this step, a certain amount of activity information and user registration data has been collected. At this point, it is necessary to extract the first activity dataset and registration dataset in batches from the target database to prepare for subsequent data analysis and push notifications. The target database is queried to obtain academic activity information captured from various disciplines (i.e., N first activity datasets). Data fields include, but are not limited to, activity title, activity release time, activity event time, activity speaker, and activity location. The target database is then queried to obtain M registration datasets. Each registration dataset corresponds to one user registration record. M registration datasets indicate that a total of M users participated in the registration. The data fields of the registration records include, but are not limited to, user ID, registered project ID, registration time, user identity category, user contact information, project publisher identity category, and project publisher contact information. It is important to note that the values of M, N, and Q are all positive integers greater than or equal to 1. Of course, the values of M, N, and Q can all be equal to 1 simultaneously. For example, when M, N, and Q are all 1, it indicates that there is at least one activity dataset, one registration dataset, and one target published project. The values of M, N, and Q do not necessarily have to be equal to 1 simultaneously; that is, any one of M, N, or Q can be equal to 1. For example, when M is 1, N can be 8, Q can be 15, etc., without limitation here. The values of M, N, and Q can also all be greater than 1 simultaneously; for example, M can be 5, N can be 9, Q can be 20, etc., without limitation here. It should also be noted that the order in which N first activity datasets and M registration datasets are obtained from the target database is not limited; alternatively, the N first activity datasets can be obtained first, followed by the M registration datasets, or vice versa.
[0023] In some embodiments, the first activity dataset and the registration dataset can be obtained in various ways. For example, an idle connection can be obtained by calling a database connection pool, and then a Structured Query Language (SQL) statement can be sent through that connection to retrieve all the required data from the data table at once. The dataset can then be processed, and finally the connection can be released. Alternatively, a paginated query approach can be used, sending multiple query requests, returning only one page of data each time, and iterating through the data until all data is retrieved. This avoids memory overflow caused by loading too much data at once. It is understood that other methods can also be used, such as stored procedures, Object-Relational Mapping (ORM) frameworks, etc., to obtain the target dataset; these are not limited here.
[0024] Step S102: Determine the target subject attribute set based on the M registration datasets, and determine the target invited subject attribute set based on the Q target published projects. The target subject attribute set is a set of subject fields pre-set by the M target users, and the target invited subject attribute set is a set of invited subject fields set by the P project publishers when publishing the Q target published projects, where P is a positive integer greater than or equal to 1.
[0025] The target subject attribute set is the set of subject areas set by M target users when registering or using the system, such as physics, chemistry, and mathematics. The target invited subject attribute set is the set of invited subject areas selected by the project publishers of Q target projects when publishing their projects. Invited subject areas refer to the target projects being more biased towards or targeted users with certain subject backgrounds. For example, a target project may be biased towards inviting target users in the fields of computer science and electronics.
[0026] Specifically, this step is performed after obtaining the registration dataset and target publishing projects to summarize the subject areas of interest for both the user side and the project publishing side, preparing for subsequent precise push matching. It iterates through M registration datasets, extracts the pre-set subject attribute field values for each registered user, removes duplicates, and summarizes them to obtain the target subject attribute set, which indicates the subject area distribution of the participating user group. It then iterates through Q target publishing projects, extracts the invited subject attribute field values for each target publishing project, removes duplicates, and summarizes them to obtain the target invited subject attribute set, which indicates which subject backgrounds the Q projects prefer to invite.
[0027] In some embodiments, there are multiple ways to determine the target subject attribute set and the target invited subject attribute set. For example, if the registration dataset already contains user subject attribute fields, the target subject attribute set can be obtained by directly traversing the M registration datasets, extracting the field values corresponding to the user subject attributes, and removing duplicates. If the registration dataset does not contain the field, the field values corresponding to the user subject attributes need to be obtained by querying the user information table using the user ID in the registration record. Alternatively, the field values and frequencies corresponding to the user subject attributes of the registered users can be preliminarily counted. After filtering out the field values with low frequency, the subject attributes corresponding to the field values with high frequency can be grouped into the target subject attribute set. This can focus on the main distribution of user subject attributes and eliminate the interference of sporadic user subject attributes. Similar to the method of determining the target subject attribute set, the difference lies in traversing the Q target publishing projects to extract the invited subject attribute fields of the target publishing projects. For example, the invited subject attribute values of the Q projects can also be counted and filtered by frequency to obtain a more representative target invited subject attribute set.
[0028] Step S103: Determine N target activity projects based on N first activity datasets, and perform target push operations on the N target activity projects and Q target release projects according to the target subject attribute set, the target invited subject attribute set, and the preset push method.
[0029] The target activities are the set of activities within the target subject area selected through analysis of the first activity dataset. The preset push method refers to pre-defined push rules to deliver activities to appropriate target users. The target push operation is the process of pushing N target activities and Q target published items to M target users based on certain push rules.
[0030] Specifically, this step begins after the target subject attribute set and the target invited subject attribute set are determined. Utilizing the acquired multi-dimensional information, the activity projects are matched, filtered, and personalized for delivery. Analyzing N datasets of the first activity, N target activity projects are identified. Taking into account the target subject attribute set, the target invited subject attribute set, and preset delivery methods, targeted delivery operations are performed on the N target activity projects and Q target published projects. Preset delivery methods include, but are not limited to, user response feedback, interaction frequency, and geographical location. Delivery results can also be persistently stored and pushed to target users at appropriate times. Targeted delivery operations can flexibly employ various personalized recommendation algorithms; the key is to fully utilize multi-dimensional information, respect user preferences, and improve the accuracy and conversion rate of the delivery.
[0031] Through the above steps, N first activity datasets corresponding to academic activities in different disciplines are obtained from the target database, along with M registration datasets corresponding to registration records. This allows for comprehensive acquisition of cross-disciplinary activity information and user subject preferences. Furthermore, the target subject attribute set for users is determined using the M registration datasets, and the target invited subject attribute set is determined using Q target published projects, thus matching user preferences with the invited subjects of project publishers. Subsequently, N target activity projects are determined based on the first activity datasets, and a target push operation is performed on the N target activity projects and Q target published projects, combining the target subject attribute set, the invited subject attribute set, and a preset push method. This achieves precise push notifications of cross-disciplinary academic activities, improves the efficiency of users obtaining information about activities of interest, promotes information interaction across different disciplines, solves the technical problem of low efficiency in cross-disciplinary information interaction in related technologies, and achieves the technical effect of improving the efficiency of cross-disciplinary information interaction.
[0032] The entity performing the above steps may be a system or platform with interdisciplinary information interaction capabilities, such as an interdisciplinary collaboration system or platform, or a device with interdisciplinary information interaction capabilities, or a controller or processor in a device or system, or a standalone controller or processor, or other processing devices or processing units with similar processing capabilities.
[0033] In an optional embodiment, before obtaining N first activity datasets from the target database, the method further includes: performing webpage structure analysis on the academic publishing website to set N target crawling configurations for the academic publishing website, wherein the academic publishing website includes N academic activity webpages, the N academic activity webpages include N academic activity-related information from different disciplines, and the N target crawling configurations correspond one-to-one with the N academic activity webpages; sequentially traversing each academic activity webpage included in the academic publishing website according to the N target crawling configurations, and performing the following operation on the currently traversed academic activity webpage: determining a first configuration item from the first crawling configuration, wherein the first configuration item is used to determine whether to crawl the activity details page data of the first academic activity webpage, the first academic activity webpage... For the currently traversed academic activity webpage, the first crawl configuration is the data crawling configuration corresponding to the first academic activity webpage, and N target crawl configurations include the first crawl configuration; if the first configuration item is yes, crawl the activity details page data and activity list page data according to the first crawl configuration to obtain the first dataset, wherein the first academic activity webpage includes the activity list page data; or, if the first configuration item is no, crawl the activity list page data according to the first crawl configuration to obtain the second dataset; if N second activity datasets are obtained, perform data optimization operations on the N second activity datasets, wherein the N second activity datasets include either the first dataset or the second dataset, and the N second activity datasets correspond one-to-one with the N academic activity webpages.
[0034] In the above embodiment, assume an academic website contains academic activity webpages for 10 different disciplines (e.g., physics, chemistry, mathematics, etc.; of course, it could also be 9, 15, or 20 different disciplines, etc., without limitation here). By analyzing the HTML (HyperText Markup Language) structure of the webpages, a crawling configuration is generated for each academic activity webpage. The configuration includes the webpage URL (Uniform Resource Locator), data crawling rules, data storage format, etc. Then, these 10 crawling configurations are loaded, and each academic activity webpage is accessed sequentially. For the currently accessed academic activity webpage, its corresponding crawling configuration is read, and the configuration item (i.e., the first configuration item) determines whether it is necessary to crawl data from both the activity list page and the activity details page simultaneously. If so, the data from the activity list page and the activity details page are crawled to form the first dataset; if not, only the activity list page data is crawled to form the second dataset. Ultimately, 10 second activity datasets were obtained, each corresponding to one of the 10 academic activity web pages. Data cleaning, deduplication, and other optimization operations were then performed on these 10 second activity datasets to generate 10 first activity datasets with standardized structure and clean content.
[0035] By analyzing the webpage structure of academic publishing websites and setting corresponding target crawling configurations for academic activity webpages in different disciplines, targeted data crawling can be achieved. By traversing each academic activity webpage using the target crawling configuration, the first configuration item determines whether to crawl data from the activity details page, flexibly controlling the granularity of data crawling. Performing data optimization operations on the acquired second activity dataset improves data quality and usability. This phased and granular data acquisition and optimization method can efficiently and accurately obtain cross-disciplinary academic activity information, providing a high-quality data foundation for subsequent information exchange.
[0036] In an optional embodiment, when N second activity data sets are obtained, data optimization operations are performed on the N second activity data sets, which specifically include: optimizing the activity data in the N second activity data sets that meet the preset optimization conditions to obtain N first activity data sets; determining the activity detail page locators of N academic activity web pages from the N first activity data sets to obtain N activity detail page locators; converting each activity detail page locator included in the N activity detail page locators into an MD5 hash value to obtain N MD5 hash values; using the N MD5 hash values as the IDs of the N first activity data sets to obtain N IDs, where the N first activity data sets and the N IDs correspond one by one; performing data detection on all the data in the target database according to the N IDs to obtain a data detection result; when the data detection result shows that the first ID already exists in the target database, determining the third data set corresponding to the first ID from the N first activity data sets, and determining the fourth data set corresponding to the first ID from the target database, where the N IDs include the first ID; performing a storage optimization operation on the N first activity data sets according to the third data set and the fourth data set, where the data optimization operation includes a storage optimization operation.
[0037] In the above embodiment, it is assumed that 500 pieces of data of academic activities are crawled from multiple academic websites to form a second activity data set. Each second activity data includes, but is not limited to, fields such as activity name, organizer, activity location, activity time, activity introduction, activity type, activity detail page URL, etc. Analyzing the second activity data set, the following several situations that need to be optimized are identified: the "activity time" field format of some academic activities is not unified, some are "YYYY-MM-DD", some are "MM / DD / YYYY", and some are "YYYY year MM month DD day", etc.; the "activity location" field of some academic activities is too brief, only writing the city name without the detailed address; there are typos or mixed case in the "organizer" field of some academic activities; the "activity detail page URL" field format of some academic activities is not standardized, with redundant characters, whitespace characters, etc.
[0038] Following the previous example, in response to the above situation, the activity data that met the preset optimization conditions were optimized as follows according to the preset optimization rules: Data with inconsistent "Activity Time" format was uniformly converted to the standard format "YYYY-MM-DDHH:MM", for example, "June 18, 2023" was converted to "2023-06-18 00:00"; Data with overly brief "Activity Location" field was obtained by calling the geocoding API (Application Programming Interface) based on the city name and activity name, for example, "Meeting Room 101, Building A, No. 1, Street, District, City"; Data with typos or mixed capitalization in the "Organizer" field was corrected and standardized by using a string similarity algorithm to match common organizer names, for example, "PekingUniversity" was standardized to "Peking". For data with abnormal formatting in the "Event Details Page URL" field, regular expressions were used to identify the valid parts of the URL, removing redundant or whitespace characters to restore it to a valid URL format. Through the above geocoding, string matching, and natural language processing optimization steps, the original second event dataset was transformed into a first event dataset with more complete content and a more standardized format, achieving a more comprehensive data optimization effect.
[0039] Through the above steps, by optimizing the second activity dataset, a higher-quality first activity dataset is obtained. The activity detail page locator of each of the N academic activity web pages is extracted and converted into an MD5 hash value, serving as a unique ID for each of the N first activity datasets, ensuring the security of the N first activity datasets. Then, data detection is performed on the target database based on the N IDs to efficiently determine the duplication of the N first activity datasets. When data detection reveals that activity data with the same ID already exists in the target database, comparing the third dataset (the datasets included in the N first activity datasets) and the fourth dataset (the existing dataset in the database) can improve data usability and storage efficiency while ensuring data quality. It also significantly enhances the efficiency of data retrieval and updating, providing strong data support for interdisciplinary information exchange.
[0040] In an optional embodiment, storage optimization operations are performed on N first active datasets based on a third dataset and a fourth dataset. Specifically, this includes: comparing the third dataset and the fourth dataset to obtain a data comparison result; if the data comparison result shows that the third dataset and the fourth dataset are different, checking the data publishing attribute and data modification attribute of the fourth dataset, wherein the data publishing attribute is used to mark whether the fourth dataset has been pushed, and the data modification attribute is used to mark whether the fourth dataset has been modified; if the data publishing attribute is checked to be yes, adjusting the data change attribute of the fourth dataset to yes, wherein the data change attribute is used to mark whether the fourth dataset has changed; or, if the data publishing attribute is checked to be no and the data modification attribute is yes, adjusting the data change attribute to yes; or, if the data publishing attribute is checked to be no and the data modification attribute is no, updating the fourth dataset in the target database to the third dataset, and storing the other datasets (excluding the third dataset) among the N first active datasets in the target database according to a preset data storage format.
[0041] In the above embodiment, assume that a difference is found between the third and fourth datasets during the data comparison process. The data publication and modification attributes of the fourth dataset are checked. If the data publication attribute is "yes" (i.e., the fourth dataset has been publicly released), the data change attribute of the fourth dataset is adjusted to "yes," indicating that the fourth dataset needs to be updated. If the data publication attribute is "no" and the data modification attribute is "yes" (i.e., the fourth dataset has been internally edited but not yet publicly released), the data change attribute of the fourth dataset is adjusted to "yes." If both the data publication and data modification attributes are "no" (i.e., the fourth dataset has not been modified since collection), the latest third dataset is used to overwrite the fourth dataset, and other entirely new datasets are written to the target database. Through these judgments, it is possible to automatically identify which published data needs to be updated, which pending data needs to be replaced, and which are newly added data, thereby achieving the updating of existing data and the addition of incremental data.
[0042] Through the above steps, by comparing the third and fourth datasets, when the newly acquired third dataset differs from the existing fourth dataset in the database, the publication and modification attributes of the fourth dataset can be checked to determine whether the existing fourth dataset has been pushed to users and whether it has been modified. If the existing data has been pushed, its change attribute is marked as changed to facilitate subsequent incremental update pushes. If the existing data has not been pushed but has been modified, its change attribute is also marked as changed for update pushes. If the existing data has neither been pushed nor modified, the original data is directly replaced with the new third dataset, and other newly added datasets are stored. This storage optimization operation can minimize data changes while ensuring timely data updates, improve storage and push efficiency, and make cross-disciplinary information interaction more intelligent and efficient.
[0043] In an optional embodiment, before obtaining M registration datasets from the target database, the method further includes: upon detecting that multiple target users have registered for multiple target publishing projects, sending reminder information to multiple project publishers, wherein the multiple target publishing projects are a set of invited subject areas set by multiple project publishers when publishing multiple target publishing projects, and the reminder information is used to notify multiple project publishers that multiple target users have registered for multiple target publishing projects, the multiple target users include M target users, the multiple target publishing projects include Q target publishing projects, and the multiple project publishers include P project publishers; obtaining a first subject attribute set corresponding to the multiple target users, and obtaining a first invited subject attribute set corresponding to the multiple target publishing projects, wherein the first subject attribute set is a set of subject areas pre-set by the multiple target users, the first subject attribute set includes a target subject attribute set, the first invited subject attribute set is a set of invited subject areas set by multiple project publishers when publishing multiple target projects, the first invited subject attribute set includes a target invited subject attribute set; matching the first subject attribute set and the first invited subject attribute set to obtain a subject attribute matching result; and displaying the first subject attribute set in the subject attribute matching result. If the subject attribute set matches the first invited subject attribute set, multiple target users are identified as project participants in multiple target release projects, and multiple registration datasets are generated, wherein the multiple registration datasets include M registration datasets; or, if the subject attribute matching result shows that the first subject attribute set and the first invited subject attribute set do not match, the registration self-description information of the first target user among the multiple target users is sent to the front-end user terminal, and the second target user other than the first target user among the multiple target users is identified as a project participant, wherein the registration self-description information is the self-description information provided by the first target user when registering to participate in multiple target release projects; a receiving instruction is received from the front-end user terminal after reviewing the registration self-description information; according to the receiving instruction, the first target user is identified as a project participant, and multiple registration datasets are generated, wherein the first target user is all users among the multiple target users whose first subject attribute does not match the first invited subject attribute set, and the second target user is all users among the multiple target users whose second subject attribute matches the target invited subject attribute set, wherein the first subject attribute set includes the first subject attribute and the second subject attribute; and the multiple registration datasets are stored in the target database.
[0044] In the above embodiment, assume that users Xiaoming, Xiaohong, and Xiaogang are registered users of the system. Xiaoming's subject attribute in the system is "Computer Science," Xiaohong's is "Bioengineering," and Xiaogang's is "Materials Science." One day, the three of them registered to attend the "ABC Academic Conference" and the "XYZ Academic Discussion Conference" within the system, respectively. The invited subject attributes for both conferences are set to "Computer Science" and "Bioengineering." After detecting that Xiaoming, Xiaohong, and Xiaogang have registered for the two conferences, a reminder message is sent to Teacher Zhang, the publisher of the "ABC Academic Conference," stating, "Users Xiaoming, Xiaohong, and Xiaogang have registered to attend the ABC Academic Conference you published." Simultaneously, a reminder message is sent to Teacher Li, the publisher of the "XYZ Academic Discussion Conference," stating, "Users Xiaoming, Xiaohong, and Xiaogang have registered to attend the XYZ Academic Discussion Conference you published." The subject attribute set of Xiaoming, Xiaohong, and Xiaogang was obtained as {"Computer Science", "Bioengineering", "Materials Science"}, and the subject attribute set of the conference's invited participants was {"Computer Science", "Bioengineering"}. Matching these sets revealed that Xiaoming and Xiaohong's subject attributes matched the conference's invited subject attributes, while Xiaogang's subject attributes did not match. Therefore, Xiaoming and Xiaohong can be directly identified as conference participants, and their registration data is generated. Xiaogang's submitted registration statement was sent to the conference organizers, Teacher Zhang and Teacher Li, for review. After reviewing Xiaogang's statement, Teacher Zhang and Teacher Li issued a receipt instruction, approving Xiaogang's participation. Based on this, Xiaogang can also be identified as a conference participant, and his registration data is generated. Finally, the registration datasets of Xiaoming, Xiaohong, and Xiaogang for the "ABC Academic Conference" and the "XYZ Academic Symposium" are stored in the target database.
[0045] Through the above steps, after detecting that multiple target users have registered for multiple target-posted projects, reminders are sent to multiple project publishers to notify that target users have registered for the projects posted by the corresponding project publishers. By comparing the first subject attributes preset by multiple target users with the first invited subject attributes set by multiple project publishers, automatic matching of users and project needs is achieved, improving the matching degree between participants and projects. For first target users whose subject attributes do not match, the corresponding project publisher reviews the registration information of the first target users for further screening. On the one hand, this respects the project publisher's autonomous decision-making right, and on the other hand, it gives potential participants whose subject attributes do not match or do not fully match an opportunity, improving the flexibility of participation. Based on the matched and approved target users, a registration dataset (i.e., multiple registration datasets) is generated and stored. This registration mechanism, which takes into account both accurate matching and flexible participation, can maximize the rational allocation of high-quality project resources while ensuring the fit between projects and participants, and build an efficient bridge for interdisciplinary information exchange.
[0046] In an optional embodiment, N target activity projects are determined based on N first activity datasets, and target push operations are performed on the N target activity projects and Q target release projects according to the target subject attribute set, the target invited subject attribute set, and a preset push method. Specifically, this includes: obtaining a fifth dataset and a sixth dataset from the target database, wherein the fifth dataset is a collection of all unpushed data stored in the target database, and the sixth dataset is a collection of all pushed and changed data stored in the target database; performing data supplementation and data validity checks on the fifth and sixth datasets to determine the N target activity projects; obtaining user setting information, wherein the user setting information consists of push parameters preset by M target users; if, based on the user setting information, a third target user among the M target users is determined to have all projects pushed, a first message body is constructed based on the Q target release projects; and if, based on the user setting information, a third target user is determined to have all projects pushed, a first message body is constructed based on the Q target release projects; and if, based on the user setting information, a third target user is determined to have all projects pushed, a target push operation is performed on the N target activity projects and Q target release projects. If a target user has a third subject attribute, a second set of invited subject attributes matching the third subject attribute is determined from the target invited subject attribute set, wherein the target subject attribute set includes the third subject attribute; a first set of publishing projects corresponding to the second set of invited subject attributes is determined from Q target publishing projects; a first detection is performed on the first message body to obtain a first detection result; if the first detection result shows that the first message body includes the first publishing project set, a first invited label is added to the first publishing project set; or, if it is determined that the first message body does not include the first publishing project set, the first publishing project set label is added to the first message body, and a first invited label is added to the first publishing project set label; if it is determined that the first message body includes the first publishing project set with the first invited label, the first message body is pushed to the third target user according to the preset push method, and the N target activity projects are pushed to each target user included in the M target users respectively.
[0047] In the above embodiment, assume that a fifth dataset and a sixth dataset are obtained from the target database. The fifth dataset contains academic activity data that has never been pushed, while the sixth dataset contains academic activity data that has been pushed but recently changed. These two datasets can be supplemented by using a geocoding API to obtain detailed addresses for missing "activity location" data and standardizing the "activity time" field using regular expressions. The legality of these two datasets can be checked by verifying the format and value range of each field according to preset rules, eliminating non-compliant data items. Ultimately, M target activity projects and Q target publication projects are determined. The push settings information of registered users in the system is obtained. One user named Xiaomei has set "all pushes" and the subject attribute "Biomedical". Since Xiaomei has set "all pushes", all Q target publication projects are added to Xiaomei's push message body. The message body of Xiaomei's push notification is inspected. If all M target activity projects with the subject attribute related to "Biomedical" are correctly included, a "Special Invitation" tag is added to the "Biomedical" related target activity projects to indicate a high match with Xiaomei's subject attribute. The email service API is then invoked to send a message body containing all published projects and the "Special Invitation" tag to Xiaomei's registered email address. It should also be noted that the above supplementary operations and the above legality judgment operations are merely exemplary embodiments, and are not limited to the examples described above.
[0048] Through the above steps, by acquiring and verifying the completeness and legality of both unpushed and pushed data that has changed, the comprehensiveness, timeliness, and reliability of the pushed content are ensured. Based on the preset push parameters for M target users, when all M target users choose to receive push notifications, a precise match is made between the user's subject attributes (i.e., the first subject attribute set) and the invited subject attributes of the activity project (i.e., the first invited subject attribute set) to construct a push message body (i.e., the first message body) that aligns with the user's professional expertise. Adding a special invitation label to the push message effectively attracts the attention of target users. This allows for differentiated and precise project pushes to different user groups according to preset push methods, greatly improving the accuracy and timeliness of information pushes, enabling each target user to easily access high-quality projects of interest, and promoting interdisciplinary exchange and collaboration.
[0049] In an optional embodiment, after obtaining user settings information, the method further includes: if, based on the user settings information, it is determined that among the M target users, there is a fourth target user who has not set up full project push notifications, then, based on the user settings information, it is determined whether the fourth target user has set up project subscription push notifications; if it is determined that the fourth target user has set up project subscription push notifications, then, it is queried about the fourth target user's subscription tags; from the Q target publishing projects, it matches K target publishing projects that match the subscription tags, and determines the third invited subject attribute set corresponding to the K target publishing projects, wherein the target invited subject attribute set includes the third invited subject attribute set, and K is a positive integer less than or equal to Q; it constructs a second message body based on the K target publishing projects; if, based on the user settings information, it is determined that the fourth target user does not have a fourth subject attribute, then, it pushes the second message body to the fourth target user according to a preset push method, wherein the target subject attribute set includes the fourth subject attribute.
[0050] In the above embodiment, assume that user Zhang San has independently set push preferences within the system. He has chosen to subscribe to course push notifications for the subjects "Mathematics" and "Physics," but has not selected "Push All." When preparing to push newly released academic projects, the system first checks user Zhang San's settings. Because Zhang San has not selected "Push All," it is necessary to further determine whether he has set up subscription push notifications for specific subjects. It is confirmed that Zhang San has set up subscription push notifications for the subjects "Mathematics" and "Physics." Zhang San's subscription tags are queried, revealing that he has subscribed to the subjects "Mathematics" and "Physics." Assume that there are 10 newly released academic projects to push, covering multiple subjects such as mathematics, physics, chemistry, and biology. From these 10 academic projects, projects matching Zhang San's subscription tags are selected. Assume that 3 projects (i.e., K is 3, 2 mathematics projects and 1 physics project) match Zhang San's subscription tags. For these 3 target release projects, the corresponding third invited subject attribute set is determined, namely "Mathematics" and "Physics." A message body (i.e., the second message body) is constructed based on the information of these 3 target release projects (title, description, link, etc.). Further examination of Zhang San's user settings confirmed that, apart from his subscriptions to "Mathematics" and "Physics," he had no additional subject attributes set (i.e., no fourth subject attribute). Since Zhang San did not set any additional subject attributes, a pre-constructed second message was pushed to him according to preset push methods (e.g., email, SMS, in-app notifications), informing him of the newly released Mathematics and Physics course projects. This demonstrated how to accurately push relevant course projects to users based on their settings, improving user experience and push efficiency.
[0051] Through the above steps, for target users who have not selected all push notifications (i.e., the fourth target user), the system automatically determines whether they have pre-subscribed to push notifications. If they have subscribed to push notifications, it matches K target publications from Q target publications based on subscription tags, determines the set of invited subject attributes corresponding to the K target publications (the third invited subject attribute set), and constructs a push message body (i.e., the second message body) that matches the user's interests and preferences. If they have not subscribed to push notifications, no push is sent, fully respecting the user's wishes. Furthermore, for subscribers who have not set subject attributes (i.e., the fourth subject attribute), there is no need to add invited tags; subscription messages can be pushed directly, eliminating the subject attribute matching step and making the push process more efficient. By respecting user choices, the system maximizes the satisfaction of users' personalized subscription needs, enabling them to easily access high-quality project resources of interest and improving their information acquisition experience, while reducing unnecessary matching steps and saving computing resources.
[0052] Obviously, the embodiments described above are only some embodiments of this application, and not all embodiments. The present application will be specifically described below with reference to specific embodiments.
[0053] This application provides a two-way selection project registration process, see below. Figure 2 , Figure 2 This is a schematic diagram of a two-way selection process in an embodiment of this application, including the following steps:
[0054] Step S201: The publisher publishes the project and it is approved.
[0055] Step S202: The applicant browses the approved projects (corresponding to the Q target projects mentioned above), is interested in the current project, and registers for the project (a reminder message is sent to the publisher when registering);
[0056] Step S203: Perform the first judgment to determine whether the project has set up invited disciplines;
[0057] Step S204: If the first judgment result is yes, a second judgment is made to determine whether the invited discipline of the project matches the discipline of the applicant.
[0058] Step S205: If the first judgment result is negative, after receiving the registration reminder, the publisher enters the system and views the registrant's registration statement.
[0059] Step S206: If the second judgment result is yes, automatically accept the applicant as a project participant; if the second judgment result is no, return to step S205.
[0060] Step S207: Perform a third judgment to determine whether the publisher accepts applicants as project participants; if the result of this third judgment is no, the project registration fails.
[0061] In step S208, if the third judgment result is yes, the publisher and the applicant are allowed to view each other's contact information.
[0062] In the user information retrieval interface for publishers and applicants, a "data permission filtering unit" exists on the backend server. This unit determines whether the publisher and applicant have successfully established a handshake based on the project. If the handshake is successful, contact information is returned; otherwise, it is not. This permission filtering ensures that unauthorized users cannot access user contact information, making the handshake process highly secure. Users do not need to perform additional operations or redirects when retrieving user information through the same data interface, resulting in a better user experience. The middle 5 digits of the student ID are hidden in the backend server's user information interface. Specifically, the first two and last three digits of the student ID are extracted, and the middle 5 digits are replaced with * (the number of digits in the student ID is not fixed, possibly ranging from 5 to 10). Hiding part of the student ID reduces the risk of information leakage; even if the frontend code is viewed or tampered with, attackers cannot easily obtain the complete student ID. For users, partially hiding sensitive information increases their trust.
[0063] After project registration is approved, both the project poster and the applicant can see each other's names, student IDs, departments, majors, grades, academic levels, supervisors, emails, mobile phones, and WeChat IDs. It's important to note that when viewing contact information after successful registration, different fields will be displayed depending on the applicant's identity; for example, teachers will not have their major, grade, or student category listed. Users can select "Participate in Projects" on the system's homepage (the information display page on the front-end) to browse published projects and filter and search by project status, type, theme, and keywords. For projects of interest, users can select "Apply," filling out a self-introduction (e.g., briefly introducing their views on the project to the project poster) to achieve efficient two-way matching. Project posters can view applicant information on the project details page and decide whether to "accept" the applicant as a collaborator. Applicants belonging to specially invited first-level disciplines can be "accepted" without poster confirmation. After a project poster accepts an applicant, a two-way selection is complete, and both parties can view each other's information and contact each other for offline communication and collaboration. When a project is completed or no new applications are accepted, the project organizer can end the project on the project details page. A project summary must be completed when ending the project. Users can select "Feedback" on the system's homepage to provide feedback to the Graduate School regarding system usage.
[0064] This application also provides a message push process, see below. Figure 3 , Figure 3 This is a schematic diagram of a message push process in an embodiment of this application, including the following steps:
[0065] Step S301: Obtain all projects that were approved yesterday (corresponding to the above N target activity projects and the above Q target release projects);
[0066] 1. Obtain Q target published projects: This refers to obtaining academic exchange project information or conference project information published by students or teachers. Academic exchange project information includes, but is not limited to, the topic of the published information, a brief description of the content, the expected end time, keywords, whether a specific discipline is invited, and project tags. After the academic exchange project information is published, information such as department, major, supervisor, and year will be displayed, but personal contact information will not be displayed.
[0067] Conference project information includes, but is not limited to, the topic of the information to be published, a brief description of the content, the conference location, the expected conference time, an image of the conference venue application approval result to be uploaded, keywords, whether a specific discipline is specially invited (e.g., first-level discipline), and project tags (same as the category on the project publishing page); after the conference project information is published, information such as department, major, supervisor, and year will be displayed, but personal contact information will not be displayed.
[0068] 2. Obtain information on N target events by using a web crawler (Python) to retrieve information about academic conferences or events published by various departments and research institutes within the university:
[0069] Data is crawled at 8:00 AM, 12:00 PM, and 4:00 PM daily (other times are also acceptable, no specific time limit is imposed here). The crawled data is stored in the database but not published immediately. Administrators log into the system daily after the data crawling is complete to supplement and publish the data.
[0070] The data scraping process is explained below:
[0071] S1, the activity posting page of each department, analyzes the HTML structure of each page one by one using the browser console tool, and analyzes the document object model (DOM) of the activity list.
[0072] S2, scrape the activity list data and obtain the activity details URL, activity title, activity release time, activity event time, activity speaker, and activity location;
[0073] 1) The activity URL is the key field for crawling. Other data that cannot be obtained can be automatically supplemented by the system and / or manually supplemented by the administrator.
[0074] 2) The event release and event duration data may not be formatted. It needs to be formatted into a uniform format based on the different syntax of each page. The actual data format might be: January 1, 2024, 01 / 01 / 2024, or 2024, January 1. The final required format can be set to YYYY-MM-DD.
[0075] S3, after scraping the data, it is necessary to first check the database based on the activity's URL address to see if the data has already been scraped;
[0076] Each time activity data is stored, an MD5 value needs to be calculated based on the UR address. This MD5 value is used to determine whether the activity information already exists in the database. If it is determined that the activity has not been crawled, it is directly stored in the database. If it is determined that the data already exists, the newly crawled data is compared with the data already in the database. If the data has changed, it is marked as changed data. This can be done automatically by the system or manually by the administrator, etc. There is no limitation here.
[0077] S4: Data is not published (i.e. pushed) immediately after being scraped. Users cannot see the newly scraped activity data at this time. It needs to be reviewed and published by the administrator.
[0078] 1) For information with relatively standardized data after scraping, the administrator can directly set it to be published;
[0079] 2) For data that is missing information after being crawled, the administrator needs to click the activity details link to supplement the activity information before publishing.
[0080] S5: After an event is published, users can view it on the page as a calendar component and search for it by event theme or speaker.
[0081] The data scraping process is explained in detail below:
[0082] It is mainly divided into three stages: preparation, automatic capture, and manual intervention.
[0083] 1. Preparation stage
[0084] 1) Collect information offline from various academic activity websites;
[0085] 2) Analyze the availability of content on each website. Unavailable websites include those with content that has not been updated for a long time, those lacking activity details pages, and those lacking specific activity times.
[0086] 3) Configure the data scraping settings for each website. Analyze the HTML structure of each website page using the browser console tool to obtain the XPath (the specific path of the field content in the page's HTML) of each field that needs to be scraped. This configuration sets how to scrape data from each website and which data to scrape. Note the "Whether to access the URL of a single activity" option. For some websites, the academic activity list page can directly scrape the required fields, but for others, the list page only contains the title, requiring you to scrape the details page of each activity to obtain the necessary fields. The configuration is shown in Table 1:
[0087] Table 1
[0088] field name Can it be empty? Example value Activity list page URL no https: / / xxxx.xxx.xxx / xshy.html Unit Code no 00001 Company Name no XXXX College Academic Activity List XPath no / / div[@class="list"] All active XPaths in the list no .li XPath of the event details page URL no .a / @href Do I need to enter the URL of a single activity? no no XPath of Activity Title yes .a / @title XPath for event time yes .a / span / text() XPath of the event speaker yes .a / div / h3 / text() XPath of the event location yes .a / div / span / text() XPath for event release time yes .a / div / p / text()
[0089] 2. Automatic capture stage
[0090] 1) Start
[0091] 2) Load the configuration module
[0092] Load all data scraping configurations and scrape content from each website one by one.
[0093] 3) Crawling Module
[0094] a. Based on the configuration, perform data scraping, and execute steps 4), 5), and 6) for each scraped academic activity data;
[0095] b. If the "Do we need to enter the URL of a single activity?" option is no, the crawling module does not need to crawl the specific data of the activity details page. The XPath of the five fields, such as the activity title and the activity time, are all for the activity list page.
[0096] c. If "Do I need to access the URL of a single event?" is yes, then the crawling module also needs to crawl the specific data of the event details page. The XPath of the five fields, such as event title and event time, are all based on the HTML structure of the event details page.
[0097] 4) Data Optimization Module
[0098] a. Optimize the data quality of the captured data fields;
[0099] b. Optimize the format of the event and release times. If the obtained time matches one of the preset regular expressions, format it as YYYY-MM-DD. The preset regular expressions are:
[0100] \d{4}-\d{1,2}-\d{1,2}
[0101] d{4}[. / ]\d{1,2}[. / ]\d{1,2}
[0102] \d{4}.*\d{1,2}.*\d{1,2}
[0103] \d{1,2}.+?\d{1,2},\s*\d{4}
[0104] Among them, the time format matched by "\d{4}-\d{1,2}-\d{1,2}" is "YYYY-MM-DD", the time format matched by "d{4}[. / ]\d{1,2}[. / ]\d{1,2}" is "YYYY.MM.DD" or "YYYY / MM / DD", the time format matched by "\d{4}.*\d{1,2}.*\d{1,2}" is "YYYY年MM月DD日", and the time format matched by "\d{1,2}.+?\d{1,2},\s*\d{4}" is "MM DD,YYYY";
[0105] c. Optimize the data of the activity title, speaker, and activity location, including: removing "【】" before and after the title, removing the label text of the data (such as "Speaker: ", etc.), removing special characters, removing blank characters, etc.
[0106] 5) Data filtering module
[0107] a. Calculate the MD5 value of the URL field of the activity details page as the unique ID, and check whether the academic activity has been saved in the database according to the ID;
[0108] b. If the activity has been saved, compare the currently crawled data with the data in the database. When comparing, check five fields: the activity theme, activity holding time, activity speaker, activity location, and activity release time. If the data has changed, check the two fields SFFB and SFSDXG;
[0109] If SFFB = yes, modify the SFBH field of the data to yes;
[0110] If SFFB = no, and SFSDXG = yes, modify the SFBH field of the data to yes;
[0111] If SFFB = no, and SFSDXG = no, automatically update the above five fields to the latest crawled data;
[0112] c. If the activity has not been saved, enter the data storage module.
[0113] 6) Data storage module
[0114] a. Save the captured academic activity information to the database. The SFFB, SFBH, and SFSDXG fields are all set to no by default.
[0115] b. The saved data format is shown in Table 2:
[0116] Table 2
[0117]
[0118]
[0119] 3. Artificial intervention stage
[0120] Idea Lab can have a management backend. During the manual intervention phase, administrators can use this backend to supplement and publish the automatically captured academic activity data.
[0121] 1) The management backend page lists all academic activity data that has been crawled but not yet published, as well as data that has been published but has changed.
[0122] 2) Manual processing of captured academic activity data:
[0123] If the SFBH field is yes, then compare the data displayed in the current management backend with the data seen by clicking the link through the HDURL field, manually determine whether the data needs to be modified, and then proceed with the subsequent publishing operation after modification;
[0124] If the data content is non-compliant, the administrator should set SFFB to no.
[0125] If the data content is compliant and the data fields are complete, and the administrator sets SFFB to yes, users can view the data after it is published.
[0126] If the data content is compliant but the data fields are incomplete, the administrator can click the link in the HDURL field to view the activity content, and then manually supplement the data. After supplementing the data, SFSDXG will be automatically set to "Yes". After manually supplementing the data, SFFB can then be set to "Yes".
[0127] Step S302: Obtain user settings information;
[0128] When users log in to Idea Lab (corresponding to the interdisciplinary collaboration system mentioned above) for the first time, they need to set up their personal information (i.e., complete their personal information and subscription preferences), including but not limited to their discipline, campus, on-campus location, email, mobile phone number, and WeChat ID. This personal information is displayed when posting or applying for projects. Users can fill in their campus, work location, email, and contact number on the "Personal Information" page. They can set the content, format, and channel for new project push notifications on the "Push Settings" page. They can select tags of interest in the "Subscription Settings." If users select to receive daily project push notifications and "Push according to subscription" in the "Push Settings," only projects related to their subscribed tags will be pushed to them.
[0129] The subject area in personal information is used to push project information based on individual settings when a project posting specifies a particular subject. The default subject values are: undergraduate students have no subject, graduate students can search for subjects based on their major, and professors can set subjects based on their research direction. The email address in personal information is set to the student ID number by default, but can be modified.
[0130] Step S303: Perform the fourth judgment to determine whether to select all push notifications;
[0131] Step S304: If the result of the fourth judgment above is yes, construct a message body based on all items;
[0132] 1) Define the message body structure, as shown in Table 3:
[0133] Table 3
[0134]
[0135] 2) Retrieve all projects that were approved yesterday, obtain the type and title of each project, and construct the msg field of the message body.
[0136] 3) Generate message text `content` based on the message body. `content` supports raw HTML values, and can be set to `msg+'`. +tips.
[0137] 4) Send messages to the user via chat app or email according to each user's settings.
[0138] Step S305: Perform the fifth judgment to determine whether the user has a subject-specific attribute;
[0139] Step S306: If the result of the fourth judgment is negative, a sixth judgment is made to determine whether to push according to the subscription; Step S307: If the result of the sixth judgment is negative, the project is not sent.
[0140] Step S308: If the result of the sixth judgment above is yes, query the user's subscription tags;
[0141] Step S309: Match items based on subscribed tags;
[0142] Step S310: Generate a message based on the matched items and return to step S305;
[0143] Step S311: If the result of the fifth judgment above is yes, the project of the invited discipline is set according to the user's subject attribute.
[0144] Figure 4 This is an entity relationship diagram of users, disciplines, and projects in the embodiments of this application, see [link / reference]. Figure 4 Each user can subscribe to 0-N tags, and each project can also set 0-N tags. Therefore, the relationship between user tags and project tags is N:N. The matching order is t_project->t_project_label->t_label->t_user_label->t_user:
[0145] 1. Student users can only have one subject attribute; faculty and staff users may have multiple subject attributes.
[0146] 2. The project allows for 0-N invited disciplines.
[0147] 3. The relationship between user disciplines and project-invited disciplines is N:N:
[0148] Two relation tables, t_user_subject and t_project_subject, are created to store the relationships between users and subjects, and between projects and invited subjects. These tables allow for the quick retrieval of users matching the invited subjects for each project. The specific matching order is as follows: first, find the relationship between the project and the invited subject from the project table; then, obtain the subject; finally, query the relationship between the subject and the user to find the user. The matching order is t_project -> t_project_subject -> t_subject -> t_user_subject -> t_user.
[0149] Step S312: If the result of the fifth judgment above is negative, send a message to the user;
[0150] Step S313: Perform the seventh judgment to determine whether there is a project in the invited discipline; if the result of the seventh judgment is no, return to step S312.
[0151] Step S314: If the result of the seventh judgment is yes, perform the eighth judgment to determine whether the invited discipline is in the generated message;
[0152] Step S315: If the result of the eighth judgment above is yes, mark the item as specially invited and return to step S312;
[0153] In step S316, if the result of the eighth judgment above is negative, add the current project to the pending messages and mark it as an invitation, then return to step S312.
[0154] It's also worth noting that the settings include an option to "Receive Invited Subject Project Push Notifications." If "Yes" is selected and the invited subject field of the project matches the user's subject, the project will be pushed. There's also an option to "Push by Subscription." If "Yes" is selected, the project's tag attributes will be pushed if they match the user's subscribed tags. The push method is scheduled direct push. Through configurable scheduled tasks, personalized messages for each user can be sent directly to their chat app or email at a predetermined time (the message sending channel can also be a public account, etc.). For example, every Monday, all registered users can receive an email push with information about academic activities published that week, displayed in the user's academic calendar. Additionally, students and teachers can subscribe to project information based on subscription tags or first-level disciplines. Project information can be pushed to subscribers daily based on their subscription information. The push time can be 7:00 AM every day, and the content can be projects newly published between 7:00 AM yesterday and 7:00 AM today. The push format is email and chat app push, with email being the default method. Both methods can be disabled. For small user groups, choosing this scheduled direct push method can ensure that each user receives personalized messages in a timely and accurate manner, while also effectively reducing system complexity and the need for real-time processing and immediate response, thereby simplifying system design and maintenance.
[0155] This application also provides a system architecture. Figure 5 This is a schematic diagram of an architecture of an interdisciplinary collaborative system as described in this application. (See attached image.) Figure 5 The management side is for graduate school administrators, who can set up system operation, query participant information, review and manage projects, supplement and review academic calendars, and query and respond to feedback in the interdisciplinary collaboration system. The user side is for students and mentors, who can set up personal push notifications, register for projects, manage academic calendars, provide feedback, and seek help in the interdisciplinary collaboration system. Mentors can also manage student activities (posting and registration). In addition, graduate school administrators can decide whether students can use the project posting and management functions.
[0156] In this embodiment, prior to precise data field extraction, the HTML structure of each target website has been analyzed, and the corresponding data extraction configuration has been saved. This allows for direct acquisition of the required information without processing large amounts of irrelevant data, thus improving the efficiency and accuracy of data extraction. In particular, key fields in academic activity notifications (such as activity dates and speakers) must be accurate to ensure the effectiveness of subsequent data analysis and use. While automatic data extraction can efficiently acquire large amounts of data, in certain special cases, due to changes in webpage structure or the specific nature of the data, automatic extraction may not fully meet the needs. In such cases, manual intervention becomes crucial. Through manual inspection and correction, the integrity and accuracy of the data can be ensured, improving data quality and making data extraction more flexible and reliable. The calendar format clearly displays the timeline of academic activities, allowing users to easily see the dates, locations, and overlaps with other activities.
[0157] The electronic device in the embodiments of this invention is described below from the perspective of hardware processing. (See attached document.) Figure 6 , Figure 6 This is a schematic diagram of the physical device structure of an electronic device in an embodiment of this application.
[0158] It should be noted that, Figure 6 The structure of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.
[0159] like Figure 6 As shown, the electronic device includes a Central Processing Unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in Read-Only Memory (ROM) 602 or a program loaded from storage portion 608 into Random Access Memory (RAM) 603, such as performing the methods described in the above embodiments. The RAM 603 also stores various programs and data required for system operation. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An Input / Output (I / O) interface 605 is also connected to the bus 604.
[0160] The following components are connected to I / O interface 605: input section 606 including audio input devices, push-button switches, etc.; output section 607 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 608 including a hard disk, etc.; and communication section 609 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 609 performs communication processing via a network such as the Internet. Drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 610 as needed so that computer programs read from them can be installed into storage section 608 as needed.
[0161] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit (CPU) 601, it performs the various functions defined in the present invention.
[0162] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.
[0164] Specifically, the electronic device in this embodiment includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the interdisciplinary information interaction method provided in the above embodiment.
[0165] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The storage medium carries one or more computer programs that, when executed by a processor of the electronic device, cause the electronic device to implement the interdisciplinary information interaction method provided in the above embodiments.
[0166] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0167] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for interdisciplinary information exchange, characterized in that, include: Obtain N first activity datasets from the target database, and obtain M registration datasets from the target database, wherein the N first activity datasets include key information extracted from academic activities in different disciplines, and the M registration datasets are sets of registration records generated when M target users successfully register for Q target published projects, and N, M, and Q are all positive integers greater than or equal to 1; The target subject attribute set is determined based on the M registration datasets, and the target invited subject attribute set is determined based on the Q target publishing projects. The target subject attribute set is a set of subject fields pre-set by the M target users, and the target invited subject attribute set is a set of invited subject fields set by P project publishers when publishing the Q target publishing projects, where P is a positive integer greater than or equal to 1. N target activity projects are determined based on the N first activity datasets, and target push operations are performed on the N target activity projects and the Q target release projects according to the target subject attribute set, the target invited subject attribute set, and the preset push method. Before obtaining N first activity datasets from the target database, data is scraped from academic publishing websites to obtain N second activity datasets; Optimize the activity data that meets the preset optimization conditions from the N second activity datasets to obtain the N first activity datasets; Determine the activity detail page locators for N academic activity web pages from the N first activity datasets; N IDs can be obtained from the MD5 hash values of N activity detail page locators; Based on the N IDs, check whether the first ID exists in the target database; If a first ID already exists in the target database, a third dataset corresponding to the first ID is determined from the N first activity datasets, and a fourth dataset corresponding to the first ID is determined from the target database, wherein the N IDs include the first ID; The third dataset and the fourth dataset are compared; If the third dataset and the fourth dataset are different, the data publishing attribute and data modification attribute of the fourth dataset are checked, wherein the data publishing attribute is used to mark whether the fourth dataset has been pushed, and the data modification attribute is used to mark whether the fourth dataset has been modified; If the data publishing attribute is found to be yes, the data change attribute of the fourth dataset is adjusted to yes, and an incremental update is pushed to the user, wherein the data change attribute is used to mark whether the fourth dataset has changed; or, if the data publishing attribute is no and the data modification attribute is yes, the data change attribute is adjusted to yes so that an update can be pushed; or, if the data publishing attribute is no and the data modification attribute is no, the fourth dataset in the target database is updated to the third dataset, and the other datasets among the N first activity datasets except the third dataset are stored in the target database according to a preset data storage format; Before retrieving M registration datasets from the target database, the method further includes: If it is detected that multiple target users have registered for multiple target release projects, a reminder message is sent to multiple project publishers. The multiple target release projects are a set of invited subject areas set by the multiple project publishers when they release the multiple target release projects. The reminder message is used to notify the multiple project publishers that the multiple target users have registered for the multiple target release projects. The multiple target users include the M target users, the multiple target release projects include the Q target release projects, and the multiple project publishers include the P project publishers. Obtain a first subject attribute set corresponding to the plurality of target users, and obtain a first specially invited subject attribute set corresponding to the plurality of target publishing projects, wherein the first subject attribute set is a set of subject fields pre-set by the plurality of target users, and the first subject attribute set includes the target subject attribute set; the first specially invited subject attribute set is a set of specially invited subject fields set by the plurality of project publishers when publishing the plurality of target projects, and the first specially invited subject attribute set includes the target specially invited subject attribute set. The first subject attribute set and the first specially invited subject attribute set are matched to obtain the subject attribute matching result; If the subject attribute matching result shows that the first subject attribute set and the first invited subject attribute set match, the multiple target users are identified as project participants of the multiple target publishing projects, and multiple registration datasets are generated, wherein the multiple registration datasets include the M registration datasets.
2. The method according to claim 1, characterized in that, Before retrieving N first active datasets from the target database, the method specifically includes: The webpage structure of the academic publishing website is analyzed to set N target crawling configurations for the academic publishing website. The academic publishing website includes N academic activity webpages, and the N academic activity webpages include N academic activity-related information from different disciplines. The N target crawling configurations correspond one-to-one with the N academic activity webpages. According to the N target crawling configurations, each academic activity webpage included in the academic publishing website is sequentially traversed, and the following operations are performed on the currently traversed academic activity webpage: a first configuration item is determined from the first crawling configuration, wherein the first configuration item is used to determine whether to crawl the activity details page data of the first academic activity webpage, the first academic activity webpage is the currently traversed academic activity webpage, the first crawling configuration is the data crawling configuration corresponding to the first academic activity webpage, and the N target crawling configurations include the first crawling configuration; if the first configuration item is yes, the activity details page data and activity list page data are crawled according to the first crawling configuration to obtain a first dataset, wherein the first academic activity webpage includes the activity list page data; or, if the first configuration item is no, the activity list page data is crawled according to the first crawling configuration to obtain a second dataset; Given N second activity datasets, perform data optimization operations on the N second activity datasets, wherein the N second activity datasets include either the first dataset or the second dataset, and the N second activity datasets correspond one-to-one with the N academic activity web pages.
3. The method according to claim 1, characterized in that, The method further includes: If the subject attribute matching result shows that the first subject attribute set and the first invited subject attribute set do not match, the registration self-description information of the first target user among the multiple target users is sent to the front-end user terminal, and the second target user other than the first target user among the multiple target users is identified as the project participant, wherein the registration self-description information is the self-description information provided by the first target user when registering to participate in the multiple target release projects; a receiving instruction is received from the front-end user terminal after reviewing the registration self-description information; the first target user is identified as the project participant according to the receiving instruction, and the multiple registration datasets are generated, wherein the first target user is all users among the multiple target users whose first subject attribute does not match the first invited subject attribute set, and the second target user is all users among the multiple target users whose second subject attribute matches the target invited subject attribute set, and the first subject attribute set includes the first subject attribute and the second subject attribute; The multiple registration datasets are stored in the target database.
4. The method according to claim 3, characterized in that, The step of determining N target activity projects based on the N first activity datasets, and performing target push operations on the N target activity projects and the Q target release projects according to the target subject attribute set, the target invited subject attribute set, and a preset push method, specifically includes: Obtain a fifth dataset and a sixth dataset from the target database, wherein the fifth dataset is a collection of all unpushed data stored in the target database, and the sixth dataset is a collection of all pushed and changed data stored in the target database; Perform data supplementation and data validity checks on the fifth and sixth datasets to determine the N target activity items; Obtain user settings information, wherein the user settings information is the push parameters pre-set by the M target users; If, based on the user settings information, it is determined that among the M target users, there is a third target user who has set up all project push notifications, then a first message body is constructed based on the Q target published projects. If it is determined that the third target user has a third subject attribute based on the user settings information, a second set of invited subject attributes that matches the third subject attribute is determined from the target invited subject attribute set, wherein the target subject attribute set includes the third subject attribute; From the Q target release projects, determine the first release project set corresponding to the second specially invited subject attribute set; Perform a first detection on the first message body to obtain a first detection result; If the first detection result shows that the first message body includes the first published project set, add a first special invitation label to the first published project set; or, if it is determined that the first message body does not include the first published project set, add the first published project set to the first message body and add the first special invitation label to the first published project set. If it is determined that the first message body includes the first set of published projects that have been marked with the first special invitation, the first message body is pushed to the third target user according to the preset push method, and the N target activity projects are pushed to each of the M target users respectively.
5. The method according to claim 4, characterized in that, After obtaining the user settings information, the method further includes: If, based on the user settings information, it is determined that among the M target users, there is a fourth target user who has not set up full push notifications for all items, then based on the user settings information, it is determined whether the fourth target user has set up item subscription push notifications. If it is determined that the fourth target user has set up a subscription push for the project, query the subscription tags of the fourth target user; From the Q target publishing projects, K target publishing projects that match the subscription tags are selected, and the third invited subject attribute set corresponding to the K target publishing projects is determined, wherein the target invited subject attribute set includes the third invited subject attribute set, and K is a positive integer less than or equal to Q; Construct a second message body based on the K target publishing projects; If it is determined from the user settings information that the fourth target user does not have a fourth subject attribute, the second message body is pushed to the fourth target user according to the preset push method, wherein the target subject attribute set includes the fourth subject attribute.
6. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-5.
7. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-5.
8. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Label classification-based campus teacher and student scientific research project matching method
CN113191924A
Dynamic social networking service system and respective methods in collecting and disseminating specialized and interdisciplinary knowledge
US20150052198A1