Federal learning-based education data hidden label generation method
By collecting interaction frequency, feature activity indicators, and coverage information from multi-source educational data, a basic batch probability mapping model is constructed. Combined with federated learning technology, the problems of multi-source data integration and privacy protection are solved, achieving efficient and accurate generation of hidden labels and data privacy protection.
Patent Information
- Application Number
- CN202510950698.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-11-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing educational data processing methods struggle to efficiently integrate multi-source data and generate accurate hidden labels, while also posing risks of data privacy breaches and failing to fully leverage the advantages of federated learning.
By collecting time-series records of interaction frequency, feature activity indicators, and coverage information from multi-source educational data, a basic batch probability mapping model is constructed. Validation parameters are dynamically monitored, and federated learning technology is used for data processing to achieve hidden label generation.
It enables efficient and accurate generation of hidden labels from multi-source educational data, protects data privacy, improves processing efficiency and quality, and complies with data privacy protection regulations.
Smart Images

Figure CN120892718A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of educational data processing, and particularly to an educational data implicit label generation method based on federated learning. BACKGROUND
[0002] In the context of rapid development of educational informatization, educational data is showing an explosive growth trend. These data contain rich information, such as students' learning behavior, interaction frequency, feature activity indicators, etc., which have important value for educational research and teaching optimization. However, educational data has characteristics such as multi-source, complexity and privacy, which bring many challenges to the effective use of data.
[0003] Traditional data processing methods often have difficulty in efficiently integrating and utilizing data from different sources when faced with multi-source educational data. For example, the data formats and standards generated by different schools and different course platforms may differ, making it difficult to directly fuse and analyze the data. At the same time, in the processing process, how to accurately extract valuable information from massive data, especially the generation of implicit labels, is a key problem. Implicit labels can reflect the potential characteristics and patterns in the data, and have important significance for student learning status evaluation, teaching effectiveness analysis, etc., but the accuracy and efficiency of generating implicit labels by traditional methods need to be improved.
[0004] In addition, educational data involves the privacy information of students and teachers, and direct sharing of raw data may cause privacy leakage risk. With the increasing strictness of data privacy protection regulations, how to realize the collaborative processing and analysis of multi-source educational data while protecting data privacy has become a problem to be solved. Traditional centralized data processing mode needs to store and process data centrally, which not only increases the risk of data leakage, but also faces obstacles in data sharing.
[0005] Federated learning, as an emerging distributed machine learning technology, provides a new way to solve the above problems. It allows participants to collaboratively build machine learning models without sharing raw data, thus realizing joint analysis and utilization of data while protecting data privacy. However, there are still many technical difficulties in applying federated learning technology to the field of educational data implicit label generation. For example, how to effectively collect and integrate the interaction frequency time series records, feature activity indicators and coverage information of multi-source educational data, how to build an accurate base batch probability mapping model to select appropriate base data batches, how to determine whether the base data batches need to be labeled or segmented according to the verification parameters, and how to efficiently determine the final number of implicit label segments, etc.
[0006] At present, the existing education data processing method lacks comprehensive consideration of the characteristics of multi-source data in generating hidden labels, and does not fully utilize the advantages of federated learning to protect data privacy and improve processing efficiency. Therefore, an education data hidden label generation method based on federated learning is urgently needed to solve the problems existing in the prior art, realize efficient processing of multi-source education data and accurate generation of hidden labels, and protect data privacy. SUMMARY
[0007] The present application aims to provide an education data hidden label generation method based on federated learning to solve the problems raised in the background art.
[0008] To achieve the above-mentioned purpose, the present application provides an education data hidden label generation method based on federated learning, which comprises:
[0009] Collecting the interaction frequency time series records, multi-period feature active indicators and coverage range information of multi-source education data to form a current interaction frequency time series set, a current feature active indicator set and a current coverage range information set;
[0010] Collecting the probability information of each candidate data batch as a basic batch, the interaction frequency time series records, multi-period feature active indicators and coverage range information of each candidate data batch, and constructing a final basic batch probability mapping model; selecting a current basic data batch according to the final basic batch probability mapping model, the current interaction frequency time series set, the current feature active indicator set and the current coverage range information set;
[0011] By collecting the verification parameters after using the current basic data batch, it is determined whether the current basic data batch needs to be labeled or segmented, and a judgment result is obtained;
[0012] According to the judgment result, the current basic data batch is segmented to obtain the final current hidden label segment number.
[0013] Preferably, the collecting of the interaction frequency time series records, multi-period feature active indicators and coverage range information of multi-source education data to form a current interaction frequency time series set, a current feature active indicator set and a current coverage range information set comprises the following operations: setting the education data scene to be processed and the corresponding data batches to form a candidate data batch set; setting a plurality of parameter categories that may affect the possibility of the data batch being used as a basic batch to form a basic batch selection influence parameter category set; setting a first current statistical period; combining the first current statistical period and the basic batch selection influence parameter category set to obtain the interaction frequency time series records, multi-period feature active indicators and coverage range information of each data batch in the candidate data batch set, and form a current interaction frequency time series set, a current feature active indicator set and a current coverage range information set.
[0014] Preferably, the collection of each candidate data batch as the basis batch probability information, the interaction frequency time series record of each candidate data batch, the multi-period feature active index and the coverage information and the construction of the final basis batch probability mapping model include the following operations:
[0015] Set the historical statistical period; combine the candidate data batch set and the basis batch selection influence parameter category set, collect the probability information of each candidate data batch as the basis batch in the historical statistical period, the interaction frequency time series record of each candidate data batch, the multi-period feature active index and the coverage information, form the historical candidate data batch basis probability information set, the historical interaction frequency time series set, the historical feature active index set and the historical coverage information set; use the historical candidate data batch basis probability information set, the historical interaction frequency time series set, the historical feature active index set and the historical coverage information set to construct the final basis batch probability mapping model and the final interaction frequency weight time series set.
[0016] Preferably, the construction of the final basis batch probability mapping model and the final interaction frequency weight time series set adopts a federal aggregation optimization strategy.
[0017] Preferably, the selection of the current basis data batch according to the final basis batch probability mapping model, the current interaction frequency time series set, the current feature active index set and the current coverage information set includes the following operations:
[0018] Set the current original basis batch; substitute each information in the current interaction frequency time series set, the current feature active index set and the current coverage information set and the final interaction frequency weight time series set into the final basis batch probability mapping model respectively to perform mapping, obtain the current candidate data batch basis probability information set; when the candidate data batch corresponding to the maximum basis probability information in the current candidate data batch basis probability information set is the same as the current original basis batch, the current basis data batch is not replaced; otherwise, the candidate data batch corresponding to the maximum basis probability information in the current candidate data batch basis probability information set is taken as the current basis data batch, and all are recorded as the current basis data batch.
[0019] Preferably, the judgment of whether the current basis data batch needs to be labeled or segmented by collecting the verification parameters after the use of the current basis data batch to obtain a judgment result includes the following operations:
[0020] Setting a second current statistical period; setting a plurality of time nodes in the second current statistical period to form a current time node set; setting a plurality of categories of verification parameters capable of reflecting the effect after using the basic data batch to form a basic data batch verification parameter category set; combining the basic data batch verification parameter category set and the current time node set, collecting the verification parameters after using the current basic data batch to form a current use verification parameter matrix; collecting the average verification parameters of a plurality of time nodes before using the current basic data batch to form a historical average verification parameter set; calculating the difference information between the historical average verification parameter set and each row of data in the current use verification parameter matrix to form a current verification difference information set; setting a first verification difference threshold and a second verification difference threshold; when there is a current verification difference information less than the second verification difference threshold in the current verification difference information set, switching the current basic data batch back to the current original basic batch again; when there is a current verification difference information greater than or equal to the second verification difference threshold and less than the first verification difference threshold in the current verification difference information set, entering a subsequent processing step; otherwise, entering the next operation; predicting the use verification parameters of future time nodes according to the current use verification parameter matrix and using a feature extraction network model to obtain a future use verification parameter matrix; calculating the difference information between the historical average verification parameter set and each row of data in the future use verification parameter matrix to form a future verification difference information set; when there is a future verification difference information less than the second verification difference threshold in the future verification difference information set, entering a subsequent processing step; otherwise, no processing is performed.
[0021] Preferably, the method performs segmented processing on the current basic data batch according to the judgment result to obtain the final current hidden label segment number, which includes the following operations: setting an initial current hidden label segment number; performing segmented processing on the current basic data batch using the initial current hidden label segment number and storing; after the segmented processing and storage of the current basic data batch are completed, collecting the current verification parameters according to the basic data batch verification parameter category set to form a current segmented verification parameter set; calculating the difference information between the current segmented verification parameter set and the historical average verification parameter set to form a current segmented difference information; when the current segmented difference information is greater than or equal to the first verification difference threshold, taking the initial current hidden label segment number as the final current hidden label segment number; otherwise, adjusting the initial current hidden label segment number until the current segmented difference information is greater than or equal to the first verification difference threshold.
[0022] Preferably, the adjustment of the initial current hidden label segment number includes the following operations:
[0023] Define the initial range of values for the current number of hidden label segments to form the current range of segment number values; construct a federated node group for adjusting the number of hidden label segments; set the maximum number of iterations and the current number of iterations for the federated node group for adjusting the number of hidden label segments, denoted as the maximum number of iterations for segment adjustment and the current number of iterations for segment adjustment, respectively; set the initial position of each node in the federated node group for adjusting the number of hidden label segments according to the current range of values for segment number values to form a second initial position set; construct the fitness function for the federated node group for adjusting the number of hidden label segments; start the iteration, setting the current number of iterations for segment adjustment to 1 before each iteration;
[0024] In each iteration, the fitness function of the federated node group with the number of hidden label segments is used to calculate the fitness value of each node position in the federated node group with the number of hidden label segments updated in the previous iteration, and the position of each node in the federated node group with the number of hidden label segments updated in the previous iteration is updated. When the current iteration number of segment adjustment reaches the maximum iteration number of segment adjustment, the iteration stops, and the second final global best fitness and the second final global best position are obtained. Otherwise, the iteration continues until the current iteration number of segment adjustment reaches the maximum iteration number of segment adjustment. The second final global best fitness is used as the difference information after the current segmentation after optimization. When the difference information after the current segmentation after optimization is greater than or equal to the first verification difference threshold, the current basic data batch is segmented and stored using the second final global best position to obtain the final number of hidden label segments. Otherwise, the iteration steps are returned to continue the iteration until the difference information after the current segmentation after optimization is greater than or equal to the first verification difference threshold.
[0025] Preferably, the step of setting the educational data scenario to be processed and the corresponding number of data batches to form a candidate data batch set includes the following operations: performing noise reduction and standardization processing on the original educational data, removing outliers and unifying the data format, dividing the data batches based on course type and student level to form a candidate data batch set.
[0026] Preferably, the collection of probability information for each candidate data batch as a base batch within the historical statistical period includes the following operations: through the collaborative statistical mechanism among federated learning nodes, the selection records of each participant in using candidate data batches as base batches within the historical statistical period are summarized, and the proportion of the number of times each candidate data batch is selected to the total number of selections is calculated to obtain the probability information for each candidate data batch as a base batch.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] At the data processing level, by collecting time-series records of interaction frequencies, characteristic activity indicators across multiple time periods, and coverage information from multi-source educational data, a corresponding set is formed, enabling comprehensive and accurate acquisition of the characteristic information of educational data. This multi-dimensional data collection method ensures that the foundational data for subsequent processing has rich connotations, providing solid data support for the generation of hidden labels. For example, when setting the educational data scenarios to be processed and the corresponding candidate data batches, the original educational data undergoes denoising and standardization processing, and data batches are divided based on course type and student level, making the division of data batches more scientific and reasonable, and better reflecting the data characteristics under different educational scenarios.
[0029] Regarding the selection of the basic data batch, a final basic batch probability mapping model is constructed, and the current basic data batch is selected by combining various current datasets, thereby improving the accuracy and adaptability of the basic data batch selection. This model employs a federated aggregation optimization strategy, fully leveraging the advantages of federated learning. Without sharing the original data, a more representative model is constructed through the collaborative efforts of all participants. When the batch corresponding to the highest probability in the current candidate data batch's basic probability information set differs from the original basic batch, the basic data batch is promptly replaced to ensure that the selected basic data batch better adapts to the current data characteristics, laying a solid foundation for subsequent latent label generation.
[0030] In the verification and processing stages, by collecting verification parameters after using the current basic data batch, it is determined whether label generation or segmentation processing is needed, thus achieving dynamic monitoring and optimization of the data processing process. Different verification difference thresholds are set, and subsequent operations are determined based on the verification difference information, making the processing more intelligent and precise. For example, when the verification difference information falls within different ranges, different measures are taken, such as switching back to the original basic batch, proceeding to subsequent processing steps, or not processing at all, ensuring the effectiveness and efficiency of data processing.
[0031] Regarding segmentation and determining the number of hidden-label segments, a federated node group is constructed to adjust the number of hidden-label segments. An iterative optimization approach, combined with a fitness function, is used to adjust the initial number of segments until a difference threshold is met, thus obtaining the final number of hidden-label segments. This method can automatically optimize the number of segments based on the actual characteristics of the data, improving the accuracy and rationality of hidden label generation. Simultaneously, by utilizing a collaborative statistical mechanism among federated learning nodes, the selection records of each participant are aggregated, and the probability information of candidate data batches as base batches is statistically analyzed. This fully leverages the advantages of distributed data, improving the reliability and comprehensiveness of the probability information.
[0032] Furthermore, this invention fully utilizes federated learning technology, ensuring that all participants do not need to share raw data throughout the data processing process, but only model parameters or intermediate results, effectively protecting data privacy. This is particularly important for the education sector, as educational data involves a large amount of personal privacy information, complies with data privacy regulations, and provides a feasible solution for the sharing and collaborative processing of educational data.
[0033] This invention achieves efficient and accurate generation of implicit labels for educational data through multi-dimensional data collection, scientific model construction, dynamic verification processing, and the application of federated learning technology, while protecting data privacy and improving the efficiency and quality of educational data processing. Attached Figure Description
[0034] Figure 1 This is a schematic diagram illustrating the working principle of the educational data latent label generation method based on federated learning described in this invention.
[0035] Figure 2 Design diagram for the collection and aggregation of multi-source educational data;
[0036] Figure 3 The design drawing selected for the current batch of basic data;
[0037] Figure 4 Design diagram for the final basic batch probability mapping model. Detailed Implementation
[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0039] Please see Figures 1-4 This invention provides a method for generating implicit labels for educational data based on federated learning, and the specific implementation steps are as follows:
[0040] Collect time-series records of interaction frequency of multi-source educational data, characteristic activity indicators of multiple time periods, and coverage information to form a current time-series set of interaction frequency, a current set of characteristic activity indicators, and a current set of coverage information.
[0041] Collect probability information of each candidate data batch as the base batch, time-series records of interaction frequency of each candidate data batch, feature activity indicators of multiple time periods, and coverage information, and construct the final base batch probability mapping model; select the current base data batch based on the final base batch probability mapping model, the current time-series set of interaction frequency, the current set of feature activity indicators, and the current set of coverage information.
[0042] By collecting and using the verification parameters after the current batch of basic data is used, it is determined whether the current batch of basic data needs to be labeled or segmented, and the judgment result is obtained.
[0043] Based on the judgment results, the current batch of basic data is segmented to obtain the final number of current hidden label segments.
[0044] Example 1:
[0045] In this embodiment, the interaction frequency time-series records of multi-source educational data, characteristic activity indicators for multiple time periods, and coverage information are collected to form a current interaction frequency time-series set, a current characteristic activity indicator set, and a current coverage information set. The specific implementation method is as follows:
[0046] It is essential to clearly define the educational data scenario to be processed and the corresponding data batches, as this forms the foundation of the entire implementation process. In practice, raw educational data often contains a large amount of noise and outliers, and the data format may be inconsistent, which can affect subsequent analysis and processing. Therefore, it is necessary to perform denoising and standardization processing on the raw educational data. Denoising primarily involves removing outliers from the data, which may be caused by errors during data collection, equipment malfunctions, or human error. For example, in student homework completion time data, there may be obviously unreasonable maximum or minimum values, which need to be identified and removed. Standardization processing involves converting data of different formats and ranges into a unified format and range to facilitate comparison and analysis. For example, converting grade data from different courses into a percentage system.
[0047] After denoising and standardization, the data needs to be batched based on course type and student level. Course types can be divided according to subject category (e.g., Chinese, Mathematics, English) and course difficulty (e.g., basic courses, advanced courses); student levels can be divided according to factors such as grade level and learning ability. In this way, educational data can be divided into multiple different candidate data batches, each batch having similar characteristics and attributes, forming a set of candidate data batches. For example, for the mathematics education data of a certain school, it can be divided into a first-grade basic course data batch, a second-grade advanced course data batch, and so on.
[0048] Several parameter categories need to be defined to influence the likelihood of a data batch being used as a base batch, forming a set of parameter categories influencing base batch selection. These parameter categories need to reflect characteristics such as the importance, stability, and applicability of the data batch. For example, interaction frequency, feature activity indicators, and coverage can all serve as parameter categories. Interaction frequency can reflect the degree of interaction between students and educational data, such as the number of times students access learning resources and submit assignments; feature activity indicators can measure the activity level of various features in the data, such as the number of times a certain knowledge point is mentioned and the time students spend on a certain knowledge point; coverage can represent the scope of content covered by the data and the scope of the student group.
[0049] Set the first current statistical period. This period needs to be determined based on the specific educational data scenario and analytical needs; it can be a day, a week, a month, or a semester, etc. For example, if you need to analyze changes in students' learning behavior over a semester, then the first current statistical period can be set to a semester.
[0050] After determining the candidate data batch set, the set of basic batch selection influencing parameter categories, and the first current statistical period, it is necessary to combine the first current statistical period and the set of basic batch selection influencing parameter categories to obtain the interaction frequency time-series record, multi-time period characteristic activity indicators, and coverage information for each data batch. For the interaction frequency time-series record, the interaction frequency changes of each data batch within the first current statistical period need to be recorded chronologically. For example, the number of times a student accesses a certain data batch is recorded daily, forming a time series. For the multi-time period characteristic activity indicators, the first current statistical period needs to be divided into multiple time periods, and the characteristic activity indicators of the data batches need to be calculated within each time period. For example, a semester can be divided into multiple months, and the number of times a certain knowledge point is mentioned can be calculated monthly. For coverage information, the scope of course content and student group covered by each data batch within the first current statistical period needs to be statistically analyzed. For example, a data batch covers the first three units of a mathematics course and is used by all third-grade students.
[0051] Through the above steps, a time series set of current interaction frequency, a set of current active feature indicators, and a set of current coverage information can be formed. These sets contain detailed information about each data batch within the first current statistical period, providing an important basis for subsequent selection of basic data batches and generation of hidden labels. In practical applications, it is necessary to ensure the accuracy of data collection and processing to guarantee the accuracy and reliability of subsequent analysis. Simultaneously, it is also necessary to continuously adjust and optimize various parameters and steps according to actual conditions to adapt to different educational data scenarios and needs. For example, if a parameter category is found to not well reflect the characteristics of a data batch, the parameter category can be adjusted or changed in a timely manner; if the first current statistical period is set unreasonably, it can be adjusted according to the actual situation. Furthermore, attention must be paid to data security and privacy. When collecting and processing educational data, relevant laws, regulations, and ethical norms must be followed to ensure that students' personal information and learning data are not leaked or misused.
[0052] Example 2:
[0053] In this embodiment, probability information of each candidate data batch as a base batch, time-series records of the interaction frequency of each candidate data batch, feature activity indicators of multiple time periods, and coverage information are collected to construct the final base batch probability mapping model. The specific implementation method is as follows:
[0054] A historical statistical period needs to be set. The determination of this period should be based on the actual application scenario and analysis needs of the educational data. It can be a month, a semester, or a longer time span. For example, the past two semesters can be selected as the historical statistical period in order to obtain enough historical data to support the subsequent model building.
[0055] After setting the historical statistical period, relevant data within the historical statistical period needs to be collected, combining the existing set of candidate data batches and the set of influencing parameter categories for the basic batch selection. This data collection work covers multiple aspects: on the one hand, it is necessary to collect the probability information of each candidate data batch as a basic batch; on the other hand, it is also necessary to collect the time-series records of the interaction frequency of each candidate data batch, the characteristic activity indicators of multiple time periods, and the coverage information.
[0056] The collection of probability information for each candidate data batch as a base batch within a historical statistical period is accomplished through a collaborative statistical mechanism among federated learning nodes. In the federated learning architecture, each participant (such as different schools or educational institutions) maintains its own selection records for candidate data batches as base batches. These selection records, scattered across various nodes, are collected and organized through a collaborative aggregation method. For example, each participant uploads the number of times it selected different candidate data batches as base batches within the historical statistical period to a unified aggregation node. Then, at the aggregation node, the proportion of each candidate data batch selected out of the total number of selections is calculated; this proportion represents the probability information of each candidate data batch as a base batch. For instance, if there are a total of 1000 base batch selection records within the historical statistical period, and candidate data batch A was selected 200 times, then the probability information of candidate data batch A as a base batch is 20%.
[0057] When collecting time-series records of interaction frequencies for each candidate data batch, it is necessary to record the interaction frequency of each candidate data batch at different points in time within the historical statistical period, in chronological order. For example, using weeks as the time unit, record the weekly interaction frequency of each candidate data batch over the past two semesters to form a time-series data. The interaction frequency here can be measured by indicators such as the number of times students accessed or performed operations on the data batch.
[0058] For collecting characteristic activity indicators across multiple time periods, it is necessary to first divide the historical statistical period into multiple different time periods, and then calculate the corresponding characteristic activity indicators within each time period. For example, the historical statistical period (two semesters) can be divided into four academic segments, with each academic segment being a time period. Within each academic segment, the activity indicators for each feature in the candidate data batch can be calculated, such as the number of discussions on a certain knowledge point or the amount of time students spend learning on related content.
[0059] The collection of coverage information involves determining the scope of course content and student group covered by each candidate data batch within the historical statistical period. For example, if a candidate data batch covers the algebra part of the mathematics course and is used by all students in the second year of junior high school, this constitutes the coverage information for that candidate data batch.
[0060] Through the above data collection process, a set of basic probability information for historical candidate data batches, a set of historical interaction frequency time series, a set of historical feature activity indicators, and a set of historical coverage information can be formed. These sets store a large amount of historical data, providing rich material for the subsequent construction of the final basic batch probability mapping model.
[0061] These historical datasets are needed to construct the final basic batch probability mapping model and the final interaction frequency weight time series set. During the construction process, a federated aggregation optimization strategy is adopted. The core idea of this strategy is to aggregate and optimize the data from each participant within the framework of federated learning to construct a globally optimal model.
[0062] Each participant trains its model and updates its parameters based on its own local historical dataset. For example, each school uses its own historical candidate data batch probability information set and historical interaction frequency time series set to train a local version of a basic batch probability mapping model. Then, through federated aggregation, the local model parameters of each participant are summarized and merged to obtain a global model parameter set. In this aggregation process, the weights and contributions of different participants' data need to be considered to ensure that the aggregated model can fully reflect the data characteristics of each participant.
[0063] When constructing the final interaction frequency weight time series set, it is also necessary to determine the interaction frequency weights at different time points based on the historical interaction frequency time series set and the federated aggregation optimization strategy. For example, according to the degree of influence of the interaction frequency at different time points in historical data on the basic batch selection, a corresponding weight is assigned to the interaction frequency at each time point, forming a weight time series set.
[0064] By leveraging historical data and a federated aggregation optimization strategy, this approach constructs a mapping model that accurately reflects the probability of candidate data batches serving as base batches, along with a reasonable time series set of interaction frequency weights. These models and sets will play a crucial role in selecting the current base data batch, helping to accurately choose suitable base data batches based on the characteristics of the current educational data, providing a reliable foundation for subsequent hidden label generation. Throughout the implementation process, it is essential to ensure the accuracy and completeness of the data, guaranteeing that the collected historical data truly reflects the actual situation. Simultaneously, the security and privacy of the federated aggregation optimization strategy must be guaranteed, avoiding the leakage of sensitive information from all participants during data aggregation. Furthermore, adjustments and optimizations to the historical statistical period, data collection indicators, and methods are necessary based on actual circumstances to ensure that the constructed models and sets better adapt to different educational data scenarios and needs.
[0065] Example 3:
[0066] In this embodiment, the current basic data batch is selected based on the final basic batch probability mapping model, the current interaction frequency time series set, the current feature activity index set, and the current coverage information set. The specific implementation method is as follows:
[0067] It is necessary to define the current original base batch. This current original base batch can be an initial base data batch selected based on historical experience, default settings, or other predetermined rules. For example, in an educational data processing scenario, the candidate data batch that was used most frequently in the previous statistical period might be used as the current original base batch by default.
[0068] After setting up the current initial base batch, each piece of information from the current interaction frequency time series set, the current feature activity index set, and the current coverage information set, along with each piece of information from the final interaction frequency weight time series set, needs to be substituted into the final base batch probability mapping model for mapping. This substitution operation must be performed according to the input format and calculation logic specified by the model.
[0069] The current interaction frequency time series set contains sequence data showing the change in interaction frequency of each candidate data batch over time within the current statistical period. For example, for candidate data batch A, its current interaction frequency time series record might be a sequence of weekly access counts within the current statistical period. The current feature activity index set contains feature activity indicators for each candidate data batch across multiple time periods within the current statistical period, such as the number of knowledge point discussions each month. The current coverage information set records the course content and student group scope covered by each candidate data batch within the current statistical period.
[0070] The final interaction frequency weight time series set is determined when constructing the final base batch probability mapping model, and it assigns corresponding weights to the interaction frequencies at different time points. For example, it may be assumed that recent interaction frequencies have a greater impact on the selection of base batches, and therefore, higher weights are assigned to interaction frequencies at recent time points.
[0071] After substituting the information from these sets into the final basic batch probability mapping model, the model calculates the probability of each candidate data batch as a basic batch based on its internal algorithm and parameters, thus obtaining the basic probability information set of the current candidate data batch. This set contains the basic probability information of each candidate data batch under the current circumstances.
[0072] We need to analyze the set of basic probability information for the current candidate data batch to find the candidate data batch corresponding to the largest basic probability information. Then, we compare this candidate data batch with the current original basic batch to determine if they are the same.
[0073] If the candidate data batch corresponding to the largest basic probability information in the current candidate data batch basic probability information set is the same as the current original basic batch, it means that under the current data analysis, the original current basic batch is still the most suitable choice. Therefore, the current basic data batch will not be changed, and the current original basic batch will continue to be used as the current basic data batch.
[0074] If the candidate data batch corresponding to the largest basic probability information is different from the current original basic batch, then the current basic data batch needs to be replaced. In this case, the candidate data batch corresponding to the largest basic probability information in the current candidate data batch's basic probability information set is taken as the current basic data batch, and this new candidate data batch is recorded as the current basic data batch.
[0075] Throughout the process of selecting the current basic data batch, it is essential to ensure the accuracy and completeness of the data. For example, data such as the current interaction frequency time series set and the current feature activity index set must be accurately collected and processed; otherwise, it will affect the output of the final basic batch probability mapping model, leading to an unreasonable selection of the current basic data batch.
[0076] The accuracy of the final base batch probability mapping model is also crucial. This model is built based on historical data and a federated aggregation optimization strategy. Therefore, when building the model, it is necessary to fully consider the representativeness of the historical data and the rationality of the federated aggregation process to ensure that the model can accurately reflect the probability of candidate data batches as base batches.
[0077] In practical applications, it may be necessary to adjust and optimize the rules for setting the initial base batch and the parameters of the final base batch probability mapping model based on the specific educational data scenario and needs. For example, if it is found that the base data batch selected by the model is not reasonable in some cases, the model's performance can be improved by adjusting the model's parameters or adding new influencing factors.
[0078] It is also important to note that the process of selecting the current basic data batch is dynamic. As time goes by and new data is continuously collected, data such as the current interaction frequency time series set and the current feature activity index set will be constantly updated. Therefore, it is necessary to reselect the data periodically or when necessary to ensure that the current basic data batch is always the most suitable choice for the current educational data characteristics.
[0079] Throughout the implementation process, each step is logically interconnected. Setting the current original base batch is the starting point; mapping the data is the crucial computational step; comparison and judgment determine whether to change the base data batch; and ultimately, determining the current base data batch is the goal of the entire process. Only by executing each step accurately can we ensure that the selected current base data batch better meets the needs of generating implicit labels for educational data, providing a reliable foundation for subsequent label generation and segmentation processing.
[0080] Example 4:
[0081] In this embodiment, by collecting and using the verification parameters after using the current basic data batch, it is determined whether label generation or segmentation of the current basic data batch is required, and the determination result is obtained. The specific implementation method is as follows:
[0082] Establish a second current statistical period. This period needs to be determined based on the application scenario of the educational data, such as a one-month period, to collect data on the effectiveness of using the current batch of basic data. Within the second current statistical period, set several time nodes to form a current time node set. For example, divide a month into four time nodes, namely the last day of each week, to periodically collect and verify parameters.
[0083] Simultaneously, several types of validation parameters are established to reflect the effects of using the basic data batch, forming a set of validation parameter categories for the basic data batch. These parameters need to be related to the actual application goals of the educational data, such as students' homework completion rate, mastery of knowledge points, and learning time. Taking online mathematics courses as an example, validation parameters might include students' accuracy rate in testing knowledge points in the current basic data batch, the number of times they submit homework, and their level of interaction and participation on the learning platform.
[0084] By combining the set of validation parameter categories from the basic data batch with the current time point set, validation parameters are collected after using the current basic data batch, forming a matrix of currently used validation parameters. For example, data for each validation parameter is collected every Friday: in the first week, students' test accuracy rate is 75%, the number of submissions is 3, and the interaction participation is 20; in the second week, the accuracy rate is 80%, the number of submissions is 4, and the interaction participation is 25, and so on. This data is arranged into a matrix according to the time point and parameter category.
[0085] The average verification parameters at several time points prior to the use of the current batch of basic data are collected to form a historical average verification parameter set. For example, the average verification parameters for each week in the month prior to the use of the current batch of basic data are selected: the average test accuracy for the first four weeks was 70%, the average number of submissions was 2.5, and the average number of interactions was 18. These averages constitute the historical average verification parameter set.
[0086] The differences between the historical average set of validation parameters and each row of the currently used validation parameter matrix are calculated to form the current set of validation difference information. These differences can be calculated using methods such as absolute value difference or percentage difference. For example, in the first week, the difference between the test accuracy and the historical average is 75% - 70% = 5%, the difference in the number of submissions is 3 - 2.5 = 0.5 times, and the difference in interaction engagement is 20 - 18 = 2 times. In the second week, the difference in accuracy is 10%, the difference in the number of submissions is 1.5 times, and the difference in interaction engagement is 7 times. These differences constitute the current set of validation difference information.
[0087] Set a first verification difference threshold and a second verification difference threshold. For example, set the first verification difference threshold to 15% (to judge significant changes in effect) and the second verification difference threshold to 5% (to judge slight changes in effect). When there are verification difference information in the current verification difference information set that is less than the second verification difference threshold, it means that the effect after using the current base data batch is very similar to the historical situation. In this case, switch the current base data batch back to the current original base batch. For example, if the interaction participation difference in a certain week is 3 times, which is less than the 4 times corresponding to the second verification difference threshold (assuming the threshold unit is the number of times), then switch back to the original batch.
[0088] If any validation difference in the current validation difference set is greater than or equal to the second validation difference threshold but less than the first validation difference threshold, it indicates that the effect has changed to some extent but is not significant, and further processing steps are required. For example, if the test accuracy difference for a certain week is 8%, falling between 5% and 15%, then subsequent analysis is initiated. If all validation differences are greater than or equal to the first validation difference threshold, it indicates that the effect has changed significantly, and the process proceeds directly to the next step.
[0089] Based on the current validation parameter matrix, a feature extraction network model is used to predict the validation parameters for future time points, resulting in a future validation parameter matrix. The feature extraction network model can be based on a deep learning framework, such as LSTM or CNN, which learns the time-series features of historical data to predict future values. For example, using data such as test accuracy and submission count from the first four weeks, parameter values for the fifth and sixth weeks can be predicted, forming the future validation parameter matrix.
[0090] The differences between the historical average set of validation parameters and each row of the future validation parameter matrix are calculated to form the future validation difference information set. For example, if the predicted test accuracy for week 5 is 85%, the difference from the historical average of 70% is 15%, and the predicted number of submissions is 5, the difference is 2.5. These predicted differences constitute the future validation difference information set.
[0091] If any future validation difference in the set of future validation difference information is less than the second validation difference threshold, it indicates that the predicted future effect is very close to the historical situation, and further processing is required. If none of the future validation difference information is less than the second validation difference threshold, no processing is required. For example, if the predicted interaction participation difference for week five is 4 times, which is equal to the second validation difference threshold, further analysis is needed to determine whether the base data batch needs to be adjusted.
[0092] Throughout the implementation process, it is crucial to ensure the timeliness and accuracy of validation parameter collection. For example, this can be achieved by using an automated data collection system on the education platform to record student behavior data in real time, avoiding errors from manual data entry. The training of the feature extraction network model must be based on sufficient historical data to guarantee the reliability of predictions. Furthermore, the setting of validation difference thresholds should be tailored to the actual needs of the educational scenario. For instance, in scenarios sensitive to teaching effectiveness, the first threshold can be set to 10% to detect adaptive changes in data batches earlier. If, during multiple assessments, the current basic data batch frequently triggers switching or segmentation, the construction of the basic batch probability mapping model needs to be re-examined for rationality, or the selection of validation parameter categories needs to be adjusted to ensure that the assessment results accurately reflect the actual effect of the data batch.
[0093] Example 5:
[0094] In this embodiment, the current basic data batch is segmented according to the judgment result to obtain the final number of current hidden label segments. The specific implementation method is as follows:
[0095] Set the initial number of hidden label segments. This number should be set based on the scale and characteristics of the educational data. For example, for a current batch of basic data containing 1,000 records, the initial number of segments can be set to 5, that is, the data is divided into 5 equal segments, each containing approximately 200 records.
[0096] The current batch of basic data is segmented and stored using the initial number of segments with hidden labels. For example, student homework data for a math course is divided into 5 segments in chronological order, with each segment corresponding to one week's homework submission records. The segmented data is stored in different files or database tables. After segmentation, the current verification parameters are collected based on the set of verification parameter categories for the basic data batch, forming the current set of verification parameters after segmentation. For example, verification parameters include the students' homework accuracy rate, completion time, and error type distribution in each segment. The accuracy rates collected for each segment are 70%, 75%, 80%, 65%, and 72%, and the average completion times are 25 minutes, 30 minutes, 28 minutes, 35 minutes, and 27 minutes, respectively, constituting the set of verification parameters.
[0097] The difference between the current set of validation parameters and the historical average set of validation parameters is calculated to form the difference information after the current segmentation. This difference information can be calculated in various ways, such as calculating the average difference for each type of parameter. Assuming the historical average accuracy is 73%, the average accuracy after the current segmentation is (70% + 75% + 80% + 65% + 72%) / 5 = 72.4%, which differs from the historical average by 73% - 72.4% = 0.6%. The historical average completion time is 29 minutes, and the average time after the current segmentation is (25 + 30 + 28 + 35 + 27) / 5 = 29 minutes, with a difference of 0. These difference values constitute the difference information after the current segmentation.
[0098] When the difference information after the current segmentation is greater than or equal to the first verification difference threshold, it means that the current number of segments has met the requirements, and the initial number of current hidden label segments is used as the final number of current hidden label segments. For example, if the first verification difference threshold is set to 1%, and the current accuracy difference is 0.6% and the completion time difference is 0, both of which are less than the threshold, then the initial number of current hidden label segments needs to be adjusted.
[0099] When adjusting the initial number of hidden-label segments, first define the range of values for the initial number of hidden-label segments, thus forming the range of values for the current number of segments. For example, considering the data scale, the range can be set to 3 to 10, meaning the number of segments can be adjusted between 3 and 10. Construct a federated node group for adjusting the number of hidden-label segments. This group consists of multiple participants, such as educational data processing nodes from different schools. Each node is responsible for exploring the optimal number of segments within its own value range.
[0100] Set the maximum and current iteration counts for adjusting the hidden label segment number of the federated node group, denoted as the maximum and current iteration counts, respectively. For example, the maximum iteration count is set to 20, and the current iteration count is initialized to 1. Set the initial position of each node based on the current segment number range, forming a second set of initial positions. For example, the initial positions of 5 nodes are set to 3, 5, 7, 8, and 10, corresponding to different segment number assumptions.
[0101] A fitness function is constructed to adjust the number of hidden-label segments in the federated node group. This function measures the quality of each node position (i.e., the number of segments). The fitness function design must be based on the difference information after the current segmentation. For example, the fitness value can be the reciprocal of the sum of the absolute values of the differences; the smaller the difference, the lower the fitness value, and vice versa. For instance, when the number of segments is 5, the sum of the absolute values of the differences is 0.6% + 0 = 0.6%, and the fitness value is 1 / 0.6% ≈ 166.67.
[0102] The iteration begins, with the segmentation adjustment and current iteration count set to 1 before each iteration. In each iteration, the fitness function is used to calculate the fitness value of each node position after the previous iteration, and the node position is updated accordingly. For example, based on a comparison between the current fitness value and the global best fitness value, the number of segments in the next round is adjusted according to a preset update rule (such as the velocity-position update formula in particle swarm optimization). If a node's initial position is 3, and its fitness value is lower than the global best, it moves towards the global best position (e.g., 7), and its position in the next round is set to 4 or 5.
[0103] When the number of iterations for segment adjustment reaches the maximum number of iterations, the iteration stops, and the second final global optimal fitness and the second final global optimal position are obtained. For example, after 20 iterations, the number of segments corresponding to the global optimal position is 8, and the fitness value is the highest at this time, indicating that the difference information under this number of segments is closest to or greater than the first verification difference threshold.
[0104] The second final global optimal fitness is used as the difference information after optimization and segmentation. It is then determined whether this difference is greater than or equal to the first verification difference threshold. For example, if the calculated difference information is 1.2% when the number of segments is 8, which is greater than the set threshold of 1%, then the second final global optimal position (i.e., 8 segments) is used to segment the current batch of basic data and store it, resulting in a final current hidden label segment number of 8. If the optimized difference information is still less than the threshold, such as 0.8% when the number of segments is 8, then the iteration step is returned to continue adjustment until a segment number that meets the conditions is found.
[0105] In practical applications, for example, processing English vocabulary learning data from a first-year high school student, the current batch of basic data contains 800 student vocabulary test records. The initial number of segments is set to 4, with 200 records per segment. Verification parameters after segmentation (such as the accuracy rate of each segment, the number of times incorrect words are repeated, etc.) are collected, and the difference from the historical average parameters is calculated. If the difference is less than a threshold, the number of segments is iteratively adjusted through the federated node group. Assuming that in the 15th iteration, when the node group determines the number of segments to be 6, the difference reaches 1.1%, meeting the threshold requirement, the final number of segments is determined to be 6, dividing the 800 records into 6 segments, with approximately 133 records per segment, to generate more refined hidden labels reflecting the students' vocabulary mastery characteristics at different learning stages. Throughout the process, it is necessary to ensure the communication security of the federated node group to avoid data leakage, and to periodically readjust the number of segments according to the data update frequency to adapt to the dynamic changes in students' learning behavior.
[0106] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0107] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for generating implicit labels for educational data based on federated learning, characterized in that, Includes the following operations: Collect time-series records of interaction frequency, feature activity indicators and coverage information of multi-source educational data to form a current time-series set of interaction frequency, a current set of feature activity indicators and a current set of coverage information. Collect probability information of each candidate data batch as the base batch, time-series records of the interaction frequency of each candidate data batch, feature activity indicators of multiple time periods, and coverage information, and construct the final base batch probability mapping model. The current basic data batch is selected based on the final basic batch probability mapping model, the current interaction frequency time series set, the current feature activity index set, and the current coverage information set. By collecting and using the verification parameters after the current batch of basic data is used, it is determined whether the current batch of basic data needs to be labeled or segmented, and the judgment result is obtained. Based on the judgment results, the current batch of basic data is segmented to obtain the final number of current hidden label segments.
2. The method for generating implicit labels for educational data based on federated learning according to claim 1, characterized in that, The process of collecting interaction frequency time-series records, multi-time period characteristic activity indicators, and coverage information from multi-source educational data to form the current interaction frequency time-series set, the current characteristic activity indicator set, and the current coverage information set includes the following operations: setting the educational data scenario to be processed and several corresponding data batches to form a candidate data batch set; setting several parameter categories that affect the likelihood of a data batch being used as a basic batch to form a basic batch selection parameter category set; setting a first current statistical period; and combining the first current statistical period and the basic batch selection parameter category set to obtain the interaction frequency time-series records, multi-time period characteristic activity indicators, and coverage information for each data batch in the candidate data batch set, thus forming the current interaction frequency time-series set, the current characteristic activity indicator set, and the current coverage information set.
3. The method for generating implicit labels for educational data based on federated learning according to claim 2, characterized in that, The process of collecting probability information of each candidate data batch as the base batch, time-series records of the interaction frequency of each candidate data batch, feature activity indicators of multiple time periods, and coverage information, and constructing the final base batch probability mapping model includes the following operations: A historical statistical period is set; combining the candidate data batch set and the set of influencing parameter categories for the basic batch selection, the probability information of each candidate data batch as the basic batch, the time series records of the interaction frequency of each candidate data batch, the feature activity indicators of multiple time periods, and the coverage information are collected within the historical statistical period to form a set of basic probability information of historical candidate data batches, a set of time series records of historical interaction frequency, a set of time series records of historical feature activity indicators, and a set of time series records of historical coverage information; the final basic batch probability mapping model and the final interaction frequency weight time series set are constructed using the set of basic probability information of historical candidate data batches, the set of time series records of historical interaction frequency, the set of time series records of historical feature activity indicators, and the set of time series records of historical coverage information.
4. The method for generating implicit labels for educational data based on federated learning according to claim 3, characterized in that, The final basic batch probability mapping model and the final interaction frequency weight time series set are constructed using a federated aggregation optimization strategy.
5. The method for generating implicit labels for educational data based on federated learning according to claim 4, characterized in that, The process of selecting the current basic data batch based on the final basic batch probability mapping model, the current interaction frequency time series set, the current feature activity index set, and the current coverage information set includes the following operations: Define the current original base batch; substitute each piece of information from the current interaction frequency time series set, the current feature activity index set, the current coverage information set, and the final interaction frequency weight time series set into the final base batch probability mapping model for mapping, to obtain the current candidate data batch base probability information set; when the candidate data batch corresponding to the largest base probability information in the current candidate data batch base probability information set is the same as the current original base batch, the current base data batch is not changed; otherwise, the candidate data batch corresponding to the largest base probability information in the current candidate data batch base probability information set is taken as the current base data batch, and all are recorded as the current base data batch.
6. The method for generating implicit labels for educational data based on federated learning according to claim 5, characterized in that, The process of determining whether label generation or segmentation of the current basic data batch is necessary by collecting and using verification parameters after the current basic data batch has been used, and obtaining the determination result includes the following operations: Set a second current statistical period; within the second current statistical period, set several time nodes to form a current time node set; set several types of verification parameters that can reflect the effect after using the basic data batch to form a basic data batch verification parameter category set; combine the basic data batch verification parameter category set and the current time node set to collect verification parameters after using the current basic data batch to form a current usage verification parameter matrix; then collect the average verification parameters of several time nodes before using the current basic data batch to form a historical average verification parameter set; calculate the difference information between the historical average verification parameter set and each row of data in the current usage verification parameter matrix to form a current verification difference information set; set a first verification difference threshold and a second verification difference threshold; when there is a current verification difference information in the current verification difference information set that is less than the second verification difference threshold, switch the current basic data batch back to the current original basic batch; when there is a current verification difference information in the current verification difference information set that is greater than or equal to the second verification difference threshold and less than the first verification difference threshold, proceed to the next processing step; Otherwise, proceed to the next step; based on the current verification parameter matrix and using a feature extraction network model, predict the verification parameters for future time nodes to obtain the future verification parameter matrix; then calculate the difference information between the historical average verification parameter set and each row of data in the future verification parameter matrix to form the future verification difference information set; if the future verification difference information set contains future verification difference information less than the second verification difference threshold, proceed to the subsequent processing steps; otherwise, no processing is performed.
7. The method for generating implicit labels for educational data based on federated learning according to claim 6, characterized in that: The step of segmenting the current basic data batch according to the judgment result to obtain the final number of current hidden label segments includes the following operations: setting an initial number of current hidden label segments; segmenting the current basic data batch using the initial number of current hidden label segments and storing the data; After the current batch of basic data is segmented and stored, the current verification parameters are collected according to the set of verification parameter categories of the basic data batch, forming the current segmented verification parameter set; the difference information between the current segmented verification parameter set and the historical average verification parameter set is calculated, forming the current segmented difference information; when the current segmented difference information is greater than or equal to the first verification difference threshold, the initial current hidden label segment number is used as the final current hidden label segment number; otherwise, the initial current hidden label segment number is adjusted until the current segmented difference information is greater than or equal to the first verification difference threshold.
8. The method for generating implicit labels for educational data based on federated learning according to claim 7, characterized in that, Adjusting the initial number of currently hidden-label segments involves the following operations: Set the initial range of values for the current number of hidden label segments to form the range of values for the current number of segments; construct a federated node group for adjusting the number of hidden label segments; set the maximum number of iterations and the current number of iterations for the federated node group for adjusting the number of hidden label segments, denoted as the maximum number of iterations for segment adjustment and the current number of iterations for segment adjustment, respectively; set the initial position of each node in the federated node group for adjusting the number of hidden label segments according to the range of values for the current number of segments to form a second set of initial positions; Construct a fitness function for adjusting the number of hidden-label segments in the federated node group; begin iteration, setting the current iteration number of segment adjustment to 1 before each iteration. In each iteration, the fitness function of the federated node group with the number of hidden label segments is used to calculate the fitness value of each node position in the federated node group with the number of hidden label segments updated in the previous iteration, and the position of each node in the federated node group with the number of hidden label segments updated in the previous iteration is updated. When the current iteration number of segment adjustment reaches the maximum iteration number of segment adjustment, the iteration stops, and the second final global best fitness and the second final global best position are obtained; otherwise, the iteration continues until the current iteration number of segment adjustment reaches the maximum iteration number of segment adjustment. The second final global best fitness is used as the difference information after the current segmentation after optimization. When the difference information after the current segmentation after optimization is greater than or equal to the first verification difference threshold, the current basic data batch is segmented and stored using the second final global best position to obtain the final number of current hidden label segments. Otherwise, the iteration step is returned to continue the iteration until the difference information after the current segmentation after optimization is greater than or equal to the first verification difference threshold.
9. The method for generating implicit labels for educational data based on federated learning according to claim 2, characterized in that, The process of setting up the educational data scenario to be processed and the corresponding data batches to form a candidate data batch set includes the following operations: performing noise reduction and standardization processing on the original educational data, removing outliers and unifying the data format, dividing the data batches based on course type and student level to form a candidate data batch set.
10. The method for generating implicit labels for educational data based on federated learning according to claim 3, characterized in that, The process of collecting the probability information of each candidate data batch as a base batch within the historical statistical period includes the following operations: through the collaborative statistical mechanism among federated learning nodes, the selection records of each participant in using candidate data batches as base batches within the historical statistical period are summarized, and the proportion of the number of times each candidate data batch is selected to the total number of selections is calculated to obtain the probability information of each candidate data batch as a base batch.