A synthetic data generation and compliance sharing method for education research and an apparatus thereof
By using multimodal seed data parsing and conditional generation mechanisms, combined with privacy protection technologies, synthetic data that meets the needs of education and scientific research is generated. This solves the problems of insufficient data form coverage, statistical consistency, and privacy risks in existing technologies, and achieves high-fidelity, multimodal, and privacy-compliant data sharing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA NORMAL UNIV
- Filing Date
- 2026-03-11
- Publication Date
- 2026-05-29
Smart Images

Figure CN122113170A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of educational research data engineering and privacy-preserving data processing technology, specifically a method and apparatus for generating and compliantly sharing synthetic data for educational research in pedagogy and sociology of education. Background Technology
[0002] With the rapid development of Large Language Models (LLM) and multi-agent frameworks in the field of educational technology, research on areas such as "classroom teaching process simulation, teaching dialogue interaction, learning process recording and analysis, and evaluation of teaching intervention strategies" is constantly emerging. Existing research proposes using LLM-driven multi-role agents to construct classroom scenarios, forming interactive data and behavioral codes that can be used for teaching analysis in teacher-student and student-student interaction links, and evaluating the interaction process through educational analysis frameworks and comparative experiments. Meanwhile, reviews and policy guidance on the application of LLM / generative AI in education and research generally point out that while generative AI can support content generation, learning support, and educational research activities, it still faces risks and challenges requiring careful governance regarding reliability, bias, fairness, and data privacy. Furthermore, in the field of educational data mining and learning analysis, due to the need for privacy compliance and open sharing, there is an increasing use of statistical models and deep generative models, such as GANs / VAEs, to synthesize and evaluate educational data, especially tabular student records, forming a practical approach represented by open-source toolchains and benchmark evaluations.
[0003] Against the backdrop of the aforementioned research and applications, educational researchers, especially those in the fields of pedagogy, sociology of education, and learning analytics, often require a large amount of controllable, reproducible, and scalable data to support different types of research designs and empirical analyses. This data requirement is not limited to "student answers or homework texts," but also commonly includes teacher interviews, classroom observation records, surveys on after-school learning time and time usage, records of home-school communication, and data on school governance and resource allocation—a variety of data sources. These data include both structured variables and a significant proportion of textual narratives. Typical research needs can be summarized as follows: 1) Educational intervention and mechanism testing research: Under similar sample structures and contextual constraints, comparing different teaching interventions, classroom organization methods, feedback strategies, or resource input methods to observe the changing patterns of key variables, supporting hierarchical models, quasi-experiments, or causal inference analyses. 2) Training and Evaluation of Educational Data Mining Models: Constructing large-scale, reproducible datasets for training, parameter tuning, and robustness evaluation of learning prediction, learning analytics, and educational models. In this area, existing work has used open-source synthetic data frameworks and generative models to synthesize and benchmark student data, and discussed its fidelity and downstream predictive utility. 3) Data Preparation for Educational Sociology and Qualitative / Mixed Research: Data such as teacher interviews, classroom interactions, and time usage logs require richer coverage of group differences and contexts to support correlation analysis of factors such as educational opportunities, resource allocation, family background, and school environment; however, real-world data is often difficult to share or reuse due to privacy and ethical constraints.
[0004] Currently, the synthesis, generation, statistical distribution alignment, and privacy compliance rectification of educational research data generally employ LLM-based educational interaction and classroom simulation technologies, educational tabular data synthesis and open-source toolchains, and privacy protection and synthetic data publishing technologies. A common technical approach in LLM educational applications involves using LLM to drive single or multi-agent agents to generate teaching dialogues, classroom interaction logs, or teaching feedback texts under predefined classroom roles such as teachers, students, and teaching assistants, and task script constraints. These texts are then used for teaching analysis or educational experimental prototype verification. In related research, multi-agent classroom simulation frameworks typically include: role definition and goal setting, dialogue / action generation, interaction flow control, and analysis and evaluation of generated interactions. Educational analysis frameworks can be incorporated to label and statistically analyze interaction types and teaching processes. This type of technology can quickly generate a certain scale of interactive texts and process records without requiring real classroom participation, facilitating preliminary comparison and prototype verification of teaching strategies. However, its output is often mainly dialogue text, and the constraints on the sample structure and statistical consistency at the dataset level are relatively weak, making it difficult to directly meet the requirements of "hierarchical structure, joint distribution, sampling ratio and long-tail group coverage" in educational social science research.
[0005] In the field of educational data mining and learning analytics, existing research has employed statistical and deep generative models, such as Copula, CTGAN, and TVAE, to synthesize student tabular data, addressing privacy protection and data sharing. Their usability is evaluated through fidelity metrics and downstream task performance. Other work utilizes open-source ecosystems, such as Synthetic Data Vault (SDV), to build synthesis and evaluation workflows, benchmarking various generators and providing reproducible pipelines. These methods and toolchains can systematically process structured tabular data, balancing statistical similarity and predictive utility to some extent, and facilitating experimental reproducibility. However, educational research data often contains substantial textual narratives and cross-level relationships (individual-class-school-region), while existing tabular synthesis methods primarily target structured fields, failing to cover data formats such as interviews, observation records, and classroom interactions. Furthermore, when research objectives emphasize group differences and long-tail scenarios, the generated data may suffer from class scarcity, joint distribution bias, or insufficient characterization of minority groups.
[0006] Technologies related to privacy protection and synthetic data publishing: Addressing the core constraint of "data non-sharing," Differential Privacy (DP), a classic formal privacy framework, proposes introducing random noise into the query or learning process to limit the impact of adding / deleting individual samples on the output, thereby reducing the risk of leakage. Related work has systematically demonstrated the relationship between noise calibration and sensitivity. Furthermore, PATE (Private Aggregation of Teacher Ensembles) proposes using noisy voting from multiple teacher models to provide supervision signals to student models, achieving privacy protection without directly disclosing training data, and has been widely discussed in subsequent research. DP and PATE provide a clear theoretical framework and engineering path for privacy risk control, and are of significant reference value in handling sensitive data, especially educational data which is often highly sensitive, particularly involving minors. However, in the multi-source and heterogeneous forms of educational and research data, especially long text narratives and cross-level related data, there is often a trade-off between privacy risk assessment and utility preservation. Moreover, different research tasks have significantly different requirements for "usability," leading to the need for more detailed task adaptation and careful evaluation of general privacy mechanisms in educational and research sharing.
[0007] Existing technologies offer multiple feasible paths for the education field to "generate data at low cost, support teaching interaction analysis and preliminary experimental reproduction, and alleviate privacy restrictions to some extent." For example, LLM and multi-agent frameworks can quickly generate classroom interaction and teaching dialogue data, supporting teaching process analysis and simulation experiments. Open-source ecosystems such as SDV, combined with models such as CTGAN / TVAE, have formed relatively mature table synthesis and evaluation processes, facilitating the generation of structured student data and benchmark comparisons. Privacy protection theories such as DP and PATE provide important methodological basis for the sharing and publication of sensitive data. However, based on the real data needs of educational and sociological research, existing technologies still have some objective shortcomings, which may limit their direct applicability in the construction and sharing of research data. These shortcomings are mainly reflected in the following aspects: 1) Insufficient coverage of data forms: Many solutions focus on single or few data forms such as "student records / answers / dialogue texts", while educational research often requires the collaborative use of multi-source heterogeneous data such as interviews, observations, time usage, school governance, and resource allocation; 2) Insufficient constraints on sample structure and statistical consistency: Educational social science research often focuses on hierarchical structures and variable relationships (such as individual-class-school-region). Generating only "seemingly reasonable" samples is not enough to guarantee the joint distribution and consistency of group differences required for the research; 3) Limited coverage of long-tail groups and contexts: When the research involves minority groups or rare contexts (long-tail samples), the existing generation process may be more likely to cluster in high-frequency patterns, resulting in insufficient characterization of long-tail phenomena; 4) Privacy and bias risks still need to be assessed: Although there are theoretical and methodological foundations such as DP / PATE, privacy risk measurement, bias monitoring, and auditable records remain objective challenges in the textual narratives and complex relational data of educational research. Therefore, for educational research, especially in educational sociology, learning analytics, and blended research scenarios, how to obtain data resources that meet the usage habits and statistical analysis requirements of research personnel remains one of the ongoing concerns in this field, especially when it is inconvenient to directly share real data.
[0008] In summary, data-driven empirical research is becoming increasingly important in the field of educational science research. Existing technologies have the following significant defects and shortcomings in data acquisition, generation and sharing: 1) The contradiction between "data silos" and privacy compliance: Educational data involves sensitive information such as the privacy of minors, family background and behavioral evaluation, which are subject to extremely strict legal and regulatory restrictions. This makes it difficult for high-quality scientific research data to circulate among different institutions or researchers, which seriously hinders the reproduction of scientific research and the improvement of algorithms.
[0009] 2) Synthetic data is modally singular and lacks semantic consistency: Existing synthetic data generation techniques are mostly limited to statistical simulation of structured tabular data, making it difficult to generate unstructured data crucial for educational research, such as classroom dialogues, interview texts, or simply large-scale model text generation. The lack of statistical distribution constraints leads to a logical disconnect between the generated text and background variables, such as ability level and socioeconomic status. 3) Insufficient coverage of long-tail groups and rare scenarios: In real-world educational data, long-tail group samples, such as those with special learning disabilities and extreme family backgrounds, are often scarce. Conventional generation methods are prone to "pattern collapse," causing the synthetic data to lose crucial diversity. 4) Lack of process data: Education is a dynamic process, and existing methods struggle to generate "process data" with time-series logic and interactive causal chains, such as multi-round questioning and teaching interaction flows. Summary of the Invention
[0010] The purpose of this invention is to address the shortcomings of existing technologies by providing a method and apparatus for generating and compliantly sharing synthetic data in educational research. Employing a standardized process, this method transforms multimodal seed data into high-fidelity and privacy-free data while strictly ensuring privacy compliance, generating a synthetic dataset that combines statistical fidelity, semantic rationality, and modal richness. This method performs variable-level structured analysis on a small amount of educational research seed data, combining conditional and scenario-based generation mechanisms to form a dataset suitable for statistical analysis and experimental reproduction in educational research. Furthermore, it optimizes the alignment of population distribution and variable association structures through statistical discrimination and distribution distance measurement. It further incorporates privacy mechanisms such as differential privacy or teacher set theory (PATE) and fairness rectification strategies to reduce the risk of re-identification and suppress bias amplification, thereby enabling data production and compliant delivery for researchers conducting educational research. This invention is applicable to common multi-source data formats in educational research, including but not limited to questionnaire surveys and scales, teacher / student / parent interview texts, classroom interaction and observation records, after-school learning time and time usage logs, and school-class-region level educational resource and management statistics, demonstrating broad application prospects.
[0011] The specific technical solution for achieving the purpose of this invention is: a method for generating and compliantly sharing synthetic data for education and scientific research, characterized by the following steps:
[0012] S1: Seed Data Acquisition and Metadata Processing
[0013] Acquire educational and scientific research seed data and its metadata, clean the seed data, unify its format and manage its version to obtain a standardized seed dataset. The seed data includes structured data or unstructured data.
[0014] S2: Seed Analysis and Variable Representation Construction
[0015] The standardized seed dataset is parsed to construct a variable representation space that maps the original field features. The variable representation space includes at least a sociological attribute vector for representing group and environmental attributes, and a cognitive behavioral state vector for representing individual states and abilities.
[0016] S3: Generate Task Configuration
[0017] Receive the generation task configuration, which determines the target data type, the subject and scene to be generated, and the target statistical distribution.
[0018] S4: Conditional Generation and Sample Expansion
[0019] Based on the variable representation space and generation task configuration, a generation context containing role and scene constraints is constructed to drive the large model to generate synthetic samples. For each sample to be generated, at least one core research variable is fixed and at least one non-core attribute variable is subject to controllable perturbation to expand sample diversity while maintaining the semantic consistency of research variables.
[0020] S5: Domain Knowledge Enhancement and Process Generation
[0021] For the target data type, an enhanced generative context is constructed by combining knowledge from the fields of education or psychology, driving the model to generate synthetic data containing procedural records or interaction details.
[0022] S6: Distribution Assessment
[0023] Calculate key statistics for the synthetic data and compare the key statistics with the target statistical distribution to obtain a measure of distribution difference.
[0024] S7: Feedback-based policy optimization
[0025] The distribution difference metric is converted into a feedback-based optimization signal, and the generation strategy is iteratively updated until the statistical characteristics of the synthesized data meet the preset distribution alignment conditions.
[0026] S8: Compliance Processing and Delivery
[0027] The optimized synthetic data undergoes privacy protection processing and fairness rectification, and is then packaged into a shareable dataset for output.
[0028] In step S1, the seed data collection covers typical modalities of educational research. Structured data specifically includes questionnaire survey data, scale assessment data, or hierarchical statistical data; unstructured data specifically includes interview texts, classroom observation records, open-ended question responses, or teaching logs. The metadata not only includes basic definitions but also further encompasses: 1) Research design information: such as cross-sectional studies, longitudinal longitudinal studies, panel structures, or quasi-experimental / controlled designs; 2) Sampling structure information: such as stratified variables, cluster variables, sample weights, post-stratified weights, time point markers, and sampling frame descriptions; 3) Measurement tool information: such as questionnaire item versions, scale dimensional structures and reverse question rules, interview outline versions, etc. Furthermore, this invention establishes a traceable identifier for each record, including data source, collection batch, and processing version, to support subsequent research replication and auditing.
[0029] In the variable representation construction of step S2, in order to ensure the logical consistency of the generated data, multi-dimensional constraint modeling was implemented when constructing the variable representation space: 1) Field type constraint: clearly distinguish the generation logic of continuous, categorical, ordinal, count, date / time, or text fields; 2) Hierarchical consistency constraint: establish nested relationships of individual-class-school-region to ensure consistent foreign key associations and consistent time sequence; 3) Response style modeling: for questionnaire data, establish response style parameters including agreement tendency, extreme reaction tendency, social expectation tendency, or random response tendency; 4) Text style modeling: for text data, establish narrative style parameters including register formality, emotional intensity, and level of detail, as well as identity parameters that distinguish the discourse characteristics of teachers, parents, students, or administrators.
[0030] In the generation task configuration of step S3, the generation task configuration not only specifies the objective, but also establishes fine-grained control rules: 1) A subject-scenario-data type mapping table is established, stipulating that the same virtual subject should share the same identity profile when generating data in different scenarios (such as classroom, home) to maintain personality consistency; 2) The source of the target distribution is clarified, which can be an estimation based on seed data, a public statistical summary, or a hypothetical distribution specified by the researcher; 3) A multi-dimensional set of evaluation indicators is defined, which, in addition to the conventional distribution, particularly emphasizes the long-tail coverage constraint, namely the requirements for the proportion of rare groups, the coverage rate of extreme quantiles, or the coverage rate of rare situations; 4) Variables are explicitly divided into a "core variable set" that must be strictly maintained and a "non-core variable set" that can be changed.
[0031] In step S4, the conditional generation process includes the following steps: First, selecting or instantiating a prompt template containing role settings, scene settings, variable constraints, and output structure. Second, sampling from the variable representation space and implementing a controllable perturbation strategy, specifically perturbing demographic attributes, scene details, narrative style, or missing patterns, thereby expanding sample diversity while keeping core research variables fixed. Then, driving the large model to simultaneously output structured fields and unstructured text. Finally, performing consistency checks (checking field validity, logical consistency, etc.) and generating sample-level meta-information containing subject identifiers, generation condition summaries, and round information for samples that pass the checks, and storing this information in the database.
[0032] In the domain knowledge enhancement step S5, the following mechanisms are specifically adopted to improve the professionalism and depth of the generated content: 1) Retrieval Enhancement (RAG): Retrieve behavioral characteristics, narrative themes, or typical event patterns from the education / psychology knowledge base based on target tags and inject them into the generation context; 2) Multi-agent interaction closed loop: Construct virtual agent pairs between researchers and interviewees, teachers and students, or parents and teachers, and generate materials containing a question-response structure through multiple rounds of interaction; 3) Process recording: Output process records containing turn numbers, speakers, triggering conditions, and timestamps to form high-dimensional data aligned with "text-process log-structured variables".
[0033] In the distribution assessment of step S6, the assessment system covers multiple levels: 1) marginal distribution and stratification ratio statistics at the individual, class, and school levels; 2) stratified joint distribution statistics of the key variable set; 3) cross-level (e.g., student and school) correlation structure statistics; 4) structured quality statistics of text data (e.g., average number of turns, topic coverage). The difference between the synthetic data and the target distribution is quantified by calculating KL divergence, Wasserstein distance, or maximum mean difference (MMD).
[0034] In the policy optimization step S7, the optimization process forms a closed-loop feedback loop, constructing an optimization signal that includes negative distribution difference values and long-tail coverage rewards; determining the policy to be updated (cue policy or sampling policy); iteratively updating the policy using a reinforcement learning algorithm (such as Proximal Policy Optimization PPO); until termination conditions such as distribution difference falling below a threshold or quality report convergence are met. Additionally, stability constraints can optionally be applied to cross-scene profile drift.
[0035] In the compliance processing of step S8, strict post-processing is implemented to ensure data security and fairness: 1) Privacy protection: Configure and implement differential privacy mechanism (injecting noise) or teacher set privacy mechanism (PATE) to ensure that the output meets the privacy budget; 2) Fairness rectification: Monitor the association between sensitive attributes (such as gender) and outcome variables, and trigger counterfactual generation (keeping sensitive attributes unchanged and resampling non-sensitive conditions) to balance the bias when the threshold is exceeded; 3) Encapsulation and delivery: The final output is a delivery package containing a synthetic data ontology, variable dictionary, sampling description, generated version information and compliance summary (including privacy risk assessment and bias check).
[0036] Furthermore, the present invention also provides a synthetic data generation and compliant sharing device for education and scientific research, characterized in that the device includes: a processor, a memory, and a computer program stored in the memory, wherein when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0037] Compared with the prior art, the present invention has the following beneficial technical effects and significant technical progress:
[0038] 1) Breaking down data barriers and achieving zero-risk sharing: By generating synthetic data that is statistically similar to real data but completely virtual at the individual level, this invention completely removes personally identifiable information (PII). Combined with differential privacy and compliant encapsulation, sensitive educational and research data that was previously unavailable for public release can circulate risk-free among research institutions, supporting large-scale research replication and algorithm competitions.
[0039] 2) Achieving bimodal alignment of "structured + unstructured": This invention can simultaneously generate questionnaires / scales (structured) and interviews / classroom records (unstructured), while ensuring logical consistency between the two (e.g., if the generated student profile is "high anxiety," the generated interview text will also reflect the language characteristics of anxiety). This provides high-quality training data for multimodal educational computing research.
[0040] 3) Enhancing long-tail coverage and sample diversity: Through the expansion strategy of "fixed core variables + perturbation of non-core variables" and the optimization mechanism based on statistical feedback, this invention can effectively enhance the sample size of rare groups (such as students with special education needs) and extreme scenarios, solve the problem of "difficulty in obtaining long-tail samples" in real data collection, and improve the robustness of downstream analysis models.
[0041] 4) Enhanced educational validity of data: Unlike general text generation, this invention introduces domain knowledge base retrieval (RAG) and multi-agent interaction, making the generated educational scenarios (such as classroom questioning and home-school conflicts) conform to educational principles and real social interaction logic, not only "similar in form" (correct format) but also "similar in spirit" (conforming to educational laws).
[0042] 5) Built-in algorithm fairness correction: Through the fairness rectification module, this invention can actively detect and eliminate potential biases against sensitive attributes such as gender and region at the source of data generation, providing standardized unbiased benchmark data for fairness research in educational AI. Attached Figure Description
[0043] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0044] This invention provides a method for synthesizing, generating, aligning statistical distributions, and rectifying privacy-compliant educational research data for educational and sociological research. It is directly applied to data construction and sharing scenarios involving common multi-source data formats in educational research (including but not limited to questionnaires and scales, teacher / student / parent interview texts, classroom interaction and observation records, after-school learning time and time usage logs, and school-class-region level educational resource and management statistics). Known and potential application areas and methods include:
[0045] 1) Construction of research datasets in education / educational sociology
[0046] Generates synthetic data that can be used for regression analysis, hierarchical linear modeling (HLM), structural equation modeling, and causal inference; outputs codebooks, variable dictionaries, sample weights, and quality reports for paper reproduction and sharing.
[0047] 2) Research on Teacher Development and Teaching Improvement
[0048] Synthesize interview transcripts, observation records, and structured data from classroom interactions; provide topic tags / behavioral tags to support qualitative analysis and classroom discourse research.
[0049] 3) Research on students' time use and learning engagement
[0050] Synthesize time-allocation sequences, questionnaires, and open-ended questions; introduce missing mechanisms and measurement error structures to match the characteristics of real-world research data collection.
[0051] 4) Education policy evaluation and "virtual experiment / simulation research"
[0052] Injecting policy variables and implementation intensity into synthetic data, constructing control / treatment groups and multi-period panel data, supports In Silico RCT or quasi-experimental analysis.
[0053] 5) Cross-institutional data sharing and open scientific data repository
[0054] Use synthetic data to replace sensitive real data for sharing; include a privacy budget / mechanism summary, a re-identification risk assessment summary, and a bias check summary to facilitate ethical review (IRB) and compliance documentation.
[0055] 6) Offline evaluation and stress testing of educational intelligent systems
[0056] Generate multi-context classroom ecosystem data (teacher strategies, home-school communication, classroom interaction, questionnaire feedback) for system robustness assessment, deviation detection, and safety testing.
[0057] 7) Research on Educational Resource Allocation and School Governance
[0058] Data on school resources, teacher structure, class size, curriculum supply, and regional availability are synthesized for research on educational equity and resource efficiency.
[0059] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention. In the description of these embodiments, it should be noted that the use of terms such as "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicates that the specific features, structures, materials, or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the present invention.
[0060] Example 1
[0061] This embodiment provides a complete, compliant method for generating synthetic data for educational scientific research, which is particularly suitable for generating multimodal complex educational datasets that include "questionnaire data + interview text".
[0062] See Figure 1 The method specifically includes the following steps:
[0063] S1: Seed Data Collection and Metadata Organization In this embodiment, researchers collected data related to "student mental health and academic performance" from 10 middle schools in a certain region as seed data. Step S1 specifically includes:
[0064] S1-1: Structured Data
[0065] We collected 2,000 data points from the "Mental Health Scale for Middle School Students (MSSMHS)" (including scores for dimensions such as anxiety and depression) and their corresponding final math exam scores.
[0066] S1-2: Unstructured data
[0067] Transcripts of psychological interviews with 50 randomly selected students were transcribed verbatim (each transcript is approximately 30 minutes long).
[0068] S1-3: Metadata Construction
[0069] Create metadata tags for the above data and define variables. (Anxiety level, range 0-5) (Score, range 0-100) (Interview text). The data hierarchy is also marked as "Students (Level 1) - Schools (Level 2)".
[0070] S2: Seed Analysis and Variable Representation Construction
[0071] The seed parsing and variable representation construction system parses seed data and constructs a high-dimensional variable representation space. Specifically, for text data, LDA topic modeling or BERT encoding is used to extract "narrative style vectors" (e.g., fluency of expression, emotional word density); for structured data, a joint probability distribution model is established. The mapping rule established in this embodiment is as follows: if a student's anxiety score in the seed data is high (>4 points), the "emotional stability" dimension in their corresponding cognitive behavioral state vector is set to a negative value, and when generating text, it is forcibly mapped to the language style features of "hesitation, repetition, and high frequency of negative words".
[0072] S3: Generate Task Configuration
[0073] Researchers input commands to generate tasks, including:
[0074] 1) Objective: To generate a new dataset containing 10,000 virtual students;
[0075] 2) Scenario: Psychological counseling after final exams;
[0076] 3) Target Distribution: The generated virtual group should include at least 15% students with severe anxiety, slightly higher than [previous percentage].
[0077] 10% of the seed data was used to simulate a high-pressure environment and to maintain the statistical regularity that "anxiety level is negatively correlated with math performance".
[0078] S4: Conditional Generation and Sample Expansion
[0079] A strategy of "fixed core variables + perturbation of non-core variables" is employed to drive large models, such as GPT-4 or a fine-tuned version of the open-source Llama-3, to generate samples. Specifically, for the... For each sample to be generated, the system first samples from the latent space to determine its core variables, such as: gender = female, anxiety level = severe. Then, the non-core variables are randomly perturbed. For example, "home address" is perturbed from a specific "street A" to "community B" with the same socioeconomic attributes; "favorite subject" is randomized. A prompt is constructed: "You are a middle school girl with severe anxiety who just failed your math test; please fill out the following scale and answer the counselor's question about 'how you are feeling right now'; please answer in a hesitant tone."
[0080] S5: Domain Knowledge Enhancement and Process Generation
[0081] To ensure that the generated interview content conforms to the principles of educational psychology, this embodiment employs Retrieval-Enhanced Generation (RAG) and multi-agent interaction, specifically including:
[0082] 1) Enhanced retrieval: Based on the tag "severe anxiety", the system retrieves typical symptom descriptions (such as "somatization symptoms, sleep disorders") from a pre-set psychology knowledge base and injects them as background knowledge into Prompt to prevent the large model from making up random information.
[0083] 2) Multi-agent interaction: Instantiate two agents—"Experienced Psychologist Agent" and "Virtual Agent"—to form a virtual agent.
[0084] The "student agent" engages in multiple rounds of dialogue (5-8 rounds); the teacher agent will ask follow-up questions based on the student's answers, such as, "You just mentioned having trouble sleeping, is that a frequent occurrence?", thereby generating procedural dialogue data with logical depth.
[0085] S6: Distribution Assessment
[0086] After generating a batch of data (e.g., 100 records) for distribution evaluation, the statistics are calculated, including:
[0087] 1) Marginal distribution: Calculate the average score and anxiety distribution histogram of the virtual students;
[0088] 2) Correlation: Calculate the Pearson correlation coefficient between anxiety level and performance;
[0089] 3) Distribution Difference: Calculate the KL divergence (Kullback-Leibler divergence) between the above statistics and the target distribution.
[0090] Divergence).
[0091] S7: Strategy Optimization
[0092] Feedback-based policy optimization: This embodiment employs closed-loop optimization based on reinforcement learning (RL), specifically including:
[0093] 1) State: The current measure of distributional dissimilarity;
[0094] 2) Action: Fine-tune the generation temperature of the large model (from 0.7 to 0.9 to increase diversity) or adjust the weight parameter of "anxiety expression" in the cue words;
[0095] 3) Reward:
[0096] ,
[0097] in, (Reward) is the reward signal (or reward function value) of the core variable and evaluation index. In the reinforcement learning framework, this value serves as a feedback signal to evaluate the quality of the synthesized data under the current generation strategy and guide the model to maximize the reward to achieve distribution alignment. The performance distribution difference measure represents the KL divergence between the synthetic data and the target seed data in the dimension of "academic performance" (such as math scores), and is used to quantify the degree of distributional fit of this structured variable; The KL divergence measure of psychological trait distribution represents the KL divergence between the synthetic data and the seed data on the core research variable of "anxiety level," ensuring the fidelity of the synthetic group in the psychometric dimension. The text quality score is a comprehensive evaluation indicator for unstructured interviews or observation records, covering narrative style parameters such as register formality, emotional intensity, semantic coherence, and whether it conforms to the character setting (e.g., "hesitant tone"). To optimize the weight parameters, The optimization priority corresponding to the distribution of academic performance The optimization priority corresponding to the distribution of mental health traits. By adjusting the weights corresponding to the rationality of the text narrative, researchers can balance the "statistical fidelity" of structured variables with the "semantic fidelity" of unstructured texts.
[0098] 4) Algorithm: Update the generated policy network using the Proximal Policy Optimization (PPO) algorithm until the distribution difference is less than a preset threshold (e.g., 0.05).
[0099] S8: Compliance Processing and Delivery
[0100] The optimized synthetic data undergoes privacy-preserving processing and fairness rectification, and is then packaged into a shareable dataset for output, specifically including:
[0101] 1) Privacy protection: The teacher collection privacy mechanism (PATE) is adopted to train 5 "teacher models" to generate data on non-overlapping subsets of seed data. Their prediction results are aggregated and voted with noise to train the final "student generator" and ensure that the records of specific real students cannot be reversed from the output.
[0102] 2) Fairness rectification: It was detected that the average math score of "girls" in the generated data was significantly lower than that of "boys" (this...).
[0103] (These are stereotypes that the model may have learned). The system triggers counterfactual generation, forcing the ability values to remain unchanged, only reversing the gender and regenerating a portion of the samples until the gender bias is eliminated.
[0104] Example 2
[0105] This embodiment provides a simplified method for generating synthetic data, suitable for scenarios with limited computing power or where only structured questionnaire data needs to be generated. The main difference between this embodiment and Embodiment 1 lies in the implementation of steps S5, S7, and S8.
[0106] S5: Simplified Generation Mechanism This embodiment does not involve multi-agent interaction; the system uses only a single large model interface. For the generation of "classroom observation records," the system directly utilizes "Chain of Thought (CoT)" technology, requiring the model to first list the three key nodes of classroom interaction (introduction, questioning, and summarizing), and then generate the complete observation text all at once, without performing multiple rounds of interactive simulation.
[0107] S7: Prompt Engineering Optimization This embodiment does not use reinforcement learning algorithms. The prompt engineering optimization specifically includes:
[0108] 1) Optimization logic: The system sets a set of predefined prompt word template variations (e.g., variation A emphasizes details, variation B emphasizes emotion);
[0109] 2) Iterative Process: After the first round of generation, if the proportion of the "long-tail group" (such as students from low-income families) in the generated data is found to be insufficient, the system automatically adds an explicit instruction to the prompt of the next round of generation: "Please increase the proportion of students whose family annual income is less than 10,000 yuan" and increases the sampling weight of this type of sample. Distribution alignment is achieved through this rule-based feedback loop.
[0110] S8: Compliance Processing and Delivery
[0111] This embodiment of differential privacy implementation does not use a complex PATE mechanism. Instead, during the statistical evaluation and output stages, Laplace noise is directly added to the output statistical table (e.g., the number of people in each score range) to satisfy... - Differential privacy requirements are used to protect privacy.
[0112] Example 3
[0113] This embodiment provides a synthetic data generation and compliant sharing device for educational scientific research. The device includes a memory and a processor. The memory stores computer programs, seed data, a knowledge base, and a constructed variable representation space. The processor is connected to the memory and executes the stored computer programs. When the processor executes the programs, it specifically implements the following functional modules:
[0114] 1) Data parsing module: used to execute S1 and S2, clean the data and construct vector representations;
[0115] 2) Generation Control Module: Used to execute S3 and S4, receive configurations, and drive the large model to perform conditional generation;
[0116] 3) Interaction Enhancement Module: Used to execute S5, call external knowledge bases, and manage the state of multi-agent dialogue;
[0117] 4) Evaluation and Optimization Module: Used to execute S6 and S7, calculate statistical differences, and adjust the generation parameters using optimization algorithms (such as PPO or evolutionary algorithms);
[0118] 5) Compliance Engine: Used to execute S8, implement differential privacy noise addition and fairness detection.
[0119] Those skilled in the art will understand that the device may be a cloud server, a high-performance workstation, or a server cluster.
[0120] The above is merely a further description of the present invention and is not intended to limit the scope of this patent. Any equivalent implementation of the present invention should be included within the scope of the claims of this patent.
Claims
1. A method for generating and compliantly sharing synthetic data for educational and scientific research, characterized in that, A statistical distribution alignment and privacy compliance rectification method is adopted to achieve synthetic data generation and compliant sharing. The method specifically includes the following steps: S1: Seed Data Acquisition and Metadata Processing Acquire seed data and its metadata for educational research, clean the seed data, unify its format and manage its version to obtain a standardized seed dataset. The seed data includes structured data or unstructured data. S2: Seed Analysis and Variable Representation Construction The standardized seed dataset is parsed to construct a variable representation space that maps the original field features. The variable representation space includes: a sociological attribute vector representing group and environmental attributes and a cognitive behavioral state vector representing individual states and abilities. S3: Generate Task Configuration Receive the generation task configuration and determine the target data type, the subject and scene to be generated, and the target statistical distribution; S4: Conditional Generation and Sample Expansion Based on the variable representation space and generation task configuration, a generation context containing role and scene constraints is constructed to drive the large model to generate synthetic samples. For each sample to be generated, at least one core research variable is fixed, and at least one non-core attribute variable is subject to controllable perturbation, thereby expanding sample diversity while maintaining the semantic consistency of research variables. S5: Domain Knowledge Enhancement and Process Generation By combining the target data type with knowledge from the fields of education or psychology, an enhanced generative context is constructed to drive the model to generate synthetic data containing procedural records or interaction details. S6: Distribution Assessment Calculate key statistics for the synthetic data and compare them with the target statistical distribution to obtain a measure of distributional difference; S7: Feedback-based policy optimization The distribution difference measure is converted into a feedback-based optimization signal, and the generation strategy is iteratively updated until the statistical characteristics of the synthetic data meet the preset distribution alignment conditions, thus obtaining the optimized synthetic data. S8: Compliance Processing and Delivery The optimized synthetic data undergoes privacy protection processing and fairness rectification, and is then packaged into a shareable dataset output.
2. The method for generating and compliantly sharing synthetic data for educational research according to claim 1, characterized in that, The seed data in step S1 covers typical modalities of educational research; the structured data includes questionnaire survey data, scale assessment data, or hierarchical statistical data; the unstructured data includes interview texts, classroom observation records, open-ended question answer texts, or teaching logs; the metadata not only includes basic definitions but also covers research design information, sampling structure information, and measurement tool information. The research design information includes: cross-sectional study, longitudinal longitudinal study, panel structure or quasi-experimental / control design identifiers; the sampling structure information includes: stratified variables, cluster variables, sample weights, post-stratified weights, time point identifiers and sampling frame descriptions; the measurement tool information includes: questionnaire item version, scale dimension structure and reverse question rules and interview outline version.
3. The method for generating and compliantly sharing synthetic data for educational research according to claim 1, characterized in that, The variable representation space construction in step S2 employs constraint modeling with field type constraints and hierarchical consistency constraints, response style modeling, and text style modeling. The field type constraints are the generation logic that distinguishes between continuous, categorical, ordinal, count, date / time, or text fields. The hierarchical consistency constraints establish nested relationships between individuals, classes, schools, and regions, ensuring consistent foreign key associations and temporal order. The response style modeling establishes response style parameters for questionnaire data, including: agreement tendency, extreme reaction tendency, social expectation tendency, or casual response tendency. The text style modeling establishes narrative style parameters for text data, including: register formality, emotional intensity, and level of detail, as well as identity parameters that distinguish the discourse characteristics of teachers, parents, students, or administrators.
4. The method for generating and compliantly sharing synthetic data for educational research according to claim 1, characterized in that, Step S3 generates the task configuration according to the following control rules: 1) A mapping table of subject-scene-data type was established, stipulating that the same virtual subject should share the same identity profile when generating data in different scenarios to maintain consistency of personality; 2) Based on the estimation of seed data, publicly available statistical summaries, or hypothetical distributions specified by the researcher, clarify the source of the target distribution; 3) Define a multi-dimensional set of evaluation indicators, as well as long-tail coverage constraints on the proportion of rare groups, coverage of extreme quantiles, or coverage of rare scenarios. 4) The variables were explicitly divided into a "core variable set" and a variable "non-core variable set".
5. The method for generating and compliantly sharing synthetic data for educational research according to claim 1, characterized in that, The specific execution process of step S4 includes: Step S4-1: Select or instantiate prompt templates for character settings, scene settings, variable constraints, and output structure; Step S4-2: Sample from the variable representation space and implement controlled perturbation strategies, including demographic attributes, scene details, narrative style, or missing patterns, to expand sample diversity while keeping the core research variables fixed. Step S4-3: Drive the large model to synchronously output structured fields and unstructured text; Step S4-4: Perform consistency verification, including checking field validity and logical consistency, and generating sample-level metadata containing subject identifier, generation condition summary and round information for samples that pass the verification and storing it in the database.
6. The method for generating and compliantly sharing synthetic data for educational research according to claim 1, characterized in that, The domain knowledge enhancement in step S5 employs the following mechanism: 1) Enhanced retrieval: Retrieve behavioral characteristics, narrative themes, or typical event patterns from the education / psychology knowledge base based on target tags, and inject them into the generated context; 2) Multi-agent interaction closed loop: Construct virtual agent pairs between researchers and interviewees, teachers and students, or parents and teachers, and generate materials containing question-response structures through multiple rounds of interaction; 3) Process Log: Outputs a process log containing turn number, speaker, triggering conditions and timestamp, forming high-dimensional data aligned with "text-process log-structured variables".
7. The method for generating and compliantly sharing synthetic data for educational research according to claim 1, characterized in that, The distribution evaluation system in step S6 covers the following multiple levels: 1) Statistical measures of marginal distribution and stratification ratio at each level of individual, class, and school; 2) Stratified joint distribution statistics of the set of key variables; 3) Cross-level correlation structure statistics; 4) Structured quality statistics of text data, and quantify the difference between synthetic data and target distribution by calculating KL divergence, Wasserstein distance or maximum mean difference.
8. The method for generating and compliantly sharing synthetic data for educational research according to claim 1, characterized in that, The policy optimization process in step S7 forms a closed-loop feedback, constructs an optimization signal containing negative distribution difference and long-tail coverage reward, determines the policy to be updated from the prompting policy or sampling policy, and uses a reinforcement learning algorithm for near-end policy optimization (PPO) to iteratively update the policy until the termination condition of distribution difference being lower than the threshold or quality report convergence is met.
9. The method for generating and compliantly sharing synthetic data for educational research according to claim 1, characterized in that, The compliance processing in step S8 includes post-processing for privacy protection, fairness rectification, and encapsulated delivery. The privacy protection configuration implements differential privacy or teacher set privacy mechanisms to ensure that the output meets the privacy budget. The fairness rectification monitors the correlation between sensitive attributes and outcome variables, and triggers counterfactual generation when the threshold is exceeded. This means keeping sensitive attributes unchanged and resampling non-sensitive conditions to balance the bias. The encapsulated delivery output includes a delivery package containing a synthetic data ontology, a variable dictionary, sampling instructions, generated version information, and a compliance summary with privacy risk assessment and bias checks.
10. An apparatus for constructing a method for generating and compliantly sharing synthetic data for educational research as described in claim 1, characterized in that, The synthetic data generation and compliance sharing device includes: a processor, a memory, and a computer program stored in the memory, wherein when the computer program is executed by the processor, the processor performs any one of the steps of the method of claim 1.