Data management architecture oriented to teaching, research and training scenes

Through the data governance architecture for teaching, scientific research and training scenarios, the problem of dispersed and poor management of data resources in the field of education, scientific research and training has been solved, data quality improvement, security enhancement and utilization efficiency have been achieved, and scientific basis for educational decision-making.

CN120068119APending Publication Date: 2025-05-30SHENZHEN INSTITUTE OF EDUCATIONAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510189502.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The data resources in the field of education, scientific research and training are scattered and lack of unified management, resulting in unstable data quality, insufficient security, and inefficient utilization efficiency.

Method used

Design a data governance architecture for teaching, scientific research and training scenarios, including data collection and preprocessing, data storage, data governance, data analysis and mining, data visualization and application, and realize the governance of teaching, scientific research and training data through the integration of advanced data processing technology, intelligent analysis tools and efficient data management strategies.

Benefits of technology

It has achieved in-depth understanding and optimization of the education and teaching process, education management, student development, etc., improved data quality, security and utilization efficiency, provided scientific decision-making basis, and promoted the optimization and development of teaching, scientific research and training business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a data governance architecture oriented to a teaching, research and training scene, and the architecture comprises a data collection and preprocessing module which is responsible for collecting teaching, research and training related data from a plurality of data sources, and carrying out the preprocessing operation, such as analysis, deduplication, format conversion, data verification, etc. The data storage module is used for storing the preprocessed mass data by adopting a distributed storage technology and establishing a data management strategy including data backup, recovery and access control, and the data management module is used for classifying and marking the stored data, implementing data security and privacy protection measures and performing data quality management. The data analysis and mining module is used for carrying out deep analysis and mining on the stored data by using intelligent analysis tools such as data mining, machine learning and deep learning, and the data visualization and application module is used for displaying analysis results in the forms of visual charts, reports and the like by using a data visualization tool. And customized data reports are generated according to requirements of different users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data governance, and in particular, to a data governance architecture for teaching, scientific research, training scenarios. Background Art

[0002] With the rapid development of information technology, a large amount of data resources have been accumulated in the field of education, scientific research and training. These data are often scattered in different systems, lacking unified management and governance, resulting in problems such as uneven data quality, insufficient data security, and low data utilization efficiency. At present, educational big data shows obvious development trends in the following aspects: continuous broadening of data sources, progress in data analysis technology, development of educational personalization, scientific education decision-making, digital transformation of the education industry, and attention to education security and privacy protection. How to analyze and mine data by making full use of educational big data on the premise of ensuring data security has become an urgent problem to be solved. Therefore, the present invention aims to provide an innovative data governance architecture applicable to the field of teaching, scientific research and training, and realizes the governance of teaching, scientific research and training data by integrating advanced data processing technologies, intelligent analysis tools and efficient data management strategies. Summary of the Invention

[0003] The present invention provides a data governance architecture for teaching, scientific research, training scenarios, which realizes in-depth understanding and optimization of aspects such as the education and teaching process, education management, and student development by collecting, integrating, analyzing and mining massive educational information. In a first aspect, the present invention provides a data governance architecture for teaching, scientific research, training scenarios, including: a data collection and preprocessing module: used to collect teaching, scientific research, training-related data from multiple data sources, and the data sources include but are not limited to teaching research documents, teaching system records, scientific research project data, training materials, etc.; and perform preprocessing operations such as parsing, duplicate removal, format conversion and data verification on the collected data to improve data quality and usability.

[0004] A data storage module: stores the preprocessed massive data by using distributed storage technology to ensure the efficient storage and management of data; at the same time, establishes a data management strategy, including data backup, recovery and access control, to ensure the security and usability of data.

[0005] A data governance module: classifies and labels the stored data, formulates classification criteria according to data characteristics and application requirements, and adds metadata tags to the data manually or automatically; implements data security and privacy protection measures, and uses data encryption technology, strict access control and data backup and recovery strategies to ensure the security and integrity of data during transmission, storage and use; performs data quality management through a dual verification method of manual + intelligent model to ensure the accuracy, consistency and integrity of data.

[0006] Data Analysis and Mining Module: Utilize intelligent analysis tools such as data mining, machine learning, and deep learning to deeply analyze and mine the stored data, discover potential values and rules in educational research, teaching, and training data, and provide a scientific basis for educational decision-making. The analysis and mining include, but are not limited to, learning analysis, teacher performance analysis, course effect analysis, etc.

[0007] Data Visualization and Application Module: Use data visualization tools to display the analysis results in intuitive charts, reports, etc., and generate customized data reports according to the needs of different users, providing decision-making support for educational administrators, researchers, trainers, etc., thereby promoting the optimization and development of educational research, teaching, and training services.

[0008] Optionally, data parsing in the data collection and preprocessing module adopts a combination of a program library and a large model. The program library includes, but is not limited to, pdfminer, python-docx, etc., and is used to parse document data in different formats.

[0009] Optionally, the distributed storage technology adopted by the data storage module includes Hadoop, Spark, etc., which can achieve distributed storage and parallel processing of data, improving data processing efficiency.

[0010] Optionally, data classification and annotation in the data governance module adopt classification methods based on rules or machine learning, which can automatically identify the categories and features of data and assign corresponding labels to them.

[0011] Optionally, the data encryption technologies adopted by the data security and privacy protection measures in the data governance module include AES, RSA, and TLS / SSL, etc., to ensure the confidentiality and integrity of data during transmission and storage; the access control restricts users' access to and operations on data by assigning roles and permissions to different user groups.

[0012] Optionally, the intelligent analysis tools in the data analysis and mining module can select appropriate algorithms and models for data mining and analysis according to different analysis objectives and data characteristics, such as association rule mining, classification algorithms, clustering analysis, etc.

[0013] Optionally, the data visualization tools in the data visualization and application module include Tableau, PowerBI, etc., which can display complex data in an intuitive and easy-to-understand graphical way, facilitating users' understanding and analysis.

[0014] Optionally, the architecture further includes a data augmentation sub-module, which is located between the data preprocessing and data governance modules and is used to augment data before it enters the storage layer and during the data governance process, including augmentation based on existing data and augmentation based on self-distillation of large models. The augmentation based on existing data adopts multi-stage augmentation methods such as key information extraction, information rewriting, and content recombination, and uses the Pangu large model and the vLLM inference acceleration framework to increase the number and diversity of documents. The augmentation based on self-distillation of large models maintains a document pool, performs iterative self-distillation operations using the Pangu large model, generates high-quality augmented documents, and performs deduplication processing to improve the richness and quality of the data.

[0015] Optionally, it further includes a data annotation sub-module, which combines automatic annotation and manual annotation. With the Pangu large model as the core, it introduces a multi-agent collaboration framework, formulates detailed annotation guidelines, and conducts multiple rounds of data annotation processes. Each round of annotation includes steps such as document collection, augmentation, screening, and automatic instruction annotation. After annotation, some data is randomly selected for the training and testing of the large model to continuously optimize the annotation effect and data quality.

[0016] On the other hand, this solution also provides a data governance method applicable to the field of education, research, training, and cultivation. This method includes the following steps: Step 1: Data collection and preprocessing module. Start data collection: Use the data collection and preprocessing module to determine the data types and sources to be collected, such as teaching and research documents, teaching system records, etc.

[0017] Execute data acquisition: Use methods such as web crawlers, data interfaces, sensor technology, or questionnaires to collect data from the determined data sources.

[0018] Data preprocessing: In the same module, perform operations such as parsing, deduplication, format conversion, and data verification on the collected data, and use specialized program libraries (such as pdfminer, python-docx) to parse documents in different formats.

[0019] Step 2: Data storage module Transfer to data storage: Transfer the preprocessed data to the data storage module.

[0020] Implement data storage: In the data storage module, use distributed storage technologies such as Hadoop and Spark for data storage, and establish data management policies, including data backup, recovery, and access control, to ensure the security and availability of the data.

[0021] Step 3: Data governance module Data Classification and Annotation: In the data governance module, classify and annotate the stored data, formulate classification criteria and add metadata tags.

[0022] Data Security and Privacy Protection: In the same module, implement data security and privacy protection measures, including data encryption (such as AES, RSA), access control, and data backup and recovery strategies.

[0023] Data Quality Management: Conduct data quality management through a dual-verification method of manual and intelligent models to ensure data accuracy, consistency, and integrity.

[0024] Step 4: Data Augmentation Sub-module Post-Data-Preprocessing Augmentation: In the data governance module, start the data augmentation sub-module and augment the data during data preprocessing and data governance.

[0025] Augmentation Based on Existing Data: Adopt multi-stage augmentation methods such as key information extraction, information rewriting, and content recombination to increase the scale and diversity of the dataset.

[0026] Augmentation Based on Large Model Self-Distillation: By maintaining a document pool, use the large model for iterative self-distillation operations to generate high-quality augmented documents and perform deduplication.

[0027] Step 5: Data Annotation Sub-module Data Annotation Initiation: In the data governance module, start the data annotation sub-module and formulate detailed annotation guidelines.

[0028] Combination of Automatic and Manual Annotation: Adopt a combination of automatic and manual annotation methods, with the Pangu large model as the core, and introduce a multi-agent collaboration framework for data annotation.

[0029] Multi-Round Data Annotation: Conduct a multi-round data annotation process, including steps such as document collection, augmentation, screening, and automatic instruction annotation, and randomly select part of the data for large model training and testing.

[0030] Step 6: Data Analysis and Mining Module Data Transmission for Analysis: Send the governed data into the data analysis and mining module.

[0031] Selection of Intelligent Analysis Tools: Select appropriate intelligent analysis tools according to the analysis objectives, such as data mining, machine learning, and deep learning tools.

[0032] Execute Analysis and Mining: Execute specific analysis and mining tasks such as learning analysis, teacher performance analysis, and course effect analysis.

[0033] Step 7: Data Visualization and Application Module Transfer to Data Visualization: Transfer the analysis results to the data visualization and application module.

[0034] Data Visualization Implementation: Use data visualization tools such as Tableau and Power BI to display complex data in an intuitive and understandable graphical way.

[0035] Customized Data Report Generation: Generate customized data reports according to the needs of different users to provide decision-making support for education administrators, researchers, and trainers.

[0036] Step 8: Data Review Sub-module Post-annotation Review of Data: After data annotation is completed, start the data review sub-module.

[0037] Format Review and Applicability Assessment: Conduct a format review and applicability assessment on the annotated data to ensure that the data can meet the requirements of specific tasks.

[0038] Manual Evaluation: By means of manual evaluation, randomly sample the annotated data samples to ensure high data quality, with accuracy and consistency both exceeding 95%. Description of the Drawings

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is the module architecture diagram of the present invention; Figure 2 It is the flowchart of the data governance method of the present invention; Figure 3 It is the flowchart of the three-stage document expansion method based on existing data of the present invention; Figure 4 It is the example diagram of information extraction of the present invention; Figure 5 It is the flowchart of the data expansion method based on large model self-distillation of the present invention; Figure 6 It is the example document diagram in the document pool of the present invention; Figure 7 It is the example diagram of the expanded document of the present invention Detailed Embodiments

[0041] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work shall fall within the protection scope of the present invention.

[0042] Unless otherwise specified, the raw materials used in the present invention are all conventional products purchased from the market.

[0043] Embodiment 1 A data governance architecture for teaching, scientific research, training and other scenarios, including the following modules: Data collection and preprocessing module: used to collect teaching, scientific research, training and other related data from multiple data sources, the data sources including but not limited to teaching research documents, teaching system records, scientific research project data, training materials, etc.; and perform preprocessing operations such as parsing, duplicate removal, format conversion and data verification on the collected data to improve data quality and usability.

[0044] Data storage module: adopts distributed storage technology to store a large amount of preprocessed data to ensure the efficient storage and management of data; at the same time, establish a data management strategy, including data backup, recovery and access control, to ensure the security and usability of data.

[0045] Data governance module: classify and label the stored data, formulate classification criteria according to data characteristics and application requirements, and add metadata tags to the data manually or automatically; implement data security and privacy protection measures, use data encryption technology, strict access control and data backup and recovery strategies to ensure the security and integrity of data during transmission, storage and use; perform data quality management through the double verification method of manual + intelligent model to ensure the accuracy, consistency and integrity of data.

[0046] Specifically, in the data governance module, start the data augmentation sub-module, and during the data preprocessing and data governance process, augment the data. The data augmentation methods include: 1. Augmentation method based on existing data There are 429 pieces of existing document data, and its data scale is small, which cannot meet the training tasks of large models in the education scenario. Therefore, the most direct data enhancement method is to augment the data based on the existing documents. In this work, a three-stage augmentation method is adopted, as Figure 3 shown.

[0047] The first stage is called the key information extraction stage. Most existing documents contain rich structured information. Taking notice documents as an example, the existing documents contain key information such as the theme, participants, time, and location of the event. Therefore, according to this feature, this work first extracts the key information from the parsed document content.

[0048] According to the type of the document, we divide the original documents into three categories: notice type, plan type, and summary type. According to the characteristics of each major category, this work designs multiple possible information fields respectively (for example, notice theme, notice participants, notice time, etc.). This work uses the Pangu large model for key information extraction. The input is the parsed rich text and the information fields to be extracted, and the output is the specific information corresponding to each field, which is output in the format of json.

[0049] This process only involves the inference of the large language model and does not involve complex training processes.

[0050] The second stage is called the information rewriting stage. According to the extracted key information, this work uses the Pangu large model to rewrite the information corresponding to each field K times. The input of the large model is the field information to be rewritten, and the output is the rewritten content information. To ensure the inference speed, we adopt the vLLM inference acceleration framework.

[0051] The third stage is called the content recombination stage. In the first two stages, each field of an original document has been rewritten. Therefore, we can randomly perform content recombination to obtain diverse document outputs. To avoid generating too many similar extended documents for the same document, we limit the number of recombined and extended documents to 5.

[0052] Through the above three-stage document extension scheme, this work further filters out highly similar duplicate samples and can obtain 2000 extended documents.

[0053] 2. Data augmentation method based on large model self-distillation This work also adopts a data augmentation method based on large model self-distillation. On the one hand, it can stimulate the knowledge reserve of the large model itself and generate long document data that meets the corresponding instructions. On the other hand, many prior experiments have found that over-relying on external data for large model training will affect the model distribution and lose some performance on general tasks. Therefore, the document data formed through the self-distillation of the large model will have very great application value.

[0054] To implement the method of self-distillation data augmentation and ensure the diversity of generated documents, this work maintains a Document Pool. Each time a generation is performed, a document is drawn from the Document Pool as a reference example to guide the generation of a new document. Therefore, this data augmentation method can be regarded as an iterative self-distillation data enhancement. See the method process for details. Figure 5 。

[0055] Perform iterative self-distillation operations using the Pangu large model. Set the number of iterative rounds to 4, and use all the documents in the Document Pool as the reference documents for each round in turn. To avoid the generated documents being too single, this work performs a deduplication operation on the generated documents. Specifically, the BLEU metric is used to calculate the similarity between the new document and each document in the Document Pool. When the metric exceeds the threshold, the generated new document is not retained.

[0056] Updating the document is implemented by a function in the iterative self-distillation process. Among them, the generate_new_document function is similar to the implementation of the above information recombination. The compare_documents function is used to compare the similarity of two documents based on the BLEU score. The sentence_bleu in nltk is used for calculation. The following table shows the total number of documents in the Document Pool after each round.

[0057] Table 1. Variation diagram of the total number of documents in the Document Pool with the iterative rounds

[0058] Figures 6 - 7 An example of data enhancement is shown, Figure 6 is the original document, Figure 7 is the document after augmentation. It can be seen that the augmented document refers to the structure of the original document and rewrites the chapter names. In terms of content, there are very big changes compared with the original document, forming a high-quality augmented document data.

[0059] Specifically, in the data governance module, the data annotation sub-module is started. The data annotation specifically formulates annotation guidelines in accordance with the data governance criteria to ensure the standardization and consistency of the annotation process. After data annotation according to the guidelines, a detailed analysis of the annotation results is carried out to ensure the data quality and integrity, and to promote the accuracy and reliability of subsequent data analysis and applications.

[0060] In terms of the three major aspects of annotation guidelines, data annotation process, and annotation result analysis, the description of this work is as follows: a. Annotation guidelines By combining automatic annotation with manual annotation, not only can the annotation efficiency and accuracy be improved, but also the diversity and quality of annotation can be ensured. By using the Pangu large model and a domestic crowdsourcing annotation platform, technical and human resources can be fully utilized to provide an efficient and reliable solution for the annotation work of various types of documents.

[0061] For automatic annotation, the Pangu large model is planned to be used as the core foundation to fully utilize its advanced capabilities in natural language processing. An execution framework of multi-agent collaboration is planned to be introduced, and multiple agents are deployed to cooperate. Each agent is responsible for different tasks, such as question instruction generation, answer output annotation, and answer content correction. Question instruction generation is performed by one agent to generate relevant questions based on the content of the original document to guide the annotation process. Output answer annotation is performed by the second agent to generate corresponding answers according to the question instructions and perform preliminary annotation. Answer content correction is performed by the third agent to correct and optimize the output answers to ensure the annotation quality.

[0062] b. The first-round data annotation process To successfully complete the training objectives of the first stage, the project team carried out data annotation in the first stage, which mainly included the following four steps: document collection, document expansion, continued pre-training document screening, and instruction automatic annotation. The following will be elaborated in detail.

[0063] 1) Document collection. In the preparatory stage of the first-stage model training, we first carried out data collection work. The goal of this step is to obtain a sufficient number of documents to ensure that the model can learn rich language features and document structures. For this purpose, we collected 531 documents from multiple sources, which mainly covered three types: notice type, plan type, and summary type. This step is the basis of the entire data construction process and provides the original materials for subsequent data expansion and model training. Among these documents, the project team selected 429 high-quality ones as candidates and screened out those with short text length, chaotic structure, and difficult to parse.

[0064] 2) Document expansion. To improve the generalization ability and robustness of the model, we adopted specific data expansion methods. Data expansion is a common technique that increases the diversity of the dataset by introducing changes, thereby helping the model better understand and handle various situations. In this step, we used the data expansion method based on existing data and the data expansion method based on large model self-distillation to expand the original 429 documents to 2772, increasing the number and coverage of the documents. This process may include techniques such as text rewriting, synonym replacement, and sentence restructuring to ensure that the expanded documents have a certain degree of novelty while maintaining the original information.

[0065] 3) Continue with pre-training document screening. After data augmentation, we carefully selected 510 documents from 2,772 documents for the continued pre-training of the large model. The purpose of this step is to ensure the quality of the documents used for training, remove possible noisy data, and improve the efficiency of model training. During the screening process, we consider factors such as the integrity, relevance, and compliance with the training objectives of the documents.

[0066] 4) Automatic instruction annotation. To enable the model to understand and execute specific instructions, we need to annotate the data to guide the learning of the large model. In this step, we adopted an automatic annotation method to annotate the instructions in 2,772 documents. Specifically, notification-type instructions mainly focus on notification time, location, specific items, and targets, etc., solution-type instructions include solution content, specific implementation content, and step content, etc., and summary-type instructions mainly include deeds, activity details, generalizations, and outlooks, etc. A total of 32,842 instructions were annotated in this process, including 13,303 notification-type instructions, 13,300 solution-type instructions, and 6,239 summary-type instructions. The automatic annotation method can process a large amount of data quickly and efficiently, but subsequent manual review and correction are also required to ensure the accuracy of the annotation. Finally, 95% of the annotated instruction data was randomly selected for the training of the large model, and the remaining 5% was used for testing.

[0067] c. The second round of data annotation process The project team first carried out data collection work, which is the first step in building an effective dataset. To meet the needs of comprehensive and knowledge-based tasks, the project team carefully selected and collected 12,335 documents. These documents may come from different fields and topics to ensure the diversity and coverage of the dataset. The collected documents will serve as the basis for subsequent data augmentation and pre-training.

[0068] After the data collection was completed, the project team decided to use all 12,335 documents as the target for data augmentation. This means that each document will be used to increase the size and diversity of the dataset. In this way, it can be ensured that the model is exposed to a wider range of information and contexts during the training process, thereby improving its generalization ability.

[0069] The project team decided to use all the collected documents for continued pre-training. This step is to ensure that the model can be exposed to a large amount of diverse data during the pre-training stage. Pre-training is an important stage for the model to learn the language foundation and document structure. By using these documents, the model can better understand and process various types of tasks.

[0070] To construct the instruction annotation data for the large model, the project team adopted an automatic construction method. This method utilizes advanced natural language processing techniques and can automatically extract and generate instruction data from documents. The automatic construction method not only improves the efficiency of data processing but also ensures the consistency and standardization of data generation. Through the automatic construction method, the project team successfully generated 86,293 instruction data. These data are crucial for model training as they directly guide the model on how to understand and execute user instructions. Specifically, comprehensive instructions focus on specific content, while knowledge-based instructions focus on knowledge points, the current state of knowledge, the context of development, and the basic points of development trends. Among these instruction data, there are 16,416 comprehensive instructions and 69,877 knowledge-based instructions. Such classification helps the model distinguish different types of instructions and thus respond more accurately to user needs in practical applications. Finally, 95% of the labeled instruction data were randomly selected for large model training, and the remaining 5% were used for testing.

[0071] d. The third round of data annotation process After two rounds of training, the project team noticed that the large model performed excellently in notice, plan, summary, and knowledge-based documents. However, when dealing with comprehensive documents, the effect was slightly insufficient, which might be due to the more diverse content of comprehensive documents. To address this issue, the project team decided to focus on expanding the data annotation work for comprehensive documents in the third round of training to enhance the large model's performance in this category. By increasing the annotation of comprehensive data, we aim to enable the model to better understand and process such complex and diverse document content, thereby comprehensively improving its performance across various types of documents and ensuring that the model can demonstrate excellent performance and adaptability in all fields.

[0072] Specifically, in the third round of the data annotation process, the project team adopted a method similar to that of the second round of data annotation to generate instruction data. The effectiveness of this method has been verified in the previous two rounds, so it is considered a reliable and efficient data annotation strategy in the third round. Through this continuous data annotation method, the project team can maintain the consistency and accuracy of data annotation. The project team generated a total of 5,381 comprehensive instructions.

[0073] Data analysis and mining module: Using intelligent analysis tools such as data mining, machine learning, and deep learning, deeply analyze and mine the stored data to discover the potential value and laws in the education, research, training, and data, and provide a scientific basis for educational decision-making. The analysis and mining include but are not limited to learning analysis, teacher performance analysis, course effect analysis, etc.

[0074] Data Visualization and Application Module: Use data visualization tools to display the analysis results in the form of intuitive charts, reports, etc., and generate customized data reports according to the needs of different users, providing decision-making support for education administrators, researchers, trainers, etc., so as to promote the optimization and development of education, research and training business. Embodiment

[0075] A data governance method applicable to the field of education, research and training includes the following steps: ① Data collection and preprocessing Data collection: Collect data from multiple sources such as student information systems, learning management systems (LMS), online test platforms, and classroom interaction tools. These data include students' basic information, grades, attendance records, learning behaviors, teachers' teaching records, etc.

[0076] Data cleaning: Process missing values, duplicate data, and outliers to ensure the accuracy and consistency of the data. For example, for teaching data, duplicate student information can be removed through data cleaning tools to ensure the uniqueness of each student's data; for research data, invalid experimental data can be removed through data cleaning tools to ensure the accuracy of the data.

[0077] Specifically, special considerations and methods for data cleaning: Handling diverse data sources: Student information system: Includes students' basic information, grades, attendance records, etc.

[0078] Learning management system (LMS): Includes course content, learning progress, online test results, etc.

[0079] Educational applications and platforms: Include data generated by online learning platforms, educational games, etc.

[0080] Social media and forums: Interaction data of students on social media and educational forums.

[0081] Handling missing values: Filling in missing values: In educational data, missing values may appear in aspects such as students' grades and attendance records. Mean, median, mode, etc. can be used to fill them, or prediction filling can be carried out according to students' historical data.

[0082] Data interpolation: For time series data, such as students' learning progress, interpolation methods can be used to fill in missing values.

[0083] Handling duplicate data: Merging duplicate records: Students may have records in multiple systems, and these records need to be merged to ensure data consistency.

[0084] Duplicate removal: Ensure the uniqueness of each student's record to avoid the impact of duplicate data on the analysis results.

[0085] Handling outliers: Identify and handle abnormal grades. For example, if a test score is extremely high or low, check for data entry errors or other reasons.

[0086] Abnormal attendance records: For example, if a student's attendance record shows consecutive days of absence, verify the accuracy of the data.

[0087] Data standardization: Unify formats: Ensure the consistency of formats for fields such as dates, times, and grades for easier subsequent analysis.

[0088] Unify encodings: For example, unify the encodings for students' grades and classes to ensure data consistency.

[0089] Data conversion: Data type conversion: Convert date strings to date types for easier time series analysis.

[0090] Feature engineering: Create new features such as students' learning habits and learning paths to improve the effectiveness of analysis and modeling.

[0091] Handling inconsistent data: Unify naming: Ensure the consistency of fields such as course names and teacher names, and avoid different naming methods for the same course or teacher.

[0092] Merge synonyms: Merge synonyms or differently spelled words into a standard vocabulary.

[0093] Data validation: Range check: Ensure that numerical data such as grades and ages are within a reasonable range.

[0094] Format check: Ensure that data such as email addresses and phone numbers conform to the expected formats.

[0095] Logical check: Ensure that the data conforms to educational business logic, such as a student's date of birth not being later than the enrollment date.

[0096] Data privacy and security: When processing student data, ensure that sensitive information (such as names and ID numbers) is anonymized to protect student privacy. Use encryption technology to protect data security during data transmission and storage.

[0097] Through these data cleaning methods tailored to the characteristics of the education industry, the quality and consistency of data can be improved, providing a reliable basis for educational data analysis and decision-making.

[0098] Data integration: Integrate data from different systems to form a unified dataset.

[0099] ② Statistical analysis Descriptive statistical analysis: Calculate basic statistics such as mean, median, standard deviation, frequency distribution, etc., to understand the basic characteristics of the data. For example, analyze the grade distribution and attendance rate of students.

[0100] Correlation analysis: Use methods such as Pearson correlation coefficient and Spearman correlation coefficient to analyze the relationship between different variables. For example, analyze the correlation between study time and grades.

[0101] Regression analysis: Establish a regression model to predict certain educational outcomes. For example, use multiple regression analysis to predict students' exam scores.

[0102] Hypothesis testing: Use methods such as t-test and chi-square test to test the impact of different teaching methods, curriculum settings, etc. on students' grades.

[0103] ③ Data visualization Static charts: Use bar charts, pie charts, line charts, scatter plots, etc. to display the results of descriptive statistical analysis. For example, display the average grades of different classes and the attendance rate of students.

[0104] Dynamic dashboards: Create interactive dashboards to display key educational indicators in real-time, such as students' learning progress and teachers' teaching effectiveness. Users can view data from different dimensions through operations such as filtering and sorting.

[0105] Geographic Information System (GIS): Display the distribution of educational resources on a map, such as the geographical locations of schools and the distribution of students' home addresses, to help optimize the allocation of educational resources.

[0106] ④ Teaching effectiveness evaluation Teaching evaluation model: Establish a teaching evaluation model to evaluate the effectiveness of different teaching methods and curriculum settings. For example, use methods such as factor analysis and structural equation modeling to evaluate teaching quality.

[0107] Analysis of student performance: Analyze students' performance under different courses and teaching methods to identify outstanding students and those in need of help. For example, use clustering analysis to divide students into different performance groups.

[0108] Teacher performance evaluation: Evaluate teachers' teaching effectiveness to help formulate teacher training and development plans. For example, analyze teachers' classroom interaction data and students' feedback.

[0109] Data application and decision support: Data application: Apply the results of analysis and mining to actual business scenarios, such as personalized teaching, teacher performance evaluation, scientific research project management, etc. Through data-driven decision-making, improve the efficiency and effectiveness of teaching, scientific research and training.

[0110] Decision support: Provide data support for education administrators and decision-makers to help them make scientific decisions. For example, through data analysis reports and dashboards, education administrators can understand the teaching quality and scientific research progress of schools in real time and adjust strategies in a timely manner.

[0111] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A data governance architecture for teaching, scientific research and training scenarios, characterized by: It includes the following modules: Data collection and preprocessing module: used to collect teaching, scientific research and training related data from various data sources, including but not limited to teaching and research documents, teaching system records, scientific research project data, training materials, etc.; and perform preprocessing operations such as parsing, deduplication, format conversion and data verification on the collected data. Data storage module: uses distributed storage technology to store massive preprocessed data to ensure efficient storage and management of data; at the same time, establish data management strategies, including data backup, recovery and access control. Data governance module: classify and label stored data, formulate classification standards according to data characteristics and application requirements, and add metadata tags to data manually or automatically; implement data security and privacy protection measures, use data encryption technology, strict access control and data backup and recovery strategies to ensure the security and integrity of data during transmission, storage and use; Data quality management is performed through the dual verification of artificial + intelligent models. Data analysis and mining module: Use intelligent analysis tools such as data mining, machine learning, and deep learning to conduct in-depth analysis and mining of stored data. The analysis and mining include but are not limited to learning analysis, teacher performance analysis, and course effect analysis. Data visualization and application module: Use data visualization tools to display the analysis results in the form of intuitive charts, reports, etc., and generate customized data reports according to the needs of different users.

2. According to the data governance architecture for teaching, scientific research and training scenarios described in claim 1, it is characterized by: The data analysis in the data collection and preprocessing module adopts a combination of program libraries and large models, wherein the program libraries include but are not limited to pdfminer, python-docx, etc., which are used to parse document data in different formats.

3. According to the data governance architecture for teaching, scientific research and training scenarios described in claim 1, it is characterized by: The distributed storage technologies adopted by the data storage module include Hadoop, Spark, etc., to achieve distributed storage and parallel processing of data.

4. According to claim 1, the data governance architecture for teaching, scientific research and training scenarios is characterized by: The data classification and labeling in the data governance module adopts a rule-based or machine learning classification method to automatically identify the categories and features of the data and assign corresponding labels to them.

5. According to the data governance architecture for teaching, scientific research and training scenarios described in claim 1, it is characterized by: The data encryption technologies adopted in the data security and privacy protection measures in the data governance module include AES, RSA and TLS / SSL, etc., to ensure the confidentiality and integrity of data during transmission and storage; the access control limits their access to and operations on data by assigning roles and permissions to different user groups.

6. According to claim 1, the data governance architecture for teaching, scientific research and training scenarios is characterized by: The intelligent analysis tools in the data analysis and mining module can select appropriate algorithms and models for data mining and analysis according to different analysis objectives and data characteristics, such as association rule mining, classification algorithms, cluster analysis, etc.

7. According to claim 1, the data governance architecture for teaching, scientific research and training scenarios is characterized by: The data visualization tools in the data visualization and application module include Tableau, Power BI, etc. The data visualization tools display complex data in an intuitive and easy-to-understand graphical manner.

8. According to claim 1, the data governance architecture for teaching, scientific research and training scenarios is characterized by: It also includes a data expansion submodule, which is located between the data preprocessing and data governance modules, and is used to expand the data before the data enters the storage layer and during the data governance process, including expansion based on existing data and data expansion based on large model self-distillation. The expansion based on existing data adopts multi-stage expansion methods such as key information extraction, information rewriting and content reorganization, and uses Pangu large model and vLLM reasoning acceleration framework to increase the number and diversity of documents. The data expansion based on large model self-distillation maintains the document pool, uses Pangu large model to perform iterative self-distillation operations, generates high-quality expanded documents, and performs deduplication processing.

9. According to claim 1, the data governance architecture for teaching, scientific research and training scenarios is characterized by: It also includes a data labeling submodule, which combines automatic labeling with manual labeling, takes the Pangu large model as the core, introduces a multi-agent collaboration framework, formulates detailed labeling guidelines, and conducts multiple rounds of data labeling. Each round of labeling includes steps such as document collection, expansion, screening, and automatic instruction labeling. After labeling, some data is randomly selected for training and testing of the large model.

10. A data governance method applicable to the field of teaching, scientific research and training, characterized in that: The method includes the following steps: Step 1: Start data collection: Use the data collection and preprocessing module to determine the type and source of data to be collected, such as teaching and research documents, teaching system records, etc. Perform data collection: Collect data from the determined data source using methods such as web crawlers, data interfaces, sensor technology or questionnaires. Data preprocessing: In the same module, perform operations such as parsing, deduplication, format conversion and data verification on the collected data, and use special program libraries (such as pdfminer, python-docx) to parse documents of different formats. Step 2: Transfer to data storage: Transfer the preprocessed data to the data storage module. Data storage implementation: In the data storage module, use distributed storage technologies such as Hadoop and Spark for data storage, and establish data management strategies, including data backup, recovery and access control. Step 3: Data classification and labeling: In the data governance module, classify and label the stored data, formulate classification standards and add metadata tags. Data security and privacy protection: In the same module, implement data security and privacy protection measures, including data encryption (such as AES, RSA), access control and data backup and recovery strategies. Data quality management: Data quality management is performed through dual verification of manual and intelligent models. Step 4: Data expansion after preprocessing: In the data governance module, the data expansion submodule is started to expand the data during data preprocessing and data governance. Expansion based on existing data: Multi-stage expansion methods such as key information extraction, information rewriting and content reorganization are used to increase the scale and diversity of the data set. Expansion based on large model self-distillation: By maintaining the document pool, the large model is used for iterative self-distillation operations to generate high-quality expanded documents and perform deduplication processing. Step 5: Data annotation startup: In the data governance module, the data annotation submodule is started to formulate detailed annotation guidelines. Combination of automatic and manual annotation: The combination of automatic annotation and manual annotation is adopted, with the Pangu large model as the core, and the multi-agent collaboration framework is introduced for data annotation. Multi-round data annotation: Multi-round data annotation process is carried out, including document collection, expansion, screening and automatic instruction annotation, and some data are randomly selected for training and testing of the large model. Step 6: Data transmission to analysis: The governed data is sent to the data analysis and mining module. Selection of intelligent analysis tools: Select appropriate intelligent analysis tools such as data mining, machine learning, and deep learning tools according to the analysis objectives. Perform analysis and mining: Perform specific analysis and mining tasks such as learning analysis, teacher performance analysis, and course effectiveness analysis. Step 7: Transfer to data visualization: Transfer the analysis results to the data visualization and application module. Data visualization implementation: Use data visualization tools such as Tableau and Power BI to display complex data in an intuitive and easy-to-understand graphical way. Customized data report generation: Generate customized data reports according to the needs of different users. Step 8: Data post-annotation review: After the data is annotated, start the data review submodule. Format review and applicability assessment: Perform format review and applicability assessment on the annotated data to ensure the requirements of the data task. Manual evaluation: Randomly sample the annotated data through manual evaluation.