Campus data governance management method and system based on machine learning

By adopting machine learning-based data governance management methods and systems in campus data governance, the problems of inefficient data governance and difficult to ensure data quality in traditional methods are solved, and the intelligent and automated processing of data is realized, adapting to business changes, and providing high-quality data support.

CN120196670AActive Publication Date: 2025-06-24MEIZHOU BAY VOCATIONAL & TECH COLLEGE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510686279.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Traditional campus data governance methods cannot effectively respond to the complex governance needs of large-scale, multi-source heterogeneous data, resulting in inefficient data governance, difficult to guarantee data quality, and lack of intelligence and automation capabilities, so it is impossible to deeply mine and analyze data.

Method used

Adopting campus data governance management methods and systems based on machine learning, the data is structured through the data middle platform, generating data governance strategies, updating data element information, and realizing automatic adaptation and efficient flow of data interfaces to build an adaptive data governance strategy system.

Benefits of technology

It has significantly improved the automation and intelligence level of campus data governance, ensured data quality and accuracy, realized in-depth data mining and analysis, adapted to business changes and data feature evolution, and provided flexible and reliable data governance support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196670A_ABST
    Figure CN120196670A_ABST
Patent Text Reader

Abstract

The invention provides a campus data governance management method and system based on machine learning, and the method comprises the steps: responding to a data obtaining request of a data terminal, and transmitting the request to a data medium station; triggering a data processing flow according to the data identification field returned by the data medium station and the target interface configuration; performing structured processing on the original campus data through machine learning to generate treatment target data; generating a governance strategy based on the governance target data and updating the data element information; if the terminal interface configuration is not deployed, associating the updated meta-information with a preset governance field to generate a meta-information set; based on the meta-information set, updating platform interface deployment parameters in the data; returning the data identification field to the terminal; and responding to a terminal assignment operation, and executing interface calling and returning governing target data through the data middle table. According to the invention, automatic management and efficient circulation of campus data can be realized, and data quality, interface compatibility and system response efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data governance and machine learning. More specifically, the present invention relates to a campus data governance management method and system based on machine learning. Background Art

[0002] In today's digital age, the construction of campus informatization is advancing rapidly, and a vast amount of data has been accumulated on campus, covering various fields such as teaching, scientific research, management, and student life. These data are not only huge in quantity but also have a wide range of sources and diverse formats, including structured data, semi-structured data, and unstructured data. For example, structured data such as student course grades and course selection records stored in the teaching management system, semi-structured data such as library borrowing records, and unstructured data such as photos and videos of campus activities. The effective governance of these data is of crucial significance for improving campus management efficiency, optimizing teaching resource allocation, promoting scientific research innovation, and enhancing students' learning experience.

[0003] Traditional campus data governance methods mainly rely on manual operations or simple data processing tools. The manual governance method is not only inefficient but also prone to errors and difficult to cope with the complex processing requirements of large-scale data. For example, when manually conducting statistical analysis on students' grade data, data entry errors may occur due to negligence, which may affect subsequent decision-making. Although simple data processing tools can achieve some basic data cleaning and sorting functions, they usually lack intelligence and automation capabilities and are unable to deeply mine and analyze data, making it difficult to meet the growing complex needs of campus data governance. For example, these tools cannot automatically identify problems such as outliers and duplicate records in the data, nor can they automatically adjust data governance strategies according to data characteristics.

[0004] With the continuous development of information technology, machine learning technology has been widely applied in the field of data processing. Machine learning algorithms can automatically discover the rules and characteristics in data through learning and analysis of a large amount of data, thereby achieving intelligent data processing and analysis. For example, in the financial field, machine learning algorithms can be used for risk assessment and fraud detection, and by learning a large amount of transaction data, they can automatically identify potential risk transaction patterns. In the medical field, machine learning algorithms can be used for disease diagnosis and treatment plan recommendation, and by learning patients' medical record data and examination data, they can provide auxiliary diagnostic suggestions for doctors. However, currently, the application of machine learning technology in the field of campus data governance is relatively less, and its potential in improving the efficiency and quality of campus data governance has not been fully exploited.

[0005] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: The traditional campus data governance method cannot effectively meet the complex governance requirements of large-scale, multi-source heterogeneous data, resulting in low data governance efficiency and difficult to guarantee data quality; The existing methods lack intelligent and automated capabilities, cannot automatically adjust governance strategies according to data characteristics, and are difficult to achieve in-depth data mining and analysis; The prior art has not fully utilized the advantages of machine learning technology and failed to fully exert its potential in improving the efficiency and quality of campus data governance. Summary of the Invention

[0006] The present invention provides a campus data governance management method and system based on machine learning.

[0007] In the first aspect of the present invention, a campus data governance management method based on machine learning is provided, including: In response to detecting a data acquisition request message of a target data set by a data terminal, sending the acquisition request message of the target data set to a data middle platform; In response to receiving the data identification field information corresponding to the target data set sent by the data middle platform, triggering a data processing process of the data middle platform according to the data identification field information and the target interface configuration information of the data middle platform; Structurally processing the original campus data corresponding to the data identification field information through the data middle platform, where the structural processing includes data feature extraction processing based on a machine learning algorithm to obtain governance target data; Generating data governance strategy information corresponding to the governance target data according to the governance target data; Performing model parameter update processing on each data element information corresponding to the target data set according to the data governance strategy information, where each data element information includes a meta-information representing data identification; In response to determining that the interface configuration information corresponding to the data terminal is not deployed, associating each updated data element information with the preset governance field configuration of the data middle platform to generate a meta-information set; Updating the interface deployment parameters of the data middle platform based on the meta-information set and the preset data interface configuration information corresponding to the governance target data; Sending the data identification field information associated with the target data set to the data terminal; In response to detecting an assignment operation of the preset governance field configuration of the preset data interface configuration information by the data terminal, converting the assigned preset governance field configuration into an interface call parameter through the data middle platform and performing a data interface call operation to return the governance target data to the data terminal.

[0008] Further, triggering the data processing process of the data middle platform includes: Extract the target field mapping rules for the corresponding target interface configuration information from the interface configuration library of the data middle platform; Inject the data identification field information into the parameter container corresponding to the target field mapping rules; Execute parameterized interface calls through the interface call engine of the data middle platform to obtain the original campus data associated with the data identification field information.

[0009] Furthermore, the structured processing includes: Perform data quality verification processing on the original campus data to generate basic data that passes the verification; Through the machine learning algorithms integrated in the data middle platform, perform feature space transformation processing on the basic data to generate governance target data with standardized data structure features.

[0010] Furthermore, the model parameter update processing includes: Generate data governance feature vectors through data governance policy information; Input the data governance feature vectors into the meta-information update module of the data middle platform to calculate the weight adjustment parameters for each data meta-information; Update the metadata storage structure of each data meta-information according to the weight adjustment parameters.

[0011] Furthermore, the association processing includes: Align the feature dimensions of the updated data meta-information with the preset governance field configuration; Through the association relationship learning module of the data middle platform, calculate the semantic similarity parameters between the data meta-information and the preset governance field configuration; Construct a meta-information set with a hierarchical structure based on the semantic similarity parameters.

[0012] Furthermore, the interface deployment parameter update includes: Convert the meta-information set into an interface description file of the data middle platform; Through the model-driven interface generator, compile the interface description file into an executable data interface configuration; Deploy the data interface configuration to the interface service gateway of the data middle platform.

[0013] Furthermore, the data quality verification processing includes: Obtain the field integrity parameters of the original campus data through the data verification engine of the data middle platform; When the field integrity parameters reach the preset verification threshold, activate the data completion module of the data middle platform; Perform interpolation calculation processing on the missing fields through the data completion module to generate basic data that passes the verification.

[0014] Furthermore, the feature space transformation processing includes: Invoke the dimensionality reduction algorithm library in the data middle platform to perform principal component analysis on the basic data; Through the standardization module in the data middle platform, perform Z-score standardization calculation on the data after principal component analysis; Write the standardization calculation result into the feature storage library of the data middle platform to generate governance target data with standardized data structure features.

[0015] Furthermore, the construction of semantic similarity parameters includes: Through the word vector generator in the data middle platform, convert the data element information and the preset governance field configuration into a set of word vectors; Based on the attention mechanism module in the data middle platform, calculate the cosine similarity matrix of the set of word vectors; Through the graph neural network in the data middle platform, generate a hierarchical meta-information set according to the cosine similarity matrix.

[0016] In the second aspect of the present invention, a campus data governance management system based on machine learning is provided, including: A detection module, configured to send the acquisition request information of the target data set to the data middle platform in response to detecting the data acquisition request information of the data terminal for the target data set; An interface trigger module, configured to trigger the data processing process of the data middle platform according to the data identification field information and the target interface configuration information of the data middle platform in response to receiving the data identification field information of the corresponding target data set sent by the data middle platform; A data processing engine, integrated in the data middle platform, configured to perform structured processing on the original campus data corresponding to the data identification field information, where the structured processing includes data feature extraction processing based on machine learning algorithms to obtain governance target data; A policy generation module, configured to generate data governance policy information corresponding to the governance target data according to the governance target data; A parameter update module, configured to perform model parameter update processing on each data element information corresponding to the target data set according to the data governance policy information, where each data element information includes meta-information representing data identification; An association processing module, configured to perform association processing on each updated data element information and the preset governance field configuration of the data middle platform in response to determining that the interface configuration information of the corresponding data terminal is not deployed, to generate a set of meta-information; An interface deployment module, configured to update the interface deployment parameters of the data middle platform based on the set of meta-information and the preset data interface configuration information corresponding to the governance target data; A field transmission module, configured to send the data identification field information associated with the target data set to the data terminal; An interface execution engine, integrated in the data middle platform, is configured to, in response to detecting an assignment operation for a preset governance field configuration of preset data interface configuration information by a data terminal, convert the assigned preset governance field configuration into interface call parameters, and execute a data interface call operation to return governance target data to the data terminal.

[0017] The above embodiments of the present invention have at least the following beneficial effects: 1. By adopting machine learning algorithms to perform automated data quality verification, feature space transformation, and structured processing on original campus data, it can effectively solve the problems existing in traditional manual data governance methods, such as low processing efficiency, strong subjectivity, and high error rates, ensure that the output data has standardized data structure characteristics and high-quality data content, and provide a reliable basis for subsequent data analysis and applications.

[0018] 2. By establishing a dynamic association mechanism between data identification fields and preset governance fields, combined with model-driven interface automatic generation technology, it can effectively solve the problems of inconsistent multi-source heterogeneous data interface standards and poor compatibility commonly existing in campus informatization systems, realize automatic adaptation and efficient call of data interfaces between different systems, and greatly improve the accuracy and real-time performance of cross-system data interaction.

[0019] 3. Through the intelligent calculation of data governance feature vectors and the dynamic adjustment of meta-information weights, combined with semantic similarity analysis and hierarchical structure construction technology, it can effectively solve the problem that traditional static data governance strategies are difficult to adapt to campus business changes and data feature evolution, enable the data governance strategy to be automatically optimized and adjusted according to actual business needs and data feature changes, and provide flexible and reliable data governance support for the continuous expansion and upgrade of campus informatization construction. Description of the Drawings

[0020] By referring to the following detailed description with reference to the drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become easy to understand. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, where: Figure 1 It is a flowchart of a method for managing campus data governance based on machine learning provided by an embodiment of the present invention; Figure 2 It is a structural diagram of a system for managing campus data governance based on machine learning provided by an embodiment of the present invention; Figure 3 It schematically shows a structural diagram of an electronic device according to an embodiment of the present invention. Detailed Embodiments

[0021] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, and do not limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to be able to fully convey the scope of the present invention to those skilled in the art.

[0022] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method or computer program product. Therefore, the present invention can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0023] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0024] The following reference Figure 1 , Figure 1 is a schematic flowchart of a campus data governance management method based on machine learning provided for an embodiment of the present invention. As Figure 1 shown, a campus data governance management method based on machine learning includes: S1. In response to detecting a data acquisition request message of a target data set by a data terminal, sending the acquisition request message of the target data set to a data middle platform; S2. In response to receiving the data identification field information corresponding to the target data set sent by the data middle platform, triggering a data processing process of the data middle platform according to the data identification field information and the target interface configuration information of the data middle platform; S3. Structurally processing the original campus data corresponding to the data identification field information through the data middle platform, where the structural processing includes data feature extraction processing based on a machine learning algorithm to obtain governance target data; S4. Generating data governance policy information corresponding to the governance target data according to the governance target data; S5. Performing model parameter update processing on each data element information corresponding to the target data set according to the data governance policy information, where each data element information includes meta-information representing data identification; S6. In response to determining that the interface configuration information corresponding to the data terminal is not deployed, associating each updated data element information with the preset governance field configuration of the data middle platform to generate a meta-information set; S7. Updating the interface deployment parameters of the data middle platform based on the meta-information set and the preset data interface configuration information corresponding to the governance target data; S8. Send the data identification field information associated with the target data set to the data terminal; S9. In response to detecting an assignment operation on the preset governance field configuration of the preset data interface configuration information by the data terminal, convert the assigned preset governance field configuration into interface call parameters through the data middle platform, and execute a data interface call operation to return the governance target data to the data terminal.

[0025] It should be noted that when a data acquisition request message from the data terminal is detected, the system will send this request message to the data middle platform. Here, the data terminal refers to various devices or systems on campus for acquiring data, such as the educational administration management system, the library management system, etc., which can initiate requests for specific data sets. The target data set refers to the specific data set requested by the terminal, such as the student information data set or the course grade data set. The data middle platform is the core platform for campus data governance, responsible for receiving requests, processing data, and returning results, etc. This process realizes the initial transfer of the data acquisition request and lays a foundation for subsequent data processing.

[0026] Specifically, the data acquisition request message includes information such as the identity identifier of the terminal, the name of the required data set, and the specific parameters of the request, etc. For example, when the educational administration management system, as a data acquisition terminal, requests the student information data set, it will carry its own system identifier, the data set name student information, and request parameters such as querying student information of a specific major. After receiving this information, the data middle platform will parse and process it according to the preset interface configuration. The data identification field information refers to the field information that can uniquely identify the target data set. For example, in the student information data set, the student ID field can be used as the data identification field. The target interface configuration information refers to the interface configuration in the data middle platform related to the target data set, including the address of the interface, the parameter format, etc. Through this information, the data middle platform can trigger the corresponding data processing process to ensure the correct acquisition and processing of data.

[0027] Preferably, when triggering the data processing process of the data middle platform, the target field mapping rule corresponding to the target interface configuration information can be extracted from the interface configuration library of the data middle platform. For example, for a request for the student information data set, the mapping relationship between fields such as student ID, name, and major and interface parameters will be stored in the interface configuration library. After injecting the data identification field information into the parameter container corresponding to the target field mapping rule, execute a parameterized interface call through the interface call engine of the data middle platform to obtain the original campus data associated with the data identification field information. This process ensures the accuracy and efficiency of data acquisition and provides a reliable data basis for subsequent data governance.

[0028] In some embodiments, triggering the data processing process of the data middle platform includes: Extract the target field mapping rules for the corresponding target interface configuration information from the interface configuration library of the data middle platform; Inject the data identification field information into the parameter container corresponding to the target field mapping rules; Execute parameterized interface calls through the interface call engine of the data middle platform to obtain the original campus data associated with the data identification field information.

[0029] It should be noted that when triggering the data processing flow of the data middle platform, it is first necessary to extract the target field mapping rules for the corresponding target interface configuration information from the interface configuration library of the data middle platform. Here, the interface configuration library is an important part of the data middle platform, which stores all configuration information related to data interfaces, including field mapping rules, interface addresses, parameter formats, etc. The target field mapping rules refer to the specific rules for mapping data identification field information to interface parameters, and these rules define how to convert the field information in the request into the parameter form that the interface can recognize. Then, inject the data identification field information into the parameter container corresponding to the target field mapping rules, and this process ensures that the data identification field information can be correctly passed to the interface call engine. Finally, execute parameterized interface calls through the interface call engine of the data middle platform to obtain the original campus data associated with the data identification field information, thereby completing the triggering of the data processing flow and providing basic data support for subsequent data governance.

[0030] Specifically, the data middle platform is a platform that integrates data storage, processing, and management. It manages the configuration information of various data interfaces through the interface configuration library. The target field mapping rules in the interface configuration library are preset according to different data interface requirements. For example, when processing the student information dataset, the target field mapping rules include mapping the student ID field to the interface parameter student_id, mapping the name field to the interface parameter name, etc. These mapping rules ensure the consistency and accuracy of data format during transmission. The parameter container is a temporary storage area for storing and passing parameters. It can receive the data identification field information and convert it into the parameter format defined by the target field mapping rules. The interface call engine is a module in the data middle platform for executing interface calls. It executes the corresponding interface call operations according to the parameters in the parameter container to obtain the required original campus data. These components work together to ensure the efficient and accurate triggering of the data processing flow.

[0031] Preferably, when performing a parameterized interface call, the interface call engine sends a request to the data storage layer according to the parameter values in the parameter container, in accordance with a preset interface protocol and parameter format. For example, if the interface protocol is RESTful API, the interface call engine constructs an HTTP request and sends the parameters in the parameter container as query parameters or the request body of the request. After obtaining the original campus data, the interface call engine performs preliminary parsing and verification on the data to ensure the integrity and accuracy of the data. If it is found that data is missing or in the wrong format, the interface call engine triggers the data completion module to perform interpolation calculation processing on the missing fields to generate basic data that passes the verification. This process not only improves the reliability of data acquisition but also provides a high-quality data foundation for subsequent data governance.

[0032] In some embodiments, the structuring process includes: Performing data quality verification processing on the original campus data to generate basic data that passes the verification; Performing feature space transformation processing on the basic data through machine learning algorithms integrated in the data middle platform to generate governance target data with standardized data structure features.

[0033] It should be noted that the structuring process is a series of processing operations performed on the original campus data, aiming to convert unstructured or semi-structured data into governance target data with standardized data structure features. This process includes two main steps: data quality verification processing and feature space transformation processing. Data quality verification processing is to check the original data to ensure the integrity and accuracy of the data and generate basic data that passes the verification. Feature space transformation processing is to use machine learning algorithms to process the basic data, extract data features, and generate standardized governance target data. Through this series of processing, the quality and usability of the data can be effectively improved, providing support for subsequent data analysis and applications.

[0034] Specifically, data quality verification and processing refer to a series of inspection operations performed on the original campus data to ensure data integrity and accuracy. This process typically includes checking the field integrity of the data, whether the data types are correct, and whether there are duplicate records. For example, for the student information dataset, check whether the student ID field is missing and whether the name field is empty. Basic data refers to the data that meets the basic quality requirements after data quality verification and processing. These data are the basis for subsequent feature space transformation processing. Feature space transformation processing refers to using machine learning algorithms to process the basic data, extract the features of the data, and transform them into governance target data with standardized data structure features. This process typically includes operations such as data dimensionality reduction and standardization. Machine learning algorithms refer to a series of algorithms used for data processing, which can automatically learn the features and rules of the data. For example, the principal component analysis (PCA) algorithm is used for dimensionality reduction, and the Z-score standardization algorithm is used for data standardization.

[0035] Preferably, during the data quality verification and processing, the field integrity parameters of the original campus data can be obtained through the data verification engine of the data middle platform. If the field integrity parameters do not reach the preset verification threshold, the data completion module will be activated to perform interpolation calculation processing on the missing fields and generate the basic data that passes the verification. For example, for the missing grade field in the student information dataset, the average value interpolation method can be used for completion. During the feature space transformation processing, the dimensionality reduction algorithm library of the data middle platform can be called to perform principal component analysis processing on the basic data and extract the main features of the data. Then, through the standardization module of the data middle platform, the Z-score standardization calculation is performed on the data after principal component analysis processing to transform the data into standard normal distribution data with a mean of 0 and a standard deviation of 1. Finally, the standardized calculation results are written into the feature storage library of the data middle platform to generate governance target data with standardized data structure features. This process not only improves the quality of the data but also provides a standardized data basis for subsequent data governance and analysis.

[0036] In some embodiments, the model parameter update processing includes: Generating a data governance feature vector through data governance policy information; Inputting the data governance feature vector into the meta-information update module of the data middle platform to calculate the weight adjustment parameters of each data meta-information; Updating the metadata storage structure of each data meta-information according to the weight adjustment parameters.

[0037] It should be noted that the model parameter update process is a process of adjusting each data element information corresponding to the target data set based on the data governance strategy information. The purpose of this process is to dynamically optimize the storage structure and weight of data elements according to the governance strategy to better reflect the importance of data and governance requirements. Among them, the data governance strategy information refers to a series of guiding information generated based on the governance target data, which is used to guide the specific operations of data governance. The data element information refers to the basic unit information that constitutes the data set, including data identification, data type, data source, etc. The model parameter update process updates the metadata storage structure of the data element information by calculating the weight adjustment parameters, thereby optimizing the effect of data governance.

[0038] Specifically, the data governance strategy information is generated through the analysis of the governance target data, and it contains a series of feature vectors and rules for guiding data governance. These feature vectors can represent the importance and relevance of data. The data governance feature vector is a vector generated based on the governance target data, which reflects the characteristics of the data and governance requirements. For example, for the student information data set, the data governance feature vectors include the importance weights of fields such as student ID, name, and grades. The meta-information update module is a module in the data middle platform, which is responsible for calculating the weight adjustment parameters of each data element information according to the data governance feature vectors. These weight adjustment parameters are used to update the metadata storage structure of the data element information, such as adjusting the storage location of the data element and optimizing the data access path.

[0039] Preferably, in the process of model parameter update, first generate data governance feature vectors through the data governance strategy information. This process can be based on the statistical analysis of data, such as calculating statistical quantities such as variance and mean of each field to determine the importance of the field. Then, input the data governance feature vectors into the meta-information update module of the data middle platform, and this module will calculate the weight adjustment parameters of each data element information according to the feature vectors. For example, for the grade field, if its variance is large, it means that the distribution of grade data is relatively scattered and requires a higher weight. Finally, update the metadata storage structure of each data element information according to the weight adjustment parameters, such as adjusting the index structure of the data element to optimize the data query efficiency. This process not only improves the flexibility of data governance but also enhances the adaptability of data governance, and can dynamically adjust the storage and management methods of data elements according to different governance requirements.

[0040] In some embodiments, the association process includes: Align the feature dimensions of the updated data element information with the preset governance field configuration; Through the association relationship learning module of the data middle platform, calculate the semantic similarity parameters between the data element information and the preset governance field configuration; Construct a hierarchical meta-information set based on the semantic similarity parameters.

[0041] It should be noted that the association process is a process of integrating the updated data element information with the preset governance field configuration, aiming to construct a hierarchical meta-information set. This process is achieved through feature dimension alignment and semantic similarity calculation to ensure semantic consistency and logical relevance between the data element information and the governance field configuration. Among them, feature dimension alignment refers to adjusting the feature dimensions of the data element information and the governance field configuration so that effective semantic similarity calculation can be performed. The semantic similarity parameter is a parameter that measures the semantic similarity degree between the data element information and the governance field configuration, and a hierarchical meta-information set can be constructed through the calculated similarity parameter.

[0042] Specifically, the updated data element information refers to the data element information after model parameter update processing, and these information include the latest weight adjustment parameters and optimized storage structures. The preset governance field configuration refers to the field configuration predefined by the data middle platform for guiding the specific operations of data governance. These configurations include information such as the name, type, and weight of the fields. Feature dimension alignment refers to adjusting the feature dimensions of the data element information and the governance field configuration so that effective semantic similarity calculation can be performed. For example, if the feature dimensions of the data element information are student ID, name, and score, and the feature dimensions of the governance field configuration are student ID, student name, and exam score, then the student ID should be aligned with the student ID, name with the student name, and score with the exam score. The semantic similarity parameter is obtained through calculation and is used to measure the semantic similarity degree between the data element information and the governance field configuration. For example, cosine similarity can be used to calculate the similarity between two vectors.

[0043] Preferably, in the association process, first, the feature dimensions of the updated data element information and the preset governance field configuration are aligned. This process can be achieved through the association relationship learning module of the data middle platform, which will automatically identify and align the feature dimensions of the data element information and the governance field configuration according to the preset rules and algorithms. Then, through the word vector generator of the data middle platform, the data element information and the governance field configuration are converted into a word vector set. Next, based on the attention mechanism module of the data middle platform, the cosine similarity matrix of the word vector set is calculated. Finally, through the graph neural network of the data middle platform, a hierarchical meta-information set is generated according to the cosine similarity matrix. This process not only improves the relevance between the data element information and the governance field configuration but also enhances the intelligent level of data governance. For example, for the student information dataset, by calculating the cosine similarity between the student ID and the student ID, name and the student name, and score and the exam score, a hierarchical meta-information set can be constructed to achieve efficient management and governance of the data.

[0044] In some embodiments, the interface deployment parameter update includes: Convert the meta - information set into an interface description file for the data middle - platform; Through a model - driven interface generator, compile the interface description file into an executable data interface configuration; Deploy the data interface configuration to the interface service gateway of the data middle - platform.

[0045] It should be noted that the interface deployment parameter update is a process of converting the meta - information set into an interface description file for the data middle - platform, compiling it into an executable data interface configuration through a model - driven interface generator, and finally deploying these configurations to the interface service gateway of the data middle - platform. The purpose of this process is to ensure that the data middle - platform can dynamically adjust the interface configuration according to the updated meta - information set, so as to better support data governance and data access requirements. Among them, the interface description file is a file that defines information such as interface functions, parameters, and return values, and it is the basis for interface generation. The model - driven interface generator is a tool that can generate specific interface code according to the interface description file. The interface service gateway is a component of the data middle - platform, responsible for managing and routing interface requests.

[0046] Specifically, the meta - information set refers to a set containing data meta - information and governance field configurations generated after association processing. It has a hierarchical structure and can reflect the logical relationship between data elements. The interface description file is a standardized file that details information such as the interface name, input parameters, output parameters, return value type, etc. For example, for a student information query interface, the interface description file will define the interface name as GetStudentInfo, the input parameter as student_id, and the output parameters as name, age, score, etc. The model - driven interface generator is a tool based on the model - driven architecture MDA that can automatically generate interface code according to the interface description file, reducing the workload and errors of manual coding. The interface service gateway is a middleware that is responsible for receiving external requests, routing the requests to the corresponding interface processing modules according to the requested interface name and parameters, and returning the processing results.

[0047] Preferably, during the process of updating the interface deployment parameters, the meta - information set is first converted into an interface description file. This process can be achieved through the conversion module of the data center. The module will generate an interface description file that conforms to the predefined format based on the hierarchical structure and content of the meta - information set. Then, through a model - driven interface generator, the interface description file is compiled into an executable data interface configuration. The interface generator will generate specific interface code according to the definitions in the interface description file, such as generating RESTful API interface code. Finally, the generated interface configuration is deployed to the interface service gateway of the data center. During the deployment process, the interface service gateway will load the interface configuration to ensure that the interface can run normally and provide services externally. For example, for the student information query interface, the interface service gateway will receive the query request according to the interface configuration, call the backend data processing module to obtain student information, and return the result to the requester. This process not only improves the efficiency of interface deployment but also enhances the maintainability and scalability of the interface.

[0048] In some embodiments, the data quality verification process includes: Obtain the field integrity parameters of the original campus data through the data verification engine of the data center; When the field integrity parameters reach the preset verification threshold, activate the data completion module of the data center; Perform interpolation calculation processing on the missing fields through the data completion module to generate basic data that passes the verification.

[0049] It should be noted that the data quality verification process is a process of checking the field integrity of the original campus data and activating the data completion module when necessary to process the missing fields. The purpose of this process is to ensure the integrity and accuracy of the data and provide high - quality basic data for subsequent data governance. Among them, the field integrity parameter is a parameter used to measure whether a data field is missing, which reflects the integrity degree of the data. The data completion module is a tool for processing missing data, which can complete the missing fields through methods such as interpolation calculation.

[0050] Specifically, data quality verification processing refers to a series of inspection operations performed on the original campus data to ensure the integrity and accuracy of the data. This process typically includes checking the field integrity of the data, whether the data types are correct, and whether there are duplicate records. The field integrity parameter is a quantitative indicator used to measure the integrity of data fields. For example, for a student information dataset containing fields such as student ID, name, and grades, the field integrity parameter can be obtained by calculating the ratio of the number of missing fields to the total number of fields. If the field integrity parameter does not reach the preset verification threshold, for example, the number of missing fields exceeds 10% of the total number of fields, the data completion module is activated. The data completion module will perform interpolation calculation processing on the missing fields according to the preset completion strategy. For example, for the missing grades field, the average value interpolation method can be used, that is, the average value of the student's other course grades is used to complete the missing grades field.

[0051] Preferably, during the data quality verification processing, first, the field integrity parameter of the original campus data is obtained through the data verification engine of the data middle platform. The data verification engine will scan each field in the dataset, count the number of missing fields, and calculate the field integrity parameter. If the field integrity parameter does not reach the preset verification threshold, for example, the number of missing fields exceeds 10% of the total number of fields, the data completion module is activated. The data completion module will select an appropriate interpolation method according to the type and distribution of the data. For example, for numerical data such as grades, the average value interpolation method can be used; for categorical data such as gender, the mode interpolation method can be used. The completed data will be checked again by the data verification engine to ensure that the completed data meets the quality requirements. This process not only improves the integrity of the data but also provides a reliable data basis for subsequent data governance and analysis.

[0052] In some embodiments, the feature space transformation processing includes: Invoking the dimensionality reduction algorithm library of the data middle platform to perform principal component analysis processing on the basic data; Performing Z-score standardization calculation on the data after principal component analysis processing through the standardization module of the data middle platform; Writing the standardization calculation result into the feature storage library of the data middle platform to generate the governance target data with the characteristics of a standardized data structure.

[0053] It should be noted that the feature space transformation process is a series of processing operations performed on the basic data, aiming to convert the data into governance target data with standardized data structure features. This process includes calling the dimensionality reduction algorithm library to perform principal component analysis on the basic data, and performing Z-score standardization calculation on the processed data through the standardization module. Through these steps, the dimensionality of the data can be effectively reduced while ensuring the standardization of the data, thereby improving the usability and processing efficiency of the data. Among them, the dimensionality reduction algorithm library refers to a collection containing multiple dimensionality reduction algorithms, which is used to reduce the dimensionality of the data; the Z-score standardization calculation is a data standardization method, which is used to convert the data into standard normal distribution data with a mean of 0 and a standard deviation of 1.

[0054] Specifically, the feature space transformation process refers to using dimensionality reduction algorithms and standardization methods to process the basic data to extract the main features of the data and reduce the dimensionality of the data. This process usually includes two main steps: dimensionality reduction processing and standardization processing. The dimensionality reduction algorithm library is a collection containing multiple dimensionality reduction algorithms, such as principal component analysis PCA, linear discriminant analysis LDA, etc. These algorithms improve the efficiency of data processing by reducing the dimensionality of the data while retaining the main features of the data. The principal component analysis process is a commonly used dimensionality reduction method, which reduces the dimensionality of the data by projecting the data onto the principal component directions. For example, for a student information dataset containing multiple features, principal component analysis can project these features onto a few principal components, thereby reducing the dimensionality of the data. The Z-score standardization calculation is a data standardization method, which converts the data into standard normal distribution data with a mean of 0 and a standard deviation of 1 by calculating the mean and standard deviation of the data. This process can eliminate the influence of the data dimension and improve the comparability of the data.

[0055] Preferably, during the feature space transformation process, first call the dimensionality reduction algorithm library of the data middle platform to perform principal component analysis on the basic data. This process can be achieved by calculating the covariance matrix of the data and finding the eigenvalues and eigenvectors of the covariance matrix. Then, select several eigenvectors with the largest eigenvalues as the principal components, and project the data onto these principal component directions to reduce the dimensionality of the data. Next, through the standardization module of the data middle platform, perform Z-score standardization calculation on the data after principal component analysis. This process includes calculating the mean and standard deviation of the data, and then subtracting the mean from each data point and dividing by the standard deviation to obtain the standardized data. Finally, write the standardized calculation result into the feature repository of the data middle platform to generate governance target data with standardized data structure features. This process not only improves the processing efficiency of the data, but also provides a high-quality data basis for subsequent data analysis and applications.

[0056] In some embodiments, the construction of the semantic similarity parameter includes: Through the word vector generator of the data middle platform, convert the data element information and the preset governance field configuration into a set of word vectors; Based on the attention mechanism module of the data middle platform, calculate the cosine similarity matrix of the set of word vectors; Through the graph neural network of the data middle platform, generate a hierarchical meta-information set according to the cosine similarity matrix.

[0057] It should be noted that the construction of the semantic similarity parameter is a process of converting the data element information and the preset governance field configuration into word vectors and calculating the cosine similarity matrix, and finally generating a hierarchical meta-information set. The purpose of this process is to ensure the semantic consistency between the data element information and the governance field configuration through semantic analysis, thereby improving the accuracy and efficiency of data governance. Among them, the word vector generator is a tool for converting text data into word vectors for subsequent semantic similarity calculation; the attention mechanism module is an algorithm module for calculating the similarity between word vectors; the graph neural network is a neural network for processing graph-structured data for constructing a hierarchical meta-information set.

[0058] Specifically, the word vector generator is a tool that converts text data into word vectors for semantic similarity calculation. For example, for the student ID in the data element information and the student ID in the preset governance field configuration, the word vector generator will convert these two texts into corresponding word vectors. These word vectors are vectors in a high-dimensional space that can represent the semantic information of the text. The cosine similarity matrix is a matrix obtained by calculating the cosine similarity between word vectors, which reflects the similarity degree between word vectors. For example, a cosine similarity value close to 1 indicates that two word vectors are very similar, while close to 0 indicates that they are not similar. The attention mechanism module is an algorithm module that generates a cosine similarity matrix by calculating the similarity between word vectors. The attention mechanism can help the model better focus on important word vectors and improve the accuracy of similarity calculation. The graph neural network is a neural network for processing graph-structured data that can generate a hierarchical meta-information set according to the cosine similarity matrix. For example, the graph neural network can group word vectors with high similarity together to form a hierarchical structure for data governance.

[0059] Preferably, in the process of constructing the semantic similarity parameter, first, through the word vector generator in the data middle platform, the data element information and the preset governance field configuration are converted into a set of word vectors. This process can be achieved through pre-trained word embedding models, such as Word2Vec or BERT, which map the text data to word vectors in a high-dimensional space. Then, based on the attention mechanism module in the data middle platform, the cosine similarity matrix of the set of word vectors is calculated. The attention mechanism module will generate a matrix according to the similarity between the word vectors, where each element represents the similarity between two word vectors. Finally, through the graph neural network in the data middle platform, a hierarchical meta-information set is generated according to the cosine similarity matrix. The graph neural network will cluster similar word vectors together according to the similarity matrix to form a hierarchical structure, thus constructing the meta-information set. This process not only improves the semantic consistency between the data element information and the governance field configuration, but also enhances the intelligent level of data governance. For example, for the student information dataset, by calculating the cosine similarity between the student ID number and the student ID, the name and the student name, and the grades and the exam scores, a hierarchical meta-information set can be constructed to achieve efficient management and governance of the data.

[0060] The above embodiments of the present invention have the following beneficial effects: 1. It can significantly improve the automation and intelligence level of campus data governance: By using machine learning algorithms to perform automated data quality verification, feature space transformation, and structured processing on the original campus data, it can effectively solve the problems existing in the traditional manual data governance method, such as low processing efficiency, strong subjectivity, and high error rate, and ensure that the output data has standardized data structure features and high-quality data content, providing a reliable basis for subsequent data analysis and applications.

[0061] 2. It can realize the intelligent adaptation and efficient transfer of campus data interfaces: By establishing a dynamic association mechanism between the data identification field and the preset governance field, combined with the model-driven interface automatic generation technology, it can effectively solve the problems of inconsistent standards and poor compatibility of multi-source heterogeneous data interfaces commonly existing in campus information systems, realize the automatic adaptation and efficient call of data interfaces between different systems, and greatly improve the accuracy and real-time performance of cross-system data interaction.

[0062] 3. It can construct an adaptive campus data governance strategy system: Through the intelligent calculation of data governance feature vectors and the dynamic adjustment of meta-information weights, combined with semantic similarity analysis and hierarchical structure construction technology, it can effectively solve the problem that traditional static data governance strategies are difficult to adapt to campus business changes and data feature evolution, enabling the data governance strategy to be automatically optimized and adjusted according to actual business needs and data feature changes, providing flexible and reliable data governance support for the continuous expansion and upgrade of campus informatization construction.

[0063] As shown Figure 2 in the figure, a machine learning-based campus data governance management system according to some embodiments includes: A detection module 201, configured to send the acquisition request information of the target data set to the data middle platform in response to detecting the data acquisition request information of the target data set by the data terminal; An interface trigger module 202, configured to trigger the data processing process of the data middle platform according to the data identification field information and the target interface configuration information of the data middle platform in response to receiving the data identification field information of the corresponding target data set sent by the data middle platform; A data processing engine 203, integrated in the data middle platform, configured to perform structured processing on the original campus data corresponding to the data identification field information, where the structured processing includes data feature extraction processing based on machine learning algorithms to obtain the governance target data; A policy generation module 204, configured to generate data governance policy information corresponding to the governance target data according to the governance target data; A parameter update module 205, configured to perform model parameter update processing on each data element information corresponding to the target data set according to the data governance policy information, where each data element information includes meta-information representing the data identification; An association processing module 206, configured to perform association processing between each updated data element information and the preset governance field configuration of the data middle platform in response to determining that the interface configuration information of the corresponding data terminal is not deployed, and generate a meta-information set; An interface deployment module 207, configured to update the interface deployment parameters of the data middle platform based on the meta-information set and the preset data interface configuration information corresponding to the governance target data; A field transmission module 208, configured to send the data identification field information associated with the target data set to the data terminal; An interface execution engine 209, integrated in the data middle platform, configured to convert the assigned preset governance field configuration into interface call parameters and execute the data interface call operation in response to detecting the assignment operation of the preset governance field configuration of the preset data interface configuration information by the data terminal, so as to return the governance target data to the data terminal.

[0064] It can be understood that the modules described in the machine learning-based campus data governance management system correspond to the steps in the machine learning-based campus data governance management method described in the reference Figure 1 description. Therefore, the operations, features, and beneficial effects described above for the machine learning-based campus data governance management method also apply to the machine learning-based campus data governance management system and the modules included therein, and will not be elaborated here.

[0065] Refer to the following Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present invention. The electronic device in some embodiments of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The terminal device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0066] As shown in Figure 3 , the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0067] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 shows the electronic device 300 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 3 Each block shown in

[0068] Further, the storage medium of the embodiments of the present application stores program instructions capable of implementing all the above methods. Among them, the program instructions can be stored in the above storage medium in the form of a software product, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0069] The above description is only some preferred embodiments of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.

Claims

1. A campus data governance management method based on machine learning, comprising: In response to detecting a data acquisition request message of a target data set by a data terminal, sending the acquisition request message of the target data set to a data middle platform; In response to receiving the data identification field information corresponding to the target data set sent by the data middle platform, triggering a data processing process of the data middle platform according to the data identification field information and the target interface configuration information of the data middle platform; Structurally processing the original campus data corresponding to the data identification field information through the data middle platform to obtain governance target data; Generating data governance strategy information according to the governance target data; Performing model parameter update processing on each data element information corresponding to the target data set including the representation data identification according to the data governance strategy information; In response to determining that the interface configuration information corresponding to the data terminal is not deployed, associating each updated data element information with a preset governance field configuration of the data middle platform to generate a meta-information set; Updating the interface deployment parameters of the data middle platform based on the meta-information set and the preset data interface configuration information corresponding to the governance target data; Sending the data identification field information associated with the target data set to the data terminal; In response to detecting an assignment operation of the preset governance field configuration to the preset data interface configuration information by the data terminal, converting the assigned preset governance field configuration into an interface call parameter through the data middle platform, and performing a data interface call operation to return the governance target data to the data terminal.

2. The method according to claim 1, wherein Triggering the data processing process of the data middle platform includes: Extracting a target field mapping rule corresponding to the target interface configuration information from an interface configuration library of the data middle platform; Injecting the data identification field information into a parameter container corresponding to the target field mapping rule; Performing a parameterized interface call through an interface call engine of the data middle platform to obtain the original campus data associated with the data identification field information.

3. The method according to claim 1, wherein The structural processing includes: Performing data quality verification processing on the original campus data to generate basic data that passes the verification; Performing feature space transformation processing on the basic data through a machine learning algorithm integrated in the data middle platform to generate governance target data with standardized data structure features.

4. The method according to claim 1, wherein The model parameter update processing includes: Generating a data governance feature vector through the data governance strategy information; Inputting the data governance feature vector into a meta-information update module of the data middle platform to calculate weight adjustment parameters for each data element information; Updating the metadata storage structure of each data element information according to the weight adjustment parameters.

5. The method according to claim 4, wherein The association processing includes: Aligning the feature dimensions of the updated data element information with the preset governance field configuration; Calculating semantic similarity parameters between the data element information and the preset governance field configuration through an association relationship learning module of the data middle platform; Constructing a meta-information set with a hierarchical structure based on the semantic similarity parameters.

6. The method according to claim 1, wherein The interface deployment parameter update includes: Converting the meta-information set into an interface description file of the data middle platform; Compiling the interface description file into an executable data interface configuration through a model-driven interface generator; Deploying the data interface configuration to an interface service gateway of the data middle platform.

7. The method according to claim 3, wherein The data quality verification processing includes: Obtain the field integrity parameters of the original campus data through the data verification engine of the data middle platform; When the field integrity parameters reach the preset verification threshold, activate the data completion module of the data middle platform; Perform interpolation calculation processing on the missing fields through the data completion module to generate the basic data that passes the verification.

8. The method according to claim 3, wherein The feature space transformation processing includes: Call the dimensionality reduction algorithm library of the data middle platform to perform principal component analysis processing on the basic data; Through the standardization module of the data middle platform, perform Z-score standardization calculation on the data after principal component analysis processing; Write the standardization calculation result into the feature storage library of the data middle platform to generate the governance target data with the characteristics of the standardized data structure.

9. The method according to claim 5, wherein The construction of semantic similarity parameters includes: Through the word vector generator of the data middle platform, convert the data element information and the preset governance field configuration into a set of word vectors; Based on the attention mechanism module of the data middle platform, calculate the cosine similarity matrix of the set of word vectors; Through the graph neural network of the data middle platform, generate a hierarchical meta-information set according to the cosine similarity matrix.

10. A campus data governance management system based on machine learning, including: A detection module, configured to send the acquisition request information of the target data set to the data middle platform in response to detecting the data acquisition request information of the target data set by the data terminal; An interface trigger module, configured to trigger the data processing process of the data middle platform according to the data identification field information and the target interface configuration information of the data middle platform in response to receiving the data identification field information of the corresponding target data set sent by the data middle platform; A data processing engine, integrated in the data middle platform, configured to perform structured processing on the original campus data corresponding to the data identification field information, where the structured processing includes data feature extraction processing based on machine learning algorithms to obtain the governance target data; A policy generation module, configured to generate data governance policy information corresponding to the governance target data according to the governance target data; A parameter update module, configured to perform model parameter update processing on each data element information corresponding to the target data set according to the data governance policy information, where each data element information includes the meta-information representing the data identification; An association processing module, configured to perform association processing on each updated data element information and the preset governance field configuration of the data middle platform in response to determining that the interface configuration information of the corresponding data terminal is not deployed, to generate a set of meta-information; An interface deployment module, configured to update the interface deployment parameters of the data middle platform based on the set of meta-information and the preset data interface configuration information corresponding to the governance target data; A field transmission module, configured to send the data identification field information associated with the target data set to the data terminal; An interface execution engine, integrated in the data middle platform, configured to convert the assigned preset governance field configuration into interface call parameters in response to detecting the assignment operation of the preset governance field configuration of the preset data interface configuration by the data terminal, and perform the data interface call operation to return the governance target data to the data terminal.

Citation Information

Patent Citations

  • Data processing method and device and storage medium

    CN112416991A

  • Stream data acquisition system and method based on edge node community network

    CN114528890A

  • Systems, methods, and apparatus for managing vehicle data collection

    CN115443637A

  • Self-learning analytical attribute and clustering segmentation system

    US20200394564A1