Personalized matching method and system based on multi-source heterogeneous data fusion

By receiving student data matching instructions, acquiring multi-source heterogeneous datasets, and using machine learning classifiers for fusion importance analysis, the problem of difficulty in adaptively adjusting fusion strategies in traditional methods is solved, thereby improving the accuracy of personalized matching and the efficiency of resource utilization.

CN121935629APending Publication Date: 2026-04-28BEIJING XINXINFU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610191099.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-10
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional methods struggle to adaptively adjust integration strategies in personalized services for smart campuses, failing to accurately capture students' dynamically changing behavioral patterns and real needs, resulting in storage redundancy and poor resource recommendation performance.

Method used

By receiving student data matching instructions, multi-source heterogeneous datasets are obtained. A machine learning classifier is used to perform fusion importance analysis, generate a fusion index, dynamically implement differentiated fusion strategies, and optimize data management.

Benefits of technology

It improves the accuracy of personalized matching and the efficiency of school database resource utilization, reduces data redundancy, and enhances the real-time performance and adaptability of matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935629A_ABST
    Figure CN121935629A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of personalized data services, and discloses a personalized matching method and system based on multi-source heterogeneous data fusion, and the method comprises the steps: obtaining a plurality of to-be-matched multi-source data sets of a target student based on a data matching period and a matched data source set, obtaining a comparison multi-source data set based on the target multi-source data set and the multi-student heterogeneous database, carrying out fusion importance analysis on the target multi-source data set by using the comparison multi-source data set to obtain fusible indexes, summarizing the fusible indexes corresponding to the target multi-source data set to obtain a plurality of fusible indexes, and carrying out fusion analysis on the fusible indexes; and fusing the target multi-source data set and the multi-student heterogeneous database according to the plurality of fusible indexes to obtain a fused heterogeneous database. According to the invention, the accuracy of personalized matching of students and the utilization efficiency of school database resources can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of personalized data service technology, and in particular to a personalized matching method and system based on the fusion of multi-source heterogeneous data. Background Technology

[0002] In the field of personalized services in smart campuses, effectively integrating heterogeneous student data from multiple sources such as academic affairs, consumption, library borrowing, and extracurricular activities is the core of achieving precise academic guidance and growth support. These data differ in format, time sequence, and semantics, and their complexity places high demands on systematic collection and integration.

[0003] Traditional methods often rely on static rules or fixed-weight aggregation strategies to build student profiles and match resources. They lack the ability to quantitatively assess the uniqueness of data and are difficult to adaptively adjust the fusion strategy. This limitation makes it impossible for the system to accurately capture students' dynamic behavioral patterns and real needs. This not only easily leads to storage redundancy but also restricts the actual effectiveness of applications such as personalized tutoring and resource recommendation. Summary of the Invention

[0004] This invention provides a personalized matching method and system based on multi-source heterogeneous data fusion, the main purpose of which is to improve the accuracy of personalized matching for students and the efficiency of school database resource utilization.

[0005] To achieve the above objectives, this invention provides a personalized matching method based on multi-source heterogeneous data fusion, comprising: Receive student data matching instructions, and identify the target student and matching data source set based on the student data matching instructions. The matching data source set includes multiple matching data sources, and the matching data sources are library data sources, academic affairs data sources, consumption data sources and activity data sources. Based on a preset data matching period and a set of matching data sources, multiple multi-source datasets to be matched for the target student are obtained. Each multi-source dataset to be matched corresponds one-to-one with a matching data source, and each multi-source dataset to be matched includes multiple multi-source datasets to be matched. A multi-student heterogeneous database was identified, which includes multiple student multi-source databases, and each student multi-source database corresponds to a student ID. Extract the multi-source datasets to be matched sequentially from multiple multi-source datasets to be matched, and denote the extracted multi-source datasets to be matched as the target multi-source dataset. Obtain the comparison multi-source dataset based on the target multi-source dataset and the heterogeneous database of multiple students. By comparing the target multi-source datasets, we can perform a fusion importance analysis on the target multi-source datasets to obtain a fusion index. By summing the fusion indices corresponding to the target multi-source datasets, we can obtain multiple fusion indices. The target multi-source dataset and multi-student heterogeneous database are fused based on multiple fusion indices to obtain a fused heterogeneous database. Student resources are then pushed based on the fused heterogeneous database to obtain the pushed resources.

[0006] Optionally, the step of obtaining the comparative multi-source dataset based on the target multi-source dataset and the multi-student heterogeneous database includes: Random student sampling is performed on a multi-student heterogeneous database based on the target multi-source dataset to obtain a comparison student group, which includes multiple comparison students; In the comparison student group, the comparison students are extracted sequentially, and the comparison multi-source database of the extracted comparison students is obtained in the multi-student heterogeneous database; Identifying and comparing multi-source data groups in a multi-source database based on data matching cycles; Merge the comparative multi-source datasets corresponding to each student to obtain the comparative multi-source dataset.

[0007] Optionally, the step of performing a fusion importance analysis on the target multi-source dataset using comparative multi-source datasets to obtain a fusion index includes: The target multi-source dataset is divided into training multi-source datasets and validation multi-source datasets. A training classification dataset is constructed based on the comparison of multi-source datasets and the training multi-source dataset; The target classifier is obtained by training a pre-built classifier using a training classification dataset. For each validation multi-source data in the validation multi-source dataset, features are constructed to obtain multiple validation data feature sets. The validation data feature sets in the multiple validation data feature sets correspond one-to-one with the validation multi-source data in the validation multi-source dataset. Multiple validation data feature sets are input into the target classifier to obtain multiple classification probabilities. The average classification probability is calculated by averaging the multiple classification probabilities and is denoted as the fusion index.

[0008] Optionally, the step of constructing a training classification dataset based on the comparison of multi-source datasets and the training multi-source dataset includes: Perform the following operation on each training multi-source data in the training multi-source dataset: Features are constructed from training multi-source data to obtain a training data feature set, which includes multiple training data features. The training data feature set is labeled based on the preset positive sample labels to obtain the labeled training feature set; In the comparison multi-source dataset, comparison multi-source data are extracted sequentially, and a comparison data feature set is constructed based on the extracted comparison multi-source data. The comparison data feature set is labeled according to the preset negative sample labels to obtain the labeled comparison feature set. Merge the labeled training feature set and the labeled contrast feature set to obtain training classification data; The training classification data corresponding to each comparison multi-source data is summarized to obtain multiple training classification data. The multiple training classification data corresponding to each training multi-source data are merged to obtain the training classification dataset.

[0009] Optionally, the step of fusing the target multi-source dataset and the multi-student heterogeneous database according to multiple fusion indices to obtain a fused heterogeneous database includes: Construct a target data storage structure, which includes multiple data source storage addresses, and each data source storage address corresponds one-to-one with the target multi-source data in the target multi-source dataset; Obtain the classification fitting threshold set, which includes: high fitting threshold and low fitting threshold; Set multiple data source weights based on multiple matching data sources; By using the weights of multiple data sources to calculate the weighted average of multiple fusionable indices, a comprehensive fit value is obtained; If the overall fit value is greater than the high fit threshold in the classification fit threshold group, then the preset empty set is recorded as the target student database. If the overall fit value is less than the low fit threshold in the classification fit threshold group, then data fusion is performed using the multi-student heterogeneous database and the target data storage structure to obtain the target student database. If the overall fit value is not greater than the high fit threshold and not less than the low fit threshold, then the target student database is generated based on the target multi-source dataset and the target data storage structure. Add the target student database to the multi-student heterogeneous database to obtain a merged heterogeneous database.

[0010] Optionally, obtaining the classification fitting threshold set includes: Based on the data matching period query, the historical fitting threshold group is queried, where the historical fitting threshold group includes: historical high fitting threshold and historical low fitting threshold. Perform data statistics on a heterogeneous database with multiple students to obtain the current data volume and the total number of students; Receive student matching weight reorganization, wherein student matching weight reorganization includes: stable matching weight and generalized matching weight; The historical fitting threshold group is adjusted based on the current data volume, total number of students, and student matching weights to obtain the classification fitting threshold group.

[0011] Optionally, the step of fusing data using a multi-student heterogeneous database and a target data storage structure to obtain a target student database includes: Extract target multi-source data sequentially from the target multi-source dataset, identify the target data source corresponding to the extracted target multi-source data, and identify the target storage address corresponding to the target data source in the target data storage structure; Based on the target data source and heterogeneous databases of multiple students, query similar data from the same source. Obtain the data storage address of similar data from the same source; The source data storage address is pointed to the target storage address to obtain the fused storage address; The merged storage addresses corresponding to each target data source are aggregated to obtain multiple merged storage addresses. The target data storage structure is then updated using these multiple merged storage addresses to obtain the target student database.

[0012] Optionally, the step of querying similar data from the same source based on the target data source and the heterogeneous database of multiple students includes: Based on the target data source, the homogeneous database of multiple students is filtered to obtain the homogeneous student dataset, which includes multiple homogeneous student datasets. Perform the following operation on each student's data in the student common-origin dataset: Construct a feature vector of the same source data based on the same source data of students, and construct a feature vector of the target data based on the extracted target multi-source data; Calculate data similarity based on the feature vectors of the source data and the feature vectors of the target data; Summarize the data similarity corresponding to each student's source data to obtain a data similarity set. Determine the maximum similarity in the data similarity set and record the source data of the student corresponding to the maximum similarity as similar source data.

[0013] Optionally, generating the target student database based on the target multi-source dataset and the target data storage structure includes: Perform the following operation on each of the multiple data source storage addresses in the target data storage structure: Create a database partition for the data source storage address to obtain a storable address, which includes metadata indexes and update logs; Based on storable addresses, homogeneous stored data were identified in the target multi-source dataset. Write the same source storage data to a storable address to obtain a valid storage address; By summarizing the valid storage addresses, multiple valid storage addresses are obtained. The target data storage structure is then updated using these multiple valid storage addresses to obtain the target student database.

[0014] To achieve the above objectives, the present invention also provides a personalized matching system based on multi-source heterogeneous data fusion, comprising: The matching instruction receiving module is used to receive student data matching instructions, and to identify the target student and the matching data source set based on the student data matching instructions. The matching data source set includes multiple matching data sources, and the matching data sources are library data sources, academic affairs data sources, consumption data sources and activity data sources. The multi-source data acquisition module is used to acquire multiple multi-source datasets to be matched for the target student based on a preset data matching period and a matching data source set. The multi-source datasets to be matched correspond one-to-one with the matching data sources, and each multi-source dataset to be matched includes multiple multi-source datasets to be matched. The heterogeneous data acquisition module is used to identify a multi-student heterogeneous database, which includes multiple student multi-source databases, and each student multi-source database corresponds to a student ID. The module sequentially extracts the multi-source datasets to be matched from multiple datasets to be matched, and records the extracted multi-source datasets to be matched as the target multi-source dataset. The module then obtains a comparison multi-source dataset based on the target multi-source dataset and the multi-student heterogeneous database. The multi-source data fusion module is used to perform fusion importance analysis on the target multi-source dataset by comparing the multi-source datasets, obtain the fusion index, summarize the fusion indices corresponding to the target multi-source datasets, obtain multiple fusion indices, fuse the target multi-source datasets and multi-student heterogeneous databases according to the multiple fusion indices, obtain the fused heterogeneous database, and push student resources based on the fused heterogeneous database to obtain the pushed resources.

[0015] To address the above problems, the present invention also provides an electronic device, the electronic device comprising: Memory, storing at least one instruction; The processor executes the instructions stored in the memory to implement the personalized matching method based on multi-source heterogeneous data fusion described above.

[0016] To address the aforementioned problems, the present invention also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor in an electronic device to implement the personalized matching method based on multi-source heterogeneous data fusion described above.

[0017] To address the problems described in the background art, this invention first receives a student data matching instruction. Based on this instruction, the target student and the matching data source set are identified. This step, by receiving the student data matching instruction and specifying the matching data source set, ensures the diversity and relevance of data sources. Compared to existing technologies that often rely on a single data source, this provides a comprehensive and accurate data foundation for subsequent multi-source heterogeneous data fusion, thereby improving the coverage and accuracy of matching. Then, based on the data matching cycle and the matching data source set, multiple multi-source datasets of the target student to be matched are obtained. This step utilizes the data matching cycle to systematically collect student data, achieving dynamic monitoring of student behavior and acquisition of time-series data. Compared to existing static or non-periodic data collection methods, this method can more effectively capture changes in student preferences, enhancing the real-time nature and adaptability of personalized matching. Next, based on the target multi-source dataset and the multi-student heterogeneous database, a comparison multi-source dataset is obtained. This step uses random student sampling and time alignment mechanisms to obtain the comparison multi-source dataset, establishing an objective comparison benchmark. Compared to subjective or direct comparison methods in existing technologies, this approach reduces human bias and provides reliable data support for subsequent fusion importance analysis. Furthermore, this solution utilizes comparative multi-source datasets to perform fusion importance analysis on the target multi-source dataset, obtaining a fusion index. By summarizing the fusion indices corresponding to the target multi-source dataset, multiple fusion indices are obtained. This step employs a machine learning classifier for fusion importance analysis, generating fusion indices and achieving a quantitative assessment of the uniqueness of student data. Compared to rule-based or experience-based methods in existing technologies, this improves the scientific rigor and automation of fusion decisions, making the matching process more accurate and efficient. Finally, based on multiple fusion indices, the target multi-source dataset and the heterogeneous database of multiple students are fused to obtain a fused heterogeneous database. This step dynamically implements differentiated fusion strategies based on the fusion indices, such as data referencing or independent storage, achieving adaptive optimization of data management. Compared to the one-size-fits-all storage method in existing technologies, this avoids data redundancy while ensuring the integrity of unique data, improving system resource utilization efficiency and matching quality. Therefore, this invention can improve the accuracy of personalized student matching and the efficiency of school database resource utilization. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a personalized matching method based on multi-source heterogeneous data fusion provided in an embodiment of the present invention. Figure 2 A functional block diagram of a personalized matching system based on multi-source heterogeneous data fusion provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device that implements the personalized matching method based on multi-source heterogeneous data fusion, according to an embodiment of the present invention.

[0019] Explanation of reference numerals in the attached figures: 10. Electronic device; 11. Processor; 12. Memory; 13. Bus.

[0020] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0022] This application provides a personalized matching method based on multi-source heterogeneous data fusion. The executing entity of the personalized matching method based on multi-source heterogeneous data fusion includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application embodiment: a server, a terminal, etc. In other words, the personalized matching method based on multi-source heterogeneous data fusion can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.

[0023] Reference Figure 1 The diagram shown is a flowchart illustrating a personalized matching method based on multi-source heterogeneous data fusion according to an embodiment of the present invention. In this embodiment, the personalized matching method based on multi-source heterogeneous data fusion includes: S1. Receive student data matching instructions, and identify the target student and matching data source set based on the student data matching instructions. The matching data source set includes multiple matching data sources, and the matching data sources are library data sources, academic affairs data sources, consumption data sources and activity data sources.

[0024] It is clear that the student data matching instruction refers to a human-initiated instruction to match data for a specific student, and the target student refers to the specific student included in the student data matching instruction. The matching data source set refers to a collection of multiple matching data sources, where a matching data source refers to a specific type of data source selected for personalized matching. Among these, the library data source refers to the data source of the school library, the academic affairs data source refers to the data source of the school's academic affairs office, the consumption data source refers to the data source of the school's consumption end (such as the canteen, convenience store, etc.), and the activity data source refers to the data source of various school activities (such as various school clubs).

[0025] S2. Based on the preset data matching period and matching data source set, obtain multiple multi-source datasets to be matched for the target student, wherein each multi-source dataset to be matched corresponds one-to-one with the matching data source, and each multi-source dataset to be matched includes multiple multi-source data to be matched.

[0026] Understandably, the data matching period refers to a fixed time interval set by the user for systematically collecting and processing student data, such as a cycle based on days, weeks, or months. The multi-source dataset to be matched refers to a collection of multiple multi-source datasets to be matched. Specifically, the multi-source dataset to be matched refers to the data set of the target student related to the matching data source within the data matching period. This multi-source dataset includes multiple multi-source datasets to be matched, and each multi-source dataset is data collected at different times related to the matching data source. For example, if the matching data source is a library data source, then the multi-source dataset to be matched consists of the student's behavioral data generated within the library's physical and digital systems using their campus identity during the data matching period, including but not limited to book borrowing history, electronic document search keywords, database access duration and frequency, and download records of subject-specific materials. If the matching data source is a consumption data source, then the multi-source dataset to be matched consists of the student's consumption behavior data generated within the data matching period through payment channels such as campus cards, including but not limited to consumption time, amount, frequency, and categories of goods consumed at different locations such as the cafeteria, supermarket, and coffee shop.

[0027] S3. Identify a multi-student heterogeneous database, which includes multiple student multi-source databases, and each student multi-source database corresponds to a student ID.

[0028] It should be explained that the aforementioned multi-student heterogeneous database is a system-level database that centrally stores and manages massive amounts of student data. Its core component is the student multi-source database, where each database is bound to a specific student by a unique student ID, used to archive historical data generated by that student across all preset data sources (such as library data sources, consumption data sources, etc.). The student multi-source database maintains the same dimensions in its data structure as the matching data source set, but its data time span is much longer than a single data matching period, containing long-term, continuously accumulated behavioral records of students. This database is a dynamically growing entity, constantly updated as new data is generated, thus providing comprehensive historical data support for student behavior analysis.

[0029] S4. Extract the multi-source datasets to be matched sequentially from multiple multi-source datasets to be matched, and denote the extracted multi-source datasets to be matched as the target multi-source dataset. Obtain the comparison multi-source dataset based on the target multi-source dataset and the heterogeneous database of multiple students.

[0030] It is clear that the aforementioned comparative multi-source dataset refers to the data set that will be used for subsequent fusion importance analysis with the target multi-source dataset. The specific method for obtaining this comparative multi-source dataset will be given later.

[0031] Specifically, the acquisition of comparative multi-source datasets based on the target multi-source dataset and multi-student heterogeneous databases includes: Random student sampling is performed on a multi-student heterogeneous database based on the target multi-source dataset to obtain a comparison student group, which includes multiple comparison students; In the comparison student group, the comparison students are extracted sequentially, and the comparison multi-source database of the extracted comparison students is obtained in the multi-student heterogeneous database; Identifying and comparing multi-source data groups in a multi-source database based on data matching cycles; Merge the comparative multi-source datasets corresponding to each student to obtain the comparative multi-source dataset.

[0032] Understandably, the "comparison student group" refers to multiple student IDs in a heterogeneous multi-student database obtained through random student sampling. Random student sampling in the heterogeneous multi-student database based on the target multi-source dataset means: after excluding the target student's own IDs from all student IDs in the heterogeneous multi-student database, a group of student IDs is randomly selected according to a preset sampling ratio or a fixed number to construct a reference group for data comparison and analysis. The "comparison multi-source database" refers to the student multi-source database corresponding to the comparison students. The "comparison multi-source data group" refers to the time-aligned and structurally consistent set of comparison data extracted from the comparison multi-source database based on the matching data source set corresponding to the target multi-source dataset and the same data matching period.

[0033] S5. Analyze the importance of fusion of the target multi-source dataset by comparing the multi-source datasets, obtain the fusion index, summarize the fusion indices corresponding to the target multi-source datasets, and obtain multiple fusion indices.

[0034] Understandably, the fusion index refers to a numerical value that quantifies the distributional differences between the target multi-source dataset and the comparative multi-source dataset from a heterogeneous multi-student database. This fusion index is the average classification probability output by the classifier on the subsequent validation multi-source dataset. The higher the fusion index, the lower the distributional consistency between the target multi-source dataset and the comparative multi-source dataset, meaning that the data patterns of the target students are more unique and more difficult to be represented by the existing student data in the heterogeneous multi-student data. Therefore, it needs to be treated differently in subsequent fusion decisions, such as creating an independent database partition for it. Conversely, a lower fusion index means high similarity, and lightweight fusion strategies such as data referencing can be adopted. This fusion index provides an objective and quantitative decision-making basis for the selection of data fusion strategies.

[0035] In detail, the method of performing fusion importance analysis on the target multi-source dataset by comparing multi-source datasets to obtain a fusion index includes: The target multi-source dataset is divided into training multi-source datasets and validation multi-source datasets. A training classification dataset is constructed based on the comparison of multi-source datasets and the training multi-source dataset; The target classifier is obtained by training a pre-built classifier using a training classification dataset. For each validation multi-source data in the validation multi-source dataset, features are constructed to obtain multiple validation data feature sets. The validation data feature sets in the multiple validation data feature sets correspond one-to-one with the validation multi-source data in the validation multi-source dataset. Multiple validation data feature sets are input into the target classifier to obtain multiple classification probabilities. The average classification probability is calculated by averaging the multiple classification probabilities and is denoted as the fusion index.

[0036] It is clear that the training multi-source dataset refers to the collection of multiple target multi-source data used to train the classifier. The validation multi-source dataset refers to the collection of multiple target multi-source data used to validate the target classifier. Dividing the target multi-source dataset means randomly or sequentially splitting the data according to a preset ratio (e.g., 7:3 or 8:2) to ensure the effectiveness and reliability of the training and validation processes. The training classification dataset refers to the collection of multiple training classification data, which includes a labeled training feature set and a labeled contrast feature set. The detailed construction methods of the labeled training feature set and the labeled contrast feature set will be given later. The classifier refers to a machine learning model capable of automatically classifying data based on input features; optionally, logistic regression, support vector machines, or lightweight neural networks may be used as the classifier. The target classifier refers to the trained classifier.

[0037] It needs to be explained that training the pre-built classifier using the training classification dataset means: inputting the training classification dataset into the initialized classifier model, and iteratively adjusting the model's internal parameters through a specific optimization algorithm (such as gradient descent) so that it can accurately distinguish between the labeled training feature set (positive sample labels) and the labeled contrastive feature set (negative sample labels). The validation data feature set refers to a collection of multiple data features describing the validation multi-source data. Feature construction of the validation multi-source data means: extracting features from the validation multi-source data through mathematical statistics, image processing, text recognition, etc., to obtain the validation data feature set. For example, if the validation multi-source data is an image of a product purchased by a student at a certain time, the feature construction method is: using a pre-trained convolutional neural network (such as ResNet) to extract the deep feature vector of the product image. If the validation multi-source data is a journal downloaded by a student from a library data source at a certain time, the feature construction method is: using a bag-of-words model or text embedding techniques (such as Word2Vec, BERT) to convert the journal's title or abstract into a numerical vector. If the verification of multi-source data is a sequence of cafeteria consumption times generated by a student in the past month, the feature is constructed by calculating the statistical features of the cafeteria consumption time sequence, such as mean, variance, maximum value, minimum value, and time series features (such as frequency domain features obtained through Fourier transform).

[0038] Furthermore, the classification probability refers to the predicted probability value that the target classifier outputs after receiving a set of validation data features, indicating that the feature set belongs to the "positive sample label" (i.e., from the target student). The higher the classification probability, the easier it is for the classifier to identify the validation data feature set as originating from the target student. In other words, the more unique the target student's features are on this data, and the more obvious the difference between them and the comparison students. Each classification probability corresponds to one set of validation data features. The average classification probability refers to the numerical average of multiple classification probabilities.

[0039] Specifically, the construction of the training classification dataset based on the comparison of multi-source datasets and the training multi-source dataset includes: Perform the following operation on each training multi-source data in the training multi-source dataset: Features are constructed from training multi-source data to obtain a training data feature set, which includes multiple training data features. The training data feature set is labeled based on the preset positive sample labels to obtain the labeled training feature set; In the comparison multi-source dataset, comparison multi-source data are extracted sequentially, and a comparison data feature set is constructed based on the extracted comparison multi-source data. The comparison data feature set is labeled according to the preset negative sample labels to obtain the labeled comparison feature set. Merge the labeled training feature set and the labeled contrast feature set to obtain training classification data; The training classification data corresponding to each comparison multi-source data is summarized to obtain multiple training classification data. The multiple training classification data corresponding to each training multi-source data are merged to obtain the training classification dataset.

[0040] Understandably, the training data feature set refers to the set of values ​​representing the features of training multi-source data after feature construction. The construction method of this training data feature set is the same as that of the training classification dataset, and will not be repeated here. The positive sample label refers to a specific numerical value (e.g., the value 1) used to identify the target student data in supervised learning. The labeled training feature set refers to the labeled training data feature set. Labeling the training data feature set based on the preset positive sample label means adding a label field to the metadata of the training data feature set and setting its value to the positive sample label (e.g., 1) to explicitly indicate that the feature set comes from the target student. The contrast data feature set refers to the set of values ​​representing the features of contrasting multi-source data after feature construction. The construction method of this contrast data feature set is the same as that of the training classification dataset. The negative sample label refers to a specific numerical value (such as 0) used to identify the data of the comparison student (non-target student) in supervised learning. The labeled comparison feature set refers to the labeled set of comparison data features. The step of labeling the comparison data feature set according to the preset negative sample label is the same as the step of labeling the training data feature set based on the preset positive sample label, and will not be repeated here. The training classification data refers to the data set labeled with the training feature set and the labeled comparison feature set.

[0041] S6. Based on multiple fusion indices, the target multi-source dataset and multi-student heterogeneous database are fused to obtain a fused heterogeneous database. Student resources are pushed based on the fused heterogeneous database to obtain the pushed resources.

[0042] It is clear that the aforementioned fusion of heterogeneous databases refers to a multi-student heterogeneous database resulting from the fusion of target multi-source datasets. The detailed fusion process will be provided in subsequent steps. The aforementioned resource push refers to pushing personalized learning resources, club activity resources, etc., to the target student based on the activity preferences, reading habits, etc., reflected in the target student database. For example, after determining that student Xiao Wang's data pattern is unique through fusion importance analysis (such as borrowing psychology books, declining grades, leaving clubs, etc.), and generating an independent target student database for Xiao Wang, precise resource matching will be performed based on this target student database: pushing a link to an online lecture on "Mental Health and Stress Coping" to Xiao Wang's campus APP, prioritizing his appointment with an academic tutor for one-on-one tutoring, and recommending participation in light social activities such as book clubs.

[0043] In detail, the process of fusing the target multi-source dataset and the multi-student heterogeneous database based on multiple fusion indices to obtain a fused heterogeneous database includes: Construct a target data storage structure, which includes multiple data source storage addresses, and each data source storage address corresponds one-to-one with the target multi-source data in the target multi-source dataset; Obtain the classification fitting threshold set, which includes: high fitting threshold and low fitting threshold; Set multiple data source weights based on multiple matching data sources; By using the weights of multiple data sources to calculate the weighted average of multiple fusionable indices, a comprehensive fit value is obtained; If the overall fit value is greater than the high fit threshold in the classification fit threshold group, then the preset empty set is recorded as the target student database. If the overall fit value is less than the low fit threshold in the classification fit threshold group, then data fusion is performed using the multi-student heterogeneous database and the target data storage structure to obtain the target student database. If the overall fit value is not greater than the high fit threshold and not less than the low fit threshold, then the target student database is generated based on the target multi-source dataset and the target data storage structure. Add the target student database to the multi-student heterogeneous database to obtain a merged heterogeneous database.

[0044] Understandably, the target data storage structure refers to a pre-defined logical database framework for the target students. This target data storage structure includes multiple data source storage addresses, and each data source storage address represents a logical pointer to the actual storage location of a specific type of data (such as library data sources, consumer data sources, etc.). The classification fitting threshold group refers to a combination of high-fit thresholds and low-fit thresholds. The high-fit threshold is a critical value used to determine that the target student data (i.e., the target multi-source dataset) is too unique and may belong to anomalies. The low-fit threshold is a critical value used to determine that the target student data is highly redundant with the existing database and lacks uniqueness. The high-fit threshold and low-fit threshold will dynamically change with the update iteration of the multi-student heterogeneous database. The data source weight refers to a numerical weight set by the user to indicate the importance of the matching data source, and the sum of the weights of multiple data sources is 1. The comprehensive fitting value refers to the value obtained after weighted calculation. The weighted calculation of multiple data source weights on multiple fusion indices refers to multiplying the data source weights with the corresponding fusion indices and adding them together. The calculated value is the comprehensive fitting value.

[0045] It should be explained that if the overall fit value is greater than the high fit threshold, it indicates that the target student's data pattern is extremely unique and differs greatly from the existing students in the multi-student heterogeneous dataset. This may be due to data collection bias or sudden changes in student behavior. Therefore, this target multi-source dataset will not be stored or fused; instead, a pre-defined empty set will be designated as the target student database, and this target multi-source dataset will be uploaded to the human end. If the overall fit value is less than the low fit threshold, it indicates that the target student's data pattern is highly similar to the existing data (i.e., the multi-student heterogeneous database), and their individual data lacks uniqueness. In this case, data fusion can be performed using the multi-student heterogeneous database and the target data storage structure. The specific data fusion method will be given in subsequent steps. If the overall fit value is neither greater than the high fit threshold nor less than the low fit threshold, it indicates that the target student's data pattern has a certain degree of uniqueness but also has a reasonable correlation with the existing data. In this case, a database that can be independently read, written, and managed can be generated for the target multi-source dataset; that is, the step of generating the target student database based on the target multi-source dataset and the target data storage structure will be executed.

[0046] Specifically, obtaining the classification fitting threshold set includes: Based on the data matching period query, the historical fitting threshold group is queried, where the historical fitting threshold group includes: historical high fitting threshold and historical low fitting threshold. Perform data statistics on a heterogeneous database with multiple students to obtain the current data volume and the total number of students; Receive student matching weight reorganization, wherein student matching weight reorganization includes: stable matching weight and generalized matching weight; The historical fitting threshold group is adjusted based on the current data volume, total number of students, and student matching weights to obtain the classification fitting threshold group.

[0047] It is clear that the historical fitting threshold group refers to the classification fitting threshold group in the previous data matching period, where the historical high fitting threshold and historical low fitting threshold refer to the high fitting threshold and low fitting threshold in the historical fitting threshold group, respectively. The current data volume refers to the total amount of data currently stored in the multi-student heterogeneous database (measured in bytes or records). The total number of students refers to the number of students in the multi-source student database that has been established in the multi-student heterogeneous database. The stable matching weight is a parameter used to adjust the student differentiation accuracy; a larger value indicates a greater tendency to finely differentiate the data patterns of different students. The generalization matching weight is a parameter used to adjust the generalization degree of the student group; a larger value indicates a greater tendency to group students with similar data patterns into one category for processing.

[0048] It should be explained that the purpose of introducing the aforementioned stable matching weights and generalized matching weights is to dynamically guide the generation of classification fitting threshold sets by adjusting these two weight parameters, thereby balancing accuracy and efficiency in personalized matching strategies. When matching accuracy is emphasized (i.e., stable matching weights dominate), the generated classification fitting threshold sets have higher discriminative power, enabling more accurate identification of students with unique data patterns. When matching efficiency and generalization ability are emphasized (i.e., generalized matching weights dominate), the generated classification fitting threshold sets are more lenient, promoting data fusion and optimizing storage and computing resources.

[0049] Importantly, the formula for calculating the high-fit threshold mentioned above is:

[0050] in, Indicates a high fit threshold. This indicates the historical high-fit threshold. Indicates stable matching weights. Indicates the generalization matching weight. This represents a logarithmic function with the natural constant as its base. Indicates the current data volume. The total number of students is represented by the formula for calculating the low-fit threshold:

[0051] in, Indicates the low-fit threshold. This represents the historical low-fit threshold.

[0052] Furthermore, the core mechanism of the high-fit threshold formula is to utilize the difference between the stable matching weight and the generalized matching weight. ) for historical high-fit threshold ( ) to perform directional adjustments when higher matching accuracy is required (i.e. When the formula is applied, the result will result in a high-fit threshold ( The increase in the logarithm in the denominator of the formula means that the target student's data needs to have a higher degree of differentiation from the comparison student's data to be considered unique, thus ensuring that the fusion decision is more cautious and accurate. ) plays a dynamic damping role: when the database size ( and As the database expands, the damping effect intensifies, and the threshold adjustment range narrows, ensuring the stability of the system on a massive data base. Conversely, in the early stages of database construction, relatively large adjustments to the threshold are allowed to adapt to changes in data distribution.

[0053] It needs to be explained that the adjustment logic of the low-fit threshold formula is the opposite of that of the high-fit threshold formula, when favoring fine-tuning ( When the fitting threshold is low ( The threshold for high-fitting data will decrease, thus widening the effective discrimination interval together with the increased high-fitting threshold, allowing for a more detailed classification of the differences in student data, when the generalization matching is more favorable. When the data volume ( ), the underfit threshold increases, narrowing the discrimination interval and promoting data merging and referencing. Similarly, the data volume ( ) ) and total number of students ( As a damping factor, this ensures that threshold changes match the overall size and maturity of the database, avoiding threshold oscillations in stable, large databases and maintaining the long-term robustness of the system. This design allows threshold settings to adaptively respond to different matching strategy preferences and the development status of the database.

[0054] In detail, the process of fusing data using a multi-student heterogeneous database and a target data storage structure to obtain a target student database includes: Extract target multi-source data sequentially from the target multi-source dataset, identify the target data source corresponding to the extracted target multi-source data, and identify the target storage address corresponding to the target data source in the target data storage structure; Based on the target data source and heterogeneous databases of multiple students, query similar data from the same source. Obtain the data storage address of similar data from the same source; The source data storage address is pointed to the target storage address to obtain the fused storage address; The merged storage addresses corresponding to each target data source are aggregated to obtain multiple merged storage addresses. The target data storage structure is then updated using these multiple merged storage addresses to obtain the target student database.

[0055] Understandably, the target storage address refers to the storage address of the data source corresponding to the target data source. The similar source data refers to the data of a student in a multi-student heterogeneous database that is highly similar to the target multi-source data. The source data storage address refers to the storage address of the similar source data. The merged storage address refers to the final data location (i.e., the source data storage address) actually associated with the target storage address after the address pointing operation is completed. Pointing the source data storage address to the target storage address means: at the database system level, establishing a pointer link from the logical target storage address to the physical source data storage address, so that the data content in the source data storage address can be actually read by accessing the target storage address without copying the data entity. Updating the target data storage structure using multiple merged storage addresses means: replacing the original target storage addresses pointing to empty storage locations in the corresponding data source of the target data storage structure one by one with newly generated merged storage addresses, thereby completing the reconstruction of the entire logical structure.

[0056] It should be explained that when the overall fit value is lower than the low fit threshold set in the classification fit threshold group, it indicates that the data pattern of the target student is highly similar to the data pattern of existing students in the multi-student heterogeneous database. This means that it is not necessary to create and maintain a separate data entity database for the target student; instead, existing data can be reused through data referencing. Therefore, this solution implements the following optimization strategy: For each data source in the target multi-source dataset, the set of existing student data most similar to it (i.e., similar datasets) is selected from the multi-student heterogeneous database. Subsequently, an address mapping relationship is established at the database logical layer, that is, the logical storage address pre-allocated to the target student database directly points to the physical storage location of the selected similar datasets. Through this address pointing mechanism, this solution constructs a virtual target student database, which is essentially a logical view composed of multiple external data references, rather than a physical collection storing entity data. This approach avoids redundant data storage and repeated processing, thereby reducing the server's storage and computational load and optimizing the overall system resource utilization efficiency.

[0057] In detail, the step of querying similar data from the same source based on the target data source and a multi-student heterogeneous database includes: Based on the target data source, the homogeneous database of multiple students is filtered to obtain the homogeneous student dataset, which includes multiple homogeneous student datasets. Perform the following operation on each student's data in the student common-origin dataset: Construct a feature vector for the same source data based on the same source data of students, and construct a feature vector for the target data based on the extracted target multi-source data; Calculate data similarity based on the feature vectors of the source data and the feature vectors of the target data; Summarize the data similarity corresponding to each student's source data to obtain a data similarity set. Determine the maximum similarity in the data similarity set and record the source data of the student corresponding to the maximum similarity as similar source data.

[0058] It is clear that the student homogeneous dataset refers to a data set in a multi-student heterogeneous database that shares the same data source as the target data source. The homogeneous data feature vector refers to a vector representing the features of the student homogeneous data. This homogeneous data feature vector is constructed as follows: first, a homogeneous data feature set is built, the construction method of which is the same as that of the training classification dataset, and will not be elaborated here. Then, the homogeneous data feature set is transformed into a vector with the same numerical value; this vector is the homogeneous data feature vector. The target data feature vector refers to a vector representing the features of the target multi-source data. The data similarity refers to the cosine value between the homogeneous data feature vector and the target data feature vector, used to represent the similarity between the student homogeneous data and the extracted target multi-source data. The maximum similarity refers to the data similarity with the highest numerical value in the data similarity set.

[0059] Specifically, the step of generating the target student database based on the target multi-source dataset and the target data storage structure includes: Perform the following operation on each of the multiple data source storage addresses in the target data storage structure: Create a database partition for the data source storage address to obtain a storable address, which includes metadata indexes and update logs; Based on storable addresses, homogeneous stored data were identified in the target multi-source dataset. Write the same source storage data to a storable address to obtain a valid storage address; By summarizing the valid storage addresses, multiple valid storage addresses are obtained. The target data storage structure is then updated using these multiple valid storage addresses to obtain the target student database.

[0060] Understandably, the storable address refers to the physical storage space independently allocated to the target student after database partitioning. Database partitioning refers to: within the database management system, allocating a dedicated, logically isolated storage area for the target student and configuring an independent metadata index and update log mechanism for this area. The homogeneous storage data refers to target multi-source data with the same data source as the storable address. The valid storage address refers to the storable address after homogeneous storage data has been written. Updating the target data storage structure using multiple valid storage addresses means: replacing the previously reserved logical addresses in the target data storage structure that are not associated with actual data with these valid storage addresses that already store actual data, thereby forming a target student database containing real data entities and capable of independent operation.

[0061] It should be explained that when the overall fit value is between the high fit threshold and the low fit threshold, it indicates that the target student's data pattern has both a certain degree of uniqueness and a reasonable correlation with the general patterns in the database. This means that the target student's data is not suitable to be simply replaced by existing data. Therefore, it is necessary to establish an independent target student database that physically stores its exclusive data. This database will not only completely store the current target multi-source dataset, but its built-in update log and other mechanisms will also ensure that new data generated by the student can be continuously collected and managed, thereby achieving long-term, independent and complete maintenance of the target student's data.

[0062] To address the problems described in the background art, this invention first receives a student data matching instruction. Based on this instruction, the target student and the matching data source set are identified. This step, by receiving the student data matching instruction and specifying the matching data source set, ensures the diversity and relevance of data sources. Compared to existing technologies that often rely on a single data source, this provides a comprehensive and accurate data foundation for subsequent multi-source heterogeneous data fusion, thereby improving the coverage and accuracy of matching. Then, based on the data matching cycle and the matching data source set, multiple multi-source datasets of the target student to be matched are obtained. This step utilizes the data matching cycle to systematically collect student data, achieving dynamic monitoring of student behavior and acquisition of time-series data. Compared to existing static or non-periodic data collection methods, this method can more effectively capture changes in student preferences, enhancing the real-time nature and adaptability of personalized matching. Next, based on the target multi-source dataset and the multi-student heterogeneous database, a comparison multi-source dataset is obtained. This step uses random student sampling and time alignment mechanisms to obtain the comparison multi-source dataset, establishing an objective comparison benchmark. Compared to subjective or direct comparison methods in existing technologies, this approach reduces human bias and provides reliable data support for subsequent fusion importance analysis. Furthermore, this solution utilizes comparative multi-source datasets to perform fusion importance analysis on the target multi-source dataset, obtaining a fusion index. By summarizing the fusion indices corresponding to the target multi-source dataset, multiple fusion indices are obtained. This step employs a machine learning classifier for fusion importance analysis, generating fusion indices and achieving a quantitative assessment of the uniqueness of student data. Compared to rule-based or experience-based methods in existing technologies, this improves the scientific rigor and automation of fusion decisions, making the matching process more accurate and efficient. Finally, based on multiple fusion indices, the target multi-source dataset and the heterogeneous database of multiple students are fused to obtain a fused heterogeneous database. This step dynamically implements differentiated fusion strategies based on the fusion indices, such as data referencing or independent storage, achieving adaptive optimization of data management. Compared to the one-size-fits-all storage method in existing technologies, this avoids data redundancy while ensuring the integrity of unique data, improving system resource utilization efficiency and matching quality. Therefore, this invention can improve the accuracy of personalized student matching and the efficiency of school database resource utilization.

[0063] like Figure 2 The diagram shown is a functional block diagram of a personalized matching system based on multi-source heterogeneous data fusion provided in an embodiment of the present invention.

[0064] The personalized matching system 100 based on multi-source heterogeneous data fusion described in this invention can be installed in an electronic device. Depending on the functions implemented, the personalized matching system 100 based on multi-source heterogeneous data fusion may include a matching instruction receiving module 101, a multi-source data acquisition module 102, a heterogeneous data acquisition module 103, and a multi-source data fusion module 104. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and which are stored in the memory of the electronic device.

[0065] The matching instruction receiving module 101 is used to receive student data matching instructions and identify the target student and matching data source set based on the student data matching instructions. The matching data source set includes multiple matching data sources, and the matching data sources are library data sources, academic affairs data sources, consumption data sources and activity data sources. The multi-source data acquisition module 102 is used to acquire multiple multi-source datasets to be matched for the target student based on a preset data matching period and a matching data source set. The multi-source datasets to be matched correspond one-to-one with the matching data sources, and each multi-source dataset to be matched includes multiple multi-source datasets to be matched. The heterogeneous data acquisition module 103 is used to identify a multi-student heterogeneous database, wherein the multi-student heterogeneous database includes multiple student multi-source databases, and each student multi-source database corresponds to a student ID. The module sequentially extracts the multi-source datasets to be matched from multiple multi-source datasets to be matched, and records the extracted multi-source datasets to be matched as the target multi-source dataset. The module then obtains a comparison multi-source dataset based on the target multi-source dataset and the multi-student heterogeneous database. The multi-source data fusion module 104 is used to perform fusion importance analysis on the target multi-source dataset by comparing the multi-source datasets, obtain a fusion index, summarize the fusion indices corresponding to the target multi-source datasets, obtain multiple fusion indices, fuse the target multi-source datasets and the multi-student heterogeneous database according to the multiple fusion indices, obtain a fused heterogeneous database, and push student resources based on the fused heterogeneous database to obtain pushed resources.

[0066] In detail, the modules in the personalized matching system 100 based on multi-source heterogeneous data fusion described in this embodiment of the invention adopt the same approach as described above when in use. Figure 1 The method uses the same techniques as the personalized matching method based on multi-source heterogeneous data fusion described in the article, and can produce the same technical effects, so it will not be elaborated here.

[0067] like Figure 3 The diagram shown is a structural schematic of an electronic device that implements a personalized matching method based on multi-source heterogeneous data fusion, according to an embodiment of the present invention.

[0068] The electronic device 1 may include a processor 10, a memory 11 and a bus 12, and may also include a computer program stored in the memory 11 and capable of running on the processor 10, such as a personalized matching method program based on multi-source heterogeneous data fusion.

[0069] The memory 11 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as the portable hard drive of the electronic device 1. In other embodiments, the memory 11 can be an external storage device of the electronic device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device 1. Furthermore, the memory 11 includes both internal storage units and external storage devices of the electronic device 1. The memory 11 can be used not only to store application software and various types of data installed on the electronic device 1, such as the code of a personalized matching method program based on multi-source heterogeneous data fusion, but also to temporarily store data that has been output or will be output.

[0070] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., personalized matching method programs based on multi-source heterogeneous data fusion), and calls data stored in the memory 11 to perform various functions of the electronic device 1 and process data.

[0071] The bus 12 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is configured to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0072] Figure 3 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 3 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0073] For example, although not shown, the electronic device 1 may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management system, thereby enabling functions such as charging management, discharging management, and power consumption management through the power management system. The power supply may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0074] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the electronic device 1 and other electronic devices.

[0075] Optionally, the electronic device 1 may further include a student interface, which may be a display or an input unit (such as a keyboard). Optionally, the student interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device 1 and to display a visual student interface.

[0076] The personalized matching method program based on multi-source heterogeneous data fusion stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When run in the processor 10, it can achieve the following: Receive student data matching instructions, and identify the target student and matching data source set based on the student data matching instructions. The matching data source set includes multiple matching data sources, and the matching data sources are library data sources, academic affairs data sources, consumption data sources and activity data sources. Based on a preset data matching period and a set of matching data sources, multiple multi-source datasets to be matched for the target student are obtained. Each multi-source dataset to be matched corresponds one-to-one with a matching data source, and each multi-source dataset to be matched includes multiple multi-source datasets to be matched. A multi-student heterogeneous database was identified, which includes multiple student multi-source databases, and each student multi-source database corresponds to a student ID. Extract the multi-source datasets to be matched sequentially from multiple multi-source datasets to be matched, and denote the extracted multi-source datasets to be matched as the target multi-source dataset. Obtain the comparison multi-source dataset based on the target multi-source dataset and the heterogeneous database of multiple students. By comparing the target multi-source datasets, we can perform a fusion importance analysis on the target multi-source datasets to obtain a fusion index. By summing the fusion indices corresponding to the target multi-source datasets, we can obtain multiple fusion indices. The target multi-source dataset and multi-student heterogeneous database are fused based on multiple fusion indices to obtain a fused heterogeneous database. Student resources are then pushed based on the fused heterogeneous database to obtain the pushed resources.

[0077] Specifically, the processor 10's implementation method for the above instructions can be found in [reference needed]. Figures 1 to 3 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.

[0078] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or system capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).

[0079] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following: Receive student data matching instructions, and identify the target student and matching data source set based on the student data matching instructions. The matching data source set includes multiple matching data sources, and the matching data sources are library data sources, academic affairs data sources, consumption data sources and activity data sources. Based on a preset data matching period and a set of matching data sources, multiple multi-source datasets to be matched for the target student are obtained. Each multi-source dataset to be matched corresponds one-to-one with a matching data source, and each multi-source dataset to be matched includes multiple multi-source datasets to be matched. A multi-student heterogeneous database was identified, which includes multiple student multi-source databases, and each student multi-source database corresponds to a student ID. Extract the multi-source datasets to be matched sequentially from multiple multi-source datasets to be matched, and denote the extracted multi-source datasets to be matched as the target multi-source dataset. Obtain the comparison multi-source dataset based on the target multi-source dataset and the heterogeneous database of multiple students. By comparing the target multi-source datasets, we can perform a fusion importance analysis on the target multi-source datasets to obtain a fusion index. By summing the fusion indices corresponding to the target multi-source datasets, we can obtain multiple fusion indices. The target multi-source dataset and multi-student heterogeneous database are fused based on multiple fusion indices to obtain a fused heterogeneous database. Student resources are then pushed based on the fused heterogeneous database to obtain the pushed resources.

[0080] In the embodiments provided by this invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and actual implementations may have other classification methods.

[0081] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0082] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0083] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A personalized matching method based on multi-source heterogeneous data fusion, characterized in that, The method includes: Receive student data matching instructions, and identify the target student and matching data source set based on the student data matching instructions. The matching data source set includes multiple matching data sources, and the matching data sources are library data sources, academic affairs data sources, consumption data sources and activity data sources. Based on a preset data matching period and a set of matching data sources, multiple multi-source datasets to be matched for the target student are obtained. Each multi-source dataset to be matched corresponds one-to-one with a matching data source, and each multi-source dataset to be matched includes multiple multi-source datasets to be matched. A multi-student heterogeneous database was identified, which includes multiple student multi-source databases, and each student multi-source database corresponds to a student ID. Extract the multi-source datasets to be matched sequentially from multiple multi-source datasets to be matched, and denote the extracted multi-source datasets to be matched as the target multi-source dataset. Obtain the comparison multi-source dataset based on the target multi-source dataset and the heterogeneous database of multiple students. By comparing the target multi-source datasets, we can perform a fusion importance analysis on the target multi-source datasets to obtain a fusion index. By summing the fusion indices corresponding to the target multi-source datasets, we can obtain multiple fusion indices. The target multi-source dataset and multi-student heterogeneous database are fused based on multiple fusion indices to obtain a fused heterogeneous database. Student resources are then pushed based on the fused heterogeneous database to obtain the pushed resources.

2. The personalized matching method based on multi-source heterogeneous data fusion as described in claim 1, characterized in that, The process of obtaining and comparing multi-source datasets based on the target multi-source dataset and multi-student heterogeneous databases includes: Random student sampling is performed on a multi-student heterogeneous database based on the target multi-source dataset to obtain a comparison student group, which includes multiple comparison students; In the comparison student group, the comparison students are extracted sequentially, and the comparison multi-source database of the extracted comparison students is obtained in the multi-student heterogeneous database; Identifying and comparing multi-source data groups in a multi-source database based on data matching cycles; Merge the comparative multi-source datasets corresponding to each student to obtain the comparative multi-source dataset.

3. The personalized matching method based on multi-source heterogeneous data fusion as described in claim 2, characterized in that, The method of performing fusion importance analysis on the target multi-source dataset by comparing multi-source datasets to obtain a fusion index includes: The target multi-source dataset is divided into training multi-source datasets and validation multi-source datasets. A training classification dataset is constructed based on the comparison of multi-source datasets and the training multi-source dataset; The target classifier is obtained by training a pre-built classifier using a training classification dataset. For each validation multi-source data in the validation multi-source dataset, features are constructed to obtain multiple validation data feature sets. The validation data feature sets in the multiple validation data feature sets correspond one-to-one with the validation multi-source data in the validation multi-source dataset. Multiple validation data feature sets are input into the target classifier to obtain multiple classification probabilities. The average classification probability is calculated by averaging the multiple classification probabilities and is denoted as the fusion index.

4. The personalized matching method based on multi-source heterogeneous data fusion as described in claim 3, characterized in that, The construction of the training classification dataset based on the comparison of multi-source datasets and the training multi-source dataset includes: Perform the following operation on each training multi-source data in the training multi-source dataset: Features are constructed from training multi-source data to obtain a training data feature set, which includes multiple training data features. The training data feature set is labeled based on the preset positive sample labels to obtain the labeled training feature set; In the comparison multi-source dataset, comparison multi-source data are extracted sequentially, and a comparison data feature set is constructed based on the extracted comparison multi-source data. The comparison data feature set is labeled according to the preset negative sample labels to obtain the labeled comparison feature set. Merge the labeled training feature set and the labeled contrast feature set to obtain training classification data; The training classification data corresponding to each comparison multi-source data is summarized to obtain multiple training classification data. The multiple training classification data corresponding to each training multi-source data are merged to obtain the training classification dataset.

5. The personalized matching method based on multi-source heterogeneous data fusion as described in claim 4, characterized in that, The process of fusing the target multi-source dataset and the multi-student heterogeneous database based on multiple fusion indices to obtain a fused heterogeneous database includes: Construct a target data storage structure, which includes multiple data source storage addresses, and each data source storage address corresponds one-to-one with the target multi-source data in the target multi-source dataset; Obtain the classification fitting threshold set, which includes: high fitting threshold and low fitting threshold; Set multiple data source weights based on multiple matching data sources; By using the weights of multiple data sources to calculate the weighted average of multiple fusionable indices, a comprehensive fit value is obtained; If the overall fit value is greater than the high fit threshold in the classification fit threshold group, then the preset empty set is recorded as the target student database. If the overall fit value is less than the low fit threshold in the classification fit threshold group, then data fusion is performed using the multi-student heterogeneous database and the target data storage structure to obtain the target student database. If the overall fit value is not greater than the high fit threshold and not less than the low fit threshold, then the target student database is generated based on the target multi-source dataset and the target data storage structure. Add the target student database to the multi-student heterogeneous database to obtain a merged heterogeneous database.

6. The personalized matching method based on multi-source heterogeneous data fusion as described in claim 5, characterized in that, The step of obtaining the classification fitting threshold set includes: Based on the data matching period query, the historical fitting threshold group is queried, where the historical fitting threshold group includes: historical high fitting threshold and historical low fitting threshold. Perform data statistics on a heterogeneous database with multiple students to obtain the current data volume and the total number of students; Receive student matching weight reorganization, wherein student matching weight reorganization includes: stable matching weight and generalized matching weight; The historical fitting threshold group is adjusted based on the current data volume, total number of students, and student matching weights to obtain the classification fitting threshold group.

7. The personalized matching method based on multi-source heterogeneous data fusion as described in claim 6, characterized in that, The process of fusing data using a multi-student heterogeneous database and a target data storage structure to obtain a target student database includes: Extract target multi-source data sequentially from the target multi-source dataset, identify the target data source corresponding to the extracted target multi-source data, and identify the target storage address corresponding to the target data source in the target data storage structure; Based on the target data source and heterogeneous databases of multiple students, query similar data from the same source. Obtain the data storage address of similar data from the same source; The source data storage address is pointed to the target storage address to obtain the fused storage address; The merged storage addresses corresponding to each target data source are aggregated to obtain multiple merged storage addresses. The target data storage structure is then updated using these multiple merged storage addresses to obtain the target student database.

8. The personalized matching method based on multi-source heterogeneous data fusion as described in claim 7, characterized in that, The step of querying similar, homogeneous data based on the target data source and multi-student heterogeneous databases includes: Based on the target data source, the homogeneous database of multiple students is filtered to obtain the homogeneous student dataset, which includes multiple homogeneous student datasets. Perform the following operation on each student's data in the student common-origin dataset: Construct a feature vector for the same source data based on the same source data of students, and construct a feature vector for the target data based on the extracted target multi-source data; Calculate data similarity based on the feature vectors of the source data and the feature vectors of the target data; Summarize the data similarity corresponding to each student's source data to obtain a data similarity set. Determine the maximum similarity in the data similarity set and record the source data of the student corresponding to the maximum similarity as similar source data.

9. The personalized matching method based on multi-source heterogeneous data fusion as described in claim 8, characterized in that, The process of generating the target student database based on the target multi-source dataset and the target data storage structure includes: Perform the following operation on each of the multiple data source storage addresses in the target data storage structure: Create a database partition for the data source storage address to obtain a storable address, which includes metadata indexes and update logs; Based on storable addresses, homogeneous stored data were identified in the target multi-source dataset. Write the same source storage data to a storable address to obtain a valid storage address; By summarizing the valid storage addresses, multiple valid storage addresses are obtained. The target data storage structure is then updated using these multiple valid storage addresses to obtain the target student database.

10. A personalized matching system based on multi-source heterogeneous data fusion, characterized in that, The system includes: The matching instruction receiving module is used to receive student data matching instructions, and to identify the target student and the matching data source set based on the student data matching instructions. The matching data source set includes multiple matching data sources, and the matching data sources are library data sources, academic affairs data sources, consumption data sources and activity data sources. The multi-source data acquisition module is used to acquire multiple multi-source datasets to be matched for the target student based on a preset data matching period and a matching data source set. The multi-source datasets to be matched correspond one-to-one with the matching data sources, and each multi-source dataset to be matched includes multiple multi-source datasets to be matched. The heterogeneous data acquisition module is used to identify a multi-student heterogeneous database, which includes multiple student multi-source databases, and each student multi-source database corresponds to a student ID. The module sequentially extracts the multi-source datasets to be matched from multiple datasets to be matched, and records the extracted multi-source datasets to be matched as the target multi-source dataset. The module then obtains a comparison multi-source dataset based on the target multi-source dataset and the multi-student heterogeneous database. The multi-source data fusion module is used to perform fusion importance analysis on the target multi-source dataset by comparing the multi-source datasets, obtain the fusion index, summarize the fusion indices corresponding to the target multi-source datasets, obtain multiple fusion indices, fuse the target multi-source datasets and multi-student heterogeneous databases according to the multiple fusion indices, obtain the fused heterogeneous database, and push student resources based on the fused heterogeneous database to obtain the pushed resources.