Multi-source heterogeneous and dynamic behavior collaborative feature mining method and system

By extracting learning cycle parameters and training feature mining models, the problem of inefficient collaborative feature mining in the existing technology is solved, efficient fusion of learners' multi-source heterogeneous data and dynamic behavior data and rapid and accurate mining of collaborative features is achieved, and the analysis and evaluation efficiency of learning results is improved.

CN119494081BActive Publication Date: 2025-05-06CHAOHU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510080580.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing collaborative feature mining methods cannot efficiently and accurately realize collaborative feature mining of multi-dimensional, multi-type and multi-state, resulting in inefficient mining of collaborative feature mining.

Method used

By querying the learner's basic learning parameters, extracting period parameters and calculating the learning cycle, obtaining multi-source heterogeneous data and dynamic behavior data, fusing them into comprehensive learning data, mining of collaborative features, and training feature mining models to realize the mining of real-time collaborative features and determining management prompts.

Benefits of technology

It realizes efficient collection and fusion of learners' multi-source heterogeneous data and dynamic behavioral data, quickly and accurately mines out collaborative features, improves the analysis and evaluation efficiency of learning results, and meets the needs of multi-dimensional, multi-type, and multi-state collaborative management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119494081B_ABST
    Figure CN119494081B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing technology, and discloses a method and system for mining collaborative features of multi-source heterogeneous and dynamic behaviors; the method comprises calculating a learning cycle of a learner, acquiring multi-source heterogeneous data and dynamic behavior data, training a feature mining model for mining collaborative features, acquiring real-time comprehensive learning data, determining whether to issue a collaborative management prompt, identifying a target feature from the collaborative feature, and marking the collaborative management data according to the target feature; compared with the prior art, the present invention can provide a reasonable and effective data basis for the accurate calculation of the learning cycle, avoid the occurrence of excessive and out-of-range multi-source heterogeneous data and dynamic behavior data, and at the same time, quickly and accurately mine collaborative features based on the integrated learning data generated by fusion, thereby achieving a combined effect of rapid screening, identification, and precise mining of complex data, and can also perform multi-dimensional, multi-type, and multi-state collaborative management operations on the learner's learning situation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and more specifically, to a method and system for mining multi-source heterogeneous and dynamic behavior collaborative features. Background Art

[0002] Multi-source heterogeneous data refers to data with different sources and different structural types. Dynamic behavior data refers to the behavior data of an individual in duration or space. In the analysis and evaluation of learners' learning outcomes, it is necessary to collect learners' multi-source heterogeneous data and dynamic behavior data, and analyze and evaluate the learning outcomes by mining the synergistic features in multi-source heterogeneous data and dynamic behavior data, so as to achieve effective analysis and evaluation of learning outcomes.

[0003] The patent application with reference publication number CN112214531A discloses a feature mining method and component across data, information, and knowledge multimodalities, including obtaining data resources to be mined, classifying data resources, and performing association fusion processing on type data, and determining the association fusion result as the feature of the data resource. Vector data includes entity positions in a coordinate system, shapes of map graphics or geographic entities, geographic action trajectories, and physical quantities with both size and direction. Range data includes continuous range data and discrete range data. In order to make the features more reliable, after obtaining the type data, the type data is processed by an association fusion method to obtain an association fusion result.

[0004] Existing collaborative feature mining methods can analyze and evaluate learners' learning outcomes by classifying and fusing relevant data of learners and mining the features with associated fusion effects after fusion. For example, in the above-mentioned patent application, it classifies data resources, performs associated fusion processing on type data, and determines the associated fusion results as the characteristics of data resources, thereby achieving collaborative feature mining effects. Since learners generate a large amount of data from different sources and different types during the learning process, and learners' learning behaviors also change dynamically, the collaborative feature mining method after fusion analysis cannot efficiently and accurately achieve multi-dimensional, multi-type and multi-state collaborative feature mining effects, thereby reducing the efficiency of collaborative feature mining.

[0005] In view of this, the present invention proposes a multi-source heterogeneous and dynamic behavior collaborative feature mining method and system to solve the above problems. Summary of the invention

[0006] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a multi-source heterogeneous and dynamic behavior collaborative feature mining method, applied to a mining server, comprising:

[0007] S1: query the learner's basic learning parameters, extract the cycle parameters from the basic learning parameters, the cycle parameters include the effective learning peak value, the learning interruption value and the learning fatigue, and calculate the learner's learning cycle;

[0008] S2: obtaining first learning data of the learner in the learning cycle, and extracting multi-source heterogeneous data from the first learning data based on a first extraction criterion, wherein the first extraction criterion is: eliminating redundant and repeated first learning data;

[0009] S3: obtaining the second learning data of the learner in the learning cycle, and extracting the dynamic behavior data from the second learning data based on the second extraction criterion, wherein the second extraction criterion is: eliminating the data to be verified whose recording time is not in the last three digits;

[0010] S4: Integrate multi-source heterogeneous data and dynamic behavior data into comprehensive learning data, mine collaborative features in the comprehensive learning data, and train a feature mining model that mines collaborative features;

[0011] S5: Obtain real-time comprehensive learning data, input it into the feature mining model to mine real-time collaborative features, and determine whether to issue collaborative management prompts;

[0012] S6: If a collaborative management prompt is issued, the target feature is identified from the collaborative features, the collaborative management data is marked according to the target feature, and the collaborative features and collaborative management data are integrated for display to the learner and the manager.

[0013] Furthermore, the method for obtaining the learning discontinuity value includes:

[0014] Through the learning management system, we can find out the learners’ learning events, and query them one by one The first recorded moment and the last recorded moment in a learning event are recorded as the starting moment and the ending moment respectively;

[0015] The duration between the end time of the previous learning event and the start time of the next learning event is recorded as the interval time, and the The length of the interval;

[0016] Remove the maximum and minimum interval durations and keep the remaining The duration of each interval is accumulated and averaged to obtain the learning interruption value;

[0017] The expression for the learning discontinuity value is:

[0018] ;

[0019] In the formula, is the learning discontinuity value, For the The interval duration.

[0020] Furthermore, the method for obtaining learning fatigue includes:

[0021] Query learners one by one The event state attributes of a learning event are obtained, and the attribute values ​​in the event state attributes are marked;

[0022] The learning events with attribute values ​​of "0-0", "0-1" or "1-0" are recorded as fatigue events, and the duration of fatigue events and learning events are queried one by one through timestamps to obtain Fatigue duration and The duration of the event;

[0023] Will After the fatigue durations are accumulated one by one, the total fatigue value is obtained, and the total fatigue value is combined with Compare the accumulated values ​​of the duration of each event to obtain the degree of learning fatigue;

[0024] The expression of learning fatigue is:

[0025] ;

[0026] In the formula, To learn fatigue, For the Fatigue duration, For the The duration of an event.

[0027] Furthermore, the learning cycle calculation method includes:

[0028] After assigning different weight factors to the effective learning peak value, learning interruption value and learning fatigue, and comparing them, the cycle added value is calculated;

[0029] The expression of periodic additional value is:

[0030] ;

[0031] In the formula, Add value to the cycle, To effectively learn the peak value, , , All are weight factors greater than 0;

[0032] The learning cycle is calculated by adding the preset standard cycle to the cycle additional value;

[0033] The expression of the learning cycle is:

[0034] ;

[0035] In the formula, For the learning cycle, The default standard cycle.

[0036] Furthermore, multi-source heterogeneous data include test score data, daily performance data, extracurricular activity data, book review data, and hands-on practice data;

[0037] Methods for extracting test score data, daily performance data, extracurricular activity data, book review data, and hands-on practice data include:

[0038] Repeating and comparing all first learning data, marking duplicate first learning data, and searching for recording times of duplicate first learning data;

[0039] Arrange the duplicate first learning data in sequence according to the order of recording time, and remove the duplicate first learning data except the last one;

[0040] Using natural language processing technology, the first semantics in the remaining first learning data are identified one by one, and the first semantics are split into a meaning part and a value part;

[0041] The first semantics whose meanings are grades, daily life, extracurricular activities, books and practical operations are recorded as test semantics, daily semantics, extracurricular semantics, book semantics and practical operation semantics respectively, and the first learning data corresponding to the test semantics, daily semantics, extracurricular semantics, book semantics and practical operation semantics are recorded as test grade data, daily performance data, extracurricular activity data, book reference data and hands-on operation data.

[0042] Furthermore, dynamic behavior data include the duration of test answering, the number of class notes, the frequency of interactive discussions, the number of searches and borrowings, and the effective practice duration;

[0043] The extraction methods of test answering time, number of class notes, frequency of interactive discussion, number of search and borrowing times, and effective practice time include:

[0044] Identify the second semantics of the second learning data one by one through natural language processing technology, and split the second semantics into a text part and a numerical part;

[0045] The textual parts of the second learning data, which are answer sheets, notes, discussions, borrowings and valid ones, are respectively recorded as the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified, and the recording times of the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified are queried one by one through the timestamps;

[0046] Arrange the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified in sequence according to the order of recording time, and remove the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified except the last three digits;

[0047] Add up the numerical parts of the remaining first data to be verified, second data to be verified, third data to be verified, fourth data to be verified and fifth data to be verified and average them to obtain the test answering time, number of class notes, frequency of interactive discussions, number of searches and borrowings and effective practice time respectively.

[0048] Furthermore, the training method of the feature mining model includes:

[0049] Collect multiple sets of comprehensive learning data and corresponding collaborative features, and convert each set of comprehensive learning data into a corresponding set of feature vectors;

[0050] The feature vector is used as the input of the machine learning model, and the collaborative features corresponding to each group of comprehensive learning data are used as the output of the machine learning model. The collaborative features actually corresponding to each group of comprehensive learning data are used as prediction targets, and minimizing the sum of prediction errors of the comprehensive learning data is used as the training target. The machine learning model is trained until the sum of prediction errors converges and the training is stopped to obtain a feature mining model that mines collaborative features.

[0051] Furthermore, the method for determining whether to issue a collaborative management prompt includes:

[0052] Count the number of collaborative features mined by the feature mining model and record it as the feature value;

[0053] When the characteristic value is less than or equal to 3, it is determined that a collaborative management prompt is issued;

[0054] When the characteristic value is greater than 3, it is determined that no collaborative management prompt is issued.

[0055] Furthermore, the tagging method for collaboratively managing data includes:

[0056] Arrange all collaborative features in order according to the mining sequence, and identify the semantic features of all collaborative features one by one through natural language processing technology;

[0057] When the feature semantics is exam, the exam score data and the exam answer time are recorded as collaborative management data;

[0058] When the feature semantics is notes, the daily performance data and the number of class notes are recorded as collaborative management data;

[0059] When the feature semantics is interactive, the extracurricular activity data and the frequency of interactive discussions are recorded as collaborative management data;

[0060] When the feature semantics is reading, the book lookup data and search borrowing times are recorded as collaborative management data;

[0061] When the feature semantics is operation, the hands-on operation data and effective practice duration are recorded as collaborative management data.

[0062] A multi-source heterogeneous and dynamic behavior collaborative feature mining system, applied to a mining server, for implementing the multi-source heterogeneous and dynamic behavior collaborative feature mining method, comprising a cycle calculation module, a first acquisition module, a second acquisition module, a model training module, a prompt determination module and a collaborative management module, wherein each module is connected via a wired or wireless network;

[0063] The cycle calculation module is used to query the basic learning parameters of the learner, extract the cycle parameters from the basic learning parameters, the cycle parameters include the effective learning peak value, the learning interruption value and the learning fatigue, and calculate the learner's learning cycle;

[0064] The first acquisition module is used to obtain the first learning data of the learner in the learning cycle, and extract multi-source heterogeneous data from the first learning data based on a first extraction criterion, wherein the first extraction criterion is: eliminating redundant and repeated first learning data;

[0065] The second acquisition module is used to obtain the second learning data of the learner in the learning cycle, and extract the dynamic behavior data from the second learning data based on the second extraction criterion, and the second extraction criterion is: eliminate the data to be verified whose recording time is not located in the last three digits;

[0066] Model training module, used to fuse multi-source heterogeneous data and dynamic behavior data into comprehensive learning data, mine collaborative features in the comprehensive learning data, and train feature mining models that mine collaborative features;

[0067] The prompt determination module is used to obtain real-time comprehensive learning data, input it into the feature mining model to mine real-time collaborative features, and determine whether to issue collaborative management prompts;

[0068] The collaborative management module is used to identify target features from collaborative features, mark collaborative management data according to the target features, and integrate the collaborative features and collaborative management data for display to learners and managers.

[0069] The technical effects and advantages of the multi-source heterogeneous and dynamic behavior collaborative feature mining method and system of the present invention are as follows:

[0070] The present invention obtains the learner's first learning data in the learning cycle by querying the learner's basic learning parameters, extracts the cycle parameters from the basic learning parameters, calculates the learner's learning cycle, extracts multi-source heterogeneous data from the first learning data based on a first extraction criterion, obtains the learner's second learning data in the learning cycle, extracts dynamic behavior data from the second learning data based on a second extraction criterion, fuses the multi-source heterogeneous data and the dynamic behavior data into comprehensive learning data, mines out collaborative features in the comprehensive learning data, trains a feature mining model for mining collaborative features, obtains real-time comprehensive learning data, inputs it into the feature mining model to mine out real-time collaborative features, determines whether to issue collaborative management prompts, identifies target features from the collaborative features, marks collaborative management data according to the target features, and integrates the collaborative features and collaborative management data to learners and managers for display; Compared with the existing technology, by extracting cycle parameters from basic learning parameters, it is possible to provide a reasonable and effective data basis for the accurate calculation of the learning cycle, thereby providing a reasonable time limit for the collection of learners' multi-source heterogeneous data and dynamic behavior data, avoiding excessive and out-of-range multi-source heterogeneous data and dynamic behavior data. At the same time, by training a feature mining model, it is possible to quickly and accurately mine collaborative features based on the integrated learning data generated by fusion, and identify target features that have a negative impact on learners' learning outcomes in the collaborative features, thereby achieving a combined effect of rapid screening, identification, and precise mining of complex data, and can also perform multi-dimensional, multi-type, and multi-state collaborative management operations on learners' learning situations, avoiding the inefficiency and error-proneness of multi-dimensional data and multi-type data in analysis and evaluation, thereby meeting the needs of learners for rapid and accurate analysis and evaluation of their learning outcomes. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 A schematic diagram of the process of the multi-source heterogeneous and dynamic behavior collaborative feature mining method provided in the first embodiment of the present invention;

[0072] Figure 2 A schematic diagram of the architecture of a multi-source heterogeneous and dynamic behavior collaborative feature mining system provided in the second embodiment of the present invention;

[0073] Figure 3 A schematic diagram of the modules of the mining server provided in the second embodiment of the present invention;

[0074] Figure 4 A schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention;

[0075] Figure 5 A schematic diagram of the structure of a computer-readable storage medium provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION

[0076] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0077] Example 1: Please refer to Figure 1 As shown, the multi-source heterogeneous and dynamic behavior collaborative feature mining method described in this embodiment is applied to a mining server, including:

[0078] S1: query the learner's basic learning parameters, extract the cycle parameters from the basic learning parameters, and calculate the learner's learning cycle;

[0079] Basic learning parameters refer to parameters that can represent the learning changes of learners during the learning process, and serve as the original parameters for subsequent analysis and evaluation of learners' learning outcomes within a reasonable period of time, thereby providing a basis for calculating learners' learning cycles;

[0080] Cycle parameters refer to various parameters that can affect the length of the learning cycle. They can also represent the changes of learners during the learning process, thereby improving the calculation accuracy of subsequent learning cycles.

[0081] The cycle parameters include effective learning peak value, learning interruption value and learning fatigue;

[0082] The effective learning peak refers to the maximum duration that the learner is in a normal learning state in the past time period, that is, it can represent the maximum value of the learner's continuous learning time in the past time period. The larger the effective learning peak, the longer the duration that represents the learner's learning situation, and the longer the learning cycle. The effective learning peak is obtained by querying the maximum value of the continuous duration that the learner is in a normal learning state in the past time period.

[0083] The learning discontinuity value refers to the interval between two adjacent learning stages of the learner in the past time period, which can be used to represent the continuity of the learner's learning in the past time period. When the learning discontinuity value is larger, the longer the time for representing the learner's learning situation is, the longer the learning cycle is.

[0084] Methods for obtaining learning discontinuity values ​​include:

[0085] Through the learning management system, we can find out the learners’ learning events, and query them one by one The first recorded moment and the last recorded moment in a learning event are recorded as the starting moment and the ending moment respectively; a learning event is used to record all the data of the learner in each independent learning stage in the past time period, that is, to record the data in each learning stage independently and completely;

[0086] The duration between the end time of the previous learning event and the start time of the next learning event is recorded as the interval time, and the The length of the interval;

[0087] Remove the maximum and minimum interval durations and keep the remaining The duration of each interval is accumulated and averaged to obtain the learning interruption value;

[0088] The expression for the learning discontinuity value is:

[0089] ;

[0090] In the formula, is the learning discontinuity value, For the The interval duration.

[0091] Learning fatigue refers to the proportion of time that learners are in a state of learning fatigue in the past time period, which can be used to indicate the degree of learning fatigue of learners in the past time period. The greater the learning fatigue, the shorter the time that indicates the learner's learning situation, and the shorter the learning cycle;

[0092] The methods of obtaining learning fatigue include:

[0093] Query learners one by one The event state attribute of each learning event is used to represent the fatigue state of the learning event, and the attribute value in the event state attribute is marked; the event state attribute is used to represent the fatigue state of the learning event as a whole, and the attribute value is the component unit of the event state attribute. Specifically, the attribute value includes 0-0, "0-1", "1-0" and "1-1";

[0094] The learning events with attribute values ​​of "0-0", "0-1" or "1-0" are recorded as fatigue events, and the duration of fatigue events and learning events are queried one by one through timestamps to obtain Fatigue duration and The duration of the event;

[0095] Will After the fatigue durations are accumulated one by one, the total fatigue value is obtained, and the total fatigue value is combined with Compare the accumulated values ​​of the duration of each event to obtain the degree of learning fatigue;

[0096] The expression of learning fatigue is:

[0097] ;

[0098] In the formula, To learn fatigue, For the Fatigue duration, For the The duration of an event.

[0099] After integrating and calculating the effective learning peak value, learning interruption value and learning fatigue, the learning cycle can be obtained, so that the learning cycle can be used as a time limit for learners to conduct subsequent learning effect evaluation and analysis data collection;

[0100] The learning cycle calculation method includes:

[0101] After assigning different weight factors to the effective learning peak value, learning interruption value and learning fatigue, and comparing them, the cycle added value is calculated;

[0102] The expression of periodic additional value is:

[0103] ;

[0104] In the formula, Add value to the cycle, To effectively learn the peak value, , , All are weight factors greater than 0;

[0105] in, , , , The setting is to balance the proportion of effective learning peak value, learning interruption value and learning fatigue in the cycle added value, so as to improve the calculation accuracy of the cycle added value and provide a good foundation for the accurate calculation of subsequent learning cycles;

[0106] The learning cycle is calculated by adding the preset standard cycle to the cycle additional value; the preset standard cycle is used to average the duration of data collected when evaluating the learning effect of different learners under normal circumstances, so as to serve as the basic value of the learning cycle and ensure that the calculation of subsequent learning cycles better meets the diverse needs of different types of learners, thereby improving the effectiveness of the learning cycle results; for example, the preset standard cycle can be one week or one month;

[0107] The expression of the learning cycle is:

[0108] ;

[0109] In the formula, For the learning cycle, The default standard cycle.

[0110] S2: obtaining first learning data of the learner in the learning cycle, and extracting multi-source heterogeneous data from the first learning data based on a first extraction criterion;

[0111] First, learning data refers to the diverse data of different sources and types generated by learners during the learning cycle, which can be used to comprehensively represent the external information that affects the learning of learners during the learning cycle;

[0112] Since the first learning data represents the multi-source and multi-type data of the learner during the learning cycle, all data that can affect the learner's learning outcomes need to be obtained. In this case, the learner's first learning data needs to be obtained from multiple different data channels and methods;

[0113] After the first learning data is obtained, the amount of the first learning data will be large, and there will also be a large amount of useless data that does not meet the subsequent learning outcome evaluation. Therefore, it is necessary to extract the useful data in the first learning data under the restriction of the first extraction criterion, so as to obtain multi-source heterogeneous data that meets the use of the subsequent learning outcome evaluation;

[0114] Multi-source heterogeneous data include test score data, daily performance data, extracurricular activity data, book reading data and hands-on practice data; among them, test score data is used to represent learners' relevant test performance during the learning cycle, daily performance data is used to represent learners' classroom learning performance during the learning cycle, extracurricular activity data is used to represent learners' extracurricular interest performance during the learning cycle, book reading data is used to represent learners' book browsing and reading performance during the learning cycle, and hands-on practice data is used to represent learners' experimental hands-on ability performance during the learning cycle;

[0115] Specifically, test score data is obtained through querying the test management system, daily performance data and extracurricular activity data are obtained through querying the learning management system, book reference data is obtained through querying the book management system, and hands-on practice data is obtained through querying the experiment management system.

[0116] The first extraction criterion is: remove redundant and repeated first learning data;

[0117] Methods for extracting test score data, daily performance data, extracurricular activity data, book review data, and hands-on practice data include:

[0118] Repeating and comparing all first learning data, marking duplicate first learning data, and searching for recording times of duplicate first learning data;

[0119] Arrange the duplicate first learning data in sequence according to the order of recording time, and remove the duplicate first learning data except the last one;

[0120] The first semantics in the remaining first learning data are identified one by one through natural language processing technology, and the first semantics are split into an interpretation part and an interpretation value part; the interpretation part is used to accurately interpret the meaning expressed by the text in the first semantics, and the interpretation value part is used to accurately interpret the meaning expressed by the numerical value in the first semantics, so that the text and numerical value of the first semantics can be effectively distinguished;

[0121] Recording the first semantics whose interpretation part is score as test semantics, and recording the first learning data corresponding to the test semantics as test score data;

[0122] Recording the first semantics whose interpretation part is daily as daily semantics, and recording the first learning data corresponding to the daily semantics as daily performance data;

[0123] The first semantics whose interpretation part is extracurricular is recorded as extracurricular semantics, and the first learning data corresponding to the extracurricular semantics is recorded as extracurricular activity data;

[0124] Recording the first semantics whose interpretation part is a book as the book semantics, and recording the first learning data corresponding to the book semantics as the book search data;

[0125] The first semantics whose interpretation part is practical operation is recorded as practical operation semantics, and the first learning data corresponding to the practical operation semantics is recorded as hands-on practical operation data.

[0126] It should be noted that by extracting unique and authentic multi-source heterogeneous data, it is possible to make a comprehensive representation of learners' learning outcomes during the learning cycle in multiple aspects and dimensions, providing a reasonable and authentic data basis for subsequent evaluation of learners' learning outcomes.

[0127] S3: obtaining second learning data of the learner in the learning cycle, and extracting dynamic behavior data from the second learning data based on a second extraction criterion;

[0128] Second, learning data refers to the diverse data of uninterrupted learning behaviors generated by learners during the learning cycle, which can represent the relevant learning behaviors of learners during the learning cycle;

[0129] Since the second learning data is an overall representation of the learner's learning behavior in the learning cycle, in order to achieve an accurate and concise representation effect, it is necessary to extract dynamic behavior data from the second learning data so that the dynamic behavior data can accurately represent the learner's learning behavior in the learning cycle;

[0130] In order to obtain accurate and reasonable dynamic behavior data, it is necessary to extract the dynamic behavior data under the restriction of the second extraction criterion to ensure that the extracted dynamic behavior data can better conform to the learner's learning behavior during the learning cycle;

[0131] Dynamic behavior data includes test answering time, number of class notes, interactive discussion frequency, search and borrowing times, and effective practice time; test answering time is used to indicate the time learners spend on completing test answers during the learning cycle; number of class notes is used to indicate the number of class notes learners take during the learning cycle; interactive discussion frequency is used to indicate the frequency of communication and discussion during extracurricular activities during the learning cycle; search and borrowing times is used to indicate the number of book inquiries and borrowings by learners during the learning cycle; and effective practice time is used to indicate the time learners spend on practical operations during the learning cycle;

[0132] Since the length of exam answering time, the number of class notes, the frequency of interactive discussions, the number of searches and borrowing times, and the effective practice time are respectively associated with the exam score data, daily performance data, extracurricular activity data, book reference data, and hands-on practice data, the length of exam answering time, the number of class notes, the frequency of interactive discussions, the number of searches and borrowing times, and the effective practice time are consistent with the channels and methods for obtaining the exam score data, daily performance data, extracurricular activity data, book reference data, and hands-on practice data, respectively.

[0133] The second extraction criterion is: remove the data to be verified except for the last three digits of the recording time;

[0134] The extraction methods of test answering time, number of class notes, frequency of interactive discussion, number of search and borrowing times, and effective practice time include:

[0135] Identify the second semantics of the second learning data one by one through natural language processing technology, and split the second semantics into a text part and a numerical part;

[0136] The textual parts of the second learning data, which are answer sheets, notes, discussions, borrowings and valid ones, are respectively recorded as the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified, and the recording times of the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified are queried one by one through the timestamps;

[0137] Arrange the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified in sequence according to the order of recording time, and remove the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified except the last three digits;

[0138] Add up the numerical parts of the remaining first data to be verified, second data to be verified, third data to be verified, fourth data to be verified and fifth data to be verified and average them to obtain the test answering time, number of class notes, frequency of interactive discussions, number of searches and borrowings and effective practice time respectively.

[0139] S4: Integrate multi-source heterogeneous data and dynamic behavior data into comprehensive learning data, mine collaborative features in the comprehensive learning data, and train a feature mining model that mines collaborative features;

[0140] After obtaining multi-source heterogeneous data and dynamic behavior data, it is necessary to fuse the multi-source heterogeneous data and dynamic behavior data so that they can form comprehensive learning data. The comprehensive learning data at this time can represent the comprehensive learning performance of learners in the learning cycle. At the same time, by extracting collaborative features from the comprehensive learning data, the features with collaborative correlation in the comprehensive learning data can be represented.

[0141] When mining collaborative features, it is first necessary to fuse multi-source heterogeneous data and dynamic behavior data together through a multidimensional data fusion algorithm and generate comprehensive learning data. At this time, the comprehensive learning data will contain features that associate multi-source heterogeneous data and dynamic behavior data, and they will be recorded as collaborative features.

[0142] After obtaining the comprehensive learning data and the collaborative features, a feature mining model that can mine the collaborative features corresponding to the comprehensive learning data can be trained based on the comprehensive learning data and the collaborative features, and the feature mining model can mine the collaborative features with the comprehensive learning data as output and the collaborative features as output;

[0143] The training methods of feature mining models include:

[0144] Collect multiple sets of comprehensive learning data and corresponding collaborative features, and convert each set of comprehensive learning data into a corresponding set of feature vectors;

[0145] The feature vector is used as the input of the machine learning model, and the collaborative features corresponding to each group of comprehensive learning data are used as the output of the machine learning model. The collaborative features actually corresponding to each group of comprehensive learning data are used as prediction targets, and minimizing the sum of prediction errors of the comprehensive learning data is used as the training target. The machine learning model is trained until the sum of prediction errors converges and the training is stopped to obtain a feature mining model that mines collaborative features.

[0146] Exemplarily, the machine learning model is any one of CNN or AlexNet;

[0147] The calculation formula for the prediction error is:

[0148] ;

[0149] In the formula, is the prediction error, is the group number of the eigenvector; For the The state value of the mining corresponding to the group feature vector, For the The actual state value corresponding to the group training data;

[0150] In the machine learning model, the feature vector is the comprehensive learning data and the state value is the collaborative feature.

[0151] S5: Obtain real-time comprehensive learning data, input it into the feature mining model to mine real-time collaborative features, and determine whether to issue collaborative management prompts;

[0152] After the feature mining model is trained, the real-time multi-source heterogeneous data and dynamic behavior data can be integrated to generate real-time comprehensive learning data. The real-time comprehensive learning data can be input into the feature mining model to mine the collaborative features corresponding to the real-time comprehensive learning data. Based on the mined collaborative features, it is determined whether to issue collaborative management prompts, and then collaborative management operations are performed on the learning outcomes of learners in the learning stage.

[0153] The determination method of whether to issue a collaborative management prompt includes:

[0154] Count the number of collaborative features mined by the feature mining model and record it as the feature value;

[0155] When the feature value is less than or equal to 3, the number of collaborative features that can be mined from the real-time comprehensive learning data is small, and the probability of inefficient and abnormal learning results of learners in the learning cycle is high, and it is determined to issue a collaborative management prompt;

[0156] When the feature value is greater than 3, the number of collaborative features that can be mined from the real-time comprehensive learning data is relatively large, and the probability of inefficient and abnormal learning outcomes of learners during the learning cycle is low, so it is determined that no collaborative management prompt will be issued.

[0157] It should be noted that collaborative management prompts will only be issued when learners have poor learning outcomes and low learning efficiency, so that collaborative management prompts can promptly remind learners and managers to take corresponding learning optimization measures, thereby improving learners' subsequent learning outcomes.

[0158] S6: If a collaborative management prompt is issued, identify the target feature from the collaborative features, mark the collaborative management data according to the target feature, and integrate the collaborative features and collaborative management data for display to learners and managers;

[0159] Target characteristics refer to specific synergistic characteristics that can have a negative impact on learners' learning outcomes during the learning cycle and serve as the object of the learners' learning outcome report during the learning cycle;

[0160] After identifying the target features, it is necessary to mark the collaborative management data from the comprehensive learning data based on the target features, integrate the obtained target features and collaborative management data, and send them to learners and managers in a unified manner, so as to provide a basis for analyzing and evaluating the learners' learning outcomes and learning status during the learning cycle;

[0161] Tagging methods for collaboratively managing data include:

[0162] All collaborative features are arranged in order according to the mining sequence, and the feature semantics of all collaborative features are identified one by one through natural language processing technology; feature semantics is used to accurately and concisely express the true meaning of collaborative features;

[0163] When the feature semantics is examination, it means that the data related to the examination in the comprehensive learning data is collaborative management data, and the examination score data and the examination answer time are recorded as collaborative management data;

[0164] When the feature semantics is notes, it means that the data related to notes in the comprehensive learning data are collaborative management data, and the daily performance data and the number of class notes are recorded as collaborative management data;

[0165] When the feature semantics is interaction, it means that the data related to interaction in the comprehensive learning data is collaborative management data, and the extracurricular activity data and the frequency of interactive discussion are recorded as collaborative management data;

[0166] When the feature semantics is reading, it means that the data related to reading in the comprehensive learning data is collaborative management data, and the book search data and search borrowing times are recorded as collaborative management data;

[0167] When the feature semantics is operation, it means that the data related to the operation in the comprehensive learning data is collaborative management data, and the hands-on operation data and effective practice time are recorded as collaborative management data.

[0168] After obtaining the target features and collaborative management data, it is necessary to integrate the target features and collaborative management data together to display the learning outcomes of learners in the learning cycle. This can be used to perform multi-dimensional, multi-type and multi-state collaborative management representation of learners' learning situations, making it easier for learners and managers to carry out subsequent learning plans more efficiently and accurately. It also enables rapid screening and identification of complex data, and accurate mining of collaborative features.

[0169] Specifically, when there is only one target feature, the target feature and the collaborative management data are compressed into one integrated data, and the integrated data is sent to learners and managers for display; when there is more than one target feature, each target feature and the corresponding collaborative management data are compressed into one integrated data, and multiple integrated data are sent to learners and managers for display at the same time, thereby providing a basis for subsequent analysis and evaluation of learners' learning outcomes, achieving the effect of multi-source heterogeneous and dynamic behavior collaborative feature mining, and meeting the needs of learners for analysis and evaluation of learning outcomes.

[0170] In this embodiment, the learner's basic learning parameters are queried, the cycle parameters are extracted from the basic learning parameters, and the learner's learning cycle is calculated, the learner's first learning data in the learning cycle is obtained, and based on the first extraction criterion, multi-source heterogeneous data is extracted from the first learning data, the learner's second learning data in the learning cycle is obtained, and based on the second extraction criterion, dynamic behavior data is extracted from the second learning data, the multi-source heterogeneous data and the dynamic behavior data are merged into comprehensive learning data, the collaborative features in the comprehensive learning data are mined, and a feature mining model that mines the collaborative features is trained, real-time comprehensive learning data is obtained, and the data is input into the feature mining model to mine real-time collaborative features, and it is determined whether to issue collaborative management prompts, the target features are identified from the collaborative features, the collaborative management data is marked according to the target features, and the collaborative features and collaborative management data are integrated to the learners and managers for display. Compared with the existing technology, by extracting cycle parameters from basic learning parameters, it can provide a reasonable and effective data basis for the accurate calculation of the learning cycle, thereby providing a reasonable time limit for the collection of learners' multi-source heterogeneous data and dynamic behavior data, avoiding excessive and out-of-range multi-source heterogeneous data and dynamic behavior data, and at the same time, by training a feature mining model, it can quickly and accurately mine collaborative features based on the integrated learning data generated by fusion, and identify target features that have a negative impact on learners' learning outcomes in the collaborative features, thereby achieving a combined effect of rapid screening, identification, and precise mining of complex data, and can also perform multi-dimensional, multi-type, and multi-state collaborative management operations on learners' learning situations, avoiding the inefficiency and error-prone problems in the analysis and evaluation of multi-dimensional data and multi-type data, thereby meeting the needs of learners for rapid and accurate analysis and evaluation of their learning outcomes.

[0171] Example 2: Please refer to Figure 2 and Figure 3 As shown, the part not described in detail in this embodiment is described in the first embodiment, and a multi-source heterogeneous and dynamic behavior collaborative feature mining system is provided, which is applied to a mining server and is used to implement a multi-source heterogeneous and dynamic behavior collaborative feature mining method, including a cycle calculation module, a first acquisition module, a second acquisition module, a model training module, a prompt determination module and a collaborative management module, wherein each module is connected via a wired or wireless network;

[0172] The cycle calculation module is used to query the basic learning parameters of the learner, extract the cycle parameters from the basic learning parameters, the cycle parameters include the effective learning peak value, the learning interruption value and the learning fatigue, and calculate the learner's learning cycle;

[0173] The first acquisition module is used to obtain the first learning data of the learner in the learning cycle, and extract multi-source heterogeneous data from the first learning data based on a first extraction criterion, wherein the first extraction criterion is: eliminating redundant and repeated first learning data;

[0174] The second acquisition module is used to obtain the second learning data of the learner in the learning cycle, and extract the dynamic behavior data from the second learning data based on the second extraction criterion, and the second extraction criterion is: eliminate the data to be verified whose recording time is not located in the last three digits;

[0175] Model training module, used to fuse multi-source heterogeneous data and dynamic behavior data into comprehensive learning data, mine collaborative features in the comprehensive learning data, and train feature mining models that mine collaborative features;

[0176] The prompt determination module is used to obtain real-time comprehensive learning data, input it into the feature mining model to mine real-time collaborative features, and determine whether to issue collaborative management prompts;

[0177] The collaborative management module is used to identify target features from collaborative features, mark collaborative management data according to the target features, and integrate the collaborative features and collaborative management data for display to learners and managers.

[0178] Example 3: Please refer to Figure 4 As shown, this embodiment discloses an electronic device, including a processor and a memory;

[0179] Wherein, the memory stores a computer program that can be called by the processor;

[0180] The processor executes the multi-source heterogeneous and dynamic behavior collaborative feature mining method by calling the computer program stored in the memory.

[0181] Since the electronic device introduced in this embodiment is an electronic device used to implement the multi-source heterogeneous and dynamic behavior collaborative feature mining method in Example 1 of this application, based on the multi-source heterogeneous and dynamic behavior collaborative feature mining method introduced in the embodiment of this application, the technical personnel in this field can understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application will not be introduced in detail here. As long as the technical personnel in this field implement the electronic device used in the multi-source heterogeneous and dynamic behavior collaborative feature mining method in the embodiment of this application, it belongs to the scope of protection of this application.

[0182] Example 4: Please refer to Figure 5 As shown, this embodiment discloses a computer-readable storage medium on which a rewritable computer program is stored;

[0183] When the computer program is executed, the multi-source heterogeneous and dynamic behavior collaborative feature mining method is implemented.

[0184] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A multi-source heterogeneous and dynamic behavior collaborative feature mining method, applied to mining servers, characterized by: include: S1: query the learner's basic learning parameters, extract the cycle parameters from the basic learning parameters, the cycle parameters include the effective learning peak value, the learning interruption value and the learning fatigue, and calculate the learner's learning cycle; Methods for obtaining learning discontinuity values ​​include: Through the learning management system, we can find out the learners’ learning events, and query them one by one The first recorded moment and the last recorded moment in a learning event are recorded as the starting moment and the ending moment respectively; The duration between the end time of the previous learning event and the start time of the next learning event is recorded as the interval time, and the The length of the interval; Remove the maximum and minimum interval durations and keep the remaining The duration of each interval is accumulated and averaged to obtain the learning interruption value; The methods for obtaining learning fatigue include: Query the learners one by one The event state attributes of a learning event are obtained, and the attribute values ​​in the event state attributes are marked; The learning events with attribute values ​​of "0-0", "0-1" or "1-0" are recorded as fatigue events, and the duration of fatigue events and learning events are queried one by one through timestamps to obtain Fatigue duration and The duration of the event; Will After the fatigue durations are accumulated one by one, the total fatigue value is obtained, and the total fatigue value is combined with Compare the accumulated values ​​of the duration of each event to obtain the degree of learning fatigue; The learning cycle calculation method includes: After assigning different weight factors to the effective learning peak value, learning interruption value and learning fatigue, and comparing them, the cycle added value is calculated; The learning cycle is calculated by adding the preset standard cycle to the cycle additional value; S2: obtaining first learning data of the learner in the learning cycle, and extracting multi-source heterogeneous data from the first learning data based on a first extraction criterion, wherein the first extraction criterion is: eliminating redundant and repeated first learning data; S3: obtaining the second learning data of the learner in the learning cycle, and extracting the dynamic behavior data from the second learning data based on the second extraction criterion, wherein the second extraction criterion is: eliminating the data to be verified whose recording time is not in the last three digits; S4: Integrate multi-source heterogeneous data and dynamic behavior data into comprehensive learning data, mine collaborative features in the comprehensive learning data, and train a feature mining model that mines collaborative features; S5: Obtain real-time comprehensive learning data, input it into the feature mining model to mine real-time collaborative features, and determine whether to issue collaborative management prompts; S6: If a collaborative management prompt is issued, the target feature is identified from the collaborative features, the collaborative management data is marked according to the target feature, and the collaborative features and collaborative management data are integrated for display to the learner and the manager.

2. The multi-source heterogeneous and dynamic behavior collaborative feature mining method according to claim 1 is characterized in that: The expression for the learning discontinuity value is: ; In the formula, is the learning discontinuity value, For the The interval duration.

3. The multi-source heterogeneous and dynamic behavior collaborative feature mining method according to claim 2 is characterized in that: The expression of learning fatigue is: ; In the formula, To learn fatigue, For the Fatigue duration, For the The duration of an event.

4. The multi-source heterogeneous and dynamic behavior collaborative feature mining method according to claim 3 is characterized in that: The expression of periodic additional value is: ; In the formula, Add value to the cycle, To effectively learn the peak value, , , All are weight factors greater than 0; The expression of the learning cycle is: ; In the formula, For the learning cycle, The default standard cycle.

5. The multi-source heterogeneous and dynamic behavior collaborative feature mining method according to claim 4 is characterized in that: Multi-source heterogeneous data include test score data, daily performance data, extracurricular activity data, book review data, and hands-on practice data; Methods for extracting test score data, daily performance data, extracurricular activity data, book review data, and hands-on practice data include: Repeating and comparing all first learning data, marking duplicate first learning data, and searching for recording times of duplicate first learning data; Arrange the duplicate first learning data in sequence according to the order of recording time, and remove the duplicate first learning data except the last one; Using natural language processing technology, the first semantics in the remaining first learning data are identified one by one, and the first semantics are split into a meaning part and a value part; The first semantics whose meanings are grades, daily life, extracurricular activities, books and practical operations are recorded as test semantics, daily semantics, extracurricular semantics, book semantics and practical operation semantics respectively, and the first learning data corresponding to the test semantics, daily semantics, extracurricular semantics, book semantics and practical operation semantics are recorded as test grade data, daily performance data, extracurricular activity data, book reference data and hands-on operation data.

6. The multi-source heterogeneous and dynamic behavior collaborative feature mining method according to claim 5 is characterized in that: Dynamic behavior data include test answering time, number of class notes, frequency of interactive discussions, number of searches and borrowing times, and effective practice time; The extraction methods of test answering time, number of class notes, frequency of interactive discussion, number of search and borrowing times, and effective practice time include: Identify the second semantics of the second learning data one by one through natural language processing technology, and split the second semantics into a text part and a numerical part; The textual parts of the second learning data, which are answer sheets, notes, discussions, borrowings and valid ones, are respectively recorded as the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified, and the recording times of the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified are queried one by one through the timestamps; Arrange the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified in sequence according to the order of recording time, and remove the first data to be verified, the second data to be verified, the third data to be verified, the fourth data to be verified and the fifth data to be verified except the last three digits; Add up the numerical parts of the remaining first data to be verified, second data to be verified, third data to be verified, fourth data to be verified and fifth data to be verified and average them to obtain the test answering time, number of class notes, frequency of interactive discussions, number of searches and borrowings and effective practice time respectively.

7. The multi-source heterogeneous and dynamic behavior collaborative feature mining method according to claim 6 is characterized in that: The training methods of feature mining models include: Collect multiple sets of comprehensive learning data and corresponding collaborative features, and convert each set of comprehensive learning data into a corresponding set of feature vectors; The feature vector is used as the input of the machine learning model, and the collaborative features corresponding to each group of comprehensive learning data are used as the output of the machine learning model. The collaborative features actually corresponding to each group of comprehensive learning data are used as prediction targets, and minimizing the sum of prediction errors of the comprehensive learning data is used as the training target. The machine learning model is trained until the sum of prediction errors converges and the training is stopped to obtain a feature mining model that mines collaborative features.

8. The multi-source heterogeneous and dynamic behavior collaborative feature mining method according to claim 7 is characterized in that: The determination method of whether to issue a collaborative management prompt includes: Count the number of collaborative features mined by the feature mining model and record it as the feature value; When the characteristic value is less than or equal to 3, it is determined that a collaborative management prompt is issued; When the characteristic value is greater than 3, it is determined that no collaborative management prompt is issued.

9. The multi-source heterogeneous and dynamic behavior collaborative feature mining method according to claim 8 is characterized in that: Tagging methods for collaboratively managing data include: Arrange all collaborative features in order according to the mining sequence, and identify the semantic features of all collaborative features one by one through natural language processing technology; When the feature semantics is exam, the exam score data and the exam answer time are recorded as collaborative management data; When the feature semantics is notes, the daily performance data and the number of class notes are recorded as collaborative management data; When the feature semantics is interactive, the extracurricular activity data and the frequency of interactive discussions are recorded as collaborative management data; When the feature semantics is reading, the book lookup data and search borrowing times are recorded as collaborative management data; When the feature semantics is operation, the hands-on operation data and effective practice duration are recorded as collaborative management data.

10. A multi-source heterogeneous and dynamic behavior collaborative feature mining system, applied to a mining server, for implementing the multi-source heterogeneous and dynamic behavior collaborative feature mining method according to any one of claims 1 to 9, characterized in that: It includes a cycle calculation module, a first acquisition module, a second acquisition module, a model training module, a prompt determination module and a collaborative management module, wherein each module is connected via a wired or wireless network; The cycle calculation module is used to query the basic learning parameters of the learner, extract the cycle parameters from the basic learning parameters, the cycle parameters include the effective learning peak value, the learning interruption value and the learning fatigue, and calculate the learner's learning cycle; The first acquisition module is used to obtain the first learning data of the learner in the learning cycle, and extract multi-source heterogeneous data from the first learning data based on a first extraction criterion, wherein the first extraction criterion is: eliminating redundant and repeated first learning data; The second acquisition module is used to obtain the second learning data of the learner in the learning cycle, and extract the dynamic behavior data from the second learning data based on the second extraction criterion, and the second extraction criterion is: eliminate the data to be verified whose recording time is not in the last three digits; Model training module, used to fuse multi-source heterogeneous data and dynamic behavior data into comprehensive learning data, mine collaborative features in the comprehensive learning data, and train feature mining models that mine collaborative features; The prompt determination module is used to obtain real-time comprehensive learning data, input it into the feature mining model to mine real-time collaborative features, and determine whether to issue collaborative management prompts; The collaborative management module is used to identify target features from collaborative features, mark collaborative management data according to the target features, and integrate the collaborative features and collaborative management data for display to learners and managers.

Citation Information

Patent Citations

  • Cross-data, information and knowledge multi-modal feature mining method and component

    CN112214531A