Multi-dimensional data-based capability identification and adaptive content recommendation method and system

By using multi-dimensional data collection and combined machine learning models, the problem of incomplete user capability assessment in existing technologies has been solved, enabling low-latency adaptive content recommendation, improving the accuracy of user profiles and interaction efficiency, and discovering and cultivating user strengths.

CN121858733APending Publication Date: 2026-04-14HOT WHEELS FIREFLY (SHANGHAI) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies have limited dimensions for user capability assessment, and the analysis models are inefficient and lack precision, resulting in high latency and poor adaptability in content recommendation. They are unable to build complete user profiles and struggle to handle multi-dimensional, heterogeneous, high-dimensional, and noisy user behavior data.

Method used

By collecting multi-dimensional behavioral data in real time, a combined machine learning model is used for data normalization, trend analysis, user profile clustering, and ability feature classification, including multiple linear regression, K-means clustering, and random forest classification models. Combined with computer vision technology, user engagement is quantified to build a low-latency closed-loop adaptive content recommendation system.

Benefits of technology

It achieves accurate identification and personalized recommendations of user capabilities, reduces system response latency, improves human-computer interaction efficiency and adaptability, can discover users' potential strengths and provide personalized training content, and solves the problem of incomplete evaluation in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858733A_ABST
    Figure CN121858733A_ABST
Patent Text Reader

Abstract

The invention provides a capability identification and recommendation method and system based on multi-dimensional data, and belongs to the technical field of artificial intelligence and data processing. The invention aims to solve the problems of inaccurate and incomplete user capability analysis, high content recommendation delay and poor adaptability caused by single evaluation dimension and insufficient efficiency and precision of an analysis model. The method comprises the steps of collecting and quantifying multi-dimensional behavior data streams of a user in real time through a sensor, wherein the multi-dimensional behavior data streams comprise task execution efficiency, quality and input degree data; a combined machine learning model sequentially comprising a trend analysis model, a user portrait clustering model and a capability feature classification model is adopted to process the data flow so as to identify user capability features; recommending content based on the identified features; and recording user behavior feedback as incremental data in real time, and feeding back the incremental data to the combined model to dynamically adjust the model and recommend. According to the method, the accuracy of capability analysis is improved, and low-delay closed-loop adaptive recommendation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and data processing technology, and in particular to a method and system for learning ability feature recognition and adaptive content recommendation based on multi-dimensional data fusion analysis. Background Technology

[0002] In adaptive learning systems, dynamically adjusting instructional content based on user behavioral feedback is key to improving the effectiveness of personalized education. Current technologies typically recommend subsequent content by analyzing user task performance. For example, some systems judge a user's mastery of knowledge points solely based on outcome metrics such as the accuracy of their answers. More advanced systems are incorporating computer vision technology, analyzing information such as facial expressions to determine a user's emotions or attention levels, and adjusting teaching strategies accordingly, thus creating a closed-loop adjustment system based on real-time user feedback. However, these existing technological solutions still have significant drawbacks: First, the limited evaluation dimensions lead to inaccurate user profiles. Relying solely on outcome-based metrics or single process indicators cannot construct a complete profile of a user's capabilities. For example, a user who completes a task quickly and accurately differs significantly in cognitive efficiency and proficiency from a user who completes the task slowly but equally accurately. However, existing systems cannot effectively distinguish between these differences, resulting in low matching accuracy for subsequent content recommendations.

[0003] Secondly, the analysis models are simple, making it difficult to balance efficiency and accuracy. When processing multi-dimensional user behavior data, existing systems often use simple rule engines or single machine learning models. These models struggle to maintain high computational efficiency while ensuring analytical accuracy when faced with complex, high-dimensional, and noisy data streams composed of heterogeneous data such as efficiency, quality, and engagement. This results in less robust and accurate identification of user capabilities.

[0004] Finally, the feedback loop suffers from high latency and poor adaptability. Many recommendation logics are based on offline or batch processing analysis results, which cannot respond quickly to changes in the user's real-time status. When a user's abilities or interests change during a learning task, the system's recommended content is not updated in a timely manner, reducing interaction efficiency and user experience.

[0005] Therefore, there is a market need for a method and system based on multidimensional data capability identification and adaptive content recommendation that can effectively integrate multidimensional data, analyze it through an efficient and accurate combination model, and achieve low-latency closed-loop adaptive content recommendation. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for capability identification and adaptive content recommendation based on multidimensional data. This solves the problems of inaccurate and incomplete user capability analysis caused by the single evaluation dimension and insufficient efficiency and accuracy of the analysis model in existing technologies, as well as the resulting technical difficulties of high content recommendation latency and poor adaptability. A learning ability feature recognition and adaptive content recommendation method based on multi-dimensional data fusion analysis, provided by the present invention, includes the following steps: The system collects and quantifies multi-dimensional behavioral data streams of users in real time when performing tasks using sensors. These multi-dimensional behavioral data streams include at least: task execution efficiency data, task execution quality data, and user task engagement data obtained by processing video streams captured by cameras using computer vision technology. A combined machine learning model is used to process the multi-dimensional behavioral data stream to identify user capability characteristics. The combined machine learning model includes: a trend analysis model for data normalization and trend analysis of the multi-dimensional behavioral data stream; a user profile clustering model for clustering user profiles based on the output of the trend analysis model; and a capability characteristic classification model for classifying user capability characteristics by combining the clustering results of the user profile clustering model and the original multi-dimensional behavioral data stream. Based on the output of the capability feature classification model, content is matched in the content library and recommended to the user.

[0007] Preferably, the task execution efficiency data includes the task processing speed obtained by calculating the time consumed by the user to complete the task; the task execution quality data includes the task accuracy rate obtained by comparing the user's output with the standard answer.

[0008] Preferably, the step of obtaining the task engagement data includes: By using facial landmark detection and head pose estimation algorithms, the deviation angle and duration of the user's gaze direction from the task area are calculated, and a quantitative score of the task engagement data is generated based on the deviation angle and duration.

[0009] Preferably, the quantification of the task engagement data also incorporates eye-tracking analysis of the user's gaze point distribution and saccade path.

[0010] Preferably, the trend analysis model is a multiple linear regression model; The user profile clustering model is either a K-means clustering model or a DBSCAN model; The capability feature classification model is a random forest classification model.

[0011] Preferably, it further includes: The system records user behavior feedback data in real time to the recommended content and feeds this feedback data as incremental data into the combined machine learning model to dynamically adjust the combined machine learning model and subsequent recommended content.

[0012] A learning ability feature recognition and adaptive content recommendation system based on multi-dimensional data fusion analysis, provided by the present invention, includes: Multi-dimensional behavioral data stream acquisition module: Real-time acquisition and quantification of multi-dimensional behavioral data streams of users when performing tasks through sensors. The multi-dimensional behavioral data streams include at least: task execution efficiency data, task execution quality data, and user task engagement data obtained by processing video streams captured by cameras through computer vision technology. Multi-dimensional behavioral data stream analysis module: Employs a combined machine learning model to process the multi-dimensional behavioral data stream to identify user capability characteristics. The combined machine learning model includes: a trend analysis model for data normalization and trend analysis of the multi-dimensional behavioral data stream; a user profile clustering model for user profile clustering based on the output of the trend analysis model; and a capability characteristic classification model for classifying user capability characteristics by combining the clustering results of the user profile clustering model and the original multi-dimensional behavioral data stream. Recommendation content generation module: Based on the output of the capability feature classification model, it matches content in the content library and recommends content to the user.

[0013] Preferably, the task execution efficiency data includes the task processing speed obtained by calculating the time consumed by the user to complete the task; The task execution quality data includes the task accuracy rate obtained by comparing user output with the standard answer; The task engagement data includes calculating the deviation angle and duration between the user's gaze direction and the task area using facial key point detection and head pose estimation algorithms, and generating a quantitative score for the task engagement data based on the deviation angle and duration. The trend analysis model is a multiple linear regression model; the user profile clustering model is a K-means clustering model or a DBSCAN model; and the capability feature classification model is a random forest classification model.

[0014] Preferably, the quantification of the task engagement data also incorporates eye-tracking analysis of the user's gaze point distribution and saccade path.

[0015] The present invention also provides an intelligent learning system, comprising: a data acquisition layer, an ability feature recognition and recommendation module, an application layer, and a data feedback link; The data acquisition layer includes a face / behavior recognition module and a textbook / homework recognition module; the face / behavior recognition module is used to process the video stream acquired by the infrared camera to identify the user's face and posture; the textbook / homework recognition module is used to process the images acquired by the high-definition camera to identify the task content and the user's answer. The capability feature recognition and recommendation module includes the learning capability feature recognition and adaptive content recommendation system based on multi-dimensional data fusion analysis; The application layer includes a content library and a user interface module; the content library pre-stores diverse task content; the user interface module is responsible for generating the interactive interface and presenting information including recommended content and analysis reports on the projection display area through a projection optical engine; The data feedback link transmits the behavioral feedback data generated by the user at the application layer back to the data acquisition layer as incremental data to participate in the next round of analysis.

[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention integrates data from three dimensions—efficiency, quality, and engagement—in task execution and uses a specific combined model for analysis. This invention overcomes the limitations of single-dimensional data analysis, enabling more accurate and robust identification of users' true capabilities and improving the accuracy and objectivity of user capability analysis.

[0017] 2. This invention constructs a complete real-time closed-loop system that can quickly adjust recommendation strategies based on the user's latest behavior, significantly reducing system response latency and achieving low-latency closed-loop adaptive recommendation, thus improving the efficiency and adaptability of human-computer interaction. The combined model architecture adopted, through layer-by-layer filtering and dimensionality reduction, effectively solves the technical problems of high computational load and susceptibility to noise interference when directly processing multidimensional sparse data, improving computational efficiency while ensuring accuracy.

[0018] 3. This invention proposes a specific technical approach to quantify user behavior captured by a camera into a focus index, transforming unstructured visual information into structured data that can be used for machine learning, thereby achieving effective utilization of process visual data and improving the depth and breadth of data utilization.

[0019] 4. This invention can also collect objective data from multiple dimensions such as students' problem-solving speed, accuracy, and concentration during the learning process, and use a combined machine learning model to conduct comprehensive analysis to discover students' potential talents and strengths. It can also automatically recommend personalized and appropriately challenging advanced training content based on the generated talent report, realizing a functional leap from the traditional "remedial" to "enrichment" and solving the pain point that parents or teachers are unable to effectively cultivate children's talents due to a lack of professional ability. Attached Figure Description

[0020] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic diagram of the structure of a smart device provided in an embodiment of the present invention; Figure 2 This is a system functional module architecture diagram provided for an embodiment of the present invention; Figure 3 A flowchart of the method provided in an embodiment of the present invention; Figure 4 This is a timing diagram of closed-loop feedback signaling interaction provided in an embodiment of the present invention. Detailed Implementation

[0021] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0022] Example 1 This embodiment provides a complete implementation scheme for learning ability feature recognition and adaptive content recommendation based on multi-dimensional data fusion analysis. The scheme uses an intelligent hardware device integrating multi-modal sensors to collect and process multi-dimensional behavioral data of users in learning tasks in real time. Then, it utilizes a specific combinatorial machine learning model for deep analysis and, based on the analysis results, provides low-latency adaptive content recommendations to users, ultimately forming a dynamically adjusted closed-loop control system.

[0023] Please see Figure 1 The figure is a schematic diagram of the structure of a smart device according to an embodiment of the present invention. In this embodiment, the method can be executed by a smart device body. The smart device body can be a desktop smart terminal designed specifically for learning scenarios, which integrates a high-performance processor, memory, and various sensors and output devices. As a specific physical implementation, the smart device body can be installed on a bookshelf or stand above a desk, with its camera module and projector facing the user's desktop work area.

[0024] The camera module is a key component for multi-dimensional data acquisition, and it can contain multiple cameras with different functions. For example, the module may include an infrared camera for capturing the user's face and posture, with performance such as 1080P resolution and 30fps frame rate, capable of stable operation under various lighting conditions; simultaneously, the module may also include a high-definition camera for recognizing task content (such as books or notebooks) on the desktop work area, with performance such as 4K resolution and 20fps frame rate, and support for autofocus to ensure clear recognition of questions and the user's written answers. The projection engine, such as an ultra-short-throw projector using digital light processing technology, functions to project system-generated interactive interfaces, recommended content, and other information onto the desktop work area, thereby forming an interactive projection display area.

[0025] Please refer to the following: Figure 2 This figure is a system functional module architecture diagram provided in an embodiment of the present invention. At the software level, the system architecture of this embodiment mainly includes a data acquisition layer, a capability feature recognition and recommendation module, and an application layer.

[0026] The data acquisition layer is responsible for acquiring raw data streams from hardware such as camera modules and performing preliminary quantization processing. This layer may include a face / behavior recognition module and a textbook / homework recognition module. The face / behavior recognition module processes the video stream acquired by the infrared camera to identify information such as faces, postures, and gazes; the textbook / homework recognition module processes images acquired by the high-definition camera, using technologies such as optical character recognition to identify task content and user responses.

[0027] The ability feature recognition and recommendation module is the core processing unit of this solution, and it internally incorporates the combined machine learning model proposed in this invention. This module, following a specific data processing order, includes a multiple linear regression module, a K-means clustering module, and a random forest classification module. These three modules work together to achieve accurate recognition of user ability features.

[0028] The application layer is responsible for interacting with the user. It contains a content library that pre-stores a large amount of learning content across different subjects, difficulties, and formats. Additionally, this layer includes a user interface module responsible for generating the interactive interface and presenting recommended content and analysis reports on the projection display area via a projection engine.

[0029] In particular, such as Figure 2 As shown, the system also includes a crucial data feedback link. This link retransmits user behavioral feedback data generated at the application layer (e.g., selection of recommended content, performance in completing new tasks, etc.) back to the data acquisition layer, where it serves as incremental data for the next round of analysis, thus forming a complete closed-loop control system.

[0030] The following will combine Figure 3 The method flowchart shown and Figure 4 The closed-loop feedback signaling interaction timing diagram shown below provides a detailed description of the working process of this embodiment. Assume a user is using the smart device of this embodiment to complete a mathematical calculation task.

[0031] The process begins with step S101: collecting and quantifying multi-dimensional behavioral data. This occurs when a user begins performing a task on their desktop workspace (corresponding to...). Figure 4 (Signaling 1) activates the system. A high-definition camera continuously captures the desktop work area, and the textbook / homework recognition module records the start and end timestamps of each question completed by the user. The time consumed to complete each question is calculated by the difference, and this time is quantified as task execution efficiency data, specifically representing task processing speed. Simultaneously, the module recognizes the user's written answers and compares them with pre-stored or real-time parsed standard answers, calculating the proportion of correct answers. This proportion is quantified as task execution quality data, specifically representing task accuracy. These two quantification methods provide basic performance indicators for subsequent analysis, supporting the specific limitations on efficiency and quality data in the dependent claims.

[0032] Meanwhile, an infrared camera continuously captures the user's facial video stream. The face / behavior recognition module runs in real time on the device's built-in AI coprocessor (a hardware unit dedicated to accelerating neural network operations). This module first uses a facial landmark detection algorithm to accurately pinpoint the positions of key points on the user's face, such as the eyes, nose, mouth, and facial contours. Subsequently, based on these key points, a head pose estimation algorithm calculates in real time the user's head rotation vector in three-dimensional space, namely the pitch angle, yaw angle, and roll angle. The system predefines the desktop work area as the standard task area, and the user's central gaze direction can be calculated from the head rotation vector. The system then calculates the deviation angle of this gaze direction from the center of the standard task area and the duration of the gaze deviation. For example, if the deviation angle lasts for more than 30 degrees for more than 3 seconds, the system determines that the user is distracted.

[0033] As a preferred implementation, to further improve the accuracy of engagement measurement, the face / behavior recognition module can also incorporate eye-tracking technology. By analyzing changes in pupil position, the distribution of the user's gaze points and saccade paths on the desktop work area can be accurately obtained. If the user's gaze points remain stable on the current task title or draft area for a long time, their engagement is considered high; conversely, if the gaze points frequently jump to areas unrelated to the task, or the saccade paths are chaotic, their engagement is considered low. The system integrates the above-mentioned multi-source information such as deviation angle, deviation duration, and gaze point distribution, and uses a weighted fusion algorithm to finally generate a standardized task engagement score that continuously varies within the interval [0, 1]. This process transforms the abstract concept of focus into a concrete, calculable quantitative indicator, namely, task engagement data.

[0034] The data for the above three dimensions (task execution efficiency, task execution quality, and task engagement) are sent in real time in time series format to the capability feature recognition and recommendation module for processing (corresponding to...). Figure 4 Signaling 2 and signaling 3 in the middle.

[0035] Next, step S102 is executed: trend analysis is performed using a multiple linear regression model. Upon receiving the three-dimensional data stream, the multiple linear regression module within the ability feature recognition and recommendation module first processes this data. It is understandable that the data in these three dimensions may experience significant fluctuations and noise in the short term. For example, a user's processing speed may decrease due to a brief pause in thought, or the accuracy of a single task may decrease due to an accidental typo. Directly using this noisy data for analysis could lead to model misjudgment. Therefore, the role of the multiple linear regression module is to fit these time-series data to filter out short-term random fluctuations and extract long-term trends reflecting the user's abilities over a period of time (e.g., the last 5 minutes or the last 10 tasks). The model outputs one or more feature vectors that represent the user's recent overall performance, such as the slope of the fitted line (representing the trend of ability change) and the intercept (representing the baseline level of ability), providing a more stable and representative input for subsequent processing.

[0036] Subsequently, step S103 is executed: user profile clustering is performed using the K-means clustering model. The stable feature vector output from the multiple linear regression module is passed to the K-means clustering module. The function of this module is to pre-classify user profiles. In this embodiment, the system presets several typical user profile categories, such as high-speed, high-quality, high-investment type, low-speed, high-quality, high-investment type, high-speed, low-quality, high-investment type (careless type), low-speed, low-quality, low-investment type (difficult type), etc. The K-means clustering module calculates the distance between the input feature vector and each preset cluster center, and assigns the user to the nearest cluster. For example, if a user's feature vector shows fast processing speed, high accuracy, and high investment score, it will be assigned to the high-speed, high-quality, high-investment user profile cluster. The value of this step is that it discretizes the continuous multidimensional feature space into a finite number of categories, reducing the search space of the subsequent accurate classification model, thereby significantly improving the computational efficiency of the entire analysis process.

[0037] Next, step S104 is executed: A random forest classification model is used for precise classification of ability features. The clustering results (i.e., user profile labels) output by the K-means clustering module, along with the raw multi-dimensional behavioral data stream from the data acquisition layer, are input into the random forest classification module. The random forest model consists of a large number of decision trees, which has a natural advantage in handling high-dimensional data and non-linear relationships, and is less prone to overfitting and has strong robustness. The task of this module is to perform more refined ability feature classification. Specifically, the system's content library predefines more specific ability labels, such as proficiency in integer addition and subtraction, mastery of decimal multiplication and division, strong high-order logic operation ability, and outstanding spatial imagination ability. The random forest classification module utilizes the input clustering labels (as a strong feature) and raw, richer detailed data (such as the speed, accuracy, and engagement fluctuations of each task) to output one or more labels that best describe the user's current specific ability through ensemble learning. For example, if a user who was previously clustered as high-speed, high-quality, and high-investment has particularly outstanding performance in handling logical reasoning questions, the random forest model may ultimately classify them as having strong high-order logic operation capabilities.

[0038] Next, step S105 is executed: generating a report and recommending content. The output of the random forest classification module (i.e., capability feature labels) is sent to the application layer (corresponding to...). Figure 4 (Signaling 4 in the code). The application layer first generates a structured capability feature data report based on the result. This report can be stored and used for long-term user capability tracking. Meanwhile, the recommendation module (corresponding to...) Figure 4Signaling 5) uses the capability feature tag as an index to search and match within the content library. For example, for users identified as having strong high-order logic operation capabilities, the system will match them with more challenging logic problems, related Olympiad math knowledge points, or extended fun math games. The matched content is then presented on the projection display area via a projector for the user to select and learn (corresponding to...). Figure 4 Signaling 6 in the middle.

[0039] Finally, step S106 is executed: user feedback is recorded and returned as incremental data. This step is crucial for achieving closed-loop dynamic adjustment. When a user begins to interact with the recommended content (corresponding to...) Figure 4 In signaling 7), for example, if a user selects a more challenging logic puzzle and begins solving it, the system will use the data feedback loop to feed this series of new user behavior data (including efficiency, quality, and engagement in completing the new task) as incremental data back to the data acquisition layer and the ability feature recognition and recommendation module (corresponding to...) in real time. Figure 4 (Signaling 8 and Signaling 9 in the code). This new data immediately participates in the next round of model calculations, used to dynamically adjust and update model parameters. For example, if a user easily completes a challenge, the model will further confirm their higher-order abilities and recommend more difficult content; if the user struggles in the challenge (e.g., slower speed, decreased engagement), the model will correspondingly lower the difficulty and recommend some basic concept reinforcement exercises. This process continuously loops, forming a low-latency closed-loop feedback system, enabling the recommendation strategy to adapt to changes in the user's abilities in real time and dynamically.

[0040] This embodiment, through the complete hardware and software integration solution described above, can accurately identify the user's ability characteristics in specific tasks and provide personalized, adaptive advanced or reinforcement content in real time. At the same time, it continuously optimizes the recommendation strategy based on the user's latest performance, thereby achieving precise guidance and efficient support for the user's learning process.

[0041] Understandably, this invention can also utilize artificial intelligence to analyze students' long-term, multi-dimensional objective learning data, enabling a more scientific and automated discovery of their potential talents and strengths, thus overcoming the limitations and biases of subjective judgment. The identification and recommendation methods of this invention not only discover talents but also automatically recommend specialized, advanced content based on the talent report to cultivate those talents, achieving a shift from "remedial" to "elite" development, addressing the pain point of not being able to effectively cultivate children's talents due to a lack of professional skills. The generated detailed talent report provides parents and students with a scientific and visualized ability development map, helping them to make more informed plans for future academic and even career development.

[0042] Example 2 This embodiment provides a variation of Embodiment 1, the main difference being the method of quantifying task engagement data. This solution aims to reduce the requirements for hardware computing power, making the technical solution of this invention applicable to devices that do not possess high-performance artificial intelligence coprocessors or dedicated eye-tracking hardware.

[0043] In this embodiment, the system's hardware configuration can be simplified; for example, it may not include a dedicated AI coprocessor. Correspondingly, the face / behavior recognition module in the software-level data acquisition layer employs a less computationally expensive attention calculation algorithm.

[0044] The specific working process is as follows: In the data quantization stage of step S101, after the system acquires the user's facial video stream through the camera module, the face / behavior recognition module calls a standard, computationally inexpensive computer vision library (such as the open-source OpenCV library) to run a head pose estimation algorithm. This algorithm can also analyze the user's head rotation vector (pitch angle, yaw angle, roll angle) in three-dimensional space and estimate the user's gaze direction accordingly.

[0045] Unlike Embodiment 1, this embodiment does not perform complex eye-tracking analysis. The system determines the user's engagement level solely based on the deviation angle between the gaze direction estimated from the head posture and the preset center of the desktop work area, as well as the duration of this deviation. The system can preset a deviation angle threshold (e.g., 30 degrees) and a duration threshold (e.g., 3 seconds). When the algorithm detects that the user's gaze deviation angle continuously exceeds 30 degrees for 3 seconds, it determines that the user has experienced a distraction event and correspondingly reduces their task engagement score. Conversely, when the user's gaze direction remains within the task area for a prolonged period, the system increases their task engagement score.

[0046] It should be noted that although this method sacrifices the higher accuracy that eye tracking can provide (e.g., it cannot distinguish whether the user is looking at a question or daydreaming), it greatly reduces the computational complexity of the algorithm, allowing the function to run smoothly using only the central processing unit without dedicated hardware acceleration.

[0047] The task engagement data quantified by this simplified scheme can also serve as a crucial third dimension of data, independent of efficiency and quality, and can be input into the subsequent capability feature identification and recommendation modules. The processing flow of the subsequent multiple linear regression, clustering, and classification modules is exactly the same as in Example 1.

[0048] The intended effect of this embodiment is to effectively quantify user task engagement at a lower cost on devices with limited computing resources. This ensures the feasibility of the core technical concept of this invention (i.e., analyzing data integrating efficiency, quality, and engagement), broadens the scope of application of this invention, and enables its deployment on a wider range of smart devices.

[0049] Example 3 This embodiment provides another variant implementation of the combined model in Embodiment 1, specifically, a replacement for the user profile clustering model. This embodiment aims to demonstrate how different clustering algorithms can be used to support the segmentation of more complex user groups.

[0050] In this embodiment, the overall architecture and process of the system are basically the same as in Embodiment 1. The main changes occur within the capability feature recognition and recommendation module. Specifically, the original K-means clustering module is replaced with a density-based noise application space clustering model.

[0051] The specific working process is as follows: In the model analysis process, after the multiple linear regression module in step S102 processes the original data and outputs stable feature vectors, these feature vectors are passed to the density-based noise application space clustering model for user profile clustering (corresponding to the modified step S103).

[0052] Unlike the K-means algorithm, which requires a pre-specified number of clusters (i.e., the K value), this density-based clustering algorithm does not require a pre-defined number of clusters. By defining two parameters—neighborhood radius and minimum number of neighboring data points—it automatically discovers clusters of arbitrary shapes based on the density of data points in the feature space, and can effectively identify outliers (i.e., noise points) that do not belong to any cluster.

[0053] For example, when processing large-scale, diverse user data, in addition to the typical user profiles mentioned in Example 1, there may be some atypical user groups, such as users who are high-speed and high-quality but whose engagement is extremely unstable. Their behavioral patterns may form an irregular, elongated cluster in the feature space. The K-means algorithm, due to its cluster center-based partitioning principle, struggles to accurately identify such non-spherical clusters and may incorrectly segment or merge them with other clusters. However, the density-based clustering algorithm used in this example can effectively identify these density-connected clusters of arbitrary shapes.

[0054] More importantly, the algorithm can identify outliers. For example, if a user's behavioral data is far from any high-density region in the feature space, this may represent a very rare or anomalous learning pattern. The model will label this user as an outlier, rather than forcibly assigning them to the nearest cluster as in K-means. This outlier label itself is valuable information, allowing the system to trigger special strategies, such as alerting a human tutor or recommending diagnostic tasks to further explore the user's specific situation.

[0055] After the density-based clustering model completes clustering, its output cluster labels (or outlier labels) along with the original data are sent to the random forest classification module in step S104 for subsequent accurate classification.

[0056] This embodiment improves the flexibility and accuracy of user profile clustering by replacing the K-means model with a density-based noise-applied spatial clustering model. It enables the system to automatically discover potential user groups of arbitrary shapes within the data and effectively identify and handle abnormal users, providing more refined and informative input for subsequent classification and recommendation steps, thereby supporting more targeted personalized teaching strategies. This fully supports the dependent claims that limit the user profile clustering model to a K-means clustering model or a density-based noise-applied spatial clustering model.

[0057] Example 4 This embodiment aims to demonstrate the broad applicability of the technical framework proposed in this invention. By applying it to the scenario of music and art learning, which is quite different from traditional disciplines, it proves the scalability of the core idea of ​​this invention.

[0058] In this embodiment, the system is used to assist users in piano practice. On the hardware side, the system can connect to an external electronic keyboard supporting the MIDI protocol via a standard interface (such as USB) as the primary data input device. On the software side, the ability feature recognition and recommendation module loads special model parameters and ability tagging systems for music discipline training, and the application layer's content library stores a large amount of music learning-related materials such as sheet music, rhythm exercises, scale exercises, harmony knowledge points, and sight-singing and ear training.

[0059] The specific working process is as follows: The user practices piano using an external MIDI keyboard, for example, by playing a designated etude.

[0060] In the data acquisition and quantification stage of step S101, the system analyzes the real-time data stream from the MIDI keyboard and combines it with visual information collected by the camera module to quantify the user's multi-dimensional behavioral data. 1. Task execution efficiency data: The system records the total time taken by the user to play a complete piece and compares it with the standard performance time of the piece (or the target time set by the teacher) to obtain a ratio or difference, which serves as a quantitative indicator of task execution efficiency. 2. Task execution quality data: This is the focus of quantification in this scenario. The system accurately compares the MIDI note sequence played by the user with the standard sheet music. By comparing the pitch of each note, the system calculates the pitch accuracy; by comparing the trigger time of each note with the beat position on the sheet music, the system calculates the rhythm accuracy. Furthermore, it can analyze whether the duration of the notes meets the requirements of the sheet music. These indicators together constitute the task execution quality data. 3. Task engagement data: The data sources for this dimension are more abundant. On the one hand, the system can still analyze the user's visual focus, i.e., whether the user is looking at the sheet music or keyboard, using the computer vision technology described in Embodiment 1 or Embodiment 2 through the camera module. On the other hand, the system can also extract new engagement metrics from MIDI data. For example, it can analyze the stability of the dynamic values ​​played by the user. Generally, a focused performer has more stable and expressive dynamic control, while when attention is scattered, the playing dynamics may exhibit unconscious, chaotic fluctuations. Therefore, the dispersion or fluctuation pattern of the playing dynamics can serve as another effective supplement to measuring task engagement.

[0061] After collecting and quantifying the multi-dimensional data applicable to music scenarios, the subsequent processing flow is similar to that in Example 1. Combined models (e.g., multiple linear regression module, K-means clustering module, random forest classification module) are used to analyze this data.

[0062] For example, the model analysis might identify a user's ability characteristics as having a strong sense of rhythm but needing to improve pitch accuracy. Based on this, in step S105, the system will generate a corresponding ability report and intelligently recommend targeted practice content from the content library, such as recommending scale exercises aimed at improving pitch discrimination or recommending some sight-singing and ear training games, rather than simply recommending another new piece of music.

[0063] In the closed-loop feedback phase of step S106, the performance data (new pitch accuracy, rhythm accuracy, etc.) of the user after completing these recommended scale exercises or sight-singing and ear training games will be recaptured by the system and fed back to the front end of the model through the data feedback link. This data is used to dynamically adjust the judgment of the user's ability profile and affect the recommendation of the next round of practice content.

[0064] This embodiment successfully transfers the multi-dimensional data acquisition-combined model analysis-closed-loop feedback recommendation framework proposed in this invention from conventional academic learning scenarios to the field of music and art learning. By adaptively adjusting the dimensions of data acquisition, it achieves quantitative evaluation and adaptive training of musical expressiveness (such as rhythm, pitch, and dynamic control), solving the technical deficiency of general learning models in evaluating artistic expressiveness, thus powerfully demonstrating the universality and broad application prospects of the technical solution of this invention.

[0065] Example 5 This embodiment further describes the optimization method of the machine learning model involved in the present invention, as well as other functions that can be integrated as a complete intelligent learning system. To ensure the accuracy and generalization ability of the combined machine learning model in the capability feature recognition and recommendation module, this embodiment employs an optimization strategy combining cross-validation and periodic iteration.

[0066] In the initial training phase of the model, cross-validation can be used. Specifically, the system randomly samples from a large-scale database of peer groups (e.g., anonymized learning data of tens of thousands of primary school students) and divides this dataset into a first dataset and a second dataset, for example, in a 7:3 ratio. The first dataset, containing 70% of the data, is used as the training set to train the parameters of the combined model (multiple linear regression, K-means clustering, random forest classification). After training, the second dataset, containing 30% of the data, is used as the validation set to evaluate the performance of the trained model and test its prediction accuracy on unseen data. By repeating this process of partitioning, training, and validating (i.e., K-fold cross-validation) multiple times, the model parameter combination with the best overall performance is selected as the initial model deployed to the intelligent learning device.

[0067] Furthermore, the model is not static. The present invention also includes a mechanism for periodic iterative updates. As the intelligent learning device is continuously used by users, the system constantly collects new user learning behavior data. The cloud server periodically (e.g., quarterly) aggregates this newly generated data and uses it to incrementally train or retrain existing machine learning models. This periodic iterative update enables the model to continuously learn new data distributions and patterns, adapting to changes in student abilities and the development of teaching content, thereby maintaining the accuracy and timeliness of its analysis over the long term.

[0068] As a complete intelligent learning system, the intelligent learning device described in Embodiment 1 of this invention, in addition to realizing the core functions of talent discovery and enhancement, also supports a series of other advanced auxiliary functions in its hardware platform and software architecture. These functions can be used as optional modules of the present invention, together forming a powerful intelligent learning ecosystem.

[0069] For example, in one optional implementation, the system also includes a refined step-by-step error location function. When the textbook / homework recognition module detects that a student is pausing for a long time or repeatedly making corrections while solving a multi-step calculation problem (such as three-digit multiplication), the system not only determines that the problem is solved incorrectly or too slowly, but also calls upon the built-in problem-solving knowledge graph to compare the student's solution steps with the standard steps, thereby accurately locating the specific step where the student may have encountered difficulty with carrying over. Accordingly, the voice interaction module will trigger a targeted explanation: "I noticed that you may have made a mistake when carrying over from the tens place to the hundreds place. Let's review this step together," thus achieving more efficient tutoring than explaining from the beginning.

[0070] Alternatively, in another embodiment, the system supports adaptive loading of subject-specific models. When the system connects to a subject-specific peripheral (such as a MIDI keyboard) via an external interface (such as USB or Bluetooth), the system automatically loads and switches to a subject-specific analysis model. For example, in a piano practice scenario, this model not only analyzes practice duration and note accuracy but also focuses on specific dimensions of musical expression, such as rhythmic stability and evenness of playing dynamics, to address the limitation of general learning models in assessing artistic skills.

[0071] As an optional implementation, the system optimizes the recognition of the "point-to-learn" interaction intent. To reduce misjudgments of unintentional finger movements by students, when the textbook / homework recognition module detects a finger pointing to a passage, the system simultaneously calls the face / behavior recognition module to analyze the student's gaze focus and lip movements. If the system determines that the student's gaze is moving rapidly between the book and the workbook without any verbal intention, the system infers that the current intention is more likely visual positioning for copying rather than requesting explanation, and therefore suppresses interactive feedback to avoid unnecessary disturbance.

[0072] In further alternative implementations, the system employs specific optimization techniques during the model training and deployment phases. For example, after training, pruning and quantization of the model can significantly compress its size without substantially reducing its accuracy. This allows complex AI models that previously required powerful cloud computing support to run more frequently on the edge processing units of intelligent learning devices. This edge computing model reduces reliance on network connectivity, significantly lowers latency caused by data transmission and cloud computing, improves the system's real-time response speed, and optimizes the device's energy efficiency.

[0073] Finally, in one optional implementation that embodies humanistic care, the system also features a companion mode outside of learning. When the face / behavior recognition module detects that a student has closed their book and is in a non-learning state, such as relaxing, doodling, or chatting with parents, the system automatically switches to companion mode. In this mode, the system invokes a large language model that has been specially fine-tuned with affective computing and open-domain dialogue capabilities. This model is designed to avoid proactively mentioning topics such as learning and exams, but instead interacts with students through empathy, listening, and sharing interesting knowledge, playing the role of an empathetic friend. This solves the problem of traditional learning-oriented AI potentially feeling didactic and utilitarian in non-learning scenarios, thus better integrating into the family environment and becoming a true growth partner for students.

[0074] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0075] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A method for learning ability feature recognition and adaptive content recommendation based on multi-dimensional data fusion analysis, characterized in that, Includes the following steps: The system collects and quantifies multi-dimensional behavioral data streams of users in real time when performing tasks using sensors. These multi-dimensional behavioral data streams include at least: task execution efficiency data, task execution quality data, and user task engagement data obtained by processing video streams captured by cameras using computer vision technology. A combined machine learning model is used to process the multi-dimensional behavioral data stream to identify user capability characteristics. The combined machine learning model includes: a trend analysis model for data normalization and trend analysis of the multi-dimensional behavioral data stream; a user profile clustering model for clustering user profiles based on the output of the trend analysis model; and a capability characteristic classification model for classifying user capability characteristics by combining the clustering results of the user profile clustering model and the original multi-dimensional behavioral data stream. Based on the output of the capability feature classification model, content is matched in the content library and recommended to the user.

2. The method according to claim 1, characterized in that, The task execution efficiency data includes the task processing speed obtained by calculating the time consumed by the user to complete the task; the task execution quality data includes the task accuracy rate obtained by comparing the user's output with the standard answer.

3. The method according to claim 1, characterized in that, The steps for obtaining the task engagement data include: By using facial landmark detection and head pose estimation algorithms, the deviation angle and duration of the user's gaze direction from the task area are calculated, and a quantitative score of the task engagement data is generated based on the deviation angle and duration.

4. The method according to claim 3, characterized in that, The quantification of task engagement data also incorporates eye-tracking analysis of the user's gaze point distribution and saccade path.

5. The method according to claim 1, characterized in that, The trend analysis model is a multiple linear regression model; The user profile clustering model is either a K-means clustering model or a DBSCAN model; The capability feature classification model is a random forest classification model.

6. The method according to claim 1, characterized in that, Also includes: The system records user behavior feedback data in real time to the recommended content and feeds this feedback data as incremental data into the combined machine learning model to dynamically adjust the combined machine learning model and subsequent recommended content.

7. A learning ability feature recognition and adaptive content recommendation system based on multi-dimensional data fusion analysis, characterized in that, include: Multi-dimensional behavioral data stream acquisition module: Real-time acquisition and quantification of multi-dimensional behavioral data streams of users when performing tasks through sensors. The multi-dimensional behavioral data streams include at least: task execution efficiency data, task execution quality data, and user task engagement data obtained by processing video streams captured by cameras through computer vision technology. Multi-dimensional behavioral data stream analysis module: Employs a combined machine learning model to process the multi-dimensional behavioral data stream to identify user capability characteristics. The combined machine learning model includes: a trend analysis model for data normalization and trend analysis of the multi-dimensional behavioral data stream; a user profile clustering model for user profile clustering based on the output of the trend analysis model; and a capability characteristic classification model for classifying user capability characteristics by combining the clustering results of the user profile clustering model and the original multi-dimensional behavioral data stream. Recommendation content generation module: Based on the output of the capability feature classification model, it matches content in the content library and recommends content to the user.

8. The system according to claim 7, characterized in that, The task execution efficiency data includes the task processing speed obtained by calculating the time consumed by the user to complete the task; The task execution quality data includes the task accuracy rate obtained by comparing user output with the standard answer; The task engagement data includes calculating the deviation angle and duration between the user's gaze direction and the task area using facial key point detection and head pose estimation algorithms, and generating a quantitative score for the task engagement data based on the deviation angle and duration. The trend analysis model is a multiple linear regression model; the user profile clustering model is a K-means clustering model or a DBSCAN model; and the capability feature classification model is a random forest classification model.

9. The system according to claim 8, characterized in that, The quantification of task engagement data also incorporates eye-tracking analysis of the user's gaze point distribution and saccade path.

10. An intelligent learning system, characterized in that, include: Data acquisition layer, capability feature identification and recommendation module, application layer and data feedback link; The data acquisition layer includes a face / behavior recognition module and a textbook / homework recognition module; the face / behavior recognition module is used to process the video stream acquired by the infrared camera to identify the user's face and posture; the textbook / homework recognition module is used to process the images acquired by the high-definition camera to identify the task content and the user's answer. The ability feature recognition and recommendation module includes the learning ability feature recognition and adaptive content recommendation system based on multi-dimensional data fusion analysis as described in any one of claims 7 to 9; The application layer includes a content library and a user interface module; the content library pre-stores diverse task content. The user interface module is responsible for generating the interactive interface and presenting information including recommended content and analysis reports on the projection display area through the projection optical engine. The data feedback link transmits the behavioral feedback data generated by the user at the application layer back to the data acquisition layer as incremental data to participate in the next round of analysis.