Attention detection method and system and electronic equipment
By fusing multimodal perception data and adjusting dynamic thresholds, the accuracy problem of attention state detection in existing technologies has been solved, enabling precise identification and personalized adaptation of user attention states, and improving the reliability and applicability of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-28
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies struggle to accurately detect user attention states, leading to wasted learning time, increased error rates, and decreased interaction satisfaction. They also fail to effectively identify continuous and subtle fluctuations in attention levels.
By acquiring multimodal perception data, including physiological state, posture and device operating status, we perform correlation fusion and feature extraction, dynamically adjust thresholds to identify attention state, use the degree of deviation for fine evaluation, and perform personalized calibration based on historical data.
It achieves precise detection of user attention state, improves detection accuracy and adaptability, and can adjust thresholds in a timely manner to adapt to individual characteristics, thereby enhancing the reliability and accuracy of attention detection.
Smart Images

Figure CN121845580A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to an attention detection method, system, and electronic device. Background Technology
[0002] In the use of electronic devices, there are various scenarios such as learning or working. In these scenarios, maintaining and improving the user's attention level is crucial to ensuring learning effectiveness, work efficiency, and a deep human-computer interaction experience. The concentration and dispersion of user attention directly determine the efficiency of information reception, the quality of task completion, and the continuity of immersive interaction. However, in existing interaction modes that lack effective state awareness and feedback, users often find it difficult to consciously detect unconscious shifts in their attention, and also struggle to adjust their state promptly and effectively after being distracted. This leads to potential wasted learning time, increased work error rates, and decreased interaction satisfaction.
[0003] Therefore, accurately detecting the user's attention state is fundamental to improving the user experience of electronic devices. Summary of the Invention
[0004] The purpose of this disclosure is to provide an attention detection method, system, and electronic device. By acquiring and fusing multimodal perception data, it can identify personalized fluctuation patterns of a target user's attention, and then perform threshold judgment on the perception data and determine the attention state based on the difference value after judgment, thereby achieving accurate detection of user attention. Furthermore, before the threshold judgment process, this application can also adaptively adjust the threshold based on perception data from historical time periods for individual users, so that the threshold judgment result is more consistent with the user's actual attention state, further improving the accuracy of attention detection.
[0005] To achieve the above objectives, the embodiments of this disclosure provide the following technical solutions: On the one hand, an attention detection method is provided. This attention detection method includes: First, acquire the target user's first multimodal perception data while using the electronic device at the current moment. This multimodal perception data includes the target user's physiological state data, posture data, and the electronic device's operational status data. Second, perform correlation fusion and feature extraction processing on the first multimodal perception data to obtain the target user's attention feature parameters. Then, acquire second multimodal perception data from historical time periods. Based on this second multimodal perception data, adjust the target user's initial attention threshold to generate a target attention threshold. The initial attention threshold is obtained by analyzing the attention data of multiple users within the target user's target group, and each target attention threshold corresponds to an attention feature parameter. Finally, determine the target user's attention state result based on the degree of deviation between the attention feature parameters and the corresponding target attention threshold.
[0006] Based on the attention detection method disclosed herein, firstly, multimodal perception data of the target user's physiological state, posture, and electronic device operating status are acquired. These data are then correlated, fused, and extracted to obtain attention feature parameters. This method can identify user attention influencing factors from multiple dimensions, including physiological, behavioral, and device interaction, overcoming the problem that a single data source cannot fully reflect the characteristics. Furthermore, based on the deviation between the attention feature parameters and a dynamic threshold, the attention state result is determined. This method utilizes the continuity of the deviation and the combination of deviations between different attention feature parameters to achieve a refined evaluation of the attention state result, rather than relying on a binary conclusion derived from a preset fixed threshold. This improves the accuracy and resolution of the state result determination. Simultaneously, this method can generate objective calibration criteria to guide threshold adjustment based on multimodal perception data from historical periods. This eliminates threshold mismatches caused by differences in physiological state or scenario, adjusting the threshold to a range that matches the user's individual characteristic patterns. This ensures the adaptability of the deviation calculation between the attention feature parameters and the dynamic threshold, resulting in a more accurate and reliable final attention state evaluation result.
[0007] In some embodiments, second multimodal perception data within a historical time period is acquired. Based on this data, the initial attention threshold of the target user is adjusted to generate a target attention threshold. This includes: determining the target group to which the target user belongs based on the target user's attribute information, and acquiring the initial attention threshold and baseline attention feature parameters corresponding to the target group. The attribute information includes: the target user's age, learning stage, learning ability quantification level, and subject type. The group is obtained by dividing different users based on the attribute information. The baseline attention feature parameters are obtained by analyzing and processing the attention feature parameters of multiple users within the target group. The second multimodal perception data is correlated, fused, and feature extracted to obtain the target user's attention feature parameters within the historical time period. The degree of deviation between the target user's attention feature parameters within the historical time period and the baseline attention feature parameters is determined. Based on the degree of feature deviation, the initial attention threshold is adjusted first to generate the target attention threshold. The first adjustment includes the following strategies: a compensation adjustment strategy based on the degree of decline in the target user's physiological state, a compensation adjustment strategy based on the working time of the electronic device, and an adaptation adjustment strategy based on the difficulty of the learning task performed by the electronic device.
[0008] In other embodiments, the runtime state data includes runtime context information of the electronic device, wherein the runtime context information includes: interaction information of the foreground application of the electronic device and the file format of the data loaded by the foreground application. Acquiring second multimodal perception data within a historical time period, and adjusting the initial attention threshold of the target user based on the second multimodal perception data to generate a target attention threshold for the target user, includes: performing semantic recognition based on runtime context information to determine the usage scenario type of the electronic device; determining a threshold adjustment strategy for a second adjustment of the initial attention threshold based on the usage scenario type and a preset mapping relationship between scenarios and threshold adjustment strategies; and performing a second adjustment on the first attention threshold based on the threshold adjustment strategy to generate the target attention threshold for the target user.
[0009] In other embodiments, second multimodal perception data within a historical time period is acquired, and the initial attention threshold of the target user is adjusted based on the second multimodal perception data to generate the target attention threshold of the target user, including: a first adjustment and / or a second adjustment.
[0010] In other embodiments, the attention state of the target user is determined based on the degree of deviation between the attention feature parameters and the corresponding target attention threshold. This includes: determining the degree of deviation between each attention feature parameter and the corresponding target attention threshold; determining the usage scenario type of the electronic device; constructing a feature combination vector from the threshold deviation and constructing a prompt message from the usage scenario type; and processing the feature combination vector based on the prompt message using an attention analysis model to obtain the attention state of the target user. The attention analysis model is trained through at least two model training phases. The model training phases include a pre-training phase based on different user attention data and a reinforcement training phase based on the target user's attention data.
[0011] In other embodiments, the target user's physiological state data includes one or more of the following: the target user's respiratory rate and heart rate variability. This physiological state data is acquired by detecting the target user using a millimeter-wave radar array. The millimeter-wave radar array includes multiple millimeter-wave radar units at different detection angles, deployed at the edges and / or sides of the electronic device's display screen. The target user's posture data includes one or more of the following: the target user's movement trajectory, the target user's movement amplitude, the sitting pressure distribution entropy value, the sitting body movement frequency, writing trajectory information, writing pressure information, and writing angle information. This posture data is acquired by detecting the target user using at least one of the following sensors: a millimeter-wave radar array, a flexible piezoelectric thin-film sensor matrix, and a stylus sensor. The flexible piezoelectric thin-film sensor matrix and the stylus sensor are communicatively connected to the electronic device. The electronic device's operational state data includes one or more of the following: the electronic device's operational context information, computational resource consumption information, and the electronic device's interaction information.
[0012] In other embodiments, the method further includes: when the target user's attention state result is less than or equal to a preset attention state result threshold, determining a multimodal attention intervention strategy and its execution order based on the attention state result, the degree of feature deviation, and the usage scenario type. The multimodal attention intervention strategy includes at least one of auditory, visual, tactile, and guided intervention strategies. The multimodal attention intervention strategy is executed according to the execution order to improve the target user's attention state result.
[0013] In other embodiments, when the target user's attention state result is less than or equal to a preset attention state result threshold, a multimodal attention intervention strategy and its execution order are determined based on the attention state result, the degree of feature deviation, and the usage scenario type. This includes: determining the attention intervention priority corresponding to the attention state result based on a preset mapping relationship between state results and priorities; determining the multimodal attention intervention strategy for the target user from a preset mapping table of intervention priorities and intervention strategies; ranking the intervention strategy set by importance based on the degree of feature deviation and the usage scenario type to determine the execution priority of each intervention strategy in the multimodal attention intervention strategy set, and generating the execution order of the intervention strategies. The attention intervention priorities, ranked from lowest to highest, include: a first attention intervention priority, a second attention intervention priority, and a third attention intervention priority. The first attention intervention priority corresponds to visual intervention strategies; the second attention intervention priority corresponds to both visual and auditory sensory strategies; and the third attention intervention priority corresponds to both tactile and auditory intervention strategies.
[0014] In other embodiments, when the target user's attention state result is less than or equal to a preset attention state result threshold, a multimodal attention intervention strategy and its execution order are determined based on the attention state result, the degree of feature deviation, and the usage scenario type. This includes: obtaining a sequence of the target user's attention state results. The attention state result sequence includes attention state results determined at different times. The attention state result sequence is processed using an attention state transition prediction model to obtain the target user's attention state change trend over a future period. Based on the attention state change trend, attention state results, usage scenario type, and the degree of feature deviation, a multimodal attention intervention strategy and its execution order are determined for the target user.
[0015] In other embodiments, the method further includes: obtaining a first attentional state result of the target user after implementing the multimodal attention intervention strategy; determining an attention improvement value relative to the attentional state result before implementing the multimodal attention intervention strategy; and, if the attention improvement value is less than or equal to an attention improvement threshold, re-determining and implementing the multimodal attention intervention strategy.
[0016] In other embodiments, the method further includes: obtaining interaction information of foreground applications from runtime context information, and determining a list of running foreground applications of the electronic device from the interaction information of the foreground applications. If the list of running foreground applications meets preset conditions and the change in attention feature parameters is less than a change threshold, then unwanted applications in the list of running foreground applications are closed or paused. The preset conditions include: the list of running foreground applications contains at least one desired application and at least one unwanted application.
[0017] In other embodiments, the method further includes: synchronizing the target user's attention data to a server, so that the server updates the initial attention threshold and / or baseline attention feature parameters of the target group based on the attention data of multiple users within the target group. The attention data includes one or more of the following: the target user's identification identifier, first multimodal perception data, attention feature parameters, and attention state results. And / or, synchronizing the target user's attention data to the device terminals of associated users of the target user. The device terminals are used to provide the target user's attention data to the associated users.
[0018] In other embodiments, the method further includes: determining the target user's usage behavior pattern of the electronic device based on the electronic device's operational status data; and generating an operational strategy for the electronic device based on the usage behavior pattern. The operational strategy is used to adjust the electronic device's content output adjustment parameters, user permissions, and application scheduling plans.
[0019] On the other hand, an attention detection system is provided. The system includes: electronic devices and sensor units.
[0020] The sensor unit communicates with or is deployed within an electronic device. The electronic device includes: a data acquisition module, a feature extraction module, a threshold generation module, and an attention determination module.
[0021] The data acquisition module is configured to acquire first multimodal perception data of the target user using the electronic device at the current moment through the sensor unit. The multimodal perception data includes: the target user's physiological state data, the target user's posture data, and the electronic device's operating status data.
[0022] The feature extraction module is configured to perform correlation fusion and feature extraction processing on the first multimodal perception data to obtain the attention feature parameters of the target user.
[0023] The threshold generation module is configured to acquire second multimodal perception data within a historical time period, adjust the initial attention threshold of the target user based on the second multimodal perception data, and generate the target attention threshold for the target user. The initial attention threshold is obtained by analyzing and processing the attention data of multiple users within the target group to which the target user belongs. Each target attention threshold corresponds to an attention feature parameter.
[0024] The attention determination module is configured to determine the attention state of the target user based on the degree of deviation between the attention feature parameters and the corresponding target attention threshold.
[0025] In some embodiments, the sensor unit includes one or more of a millimeter-wave radar array, a flexible piezoelectric thin-film sensor matrix, and a stylus sensor. The millimeter-wave radar array includes multiple millimeter-wave radar units with different detection angles. Each millimeter-wave radar unit is deployed at the edge and / or side of the electronic device's display screen. The flexible piezoelectric thin-film sensor matrix and the stylus sensor are communicatively connected to the electronic device.
[0026] The stylus sensors include: a three-axis speed sensor, a pressure sensor, and a tilt angle sensor.
[0027] The millimeter-wave radar array is configured to collect physiological state data of the target user and posture data related to the target user's actions.
[0028] A flexible piezoelectric thin-film sensor matrix is configured to acquire posture data related to the sitting posture of the target user.
[0029] The stylus sensor is configured to collect posture data related to the writing posture of the target user.
[0030] The electronic devices are also configured to collect operational status data during operation.
[0031] In another aspect, an electronic device is provided. It includes a memory and one or more processors, the memory being coupled to the processors; wherein the memory stores computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the electronic device performs the attention detection method and its embodiments described above.
[0032] The attention detection system and electronic device described above have the same beneficial technical effects as the attention detection method and its embodiments described above, and will not be repeated here. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in this disclosure, the accompanying drawings used in some embodiments of this disclosure will be briefly described below. Obviously, the drawings described below are only drawings of some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings. In addition, the drawings described below can be regarded as schematic diagrams and are not intended to limit the actual size of the product, the actual flow of the method, the actual timing of the signals, etc. involved in the embodiments of this disclosure.
[0034] Figure 1 This is a schematic diagram of the architecture of an attention detection system according to some embodiments. Figure 1 ; Figure 2 This is a functional schematic diagram of a device status detection module according to some embodiments; Figure 3 A flowchart illustrating an attention detection method according to some embodiments. Figure 1 ; Figure 4 This is a flowchart illustrating a threshold adjustment method according to some embodiments. Figure 1 ; Figure 5 This is a flowchart illustrating a threshold adjustment method according to some embodiments. Figure 2 ; Figure 6 A flowchart illustrating an attention detection method according to some embodiments. Figure 2 ; Figure 7 This is a flowchart illustrating the feedback regulation mechanism of an attention intervention strategy according to some embodiments; Figure 8 This is a schematic diagram of the attention state transition path according to some embodiments; Figure 9 This is a schematic diagram of the data processing flow of a hybrid attention prediction model according to some embodiments; Figure 10 This is a flowchart illustrating a learning mode switching strategy according to some embodiments; Figure 11 This is a schematic diagram of the architecture of an attention detection system according to some embodiments. Figure 2 ; Figure 12 This is a schematic diagram of the structure of an electronic device according to some embodiments. Detailed Implementation
[0035] The technical solutions in some embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments provided in this disclosure are within the scope of protection of this disclosure.
[0036] Unless the context otherwise requires, throughout the specification and claims, the term "comprising" is interpreted as open-ended and encompassing, meaning "including, but not limited to." In the description of the specification, terms such as "one embodiment," "some embodiments," "exemplary embodiment," "example," or "some examples" are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this disclosure. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics mentioned may be included in any suitable manner in any one or more embodiments or examples.
[0037] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.
[0038] In describing some embodiments, the terms "coupled" and "connected" and their derivative expressions may be used. The term "connected" should be interpreted broadly; for example, "connected" can be a fixed connection, a detachable connection, or an integral part; it can be a direct connection or an indirect connection through an intermediate medium. The embodiments disclosed herein are not necessarily limited to the content of this document.
[0039] "At least one of A, B and C" has the same meaning as "at least one of A, B or C", both including the following combinations of A, B and C: only A, only B, only C, combinations of A and B, combinations of A and C, combinations of B and C, and combinations of A, B and C.
[0040] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.
[0041] As used herein, depending on the context, the term “if” may optionally be interpreted as meaning “when”, “in the event of”, “in response to determination”, or “in response to detection”. Similarly, depending on the context, the phrase “if it is determined that…” or “if [the stated condition or event] is detected” may optionally be interpreted as meaning “in the event of determination that…”, “in response to determination that…”, “when [the stated condition or event] is detected”, or “in response to the detection of [the stated condition or event]”.
[0042] The use of “applies to” or “configured to” in this article implies an open and inclusive language that does not preclude applicability to or configuration to devices that perform additional tasks or steps.
[0043] In addition, the use of “based on” implies openness and inclusivity, because processes, steps, calculations or other actions “based on” one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated.
[0044] As used herein, “about,” “approximately,” or “approximately” includes the stated value and the average value within an acceptable range of deviation from the given value, wherein the acceptable range of deviation is determined by a person skilled in the art taking into account the measurement under discussion and the error associated with the measurement of the given quantity (i.e., the limitations of the measurement system).
[0045] In the use of electronic devices, there are various scenarios such as learning or working. In these scenarios, maintaining and improving the user's attention level is crucial to ensuring learning effectiveness, work efficiency, and a deep human-computer interaction experience. Therefore, how to accurately detect the user's attention state is an important problem that needs to be solved.
[0046] In one technical solution, electronic devices rely on preset fixed thresholds to determine attentional state. However, this static, fixed threshold-based approach results in a binary outcome (e.g., only distinguishing between "focused" and "distracted"), which fails to accurately reflect the continuous, subtle fluctuations in the user's attention level. It cannot effectively differentiate between different degrees of focused or distracted attention, leading to overly vague and coarse evaluation results that cannot adequately define and describe the user's true attentional state.
[0047] In view of this, this disclosure provides an attention detection method, the inventive concept of which is to construct a comprehensive representation system by fusing multimodal perception data and introduce a dynamic threshold calibration mechanism based on historical data, and to perform refined dynamic evaluation based on the degree of continuous deviation. Compared with binary static threshold judgment, this method can achieve more accurate and personalized identification and description of the user's attention state.
[0048] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0049] The attention detection method provided in this application embodiment can be applied to an attention detection system (hereinafter referred to as the system).
[0050] like Figure 1 As shown, the system includes an electronic device 110 and a sensor unit 120. The sensor unit 120 serves as the sensing front-end of the system, providing multimodal sensing data to the electronic device 110 for attention detection. The electronic device 110, as the data processing back-end of the system, processes the multimodal sensing data to obtain the attention detection result of the target user.
[0051] The electronic device 110 includes: a data acquisition module 111, a feature extraction module 112, a threshold generation module 113, and an attention determination module 114. The electronic device can be a learning machine, tablet computer, mobile phone, or mobile computer, or other terminal device with data processing capabilities. This application embodiment does not limit the specific form of the electronic device. Specifically, it includes: The data acquisition module 111 is configured to acquire first multimodal perception data of the target user using the electronic device at the current moment through the sensor unit. The multimodal perception data includes: the target user's physiological state data, the target user's posture data, and the electronic device's operational status data.
[0052] The feature extraction module 112 is configured to perform correlation fusion and feature extraction processing on the first multimodal perception data to obtain the attention feature parameters of the target user.
[0053] The threshold generation module 113 is configured to acquire second multimodal perception data within a historical time period, adjust the initial attention threshold of the target user based on the second multimodal perception data, and generate the target attention threshold for the target user. The initial attention threshold is obtained by analyzing and processing the attention data of multiple users within the target group to which the target user belongs. Each target attention threshold corresponds to an attention feature parameter.
[0054] The attention determination module 114 is configured to determine the attention state result of the target user based on the degree of threshold deviation between the attention feature parameters and the corresponding target attention threshold.
[0055] The sensor unit 120 includes several components: a millimeter-wave radar array 121, a flexible piezoelectric thin-film sensor matrix 122, a stylus sensor 123, and a device status detection module 124. The millimeter-wave radar array 121 and the device status detection module 124 are deployed on the electronic device, and the flexible piezoelectric thin-film sensor matrix 122 and the stylus sensor 123 are communicatively connected to the electronic device (e.g., local area network connection or Bluetooth connection, etc.; this embodiment does not limit the connection method).
[0056] The millimeter-wave radar array 121 includes millimeter-wave radar units with different detection angles, and multiple millimeter-wave radar units are deployed on the edge and / or side of the electronic device display screen.
[0057] The millimeter-wave radar array 121 is configured to collect physiological state data of the target user and posture data related to the target user's actions.
[0058] Specifically, the millimeter-wave radar array 121, through multi-node collaborative detection and data fusion, can construct a three-dimensional micro-motion vector field in the space in front of the electronic device (i.e., the surface facing the user). This vector field can resolve sub-millimeter-level displacement and attitude angle changes of key parts of the user's body (such as the head and chest), thereby extracting refined physiological and behavioral features, including but not limited to respiratory rate, heart rate variability, and continuous posture trajectory. Through array-based deployment and spatial coverage optimization, the system effectively overcomes the limitations of a single sensor's limited field of view and susceptibility to occlusion, achieving robust perception of the user's multi-dimensional state information in learning scenarios, and providing a high signal-to-noise ratio raw data foundation for subsequent attention feature calculation.
[0059] For example, the specifications of the millimeter-wave radar unit can be 9 channels at 77 GHz, or 4 channels at 60 GHz, etc. This application does not limit the specifications of the millimeter-wave radar unit.
[0060] In some embodiments, the physiological state data collected by the millimeter-wave radar array 121 includes respiratory rate and heart rate variability.
[0061] Among them, changes in respiratory rate can reflect a user's emotional fluctuations and cognitive load level. For example, breathing may tend to be slow and stable when focusing on thinking. Heart rate variability is related to the activity of the autonomic nervous system and can characterize a user's stress level and psychological arousal level, providing a physiological basis for judging their attention maintenance ability.
[0062] For example, the normal reference range for respiratory rate can be set to 12-20 breaths per minute. When the detected value consistently exceeds this range and is accompanied by specific postural changes, it may indicate that the state of attention may have changed.
[0063] In some embodiments, the action-related attitude data collected by the millimeter-wave radar array 121 includes: the action trajectory of the target user and the action amplitude of the target user.
[0064] Among these, motion trajectory can be used to analyze the movement path and patterns of a user's body or head in space; frequent and irregular trajectory changes are often associated with distraction. Movement amplitude can be used to quantify the intensity of a user's physical activity; excessively small or large amplitudes may correspond to states of lethargy or restlessness, respectively.
[0065] The flexible piezoelectric thin film sensor matrix 122 is configured to acquire posture data related to the sitting posture of the target user.
[0066] The flexible piezoelectric thin film sensor matrix 122 is composed of multiple piezoelectric sensing units arranged in an array, such as in rows and columns of 16×16 or 18×18, which can collect and output pressure distribution image data of the user's contact area in real time.
[0067] In some embodiments, the flexible piezoelectric thin-film sensor matrix 122 can be specifically deployed in a cushion, backrest, or floor mat that is communicatively connected to an electronic device. This application embodiment does not limit the carrier tool for the flexible piezoelectric thin-film sensor matrix 122.
[0068] In some embodiments, the posture data related to sitting posture collected by the flexible piezoelectric thin film sensor matrix 122 includes: sitting pressure distribution entropy value and sitting body movement frequency.
[0069] Among them, the entropy value of sitting pressure distribution is used to quantify the uniformity and stability of pressure distribution in the buttocks. An increased entropy value often indicates unstable sitting posture and frequent adjustments, which may be a sign of distracted attention. The sitting body movement frequency refers to the number of micro-movements a user makes per unit of time by adjusting their sitting posture. An abnormally high frequency usually reflects that the user is in a state of anxiety, impatience, or fatigue.
[0070] The stylus sensor 123 is configured to collect posture data related to the writing posture of the target user.
[0071] In this example, the stylus sensor 123 includes a three-axis velocity sensor, a pressure sensor, and a tilt angle sensor. The three-axis accelerometer is used to detect the pen's motion state in real time during writing, acquiring multi-dimensional motion parameters including instantaneous acceleration, rate of change of acceleration, and smoothness of the motion trajectory. The pressure sensor detects the real-time writing pressure on the pen tip and outputs the pressure value and pressure fluctuation rate. The tilt angle sensor detects the pen grip posture and outputs the real-time tilt angle and angle fluctuation amplitude.
[0072] In some embodiments, the stylus sensor 123 may be deployed with a preset long short-term memory (LSTM) network model. By fusing and analyzing the time-series data output by the sensor through the LSTM model, including jointly modeling the periodicity of pressure fluctuation rate, the coherence of acceleration trajectory, and the stability of angle change, perceptual data representing writing posture can be extracted from real-time writing behavior.
[0073] In some embodiments, the writing posture data collected by the stylus sensor 123 includes: writing trajectory information, writing pressure information, and writing angle information.
[0074] The writing trajectory information includes the pen tip's movement path, speed, and acceleration changes; a continuous and regular trajectory usually corresponds to a focused writing state. Writing pressure information reflects the magnitude and changes in the force applied by the pen tip to the writing surface; stable pressure output is one indicator of writing engagement. Writing angle information refers to the pen body's tilt angle relative to the writing plane and its changes; maintaining a relatively stable angle contributes to controllability and focus during writing.
[0075] For example, the specific content of this feature information can be found in Table 1.
[0076] Table 1
[0077] As shown in Table 1, by analyzing the writing posture data obtained from the stylus sensor, different patterns representing writing behavior with high attentional input and casual behavior with low attentional input can be distinguished.
[0078] In this example, the acquisition of operating status data by the electronic device can be specifically achieved through the device status detection module 124, that is: The device status detection module 124 is configured to collect operating status data of electronic equipment.
[0079] In some embodiments, the operating status data collected by the device status detection module 124 includes: operating context information of the electronic device, computing resource consumption information, and interaction information of the electronic device.
[0080] The runtime context information includes foreground application interaction information, runtime, network connection status, etc., which is used to determine the type of task currently being performed by the device (such as expected type, such as work application, study application, etc., and unexpected type, such as entertainment application, movie application, etc.).
[0081] Information on computing resource consumption includes real-time usage of CPU, memory, and battery. Abnormal resource usage patterns may be related to non-learning activities or system anomalies.
[0082] Interaction information includes: touch behavior (click or swipe, etc.), task switching frequency, and input information type.
[0083] It should be understood that the aforementioned multimodal perception data system, composed of physiological state data, posture data, and device operation status data, can overcome the limitations in dimensionality and insufficient anti-interference capabilities inherent in relying on a single data source for attention detection. By integrating the intrinsic state representation of physiological signals, the external manifestation of behavioral actions, and the contextual information of device interaction, this system provides a complementary and rich foundation for comprehensively and robustly assessing user attention states through joint observation.
[0084] In other embodiments, the device status detection module 124 is also configured to implement application management and attention intervention strategies.
[0085] like Figure 2 As shown, the functions of the device status detection module 124 include: collecting operating status data, application management, and executing attention intervention strategies.
[0086] The runtime status data acquisition function is used to collect: runtime context analysis, computing resource consumption analysis, and interaction information analysis.
[0087] Among them, runtime context analysis is used to obtain the above runtime context information, including: application runtime detection and task difficulty analysis.
[0088] The computational resource consumption analysis is used to obtain the above-mentioned computational resource consumption information, including CPU usage analysis and memory usage analysis.
[0089] Interaction information analysis is used to obtain the above-mentioned interaction information, including: touch behavior records, task switching frequency, and input patterns.
[0090] Application management includes maintaining an application whitelist (also known as a list of desired applications), which includes the application's identification information. When an electronic device is detected running an application that is not on the application whitelist (i.e., an unwanted application), restrictions are imposed on the unwanted application.
[0091] The execution of attention intervention strategies includes: attention intervention strategies responsible for implementing system decisions. For a detailed introduction to attention intervention strategies, please refer to S501-S502 below, which will not be elaborated here.
[0092] like Figure 3 As shown, when the attention detection method provided in this application is applied to the above-mentioned attention detection system, it specifically includes the following steps S101-S104: S101. Obtain the first multimodal perception data of the target user when using the electronic device at the current moment.
[0093] Multimodal sensing data refers to sensing data collected by multiple sensors integrated into or connected to electronic devices. For a detailed description of the sensors, please refer to the above text. Figure 1 This will not be elaborated upon here.
[0094] Specifically, multimodal sensing data includes: the target user's physiological state data, the target user's posture data, and the electronic device's operational status data. For an introduction to the acquisition and explanation of multimodal sensing data, please refer to the above text. Figure 1 Its detailed description is omitted here.
[0095] In one technical solution, electronic devices typically rely on a single type of data (such as EEG signals, vision, etc.) for attention perception. Such methods have limited data analysis dimensions and inherent limitations: EEG signals require additional detection equipment, making widespread application difficult; visual data depends on ambient lighting conditions and is easily obstructed. Therefore, this application constructs a complementary perception system by simultaneously collecting multi-dimensional data, including physiological, posture, and device operation data, which overcomes the shortcomings of incomplete representation and susceptibility to interference from single data sources.
[0096] In some embodiments, the timing of the system executing S101 can be timed acquisition, event-triggered acquisition (e.g., when the user starts a learning task, the acquisition begins), or it can be manually started by the user. This application embodiment does not limit the specific timing of the system acquiring multimodal perception data.
[0097] S102. Perform correlation fusion and feature extraction processing on the first multimodal perception data to obtain the attention feature parameters of the target user.
[0098] Among these, correlation fusion refers to the correlation analysis of perceptual data from different modalities, and through the complementary synergy of different data, to uncover more comprehensive and robust attention representation data. Feature extraction refers to identifying and quantifying key indicators that can distinguish user attention states from the fused data.
[0099] Specifically, since multimodal data differ in their acquisition sources, signal characteristics, and noise performance, direct processing may introduce information redundancy and interference. Therefore, by associating and fusing multimodal sensing data and extracting features, the system can enhance effective information, suppress redundant noise, and generate more discriminative attention feature parameters.
[0100] In some embodiments, attentional characteristic parameters may include: the variance of respiratory rate, the ratio of low to high frequency in heart rate variability, micro-movement splitting, micro-movement amplitude, entropy value of sitting pressure distribution, mean trunk forward tilt angle, writing action pattern, device application switching frequency, and device interaction behavior pattern. Further, they may include one or more of the following: respiratory-heart rate coordination index, posture-writing stability coupling coefficient, and physiological-interaction response delay matching degree. This application does not limit the specific types of attentional characteristic parameters.
[0101] In some embodiments, since multimodal sensing data is heterogeneous and the data modes differ in dimensions, sampling rates and signal characteristics, the system can preprocess the multimodal sensing data before performing data association and fusion.
[0102] This preprocessing can include: data cleaning (such as filtering and noise reduction), format standardization, temporal alignment, and spatial registration. Specifically, data cleaning removes noise and outliers from the original signal; format standardization converts heterogeneous data into feature vectors with uniform dimensions and numerical ranges; temporal alignment uses the system clock as a reference to timestamp and synchronize data streams with different sampling frequencies; and spatial registration uses a coordinate transformation matrix to map sensor data from different spatial locations to a unified coordinate system.
[0103] One possible implementation is for radar arrays and flexible piezoelectric film sensor matrices, which include multiple sensors deployed in different spatial locations. In order to integrate multi-view data to form a unified and coherent spatial sensing field, the system needs to spatially register the acquired data.
[0104] For example, taking a radar array as an example, the system constructs a unified three-dimensional coordinate system (e.g., with the center of the learning machine screen as the origin, the X-axis horizontally to the right, the Y-axis vertically upward, and the Z-axis perpendicular to the screen and outward). Using a coordinate transformation matrix calibrated based on the physical installation positions of each radar sensor, local observation data from radars at different locations (such as the top and sides) are mapped to this same coordinate system, resulting in normalized three-dimensional spatial point clouds or micro-motion vector field data. This normalized spatial data can be used for quantitative analysis of the spatial motion trajectory, posture changes, and micro-motion patterns of user body parts (such as the head and torso).
[0105] For example, consider a flexible piezoelectric thin-film sensor matrix, which typically consists of multiple discrete pressure sensing units distributed on or inside the surface of the support tool. After spatially registering the data collected by each pressure sensing unit, the system can use a weighted fusion algorithm to collaboratively process the synchronous pressure values from each sensing unit to obtain a robust overall posture pressure characterization.
[0106] Let the number of effective sensing units be n, and the output pressure value of the i-th pressure sensor unit be... Signal quality weights are (Based on the implementation of signal-to-noise ratio assessment) and the weight of the contribution of the acquisition angle (calculated based on the angle between the normal direction of the sensing unit and the pressure direction). The process of calculating the pressure value after fusion satisfies equation (1):
[0107] In some embodiments, the system may employ a pre-defined feature fusion model (such as a rule engine or a multi-branch processing module) for correlation fusion. This model receives pre-processed multimodal perception data as input and performs fusion analysis on the multimodal perception data according to physiological, action / posture, and device operating state dimensions respectively, obtaining intermediate feature sets corresponding to each dimension. Subsequently, the model comprehensively weights and logically correlates these three intermediate feature sets using pre-defined cross-dimensional association rules or weights to obtain at least one attention feature parameter.
[0108] For example, taking heart rate variability in physiological perception data and head movement trajectory in motion posture perception data as examples, the attention feature parameters extracted from these two data are shown in Table 2.
[0109] Table 2
[0110] S103. Obtain the second multimodal perception data within the historical time period, adjust the initial attention threshold of the target user based on the second multimodal perception data, and generate the target attention threshold of the target user.
[0111] The attention threshold can be understood as the critical or boundary value of the attention feature parameter, and each target attention threshold corresponds to an attention feature parameter.
[0112] The initial attention threshold is a value obtained by analyzing and processing the attention data of multiple users within the target group to which the target user belongs. It is used to provide an initial reference benchmark based on the commonalities of the group for each attention feature parameter.
[0113] The second multimodal sensing data within a historical period refers to the time series of multiple batches of multimodal sensing data generated by the target user within a preset time period (e.g., the most recent 15 minutes) and arranged in chronological order.
[0114] Specifically, acquiring the second multimodal perception data within a historical timeframe is crucial because this data sequence continuously records users' attention-related behaviors and physiological performance over a period of time. By analyzing and processing this data, the system can learn and quantify the stable baseline level, typical fluctuation range, and personalized change patterns of each user's individual attention. This allows the system to extract quantitative patterns representing individual user characteristics, providing a direct basis for adjusting the target threshold, which is based on a general initial threshold derived from group statistics, to better reflect the user's actual state.
[0115] In some embodiments, the system's adjustment of the initial attention threshold may include: a first adjustment based on group and individual differences, and / or a second adjustment based on the current usage scenario of the user's electronic device.
[0116] The first adjustment refers to the system learning the baseline level and fluctuation pattern of the target user's personal attention characteristics based on their historical multimodal perception data, and then personalizing the initial threshold based on group statistics. The second adjustment refers to the system dynamically adapting the corresponding threshold based on the identified current usage scenario type (such as in-depth reading or interactive problem-solving) and the preset attention performance pattern requirements of different scenarios.
[0117] In some embodiments, before adjusting the initial attention threshold, the system may also pre-calibrate and adjust the initial threshold based on the user's usage preferences set in the electronic device being used (such as a user-defined level of focus sensitivity or a comfort zone marked in historical feedback) to provide an initial benchmark that is closer to the user's subjective feelings before personalized learning begins.
[0118] In some embodiments, the timing of threshold adjustment can be periodic, such as triggered after a certain period of accumulated usage; it can also be based on scene changes, such as triggered when the system detects a change in user task type or application environment; it can be session-based, such as batch updating based on the overall performance of attention state during the previous one or several usage periods each time the user starts using the electronic device; or it can be event-triggered, such as triggered when the system detects a significant drift in the distribution of user attention features or when the user actively calibrates their state. This application embodiment does not limit the timing of threshold adjustment.
[0119] S104. Determine the target user's attention state based on the degree of deviation between the attention feature parameters and the corresponding target attention threshold.
[0120] The threshold deviation refers to the quantitative difference calculated by comparing the target user's attention feature parameters at the current moment with the corresponding adjusted target attention threshold.
[0121] For example, attention state results may include continuous attention concentration scores (such as values from 0 to 100, with attention ranging from low to high) or discrete attention levels (such as levels like "highly focused", "moderately focused", "slightly distracted", "severely distracted", etc.). The specific form of attention state results is not limited in the embodiments of this application.
[0122] Specifically, if a binary comparison is directly performed between the target attention threshold and feature parameters (e.g., determining whether a threshold is exceeded), this method can only provide isolated yes / no judgments when multiple feature parameters and multiple thresholds coexist, failing to comprehensively measure the overall deviation across various dimensions. Furthermore, different attention feature parameters often exhibit correlations and complementarities; relying solely on a single threshold for exceeding limits ignores these correlations and cannot quantify the severity of the deviation, making it difficult to accurately assess complex attention states. Therefore, by calculating and integrating the threshold deviations of various feature parameters, the system can transform multidimensional and heterogeneous state information into a unified, weighted quantitative indicator, providing a data foundation for comprehensive analysis of attention states.
[0123] In some embodiments, the system can pre-define an attention analysis model. This model is pre-trained based on a large amount of attention data from users with different attribute characteristics (such as age group, task type). In actual use, the system automatically adjusts the model parameters based on the target user's baseline behavior data during initial use, as well as the historical attention feature parameter sequences and corresponding state labels (if any) accumulated during subsequent use, through online learning or incremental learning algorithms, thereby constructing and continuously optimizing the attention analysis model for that user. This model takes the threshold deviation of the real-time acquired attention feature parameters as input, and after model inference, outputs an attention state result that more closely matches the user's actual attention state.
[0124] Based on the attention detection method disclosed herein, firstly, multimodal perception data of the target user's physiological state, posture, and electronic device operating status are acquired. These data are then correlated, fused, and extracted to obtain attention feature parameters. This method can identify user attention influencing factors from multiple dimensions, including physiological, behavioral, and device interaction, overcoming the problem that a single data source cannot fully reflect the characteristics. Furthermore, based on the deviation between the attention feature parameters and a dynamic threshold, the attention state result is determined. This method utilizes the continuity of the deviation and the combination of deviations between different attention feature parameters to achieve a refined evaluation of the attention state result, rather than relying on a binary conclusion derived from a preset fixed threshold. This improves the accuracy and resolution of the state result determination. Simultaneously, this method can generate objective calibration criteria to guide threshold adjustment based on multimodal perception data from historical periods. This eliminates threshold mismatches caused by differences in physiological state or scenario, adjusting the threshold to a range that matches the user's individual characteristic patterns. This ensures the adaptability of the deviation calculation between the attention feature parameters and the dynamic threshold, resulting in a more accurate and reliable final attention state evaluation result.
[0125] In some embodiments, regarding the process of adjusting the initial attention threshold based on individual differences in S103 above, such as... Figure 4 As shown, S103 specifically includes steps S201-S204: S201. Based on the attribute information of the target user, determine the target group to which the target user belongs, and obtain the initial attention threshold and baseline attention feature parameters corresponding to the target group.
[0126] Among these, attribute information refers to information used to describe the individual characteristics of the target user, including age, learning stage, quantitative level of learning ability, and type of learning subject. Age and learning stage together reflect the stage of cognitive development and knowledge background, the learning ability level provides a quantitative reference for their ability performance, and the type of learning subject indicates the specific area of their current learning activity.
[0127] Specifically, attribute information reflects the basic situation of target users in terms of cognitive development, knowledge reserves, and task background. Grouping users by attributes is because users with similar attributes exhibit common patterns in their attentional performance. By categorizing target users into groups with similar attributes, the system can provide an initial reference benchmark based on group commonalities, thereby improving the reasonableness of the initial assessment when individual historical data is lacking.
[0128] Among them, the baseline attention feature parameter refers to the statistical quantity obtained by analyzing and processing the historical attention feature parameters of a large number of users in the target group, which is used to characterize the typical level of the group in each dimension of attention.
[0129] For example, baseline attentional characteristics could be the mean or median of historical data for users within the group. For instance, for the "10-12 year old math learning" group, baseline attentional characteristics might include: a mean of 0.75 for writing trajectory smoothness and a variance of 1.2 for posture stability.
[0130] In some embodiments, the system can determine the group to which a target user belongs by querying predefined group segmentation rules. Based on the user's attribute information, the system matches the corresponding group identifier, then accesses the associated data storage area to read the pre-calculated and stored initial attention threshold and baseline attention feature parameters for that group.
[0131] In other embodiments, the system may also generate an initial attention threshold and baseline attention feature parameters in real time through dynamic analysis and processing based on locally stored historical data of anonymous users belonging to the same target group.
[0132] One possible implementation involves the system maintaining a local pool of anonymous group data. When a group baseline needs to be generated, the system filters relevant historical data based on group attributes and analyzes each feature dimension separately. For the initial attention threshold, a quantile covering a specific probability density can be set as the threshold based on a probability distribution model (such as a Gaussian mixture model). For the baseline attention feature parameters, robust central tendency statistics of the historical data, such as the median, can be calculated.
[0133] For example, the composition of the initial attention threshold is shown in Table 3: Table 3
[0134] S202. Perform correlation fusion and feature extraction processing on the second multimodal perception data to obtain the attention feature parameters of the target user in the historical time period.
[0135] Among them, the second multimodal sensing data refers to the heterogeneous data set formed by continuous collection by sensor units and arranged in chronological order within a historical period.
[0136] For the correlation fusion and feature extraction processing of S202, please refer to the above S102 and it will not be repeated here.
[0137] S203. Determine the degree of deviation between the target user's attention feature parameters and the baseline attention feature parameters during the historical period.
[0138] Among them, the degree of feature deviation refers to the quantitative difference between the attention feature parameters of the target user and the baseline attention feature parameters, which characterizes the individual characteristics of the user in terms of attention performance that are different from the group level.
[0139] Specifically, these individual differences stem from two main sources: firstly, the user's inherent physiological characteristics, such as their resting heart rate and respiratory rhythm, which differ from the group baseline; and secondly, their behavioral patterns, such as their study duration, task switching frequency, or adaptability to task difficulty, which differ from the group average. The purpose of calculating the degree of characteristic deviation is to quantify these individual differences.
[0140] In some embodiments, the degree of deviation of a feature may include the following: deviation of physiological rhythm, such as the difference between an individual's average respiratory rate and a baseline; deviation of behavioral endurance, such as the ratio of the rate of decay of an individual's attentional characteristics during a continuous learning period to a baseline; and deviation of task difficulty response, such as the difference between the fluctuation range of an individual's characteristic parameters and a baseline when facing tasks of different difficulty.
[0141] One possible implementation involves the system employing dimensional comparison and standardization to calculate the degree of feature deviation. The system calculates the difference between the statistical value of a user's individual feature parameter and its corresponding benchmark value, then standardizes this difference by dividing it by the standard deviation of the benchmark value, resulting in a dimensionless deviation score. The deviation scores from different feature dimensions collectively constitute a multidimensional feature deviation vector.
[0142] S204. Based on the degree of feature deviation, the initial attention threshold is adjusted to generate the target attention threshold.
[0143] The first adjustment is the process of adjusting the initial attention threshold based on the individual differences of the target users.
[0144] Specifically, the system uses the calculated degree of feature deviation to identify the inherent characteristics of users that differ from the average level of the group in terms of physiological baseline, behavioral endurance, and task response patterns. It then uses this as the core basis to adjust the initial threshold, making the threshold setting more in line with the user's personal normal state.
[0145] In some embodiments, the first adjustment includes the following strategies: a compensation adjustment strategy based on the degree of decline in the target user's physiological state, a compensation adjustment strategy based on the working time of the electronic device, and an adaptation adjustment strategy based on the difficulty of the learning task performed by the electronic device.
[0146] Among these strategies, the physiological state-based compensation adjustment strategy involves the system relaxing relevant thresholds based on the user's physiological fatigue trend. The work duration-based compensation adjustment strategy involves the system relaxing the judgment criteria according to the degree of attenuation based on device usage time. The task difficulty-based adaptation adjustment strategy involves the system identifying the task difficulty level and dynamically adjusting the strictness of the thresholds to distinguish between focused thinking and inattentiveness.
[0147] The degree of deviation from the baseline determines the individual's physiological baseline offset. When a further decrease in heart rate variability from this baseline is detected, the system proportionally relaxes the body rate threshold based on this degree of deviation.
[0148] For example, if a user's heart rate variability is -0.2 (the individual baseline is 20% lower than the group baseline), when a further 10% decrease in heart rate variability is detected, the system will relax the body movement frequency threshold from 2 beats / minute to 2.4 beats / minute.
[0149] One possible implementation involves a compensation adjustment strategy based on working duration. The system adjusts the attenuation of attention feature parameters according to preset rules based on the continuous working duration obtained from operational status data. Simultaneously, the attenuation magnitude is personalized by scaling the adjustment based on the degree of feature deviation.
[0150] For example, when the continuous learning time reaches 45 minutes, the system adjusts the writing error rate threshold by a basic decay based on preset rules (e.g., increasing it from 15% to 17%). At the same time, if the user's writing stability characteristic deviates by -0.3 (stability is better than the baseline by 30%), the final threshold will only be adjusted to 16.4%.
[0151] One possible implementation is that, for an adaptation strategy based on task difficulty, the system obtains a reference value corresponding to the standard difficulty from the baseline attention feature parameters and makes adaptation adjustments based on the degree of deviation of user features.
[0152] For example, in high-difficulty task scenarios, the standard benchmark sets the threshold for the duration of head stillness to 70 seconds. If a user's focus characteristic deviates by +0.25 in this task (focus performance is 25% better than the benchmark), the system will adjust the threshold to 77.5 seconds to match their higher focus tolerance.
[0153] It should be understood that S201-S204 quantifies individual and group differences by the degree of feature deviation, and dynamically adjusts the threshold by combining real-time physiological state, working hours and task difficulty, thereby realizing the adjustment from group benchmark to personalized threshold, so that the attention threshold can be adapted to the target user's personal physiological baseline, behavioral endurance characteristics and contextual cognitive load pattern.
[0154] In some embodiments, the scene adaptive adjustment process in S103 described above, based on the user's current task scenario using the electronic device, is as follows: Figure 5 As shown, S103 specifically includes steps S301-S303: S301. Based on the runtime context information, perform semantic recognition to determine the usage scenario type of the electronic device.
[0155] The runtime context information includes: the interaction information of the foreground application of the electronic device and the file format of the data loaded by the foreground application. This information reflects the identity of the currently running task, the user's operation behavior, and the data characteristics of the processed content, and is used to identify the currently executed task content.
[0156] Specifically, the scenario type is used to indicate the cognitive attributes and behavioral patterns of an activity. Different scenarios have different requirements for attentional resources; for example, "deep reading" requires physical stability, while "interactive problem-solving" allows for more actions. Therefore, a uniform threshold cannot accurately assess attention across scenarios. By identifying scenario types, the system provides contextual basis for threshold adjustment, enabling the evaluation criteria to dynamically adapt to scenario characteristics and improve the reasonableness and effectiveness of the results.
[0157] In some embodiments, the system employs multi-source data fusion for identification. The system analyzes runtime context information (such as file type .ppt / .pdf, application name) and integrates micro-motion characteristic data from millimeter-wave radar arrays (such as low head micro-motion frequency during reading and frequent hand movements during problem-solving) to assist in judgment, thereby more accurately identifying the usage scenario type.
[0158] One possible implementation involves constructing a multimodal scene recognition model. This model uses runtime context information and radar-sensed user posture micro-motion features as joint inputs. By learning the correlation between device tasks and user behavior in different scenarios, it outputs a comprehensively determined scene type.
[0159] For example, the system detects that the foreground is "programming learning software," loads .py files, and the interaction is mainly based on keyboard input and debugging operations. At the same time, radar detects high-frequency, regular hand movements in the keyboard area. According to the multimodal model, this kind of evidence is collectively mapped to an "interactive practice scenario."
[0160] S302. Based on the usage scenario type and the preset mapping relationship between the scenario and the threshold adjustment strategy, determine the threshold adjustment strategy for the second adjustment of the initial attention threshold.
[0161] The scenario type represents the user's cognitive needs and typical behavioral patterns for the current task. The pre-defined mapping relationship between scenarios and threshold adjustment strategies defines the specific threshold adjustment logic to be used in different scenarios. The second adjustment is a further adaptation correction based on the scenario type to the first attention threshold that has undergone preliminary personalized correction.
[0162] For example, the mapping relationship between the preset scenario and the threshold adjustment strategy can be specifically represented by a mapping table, as shown in Table 4.
[0163] Table 4
[0164] One possible implementation is that the system determines the threshold adjustment strategy by querying a pre-set mapping table. The system maintains a mapping table locally, as shown in Table 4. After the system determines the current usage scenario type (e.g., "mathematical problem solving") in S301, it performs a matching query in this mapping table to determine the threshold adjustment strategy corresponding to the current usage scenario type.
[0165] S303. Based on the threshold adjustment strategy, the first attention threshold is adjusted a second time to generate the target attention threshold for the target user.
[0166] Specifically, the second adjustment is the process of adjusting the initial attention threshold based on the type of use case.
[0167] One possible implementation is that the system parses the threshold adjustment strategy into mathematical operations or numerical substitution instructions. For example, if the strategy is "reduce the writing pressure volatility threshold to 20%", the system updates the first threshold to this value; if the strategy is "increase the head rotation frequency threshold by 50%", the system calculates the new threshold = the original threshold. 1.5. By executing such instructions, a target threshold for scene alignment is generated.
[0168] In other embodiments, the system employs a dynamic fusion method of policy weights. When a policy contains multiple composite instructions, the system transforms each policy into an influence weight and adjustment direction on a specific threshold, and performs a comprehensive adjustment of the first threshold through weighted calculation, outputting a target threshold set that balances the requirements of multiple dimensions.
[0169] It should be understood that S301-S303, by identifying the device usage scenario type and dynamically adapting the already individually corrected first attention threshold based on scenario characteristics, achieves a key shift in attention assessment standards from static and general to dynamic and scenario-specific. This mechanism ensures that the final generated target attention threshold not only reflects the user's personal baseline but also accurately matches the specific attention requirements of the current task, thereby significantly improving the contextual rationality and accuracy of attention state assessment in different learning activities.
[0170] In some embodiments, regarding the process of determining the attention state based on the threshold deviation described in S104, the system can also introduce the current usage scenario type as a prompt during the model inference process, enriching the input information of the attention analysis model, thereby enabling the attention state evaluation to achieve refined decision-making based on scenario perception. This process is as follows: Figure 6 As shown, S104 may specifically include steps S401-S404: S401. Determine the degree of deviation between each attention feature parameter and the corresponding target attention threshold.
[0171] One possible implementation is that the system determines the degree of threshold deviation by calculating the relative deviation. For each attention feature parameter, the system compares its real-time measurement with the corresponding target attention threshold and calculates the relative difference between the two (e.g., (real-time value - target threshold) / target threshold) to obtain a standardized deviation value.
[0172] For example, when the target threshold for “writing pressure volatility” is 20% and the real-time measurement is 28%, the degree of deviation from the threshold is calculated as (28%-20%) / 20% = 0.4 (i.e., a positive deviation of 40%).
[0173] S402. Determine the usage scenario type of the electronic device.
[0174] For details on the specific implementation process of this step, please refer to S301 above, which will not be repeated here.
[0175] S403, construct the threshold deviation as a feature combination vector and the usage scenario type as a prompt message.
[0176] The feature combination vector refers to the structured numerical vector formed by arranging and standardizing the threshold deviations of each attention feature parameter calculated in S401 according to a preset dimensional order.
[0177] The prompt information refers to the conditional input formed after encoding and converting the electronic device usage scenario type determined in S402. The prompt information provides high-level semantic guidance on the current task background for subsequent analysis models, and its form may include one-hot encoded vectors, word embedding vectors, or pre-trained scene semantic feature vectors, etc. This application embodiment does not limit the specific form of the prompt information.
[0178] One possible implementation involves the system constructing two types of information through a feature vectorization module and a scene encoder, respectively. The feature vectorization module receives the threshold deviation values of all features, performs normalization to eliminate dimensional differences, and then concatenates them into a one-dimensional numerical vector in a fixed order. Simultaneously, the scene encoder maps the scene type name (such as "mathematical problem solving") to a low-dimensional dense vector by querying a pre-trained word embedding table or scene feature mapping table.
[0179] For example, the system calculates the heart rate variability deviation as 0.1, the writing pressure deviation as -0.2, and the head rotation deviation as 0.3. After min-max normalization, these deviations are concatenated in a predetermined order to obtain the feature combination vector [0.6, 0.2, 0.8, ...]. Simultaneously, the scene encoder maps "deep reading" to a cue vector [0.12, -0.05, 0.33, 0.21, ...].
[0180] S404. By using the attention analysis model and processing feature combinations based on prompt information, the attention state results of the target user are obtained.
[0181] The attention analysis model is trained through at least two training phases: a pre-training phase based on different user attention data and a reinforcement training phase based on target user attention data.
[0182] Specifically, the attention analysis model processes feature combination vectors and cue information, with its core being the construction of a decision function conditioned on the scenario. The model uses the cue information as a conditional variable and learns the joint probability distribution P(state|feature vector.scenario), allowing the decision boundary to dynamically adjust with the scenario. This achieves an explicit dependence of state decisions on the task context, ensuring that the output conforms to both data characteristics and scenario consistency.
[0183] Specifically, the two-stage training aims to balance model generality and individual adaptability. Pre-training learns common patterns from cross-user data to establish a baseline model. Reinforcement training uses historical data from the target user to fine-tune the model, adapting it to the user's unique physiological baseline, behavioral patterns, and attentional patterns, thereby improving the accuracy and reliability of individual state determination.
[0184] In some embodiments, the attention analysis model can be various machine learning models capable of processing structured feature vectors and conditional information. For example, it can be a gradient boosting decision tree model that processes feature combination vectors and prompts through feature concatenation or conditional branching, or a random forest model that processes feature combination vectors and prompts through feature grouping and conditional weighted voting.
[0185] One possible implementation involves employing a conditional augmentation model based on gradient boosting decision trees. During pre-training, the model concatenates scene cue encodings with feature combination vectors and trains a tree set on cross-user data to learn a general discriminant function. In the reinforcement training phase, the model is incrementally trained using target user data, or the final tree layers are fine-tuned to adjust the output values of leaf nodes, achieving personalized correction. During inference, the concatenated vector is input into the model, and the final attention state result is obtained through the ensemble output of the tree set.
[0186] For example, in a "mathematical problem-solving" scenario, after receiving the scenario's prompt, the model assigns higher decision weights to the threshold deviation of features such as "frequency of writing interruptions." Therefore, even if the deviation of other features (such as head turning frequency) is low, if the deviation of "frequency of writing interruptions" is significantly positive, the model may still determine it as "distracted attention." Conversely, in a "Chinese reading" scenario, the model reduces its sensitivity to the deviation of features such as "frequency of head turning." Even if there is some positive deviation, as long as the deviation of features such as "frequency of content interaction" remains within a good range, the model may still determine it as "focused attention."
[0187] It should be understood that, through S401-S404, the system constructs an attention state decision-making mechanism based on scene conditions, using the usage scene type as a cue for the attention analysis model. This mechanism dynamically adjusts the classification boundary of attention states according to the input scene information by learning the joint distribution of features and scenes. This ensures that the model can output attention state judgments that are consistent with the cognitive requirements of the current task, effectively solving the evaluation bias problem caused by ignoring task context under a fixed threshold system, thereby improving the accuracy of attention detection in different scenes.
[0188] In some embodiments, after determining the attention state result, the system can further apply attention intervention strategies to the target user based on the attention state result, thereby transforming the evaluation result into a closed-loop regulatory behavior to improve the user's attention maintenance and learning efficiency. In this case, the method further includes steps S501-S502: S501. When the target user's attention state result is less than or equal to the preset attention state result threshold, determine the multimodal attention intervention strategy and the order of execution of the intervention strategy for the target user based on the attention state result, the degree of feature deviation and the type of use scenario.
[0189] Among them, multimodal attention intervention strategies include at least one of the following: auditory intervention strategies, visual intervention strategies, tactile intervention strategies, and guided intervention strategies.
[0190] The execution sequence of intervention strategies indicates the triggering priority and combination logic of different intervention strategies, aiming to effectively guide users' attention back in a gradual and coordinated manner.
[0191] Specifically, attention state results provide the overall intensity requirement for the intervention; the degree of feature deviation reveals the specific behavioral or physiological dimensions that lead to decreased attention (such as writing interruption or postural instability), thus guiding the selection of targeted intervention modalities (such as triggering tactile reminders for postural instability); and the type of use scenario constrains the applicability and interference limits of the intervention method (such as avoiding the use of strong auditory interventions in scenarios requiring quiet listening). These three types of information work together to ensure that the intervention strategy is targeted, effective, and appropriate for the scenario.
[0192] In some embodiments, auditory intervention strategies include one or more of the following: playing preset reminder audio through an electronic device and adjusting the audio playback parameters of the electronic device.
[0193] For example, the system controls electronic devices to emit different types of sound alerts, such as gentle prompts or interesting animated voices, to attract the user's attention in a non-intrusive way.
[0194] Visual intervention strategies include one or more of the following: playing preset reminder videos or images through electronic devices, adjusting the screen display parameters of electronic devices, and converting the displayed content of electronic devices into audio or video formats for playback.
[0195] For example, the system can display personalized encouraging images or gamified learning tasks on electronic device screens to attract users' attention; or dynamically adjust the presentation of learning content, such as converting text content into more vivid forms like animation or audio for playback.
[0196] Tactile intervention strategies include one or more of the following: providing vibration alerts through the vibration module of electronic devices and adjusting the intensity or frequency of interactive vibrations of electronic devices.
[0197] For example, the system can provide vibration alerts through the vibration module of electronic devices and can dynamically adjust the intensity or frequency of the vibration based on the user's status.
[0198] The system executes pre-defined guided interactive tasks via electronic devices. These tasks consist of multiple guided steps, which instruct the target user to perform interactive actions.
[0199] For example, the system can launch a short gamified interactive task on the screen of an electronic device (such as clicking to follow a moving target) and guide the user through the operation with step-by-step instructions, thereby helping them refocus on the learning content.
[0200] It should be understood that the aforementioned multimodal intervention strategy constructs a non-invasive attention guidance method by integrating multiple sensory and interactive channels such as auditory, visual, tactile, and guided interaction. This design enhances the perceptual reliability of the intervention signal through multi-channel complementarity.
[0201] In some embodiments, the system can employ a rule-matching and priority-scoring approach to determine the final intervention strategy and execution order. Specifically, this can be implemented as follows: 1. Based on the preset mapping relationship between state results and priorities, determine the attention intervention priority corresponding to the attention state results.
[0202] Among them, attention intervention priority is a level indicator used to quantify the urgency of intervention required for attention state outcomes. It reflects the degree to which the current attention state deviates from the ideal level, and a higher level usually indicates that more immediate or intensive intervention is needed.
[0203] Multimodal attention intervention strategies refer to a complete approach that combines interventions through different sensory channels, such as auditory, visual, tactile, and guided interaction, to guide users' attention back. It represents a series of specific actions that computing devices are planned to take.
[0204] 2. Determine the multimodal attention intervention strategy for the target user from the preset intervention priority and intervention strategy mapping table.
[0205] 3. Based on the degree of feature deviation and the type of use scenario, the set of intervention strategies is sorted by importance to determine the execution priority of each intervention strategy in the multimodal attention intervention strategy and generate the execution order of intervention strategies.
[0206] One possible implementation is that the attention intervention priorities, arranged from lowest to highest, include: a first attention intervention priority, a second attention intervention priority, and a third attention intervention priority. The first attention intervention priority corresponds to visual intervention strategies; the second attention intervention priority corresponds to both visual and auditory intervention strategies; and the third attention intervention priority corresponds to both tactile and auditory intervention strategies. This application does not limit the priority division method or the corresponding intervention strategies. In other implementations, more or fewer priority arrangements or more corresponding intervention strategies can be implemented.
[0207] The execution sequence of the intervention strategy defines the order in which the various sub-strategies or actions within the determined multimodal attention intervention strategy are executed. It reflects the sequential logic and synergistic relationship between different intervention methods triggered in a composite intervention program.
[0208] Specifically, through three progressive steps, the abstract attention state results, individual characteristic differences, and environmental context are transformed into concrete, ordered, and adaptive action instructions. First, the computing device maps a basic intervention urgency (priority) based on the severity of the state results. Then, a preliminary set of strategies is selected based on this priority. Finally, by combining the user's personalized feature deviation patterns with the current task scenario, the strategy set is finely sorted and adjusted to generate the final execution sequence.
[0209] One possible implementation involves the computing device first performing priority mapping and initial policy screening. The computing device internally maintains a mapping table between state results and attention intervention priorities. Upon obtaining an attention state result, the computing device queries this table and maps it to a specific priority level. Subsequently, the computing device accesses another pre-defined mapping table between intervention priorities and intervention policies, and based on the newly determined priority, retrieves one or more multimodal attention intervention policies associated with that level, forming an initial policy set.
[0210] For example, the mapping table might specify that the "severely distracted" state is mapped to "priority 3," and "priority 3" is associated with policies such as "strong haptic vibration alert" and "forced interactive task" in the policy mapping table. The computing device then incorporates these policies into the initial set.
[0211] Another possible approach is that, when determining the final intervention strategy and execution order, the computing device can introduce a secondary sorting mechanism based on personalized features and scenarios, in addition to the initial screening.
[0212] One possible implementation involves the computing device then performing a refined ranking of features and scenarios. The device acquires the degree of feature deviation representing individual and group differences among users, as well as the current usage scenario of the electronic device. The device evaluates each intervention strategy in the initial strategy set according to a set of predefined ranking rules. These rules consider the dominant distraction dimension revealed by the degree of feature deviation (e.g., increasing the ranking weight of tactile interventions if feature deviation indicates poor posture is the primary cause), and the limitations imposed by the usage scenario type on the intervention method (e.g., decreasing the weight of loud auditory interventions in a "library reading" scenario). By calculating a comprehensive score for each strategy, the device ranks all strategies by importance, thereby determining the final execution priority of each strategy and generating the intervention strategy execution order accordingly.
[0213] For example, even if the "strong alert tone" and "screen flashing" strategies are initially selected under the "severe distraction" priority, if the current scenario is identified as an "online meeting", the computing device will significantly reduce the score of the "strong alert tone" according to the scenario rules, making it lower in the final execution order or not adopted at all, and may instead increase the priority of the "private haptic reminder" strategy.
[0214] It should be understood that by breaking down the decision-making process into two stages—"initial priority screening" and "personalized scenario-based fine-tuning"—this method first ensures that the intensity of intervention is basically matched with the severity of the problem. Then, by incorporating individual user characteristics and real-time scenario information, it achieves precise customization and optimization of intervention strategies. This effectively avoids a single, rigid strategy matching logic, improves the pertinence, acceptability, and ultimate effectiveness of intervention measures, and makes attention guidance support more intelligent and humanized.
[0215] The system maintains a rule base, which defines the range of different state outcomes, deviation patterns of specific features (such as "high-frequency gaze wandering"), and applicability scores and combination rules for various intervention strategies under different scenario types. The system matches the current user's data with the rule base, calculates and sorts the comprehensive scores of each candidate strategy, and generates an ordered sequence of intervention strategies.
[0216] S502. Execute multimodal attention intervention strategies based on the order of intervention strategy execution to improve the attention state of the target user.
[0217] Specifically, the system executes predefined strategy logic, transforming abstract strategy instructions into specific control commands for electronic devices and their associated components, thereby achieving targeted intervention of the user's senses and guiding the user's attention state to improve in the expected direction.
[0218] One possible implementation involves the system sequentially executing intervention strategies through a strategy execution engine. This engine reads the strategy sequence determined by S501, parses each strategy item in turn, and generates corresponding device control instructions based on the strategy type and parameters. For example, for the "auditory intervention: play a gentle prompt tone" strategy, the engine will call the audio module and specify the audio file and volume parameters; for the "visual intervention: display an encouraging image" strategy, the engine will call the display module and load the specified image resource. Through this mechanism, the system transforms the logical sequence of intervention strategies into a series of ordered, hardware-executable operations.
[0219] In other embodiments, the system may also employ a feedback adjustment mechanism based on strategy effectiveness. Specifically, after the intervention strategy is executed, the system will evaluate the actual improvement effect of the strategy on the target user's attention state in real time or near real time, and dynamically adjust the execution parameters, sequence, or whether to trigger additional interventions for subsequent intervention strategies based on the evaluation results.
[0220] One possible implementation method, the specific process of the feedback adjustment mechanism is as follows: obtain the first attention state result of the target user after the execution of the multimodal attention intervention strategy; determine the attention improvement value of the first attention state result relative to the attention state result before the execution of the multimodal attention intervention strategy; if the attention improvement value is less than or equal to the attention improvement threshold, redetermine and execute the multimodal attention intervention strategy.
[0221] In an exemplary embodiment, the implementation process of this method is as follows: Figure 7 As shown, specifically, the process includes the following steps: Implement multimodal attention intervention strategies: The system executes the multimodal attention intervention strategy sequence determined by S501, such as playing prompts and displaying encouraging animations in sequence.
[0222] Obtaining the first attention state result: After the intervention strategy is completed, the system immediately or after a short delay re-collects the user's multimodal perception data, and then processes it again through the S101 to S104 process flow to obtain the attention state evaluation result after the intervention, i.e., the first attention state result.
[0223] Determine the attention improvement value: The system compares the first attention state result with the original attention state result before the intervention, and quantifies the intervention effect through preset calculation rules (such as difference calculation, ratio calculation or mapping calculation based on state level) to obtain the attention improvement value.
[0224] Attention boost value less than or equal to attention boost threshold: The system compares the calculated attention boost value with the preset attention boost threshold.
[0225] If so, it is determined that the intervention did not achieve the expected results, and the attention intervention strategy should be redefined.
[0226] If not, stop distracting yourself.
[0227] Redefine attention intervention strategies: Go through S501 again to determine new attention intervention strategies and execution order, and re-execute the multimodal attention intervention strategies.
[0228] For example, the system detects that the user's attention state is "mildly distracted" and executes a preset intervention sequence accordingly: first playing a prompting sound, then displaying an encouraging animation. After the intervention, the system immediately re-evaluates and obtains a new attention state result of "mildly distracted," with the calculated improvement value close to zero, below the effective improvement threshold. At this point, the system determines that the initial intervention was ineffective and therefore makes a new decision, generating and executing a second intervention strategy that includes stronger tactile vibrations and guided interactive tasks, aiming to achieve a better improvement.
[0229] It should be understood that the feedback adjustment mechanism creates a dynamic closed loop in the intervention process. The system can evaluate the intervention effect based on the user's real-time response, and adaptively adjust or strengthen the strategy when the effect does not meet expectations, thereby improving the accuracy and success rate of the intervention and ensuring effective improvement in attention status.
[0230] It should be understood that, through S501-S502, the system achieves a closed-loop decision-control process from attention state assessment to personalized intervention execution. This mechanism dynamically generates and executes an appropriate sequence of multimodal intervention strategies based on real-time assessment results, the degree of deviation from specific characteristics, and current scenario information. This process maps abstract attention states into concrete, executable environmental control instructions, achieving automated connection from perception to intervention, thereby improving the accuracy and timeliness of attention control and optimizing the overall system performance through a closed-loop optimization.
[0231] In some embodiments, in step S501 above, the system can also proactively optimize the intervention strategy based on the temporal trend of the user's historical attention state. S501 specifically includes the following steps S601-S604: S601. Obtain the sequence of attention state results for the target user.
[0232] The attention state result sequence includes attention state results determined at different times.
[0233] Specifically, the system records and stores the user's attention state results (such as focus scores or discrete state labels) generated periodically or event-triggered by step S104 within a preset time window (e.g., the past 15-30 minutes), and sorts them by timestamp to form a sequence. This sequence provides a time-series data foundation for analyzing the user's attention change patterns.
[0234] S602. Process the attention state result sequence through the attention state transition prediction model to obtain the trend of the target user's attention state change in the future.
[0235] The attention state transition prediction model is a trained time-series prediction model used to predict future state evolution based on historical state sequences. The attention state change trend includes a quantitative prediction of the probability of different attention state outcomes occurring at various future time points.
[0236] Specifically, the system inputs the acquired state sequence into the predictive model. By analyzing patterns in the sequence (such as state persistence, transition probability, and decay rate), the model outputs predictions about the user's future attention state, such as "attention will continue to decline," "the user will enter a slightly distracted state in one minute," or "the state will tend to stabilize." This forward-looking prediction provides a crucial basis for predictive intervention.
[0237] In some embodiments, the prediction model can be based on an LSTM architecture. This model uses historically accumulated attention state result sequences as training data to learn the temporal transition patterns between states. In application, the model receives recent attention state result sequences of the target user as input, and through its internally learned sequence evolution patterns, outputs a prediction of the attention state sequence for a future period (such as the next 5 minutes), thereby obtaining the specific trend of attention state changes.
[0238] In other embodiments, the prediction model may also be built based on a hidden markov model (HMM).
[0239] Hidden Markov Models (HMMs) are statistical models used to model time-series data. Their core idea is to assume that a discrete sequence of states (i.e., hidden states) that cannot be directly observed evolves according to Markov properties and can be inferred from a series of probabilistically related observational data sequences. In this application, HMMs are used to model and predict the evolution of a user's attention state, and their parameters are defined as follows: Hidden state set Q: The set of attention state results corresponding to the system, for example Q={S1, S2, S3, S4, S5}, which represent different state levels such as "high focus" and "moderate focus".
[0240] The observation state set O corresponds to the feature vector space formed by the attention feature parameters extracted from the multimodal perception data.
[0241] Model parameters: The system obtains the following parameters based on historical data training: setting the initial state probability distribution. This represents the probability that the user is in each attention state at the initial moment; the state transition probability matrix A has elements... =P( = | = Defines the state from the current state. Transition to state The pattern; the observation probability matrix B, its elements =P( = | = The definition is given for when the object is in a hidden state. The probability of observing a specific feature vector o.
[0242] Specifically, the process by which the system obtains the trend of attention state changes through the trained HMM is as follows: Input observation sequence: Input the recent attention feature parameter sequence (i.e., observation sequence) of the target user sorted by time into the model.
[0243] Decoding the most likely state path: Using the Viterbi algorithm, the hidden state sequence most likely to produce the observation sequence is calculated based on the model parameters (π, A, B), which is the evolution path of the user's attention state over a period of time.
[0244] Predicting future trends: Based on the decoded most likely hidden state at the current moment and according to the state transition probability matrix A, the probability distribution of the user's state in each hidden state after one or more time steps can be calculated, or the probability of future observation sequences can be directly calculated through a forward algorithm. This probability distribution or the most likely state sequence is the predicted trend of attention state changes. For example, the model can output "The probability that the user will transition from the current S3 state to the S4 state within the next 2 minutes is greater than 60%".
[0245] For example, when modeling and predicting user attention states based on the above HMM model, each hidden state (attention level) can be associated with typical multimodal observation feature patterns to assist model training and result interpretation, as shown in Table 5.
[0246]
[0247] In an exemplary embodiment, such as Figure 8 As shown, the system constructs a state transition network using five discrete attention states (S1: high focus, S2: moderate focus, S3: general focus, S4: pre-distraction aura, S5: severe distraction) as nodes. Specifically, it includes: The S1 state can transition to the S2 state due to a decline in physiological state.
[0248] The S2 state can be downgraded and transitioned to the S3 state as the task difficulty increases.
[0249] State S3 can transition to state S4 due to increased cumulative duration. Alternatively, applying attention intervention 3 can transition to state S1.
[0250] State S4 can transition to state S5 without attentional intervention. Alternatively, state S2 can be transitioned to state S2 with attentional intervention 2.
[0251] The S5 state can be shifted to S3 by applying attention intervention 1.
[0252] Among them, attention intervention 1, attention intervention 2, and attention intervention 3 represent different types, intensities, or systemic intervention strategies targeting different sources of distraction.
[0253] It should be understood that the transition conditions marked in the diagram (such as "decline in physiological state") are operational summaries of changes in multimodal perception data patterns or system intervention actions. They are logically mapped to multimodal perception data, attention feature parameters, and intervention strategies, providing an interpretable operational expression for the HMM transition probability matrix. The diagram illustrates the directionality of state transitions (degeneration driven by negative conditions and recovery driven by intervention), the stages of state evolution (the progressive process from S1 to S5), and shows that external intervention is a key controllable node that interrupts negative evolution and guides the state towards a positive direction.
[0254] In some other embodiments, the prediction model can be a hybrid attention prediction model that combines LSTM and HMM in parallel, such as... Figure 9 As shown, this hybrid prediction model employs a dual-path parallel processing architecture: The model receives the same time-series data as input, which is a sequence of attention feature parameters (i.e., a set of attention feature parameters arranged in chronological order, which may include quantitative indicators of multiple dimensions such as physiology, posture, and device interaction) obtained by associating, fusing and extracting features from multimodal perception data according to the method described in S102.
[0255] LSTM Path. The sequence is directly input into the LSTM network. Through its internal gating mechanism, the LSTM learns the complex long-term dependencies and nonlinear dynamic evolution patterns in the feature sequence, and outputs LSTM temporal features.
[0256] The Hidden Markov Model (HMM) pathway involves simultaneously inputting the same feature sequences into a pre-trained model. The HMM uses a forward-backward algorithm to calculate the posterior probability distribution for each hidden attention state (e.g., S1 to S5) at each time point, outputting the HMM state posterior probability feature. This feature exhibits good robustness to noise and transient anomalies in the data.
[0257] Feature fusion. The system fuses the LSTM temporal features with the HMM state posterior probability features along the feature dimension (e.g., by concatenation or weighted combination) to form a more comprehensive and robust fused feature vector.
[0258] Final prediction. The fused feature vector is input into a prediction layer (e.g., a fully connected neural network classifier or regressor). The prediction layer combines the advantages of both types of features and outputs a prediction of the target user's attention state trend over a future period, which is more robust and accurate than a single model.
[0259] It should be understood that this architecture achieves complementary advantages and improves prediction performance by extracting and fusing the continuous dynamic modeling capability of LSTM and the discrete state probability inference capability of HMM in parallel.
[0260] S603. Based on the trend of attention state changes, attention state results, usage scenario type, and degree of feature deviation, determine the multimodal attention intervention strategy and the order of execution of the intervention strategy for the target user.
[0261] One possible implementation involves the system maintaining a state change trend-intervention strategy mapping table (as shown in Table 6). This table defines the warning levels and intervention strategies corresponding to different state prediction probability conditions. The system performs a dual-path decision: matching the predicted trend with the table to generate a set of trend-suggested strategies; simultaneously generating a set of current-state-suggested strategies based on attention state results and feature deviation degrees (as in the embodiment in S501 above). The two sets are then fused, using either an intersection or union mode, and the execution intensity and priority are adjusted according to the warning level. Finally, the fused strategies are sorted based on the usage scenario type to generate the execution order.
[0262] Table 6
[0263] It should be understood that, through S601-S603, the system introduces predictive analysis capabilities based on time-series state sequences. This mechanism, by establishing a model of the user's attention state evolution, achieves a crucial shift from passively responding to the current state to proactively predicting future trends. This enables the system to identify potential attention decay risks in advance, thereby shifting the intervention timing from "after the problem occurs" to "before the problem occurs," providing a basis for decision-making regarding preventative and guiding interventions, and significantly enhancing the proactivity and foresight of attention maintenance support.
[0264] In other embodiments, the attention state results, in addition to indicative of decisions regarding attention intervention strategies, can also be used by the system to instruct the electronic device to switch learning task modes, adapting the current learning task to the current user's attention state. Specifically, the process is as follows: Figure 10 As shown, it specifically includes: The system first quantifies the acquired attention state results (such as the output of S104) into an attention level (e.g., discretizing continuous values according to a preset interval). Then, based on this attention level, the system enters a pattern decision process: First, obtain the attention detection results and quantify them into an attention level (e.g., discretize continuous values according to a preset interval).
[0265] Secondly, based on attention levels, the learning operation mode of electronic devices is determined, including: Attention level ≤ 1 (highly focused) corresponds to the extended mode, which pushes extended learning tasks to the user; Attention level 2-3 (medium stability) corresponds to the reinforcement learning mode, which pushes standard courses to users; For users with an attention level of ≥4 (significantly distracted), the corresponding review mode will push lightweight review content such as knowledge cards to the user.
[0266] Then, schedule timed checks to monitor task completion for each mode.
[0267] One possible implementation involves the system specifically detecting the interactive behavior and status logs of electronic devices, including recording task start / end time, duration of stay on a single question / page, frequency of answer submission, number of times content is jumped and applications are switched out, collecting user operation behavior characteristics, and combining them with real-time attention status feedback to determine whether attention is continuously maintained within the target range of the current mode during task execution.
[0268] Then, an effectiveness evaluation is conducted based on the detection results of task completion.
[0269] One possible implementation is that the system can comprehensively evaluate the performance based on multi-dimensional behavioral data and changes in attention state of electronic devices, including calculating task completion, analyzing attention stability, assessing learning quality, and combining negative behavioral feedback to determine whether the mode execution effect is effective.
[0270] If the effectiveness evaluation of the currently determined operating mode is valid, the current mode shall be maintained.
[0271] If the effectiveness evaluation indicators of the currently determined operating mode are invalid, adjust the current mode. For example, adjusting the task difficulty and presentation within the same mode, or switching to another mode based on a reassessment of attention levels after certain conditions are met.
[0272] In some embodiments, to address the potential attention distraction that may occur when users perform split-screen or multitasking operations on electronic devices (e.g., simultaneously opening a learning application and watching a video), the system can manage the foreground application after determining the attention state and identifying the usage scenario type. In this case, the method further includes steps S701-702: S701. Obtain the interaction information of the foreground application in the running context information, and determine the list of foreground applications running on the electronic device from the interaction information of the foreground application.
[0273] Interaction information includes data that characterizes the foreground state of an application, such as window stacking order, input focus, and the process to which user interaction events belong. The foreground application running list refers to a set of identifiers identifying one or more applications currently running simultaneously in the device's foreground (visible and interactive) by analyzing the above information.
[0274] S702. If the list of running foreground applications meets the preset conditions and the change in attention feature parameters is less than the change threshold, close or pause the unwanted applications in the list of running foreground applications.
[0275] The preset conditions include: the presence of at least one desired application and at least one unwanted application in the foreground application running list. Desired applications are those that meet the preset learning objectives and are included in the application whitelist (such as learning software and document tools). Undesired applications are those that may interfere with concentration (such as audio-visual entertainment and social applications). These preset conditions can identify whether the user is currently in a split-screen or multitasking state.
[0276] Specifically, if the change in attention characteristic parameters is less than the threshold, it indicates that the user is in a low-involvement "shallow multitasking" state, such as relaxed body posture, non-interactive micro-movements of the hands, and frequent but irregular small-scale head movements between the corresponding areas of the multitasking window. In this case, the system will close or pause unintended applications to guide a return to focus. If the change is greater than or equal to the threshold, and is accompanied by task-related positive posture adjustments (such as leaning forward when solving problems) and purposeful micro-movements (such as writing or clicking), it indicates that the user may be effectively handling multitasking, and the system will maintain the status quo without forced intervention.
[0277] One possible implementation, S701-S702, is as follows: The system extracts runtime context information from the device's runtime status data, analyzes the process and window states of this information, identifies the process identifiers of all currently running applications in the foreground, and thus obtains a list of running foreground applications. Subsequently, the system queries a preset application category list (which pre-labels applications as "expected" or "unexpected") and performs category matching and labeling for each application in the running list.
[0278] When the system detects that the running list contains at least one desired application and at least one unwanted application simultaneously, it determines that it has entered split-screen multitasking monitoring mode. In this mode, the system synchronously retrieves the user's attention feature parameters (such as micro-motion features related to posture and writing) within a preset time window and calculates their change magnitude. If the change magnitude is lower than a preset threshold, and combined with specific behavioral characteristics (such as relaxed posture and lack of task-related micro-motions), the system determines that the user is in an inefficient "shallow multitasking" state and automatically triggers resource management strategies to close, pause, or forcibly switch the identified unwanted applications to the background.
[0279] It should be understood that S701-S702 addresses the issue of unwanted applications interfering with user attention during split-screen operations. This mechanism first identifies the split-screen scenario using preset conditions, then analyzes the fluctuation levels of attention characteristic parameters to confirm whether unwanted applications actually cause attention distraction. Based on this, restricting unwanted applications can reduce attention interference caused by parallel non-task applications.
[0280] In some embodiments, to improve the group adaptability of the attention detection model and the remote superviseability of the evaluation results, the system can synchronize and remotely share the attention data generated locally, specifically implemented as S801 and / or S802: S801. Synchronize the attention data of the target users to the server so that the server can update the initial attention threshold and / or baseline attention feature parameters of the target group based on the attention data of multiple users in the target group.
[0281] The attention data includes one or more of the following: the target user's identification identifier, first multimodal perception data, attention feature parameters, and attention state results.
[0282] The server can be a distributed cluster consisting of one or more virtual server instances deployed on a cloud computing platform. This server provides elastic computing and storage resources, and is responsible for receiving and storing anonymized attention data from various attention detection systems (usually bound to electronic devices), and performing aggregation analysis of group attention data and calculation of model parameters.
[0283] S802. Synchronize the target user's attention data to the device terminal of the target user's associated users.
[0284] Among them, the device terminal is used to provide the target user's attention data to the associated users.
[0285] Associated users can be supervisors (such as parents), educators (such as teachers), or learning partners of the target user. This application does not restrict the specific relationship between the associated user and the target user. The device terminal can be a smartphone, tablet, personal computer, or a dedicated home-school management platform client used by the associated user.
[0286] In some embodiments, associated users can remotely view the real-time or historical trends of attentional state changes and specific distraction characteristics of target users through the received attention data, and provide the ability to provide remote voice encouragement, task guidance, or directly trigger specific intervention strategies (such as locking non-learning applications) through the device terminal.
[0287] It should be understood that this synchronization mechanism optimizes attention detection from two dimensions: first, it synchronizes attention data to the server, enabling dynamic updates of the initial attention threshold and baseline attention feature parameters based on data from multiple users within the target group, thus improving the adaptability and timeliness of the initial attention threshold; second, it synchronizes attention data to the devices of associated users, establishing an external supervision and intervention channel. By allowing associated users to view and provide feedback on the attention status results of the target users, the credibility of the system evaluation and the actual effectiveness of intervention measures are enhanced. These two mechanisms work synergistically to improve the accuracy of attention status assessment, its adaptability to changes in groups and scenarios, and the operability of the assessment results in practical applications.
[0288] In some embodiments, the attention detection system can also manage the operating strategy of electronic devices based on user behavior analysis to help maintain or improve the user's attention state. This management process specifically includes steps S901-S902: S901. Based on the operating status data of electronic devices, determine the usage behavior patterns of target users when using electronic devices.
[0289] Among them, usage behavior patterns refer to the combinations of device operation habits or preferences that are statistically regular and repeatable, identified by clustering or sequence pattern mining of multi-dimensional features such as interaction timing, application usage duration, task switching frequency, and time period activity extracted from device operation status data.
[0290] S902. Generate operating strategies for electronic devices based on usage behavior patterns.
[0291] The operating strategy refers to the set of instructions used to dynamically adjust the hardware and software resources and interactive interface of electronic devices, aiming to match the device behavior with the user's current attention maintenance needs.
[0292] Operating policies are used to adjust the content output regulation parameters of electronic devices, user access permissions, and application scheduling plans.
[0293] Specifically, content output adjustment parameters include adaptive adjustments to screen and audio output, such as automatically optimizing screen brightness and color temperature based on ambient light and usage duration, or switching system notification sounds to softer, non-intrusive sound effects depending on the scenario; user access permissions include time-based or conditional management of access to specific applications, such as suspending notification pushes from non-learning applications during learning tasks, or restricting the launch of entertainment applications when a decline in attention is detected; application scheduling plans include prediction-based resource allocation and process management, such as preloading relevant learning tools into memory before users typically begin learning, or prioritizing the performance of foreground learning applications when system resources are scarce.
[0294] One possible implementation is as follows: The system first establishes a user behavior pattern library by analyzing historical operational data. When real-time operational data matches a pattern in the library, the system triggers a corresponding predefined strategy generation process. This process instantiates and assigns values to parameters in a preset strategy template (such as content output parameter thresholds, permission control rules, and application scheduling priorities) based on the current attention state and usage scenario, forming an executable set of operational strategy instructions. After the strategy is executed, the system evaluates the strategy's effectiveness by comparing the directional changes in attention characteristic parameters before and after the strategy implementation (such as the decrease in specific distraction indicators). If the evaluation result does not meet the preset improvement target, the system will automatically adjust the strategy strength parameters or switch to other alternative strategy sequences for that distraction pattern, depending on the deviation type of the characteristic parameters.
[0295] It should be understood that, through S901-S902, the system achieves a closed-loop dynamic adjustment of device parameters based on behavioral patterns. This mechanism extracts user habit patterns from historical data and generates device resource adjustment commands in real time accordingly, making targeted adjustments to system-level parameters such as screen output, application permissions, and process scheduling. This enhances the system's ability to reduce environmental interference and adapt to individual usage rhythms, providing system-level automated control support for attention maintenance.
[0296] In an exemplary embodiment, in Figure 1 Based on the above embodiments, such as Figure 11 As shown, the attention detection system provided in this application embodiment also includes: an attention intervention module 115, an attention state transfer prediction module 116, a data interaction module 117, and a behavior analysis module 118.
[0297] Attention intervention module 115 is configured to perform the above S501-S502.
[0298] The attention state shift prediction module 116 is configured to execute the above-described S601-S602, and the attention intervention module 115 is also configured to execute the above-described S603.
[0299] The data interaction module 117 is configured to execute the above-described S801 and / or S802.
[0300] Behavior analysis module 118 is configured to execute the above S901-S902.
[0301] For specific implementation examples, please refer to the above content, which will not be repeated here.
[0302] This disclosure also provides an electronic device, such as... Figure 12 The diagram shown is a structural schematic of an electronic device. Figure 12 As shown, the electronic device 1200 includes one or more processors 1210, a memory 1220, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device 1200, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interface). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices 1200 can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 12 Take the 1210 processor as an example.
[0303] Processor 1210 may be a central processing unit, a network processor, or a combination thereof. Processor 1210 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPRS), or any combination thereof.
[0304] The memory 1220 stores instructions executable by at least one processor 1210 to cause the at least one processor 1210 to perform the method shown in the above embodiments.
[0305] The memory 1220 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function. The data storage area may store data created based on the use of the electronic device 1200. Furthermore, the memory 1220 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device.
[0306] In some alternative implementations, memory 1220 may optionally include memory remotely located relative to processor 1210, and this remote memory may be connected to the electronic device 1200 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0307] The memory 1220 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 1220 may also include a combination of the above types of memory.
[0308] The electronic device 1200 also includes an input device 1230 and an output device 1240. The processor 1210, memory 1220, input device 1230, and output device 1240 can be connected via a bus 1250 or other means. Figure 12 Taking the connection between China and Israel via bus 1250 as an example.
[0309] Input device 1230 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device 1200, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1240 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touch screen.
[0310] The electronic device 1200 also includes a communication interface 1260 for communicating with other devices or communication networks.
[0311] Some embodiments of this disclosure provide a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) storing computer program instructions that, when executed on a computer (e.g., a receiving node), cause the computer to perform a task allocation method as described in any of the embodiments above.
[0312] For example, the computer-readable storage media described above may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes), optical discs (e.g., CDs (compact disks), DVDs (digital versatile disks), etc.), smart cards and flash memory devices (e.g., EPROMs (erasable programmable read-only memory), cards, sticks, or key drives, etc.). The various computer-readable storage media described in this disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.
[0313] Some embodiments of this disclosure also provide a computer program product, for example, stored on a non-transitory computer-readable storage medium. The computer program product includes computer program instructions that, when executed on a computer (e.g., a receiving node), cause the computer to perform the task allocation method as described in the above embodiments.
[0314] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. An attention detection method, characterized in that, The method includes: Acquire first multimodal perception data of the target user when using the electronic device at the current moment; the multimodal perception data includes: the target user's physiological state data, the target user's posture data, and the electronic device's operating state data; The first multimodal perception data is correlated, fused, and feature extracted to obtain the attention feature parameters of the target user; Second multimodal perception data within a historical time period is acquired. Based on the second multimodal perception data, the initial attention threshold of the target user is adjusted to generate the target attention threshold of the target user. The initial attention threshold is obtained by analyzing and processing the attention data of multiple users within the target group to which the target user belongs. Each target attention threshold corresponds to an attention feature parameter. The attention state of the target user is determined based on the degree of deviation between the attention feature parameters and the corresponding target attention threshold.
2. The method according to claim 1, characterized in that, The process of acquiring second multimodal perception data within a historical time period, adjusting the initial attention threshold of the target user based on the second multimodal perception data, and generating the target attention threshold of the target user includes: Based on the target user's attribute information, the target group to which the target user belongs is determined, and the initial attention threshold and baseline attention feature parameters corresponding to the target group are obtained; wherein, the attribute information includes: the target user's age, the target user's learning stage, the target user's learning ability quantification level, and the target user's learning subject type; the group is obtained by dividing different users based on the attribute information; the baseline attention feature parameters are obtained by analyzing and processing the attention feature parameters of multiple users within the target group; The second multimodal perception data is correlated, fused, and feature extracted to obtain the attention feature parameters of the target user in historical time periods; Determine the degree of deviation between the target user's attention feature parameters and the baseline attention feature parameters within a historical time period; Based on the degree of deviation of the features, the initial attention threshold is adjusted in the first way to generate the target attention threshold; the first adjustment includes the following strategies: a compensation adjustment strategy based on the degree of decline in the target user's physiological state, a compensation adjustment strategy based on the working time of the electronic device, and an adaptation adjustment strategy based on the difficulty of the learning task performed by the electronic device.
3. The method according to claim 1, characterized in that, The operating status data includes the operating context information of the electronic device; The runtime context information includes: the interaction information of the foreground application of the electronic device and the file format of the data loaded by the foreground application; The process of acquiring second multimodal perception data within a historical time period, adjusting the initial attention threshold of the target user based on the second multimodal perception data, and generating the target attention threshold of the target user includes: Based on the runtime context information, semantic recognition is performed to determine the usage scenario type of the electronic device; Based on the usage scenario type and the preset mapping relationship between scenario and threshold adjustment strategy, a threshold adjustment strategy for making a second adjustment to the initial attention threshold is determined; Based on the threshold adjustment strategy, the first attention threshold is adjusted a second time to generate the target attention threshold for the target user.
4. The method according to claim 1, characterized in that, The step of determining the attention state result of the target user based on the degree of threshold deviation between the attention feature parameters and the corresponding target attention threshold includes: Determine the degree of deviation of each attention feature parameter from the corresponding target attention threshold; Determine the usage scenario type of the electronic device; The degree of threshold deviation is constructed into a feature combination vector, and the usage scenario type is constructed into a prompt message; The attention analysis model processes the feature combination vector based on the prompt information to obtain the attention state result of the target user; wherein, the attention analysis model is trained through at least two model training stages; the model training stages include: a pre-training stage based on different user attention data and a reinforcement training stage based on the target user attention data.
5. The method according to claim 1, characterized in that, The physiological state data of the target user includes one or more of the following: the target user's respiratory rate and heart rate variability; The physiological state data is obtained by detecting and collecting data from the target user using a millimeter-wave radar array; the millimeter-wave radar array includes multiple millimeter-wave radar units with different detection angles, and the multiple millimeter-wave radar units are deployed on the edge and / or side of the electronic device's display screen; The target user's posture data includes one or more of the following: the target user's motion trajectory, the target user's motion amplitude, the sitting pressure distribution entropy value, the sitting body movement frequency, writing trajectory information, writing pressure information, and writing angle information; the posture data is obtained by detecting and collecting the target user through at least one of the following sensors: a millimeter-wave radar array, a flexible piezoelectric thin film sensor matrix, and a stylus sensor; the flexible piezoelectric thin film sensor matrix and the stylus sensor are communicatively connected to the electronic device. The operating status data of the electronic device includes one or more of the following: the operating context information of the electronic device, the computing resource consumption information, and the interaction information of the electronic device.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: If the target user's attention state result is less than or equal to a preset attention state result threshold, a multimodal attention intervention strategy and the order of execution of the intervention strategy are determined based on the attention state result, the degree of feature deviation, and the usage scenario type. The multimodal attention intervention strategy includes at least one of auditory intervention strategy, visual intervention strategy, tactile intervention strategy, and guided intervention strategy. The multimodal attention intervention strategy is executed according to the execution sequence of the intervention strategy to improve the attention state of the target user.
7. The method according to claim 6, characterized in that, When the attention state result of the target user is less than or equal to a preset attention state result threshold, based on the attention state result, the degree of feature deviation, and the usage scenario type, a multimodal attention intervention strategy and the execution order of the intervention strategy for the target user are determined, including: Based on the preset mapping relationship between state results and priorities, the attention intervention priority corresponding to the attention state result is determined; From the preset intervention priority and intervention strategy mapping information, determine the multimodal attention intervention strategy for the target user; The set of intervention strategies is ranked by importance based on the degree of feature deviation and the type of use scenario, the execution priority of each intervention strategy in the multimodal attention intervention strategy is determined, and the execution order of the intervention strategies is generated. The attention intervention priorities, ranked from lowest to highest, include: first attention intervention priority, second attention intervention priority, and third attention intervention priority; The first attentional intervention priority corresponds to a visual intervention strategy; the second attentional intervention priority corresponds to both a visual intervention strategy and an auditory sensory strategy; and the third attentional intervention priority corresponds to both a tactile intervention strategy and an auditory intervention strategy.
8. The method according to claim 6, characterized in that, When the attention state result of the target user is less than or equal to a preset attention state result threshold, based on the attention state result, the degree of feature deviation, and the usage scenario type, a multimodal attention intervention strategy and the execution order of the intervention strategy for the target user are determined, including: Obtain the attention state result sequence of the target user; the attention state result sequence includes the attention state results determined at different times; The attention state result sequence is processed by the attention state transition prediction model to obtain the trend of the target user's attention state change over a future period of time. Based on the trend of attention state changes, the results of attention state, the type of usage scenario, and the degree of feature deviation, a multimodal attention intervention strategy and the order of execution of the intervention strategy are determined for the target user.
9. The method according to claim 6, characterized in that, The method further includes: Obtain the first attention state result of the target user after implementing the multimodal attention intervention strategy; Determine the attention improvement value of the first attention state result relative to the attention state result before implementing the multimodal attention intervention strategy; If the attention enhancement value is less than or equal to the attention enhancement threshold, the multimodal attention intervention strategy is redefined and executed.
10. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtain the interaction information of the foreground application from the running context information, and determine the list of foreground applications running on the electronic device from the interaction information of the foreground application; If the foreground application running list meets preset conditions and the change magnitude of the attention feature parameter is less than the change magnitude threshold, the unwanted application in the foreground application running list is closed or paused; wherein, the preset conditions include: the foreground application running list contains at least one desired application and at least one unwanted application.
11. The method according to any one of claims 1-5, characterized in that, The method further includes: The attention data of the target user is synchronized to the server so that the server updates the initial attention threshold and / or baseline attention feature parameters of the target group based on the attention data of multiple users in the target group; wherein, the attention data includes one or more of the following: the identification identifier of the target user, the first multimodal perception data, the attention feature parameters, and the attention state result; and / or; The attention data of the target user is synchronized to the device terminal of the associated user of the target user; the device terminal is used to provide the attention data of the target user to the associated user.
12. The method according to any one of claims 1-5, characterized in that, The method further includes: Based on the operating status data of the electronic device, determine the target user's usage behavior pattern of the electronic device; Based on the usage behavior pattern, an operating strategy is generated for the electronic device; wherein the operating strategy is used to adjust the content output adjustment parameters, user access permissions, and application scheduling plan of the electronic device.
13. An attention detection system, characterized in that, The system includes: electronic devices and sensor units; The sensor unit is communicatively connected to or deployed within the electronic device; the electronic device includes: a data acquisition module, a feature extraction module, a threshold generation module, and an attention determination module; The data acquisition module is configured to acquire first multimodal perception data of the target user using the electronic device at the current moment through the sensor unit; the multimodal perception data includes: the target user's physiological state data, the target user's posture data, and the electronic device's operating state data; The feature extraction module is configured to perform correlation fusion and feature extraction processing on the first multimodal perception data to obtain the attention feature parameters of the target user; The threshold generation module is configured to acquire second multimodal perception data within a historical time period, adjust the initial attention threshold of the target user based on the second multimodal perception data, and generate a target attention threshold for the target user; the initial attention threshold is obtained by analyzing and processing the attention data of multiple users within the target group to which the target user belongs; each target attention threshold corresponds to an attention feature parameter; The attention determination module is configured to determine the attention state result of the target user based on the degree of threshold deviation between the attention feature parameters and the corresponding target attention threshold.
14. The system according to claim 13, characterized in that, The sensor unit includes one or more of the following: a millimeter-wave radar array, a flexible piezoelectric thin film sensor matrix, and a stylus sensor. The millimeter-wave radar array includes multiple millimeter-wave radar units with different detection angles; each millimeter-wave radar unit is deployed on the edge and / or side of the electronic device's display screen; the flexible piezoelectric thin-film sensor matrix and the stylus sensor are respectively communicatively connected to the electronic device; The stylus sensor includes: a three-axis speed sensor for the stylus, a pressure sensor for the stylus, and a tilt angle sensor for the stylus; The millimeter-wave radar array is configured to collect physiological state data of the target user and posture data related to the target user's actions; The flexible piezoelectric thin film sensor matrix is configured to collect posture data related to the sitting posture of the target user; The stylus sensor is configured to collect posture data related to the writing posture of the target user; The electronic device is also configured to collect operational status data during operation.
15. An electronic device, characterized in that, The electronic device includes: A memory and one or more processors, the memory being coupled to the processors; wherein the memory stores computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1-12.
Citation Information
Cited By
A method and system for intelligent autonomy of a multimodal fusion core network based on images and text.
CN122137454A
A picture-text multi-modal fusion core network intelligent autonomy method and system
CN122137454B