Multi-modal user authentication method and system for television-based education applications

By monitoring the user's facial orientation and voice interaction frequency, the authentication process is dynamically adjusted to achieve multimodal biometric weighted fusion calculation, which solves the problem of the disconnect between authentication strategies and learning status in TV-based educational applications, and improves authentication accuracy and user experience.

CN120956974BActive Publication Date: 2025-12-23CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511486784.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-12-23
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

In TV-based educational applications, user authentication strategies cannot adapt to the dynamic changes in learning states, leading to decreased authentication accuracy and a poor user learning experience, especially with high security risks during focused learning and excessively high authentication requirements during distracted learning.

Method used

By monitoring the user's facial orientation deviation and voice interaction frequency through IPTV set-top boxes, learning participation indicators are calculated, the strictness of the authentication process is dynamically adjusted, and multimodal biometric weighted fusion calculation is used to generate identity confidence scores associated with the learning state, thereby achieving intelligent matching between authentication strategies and user learning states.

Benefits of technology

It improves the accuracy of authentication and user experience, ensures optimized configuration of authentication accuracy in different learning states, avoids misidentification and excessively high authentication thresholds, meets security requirements and does not interfere with normal learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956974B_ABST
    Figure CN120956974B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses a television terminal education application multi-modal user authentication method and system. The method comprises the following steps: obtaining a learning participation index by monitoring a user facial orientation offset degree and a voice interaction frequency through an IPTV set-top box; performing focused learning or distraction state authentication process according to comparison of the index and a preset threshold value 0.7; dynamically adjusting high sensitivity or low sensitivity authentication parameters through an authentication strictness adjuster; performing weighted fusion calculation on multi-modal biological characteristics of a user to obtain a learning state associated identity confidence score; and matching and verifying the identity confidence score with a dynamic authentication threshold value to obtain a user identity authentication result based on learning concentration. The application solves the problem that a user authentication strategy in a television terminal education application cannot adapt to dynamic changes in a learning state, and improves multi-modal user identity authentication accuracy based on learning concentration and adaptability of a user learning experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a multi-modal user authentication method and system for an education application on a television terminal. BACKGROUND

[0002] The user authentication technology of the education application on the television terminal mainly adopts a single-modal identity verification method, including password authentication based on remote control input, contact authentication based on a smart card, and single-modal biometric authentication based on voice recognition or facial recognition. These existing technologies are widely used in user identity recognition and personalized content pushing in the IPTV education platform. The voice recognition technology verifies the identity by analyzing the voiceprint characteristics of the user, the facial recognition technology confirms the identity by comparing the facial features of the user with a pre-stored template, and the traditional password input method relies on the user to manually input a preset combination of numbers or characters to complete the authentication process.

[0003] The existing technology has significant deficiencies. The single-modal authentication method is easily disturbed by factors such as light changes, background noise, and user posture changes in the complex environment of the television terminal, resulting in a decrease in authentication accuracy. The fixed threshold authentication strategy cannot adapt to the changes in the behavior characteristics of the user in different learning states. When the user is in a focused learning or distracted state, the differences in their biometric characteristics are large, but the existing technology uses a unified authentication standard, which can easily cause misidentification or rejection. In addition, the existing authentication method lacks a deep association with the user's learning behavior and the characteristics of the education content, and cannot adjust the strictness of the authentication strategy according to the dynamic changes in the learning scene.

[0004] Based on the above analysis, the core problem of the existing technology is the disconnection between the authentication strategy and the user's learning state. The difference in the user's concentration when watching different important education content significantly affects the stability and recognizability of their biometric characteristics. A fixed authentication threshold cannot adapt to this dynamic change. Furthermore, the lack of a learning state perception mechanism makes it impossible to distinguish between the focused learning state and the distracted state of the user during the authentication process, dynamically adjust the weight distribution of multi-modal features according to the learning concentration, and ultimately lead to the technical dilemma of too low authentication requirements in the focused state, which poses a security risk, and too high authentication requirements in the distracted state, which affects the user's learning experience. SUMMARY

[0005] The present application provides a multi-modal user authentication method and system for an education application on a television terminal, which solves the problem that the user authentication strategy in the education application on the television terminal cannot adapt to the dynamic changes in the learning state, and improves the accuracy of multi-modal user identity authentication based on learning concentration and the adaptability of user learning experience.

[0006] In a first aspect, the present application provides a multi-modal user authentication method for an education application on a television terminal, which comprises:

[0007] Step S1: synchronously monitoring the face orientation deviation and voice interaction frequency of the user when watching the educational content through the IPTV set-top box to obtain a learning participation index;

[0008] Step S2: comparing the learning participation index with a preset participation threshold value, if the learning participation index is greater than 0.7, executing a focused learning authentication process, if the learning participation index is less than or equal to 0.7, executing a distraction state authentication process;

[0009] Step S3: dynamically adjusting the focused learning authentication process or the distraction state authentication process through an authentication strictness adjuster to obtain a high-sensitivity authentication parameter or a low-sensitivity authentication parameter;

[0010] Step S4: performing weighted fusion calculation processing on the multi-modal biological features of the user according to the high-sensitivity authentication parameter or the low-sensitivity authentication parameter to obtain a learning state-related identity confidence score;

[0011] Step S5: matching and verifying the learning state-related identity confidence score with a dynamic authentication threshold value to obtain a user identity authentication result based on learning concentration.

[0012] In a second aspect, the present application provides a television-side educational application multi-modal user authentication system, which comprises:

[0013] A monitoring module is configured to synchronously monitor the face orientation deviation and voice interaction frequency of the user when watching the educational content through the IPTV set-top box to obtain a learning participation index;

[0014] A judgment module is configured to compare the learning participation index with a preset participation threshold value, if the learning participation index is greater than 0.7, execute a focused learning authentication process, if the learning participation index is less than or equal to 0.7, execute a distraction state authentication process;

[0015] An adjustment module is configured to dynamically adjust the focused learning authentication process or the distraction state authentication process through an authentication strictness adjuster to obtain a high-sensitivity authentication parameter or a low-sensitivity authentication parameter;

[0016] A fusion module is configured to perform weighted fusion calculation processing on the multi-modal biological features of the user according to the high-sensitivity authentication parameter or the low-sensitivity authentication parameter to obtain a learning state-related identity confidence score;

[0017] A verification module is configured to match and verify the learning state-related identity confidence score with a dynamic authentication threshold value to obtain a user identity authentication result based on learning concentration.

[0018] In a third aspect, a television-side education application multi-modal user authentication device is provided, comprising: a memory and at least one processor, the memory storing instructions; the at least one processor invoking the instructions in the memory to cause the television-side education application multi-modal user authentication device to perform the television-side education application multi-modal user authentication method described above.

[0019] In a fourth aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the television-side education application multi-modal user authentication method described above.

[0020] In the technical solutions provided in the present application, the learning participation index is obtained by synchronously monitoring and processing the face orientation offset degree and the voice interaction frequency of the user when watching the education content through the IPTV set-top box, the technical problem that the user learning state cannot be accurately evaluated in the prior art is solved, the focused learning authentication process and the distracted state authentication process can be accurately distinguished after the learning participation index is compared and judged with the preset participation threshold, the misrecognition phenomenon caused by the single standard in the traditional authentication method is avoided, the high-sensitivity authentication parameter or the low-sensitivity authentication parameter is obtained by dynamically adjusting and processing the security level through the authentication strictness adjuster in the focused learning authentication process and the distracted state authentication process, the intelligent matching of the authentication strategy and the user learning state is realized, the learning state associated identity confidence score is obtained by weighted fusion calculation and processing of the multi-modal biological characteristics of the user according to the high-sensitivity authentication parameter or the low-sensitivity authentication parameter, the optimal configuration of the authentication accuracy under different learning states is ensured, the user identity authentication result based on the learning concentration degree is obtained by matching and verifying the learning state associated identity confidence score with the dynamic authentication threshold, and the technical defect that the fixed authentication threshold in the prior art cannot adapt to the change of the learning scene is completely solved.

[0021] The learning participation index calculation algorithm can perceive the learning input degree of the user in real time, and provides accurate state basis for dynamic adjustment of subsequent authentication strategies. The security level mapping algorithm in the authentication strictness adjuster dynamically adjusts the authentication requirements according to the learning state and the importance of the education content, avoids the technical problems of setting too low security standards for important education content or setting too high authentication threshold for general content, the multi-modal biological feature weighted fusion algorithm adaptively adjusts the weight distribution of face recognition, voice recognition and gesture recognition according to the learning state, strengthens the role of the main authentication mode in the focused state, balances the contribution of multi-modal in the distracted state, reduces the risk of single mode failure, the dynamic authentication threshold generation algorithm calculates the most suitable authentication standard combined with the current learning state and the importance of the education content, ensures that the authentication process meets the security requirements and does not interfere with normal learning, and the synergistic effect of these algorithm characteristics enables the whole authentication scheme to realize intelligent authentication of learning state perception in the television education application scene, significantly improves the authentication experience and system security of the user in different learning states. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without creative labor.

[0023] Figure 1 An embodiment of the television education application multi-modal user authentication method in the embodiment of the present application is shown in the figure.

[0024] Figure 2 An embodiment of the television education application multi-modal user authentication system in the embodiment of the present application is shown in the figure.

[0025] Figure 3 The structure schematic diagram of the television education application multi-modal user authentication device in the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0026] The television-side education application multi-modal user authentication method and system are provided in the embodiments of the present application. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not have to be used to describe a particular order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "comprising" or "having" and any variation thereof is intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.

[0027] For ease of understanding, the specific processes of the embodiments of the present application are described below. Please refer to Figure 1 One embodiment of the television-side education application multi-modal user authentication method in the embodiments of the present application includes the following steps.

[0028] Step S1: synchronously monitoring and processing the face orientation offset degree and the voice interaction frequency of a user when watching education content through an IPTV set-top box, to obtain a learning participation index;

[0029] Step S2: comparing and judging the learning participation index with a preset participation threshold value, if the learning participation index is greater than 0.7, executing a focused learning authentication process, and if the learning participation index is less than or equal to 0.7, executing a distraction state authentication process;

[0030] Step S3: dynamically adjusting and processing the security level of the focused learning authentication process or the distraction state authentication process through an authentication strictness adjuster, to obtain a high-sensitivity authentication parameter or a low-sensitivity authentication parameter;

[0031] Step S4: performing weighted fusion calculation and processing on multi-modal biological features of a user according to the high-sensitivity authentication parameter or the low-sensitivity authentication parameter, to obtain an identity confidence score associated with a learning state;

[0032] Step S5: performing matching and verification processing on the identity confidence score associated with the learning state and a dynamic authentication threshold value, to obtain a user identity authentication result based on learning concentration.

[0033] It can be understood that the execution subject of the present application can be a television-side education application multi-modal user authentication system, and can also be a terminal or a server, which is not limited here. The embodiments of the present application take the server as an example for illustration.

[0034] Specifically, the real-time monitoring of the user learning state is realized by the hardware integration capability of the IPTV set-top box. The face orientation offset degree is continuously collected by the set-top box integrated camera through the face image sequence of the user, the 68 standard face key points are located by the face key point detection algorithm, the displacement change of the key points between adjacent frames is calculated to quantify the offset degree of the user's head orientation, and the voice interaction frequency is identified by the voice activity detection algorithm after the voice signal is collected by the set-top box microphone to recognize the effective voice segment and count the number of interactions per unit time. Finally, the face orientation offset degree value and the voice interaction frequency parameter are fused by weighting to calculate the comprehensive learning participation index value.

[0035] When the learning participation index is obtained, the system compares the index with the preset participation threshold 0.7 in numerical value, and corrects the learning state in combination with the difficulty coefficient of the currently played educational content. When the corrected learning state evaluation value is greater than 0.7, the focused learning authentication process is activated and the user attention concentration time parameter is recorded. When the evaluation value is less than or equal to 0.7, the distraction state authentication process is activated and the biological feature collection frequency is reduced to half of the standard frequency. Then, the selected authentication process is matched and analyzed with the user's historical learning behavior pattern to obtain the personalized authentication strategy configuration parameter, and the authentication process dispatcher is used to prioritize and arrange the execution order of these parameters.

[0036] The authentication strictness adjuster is responsible for dynamically adjusting the security level according to different learning states. After receiving the focused learning authentication process or the distraction state authentication process, the process type is identified and the security level is mapped. In combination with the importance coefficient of the current educational content, the security level is dynamically calibrated. When the calibrated security level value is greater than 0.8, the high sensitivity authentication parameters are set, including the face recognition weight 0.6 and the voice recognition weight 0.4. When the value is less than or equal to 0.8, the low sensitivity authentication parameters are set, including the face recognition weight 0.4, the voice recognition weight 0.3 and the gesture recognition weight 0.3.

[0037] In the multi-modal biometric feature processing stage, the IPTV set-top box collects the user's facial features, voice features and gesture features in real time to form a multi-modal biometric feature dataset, a facial recognition algorithm performs feature point positioning and feature vector encoding processing on the key areas of the face to generate a standardized facial feature vector, the vector and the user's pre-registered face template are matched through vector distance calculation and cosine similarity to obtain a face recognition original confidence score, a voice recognition algorithm performs frequency domain transformation and acoustic feature parameter extraction on the voice signal to obtain a voice feature parameter set, the set and the user's pre-registered voiceprint template are matched through statistical matching to calculate the similarity score and obtain a voice recognition original confidence score, a gesture recognition algorithm performs geometric feature analysis and motion pattern recognition on the hand movement trajectory to obtain a gesture feature descriptor, the descriptor and the preset gesture action template library are matched through a similarity calculation function to evaluate the matching degree and obtain a gesture recognition original confidence score, and then the original confidence scores of each modality are weighted and summed according to the high-sensitivity authentication parameter or the low-sensitivity authentication parameter to obtain a comprehensive confidence score after weighted fusion.

[0038] The identity authentication verification process is realized through a dynamic threshold adjustment mechanism. The system evaluates the authentication security requirement level according to the current learning state of the user and the importance of the educational content, generates a corresponding dynamic authentication threshold, compares the identity confidence score associated with the learning state with the dynamic authentication threshold in terms of numerical value, and judges whether the score is higher than the threshold. When the score is higher than the threshold, the user's identity information and learning state identifier are recorded to generate an authentication success identifier. When the score is lower than the threshold, a re-authentication process is triggered or a degraded access permission is set to output the user identity authentication result based on learning concentration. For example, when a primary school student Zhang watches a mathematics course, the camera detects that his face orientation deviation is small and the voice interaction frequency is high. The system calculates that the learning participation index is 0.85, which exceeds the threshold 0.7 and enters the focused learning authentication process. The authentication strictness adjuster sets the security level to high sensitivity mode, and the weights of facial recognition and voice recognition are set to 0.6 and 0.4 respectively. When the system collects the biometric features of Zhang, it calculates the facial recognition confidence score 0.92 and the voice recognition confidence score 0.88 respectively. Through weighted fusion, the comprehensive confidence score is 0.904. After comparing this value with the dynamic authentication threshold 0.85, it is confirmed that the authentication is passed, and the system outputs the identity authentication success result based on learning concentration.

[0039] In an embodiment, step S1 comprises:

[0040] The IPTV set-top box integrates a camera to continuously collect and process frame images of the user's face area to obtain face image sequence data;

[0041] The face key point detection algorithm is used to perform key point positioning and inter-frame displacement calculation on the face image sequence data to obtain the face orientation deviation;

[0042] The voice signal collected by the microphone of the IPTV set-top box is input into a voice activity detection algorithm for effective voice segment recognition and frequency statistics processing to obtain a voice interaction frequency;

[0043] The learning participation index is obtained through comprehensive evaluation processing by weighted fusion calculation according to the facial orientation offset degree and the voice interaction frequency.

[0044] Specifically, the IPTV set-top box integrates a camera to continuously capture image data of the user's facial region at a rate of 30 frames per second, and the resolution of each frame of image is set to 1920x1080 pixels. The effective shooting distance of the camera covers a television viewing area of 2 to 5 meters. The image data generated during the continuous acquisition process is arranged in chronological order to form facial image sequence data, which contains all facial state change information of the user during the viewing of educational content. The automatic focusing function of the camera ensures that clear facial images can be obtained at different distances, and the low light compensation mechanism of the image sensor ensures that it can work normally under various light conditions in the living room environment. After receiving the facial image sequence data, the facial key point detection algorithm first pre-processes each frame of image, including grayscale conversion and noise filtering. Then the algorithm locates 68 standard facial key points in each frame of image, which are distributed in the regions of eyebrows, eyes, nose, mouth and facial contour. The algorithm identifies the accurate coordinate positions of these key points through a trained convolutional neural network model, and then calculates the displacement change of the same key points between adjacent frames. The specific calculation method is to subtract the coordinate of the corresponding key point in the previous frame from the coordinate of the current frame to obtain a displacement vector, and then calculate the average value of all key point displacement vectors as the quantification index of the facial orientation offset degree through the Euclidean distance formula. This index reflects the degree of offset of the user's head relative to the center position of the screen.

[0045] The digital microphone array built-in in the IPTV set-top box continuously collects the voice signals emitted by the user, and the sampling frequency is set to 48 kHz to ensure high fidelity of the voice signals. The pickup range of the microphone is designed to be omnidirectional to cover the entire living room space. The collected analog voice signals are converted into digital audio data streams through an analog-to-digital converter. After receiving these digital audio data, the voice activity detection algorithm first performs pre-emphasis processing to balance the high and low frequency components, and then identifies the effective voice segment through a short-time energy and zero-crossing rate double detection mechanism. The short-time energy detection distinguishes between voice and silence segments by calculating the energy value of the audio frame. The zero-crossing rate detection identifies voice features by counting the number of jumps near zero of the audio signal. When the short-time energy exceeds the preset threshold and the zero-crossing rate is within the voice range, the algorithm marks this time segment as an effective voice segment. Subsequently, the algorithm counts the number of effective voice segments detected within a fixed time window to obtain the voice interaction frequency parameter. The high or low value of this parameter directly reflects the active degree of the user's voice interaction with the educational content.

[0046] The weighted fusion calculation process first normalizes the facial orientation offset value to control the value range between 0 and 1, the normalization method is to divide the current offset by the preset maximum offset threshold, similarly, the speech interaction frequency parameter is also normalized by dividing the preset maximum interaction frequency, the two normalized parameters are multiplied by the corresponding weight coefficients respectively, wherein the weight coefficient of the facial orientation offset is set to 0.6 to reflect its dominant position in the learning participation evaluation, and the weight coefficient of the speech interaction frequency is set to 0.4 to reflect its auxiliary evaluation function, the specific calculation of the weighted fusion is to multiply the normalized facial orientation offset by 0.6 and add the normalized speech interaction frequency multiplied by 0.4 to obtain the comprehensive learning participation index value, the value range of the value is between 0 and 1, the higher the value, the stronger the learning participation of the user. For example, when middle school student Li watches an English course, the camera continuously collects his facial images and finds that his head basically keeps a posture facing the screen, the facial key point detection algorithm calculates that the average displacement of the key points between adjacent frames is small, and after normalization, the facial orientation offset is 0.1. At the same time, the microphone detects that Li often makes reading sounds during the course playing process, and the speech activity detection algorithm identifies the frequent valid speech segments, and after normalization, the speech interaction frequency is 0.8. The weighted fusion calculation multiplies 0.1 by 0.6 and adds 0.8 by 0.4 to obtain a learning participation index of 0.38, which indicates that Li is currently in a medium learning participation state.

[0047] In a specific embodiment, step S2 comprises:

[0048] The learning participation index is compared with the preset participation threshold value 0.7, and the learning state correction process is performed in combination with the difficulty coefficient of the currently played educational content to obtain a corrected learning state evaluation value;

[0049] When the corrected learning state evaluation value is greater than 0.7, the focused learning authentication process is activated, and the user attention concentration duration parameter is recorded to obtain the focused learning authentication process;

[0050] When the corrected learning state evaluation value is less than or equal to 0.7, the distraction state authentication process is activated, and the biological feature collection frequency is reduced to 0.5 times the standard frequency to obtain the distraction state authentication process;

[0051] Based on the focused learning authentication process or the distraction state authentication process, the matching analysis process is performed on the user's historical learning behavior mode to obtain a personalized authentication strategy configuration parameter;

[0052] The personalized authentication strategy configuration parameter is input into the authentication process scheduler for process priority sorting and execution order arrangement processing to obtain an authentication execution scheme.

[0053] Specifically, the learning engagement indicator value is directly compared with the preset engagement threshold value 0.7. When the learning engagement indicator value is greater than 0.7, it is determined that the user is in a relatively focused learning state. When the value is less than or equal to 0.7, it is determined that the user is in a distracted or inattentive state. Then, the data processing module reads the difficulty coefficient parameter of the currently played educational content. The difficulty coefficient is pre-set according to factors such as subject type, knowledge point complexity, and applicable age range of the educational content. The difficulty coefficient of science subjects such as mathematics and physics is usually set to a higher value, while the difficulty coefficient of liberal arts subjects such as language and art is relatively low. The learning state correction process adjusts the value by multiplying the original learning engagement indicator by the reciprocal of the difficulty coefficient. When the educational content is difficult, the difficulty coefficient is large and its reciprocal is small. The corrected result after multiplication will appropriately reduce the original engagement indicator. Conversely, when the educational content is easy, the corrected result will be correspondingly improved. The corrected learning state evaluation value obtained after correction more accurately reflects the user's real learning state under specific educational content.

[0054] When the corrected learning state evaluation value is greater than 0.7, the authentication process manager immediately activates the focused learning authentication process. The activation process of this process includes setting the user state identifier to focused learning mode and starting the recording function of the attention concentration duration parameter. The attention concentration duration parameter is calculated by continuously monitoring the stability of the user's facial orientation deviation and speech interaction frequency. When these two parameters remain relatively stable and the values are within the focused range within a continuous time period, the timer starts to accumulate the duration. The data structure of the focused learning authentication process includes user identifier, learning content identifier, focused start timestamp, and real-time updated attention concentration duration. When the corrected learning state evaluation value is less than or equal to 0.7, the authentication process manager activates the distracted state authentication process. After the process is activated, it immediately sends an instruction to the hardware control module to reduce the biometric feature collection frequency from the standard 30 times per second to 15 times per second, which is 0.5 times the standard frequency. The frequency reduction is achieved by adjusting the data sampling interval of the camera and microphone. The original 33 ms data collection interval is adjusted to 66 ms. The data structure of the distracted state authentication process includes basic user and content identifier information, as well as special fields such as frequency reduction timestamp and collection frequency adjustment record.

[0055] The user historical learning behavior pattern matching analysis process first retrieves the user's past learning records from the user database. These historical data include the user's learning engagement variation curve under different educational content, the duration distribution of focused and distracted states, and the corresponding authentication success rate statistics. The matching analysis algorithm compares the current focused learning authentication process or distracted state authentication process with similar scenarios in the historical data. By calculating the similarity between the current learning state features and the state features in the historical records, the most matching historical pattern is found. The similarity calculation considers multiple dimensions such as the matching degree of learning content type, the consistency of learning time period, and the relevance of user behavior features. Based on the matching results, personalized authentication strategy configuration parameters are generated. This parameter set includes recommended authentication modality weight distribution, suitable authentication threshold setting, and expected authentication time interval, among other key configuration information.

[0056] The authentication process scheduler receives the personalized authentication strategy configuration parameters and starts the process priority sorting and execution order arrangement processing. The scheduler first analyzes the importance weights of each authentication modality in the configuration parameters. Higher-weighted authentication modalities such as facial recognition are arranged as priority execution items, while lower-weighted modalities such as gesture recognition are arranged as subsequent execution or backup options. The execution order arrangement also needs to consider hardware resource occupation and processing time optimization. The scheduler allocates reasonable execution time windows for each authentication modality through time slice allocation mechanism to ensure that resource conflicts do not occur between multiple authentication tasks. The final generated authentication execution scheme contains detailed task execution schedules, resource allocation schemes, and exception handling plans, among other complete information. For example, when high school student Wang watches a chemistry experiment course, his learning engagement index is 0.6, which is less than the threshold value of 0.7. Since the chemistry experiment belongs to high-difficulty content, its difficulty coefficient is 1.5. The corrected learning state evaluation value is 0.4 after dividing 0.6 by 1.5. This value triggers the activation of the distracted state authentication process. The authentication process manager reduces the biometric feature collection frequency from 30 times per second to 15 times. At the same time, the historical behavior pattern matching finds that Wang usually needs more adaptation time when watching science content. The personalized authentication strategy configuration parameters suggest extending the authentication time interval and increasing the weight of voice interaction. The authentication process scheduler arranges the execution order based on these parameters, with facial recognition as the priority, followed by voice recognition, and gesture recognition as the auxiliary.

[0057] In a specific embodiment, step S3 comprises:

[0058] The focused learning authentication process or distracted state authentication process is input into the authentication strictness adjuster for process type identification and security level mapping processing, obtaining an initial security level identifier;

[0059] According to the initial security level identification, the security level dynamic calibration processing is combined with the importance coefficient of the current education content to obtain the calibrated security level value.

[0060] When the calibrated security level value is greater than 0.8, high sensitivity authentication parameters are set, including a face recognition weight of 0.6 and a voice recognition weight of 0.4, to obtain the high sensitivity authentication parameters.

[0061] When the calibrated security level value is less than or equal to 0.8, low sensitivity authentication parameters are set, including a face recognition weight of 0.4, a voice recognition weight of 0.3, and a gesture recognition weight of 0.3, to obtain the low sensitivity authentication parameters.

[0062] Specifically, the authentication strictness adjuster receives input data of the focused learning authentication process or the distraction state authentication process, and first performs process type identification processing. The identification process distinguishes whether the input is the focused learning authentication process or the distraction state authentication process by reading the process identification field in the authentication process data structure. The identification code of the focused learning authentication process is set to 1, and the identification code of the distraction state authentication process is set to 0. The identification module directly determines the process type according to the value of the identification code. Then, the security level mapping processing module converts different process types into corresponding initial security level identifiers according to the preset mapping rule. The focused learning authentication process is mapped to a high initial security level identifier value of 0.9 because it requires high security protection due to the user's attention. The distraction state authentication process is mapped to a low initial security level identifier value of 0.4 because it requires relatively relaxed authentication due to the user's distraction. The mapping rule is stored in the security level mapping table, which contains a one-to-one correspondence between the process type and the security level.

[0063] The security level dynamic calibration processing first obtains the importance coefficient of the current playing content from the education content database. This coefficient quantitatively evaluates the importance of the education content for the learner's knowledge system construction. The importance coefficient of core basic courses such as mathematical basic operations is set to a higher value, and the importance coefficient of extension courses such as art appreciation is set to a lower value. The calibration processing calculates the calibrated security level value by multiplying the initial security level identifier and the importance coefficient. When the education content is important, its coefficient is large, which further amplifies the initial security level. When the education content is less important, its coefficient is small, which appropriately reduces the initial security level. The calibration calculation ensures that the security level setting matches the actual value of the education content, avoiding setting too high security requirements for unimportant content or too low security standards for important content.

[0064] The comparison judgment process between the calibrated security level value and the threshold value 0.8 determines the setting strategy of the authentication parameter through a numerical size comparison operation. When the calibrated security level value is greater than 0.8, it is determined that the current scene needs a high security protection level, and the authentication parameter configuration module immediately sets a high sensitivity authentication parameter. The parameter configuration sets the weight of the face recognition in the multi-modal authentication to 0.6, which reflects the important status of the face recognition as the main authentication means. The weight of the voice recognition is set to 0.4 as an important auxiliary authentication means. The weight distribution of the high sensitivity authentication parameter follows the principle of clear primary and secondary to ensure that the key authentication mode obtains sufficient influence. This configuration is suitable for the scene where the user is focused on learning and the educational content is important, and the user's identity needs to be strictly verified.

[0065] When the calibrated security level value is less than or equal to 0.8, the authentication parameter configuration module sets a low sensitivity authentication parameter. The parameter configuration reduces the weight of the face recognition to 0.4 to avoid imposing too high a face recognition requirement on the distracted user. The weight of the voice recognition is set to 0.3 as a secondary authentication means, and the weight of the gesture recognition is also set to 0.3 as a supplementary authentication method. The low sensitivity authentication parameter adopts a three-modal balanced distribution strategy to reduce the dependence on a single mode. This configuration disperses the authentication pressure so that even if the user performs poorly in a certain mode, the user can complete the authentication through other modes. The sum of the weight distribution always remains 1.0 to ensure the mathematical consistency of the authentication algorithm. After the parameter configuration is completed, the relevant data is written to the authentication parameter storage area for subsequent identity verification calculation. For example, when a junior high school student Chen watches an important mathematics foundation course and is in a focused learning state, the authentication strictness adjuster recognizes the identification code 1 of the focused learning authentication process and maps it to the initial security level identification 0.9. The importance coefficient of the mathematics foundation course is 1.2. The calibrated security level value obtained by multiplying 0.9 by 1.2 is 1.08, which is greater than the threshold value 0.8, triggering the setting of the high sensitivity authentication parameter. The weight of the face recognition is configured to 0.6, and the weight of the voice recognition is configured to 0.4. On the contrary, when Chen watches an art appreciation course and is in a distracted state, the identification code 0 of the distracted state authentication process is mapped to the initial security level identification 0.4. The importance coefficient of the art appreciation is 0.8. The calibrated security level value obtained by the calibration calculation is 0.32, which is less than 0.8, triggering the setting of the low sensitivity authentication parameter. The weights of the face recognition, the voice recognition, and the gesture recognition are configured to 0.4, 0.3, and 0.3, respectively.

[0066] In a specific embodiment, step S4 comprises:

[0067] The IPTV set-top box collects and processes the user's face features, voice features, and gesture features in real time to obtain a user multi-modal biometric feature data set.

[0068] The user multi-modal biometric feature dataset is input into a face recognition algorithm, a speech recognition algorithm and a gesture recognition algorithm respectively for feature matching and similarity calculation processing to obtain original confidence scores of each modality;

[0069] The original confidence scores of each modality are subjected to weight distribution and weighted summation calculation processing according to high-sensitivity authentication parameters or low-sensitivity authentication parameters to obtain a comprehensive confidence score after weighted fusion;

[0070] Based on the comprehensive confidence score, an association marking process is performed on the current learning state of the user to obtain an identity confidence score associated with the learning state.

[0071] Specifically, the IPTV set-top box collects the user's biometric features in all directions through an integrated multi-sensor module. The face feature collection captures image data of the user's face area through a high-definition camera and extracts key feature point coordinates, face contour information and texture features, etc. The speech feature collection records audio waveform data when the user speaks through a digital microphone array and analyzes acoustic parameters such as frequency spectrum distribution, fundamental frequency change and formant position. The gesture feature collection tracks the user's hand movement trajectory through a depth sensor or an infrared sensor and records motion parameters such as joint angle change, movement speed and spatial position. The three feature collection processes are synchronized to ensure the consistency of the time stamp. The collected raw data is preprocessed, including noise filtering, data format standardization and feature vector normalization, etc. Finally, the user multi-modal biometric feature dataset containing face feature vectors, speech feature vectors and gesture feature vectors is formed. The data structure of this dataset adopts a unified feature descriptor format to facilitate subsequent algorithm processing.

[0072] The face recognition algorithm receives the face feature data in the user multi-modal biometric feature dataset, and first performs feature point detection and face region segmentation. The algorithm encodes the face image through a deep neural network model to generate a high-dimensional feature vector. The feature vector is compared with the face template stored during the user's pre-registration. The cosine similarity measurement method is used to calculate the cosine value of the angle between the two feature vectors to determine the similarity. The closer the similarity value is to 1, the higher the matching degree. The calculation result is converted into the original confidence score of face recognition. The speech recognition algorithm performs frequency domain analysis and time domain analysis on the speech feature data. The algorithm extracts the Mel frequency cepstrum coefficient of the audio signal as the voiceprint feature. The feature is matched with the pre-stored voiceprint template through the dynamic time warping algorithm. This algorithm can handle the nonlinear alignment problem of the speech signal on the time axis. The calculation process generates the optimal path distance between the two voiceprint feature sequences. The distance value is converted into the original confidence score of speech recognition. The gesture recognition algorithm analyzes the geometric features of the hand motion trajectory, including trajectory length, curvature change, and motion direction parameters. The algorithm combines these parameters into a gesture feature descriptor. The descriptor is matched with the standard patterns in the preset gesture template library. The matching process calculates the Euclidean distance between the feature descriptor and each standard pattern. The minimum distance corresponds to the matching degree, which is converted into the original confidence score of gesture recognition.

[0073] The weight assignment and weighted summation calculation process performs mathematical operations on the original confidence scores of each modality based on the high-sensitivity authentication parameters or low-sensitivity authentication parameters obtained in the previous steps. When using high-sensitivity authentication parameters, the original confidence score of face recognition is multiplied by a weight of 0.6, and the original confidence score of speech recognition is multiplied by a weight of 0.4. Gesture recognition does not participate in the calculation because the weight is 0. The weighted summation process adds the weighted confidence scores of each modality to obtain the comprehensive confidence score. When using low-sensitivity authentication parameters, the original confidence score of face recognition is multiplied by a weight of 0.4, the original confidence score of speech recognition is multiplied by a weight of 0.3, and the original confidence score of gesture recognition is multiplied by a weight of 0.3. The three weighted values are added to form the comprehensive confidence score. This calculation process ensures the reasonable allocation of the contribution of each modality under different sensitivity modes.

[0074] The association mark processing generates identity confidence data with learning state information based on the corresponding relationship between the comprehensive confidence score and the current learning state of the user. The processing process first reads the current learning state mark of the user, including specific marks of focused learning state or distracted state, and then binds the comprehensive confidence score with the learning state mark to form a composite data structure. The binding process adds metadata information such as learning state identifier, timestamp, and confidence validity period on the basis of the comprehensive confidence score. The association mark ensures the corresponding relationship between the identity authentication result and the specific learning context. The learning state associated identity confidence score generated by the mark processing includes the original confidence value and the associated learning state context information. For example, when the college student Zhou views the higher mathematics course, the IPTV set-top box simultaneously collects his face image, voice recording, and gesture action to form a multi-modal biometric feature data set. The face recognition algorithm obtains the face recognition original confidence score by comparing the similarity between the current face features of Zhou and the registered template. The voice recognition algorithm obtains the voice recognition original confidence score by analyzing the matching degree between the pronunciation of reading and the personal voiceprint template. The gesture recognition algorithm obtains the gesture recognition original confidence score by recognizing the writing and pointing actions. Since Zhou is in a focused learning state and the importance of the mathematics course is high, the high sensitivity authentication parameter is triggered. The weight distribution multiplies the face recognition confidence score by 0.6 and the voice recognition confidence score by 0.4 for weighted summation to obtain the comprehensive confidence score. The association mark processing binds the comprehensive confidence score with the focused learning state mark to generate the final learning state associated identity confidence score.

[0075] In a specific embodiment, the execution step of inputting the user multi-modal biometric feature data set into the face recognition algorithm, the voice recognition algorithm, and the gesture recognition algorithm for feature matching and similarity calculation processing can specifically include the following steps:

[0076] Extracting face image data from the user multi-modal biometric feature data set, positioning feature points in the key areas of the face, and encoding feature vectors through the face recognition algorithm to obtain standardized face feature vectors;

[0077] Calculating the vector distance between the standardized face feature vectors and the user pre-registered face template, and quantifying the matching degree through the cosine similarity formula to obtain the face recognition original confidence score;

[0078] Extracting voice audio data from the user multi-modal biometric feature data set, performing frequency domain transformation and acoustic feature parameter extraction processing on the voice signal through the voice recognition algorithm to obtain a set of voice feature parameters;

[0079] Comparing the set of voice feature parameters with the user pre-registered voiceprint template, and calculating the similarity score through the statistical matching algorithm to obtain the voice recognition original confidence score;

[0080] Gesture action data is extracted from the user multi-modal biometric data set, and the hand movement trajectory is analyzed for geometric features and action pattern recognition through a gesture recognition algorithm to obtain gesture feature descriptors;

[0081] The gesture feature descriptors are matched with a preset gesture action template library, and the matching degree is evaluated through a similarity calculation function to obtain the gesture recognition original confidence score.

[0082] Specifically, the face image data extraction process separates the special face image information from the user multi-modal biometric data set, which contains a sequence of continuous face image frames during the user's viewing of educational content. After the face recognition algorithm receives these image data, it first pre-processes each image, including grayscale conversion, contrast adjustment, and noise removal. Then the algorithm locates the face bounding box and extracts the key regions in the pre-processed image. The face key regions include eye, nose, mouth, and face contour, etc. important feature regions. The feature point positioning process identifies 68 standard feature points of the face through a trained deep learning model. The coordinate positions of these feature points reflect the geometric structure and expression state of the face. The feature vector encoding process integrates the coordinate information, regional texture features, and geometric relationship parameters of all feature points into a high-dimensional numerical vector. The encoding process uses normalization to ensure the numerical range of the feature vector uniform. The standardized face feature vector generated finally has a fixed dimension and numerical range, which is convenient for subsequent comparison and calculation. The vector distance calculation process performs mathematical operations on the standardized face feature vector generated currently and the face template vector stored during the user's pre-registration. The calculation process first ensures that the two vectors have the same dimension structure, and then calculates the cosine value of the included angle between the two vectors through the cosine similarity formula. The formula calculates the dot product of the two vectors divided by the product of their lengths. The cosine similarity value ranges from negative one to positive one. The closer the value is to positive one, the more similar the two vectors are. The matching degree quantification process converts the cosine similarity result into a confidence score ranging from 0 to 1. The conversion process adds one to the cosine similarity and then divides by two to obtain the face recognition original confidence score.

[0083] The voice audio data extraction separates the pure voice signal from the user's multi-modal biometric feature data set, which contains audio information such as the user's pronunciation, reading and voice interaction during the learning process. The voice recognition algorithm performs digital signal processing on the extracted audio data. The frequency domain transformation process converts the time-domain audio signal into a frequency-domain representation through fast Fourier transform. The transformed frequency spectrum data shows the energy distribution of different frequency components. The acoustic feature parameter extraction process calculates key parameters such as mel-frequency cepstral coefficients, linear predictive coding coefficients and fundamental frequency variation trajectories from the frequency spectrum data. These parameters constitute the voice feature parameter set, which reflects the user's personal voiceprint characteristics. The parameter set is standardized to ensure the consistency and comparability of the numerical values. The feature comparison process compares the current extracted voice feature parameter set with the user's pre-registered voiceprint template item by item. The voiceprint template contains the standard acoustic feature parameter benchmark values of the user. The statistical matching algorithm calculates the statistical correlation and difference between the two sets of parameters. The algorithm analyzes the mean difference, variance comparison and distribution similarity of each parameter. The similarity score calculation integrates the matching degree of all parameters into a single numerical score. The scoring process uses a weighted average method to distinguish the importance of different parameters. Important parameters such as fundamental frequency and formant receive higher weights, while auxiliary parameters such as energy variation receive lower weights. Finally, the original confidence score of voice recognition is calculated.

[0084] The gesture action data extraction separates the trajectory information of the hand movement from the user's multi-modal biometric feature data set, which contains the spatial coordinate sequence of the user's pointing, writing, waving, and other actions while watching the educational content. The gesture recognition algorithm analyzes the geometric features of the movement trajectory data, including the length, curvature, turning angle, and movement speed of the trajectory. The action pattern recognition process identifies specific gesture types by analyzing the temporal characteristics and spatial distribution patterns of the trajectory. The recognition process divides the continuous movement trajectory into meaningful action units, each corresponding to a basic gesture action. The gesture feature descriptor integrates all geometric parameters and action pattern information into a multi-dimensional feature vector. The descriptor uses a unified data format to ensure compatibility with the template library. The pattern matching process compares the gesture feature descriptor with the standard patterns in the preset gesture action template library. The template library contains standard feature descriptors of common gesture actions. The comparison process calculates the feature distance between the descriptor and each template. The distance calculation considers the numerical difference of geometric parameters and the structural similarity of action patterns. The similarity calculation function converts the feature distance into a similarity value. The function uses an exponential decay model to make the similarity higher when the distance is smaller. The matching degree evaluation process selects the template with the highest similarity as the matching result and converts the evaluation result into the original confidence score of gesture recognition. For example, when a vocational school student Liu watches a mechanical drawing course, the IPTV set-top box camera captures his facial image and locates 68 feature points such as eye corners and nose tips through the facial recognition algorithm. The algorithm encodes these feature point coordinates into a standardized facial feature vector. This vector and the facial template of Liu registered at the time are compared through cosine similarity to obtain the original confidence score of facial recognition. At the same time, the microphone collects the speech of Liu reading professional terms, and the speech recognition algorithm extracts the voiceprint feature parameters of the audio through frequency domain transformation. These parameters are compared with the pre-stored voiceprint template through statistical matching algorithm to obtain the original confidence score of speech recognition. In addition, the depth sensor tracks the action trajectory of Liu drawing a figure in the air with his fingers, and the gesture recognition algorithm analyzes the geometric features of the trajectory to generate a gesture feature descriptor. This descriptor is matched with the drawing gesture template library to obtain the original confidence score of gesture recognition.

[0085] In a specific embodiment, step S5 comprises:

[0086] According to the user's current learning state and the importance of the educational content, the authentication security requirements are evaluated by threshold value adjustment processing to obtain a dynamic authentication threshold value;

[0087] The identity confidence score associated with the learning state is compared with the dynamic authentication threshold value to determine whether the authentication is passed or failed.

[0088] When the preliminary determination result is that the authentication is passed, record the user identity information and learning state identifier, and obtain an authentication success identifier through identity confirmation data encapsulation processing;

[0089] When the preliminary determination result is that the authentication is failed, trigger a re-authentication process or a degraded access permission setting, and obtain a user identity authentication result based on learning concentration through authentication state output processing.

[0090] Specifically, the level evaluation processing first analyzes the specific type of the current learning state of the user, including the detailed classification of the focused learning state or the distracted state. The focused learning state is subdivided into three sub-levels of high focus, medium focus and low focus according to the degree of attention concentration, and the distracted state is subdivided into three sub-levels of slight distraction, obvious distraction and serious distraction according to the degree of distraction. Each learning state sub-level corresponds to different basic security requirement values. The basic security requirement value of the high focus state is set to a higher value to reflect the need for strict authentication, and the basic security requirement value of the serious distraction state is set to a lower value to avoid excessive authentication interference. Then, the evaluation processing reads the importance coefficient of the current education content. The coefficient is quantified according to the core position of the course in the entire education system. The importance coefficient of the core compulsory course such as basic mathematics and language is set to a high value, and the importance coefficient of the elective extension course such as interest and hobby is set to a low value. The threshold value adjustment processing multiplies the basic security requirement value corresponding to the learning state by the importance coefficient of the education content to obtain a dynamic authentication threshold. The adjustment process ensures that important education content maintains a higher authentication standard in any learning state, and general education content appropriately raises the authentication requirement in the focused state and appropriately reduces the requirement in the distracted state.

[0091] The numerical value comparison and determination processing receives the identity confidence score associated with the learning state and the dynamic authentication threshold generated in the foregoing steps. The comparison process determines whether the identity confidence score meets the authentication requirement through a simple mathematical size determination operation. When the identity confidence score associated with the learning state is greater than the dynamic authentication threshold, it is determined that the authentication is passed, indicating that the user identity verification is successful. When the identity confidence score is less than or equal to the dynamic authentication threshold, it is determined that the authentication is failed, indicating that the user identity verification does not meet the standard. The result of the comparison and determination directly affects the subsequent authentication process. The determination result is stored in the temporary cache, including the determination type, the comparison value and the determination timestamp, etc. The preliminary determination result is an important basis for authentication decision to guide the execution direction of the subsequent processing steps.

[0092] The identity confirmation data encapsulation process performs a recording operation of user identity information and learning state identification immediately after receiving the preliminary determination result of authentication success, the user identity information including detailed data such as a user unique identifier, an authentication timestamp, an authentication confidence score and a used authentication modal combination, and the learning state identification containing context information such as a current specific learning state type, a learning content identification and a state duration, the encapsulation process integrates these scattered data into a unified data structure, the data structure adopts a standardized format to ensure compatibility with other modules, the encapsulated authentication success identification contains complete authentication session information and state context, the identification is stored in an authentication record database as a credential of user successful identity verification, and the generation of the authentication success identification marks the successful completion of the entire authentication process.

[0093] The authentication state output process determines a subsequent processing strategy according to the specific reason and degree of failure after receiving the preliminary determination result of authentication failure, triggers a re-authentication process when the authentication failure is due to a confidence score slightly lower than a threshold value and a small gap, the re-authentication process including measures such as re-acquiring user biometric data, adjusting authentication parameter settings and extending an authentication time window, triggers a degraded access permission setting when the authentication failure is due to a confidence score significantly lower than a threshold value or continuous multiple failures, the degraded process degrades the user's access permission from full access to limited access, the limited access including measures such as limiting viewing of advanced education content, prohibiting modification of personalized settings and adding an additional authentication link, and the authentication state output process generates a corresponding user identity authentication result based on learning concentration according to different failure conditions, the result containing complete information such as authentication state, failure reason, recommended measures and validity period.

[0094] The above describes the television-side education application multi-modal user authentication method in the embodiments of the present application, and the following describes the television-side education application multi-modal user authentication system in the embodiments of the present application, please refer to Figure 2 An embodiment of the television-side education application multi-modal user authentication system in the embodiments of the present application includes:

[0095] The monitoring module is configured to perform synchronous monitoring and processing on the face orientation deviation and voice interaction frequency of the user when watching the education content through the IPTV set-top box, and obtain a learning participation index.

[0096] The judgment module is configured to compare and judge the learning participation index with a preset participation threshold value, if the learning participation index is greater than 0.7, execute the focused learning authentication process, and if the learning participation index is less than or equal to 0.7, execute the distraction state authentication process.

[0097] An adjusting module is configured to perform security level dynamic adjustment processing on the focused learning authentication process or the distraction state authentication process through an authentication strictness adjuster, so as to obtain a high-sensitivity authentication parameter or a low-sensitivity authentication parameter.

[0098] A fusion module is configured to perform weighted fusion calculation processing on the multi-modal biological features of the user according to the high-sensitivity authentication parameter or the low-sensitivity authentication parameter, so as to obtain a learning state associated identity confidence score.

[0099] A verification module is configured to perform matching verification processing on the learning state associated identity confidence score and a dynamic authentication threshold, so as to obtain a user identity authentication result based on learning concentration.

[0100] The above Figure 2 The television terminal education application multi-modal user authentication system in the embodiment of the application is described in detail from the perspective of a modular functional entity, and the television terminal education application multi-modal user authentication device in the embodiment of the application is described in detail from the perspective of hardware processing.

[0101] Reference Figure 3 In the embodiment of the application, a television terminal education application multi-modal user authentication device is also provided, which can be a server, and the internal structure thereof can be as shown in Figure 3 The television terminal education application multi-modal user authentication device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. The processor of the computer is configured to provide computing and control capabilities. The memory of the television terminal education application multi-modal user authentication device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the television terminal education application multi-modal user authentication device is configured to store the corresponding data in the embodiment. The network interface of the television terminal education application multi-modal user authentication device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the above method.

[0102] Those skilled in the art can understand Figure 3 The structure shown in the above

[0103] The application further provides a computer readable storage medium, which can be a nonvolatile computer readable storage medium or a volatile computer readable storage medium, and the computer readable storage medium stores instructions, and the instructions make a computer execute the steps of the television-side education application multi-modal user authentication method when the instructions are run on the computer.

[0104] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, system and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described here.

[0105] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the application or the whole or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for making a television-side education application multi-modal user authentication device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk and various program code storage media.

[0106] The above embodiments are only used to illustrate the technical solutions of the application, rather than limit them; although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the application.

Claims

1. A multimodal user authentication method for television-based educational applications, characterized in that, The method includes: Step S1: Simultaneously monitor and process the facial orientation deviation and voice interaction frequency of users when watching educational content through the IPTV set-top box to obtain the learning participation index. Step S2: Compare the learning engagement index with the preset engagement threshold. If the learning engagement index is greater than 0.7, execute the focused learning certification process; if it is less than or equal to 0.7, execute the distraction certification process. Step S3: The focused learning authentication process or the distracted state authentication process is dynamically adjusted in terms of security level through an authentication severity regulator to obtain high-sensitivity authentication parameters or low-sensitivity authentication parameters. This includes: inputting the focused learning authentication process or the distracted state authentication process into the authentication severity regulator for process type identification and security level mapping to obtain an initial security level identifier, wherein the focused learning authentication process is mapped to a high initial security level identifier value; and the distracted state authentication process is mapped to a low initial security level identifier value. Based on the initial security level identifier and the importance coefficient of the current educational content, the security level is dynamically calibrated to obtain the calibrated security level value. The high-sensitivity authentication parameters use facial recognition as the primary authentication method, and the weight allocation of the high-sensitivity authentication parameters follows the principle of clear distinction between primary and secondary parameters to ensure that the key authentication modality has sufficient influence. The low-sensitivity authentication parameters adopt a three-modal balanced allocation strategy to reduce the dependence on a single modality. Step S4: Perform weighted fusion calculation on the user's multimodal biometric features based on the high-sensitivity authentication parameters or low-sensitivity authentication parameters to obtain the identity confidence score associated with the learning state; Step S5: Match and verify the identity confidence score associated with the learning state with the dynamic authentication threshold to obtain the user identity authentication result based on learning focus.

2. The multimodal user authentication method for television-based educational applications according to claim 1, characterized in that, Step S1 includes: The user's facial area is captured and processed continuously using the camera integrated into the IPTV set-top box to obtain facial image sequence data. Based on the facial key point detection algorithm, the facial image sequence data is processed to locate key points and calculate inter-frame displacement to obtain the facial orientation offset. The voice signal collected by the microphone of the IPTV set-top box is input into the voice activity detection algorithm for effective voice segment recognition and frequency statistics processing to obtain the voice interaction frequency; The learning engagement index is obtained by comprehensively evaluating the facial orientation offset and voice interaction frequency through weighted fusion calculation.

3. The multimodal user authentication method for television-based educational applications according to claim 1, characterized in that, Step S2 includes: The learning participation index is compared with the preset participation threshold of 0.7, and the learning status is corrected by combining the difficulty coefficient of the currently played educational content to obtain the corrected learning status evaluation value. The focused learning certification process is activated when the corrected learning status evaluation value is greater than 0.7, and the user's attention duration parameter is recorded to obtain the focused learning certification process. The distraction state authentication process is activated when the corrected learning state evaluation value is less than or equal to 0.7, and the biometric collection frequency is reduced to 0.5 times the standard frequency to obtain the distraction state authentication process. Based on the matching and analysis of the focused learning authentication process or the distracted state authentication process with the user's historical learning behavior patterns, personalized authentication strategy configuration parameters are obtained. The personalized authentication strategy configuration parameters are input into the authentication process scheduler for process priority sorting and execution order arrangement to obtain the authentication execution plan.

4. The multimodal user authentication method for television-based educational applications according to claim 1, characterized in that, Step S3 includes: Based on the calibration of the security level value being greater than 0.8, high-sensitivity authentication parameters are set, including a face recognition weight of 0.6 and a voice recognition weight of 0.4, to obtain the high-sensitivity authentication parameters; When the calibrated security level value is less than or equal to 0.8, low-sensitivity authentication parameters are set, including a face recognition weight of 0.4, a voice recognition weight of 0.3, and a gesture recognition weight of 0.3, to obtain the low-sensitivity authentication parameters.

5. The multimodal user authentication method for television-based educational applications according to claim 1, characterized in that, Step S4 includes: The user's facial features, voice features, and gesture features are collected and processed in real time through IPTV set-top boxes to obtain a user multimodal biometric dataset; The user's multimodal biometric dataset is input into the facial recognition algorithm, speech recognition algorithm, and gesture recognition algorithm respectively for feature matching and similarity calculation to obtain the original confidence score of each modality; The original confidence scores of each modality are weighted and summed according to the high-sensitivity authentication parameters or low-sensitivity authentication parameters to obtain the weighted fusion comprehensive confidence score. Based on the comprehensive confidence score and the user's current learning state, an identity confidence score associated with the learning state is obtained.

6. The multimodal user authentication method for television-based educational applications according to claim 5, characterized in that, The process of inputting the user's multimodal biometric dataset into facial recognition, speech recognition, and gesture recognition algorithms for feature matching and similarity calculation to obtain the original confidence score for each modality includes: Facial image data is extracted from the user's multimodal biometric dataset, and feature point localization and feature vector encoding are performed on key facial regions using a facial recognition algorithm to obtain a standardized facial feature vector. The standardized facial feature vector is compared with the user's pre-registered facial template by calculating the vector distance. The matching quantification is then performed using the cosine similarity formula to obtain the original confidence score for facial recognition. Speech audio data is extracted from the user's multimodal biometric dataset, and the speech signal is processed by frequency domain transformation and acoustic feature parameter extraction using a speech recognition algorithm to obtain a set of speech feature parameters; The set of speech feature parameters is compared with the user's pre-registered voiceprint template, and a similarity score is calculated using a statistical matching algorithm to obtain the original confidence score for speech recognition. Gesture action data is extracted from the user's multimodal biometric dataset, and geometric feature analysis and action pattern recognition processing of hand movement trajectory are performed using a gesture recognition algorithm to obtain gesture feature descriptors; The gesture feature descriptor is matched with a preset gesture action template library, and the matching degree is evaluated by a similarity calculation function to obtain the original confidence score of gesture recognition.

7. The multimodal user authentication method for television-based educational applications according to claim 1, characterized in that, Step S5 includes: The authentication security requirements are assessed based on the user's current learning status and the importance of the educational content. The dynamic authentication threshold is obtained by adjusting the threshold value. The identity confidence score associated with the learning state is compared with the dynamic authentication threshold to obtain a preliminary determination result of whether the authentication is successful or failed. Based on the preliminary judgment result that user identity information and learning status identifier are recorded when authentication is successful, an authentication success identifier is obtained through identity confirmation data encapsulation and processing. Based on the preliminary judgment result that a re-authentication process or downgraded access permissions is triggered when authentication fails, the user identity authentication result based on learning focus is obtained through authentication status output processing.

8. A multimodal user authentication system for television-based educational applications, characterized in that, For implementing the multimodal user authentication method for television-based educational applications as described in any one of claims 1-7, the multimodal user authentication system for television-based educational applications comprises: The monitoring module is used to simultaneously monitor and process the facial orientation deviation and voice interaction frequency of users when watching educational content through the IPTV set-top box, and obtain the learning participation index. The judgment module is used to compare the learning participation index with the preset participation threshold. If the learning participation index is greater than 0.7, the focused learning certification process is executed; if it is less than or equal to 0.7, the distraction state certification process is executed. The adjustment module is used to dynamically adjust the security level of the focused learning authentication process or the distracted state authentication process through the authentication severity regulator to obtain high-sensitivity authentication parameters or low-sensitivity authentication parameters. The fusion module is used to perform weighted fusion calculation on the user's multimodal biometrics based on the high-sensitivity authentication parameters or low-sensitivity authentication parameters to obtain the identity confidence score associated with the learning state. The verification module is used to match and verify the identity confidence score associated with the learning state with the dynamic authentication threshold to obtain the user identity authentication result based on learning focus.

9. A multimodal user authentication device for television-based educational applications, characterized in that, The device includes a memory and a processor, the memory storing a computer program that can run on the processor, and the processor executing the computer program to implement the multimodal user authentication method for television-based educational applications as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it causes the processor to execute the multimodal user authentication method for television-based educational applications as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Screen interaction method and system based on face recognition

    CN120045076A

  • Multi-modal data fusion identity authentication and security monitoring system based on deep learning

    CN120048010A