Non-inductive user identification and adaptive training method and system based on AI display screen
By capturing user images and posture videos through an AI display screen and combining facial texture features and skeletal key point data, the system achieves seamless user recognition and adaptive training for smart sports devices. This solves the problems of low identity recognition and adaptability in existing devices and improves training effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DONGGUAN BOQUN ELECTRONIC SCI & TECH CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-29
AI Technical Summary
Existing smart sports devices cannot automatically identify users when they come into contact with them, and lack the ability to monitor users' real-time physiological state and movement posture without being noticed. This results in low adaptability during training and affects training effectiveness.
By using the visual sensors of the AI display screen to collect frontal facial images and human posture videos of users, facial texture features and skeletal key point data are extracted. Combined with physiological feedback signals and behavioral performance data, user identification, action standardization analysis and physiological fatigue assessment are performed to achieve adaptive adjustment of training parameters.
It improves the accuracy of user identification and the adaptability of training, enhances the personalized adaptation capability of the training process, and improves training results.
Smart Images

Figure CN122097928A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a non-contact user recognition and adaptive training method and system based on an AI display screen, belonging to the field of artificial intelligence technology. Background Technology
[0002] Most smart sports devices on the market today use pre-set training programs and fixed rules to control operation. Users need to manually select the course mode and adjust the device parameters themselves. The data is then synchronized to the mobile terminal to generate a static report after the exercise is completed.
[0003] However, this type of device cannot automatically identify the user's identity upon contact and requires login via account, card, or other means. During training, it lacks real-time, undetectable, and synchronous monitoring of the user's physiological state and movement posture, such as heart rate, breathing, and posture. It cannot perceive the user's immediate load and movement quality, making it impossible to analyze the user's physical condition and fatigue level for the day. Although some smart sports devices on the market have certain facial recognition, posture detection, or heart rate monitoring functions, these functions often operate in isolation and cannot work together, resulting in low adaptability during training and reducing the user's training effect. Summary of the Invention
[0004] This invention provides a method and system for contactless user recognition and adaptive training based on an AI display screen, the main purpose of which is to improve the accuracy of contactless user recognition and adaptive training based on an AI display screen.
[0005] To achieve the above objectives, the present invention provides a non-contact user recognition and adaptive training method based on an AI display screen, comprising:
[0006] The AI display of the motion device seamlessly captures images of the user's frontal face and videos of their posture using its visual sensors. After extracting the component-related descriptors of the facial contour region in the frontal face image, the facial texture features corresponding to the user are generated. The user identity is determined by combining the facial texture features with the preset user pre-stored feature library. Extract skeletal key point data from the human posture video to determine the user's joint position and limb angle, retrieve the training template related to the user's identity, and analyze the user's movement standardization by combining the joint position, limb angle and training template. The auxiliary sensors in the AI display screen are used to collect the user's physiological feedback signals and behavioral performance data in real time. The physiological fatigue level of the user is analyzed by combining the physiological feedback signals and behavioral performance data. By combining the physiological fatigue level and the degree of proper movement, the motion parameters of the exercise equipment are adaptively adjusted to obtain the adjustment result.
[0007] Optionally, after extracting the component-related descriptors of the facial contour region in the frontal face image, generating the facial texture features corresponding to the user further includes: Key points are located on the face in the frontal face image to obtain facial key points; Based on the facial key points, predefined components are divided within the facial contour area of the frontal facial image; Extract component-related descriptors from the predefined components, and analyze the contribution of the predefined components to identity determination based on the component-related descriptors; Based on the contribution, the component-related descriptors are subjected to weighted pooling to obtain the component-weighted texture vector; Based on the weighted texture vector of the component, the facial texture features corresponding to the user are generated.
[0008] Optionally, the step of analyzing the contribution of the predefined component to identity determination based on the component-related descriptor includes: Calculate the inter-class and intra-class differences corresponding to the component-related descriptors; By combining the inter-class difference value and the intra-class difference value, the discrimination ability score corresponding to the predefined component is calculated; Based on the discrimination ability score, the contribution of the predefined components to identity discrimination is analyzed.
[0009] Optionally, the step of performing weighted pooling on the component-related descriptors based on the contribution to obtain a component-weighted texture vector includes: The component-related descriptors are standardized to obtain standard coded values; Count the number of valid pixels corresponding to the predefined component in the frontal face image; Combining the contribution, the standard encoding value, and the number of effective pixels, the component-related descriptors are weighted and pooled using the following formula to obtain the component-weighted texture vector:
[0010] in, This represents the weighted texture vector of the component. This represents the number of valid pixels in the k-th component of the predefined components, where i represents the sequence number of the pixel. This represents the contribution of the k-th component among the predefined components. This represents the i-th pixel within the k-th component. The corresponding standard encoding value.
[0011] Optionally, determining the user identity by combining the facial texture features with a preset user pre-stored feature library includes: Identify the feature identifier corresponding to the facial texture features, and extract the registration feature vector corresponding to the feature identifier from a preset user-stored feature library; Construct the correlation matrices corresponding to the facial texture features and the registered feature vectors respectively to obtain the texture correlation matrix and the registration correlation matrix; The maximum similarity between the texture association matrix and the registration association matrix is calculated using the following formula:
[0012] Where A represents the maximum similarity between the texture association matrix and the registration association matrix. Texture correlation matrix representing facial texture features. This represents the i-th registration association matrix in the user's pre-stored feature library. The matrix norm corresponding to the texture correlation matrix representing facial texture features. Let the matrix norm of the i-th registration association matrix be denoted as . Represents matrix norm operations; When the maximum similarity is greater than a preset threshold, the user identity corresponding to the user is determined based on the registration feature vector.
[0013] Optionally, extracting skeletal keypoint data from the human posture video includes: The human posture video is segmented into frames to obtain posture video frames; Extract the joint heatmap, joint location association map, root node depth map, and location relative depth map from the posture video frame; Two-dimensional joint node coordinates are detected from the joint heat map, and limb vector fields are extracted from the joint location association map. By combining the coordinates of the two-dimensional joint nodes and the limb vector field, a two-dimensional human skeleton of the user is constructed; By combining the absolute depth information of the root node depth map and the relative limb depth information of the part relative depth map, the two-dimensional human skeleton is subjected to three-dimensional generation processing to obtain skeletal key point data.
[0014] Optionally, the step of analyzing the user's movement standardization by combining the joint point positions, limb angles, and training templates includes: The joint point positions and limb angles are time-synchronized with the training template to obtain aligned joint point positions and aligned limb angles. Calculate the distance deviation between the alignment joint node position and each node in the training template; Calculate the cosine similarity between the aligned limb angle and the limb angle in the training template; By combining the distance deviation value and the cosine similarity, the user's behavior regularity is analyzed.
[0015] Optionally, the step of combining the physiological feedback signals and the behavioral performance data to analyze the user's physiological fatigue includes: Extract the cardiopulmonary function (CPF) features from the physiological feedback signals, and determine the user's exercise physiological load based on the CPF features; Based on the behavioral performance data, analyze the quality of the user's actions; By combining the exercise physiological load and the quality of the action execution, the user's current exercise status score is evaluated; Based on the current exercise status score, the user's physiological fatigue level is analyzed.
[0016] Optionally, analyzing the user's action execution quality based on the behavioral performance data includes: The behavioral performance data is segmented into homogeneous action behavior data, and homogeneous behavioral features are extracted from the homogeneous action behavior data; Calculate the behavioral deviation rate of each behavioral feature in the homologous behavioral features over multiple periods; Based on the behavior deviation rate, the quality of the user's action execution is analyzed.
[0017] To address the aforementioned problems, this invention also provides a non-contact user recognition and adaptive training system based on an AI display screen, the system comprising: The data acquisition module is used to non-intrusively acquire frontal facial images and human posture videos of users using the visual sensors of the AI display screen of the motion device; The identity determination module is used to extract component-related descriptors of the facial contour region in the frontal face image, generate facial texture features corresponding to the user, and combine the facial texture features with a preset user pre-stored feature library to determine the user's identity. The standardization analysis module is used to extract skeletal key point data from the human posture video to determine the user's joint position and limb angle, retrieve the training template related to the user's identity, and analyze the user's movement standardization by combining the joint position, the limb angle and the training template. The fatigue analysis module is used to collect the user's physiological feedback signals and behavioral performance data in real time using the auxiliary sensors in the AI display screen, and to analyze the user's physiological fatigue level by combining the physiological feedback signals and behavioral performance data. The training adjustment module is used to adaptively adjust the motion parameters of the exercise equipment by combining the physiological fatigue level and the motion standardization, and obtain the adjustment result.
[0018] Compared to the problems described in the background art, this invention utilizes the visual sensor of an AI display screen in a motion device to simultaneously record the user's facial image and body posture video while the user is viewing screen content, providing data support for subsequent processing. This invention extracts component-related descriptors from the facial contour region of the frontal facial image to generate facial texture features corresponding to the user, transforming the texture information of the facial image into representative features. This yields more discriminative regions within the facial edges, thereby improving the accuracy of subsequent user identification. Furthermore, this invention extracts skeletal keypoint data from the human posture video to determine the user's joint positions and limb angles, thus obtaining the user's... The invention first obtains kinematic information about the user in space, thus providing data support for the subsequent analysis of the standardization of movements. Next, by combining the physiological feedback signals and behavioral performance data, the invention analyzes the user's physiological fatigue level, understanding the degree of physical fatigue during exercise, and providing a reference for subsequent training adjustments to avoid overtraining. Then, by combining the physiological fatigue level and the standardization of movements, the invention adaptively adjusts the motion parameters of the exercise device, allowing the training content of the exercise device to be adapted to the user's physical state, thereby improving the user's training efficiency. Therefore, the invention can improve the accuracy of seamless user recognition and adaptive training based on AI displays. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating a method for contactless user recognition and adaptive training based on an AI display screen, according to an embodiment of the present invention. Figure 2 Reference diagrams corresponding to skeletal key point data for a non-intrusive user recognition and adaptive training method based on an AI display screen provided in the embodiments of this application; Figure 3 A root node depth map diagram for an embodiment of this application regarding a non-intrusive user recognition and adaptive training method based on an AI display screen; Figure 4 This is a schematic diagram of a module for implementing a non-contact user recognition and adaptive training method based on an AI display screen, according to an embodiment of the present invention. Figure 5A schematic diagram of a computer device for a non-contact user identification and adaptive training method based on an AI display screen, according to an embodiment of the present invention; The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0021] This application provides a method for seamless user identification and adaptive training based on an AI display screen. The execution entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0022] Reference Figure 1 The diagram shown is a flowchart illustrating a method for seamless user recognition and adaptive training based on an AI display screen, according to an embodiment of the present invention. In this embodiment, the method for seamless user recognition and adaptive training based on an AI display screen includes: S1. Use the visual sensor on the AI display of the motion device to collect the user's frontal facial image and human posture video without being noticed.
[0023] This invention utilizes the visual sensor of the AI display screen of a motion device to simultaneously record the user's facial image and body posture video while the user is viewing the screen content, providing data support for subsequent processing.
[0024] The AI display screen is a display unit device with local intelligent processing capabilities in the motion device. The visual sensor is an optical sensing device embedded in the AI display screen for capturing images of the scene in front of the screen. The frontal face image is a static image with the user's face as the main focus. The human posture video is a collection of continuous images of the user's upper body movements.
[0025] S2. After extracting the component-related descriptors of the facial contour region in the frontal face image, generate the facial texture features corresponding to the user, and combine the facial texture features with the preset user pre-stored feature library to determine the user identity.
[0026] This invention extracts component-related descriptors from the facial contour region of a frontal face image to generate facial texture features corresponding to the user. This transforms the texture information of the facial image into representative features, thereby obtaining more discriminative regions within the facial edges and improving the accuracy of subsequent user identification. The component-related descriptors are feature sets corresponding to texture details in the facial contour region of the frontal face image. Examples include the direction of fine lines at the corner of the left eye and the roughness of the skin texture on both sides of the nose—features that reflect local texture differences on the face. These features all belong to component-related descriptors.
[0027] As an embodiment of the present invention, after extracting the component-related descriptors of the facial contour region in the frontal facial image, generating the facial texture features corresponding to the user further includes: Key points are located on the face in the frontal face image to obtain facial key points; Based on the facial key points, predefined components are divided within the facial contour area of the frontal facial image; Extract component-related descriptors from the predefined components, and analyze the contribution of the predefined components to identity determination based on the component-related descriptors; Based on the contribution, the component-related descriptors are subjected to weighted pooling to obtain the component-weighted texture vector; Based on the weighted texture vector of the component, the facial texture features corresponding to the user are generated.
[0028] The facial key points are pixel coordinates located in the facial contour and facial features area, with clear anatomical coordinates; the predefined components refer to several fixed local regions divided on the image plane based on the topological relationship and physiological structure between the facial key points; the contribution degree is a weight used to quantify the discriminative ability of each predefined component in distinguishing different people's identities; the component weighted texture vector is a feature vector representing the texture attributes of a component, formed by aggregating the contribution weight of the corresponding component onto all its component-related descriptors.
[0029] Furthermore, key points of the face in the frontal face image can be located using a constrained local model, such as the Z-constrained local model. Based on these key points, predefined components within the facial contour region of the frontal face image can be divided by connecting specific key points to form polygonal regions, such as dividing the face into six basic components: left eye, right eye, nose, mouth, and both cheeks. Component-related descriptors can be extracted from these predefined components using a Gabor filter bank-based histogram of orientation or a local binary mode operator. Based on the component-weighted texture vectors, the facial texture features corresponding to the user are generated by concatenating the feature vectors of all components in spatial order.
[0030] Optionally, the step of analyzing the contribution of the predefined component to identity determination based on the component-related descriptor includes: Calculate the inter-class and intra-class differences corresponding to the component-related descriptors; By combining the inter-class difference value and the intra-class difference value, the discrimination ability score corresponding to the predefined component is calculated; Based on the discrimination ability score, the contribution of the predefined components to identity discrimination is analyzed.
[0031] The inter-class difference value and the intra-class difference value are respectively the degree of separation of component feature distribution among different identity samples corresponding to the component-related descriptor and the degree of compactness of component feature distribution within the same identity sample; the discrimination ability score is a quantitative value of the ability to distinguish identity differences corresponding to the predefined component by combining the inter-class difference value and the intra-class difference value. The higher the score, the more effectively the component can distinguish different identities.
[0032] Furthermore, the inter-class and intra-class variance values corresponding to the relevant descriptors of the components can be calculated using chi-square distance combined with class center statistics. Combining these inter-class and intra-class variance values, the discriminative ability score corresponding to the predefined components can be calculated using the Fisher criterion formula. Specifically, the square of the inter-class variance value is used as the numerator, and the sum of the squares of the intra-class variance values is used as the denominator. Then, these squares are substituted into the Fisher criterion formula to calculate the discriminative ability score. Based on the magnitude of the discriminative ability score, the contribution of the predefined components to identity discrimination is analyzed. For example, eye components (score ≥ 20) in the top 30% of the discriminative ability score contribute more than 25% to identity discrimination, while forehead components (score ≤ 10) in the bottom 30% contribute less than 8%. Components with higher scores have a more significant supporting role in distinguishing identity differences.
[0033] Optionally, the step of performing weighted pooling on the component-related descriptors based on the contribution to obtain a component-weighted texture vector includes: The component-related descriptors are standardized to obtain standard coded values; Count the number of valid pixels corresponding to the predefined component in the frontal face image; Combining the contribution, the standard encoding value, and the number of effective pixels, the component-related descriptors are weighted and pooled using the following formula to obtain the component-weighted texture vector:
[0034] in, This represents the weighted texture vector of the component. This represents the number of valid pixels in the k-th component of the predefined components, where i represents the sequence number of the pixel. This represents the contribution of the k-th component among the predefined components. This represents the i-th pixel within the k-th component. The corresponding standard encoding value.
[0035] The standard encoding value is the encoding value obtained after eliminating the differences between the component-related descriptors; the effective pixel refers to a pixel located within the predefined component polygonal region and confirmed by face segmentation to belong to skin or facial texture; furthermore, the component-related descriptors can be standardized using the Z-score normalization method to obtain the standard encoding value; the number of effective pixels corresponding to the predefined component in the frontal face image can be counted using skin color segmentation technology based on the YCbCr color space.
[0036] This invention determines the user's identity by combining the facial texture features with a pre-stored user feature library. It compares and matches the collected user facial texture features with pre-stored known user features in the system, thereby completing identity identification and confirmation. The pre-stored user feature library is a feature template library constructed during the user registration phase by collecting multiple frontal facial images and extracting corresponding facial texture features. The user identity refers to the unique identifier identified through the above comparison process that matches the user's pre-registration information.
[0037] As an embodiment of the present invention, determining the user identity corresponding to the user by combining the facial texture features and a preset user pre-stored feature library includes: Identify the feature identifier corresponding to the facial texture features, and extract the registration feature vector corresponding to the feature identifier from a preset user-stored feature library; Construct the correlation matrices corresponding to the facial texture features and the registered feature vectors respectively to obtain the texture correlation matrix and the registration correlation matrix; The maximum similarity between the texture association matrix and the registration association matrix is calculated using the following formula:
[0038] Where A represents the maximum similarity between the texture association matrix and the registration association matrix. Texture correlation matrix representing facial texture features. This represents the i-th registration association matrix in the user's pre-stored feature library. The matrix norm corresponding to the texture correlation matrix representing facial texture features. Let the matrix norm of the i-th registration association matrix be denoted as . Represents matrix norm operations; When the maximum similarity is greater than a preset threshold, the user identity corresponding to the user is determined based on the registration feature vector.
[0039] Wherein, the feature identifier is a unique identity code corresponding to the facial texture feature, used to index the corresponding registration feature vector in the user's pre-stored feature library; the registration feature vector is a high-dimensional vector corresponding to the feature identifier extracted from the preset user's pre-stored feature library; the texture association matrix and the registration association matrix are square matrices composed of the facial texture feature and the registration feature vector, respectively; the maximum similarity is a similarity measure value calculated by normalizing the difference between the texture association matrix and the registration association matrix; the preset threshold is a pre-set similarity threshold value, which can be 0.8 or can be set according to the actual application scenario.
[0040] Furthermore, the feature identifier corresponding to the facial texture features can be identified by a recognition function, and the registration feature vector corresponding to the feature identifier can be extracted from a preset user pre-stored feature library. The recognition function can be compiled using the JAVA language. The correlation matrix corresponding to the facial texture features and the registration feature vector can be constructed by a matrix function, such as the zero matrix function. When the maximum similarity is greater than a preset threshold, the user identity corresponding to the user is determined according to the registration feature vector. For example, when the similarity exceeds 0.8, the user is determined to be the account associated with the registration feature vector and is authorized to log in.
[0041] S3. Extract skeletal key point data from the human posture video to determine the user's joint position and limb angle, retrieve the training template related to the user's identity, and analyze the user's movement standardization by combining the joint position, the limb angle and the training template.
[0042] This invention extracts skeletal keypoint data from human posture videos to determine the user's joint positions and limb angles, thereby obtaining the user's kinematic information in space. This provides data support for subsequent analysis of movement accuracy. The skeletal keypoint data is a set of three-dimensional spatial coordinates of core body parts in the human posture video that characterize human movement postures. For details, please refer to the following... Figure 2 This is a reference diagram corresponding to the skeletal key point data in an embodiment of the present invention. The diagram presents a comparison between a complex human skeleton model and a simplified human skeleton model. The complex model on the left contains 25 joints, while the simplified model on the right removes joints that contribute little to motion analysis. When extracting user skeletal key points, this simplified model is referenced, which can reduce the computational load on the AI display screen, improve the efficiency of posture analysis, and focus on core joint data. This provides accurate skeletal information support for the subsequent comparison of the training movements with the standard template. The joint point position and the limb angle are the three-dimensional spatial coordinates of the user's joint connection and the angle between the limbs formed by the lines connecting adjacent joints, respectively. Furthermore, based on the skeletal key point data, the user's joint point position and limb angle are determined. A posture estimation algorithm is used to extract the coordinates of key points such as the user's shoulder, elbow, and wrist in the video frame, thereby obtaining the joint node position. Based on these three-dimensional coordinates, the real-time angle between the upper arm and forearm is calculated. The posture estimation algorithm includes the OpenPose algorithm.
[0043] As an embodiment of the present invention, the extraction of skeletal key point data from the human posture video includes: The human posture video is segmented into frames to obtain posture video frames; Extract the joint heatmap, joint location association map, root node depth map, and location relative depth map from the posture video frame; Two-dimensional joint node coordinates are detected from the joint heat map, and limb vector fields are extracted from the joint location association map. By combining the coordinates of the two-dimensional joint nodes and the limb vector field, a two-dimensional human skeleton of the user is constructed; By combining the absolute depth information of the root node depth map and the relative limb depth information of the part relative depth map, the two-dimensional human skeleton is subjected to three-dimensional generation processing to obtain skeletal key point data.
[0044] The posture video frames are continuous image frames extracted from human posture videos at a fixed frame rate; the joint heatmap is a heat map representing the probability of each joint's position on a two-dimensional plane; the joint association map is a two-dimensional vector field encoding the connection direction and probability of limbs, with one association map corresponding to each limb; the root node depth map is an image reflecting the absolute depth of the root node in the middle of the human waist. For details, please refer to the following...Figure 3 This is a schematic diagram of the root node depth map according to an embodiment of the present invention. The map labels the precise depth information of each user's root node, is not limited by the number of trainees, and does not require depth data from the entire image, directly providing core depth information for 3D generation processing. The relative depth map of the body part is an image representing the depth difference between non-root joints and their parent joints. The two-dimensional joint node coordinates are the pixel coordinates corresponding to the probability peaks extracted from the joint heatmap. The limb vector field is a set of unit vectors pointing to the joints at both ends of the limb in the joint part association map. The two-dimensional human skeleton is a tree structure formed by grouping and connecting joints. The absolute depth information is the true depth value of the middle root node of the human waist in the three-dimensional space in the root node depth map. The relative depth information of the limb is the depth difference between non-root joints and their corresponding parent joints in the relative depth map of the body part.
[0045] Furthermore, the human posture video can be segmented into frames using OpenCV-based video frame sampling technology to obtain posture video frames. A lightweight multi-branch hourglass stacked convolutional neural network can be used to extract joint heatmaps, joint location association maps, root node depth maps, and relative location depth maps from the posture video frames. Two-dimensional joint coordinates can be detected from the joint heatmap using non-maximum suppression and Gaussian center localization methods. Limb vector fields in the joint location association map can be extracted by directly reading pre-encoded vector information from the joint location association map. Combining the two-dimensional joint coordinates and the limb vector fields, a depth-aware location association algorithm is used to construct the user's two-dimensional human skeleton. Combining the absolute depth information of the root node depth map and the relative limb depth information of the relative location depth map, a coordinate back-projection and depth fusion algorithm based on a camera perspective model is used to perform three-dimensional generation processing on the two-dimensional human skeleton to obtain skeletal keypoint data.
[0046] This invention analyzes the user's movement standardization by combining the joint point positions, limb angles, and training templates. This allows for an understanding of the user's fitness movement standardization during exercise, providing a basis for subsequent adaptive training adjustments. The training template is a training reference standard movement corresponding to the user's identity; the movement standardization indicates the degree of consistency between the user's real-time movement posture and the standard template movement. Furthermore, training templates related to the user's identity can be retrieved from a preset fitness template database. This database stores standard templates for fitness program categories bound to the user's identity, including joint point positions, limb angle baseline data, and corresponding video frame features of the coach's standard movements.
[0047] As an embodiment of the present invention, the step of analyzing the user's movement standardization by combining the joint point position, the limb angle, and the training template includes: The joint point positions and limb angles are time-synchronized with the training template to obtain aligned joint point positions and aligned limb angles. Calculate the distance deviation between the alignment joint node position and each node in the training template; Calculate the cosine similarity between the aligned limb angle and the limb angle in the training template; By combining the distance deviation value and the cosine similarity, the user's behavior regularity is analyzed.
[0048] Wherein, the aligned joint node position refers to the node position after the joint point position and the standard joint point coordinates of the training template are synchronized in time; the aligned limb angle refers to the limb angle after the limb angle and the corresponding standard limb angle in the training template are synchronized in time; the distance deviation value refers to the three-dimensional Euclidean distance between the aligned user joint point and the template standard joint point; and the cosine similarity refers to the cosine value of the angle between the aligned user limb vector and the template standard limb vector.
[0049] Furthermore, the joint positions and limb angles can be synchronized with the training template by optimizing the dynamic time warping algorithm to obtain aligned joint node positions and aligned limb angles, such as the constrained dynamic time warping algorithm based on root node trajectory. The distance deviation between the aligned joint node positions and each node in the training template can be calculated using the Euclidean distance algorithm. The cosine similarity between the aligned limb angles and the limb angles in the training template can be calculated using the cosine similarity algorithm. Combining the distance deviation and the cosine similarity, the user's comprehensive posture similarity score is calculated. The user's action standardization is analyzed based on the comprehensive posture similarity score, such as a score ≥90 as "excellent", 80-89 as "good", 60-79 as "qualified", and <60 as "needs improvement".
[0050] S4. Use the auxiliary sensors in the AI display screen to collect the user's physiological feedback signals and behavioral performance data in real time, and combine the physiological feedback signals and behavioral performance data to analyze the user's physiological fatigue level.
[0051] This invention analyzes the user's physiological fatigue level by combining the physiological feedback signals and behavioral performance data. This allows for an understanding of the user's physical fatigue level during exercise, providing a reference for subsequent training adjustments and preventing overtraining. The auxiliary sensors are devices within the AI display screen used to collect physiological data, such as a monocular RGB camera for collecting the user's posture data during exercise; an infrared camera for collecting the user's heart rate data; and a millimeter-wave radar for collecting the user's breathing rhythm data. The physiological feedback signals are physiological records of the user during exercise, such as heart rate and breathing rhythm. The behavioral performance data are records of the user's posture during exercise. The physiological fatigue level represents the degree of physical fatigue experienced by the user during exercise.
[0052] As an embodiment of the present invention, the step of analyzing the user's physiological fatigue level by combining the physiological feedback signal and the behavioral performance data includes: Extract the cardiopulmonary function (CPF) features from the physiological feedback signals, and determine the user's exercise physiological load based on the CPF features; Based on the behavioral performance data, analyze the quality of the user's actions; By combining the exercise physiological load and the quality of the action execution, the user's current exercise status score is evaluated; Based on the current exercise status score, the user's physiological fatigue level is analyzed.
[0053] Wherein, the exercise cardiopulmonary characteristics are the representations of heart rate and respiration in the physiological feedback signals; the exercise physiological load refers to the quantitative value reflecting the level of internal stress experienced by the user's body, calculated based on the exercise cardiopulmonary characteristics; the action execution quality is a comprehensive level reflecting the user's neuromuscular control ability and action economy; and the current exercise state score represents the quantitative value of the user's exercise state performance.
[0054] Furthermore, cardiopulmonary features such as heart rate and respiratory rate can be extracted from the physiological feedback signal through direct extraction. These cardiopulmonary features are then compared with the corresponding load feature mapping curve to determine the user's physiological workload. The load feature mapping curve is a curve constructed from a large amount of statistical exercise data, showing the mapping relationship between cardiopulmonary features and the corresponding load. The physiological workload and the quality of the movement are normalized to obtain normalized physiological workload and movement quality values. A coordination efficiency value is calculated using the formula: movement quality value / (1 + physiological workload value). This coordination efficiency value is converted into a percentage to obtain the user's current exercise state score. Based on the current exercise state score, the user's physiological fatigue level is analyzed. If the score is below a preset threshold of 70, the user is considered to be in a "mild fatigue" state.
[0055] Optionally, analyzing the user's action execution quality based on the behavioral performance data includes: The behavioral performance data is segmented into homogeneous action behavior data, and homogeneous behavioral features are extracted from the homogeneous action behavior data; Calculate the behavioral deviation rate of each behavioral feature in the homologous behavioral features over multiple periods; Based on the behavior deviation rate, the quality of the user's action execution is analyzed.
[0056] Wherein, the source action behavior data is a set of repeatedly executed behavior data in the behavior performance data, the source behavior feature is the core quantitative feature reflecting the consistency and standardization of the action in the source action behavior data, and the behavior deviation rate represents the degree of deviation of each behavior feature in the source behavior features over multiple periods.
[0057] Furthermore, based on the behavioral labels in the behavioral performance data, homogeneous action behavior data, such as arm swinging behavior, is segmented. Homogeneous behavioral features in the homogeneous action behavior data can be extracted using feature decoupling algorithms. For example, dynamic time warping algorithms are used to achieve time synchronization alignment of multiple action cycles. Then, principal component analysis algorithms are used to separate independent feature components representing the core action pattern from the multi-dimensional raw data. Dynamic time warping algorithms include the classic DTW algorithm, and principal component analysis algorithms include the kernel PCA algorithm. By calculating the feature deviation degree of each behavioral feature in the homogeneous behavioral features over multiple cycles, the behavior deviation rate is determined based on the feature deviation degree. For example, by calculating the feature deviation degree of each behavioral feature over multiple cycles, the average value of the feature deviation degree is finally calculated to obtain the behavior deviation rate. Based on the behavior deviation rate, the user's action execution quality is analyzed. If the behavior deviation rate is less than 5%, the action execution quality is judged as excellent; 5%~15% is good; 15%~30% is qualified; and above 30% is judged as unqualified.
[0058] S5. Combining the physiological fatigue level and the standardization of the movement, adaptively adjust the motion parameters of the exercise equipment to obtain the adjustment result.
[0059] This invention combines the physiological fatigue level and the degree of proper movement to adaptively adjust the motion parameters of the exercise equipment, thereby adapting the training content of the exercise equipment to the user's physical condition and improving the user's training efficiency.
[0060] Furthermore, combining the physiological fatigue level and the movement standardization, the motion parameters of the exercise equipment are adaptively adjusted for training. The corresponding adjustment steps are as follows: If both the physiological fatigue level and the movement standardization are within the preset "good" threshold range, it indicates that the user's current physical state is stable and the movement control is accurate. It is necessary to appropriately increase the resistance or speed of the exercise equipment to increase the training load and stimulate potential. If both the physiological fatigue level and the movement standardization are below the preset "acceptable" threshold, it indicates that the user is currently over-fatigued and the movement has become deformed. It is necessary to immediately reduce the equipment resistance and shorten the current training set duration to ensure safety and prevent injury. If the physiological fatigue level and the movement standardization show an asymmetrical state of one being high and the other low, targeted adjustments are made based on which indicator is lower: for example, when the fatigue level is low and the standardization is low, the focus is on providing assistance through the equipment to correct the movement; when the fatigue level is high and the standardization is still high, the focus is on reducing the load to maintain the quality of the movement.
[0061] like Figure 4 The diagram shown is a functional block diagram of the AI-based contactless user recognition and adaptive training system of the present invention.
[0062] The AI-based seamless user recognition and adaptive training system 400 described in this invention can be installed in an electronic device. Depending on the functions implemented, the AI-based seamless user recognition and adaptive training system 400 may include a data acquisition module 401, an identity verification module 402, a standardization analysis module 403, a fatigue analysis module 404, and a training adjustment module 405. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0063] In this embodiment of the invention, the functions of each module / unit are as follows: The data acquisition module 401 is used to collect the user's frontal face image and human posture video without being noticed by the visual sensor of the AI display screen of the motion device. The identity determination module 402 is used to extract the component-related descriptors of the facial contour region in the frontal face image, generate the facial texture features corresponding to the user, and combine the facial texture features with the preset user pre-stored feature library to determine the user identity corresponding to the user. The standardization analysis module 403 is used to extract skeletal key point data from the human posture video to determine the user's joint position and limb angle, retrieve the training template related to the user's identity, and analyze the user's movement standardization by combining the joint position, the limb angle and the training template. The fatigue analysis module 404 is used to collect the user's physiological feedback signals and behavioral performance data in real time using the auxiliary sensors in the AI display screen, and to analyze the user's physiological fatigue by combining the physiological feedback signals and behavioral performance data. The training adjustment module 405 is used to adaptively adjust the motion parameters of the exercise equipment by combining the physiological fatigue level and the motion standardization, and obtain the adjustment result.
[0064] In detail, the modules in the AI-based seamless user recognition and adaptive training system 500 described in this embodiment of the invention employ the same methods as described above. Figure 1 The AI-based seamless user recognition and adaptive training method described herein uses the same technical means and can produce the same technical effect, so it will not be elaborated here.
[0065] In one embodiment, a computer device is provided, which may be a server or a client, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements functions or steps on the server or client side of an AI-based display-based contactless user recognition and adaptive training method.
[0066] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0067] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0068] Finally, it should be noted that in the above embodiments, each embodiment can be combined with each other or independent. Deleting any one of them will not affect the technical implementation of other embodiments. The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for seamless user recognition and adaptive training based on an AI display screen, characterized in that, The method includes: The AI display of the motion device seamlessly captures images of the user's frontal face and videos of their posture using its visual sensors. After extracting the component-related descriptors of the facial contour region in the frontal face image, the facial texture features corresponding to the user are generated. The user identity is determined by combining the facial texture features with the preset user pre-stored feature library. Extract skeletal key point data from the human posture video to determine the user's joint position and limb angle, retrieve the training template related to the user's identity, and analyze the user's movement standardization by combining the joint position, limb angle and training template. The auxiliary sensors in the AI display screen are used to collect the user's physiological feedback signals and behavioral performance data in real time. The physiological fatigue level of the user is analyzed by combining the physiological feedback signals and behavioral performance data. By combining the physiological fatigue level and the degree of proper movement, the motion parameters of the exercise equipment are adaptively adjusted to obtain the adjustment result.
2. The seamless user recognition and adaptive training method based on an AI display screen as described in claim 1, characterized in that, After extracting the component-related descriptors of the facial contour region in the frontal face image, generating the facial texture features corresponding to the user further includes: Key points are located on the face in the frontal face image to obtain facial key points; Based on the facial key points, predefined components are divided within the facial contour area of the frontal facial image; Extract component-related descriptors from the predefined components, and analyze the contribution of the predefined components to identity determination based on the component-related descriptors; Based on the contribution, the component-related descriptors are subjected to weighted pooling to obtain the component-weighted texture vector; Based on the weighted texture vector of the component, the facial texture features corresponding to the user are generated.
3. The method for seamless user recognition and adaptive training based on an AI display screen as described in claim 2, characterized in that, The step of analyzing the contribution of the predefined component to identity determination based on the component-related descriptor includes: Calculate the inter-class and intra-class differences corresponding to the component-related descriptors; By combining the inter-class difference value and the intra-class difference value, the discrimination ability score corresponding to the predefined component is calculated; Based on the discrimination ability score, the contribution of the predefined components to identity discrimination is analyzed.
4. The method for seamless user recognition and adaptive training based on an AI display screen as described in claim 2, characterized in that, The step of performing weighted pooling on the component-related descriptors based on the contribution to obtain a component-weighted texture vector includes: The component-related descriptors are standardized to obtain standard coded values; Count the number of valid pixels corresponding to the predefined component in the frontal face image; Combining the contribution, the standard encoding value, and the number of effective pixels, the component-related descriptors are weighted and pooled using the following formula to obtain the component-weighted texture vector: in, This represents the weighted texture vector of the component. This represents the number of valid pixels in the k-th component of the predefined components, where i represents the sequence number of the pixel. This represents the contribution of the k-th component among the predefined components. This represents the i-th pixel within the k-th component. The corresponding standard encoding value.
5. The seamless user recognition and adaptive training method based on an AI display screen as described in claim 1, characterized in that, The step of combining the facial texture features with a pre-stored user feature library to determine the user's identity includes: Identify the feature identifier corresponding to the facial texture features, and extract the registration feature vector corresponding to the feature identifier from a preset user-stored feature library; Construct the correlation matrices corresponding to the facial texture features and the registered feature vectors respectively to obtain the texture correlation matrix and the registration correlation matrix; The maximum similarity between the texture association matrix and the registration association matrix is calculated using the following formula: Where A represents the maximum similarity between the texture association matrix and the registration association matrix. The texture association matrix represents the facial texture features. This represents the i-th registration association matrix in the user's pre-stored feature library. The matrix norm corresponding to the texture correlation matrix representing facial texture features. Let the matrix norm of the i-th registration association matrix be denoted as . Represents matrix norm operations; When the maximum similarity is greater than a preset threshold, the user identity corresponding to the user is determined based on the registration feature vector.
6. The method for seamless user recognition and adaptive training based on an AI display screen as described in claim 1, characterized in that, The extraction of skeletal key point data from the human posture video includes: The human posture video is segmented into frames to obtain posture video frames; Extract the joint heatmap, joint location association map, root node depth map, and location relative depth map from the posture video frame; Two-dimensional joint node coordinates are detected from the joint heat map, and limb vector fields are extracted from the joint location association map. By combining the coordinates of the two-dimensional joint nodes and the limb vector field, a two-dimensional human skeleton of the user is constructed. By combining the absolute depth information of the root node depth map and the relative limb depth information of the part relative depth map, the two-dimensional human skeleton is subjected to three-dimensional generation processing to obtain skeletal key point data.
7. The method for seamless user recognition and adaptive training based on an AI display screen as described in claim 1, characterized in that, The analysis of the user's movement standardization, combining the joint point positions, limb angles, and training templates, includes: The joint point positions and limb angles are time-synchronized with the training template to obtain aligned joint point positions and aligned limb angles. Calculate the distance deviation between the alignment joint node position and each node in the training template; Calculate the cosine similarity between the aligned limb angle and the limb angle in the training template; By combining the distance deviation value and the cosine similarity, the user's behavior regularity is analyzed.
8. The method for seamless user recognition and adaptive training based on an AI display screen as described in claim 1, characterized in that, The analysis of the user's physiological fatigue level by combining the physiological feedback signals and the behavioral performance data includes: Extract the cardiopulmonary function (CPF) features from the physiological feedback signals, and determine the user's exercise physiological load based on the CPF features; Based on the behavioral performance data, analyze the quality of the user's actions; By combining the exercise physiological load and the quality of the action execution, the user's current exercise status score is evaluated; Based on the current exercise status score, the user's physiological fatigue level is analyzed.
9. The seamless user recognition and adaptive training method based on an AI display screen as described in claim 8, characterized in that, The analysis of the user's action execution quality based on the behavioral performance data includes: The behavioral performance data is segmented into homogeneous action behavior data, and homogeneous behavioral features are extracted from the homogeneous action behavior data; Calculate the behavioral deviation rate of each behavioral feature in the homologous behavioral features over multiple periods; Based on the behavior deviation rate, the quality of the user's action execution is analyzed.
10. A seamless user recognition and adaptive training system based on an AI display screen, characterized in that, The system includes: The data acquisition module is used to non-intrusively acquire frontal facial images and human posture videos of users using the visual sensors of the AI display screen of the motion device; The identity determination module is used to extract component-related descriptors of the facial contour region in the frontal face image, generate facial texture features corresponding to the user, and combine the facial texture features with a preset user pre-stored feature library to determine the user's identity. The standardization analysis module is used to extract skeletal key point data from the human posture video to determine the user's joint position and limb angle, retrieve the training template related to the user's identity, and analyze the user's movement standardization by combining the joint position, the limb angle and the training template. The fatigue analysis module is used to collect the user's physiological feedback signals and behavioral performance data in real time using the auxiliary sensors in the AI display screen, and to analyze the user's physiological fatigue level by combining the physiological feedback signals and behavioral performance data. The training adjustment module is used to adaptively adjust the motion parameters of the exercise equipment by combining the physiological fatigue level and the motion standardization, and obtain the adjustment result.