Wearable device-based motion gesture monitoring method, system, and electronic device

CN122701318APending Publication Date: 2026-09-08SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610869122.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

[0005]本申请实施例提供了一种基于可穿戴设备的运动姿态监测方法、系统和电子设备,可以解决相关技术中的可穿戴设备缺乏教学场景自适应功能,教学辅助效果差的技术问题

Benefits of technology

[0013]The beneficial effects of this application embodiment compared with the prior art are as follows: An image processing unit and an inertial measurement unit are configured on the wearable device. The wearable device is configured with two monitoring modes, such as saccade mode and gaze mode. In actual use, the monitoring intention of the wearer is identified, and then the target monitoring mode of the wearable device is determined based on the monitoring intention, and automatically switched to the target monitoring mode. For example, when the target monitoring mode is saccade mode, a coarse analysis of motion posture is performed based on the scene image sequence and IMU data, and a motion posture heatmap of all objects in the scene image sequence is output. When the target monitoring mode is gaze mode, a coarse analysis of motion posture is performed based on the scene image... The sequence and the IMU data are used to perform fine analysis of the motion posture of the target object being observed in the scene image sequence, and output the motion posture of the target object. That is, in the scenario where the coach wears a wearable device to assist teaching, when the coach scans and observes the students, the wearable device is controlled to switch to scanning mode. By coarsely analyzing the motion posture of all students, the overall teaching results are quickly output. When the coach focuses on observing a certain student, the wearable device is controlled to switch to gaze mode to finely analyze the student's motion posture, which makes it easier for the coach to correct the student. This adds an adaptive function to the teaching scenario, meets different monitoring needs in different teaching scenarios, and improves the teaching assistance effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122701318A_ABST
    Figure CN122701318A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a motion posture monitoring method and system based on a wearable device and an electronic device, the wearable device comprising an image acquisition unit and an inertial measurement unit, the method comprising: acquiring a scene image sequence captured by the image acquisition unit and IMU data captured by the inertial measurement unit, and correcting the scene image sequence based on the IMU data; determining a monitoring intention of a wearer of the wearable device, and determining a target monitoring mode of the wearable device according to the monitoring intention; when the target monitoring mode is a saccade mode, performing rough motion posture analysis processing based on the corrected scene image sequence, and outputting a motion posture quality distribution heat map of all objects in the scene image sequence; and when the target monitoring mode is a fixation mode, performing fine motion posture analysis processing based on a target object being fixated in the corrected scene image sequence, and outputting a motion posture of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and in particular relates to a motion posture monitoring method, system and electronic device based on wearable devices. Background Technology

[0002] Currently, wearable devices are portable and intelligent, and are widely used in various fields. For example, in the field of sports teaching, wearable devices can be used as teaching aids to monitor students' movement postures, facilitating the analysis and correction of these postures.

[0003] Existing technologies, such as wearable devices used as teaching aids, collect image data of teaching scenarios and perform posture analysis on this data to monitor students' movements. However, the monitoring methods of existing wearable devices are fixed, while actual teaching scenarios are flexible and varied, and existing wearable devices cannot meet the different monitoring needs of different teaching scenarios.

[0004] Therefore, existing wearable devices lack adaptive functionality for teaching scenarios, resulting in poor teaching assistance. Summary of the Invention

[0005] This application provides a motion posture monitoring method, system, and electronic device based on wearable devices, which can solve the technical problem in related technologies that wearable devices lack adaptive functions for teaching scenarios and have poor teaching assistance effects.

[0006] In a first aspect, embodiments of this application provide a motion posture monitoring method based on a wearable device, wherein the wearable device includes an image acquisition unit and an inertial measurement unit, and the motion posture monitoring method includes: The scene image sequence acquired by the image acquisition unit and the IMU data acquired by the inertial measurement unit are obtained, and the scene image sequence is corrected based on the IMU data; The monitoring intention of the wearer of the wearable device is determined, and the target monitoring mode of the wearable device is determined based on the monitoring intention. The monitoring mode of the wearable device includes saccade mode and gaze mode. When the target monitoring mode is the scanning mode, a rough motion posture analysis is performed based on the corrected scene image sequence, and a heat map of the motion posture quality distribution of all objects in the scene image sequence is output. When the target monitoring mode is the gaze mode, the motion posture of the gazed target object in the corrected scene image sequence is finely analyzed and processed, and the motion posture of the target object is output.

[0007] In one possible implementation of the first aspect, the wearable device further includes an eye-tracking unit for acquiring eye image sequences, wherein determining the monitoring intent of the wearer of the wearable device includes: The eye image sequence acquired by the eye-tracking unit is corrected based on the IMU data; Eye-tracking data is extracted based on the corrected eye image sequence; Based on the eye-tracking data, the dwell time and speed of the wearer's gaze are determined. Based on the dwell time and the speed of eye movement, the wearer's monitoring intention is identified, and the target monitoring mode of the wearable device is determined based on the monitoring intention.

[0008] In one possible implementation of the first aspect, determining the target monitoring mode of the wearable device based on the dwell time and the gaze movement speed includes: When the dwell time is greater than or equal to a preset time threshold and the gaze movement speed is less than or equal to a preset speed threshold, the target monitoring mode is determined to be the gaze mode. If the dwell time is less than the preset time threshold, or the gaze movement speed is greater than the preset speed threshold, the target monitoring mode is determined to be the scanning mode.

[0009] In one possible implementation of the first aspect, the step of performing a coarse motion pose analysis based on the corrected scene image sequence to output a heatmap of the motion pose quality distribution of all objects in the scene image sequence includes: Identify all objects in the corrected scene image sequence and divide all objects into at least two object groups; Based on preset standard motion postures, calculate the quality score of the motion posture of each group of objects; The target color corresponding to the quality score is rendered at the location of the object group to generate a heatmap of motion posture quality distribution.

[0010] In one possible implementation of the first aspect, calculating the quality score of the motion posture of each group of objects based on a preset standard motion posture includes: Extract the skeletal features of objects in the object group, and identify the motion posture of the objects based on the skeletal features; Based on the standard motion posture, determine the standard score of the object's motion posture; The quality score for each group of objects is determined based on the average of the standard scores of all objects in the object group.

[0011] In one possible implementation of the first aspect, the wearable device further includes an eye-tracking unit for acquiring a sequence of eye images. The step of performing fine-grained motion posture analysis on the gazed target object in the corrected scene image sequence and outputting the motion posture of the target object includes: Based on the corrected eye image sequence, the point where the wearer's gaze falls in the scene image plane at the same moment is determined; The target object being observed is determined based on the point of gaze. Based on the corrected scene image sequence, extract the skeletal and joint features of the target object, and analyze the motion posture of the target object based on the skeletal and joint features; Identify the target part of the target object whose deviation from the preset standard motion posture is greater than or equal to the preset deviation; Output the motion posture of the target object and display correction prompts at the target location.

[0012] In one possible implementation of the first aspect, the step of extracting skeletal and joint features of the target object based on the modified scene image sequence, and analyzing the motion posture of the target object based on the skeletal and joint features, includes: Based on the corrected scene image sequence, skeletal features and joint features of the target object are extracted, wherein the joint features include joint position and joint confidence. When the joint confidence corresponding to the joint position in multiple consecutive frames of the first scene image is less than a preset confidence, the motion trend of the target object is predicted based on the second scene image in the scene image sequence other than the first scene image. The joint position is adjusted according to the movement trend; Based on the corrected joint positions and skeletal features, the motion posture of the target object is determined.

[0013] The beneficial effects of this application embodiment compared with the prior art are as follows: An image processing unit and an inertial measurement unit are configured on the wearable device. The wearable device is configured with two monitoring modes, such as saccade mode and gaze mode. In actual use, the monitoring intention of the wearer is identified, and then the target monitoring mode of the wearable device is determined based on the monitoring intention, and automatically switched to the target monitoring mode. For example, when the target monitoring mode is saccade mode, a coarse analysis of motion posture is performed based on the scene image sequence and IMU data, and a motion posture heatmap of all objects in the scene image sequence is output. When the target monitoring mode is gaze mode, a coarse analysis of motion posture is performed based on the scene image... The sequence and the IMU data are used to perform fine analysis of the motion posture of the target object being observed in the scene image sequence, and output the motion posture of the target object. That is, in the scenario where the coach wears a wearable device to assist teaching, when the coach scans and observes the students, the wearable device is controlled to switch to scanning mode. By coarsely analyzing the motion posture of all students, the overall teaching results are quickly output. When the coach focuses on observing a certain student, the wearable device is controlled to switch to gaze mode to finely analyze the student's motion posture, which makes it easier for the coach to correct the student. This adds an adaptive function to the teaching scenario, meets different monitoring needs in different teaching scenarios, and improves the teaching assistance effect.

[0014] Secondly, embodiments of this application provide a motion posture monitoring system based on a wearable device, comprising: an image acquisition unit and an inertial measurement unit, and further comprising: The acquisition module is used to acquire the scene image sequence acquired by the image acquisition unit and the IMU data acquired by the inertial measurement unit, and to correct the scene image sequence based on the IMU data; A determination module is used to determine the monitoring intention of the wearer of the wearable device, and determine the target monitoring mode of the wearable device based on the monitoring intention. The monitoring modes of the wearable device include saccade mode and gaze mode. The first processing module is used to perform a rough motion posture analysis based on the corrected scene image sequence when the target monitoring mode is the scanning mode, and output a heat map of the motion posture quality distribution of all objects in the scene image sequence. The second processing module is used to perform fine motion posture analysis on the target object being watched in the corrected scene image sequence when the target monitoring mode is the gaze mode, and output the motion posture of the target object.

[0015] Thirdly, embodiments of this application provide an electronic device, which further includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any one of the first aspects above.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any one of the first aspects above.

[0017] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the method described in any one of the first aspects above.

[0018] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a system schematic diagram of a wearable device provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a motion posture monitoring method based on a wearable device according to an embodiment of this application; Figure 3 This is a flowchart illustrating a motion posture monitoring method based on a wearable device, provided in another embodiment of this application. Figure 4 This is a schematic diagram of the structure of a motion posture monitoring system based on a wearable device provided in an embodiment of this application. Detailed Implementation

[0021] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0022] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0023] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0024] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0025] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0026] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0027] Definitions: IMU stands for Inertial Measurement Unit (or Inertial Sensor), which is a sensor module that integrates an accelerometer and a gyroscope. IMU data includes at least one of angular velocity, acceleration, and heading angle.

[0028] DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise.

[0029] Wearable devices refer to head-mounted devices such as augmented reality (AR) terminals and smart glasses. Alternatively, wearable devices can also be electronic devices worn on clothing or worn on the hand. This application uses an AR terminal as an example of a wearable device for illustration.

[0030] With the development of smart technology, wearable devices are becoming increasingly intelligent. Due to their portability, wearable devices have been widely used in various fields. For example, in the field of sports teaching, wearable devices can be used as teaching aids to monitor students' movement postures, facilitating the analysis and correction of these postures.

[0031] Wearable devices, as teaching aids, typically employ relatively fixed methods for monitoring instruction. For example, in a teaching setting, wearable devices collect image data of the teaching environment and analyze student postures based on this data to monitor their movements. However, actual teaching scenarios are highly variable, and instructors' monitoring needs may differ across these scenarios. In some scenarios, instructors may want to monitor the overall movement of all students to understand their overall performance. In others, they may want to monitor a specific student to accurately correct their posture. Therefore, the information needed by instructors varies across different scenarios, leading to different requirements for wearable devices in terms of instructional support and, consequently, different information requirements.

[0032] Current wearable devices cannot meet the diverse monitoring needs in teaching scenarios. Therefore, these devices lack adaptive functionality for teaching scenarios, resulting in poor teaching support.

[0033] Based on this, embodiments of this application provide a motion posture monitoring method based on a wearable device. During image acquisition, the user's monitoring intention is identified. If the intention is to inspect all students, the target monitoring mode is determined to be a saccade mode, and the wearable device is controlled to process the scene image sequence and IMU data according to the saccade mode. If the intention is to focus on (or gaze at) a specific student, the target monitoring mode is determined to be a gaze mode, and the wearable device is controlled to process the scene image sequence and IMU data according to the gaze mode. Therefore, the motion posture monitoring method provided by embodiments of this application enables the provision of different monitoring modes and output of different monitoring information in different teaching scenarios, giving the wearable device an adaptive function for teaching scenarios and improving the teaching assistance effect.

[0034] The following detailed description of the implementation of the motion posture monitoring method based on wearable devices is provided through various embodiments.

[0035] Figure 1 A system schematic diagram of a wearable device is shown, such as... Figure 1 As shown, it includes: a wearable device 101 and a terminal device 102, with the wearable device 101 connected to the terminal device 102.

[0036] Wearable device 101 includes AR terminals, smart glasses, etc. Wearable device 101 includes an image acquisition unit and an inertial measurement unit (IMU). In some embodiments, wearable device 101 also includes an eye-tracking unit. The image acquisition unit is used to acquire scene image sequences, and multiple scene image sequences are stitched together in time to generate a scene video stream. The inertial measurement unit is used to acquire IMU data, including head angular velocity, acceleration, and yaw angle. The eye-tracking unit is used to acquire eye image sequences, and multiple eye image sequences are stitched together in time to generate an eye video stream.

[0037] In some embodiments, the wearable device 101 further includes a processing unit connected to the image acquisition unit, the inertial measurement unit, and the eye tracking unit. The processing unit preprocesses the scene image sequence, the eye image sequence, and the IMU data, and then transmits the preprocessed scene image sequence, eye image sequence, and IMU data to the terminal device 102 for motion posture analysis. Alternatively, in some embodiments, the processing unit can also perform motion posture analysis on the preprocessed scene image sequence, eye image sequence, and IMU data without requiring processing by the terminal device 102.

[0038] In some embodiments, the wearable device 101 further includes a display unit for displaying motion posture analysis results, such as displaying a motion posture mass distribution heatmap, or displaying the motion posture of a target object.

[0039] In some embodiments, the terminal device 102 is used to process or store data, and the terminal device 102 includes, but is not limited to, a computer or the cloud.

[0040] Figure 2 The illustration shows a schematic flowchart of a motion posture monitoring method based on a wearable device according to an embodiment of this application. This is illustrative and not limiting; the method can be applied to the aforementioned wearable device or to the aforementioned terminal device connected to the wearable device. The motion posture monitoring method includes: S201, acquire the scene image sequence acquired by the image acquisition unit and the IMU data acquired by the inertial measurement unit, and correct the scene image sequence based on the IMU data.

[0041] In some embodiments, the wearable device includes an image acquisition unit and an inertial measurement unit.

[0042] In other embodiments, the wearable device includes an image acquisition unit, an eye-tracking unit, and an inertial measurement unit.

[0043] Taking AR glasses as an example of a wearable device, the AR glasses include an inner side facing the eyes when worn and an outer side opposite to the inner side. An image acquisition unit, an eye-tracking unit, and an inertial measurement unit are all housed within the AR glasses. The image acquisition unit's viewing angle faces the outer side of the AR glasses, while the eye-tracking unit's viewing angle faces the inner side. A first transformation matrix between the inertial measurement unit and the image acquisition unit is pre-calibrated, as is a second transformation matrix between the inertial measurement unit and the eye-tracking unit.

[0044] The image acquisition unit is used to acquire scene image sequences, the inertial measurement unit is used to acquire IMU data, and the eye tracking unit is used to acquire eye image sequences.

[0045] The image acquisition unit includes a wide-angle RGB camera. Optionally, the wide-angle RGB camera acquires a sequence of scene images at a frequency of 1080P / 60fps.

[0046] The inertial measurement unit includes an integrated 9-axis IMU sensor (e.g., accelerometer + gyroscope + magnetometer) for high-frequency acquisition of IMU data from the head.

[0047] The eye-tracking unit includes an infrared camera, and optionally, the eye-tracking unit acquires eye image sequences at an acquisition frequency of 120 Hz or higher.

[0048] Since wearable devices are typically worn on the body, and people are constantly moving, the scene image sequences and eye image sequences captured by wearable devices suffer from perspective distortion. Therefore, after acquiring the scene image sequences and eye image sequences, IMU data is used to perform geometric correction on the scene image sequences and eye image sequences.

[0049] It is understandable that IMU data includes the head's angular velocity ω and acceleration a. After acquiring the angular velocity ω and acceleration a, the rotation matrix R of the image acquisition unit relative to world coordinates (gravity direction) is calculated by quaternion solving or Kalman filtering, combined with the first transformation matrix (the transformation matrix between the inertial measurement unit and the image acquisition unit); or, after acquiring the angular velocity ω and acceleration a, the rotation matrix R of the eye tracking unit relative to world coordinates (gravity direction) is calculated by quaternion solving or Kalman filtering, combined with the second transformation matrix (the transformation matrix between the inertial measurement unit and the eye tracking unit).

[0050] After calculating the rotation matrix R, inverse perspective transformation is performed on the scene image sequence and the eye image sequence. For example, based on the rotation matrix R, an inverse perspective mapping function is constructed to perform geometric transformation on the original image sequence. For instance, the inverse perspective mapping function is: P_corrected = K·R1 (-1)·K (-1)·P_raw; Where P_raw is the pixel coordinates of the original image sequence (such as a scene image sequence or an eye image sequence), K is the intrinsic parameter matrix of the image acquisition unit or eye tracking unit, R is the rotation matrix of the camera (such as the image acquisition unit or eye tracking unit) relative to world coordinates, and P_corrected is the pixel coordinates of the corrected image sequence.

[0051] By using IMU data to correct scene image sequences and eye image sequences, the object's body posture is presented as a standard orthogonal view in the scene image sequence, avoiding perspective distortion in the scene image sequence that affects the accuracy of motion posture analysis. Alternatively, the eye position is presented as a standard orthogonal view in the eye image sequence, improving the accuracy of identifying the pupil and corneal positions in the eye image sequence.

[0052] S202, determine the monitoring intent of the wearer of the wearable device, and determine the target monitoring mode of the wearable device based on the monitoring intent. The monitoring modes of the wearable device include saccade mode and gaze mode.

[0053] It can be understood that wearable devices are configured to monitor the motion and posture of objects within a given viewpoint based on saccade and gaze modes. In other words, the monitoring modes of wearable devices include saccade and gaze modes.

[0054] There is a correspondence between monitoring intent and monitoring mode. For example, monitoring intent refers to the wearer's monitoring needs for a scene. For instance, monitoring intent includes the wearer's intent to quickly scan the scene and the wearer's intent to gaze at a specific object in the scene. The intent to quickly scan the scene corresponds to the scanning mode, and the intent to gaze at a specific object in the scene corresponds to the gazing mode.

[0055] In some embodiments, an eye-tracking unit acquires a sequence of eye images, and then tracks the coach's gaze based on the eye image sequence to identify the coach's monitoring intention. For example, the wearer's gaze point can be predicted based on the eye image sequence, and then the wearer's monitoring intention can be predicted based on the gaze point.

[0056] For example, eye-tracking data can be extracted based on the corrected eye image sequence. Based on the eye-tracking data, the wearer's gaze point, the dwell time at the gaze point, and the gaze movement speed can be identified. Then, based on the dwell time and gaze movement speed, the wearer's monitoring intention can be determined, and the target monitoring mode of the wearable device can be determined based on the monitoring intention.

[0057] As an example, eye-tracking data includes the pupil center position and the position of the corneal reflective bright spot. Based on the pupil center position and the corneal reflective bright spot position, the wearer's gaze point is determined. The positional difference between adjacent gaze points is determined based on the time sequence. If the positional difference is less than a preset threshold, it is determined that the wearer's gaze remains at that gaze point. The duration of lingering at that gaze point is the wearer's gaze dwell time. Based on the positional difference (such as the distance between adjacent gaze points) and the time interval between adjacent gaze points, the gaze movement speed is calculated.

[0058] When the wearer's gaze lingers on a specific point for a duration greater than or equal to a preset time threshold, and the gaze movement speed is less than or equal to a preset speed threshold, the target monitoring mode is determined to be gaze mode. For example, in a sports teaching scenario, when a coach wants to monitor a student, they typically gaze at that student for a period of time to observe their posture. Therefore, the coach's gaze lingers on the student for a certain duration, and the gaze movement speed is slow. Thus, by monitoring whether the gaze lingers on a specific point for a duration greater than or equal to the preset time threshold, and whether the gaze movement speed is less than or equal to the preset speed threshold, the wearer's monitoring intention can be predicted, thereby determining the target monitoring mode of the wearable device.

[0059] When the wearer's gaze lingers on a point less than a preset time threshold, or the gaze movement speed exceeds a preset speed threshold, the target monitoring mode is determined to be a scanning mode. For example, in a sports teaching scenario, when a coach wants to monitor the overall movement posture of the students, they typically quickly scan the students to observe the uniformity of their posture or quickly locate students with obvious errors. Therefore, the coach's gaze will not linger on any specific student, and the gaze movement speed is fast. Thus, by monitoring whether the gaze lingers for less than a preset time threshold and whether the gaze movement speed exceeds a preset speed threshold, the wearer's monitoring intention can be predicted, thereby determining the target monitoring mode of the wearable device.

[0060] In this embodiment, compared to the target monitoring mode based on the dwell time of the gaze point, the accuracy of identifying the monitoring intent is higher by combining the dwell time of the gaze point and the speed of gaze movement.

[0061] In some embodiments, the monitoring intent can also be determined by recognizing gestures. For example, based on the scene image sequence acquired by the image acquisition unit, when the wearer's gesture is recognized as matching a preset gaze gesture, it is determined that the wearer's monitoring intent is to gaze at a specific object in the scene, that is, the corresponding target monitoring mode is gaze mode; when the wearer's gesture is recognized as matching a preset scanning gesture, it is determined that the wearer's monitoring intent is to quickly scan the scene, and the corresponding target monitoring mode is scanning mode.

[0062] In some embodiments, the monitoring intent can also be identified via voice commands. For example, the wearer's voice information can be identified, and the monitoring intent can be determined based on the voice information.

[0063] Alternatively, in some embodiments, the wearable device is provided with a monitoring mode switching button or switching control, which the wearer can touch to switch the monitoring mode.

[0064] Alternatively, in some embodiments, the monitoring intent is determined by calculating the wearer's head movement amplitude. For example, based on IMU data collected by the inertial measurement unit, the wearer's head movement amplitude within a preset time is determined. If the head movement amplitude is less than or equal to a preset amplitude threshold, the monitoring intent is determined to be a gaze at a specific object in the scene, and the corresponding monitoring mode is gaze mode; if the head movement amplitude is greater than the preset amplitude threshold, the monitoring intent is determined to be a rapid saccade scene, and the corresponding target monitoring mode is saccade mode.

[0065] S203, when the target monitoring mode is scan mode, performs coarse motion posture analysis based on the corrected scene image sequence and outputs a heat map of the motion posture quality distribution of all objects in the scene image sequence.

[0066] It is understandable that coarse analysis processing refers to the quality assessment of the overall motion posture of all objects in a scene image sequence, or the quality assessment of motion posture by group / region.

[0067] This embodiment uses a coarse grouping analysis of motion posture quality as an example: For instance, when the target monitoring mode is determined to be a scan mode, all objects (such as people) in the corrected scene image sequence are identified. All objects are divided into at least two groups according to preset rules, with objects in each group being adjacent, such as a group of 5 or 10 adjacent people. Then, based on preset standard motion postures, the motion posture quality score of each group is calculated. According to the quality score of the group, the corresponding color is rendered at the location of the group, generating a motion posture quality distribution heatmap. This heatmap is then displayed on the wearable device's display unit, allowing the wearer to understand the motion posture quality of each group and quickly locate the groups with low quality scores.

[0068] It is understandable that all objects in a scene image sequence can be identified through skeleton recognition and tracking. Skeleton recognition and tracking is the same as the related skeleton recognition technology, which will not be described in detail here.

[0069] In some embodiments, the DBSCAN clustering algorithm can be used to automatically divide adjacent objects into several object groups G={G_1, G_2, ...}, evaluate the standard score of the motion posture of the objects in each object group relative to the standard motion posture, and determine the quality score of the object group by the mean of the standard scores of all objects in the object group.

[0070] For example, the formula for calculating the quality score of an object group is: Q_group=(1 / n)ΣScore(S_i); Where Q_group is the quality score of the object group, Score(S_i) is the standard score of the motion posture of a single object, and n is the number of objects in the object group.

[0071] In some embodiments, the motion posture of an object is identified by its skeletal features, and then scored based on the deviation of each object's motion posture from a preset standard motion posture. For example, skeletal features of objects in a group are extracted from a corrected scene image sequence. The motion posture of the object is identified based on these skeletal features, and then a standard score for each object's motion posture is determined by comparing its similarity to a standard motion posture. For example, taking a score of 100 as a perfect match, a standard motion posture corresponding to the sport is determined. If the similarity between an object's motion posture and this standard motion posture is 80%, the corresponding standard score is 80. The similarity between the motion posture and the standard motion posture is determined by comprehensively considering the similarity of the object's key joints, joint angles, and timing. In some embodiments, similarity can also be determined by comparing the deviation values ​​between the motion posture and the standard motion posture.

[0072] It's understandable that a pre-defined correspondence between quality scores and target colors is established, with different quality scores corresponding to different target colors. For example, the quality score is inversely proportional to the depth of the target color; a lower quality score results in a deeper target color, and a higher quality score results in a lighter target color. Alternatively, at least two target colors can be set, such as a first color and a second color. When the quality score is less than or equal to a preset threshold, the corresponding target color is the first color; that is, in the location of object groups with a quality score less than or equal to the preset threshold, the first color swatch is rendered. When the quality score is greater than the preset threshold, the corresponding target color is the second color; that is, in the location of object groups with a quality score greater than the preset threshold, the second color swatch is selected. For example, the first color is dark red and the second color is green. Dark red is rendered in areas with low quality scores to indicate to the wearer that the overall motion and posture quality of the object group corresponding to the "dark red" area is poor, guiding the wearer to pay attention to the corresponding group and meeting the wearer's need for overall monitoring. Green is rendered in areas with high quality scores to indicate to the wearer that the overall motion and posture quality of the object group corresponding to the "green" area is good, allowing them to pay more attention to object groups in other areas.

[0073] When the target monitoring mode is in scan mode, the coach focuses more on the overall state. This embodiment performs a rough analysis of the motion posture of the corrected scene image sequence, quickly generating a heatmap of the motion posture quality distribution of all objects in the scene image sequence. This allows the coach to understand the overall motion posture quality of all objects, and at the same time, quickly locates the object group with poor motion posture quality, prompting the coach to pay special attention to the object group with poor motion posture quality. This meets the coach's teaching scenario needs and assists in better teaching effectiveness.

[0074] In some embodiments, after outputting the motion posture quality distribution heatmap of all objects in the scene image sequence, the method further includes: obtaining a selection gesture, and when it is determined that the selected area corresponding to the selection gesture is an area or object group with a low quality score in the motion posture quality distribution heatmap, switching the target monitoring mode to gaze mode.

[0075] In other words, wearable devices have gesture recognition capabilities. When the wearable device is in scanning mode, a heatmap of the motion posture quality distribution of all objects is displayed on the display unit. At this time, the user can use gestures to select and focus on areas or groups of objects with low quality scores in the motion posture quality distribution heatmap. When the selection gesture is recognized, the device automatically switches to gaze mode (e.g., jumps to step S204). There is no need to switch monitoring modes by looking at the target object; the device directly switches from scanning mode to gaze mode, making the operation convenient and quick.

[0076] Alternatively, you can switch the low-quality area to gaze mode via voice, toggle button, or toggle control, perform targeted error correction in the low-quality area, and generate classroom notes or after-class training suggestions.

[0077] S204, when the target monitoring mode is gaze mode, performs fine analysis of the motion posture of the gazed target object in the corrected scene image sequence and outputs the motion posture of the target object.

[0078] When the target monitoring mode is gaze mode, the wearable device switches to focus on a single object mode, that is, to perform fine analysis and processing of the motion posture of a single object.

[0079] It can be understood that fine-grained analysis and processing refers to analyzing the contours, motion postures, etc. of individual objects in a scene image sequence to generate the specific motion posture of a single object. For example, fine-grained analysis and processing includes extracting skeletal features and joint features of the target object, and then refining these features. For instance, the motion posture of the target object is generated based on skeletal and joint features; skeletal connections are generated based on skeletal features; joint angles are generated based on joint features; and finally, information such as skeletal features, skeletal connections, joint features, and joint angles is used to generate and output the motion posture of the target object.

[0080] In some embodiments, when the target detection mode is gaze mode, the target object to be gazed upon is predetermined. For example, based on a corrected sequence of eye images, the gaze point in the scene image plane corresponding to the wearer's gaze at the same moment is determined, and then the target object to be gazed upon is determined based on the gaze point.

[0081] It is understandable that, since both the eye image sequence and the scene image sequence are based on temporal order, they can be aligned using timestamps. If, based on the eye image sequence, the wearer's gaze lingers at a point within a first time period for a duration greater than or equal to a preset time threshold, and the gaze movement speed is less than or equal to a preset speed threshold, then the scene image corresponding to the first time period is determined as the scene image corresponding to the wearer's gaze at that same moment. The gaze point is then mapped onto the scene image to obtain the gaze point in the scene image plane.

[0082] As an example, the distance between the gaze point and the center of the skeleton of each object in the scene image can be calculated to determine the object being watched. For instance, the object with the smallest distance is selected as the object being watched corresponding to the gaze point. Alternatively, the object with the smallest distance, and whose distance is less than a preset distance, is selected as the object being watched corresponding to the gaze point, improving recognition accuracy. Taking the distance between the gaze point and the center of the object's skeleton as D_i, if D_k = min(D_i) and D_k < preset distance, then object k is determined to be the target object.

[0083] After identifying the target object, other objects or their detailed features in the scene image sequence are filtered out, while retaining their contour features. Then, the skeletal and joint features of the target object are extracted, and the target object's motion posture is analyzed based on these features.

[0084] For example, joint features include joint position and joint confidence, and joint angles are calculated from joint position. Skeletal features include bone length and bone vector, and the motion posture of the target object is analyzed in detail based on bone vector, bone length and joint angle.

[0085] In some embodiments, after initially analyzing and obtaining the motion posture of the target object through the above method, a target part of the target object's motion posture that deviates from the preset standard motion posture by more than or equal to a preset deviation is identified. The preset deviation is pre-set, for example, 15%. If there is a part of the target object's motion posture that deviates from the standard motion posture by more than or equal to the preset deviation, then a correction prompt message for that target part is displayed on the motion posture. For example, the target part is displayed with a red arrow, or the target part is displayed with the corrected posture or outline, etc., intuitively prompting the wearer where the target object's motion posture is wrong, or guiding the wearer to correct the target object's motion posture.

[0086] In this embodiment, an image processing unit, an inertial measurement unit, and an eye-tracking unit are configured on the wearable device. The wearable device is equipped with two monitoring modes, such as saccade mode and gaze mode. In actual use, after acquiring scene image sequences and eye image sequences, the scene image sequences and eye image sequences are corrected based on IMU data to reduce perspective distortion and ensure the accuracy of feature information in the scene image sequences and eye image sequences. Based on the eye image sequences acquired by the eye-tracking unit, the monitoring intention of the wearer can be identified, and then the target monitoring mode of the wearable device is determined according to the monitoring intention, and the device is automatically switched to the target monitoring mode. For example, when the target monitoring mode is saccade mode, a rough analysis of motion posture is performed based on the scene image sequence. The system processes and outputs a heatmap of the motion postures of all objects in the scene image sequence. When the target monitoring mode is gaze mode, it performs fine analysis of the motion postures of the gazed target objects in the scene image sequence based on the scene image sequence, and outputs the motion postures of the target objects. That is, in the scenario where the coach wears a wearable device to assist teaching, when the coach scans and observes the students, the wearable device is controlled to switch to scan mode. By coarsely analyzing the motion postures of all students, the overall teaching results are quickly output. When the coach gazes and observes a particular student, the wearable device is controlled to switch to gaze mode, and the motion posture of that student is finely analyzed, which makes it easier for the coach to correct the student. This adds an adaptive function to the teaching scenario, meets different monitoring needs in different teaching scenarios, and improves the teaching assistance effect.

[0087] Figure 3 A schematic flowchart of a motion posture monitoring method based on a wearable device according to another embodiment of this application is shown. Based on the aforementioned implementation, this embodiment further corrects the occluded motion posture of the target object in gaze mode, ensuring the complete motion posture of the target object is monitored. Figure 3 As shown, it also includes: S301, extract the skeletal and joint features of the target object based on the corrected scene image sequence. The joint features include joint position and joint confidence.

[0088] It is understandable that the extraction of skeletal and joint features can be based on conventional neural network models, which will not be described in detail here.

[0089] S302, when the joint confidence corresponding to the joint position in multiple consecutive frames of the first scene image is less than the preset confidence, predict the motion trend of the target object based on the second scene image other than the first scene image in the scene image sequence.

[0090] For example, a scene image sequence includes multiple frames of scene images arranged in a time sequence, and the pose of the target object in these multiple frames forms a motion pose. Since the scene images are captured during movement, or during the capture process, the wearer's head may be rotating, and the target object may be partially occluded from the image acquisition unit's perspective (e.g., occluded by other objects or objects). Taking the scene image captured when the target object is partially occluded as the first scene image as an example, the confidence of joints at occluded positions in the joint features extracted from the first scene image is often relatively low, and these joint positions are usually discarded, resulting in low accuracy in analyzing the motion pose of the target object. Therefore, in this embodiment, after extracting the joint features of each scene image, it is determined whether the confidence of the joint corresponding to the joint position in the scene image is less than a preset confidence level. If there is a joint position with a confidence level less than the preset confidence level, it is determined whether the confidence level of the joint corresponding to that joint position in N consecutive frames of the first scene image is less than the preset confidence level. If so, it is determined that the joint position is occluded, where N is greater than 1.

[0091] When a joint position is occluded, the motion trend of the target object is predicted based on the sequence of unoccluded scene images. For example, the motion trend of the target object is predicted based on a second scene image other than the first scene image.

[0092] As an example, a time-series prediction module (Kalman filter) is used to predict the motion trend of a target object.

[0093] S303 adjusts joint position based on movement trends.

[0094] Based on the motion trend of the target object, the trajectory of joint changes is obtained, and then the actual joint position of the occluded joint is determined. The actual joint position is the corrected joint position.

[0095] S304 determines the motion posture of the target object based on the corrected joint positions and skeletal features.

[0096] In this embodiment, when the target object is occluded, the position of the occluded joint can be corrected by predicting the movement trend of the target object, thereby improving the accuracy of the target object's movement posture recognition.

[0097] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0098] Corresponding to the wearable device-based motion posture monitoring method in the above embodiments, Figure 4 The diagram shows a structural block diagram of a motion posture monitoring device based on a wearable device according to an embodiment of this application. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0099] Reference Figure 4 The motion posture monitoring device 40 based on wearable devices includes an image acquisition unit 401 and an inertial measurement unit 402. The motion posture monitoring device 40 also includes: The acquisition module 403 is used to acquire the scene image sequence acquired by the image acquisition unit and the IMU data acquired by the inertial measurement unit, and to correct the scene image sequence based on the IMU data; The determination module 404 is used to determine the monitoring intention of the wearer of the wearable device, and determine the target monitoring mode of the wearable device based on the monitoring intention. The monitoring modes of the wearable device include saccade mode and gaze mode. The first processing module 405 is used to perform a rough analysis of motion posture based on the corrected scene image sequence when the target monitoring mode is a scanning mode, and output a heat map of the motion posture quality distribution of all objects in the scene image sequence. The second processing module 406 is used to perform fine analysis of the motion posture of the gazed target object in the corrected scene image sequence when the target monitoring mode is gaze mode, and output the motion posture of the target object.

[0100] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0101] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0102] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps described in the various method embodiments above.

[0103] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.

[0104] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographic device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0105] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0106] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0107] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for motion posture monitoring based on wearable devices, characterized in that, The wearable device includes an image acquisition unit and an inertial measurement unit, and the motion posture monitoring method includes: The scene image sequence acquired by the image acquisition unit and the IMU data acquired by the inertial measurement unit are obtained, and the scene image sequence is corrected based on the IMU data; The monitoring intention of the wearer of the wearable device is determined, and the target monitoring mode of the wearable device is determined based on the monitoring intention. The monitoring mode of the wearable device includes saccade mode and gaze mode. When the target monitoring mode is the scanning mode, a rough motion posture analysis is performed based on the corrected scene image sequence, and a heat map of the motion posture quality distribution of all objects in the scene image sequence is output. When the target monitoring mode is the gaze mode, the motion posture of the gazed target object in the corrected scene image sequence is finely analyzed and processed, and the motion posture of the target object is output.

2. The method according to claim 1, characterized in that, The wearable device further includes an eye-tracking unit for acquiring eye image sequences. Determining the monitoring intent of the wearer of the wearable device includes: The eye image sequence acquired by the eye-tracking unit is corrected based on the IMU data; Eye-tracking data is extracted based on the corrected eye image sequence; Based on the eye-tracking data, the dwell time and speed of the wearer's gaze are determined. Based on the dwell time and the speed of eye movement, the wearer's monitoring intention is identified, and the target monitoring mode of the wearable device is determined based on the monitoring intention.

3. The method according to claim 2, characterized in that, Determining the target monitoring mode of the wearable device based on the dwell time and the speed of eye movement includes: When the dwell time is greater than or equal to a preset time threshold and the gaze movement speed is less than or equal to a preset speed threshold, the target monitoring mode is determined to be the gaze mode. If the dwell time is less than the preset time threshold, or the gaze movement speed is greater than the preset speed threshold, the target monitoring mode is determined to be the scanning mode.

4. The method according to any one of claims 1 to 3, characterized in that, The process involves performing a coarse motion posture analysis based on the corrected scene image sequence, outputting a heatmap of the motion posture quality distribution of all objects in the scene image sequence, including: Identify all objects in the corrected scene image sequence and divide all objects into at least two object groups; Based on preset standard motion postures, calculate the quality score of the motion posture of each group of objects; The target color corresponding to the quality score is rendered at the location of the object group to generate a heatmap of motion posture quality distribution.

5. The method according to claim 4, characterized in that, The process of calculating the quality score of the motion posture of each group of objects based on a preset standard motion posture includes: Extract the skeletal features of objects in the object group, and identify the motion posture of the objects based on the skeletal features; Based on the standard motion posture, determine the standard score of the object's motion posture; The quality score for each group of objects is determined based on the average of the standard scores of all objects in the object group.

6. The method according to any one of claims 1 to 3, characterized in that, The wearable device further includes an eye-tracking unit for acquiring eye image sequences. The process of performing fine-grained motion posture analysis on the target object being gazed upon in the corrected scene image sequence and outputting the motion posture of the target object includes: Based on the corrected eye image sequence, determine the point of gaze landing on the scene image plane corresponding to the wearer's gaze at the same moment; The target object being observed is determined based on the point of gaze; Based on the corrected scene image sequence, extract the skeletal and joint features of the target object, and analyze the motion posture of the target object based on the skeletal and joint features; Identify the target part of the target object whose deviation from the preset standard motion posture is greater than or equal to the preset deviation; Output the motion posture of the target object and display correction prompts at the target location.

7. The method according to claim 6, characterized in that, The step of extracting skeletal and joint features of the target object based on the corrected scene image sequence, and analyzing the motion posture of the target object based on the skeletal and joint features, includes: Based on the corrected scene image sequence, skeletal and joint features of the target object are extracted, wherein the joint features include joint position and joint confidence. When the joint confidence corresponding to the joint position in multiple consecutive frames of the first scene image is less than a preset confidence, the motion trend of the target object is predicted based on the second scene image in the scene image sequence other than the first scene image. The joint position is adjusted according to the movement trend; Based on the corrected joint positions and skeletal features, the motion posture of the target object is determined.

8. A motion posture monitoring system based on wearable devices, characterized in that, Including an image acquisition unit and an inertial measurement unit, it also includes: The acquisition module is used to acquire the scene image sequence acquired by the image acquisition unit and the IMU data acquired by the inertial measurement unit, and to correct the scene image sequence based on the IMU data; A determination module is used to determine the monitoring intention of the wearer of the wearable device, and determine the target monitoring mode of the wearable device based on the monitoring intention. The monitoring modes of the wearable device include saccade mode and gaze mode. The first processing module is used to perform a rough motion posture analysis based on the corrected scene image sequence when the target monitoring mode is the scanning mode, and output a heat map of the motion posture quality distribution of all objects in the scene image sequence. The second processing module is used to perform fine motion posture analysis on the target object being watched in the corrected scene image sequence when the target monitoring mode is the gaze mode, and output the motion posture of the target object.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.