Foot health state detection method and device based on video, medium and product
By combining rear and side view videos to generate multimodal human posture key point change sequences, gait cycle, kinematic and dynamic analysis is performed, and a deep learning model is used to output foot health status, which solves the problem of low detection accuracy in existing technologies and achieves higher accuracy foot health detection.
Patent Information
- Application Number
- CN202511280753.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2026-02-17
AI Technical Summary
Existing technologies for foot health detection have low precision and lack detailed analysis of the gait process, resulting in inaccurate test results.
By acquiring rear-view and lateral-view videos, a multimodal sequence of key point changes in human posture is generated. Combined with gait cycle, kinematic and dynamic analysis, a deep learning model is used to output the foot health status.
It improves the accuracy of foot health detection, reflecting changes in the feet more in detail during walking, and enhances the accuracy and convenience of detection.
Smart Images

Figure CN121545209A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of foot health detection technology, and in particular to a video-based method, device, medium, and product for detecting foot health status. Background Technology
[0002] Foot health is an important part of overall health. Therefore, if you experience foot discomfort or other illnesses that may be caused by unhealthy feet, you should have your feet checked promptly and receive timely treatment.
[0003] Currently, foot health is detected using the following methods: capturing gait videos of the target individual using cameras or other acquisition devices; determining parameters such as stride length and stride speed based on the walking videos; and then determining the foot health status based on these parameters. However, existing technologies still suffer from low detection accuracy. Summary of the Invention
[0004] The purpose of this application is to provide a video-based method, device, medium, and product for detecting foot health status, in order to solve the problem of low accuracy in foot health detection described in the background art.
[0005] To achieve the above objectives, this application provides the following solution:
[0006] Firstly, this application provides a video-based method for detecting foot health status, including:
[0007] Acquire rear-view and side-view videos of the target person walking;
[0008] Based on the rear-view video and the side-view video, a multimodal human posture key point change sequence of the target person is generated when walking;
[0009] Gait analysis was performed on the multimodal human posture key point change sequence to obtain a time series of gait analysis results;
[0010] The gait analysis results time series are sent to a deep learning model so that the deep learning model can output the foot health status of the target person.
[0011] Optionally, the step of "generating a multimodal human posture key point change sequence of the target person while walking based on the rear view video and the side view video" is achieved by the following method:
[0012] The lateral view video is sent to a 3D-based human pose estimation model to output a 3D human pose key point change sequence.
[0013] The rear-view video is sent to a 2D-based human pose estimation model to output a 2D sequence of human pose key point changes.
[0014] The multimodal human posture key point change sequence is obtained by fusing the 3D human posture key point change sequence and the 2D human posture key point change sequence.
[0015] Optionally, the gait analysis includes gait periodicity analysis, gait kinematics analysis, and gait dynamics analysis. The step of "performing gait analysis on the multimodal human posture key point change sequence to obtain a time series of gait analysis results" is achieved through the following method:
[0016] Gait cycle analysis was performed on the multimodal human posture key point change sequence to obtain a time series of gait cycle analysis results;
[0017] Gait kinematics analysis was performed on the multimodal human posture key point change sequence to obtain a time series of gait kinematics analysis results;
[0018] Gait dynamics analysis was performed on the sequence of changes in key points of the multimodal human posture to obtain a time series of gait dynamics analysis results.
[0019] Optionally, the step of "sending the gait analysis result time series to the deep learning model" is achieved by sending the gait cycle analysis result time series, the gait kinematics analysis result time series, and the gait dynamics analysis result time series to the deep learning model, so that the deep learning model outputs the foot health status of the target person.
[0020] Optionally, after acquiring the rear-view and side-view videos of the target person walking, the method further includes:
[0021] The side-view and rear-view videos are filtered to obtain the filtered side-view and rear-view videos;
[0022] The selected side-view and rear-view videos are time-aligned to obtain new side-view and rear-view videos.
[0023] Optionally, the foot health condition includes a number of foot health conditions such as flat feet, high arches, hallux valgus, inversion foot, eversion foot, plantar fasciitis, and Achilles tendinitis.
[0024] Optionally, the rear-view video is captured by an RGB camera, and the side-view video is captured by a binocular depth camera.
[0025] In a second aspect, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described in any one of the first aspects above.
[0026] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any one of the first aspects above.
[0027] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any one of the first aspects above.
[0028] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0029] The video-based foot health status detection method provided in this application acquires rear-view video using an RGB camera and lateral-view video using a binocular depth camera; it then generates a multimodal human posture key point change sequence from these two videos; furthermore, it obtains a gait analysis result time series from the multimodal human posture key point change sequence; finally, it sends the gait analysis result time series to a deep learning model to determine the foot health status. Compared to existing solutions, this application not only focuses on conventional gait parameters such as stride length and stride speed, but also emphasizes the analysis of kinematic and dynamic parameters during the gait process; the method proposed in this application focuses more on the microscopic parameters of human posture key points, which can reflect the changes in the foot during walking in more detail and improve the accuracy of foot health detection. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is an application environment diagram of a video-based foot health status detection method in one embodiment of this application;
[0032] Figure 2 A schematic flowchart illustrating a video-based foot health status detection method provided in an embodiment of this application;
[0033] Figure 3 A flowchart illustrating a method for generating a multimodal human pose key point change sequence according to an embodiment of this application;
[0034] Figure 4 A schematic flowchart illustrating a method for filtering and aligning lateral view videos and rear view videos according to an embodiment of this application;
[0035] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0036] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0037] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the contents of this application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0038] The video-based foot health status detection method provided in this application can be applied to, for example... Figure 1 In the application environment shown, the application environment includes a terminal and a server. The terminal communicates with the server via a network. A data storage system stores the data that the server needs to process. The data storage system can be set up independently, integrated into the server, or located in the cloud or on another server. The terminal can send rear-view and lateral-view videos to be processed to the server. After receiving the videos, the server can store them first and retrieve them from the storage location when processing is needed, or it can perform the processing task while storing the videos. The server can provide feedback on the detection results obtained from the videos to the terminal. Furthermore, in some embodiments, the video-based foot health status detection method can also be implemented independently by the server or the terminal. For example, the terminal can directly process the rear-view and lateral-view videos, or the server can retrieve the videos to be processed from the data storage system and perform foot health status detection on the videos.
[0039] The terminals can be, but are not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. Servers can be implemented using independent servers, server clusters composed of multiple servers, or cloud servers.
[0040] In one exemplary embodiment, see Figure 2 As shown, a video-based method for detecting foot health status is provided. This method is executed by a computer device, specifically a terminal or server, or both. In this embodiment, the method is applied to... Figure 1 Taking the server in the example, the following steps 101 to 104 are used as an example:
[0041] Step 101: Acquire rear-view and side-view videos of the target person walking. The rear-view video is captured by an RGB camera, and the side-view video is captured by a binocular depth camera.
[0042] The target individuals are those who need to have their foot health checked.
[0043] The shooting process for the rear-view and side-view videos is as follows: the RGB camera is placed behind the person, at a distance and height that allows the whole body to be captured; the binocular depth camera is placed to the side of the person's walking direction, at a predetermined distance from the target person, at the same height as the RGB camera, so that the target person has 1 to 3 complete gait cycles within the binocular depth camera's shooting range during walking; then the target person begins to walk forward, and the two cameras begin to synchronously shoot videos.
[0044] Furthermore, when shooting, please pay attention to the following: the RGB camera and the stereo depth camera should remain in a fixed position during the recording process, and both should be at the same height and their relative positions should remain fixed.
[0045] The binocular depth camera captures 3D video containing depth information, while the RGB camera captures higher-resolution 2D video containing analytical details. The fusion of these two technologies provides more accurate and richer gait data. Therefore, the gait video data in this application is multimodal data.
[0046] Step 102: Generate a multimodal human posture key point change sequence of the target person when walking based on the rear view video and the side view video;
[0047] Among them, the sequence of key points of human posture change represents the movement trajectory of each key point of the human body during the walking process in the gait video, which is represented by a sequence of data points with time attributes.
[0048] The key points of human posture refer to the joints of the human skeleton (head, neck, shoulder, elbow, wrist, hip, knee, ankle, etc.) and other body parts related to gait data analysis (such as calcaneus, distal metatarsus, etc.).
[0049] The sequence of key point changes in human pose includes the spatial coordinates of each key point at different time points (X, Y, and Z coordinates in 3D; X and Y coordinates in 2D), as well as other possible attributes (such as the visibility and confidence of the key points).
[0050] As the target person walks for a period of time, the coordinates of these key points will change, thus forming a sequence that records the entire process of the human body transitioning from one posture to another.
[0051] Because everyone's foot health is different, the key points of movement during walking also differ from person to person. Furthermore, as walking time progresses, the sequence of key point changes in human posture varies depending on the individual's foot health. This sequence of key point changes can reflect the walking pattern of a target individual and can therefore be used to assess foot health.
[0052] Step 103: Perform gait analysis on the multimodal human posture key point change sequence to obtain a time series of gait analysis results;
[0053] Gait analysis time series is a quantitative description of the changes in gait characteristics over time during human walking. It is generated after in-depth analysis of the change sequence of key points in human posture and is used to reflect the walking state of the target person.
[0054] Multimodal human posture key point change sequences can reflect gait changes during human walking, and different gait changes reflect the health of a person's feet.
[0055] Step 104: Send the gait analysis result time series to the deep learning model so that the deep learning model can output the foot health status of the target person.
[0056] Among them, the deep learning model is pre-trained and used to determine the foot health status based on the time series of gait analysis results.
[0057] The video-based foot health status detection method provided in this application acquires rear-view video using an RGB camera and lateral-view video using a binocular depth camera; it then generates a multimodal human posture key point change sequence from these two videos; furthermore, it obtains a gait analysis result time series from the multimodal human posture key point change sequence; finally, it sends the gait analysis result time series to a deep learning model to determine the foot health status. Compared to existing solutions that focus more on macroscopic gait parameters, such as stride length and gait speed, the method proposed in this application focuses more on the microscopic parameter of human posture key points, which can reflect the changes in the foot during walking in more detail and improve the accuracy of foot health detection.
[0058] Furthermore, this application uses a binocular depth camera and an RGB camera as video acquisition devices, eliminating the need for traditional wearable devices and offering greater convenience. The device itself is simple in design, maintains accurate measurements even at long distances, and exhibits strong stability. Additionally, using both a binocular depth camera and an RGB camera to acquire video data allows for cross-verification, ensuring data accuracy and improving the overall accuracy of foot health detection.
[0059] In addition, with the rapid rise of computer vision technology, optical motion capture solutions have also emerged. Although they offer high accuracy, they are costly, limited to laboratory settings, and involve cumbersome processes. Pure vision-based solutions have emerged with two main approaches: monocular RGB cameras and simplified depth cameras. However, monocular RGB cameras lack depth information and cannot perform dynamic parameter calculations (such as joint torques). Simplified depth cameras (such as Kinect) have limitations in resolution and field of view, making it impossible to guarantee computational accuracy.
[0060] Meanwhile, existing technologies also suffer from a lack of gait analysis dimensions: existing solutions focus more on macroscopic gait parameters (step length / step speed), lacking joint analysis of lower limb dynamic chains (such as the influence of pelvic tilt on foot pressure) and posterior foot morphological features (such as hallux valgus), resulting in single discriminative features; and insufficient prediction of multiple diseases: traditional methods often target single diseases (such as judging flat feet based solely on the arch), lacking a generalized predictive model that associates gait dynamics-kinematic parameters with multiple foot diseases.
[0061] Existing solutions have significant shortcomings in terms of multi-dimensional data acquisition accuracy in unrestricted scenarios, spatiotemporal analysis capabilities for complex foot and ankle movements, and generalization of cross-disease prediction based on gait parameters. This application constructs a non-invasive assessment system that can be deployed on lightweight terminals by simultaneously acquiring multimodal data using a binocular depth camera and an RGB camera, fusing 3D / 2D posture, and performing full-stack analysis of gait dynamics and kinematics, thus achieving a leap from "precise laboratory measurement" to "scenario-based universal screening".
[0062] Optionally, see Figure 3 In another exemplary embodiment of this application, step 102 is implemented through the following steps 201-203:
[0063] Step 201: Send the lateral view video to a 3D-based human pose estimation model to output a 3D human pose key point change sequence.
[0064] The 3D human pose estimation model is a pre-trained model that can output a sequence of 3D human pose key point changes.
[0065] During training, a large number of lateral view videos are acquired as input samples, and the corresponding correct 3D human pose key point change sequences are obtained as labels for training. Other related training content is the same as existing technology and will not be repeated here.
[0066] Among them, the 3D-based human pose estimation model is existing technology, and relevant content can be found in the existing technology, which will not be introduced in this application.
[0067] Step 202: Send the rear-view video to a 2D-based human pose estimation model to output a 2D human pose key point change sequence.
[0068] The 2D human pose estimation model is a pre-trained model that can output a sequence of 2D human pose key point changes.
[0069] During training, a large number of rear-view videos are acquired as input samples, and the corresponding correct 2D human pose key point change sequences are obtained as labels for training. Other related training content is the same as existing technology and will not be repeated here.
[0070] Among them, the 2D-based human pose estimation model is existing technology, and relevant content can be found in the existing technology, which will not be described in this application.
[0071] Step 203: Merge the 3D human posture key point change sequence and the 2D human posture key point change sequence to obtain the multimodal human posture key point change sequence.
[0072] Multimodal human posture key point change sequences make the data more robust and improve the accuracy of foot health status detection compared to individual 3D or 2D human posture key point sequences.
[0073] The fusion of 3D and 2D human pose keypoint change sequences requires three steps: spatiotemporal alignment, multimodal feature fusion, and dynamic compensation. The core objective is to leverage the depth information of 3D data and the high-resolution details of 2D data to improve the robustness and accuracy of pose estimation. Specific fusion processes can be found in existing technologies and will not be detailed here.
[0074] Optionally, in another exemplary embodiment of this application, the gait analysis includes gait cycle analysis, gait kinematic analysis, and gait dynamics analysis, and step 103 includes steps 401 to 403:
[0075] Step 401: Perform gait period analysis on the multimodal human posture key point change sequence to obtain a time series of gait period analysis results;
[0076] Among them, gait periodicity analysis is the temporal evolution of macroscopic gait parameters.
[0077] Gait period analysis includes the calculation and analysis of stride length, stride frequency, stride speed, support phase duration, swing phase duration, and other gait period parameters.
[0078] A gait cycle refers to the process from the heel of one foot striking the ground to the heel striking the ground again on the same side during walking. A gait cycle can be divided into a stance phase and a swing phase. The stance phase refers to the time the lower limb contacts the ground and bears weight, while the swing phase refers to the time between the foot leaving the ground and taking a step forward, and then landing again. Stride length is the longitudinal straight-line distance between two points when the left and right heels or toes strike the ground successively. Step frequency is the number of steps taken per unit of time.
[0079] For example, if a 30-second walking video is recorded, the gait cycle analysis result time series will contain continuous values of the above parameters within each gait cycle, forming a time series such as [step length 1, step frequency 1, ..., step length 2, step frequency 2, ...].
[0080] Step 402: Perform gait kinematic analysis on the multimodal human posture key point change sequence to obtain a time series of gait kinematic analysis results;
[0081] Among them, gait kinematics analysis is the time-varying changes in joint angles and postures.
[0082] Among them, gait kinematic analysis includes the analysis of pelvic tilt angle, pelvic rotation angle, and pelvic deviation angle; the analysis of hip joint flexion angle, adduction / abduction angle, and anteroposterior rotation angle; the analysis of knee joint flexion angle, adduction / abduction angle, and anteroposterior rotation angle; and the calculation and analysis of parameters such as foot and ankle joint dorsiflexion angle, eversion / inversion angle, and foot forward angle.
[0083] For example, the time series of gait kinematics analysis results records the dynamic changes of joint angles in each frame of video, such as [hip flexion angle 1, knee flexion angle 1, ..., hip flexion angle 2, knee flexion angle 2, ...].
[0084] Step 403: Perform gait dynamics analysis on the multimodal human posture key point change sequence to obtain a time series of gait dynamics analysis results;
[0085] Among them, gait dynamics analysis is the time-varying changes of joint torque and load.
[0086] The dynamic analysis includes the analysis and calculation of parameters such as the flexion-extension torque and abduction-inversion torque of the hip, knee, and ankle joints.
[0087] For example, the time series of gait kinematics analysis results includes fluctuations in joint forces over time, such as [hip flexion-extension torque 1, ankle abduction-inversion torque 1, ...].
[0088] In this application, geometric calculations (such as distance and angle calculations) are performed on the coordinates of key points within each period. Dynamic parameters (such as torque) are derived using a biomechanical model. Time alignment: All parameters are synchronized according to timestamps to form a multi-parameter time series.
[0089] In this application, gait cycle analysis, dynamic analysis, and kinematic analysis are performed on the sequence of changes in key points of human posture during walking. Therefore, the analysis data is comprehensive and can obtain health prediction results for a variety of foot health conditions, including flat feet, high arches, hallux valgus, inversion foot, eversion foot, plantar fasciitis, and Achilles tendinitis. Thus, the detection range is wide and the application scenarios are rich.
[0090] Optionally, in another exemplary embodiment of this application, step 104 is implemented by the following step 501: sending the time series of gait cycle analysis results, the time series of gait kinematics analysis results, and the time series of gait dynamics analysis results to a deep learning model, so that the deep learning model outputs the foot health status of the target person.
[0091] The foot health status mentioned includes various foot conditions such as flat feet, high arches, hallux valgus, inversion foot, eversion foot, plantar fasciitis, and Achilles tendinitis. Of course, the foot health status also includes "healthy," meaning the feet are free of disease.
[0092] In the construction and training phase of the deep learning model, the time series of gait analysis results are used as the input of the deep learning model, and the results of foot health status (including flat feet, high arches, hallux valgus, inversion foot, valgus foot, plantar fasciitis, Achilles tendinitis and other foot health diseases) are used as labels to construct an initial deep learning model and train it to obtain a deep learning model that can predict foot health status through the time series of gait analysis results.
[0093] After training with large amounts of data to obtain a stable deep learning model, the time series of gait analysis results can be used as input to the deep learning model to predict the health status of a person's feet.
[0094] For further information on model training, please refer to existing materials on model training; this application will not elaborate further.
[0095] Furthermore, foot health status analysis reports can be generated based on the prediction results for users to review.
[0096] Optionally, see Figure 4In another exemplary embodiment of this application, after step 101, the method further includes steps 601 and 602:
[0097] Step 601: Filter the side view video and the rear view video to obtain the filtered side view video and rear view video.
[0098] Step 602: Align the selected side-view and rear-view videos by time points to obtain new side-view and rear-view videos.
[0099] The selection of side-view videos (captured by a binocular depth camera) and rear-view videos (captured by an RGB camera) should be based on three core dimensions: data quality, synchronization, and completeness of key events.
[0100] The following are the specific screening process and criteria:
[0101] 1. Data quality screening
[0102] (1) Clarity and completeness
[0103] Side-view video (depth):
[0104] Remove missing depth data segments caused by camera shake or occlusion (such as clothing covering the image).
[0105] Check if the depth map resolution meets the standard (e.g., the nominal resolution of a stereo camera is ≥640×480). If the resolution of a local area drops by more than 20%, it is marked as invalid.
[0106] Rear-view video (RGB):
[0107] Exclude frames with blurred human body outlines caused by excessively strong or dim lighting (such as backlighting or shadows).
[0108] Verify the color reproduction of the RGB image; discard any image with severe color cast (such as skin tones appearing reddish / blueish).
[0109] (2) Stability verification
[0110] Check for frame rate fluctuations in the two-view videos (e.g., nominally 60fps but actual fluctuation > ±2fps).
[0111] Remove any stuttering or frame skipping caused by device synchronization issues.
[0112] 2. Time Synchronization Filtering
[0113] (1) Initial time alignment
[0114] Method: Using the system clocks of the binocular depth camera and the RGB camera as a reference, record the absolute timestamps (such as UTC time) of the two devices when they start shooting.
[0115] Standard: The start time difference between two videos must be ≤50ms, otherwise they are considered unsynchronized.
[0116] (2) Dynamic time alignment
[0117] step:
[0118] Extract dynamic events that occur in both videos (such as the moment a foot touches the ground).
[0119] The event time offset is calculated by cross-correlation analysis, and the video timeline is adjusted to align the events.
[0120] Tools: Use OpenCV's cv2.matchTemplate function to match dynamic event frames.
[0121] (3) Synchronization accuracy verification
[0122] Two gait cycles were randomly selected, and the time difference between the foot's ground contact time in both perspectives was checked to see if it was ≤1 frame (approximately 16ms@60fps).
[0123] If the number of out-of-tolerance frames is greater than 2, then resynchronize or remove the segment.
[0124] 3. Critical event completeness screening
[0125] (1) Gait cycle coverage
[0126] Ensure the video contains at least two complete gait cycles (from left foot strike to the next left foot strike).
[0127] Remove gait cycle segments that are interrupted due to shooting distance being too close or too far.
[0128] (2) Key motion capture
[0129] Lateral view: The movements of the foot and ankle joints, such as dorsiflexion / plantar flexion and knee flexion and extension, must be clearly shown.
[0130] Rearward view: It is necessary to capture the pelvic rotation, hip adduction / abduction and other movements in their entirety.
[0131] Standard: If the number of missing frames for critical actions is greater than 10% (e.g., more than 0.5 seconds are missing in a 5-second video), then it is discarded.
[0132] 4. Post-screening processing
[0133] (1) Fragment extraction
[0134] Based on the filtering results, retain video clips that simultaneously meet the following conditions:
[0135] Data quality meets standards (clarity, stability).
[0136] Time synchronization error ≤ 1 frame.
[0137] It contains ≥2 complete gait cycles.
[0138] (2) Metadata annotation
[0139] Add metadata to the filtered videos, including:
[0140] Shooting device type (Binocular Depth / RGB).
[0141] Time synchronization offset (ms).
[0142] Key event timestamps (such as the start of a gait cycle).
[0143] 5. Alternative solutions (if there is insufficient data after filtering)
[0144] Data augmentation: Mirror the preserved segments and adjust the speed (±10%) to generate synthetic data.
[0145] Re-acquisition: If the critical gait cycle loss rate is >30%, a re-acquisition process is triggered.
[0146] Technical effect correlation
[0147] The relationship between filtering and portability: By eliminating low-quality data and reducing reliance on complex wearable devices, high-precision synchronous acquisition can be achieved with just a binocular depth + RGB camera.
[0148] Relationship between screening and detection range: The capture of complete gait cycles and key movements provides multi-dimensional data for kinematic / dynamic analysis, supporting the prediction of 12 foot diseases such as flat feet and hallux valgus.
[0149] The above screening process ensures that the video data input to the gait analysis module meets the requirements of quality, synchronization, and completeness, providing a reliable foundation for subsequent human posture key point detection and foot health prediction.
[0150] To align the time points of the lateral view video (captured by a binocular depth camera) and the rear view video (captured by an RGB camera), three steps are required: hardware synchronization, dynamic event matching, and timestamp calibration. This ensures that the time difference of the same biomechanical event (such as foot contact with the ground) in the two views is ≤1 frame (approximately 16ms@60fps).
[0151] The specific technical solution is as follows:
[0152] 1. Hardware-level synchronization (preprocessing stage)
[0153] (1) External triggering synchronization
[0154] Equipment configuration:
[0155] The same signal generator is used to generate TTL pulses, which simultaneously trigger the exposure of both the stereo depth camera and the RGB camera.
[0156] Ensure that the frame rates of the two cameras are consistent (e.g., both are set to 60fps) and that the exposure time overlap is ≥90%.
[0157] Effect: Eliminates initial time offset, ensuring that the time difference between the starting frames of two videos is ≤1ms.
[0158] (2) GPS timestamp synchronization
[0159] Applicable scenarios: When external triggers cannot be used (such as consumer-grade cameras).
[0160] step:
[0161] Equipped with two camera GPS modules to record UTC timestamps during shooting.
[0162] Align the video start time with the GPS timestamp.
[0163] Accuracy: Up to ±10ms (affected by GPS signal quality).
[0164] 2. Dynamic Event Matching (Core Alignment Phase)
[0165] (1) Definition of Critical Events
[0166] Select events with significant transients and co-occurrence of viewpoints within the gait cycle as alignment anchors, for example:
[0167] At the moment of foot contact with the ground: from a side view, the sole of the foot contacts the ground; from a rear view, the outline of the heel changes abruptly.
[0168] At the moment the toes leave the ground: in a side view, the toes lift up, and in a rear view, the length of the foot shortens.
[0169] (2) Event Detection Algorithm
[0170] Side view (depth video):
[0171] Depth gradient analysis was used to detect abrupt changes in the contact area between the foot and the ground.
[0172] Rearward view (RGB video):
[0173] Foot movement mutations can be detected using optical flow methods, or key points on the heel / toe can be identified using models such as YOLOv8.
[0174] (3) Cross-perspective event matching
[0175] Method: The Dynamic Time Warping (DTW) algorithm is used to match the temporal patterns of event sequences in two perspectives.
[0176] step:
[0177] Extract the timestamp sequence of key events from the two videos (e.g., [touchdown 1, touchdown 2, ...]).
[0178] Calculate the DTW distance matrix to find the optimal time alignment path.
[0179] Calculate the time offset Δt based on the path.
[0180] 3. Timestamp calibration (post-processing stage)
[0181] (1) Frame-level alignment
[0182] Adjust the backward video timeline based on Δt calculated by DTW:
[0183] If Δt>0: Insert Δt frames before the back video (repeat the first frame or interpolate).
[0184] If Δt < 0: Insert |Δt| frames before the lateral video.
[0185] 4. Alignment effect verification
[0186] (1) Quantitative assessment
[0187] Two gait cycles are randomly selected, and the time difference of foot ground contact time in the two viewpoints is calculated:
[0188] Acceptable criteria: mean time difference ≤ 1 frame (33ms), standard deviation ≤ 10ms.
[0189] (2) Qualitative verification
[0190] Visualize the alignment results:
[0191] Play the two videos simultaneously and observe whether the foot movements are consistent.
[0192] 5. Exception Handling
[0193] (1) Missing event
[0194] If event detection fails in a certain viewpoint (e.g., feet are occluded):
[0195] Timestamps are estimated using interpolation of adjacent events.
[0196] This segment is marked as "low confidence" and will be weighted less in subsequent analyses.
[0197] (2) Synchronization loss
[0198] If dynamic matching fails (e.g., DTW distance exceeds the threshold):
[0199] Revert to the hardware synchronization results and trigger the re-acquisition process.
[0200] Technical effect
[0201] Synchronization accuracy: up to ±1ms (hardware synchronization) or ±10ms (software synchronization), meeting the requirements of gait analysis.
[0202] Robustness: It adapts to different lighting and occlusion scenarios through multi-event matching and anomaly handling.
[0203] Compatibility: Supports consumer cameras (such as Intel RealSense D435i) and professional devices (such as Vicon motion capture systems).
[0204] The above method can achieve high-precision time alignment of lateral and rearward view videos, providing a reliable time reference for subsequent human posture key point fusion and foot health prediction.
[0205] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram can be found in [reference needed]. Figure 5 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores relevant data for a video-based foot health status detection method. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it can implement a video-based foot health status detection method.
[0206] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0207] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0208] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0209] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0210] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data that have been agreed to by the user or have been fully agreed to by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0211] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. In the embodiments provided in this application, any reference to memory, database, or other media can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0212] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units, etc., and are not limited to these.
[0213] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0214] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A video-based foot health status detection method, characterized in that, The method comprises: obtaining a rear-view video and a side-view video of a target person walking; generating a multi-modal human posture key point change sequence of the target person walking based on the rear-view video and the side-view video; performing gait analysis on the multi-modal human posture key point change sequence to obtain a gait analysis result time sequence; sending the gait analysis result time sequence to a deep learning model to make the deep learning model output a foot health status of the target person.
2. The video-based foot health status detection method of claim 1, wherein, The "generating a multi-modal human posture key point change sequence of the target person walking based on the rear-view video and the side-view video" is realized by the following method: sending the side-view video to a 3D-based human posture estimation model to output a 3D human posture key point change sequence; sending the rear-view video to a 2D-based human posture estimation model to output a 2D human posture key point change sequence; fusing the 3D human posture key point change sequence and the 2D human posture key point change sequence to obtain the multi-modal human posture key point change sequence.
3. The video-based foot health status detection method of claim 1, wherein, The gait analysis includes gait cycle analysis, gait kinematics analysis, and gait dynamics analysis, and the "performing gait analysis on the multi-modal human posture key point change sequence to obtain a gait analysis result time sequence" is realized by the following method: performing gait cycle analysis on the multi-modal human posture key point change sequence to obtain a gait cycle analysis result time sequence; performing gait kinematics analysis on the multi-modal human posture key point change sequence to obtain a gait kinematics analysis result time sequence; performing gait dynamics analysis on the multi-modal human posture key point change sequence to obtain a gait dynamics analysis result time sequence.
4. The video-based foot health status detection method of claim 3, wherein, The "sending the gait analysis result time sequence to a deep learning model" is realized by the following method: sending the gait cycle analysis result time sequence, the gait kinematics analysis result time sequence, and the gait dynamics analysis result time sequence to the deep learning model to make the deep learning model output the foot health status of the target person.
5. The video-based foot health status detection method of claim 1, wherein, After the "obtaining a rear-view video and a side-view video of a target person walking", the method further comprises: screening the side-view video and the rear-view video to obtain screened side-view video and rear-view video; aligning the time points of the screened side-view video and the rear-view video to obtain new side-view video and rear-view video.
6. The video-based foot health status detection method of claim 1, wherein, The foot health status includes: flat feet, high arches, hallux valgus, varus foot, valgus foot, plantar fasciitis, Achilles tendonitis, and other foot health diseases.
7. The video-based foot health status detection method according to any one of claims 1-6, wherein, The rear-view video is shot by an RGB camera, and the side-view video is shot by a binocular depth camera.
8. A computer device comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to implement the video-based foot health status detection method of any one of claims 1-7.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processor, implements the video-based foot health state detection method according to any one of claims 1-7.
10. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the video-based foot health state detection method according to any one of claims 1-7.