Foot stepping recognition method and apparatus, and device and storage medium

By distinguishing multiple target parts where the foot contacts the ground and identifying the stepping status and events of each part, the limitations of foot stepping recognition in existing technologies are solved, and accurate information interaction in a virtual reality environment is achieved.

WO2025067087A9PCT designated stage expired Publication Date: 2025-09-11BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/120218
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-26
Filing Date
2024-09-20
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Existing foot pedaling recognition methods usually directly analyze the entire foot pedaling action as a whole, resulting in recognition limitations and the inability to accurately recognize diverse pedaling actions.

Method used

By distinguishing at least two target parts where the user's foot contacts the ground, determining the pedaling state of each target part, and identifying pedaling events based on state changes, the pedaling events of multiple target parts are comprehensively analyzed to identify the overall pedaling action of the foot.

Benefits of technology

It achieves accurate recognition of foot stepping movements, ensures the diversity and comprehensiveness of stepping recognition, and supports precise information interaction in a virtual reality environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024120218_12092025_PF_FP_ABST
    Figure CN2024120218_12092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a foot stepping recognition method and apparatus, and a device and a storage medium. The method comprises: determining stepping states of at least two target parts of the feet of a user, wherein the target parts are contact parts between the feet and the ground when the feet make contact with the ground; and on the basis of a change in the stepping state of each target part, determining a stepping event of the target part. By means of the embodiments of the present application, stepping events of at least two target parts of feet can be accurately recognized, and accurate recognition of various stepping actions that are executed by the feet by means of different target parts is supported, thereby ensuring the diversity of foot stepping recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Foot stepping recognition method, device, equipment and storage medium

[0001] This application claims priority to the Chinese invention patent application entitled “Foot stepping identification method, device, equipment and storage medium” and application number 202311253633.9 filed on September 26, 2023. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of data processing technology, and in particular to a method, device, equipment and storage medium for identifying foot stepping. Background Art

[0003] In various virtual reality scenarios supported by Extended Reality (XR) technology, users can often interact with information by performing corresponding footstepping movements. For example, users can experience dance machines, music-based stepping games, and other similar experiences by performing footstepping movements in a virtual reality scenario.

[0004] Currently, existing foot stepping recognition methods usually directly analyze the stepping states represented by various stepping actions performed by the entire foot as a whole, which results in certain limitations in foot stepping recognition.

[0005] Summary of the Invention

[0006] The embodiments of the present application provide a foot stepping recognition method, apparatus, device and storage medium to achieve accurate recognition of stepping events at at least two target parts of the foot, ensuring diversity in foot stepping recognition.

[0007] In a first aspect, an embodiment of the present application provides a method for identifying footstepping, the method comprising:

[0008] determining the pedaling status of at least two target parts of the user's foot based on the pedaling motion data of the user's foot, wherein the target parts are parts of the user's foot that can contact the ground when the foot lands;

[0009] According to the change of the trampling state of each target part, the trampling event of the target part is determined.

[0010] In a second aspect, an embodiment of the present application provides a foot stepping recognition device, which includes:

[0011] a pedaling state determination module, configured to determine the pedaling states of at least two target parts of the user's foot based on the pedaling motion data of the user's foot, wherein the target parts are parts of the user's foot that can contact the ground when the foot lands;

[0012] The foot stepping determination module is used to determine the stepping event of each target part according to the stepping state change of the target part.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising:

[0014] A processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the foot stepping recognition method provided in the first aspect of the present application.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program, which enables a computer to execute the foot stepping recognition method provided in the first aspect of the present application.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which enables a computer to execute the foot stepping recognition method provided in the first aspect of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] FIG1 is a flow chart of a foot stepping recognition method provided by an embodiment of the present application;

[0019] FIG2 is an exemplary schematic diagram of a target part of a foot provided in an embodiment of the present application;

[0020] FIG3 is an exemplary schematic diagram of a stampede event at a target site provided by an embodiment of the present application;

[0021] FIG4 is a schematic diagram showing the principle of a process for determining an overall foot stamping event according to an embodiment of the present application;

[0022] FIG5 is a schematic structural diagram of a stepping recognition system provided in an embodiment of the present application;

[0023] FIG6a, FIG6b, FIG6c and FIG6d are respectively exemplary model structure diagrams of four different structures of the stepping recognition model provided in an embodiment of the present application;

[0024] FIG7 is a schematic structural diagram of a training data acquisition system for a stepping recognition model provided in an embodiment of the present application;

[0025] FIG8 is a functional block diagram of a foot stepping recognition device provided in an embodiment of the present application;

[0026] FIG9 is a schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0029] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or solution described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or solutions. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0030] Before introducing the specific technical solutions of this application, the application scenarios of this application are first described accordingly:

[0031] The implementation scheme of the present application can be applied to any scenario that can support the user's feet to trigger corresponding information interaction by performing various stepping movements. Among them, the application scenario may include triggering information interaction supported by a relevant device in the real environment in which the user is currently located through the stepping movement performed by the user's feet, such as a dance machine in a real environment. The application scenario can also include the user entering a virtual reality environment through an XR device, and triggering corresponding information interaction in the virtual reality environment through the stepping movement performed by the user's feet, such as a dance machine in a virtual reality environment, a music stepping game, etc. This application does not limit the specific application scenario, and can support the user's feet to trigger corresponding information interaction by performing various stepping movements.

[0032] Taking the virtual reality environment as an example, the application scenarios of this application can be exemplified as follows:

[0033] In a virtual reality environment, users can perform various stepping movements with their feet to trigger corresponding information interactions, such as dance machines and music stepping games in a virtual reality environment. In order to achieve foot stepping interactions in a virtual reality environment, the present application can display the corresponding virtual reality environment through the display of any electronic device that communicates with a display and one or more input devices. The electronic device can be any extended reality (XR) device, specifically including virtual reality (VR) devices, augmented reality (AR) devices, and mixed reality (MR) devices, etc., which are not limited in the present application.

[0034] The display may be any display screen that has a communication connection with the electronic device. For example, the display may be a display screen configured for a head-mounted display device on a VR device, an AR device, or an MR device, and this application does not limit this.

[0035] Moreover, in order to enable normal user interaction in a virtual reality environment, the present application can initiate corresponding interactive operations on the virtual reality environment displayed on the display through one or more input devices that communicate with the above-mentioned electronic devices, thereby supporting users to perform various interactions in the virtual reality environment.

[0036] The one or more input devices may be any control device or information collection device that has established a communication connection with the above-mentioned electronic device. For example, the one or more input devices may be a handle configured on a VR device, an AR device, or an MR device, or a collection module for detecting hand operations and eye movement information, or a voice collector for collecting user voice information, etc., and this application does not limit this.

[0037] In the present application, in order to accurately identify various pedaling actions performed by the user's feet, one or more input devices may be a collection module for detecting pedaling motion data when the user's feet perform various pedaling actions.

[0038] At present, in order to solve the problem of certain limitations in foot stepping recognition when directly analyzing the overall stepping action of the entire foot through a binary classification model to identify the stepping state of the foot, the inventive concept of this application is: based on the contactable parts between the user's foot and the ground when the user's foot lands, at least two target parts can be distinguished on the user's foot, and the stepping state of each target part can be determined. According to the change in the stepping state of each target part, the stepping event of the target part can be determined. Then, by comprehensively referring to the stepping events of at least two target parts of the foot, the overall stepping event of the foot can be determined, thereby achieving accurate identification of the foot stepping event. By comprehensively referring to the stepping events of multiple target parts of the foot, it supports accurate identification of various stepping actions performed by the foot through different target parts, ensuring the comprehensiveness of foot stepping recognition.

[0039] Figure 1 is a flow chart of a foot stepping recognition method provided in an embodiment of the present application. The method can be applied to XR devices, but is not limited thereto. The method can be executed by the foot stepping recognition device provided in the present application, wherein the foot stepping recognition device can be implemented by any software and / or hardware. For example, the foot stepping recognition device can be configured in an electronic device such as AR / VR / MR that can simulate virtual reality scenes. The present application does not impose any restrictions on the specific type of electronic device.

[0040] Through the technical solution of this application, the stepping state of at least two target parts of the ground that can be contacted by the user's foot when landing can be determined. Then, based on the change in the stepping state of each target part, the stepping event of the target part can be determined, thereby achieving accurate recognition of stepping events of at least two target parts of the foot, supporting accurate recognition of various stepping actions performed by the foot through different target parts, and ensuring diversity in foot stepping recognition.

[0041] Specifically, as shown in FIG1 , the method may include the following steps:

[0042] S110 , determining the pedaling states of at least two target parts of the user's foot according to the pedaling motion data of the user's foot.

[0043] In a virtual reality environment, information interaction can often be achieved by the user performing footstepping movements. To ensure accurate user interaction within the virtual reality environment, it is often necessary to accurately determine whether the user's foot is touching the ground within the virtual reality environment. This task, known as footstepping recognition, enables various interactions within the virtual reality environment.

[0044] Considering that a user's foot is not a simple point but consists of multiple different parts, for the footstepping recognition task, the user's foot can step on the ground through multiple different parts that can contact the ground to perform the corresponding stepping action. Since stepping on the ground through a part of the foot that can contact the ground when landing does not need to consider the state of other parts of the foot and the ground, this shows that the user's foot stepping actions are diverse.

[0045] Exemplarily, the stepping actions performed by the user's feet may include the following: stepping on the ground with the toes when the heels touch the ground, stepping on the ground with the heels when the toes touch the ground, stepping on the ground with the toes when the heels are suspended in the air, stepping on the ground with the heels when the toes are suspended in the air, stepping straight up and down with the toes and heels when the foot is parallel to the ground, etc.

[0046] Therefore, to ensure comprehensive footstepping recognition, the present application can determine at least two target foot parts by analyzing the contact areas between the user's foot and the ground when the foot lands, thereby supporting various footstepping actions performed using the target foot parts. The contact areas between the user's foot and the ground when the foot lands may include, but are not limited to, the toes, heel, toes, forefoot, and side soles.

[0047] For example, in order to support the convenient execution of the user's foot stepping action, as shown in FIG2 , the target parts in the present application may be the toes and heels of the foot.

[0048] It is understood that the user's foot can be a single foot or two feet, and this application does not limit this. Therefore, for a single foot or two feet, each single foot can include at least two target parts.

[0049] Then, when the user's foot performs the corresponding pedaling action, the present application can use the corresponding data acquisition equipment to obtain various pedaling motion data of the foot in real time. The pedaling motion data may include but is not limited to at least one of the following information: the motion position, direction, posture, acceleration, angular velocity, etc. of the user's foot when performing various pedaling actions.

[0050] Then, for each target part of the user's foot, by performing corresponding motion analysis on the various pedaling motion data of the foot, it is possible to judge whether each target part is in contact with the ground or has left the ground, thereby determining the pedaling state of each target part.

[0051] It can be seen that the stepping state of each target part can include two states: a stepping state and a lifting state. Among them, the distinguishing criterion between the stepping state and the lifting state can usually be the height of the target part of the foot from the ground. If the height of a target part from the ground is less than a preset threshold (for example, 5 cm), it means that the target part is approximately in contact with the ground, and it can be determined that the target part is in the stepping state. If the height of a target part from the ground is greater than or equal to a preset threshold (for example, 5 cm), it means that the target part is approximately away from the ground, and it can be determined that the target part is in the lifting state.

[0052] S120 , determining a trampling event of each target part according to a trampling state change of the target part.

[0053] Since the pedaling action performed by the user's foot is mainly the continuous movement of the user's foot stepping on and off the ground, that is, the pedaling state of the user's foot is constantly changing. Therefore, in order to accurately identify each pedaling action performed by the user's foot, the present application can characterize the change of the foot pedaling state through a pedaling event.

[0054] Therefore, after determining the pedaling states of at least two target parts of the user's foot, the present application can determine the pedaling event represented by each target part when the pedaling state changes by analyzing the changes in the pedaling state of the target part. Thus, in the same manner as above, the pedaling event of each target part can be determined.

[0055] As an optional implementation scheme in the present application, since the stepping state of each target part of the user's foot can include a stepping state and a lifting state, as shown in FIG3 , for each target part of the foot, the stepping event of the target part can be divided into the following two situations:

[0056] Case 1: If the target part changes from the grounded state to the lifted state, a lift event of the target part is generated.

[0057] That is to say, in the process of the user performing the corresponding stepping action through a certain target part of the foot, if the target part changes from the stepping state to the lifting state, such as from the stepping state represented by state "1" to the lifting state represented by state "0" in Figure 3, it means that the target part leaves the ground at the current transformation moment, thereby generating a lifting event of the target part.

[0058] Case 2: If the target part changes from the lifted state to the stepped-on state, a stepped-on event of the target part is generated.

[0059] That is to say, in the process of the user performing a corresponding stepping action through a certain target part of the foot, if the target part changes from a raised state to a stepped state, such as from the raised state represented by state "0" to the stepped state represented by state "1" in Figure 3, it means that the target part contacts the ground at the current moment of change, thereby generating a stepped event of the target part.

[0060] Based on the above, after determining the stepping events of at least two target parts of the user's foot, the present application can use the stepping events of any one or more of the target parts as trigger conditions for a certain interactive operation in the virtual reality environment. Then, after detecting the stepping events of one or more target parts, the corresponding interactive operation can be executed in the virtual reality environment.

[0061] The technical solution provided by the embodiments of the present application can determine the pedaling state of at least two target areas of the ground that are contactable by the user's foot when landing. Then, based on the pedaling state changes of each target area, a pedaling event for that target area is determined, thereby accurately identifying pedaling events for at least two target areas of the foot. This supports accurate recognition of various pedaling actions performed by the foot through different target areas, ensuring diversity in foot pedaling recognition.

[0062] In this application, in order to ensure comprehensive identification of foot stepping, after determining the stepping events of at least two target parts of the foot, this application can also determine the overall stepping event of the foot based on the stepping events of at least two target parts of the foot.

[0063] That is, after determining the stampede events of each target part of the foot, since each stampede event of the target part of the foot can represent a stampede event generated by the foot performing a stampede action, in order to ensure comprehensive identification of the user's stampede events, the present application can perform a comprehensive analysis of the stampede events of at least two target parts of the foot to obtain a complete stampede event, which is used as the overall stampede event of the foot in the present application.

[0064] In some possible implementations, in order to avoid repetition and redundancy of the overall pedaling event of the foot, the present application can adopt the following steps to determine the overall pedaling event of the foot: merge the pedaling events of each target part of the foot to obtain a pedaling event combination of the foot; filter the pedaling event combination according to the preset pedaling logic of the foot to obtain the overall pedaling event of the foot.

[0065] After determining the stampede events of each target part, the present application can combine the stampede events of each target part according to the execution time sequence of the stampede events of each target part to obtain a stampede event combination of the foot.

[0066] Since different foot-stepping application scenarios have different recognition requirements for foot-stepping tasks, the present application can pre-set a suitable foot-stepping logic in each foot-stepping application scenario according to the different foot-stepping recognition requirements in the present application, as the preset foot-stepping logic in the present application.

[0067] Therefore, for the foot stepping event combination, this application can adopt the preset stepping logic suitable for the current foot stepping application scenario, filter out some stepping events in the foot stepping event combination that do not conform to the preset stepping logic, and delete the filtered part of the stepping events from the stepping event combination to obtain the overall foot stepping event.

[0068] Taking the toes and heels as an example, as shown in FIG4 , the present application can respectively determine the toe stepping state and the heel stepping state, thereby determining the toe stepping event and the heel stepping event. In FIG4 , the upward arrow can represent a lifting event, and the downward arrow can represent a stepping event.

[0069] Then, according to the execution time sequence of each toe stamping event and each heel stamping event, the present application can merge each toe stamping event and each heel stamping event together to obtain a foot stamping event combination.

[0070] Assume that the preset pedaling logic of the foot is "the pedaling event and the lifting event of the foot appear alternately in sequence" and "the time interval between adjacent pedaling events is not less than the preset interval threshold (for example, 1s)". Then, according to the preset pedaling logic of "the pedaling event and the lifting event of the foot appear alternately in sequence", in the pedaling event combination of the foot, for the same pedaling event (such as the pedaling event or the lifting event) that appears continuously, this application can only retain the first pedaling event that appears, and delete the subsequent repeated pedaling events that appear continuously, such as the pedaling event and the lifting event marked in the pedaling event combination in Figure 4, so as to obtain the pedaling event combination after preliminary filtering.

[0071] Then, in the stampede event combination after preliminary filtering, according to the preset stampede logic of "the time interval between adjacent stampede events is not less than the preset interval threshold (for example, 1s)", this application can continue to delete some stampede events whose interval time with the previous stampede event is less than the preset interval threshold, such as the stepping event and lifting event marked in the stampede event combination after preliminary filtering in Figure 4, so as to obtain the overall stampede event of the foot.

[0072] Based on the above content, after determining the overall stepping event of the user's foot, this application can use the overall stepping event of the foot as a trigger condition for a certain interactive operation in the virtual reality environment. Then, after detecting the overall stepping event of the foot, the corresponding interactive operation can be performed in the virtual reality environment.

[0073] The technical solution provided in the embodiment of the present application determines the overall stepping event of the foot through a comprehensive analysis of the stepping events of at least two target parts of the foot, thereby achieving accurate identification of the foot stepping event. By comprehensively referring to the stepping events of multiple target parts of the foot, it supports accurate identification of various stepping actions performed by the foot through different target parts, thereby ensuring the comprehensiveness of foot stepping identification.

[0074] As an optional implementation scheme in the present application, in order to ensure accurate identification of the pedaling state of each target part of the user's foot, the present application can pre-train a pedaling recognition model, which can be used to accurately predict the pedaling state of each target part of the foot.

[0075] Then, for the pedaling state of each target part of the foot, the present application can input the pedaling motion data of the user's foot into a pre-built pedaling recognition model, and output the pedaling states of at least two target parts of the foot.

[0076] In this application, when a user's foot performs various pedaling motions, the application may use corresponding data acquisition equipment to obtain various pedaling motion data of the foot in real time. The pedaling motion data may include, but is not limited to, at least one of information such as the motion position, direction, posture, acceleration, and angular velocity of the user's foot when performing various pedaling motions.

[0077] In some implementations, the present application may pre-build a stepping recognition system to obtain the stepping motion data of the foot through the stepping recognition system.

[0078] As shown in FIG. 5 , the stepping recognition system in the present application may include an electronic device 510 and at least one inertial sensor 520 communicatively connected to the electronic device 510 .

[0079] Among them, the electronic device 510 can be an XR device, specifically including an AR device, a VR device, an MR device, etc., so as to present a corresponding virtual reality environment to the user.

[0080] Specifically, the number of inertial sensors 520 can be equal to the number of the user's feet, and each inertial sensor 520 can be worn at a corresponding location on the user's foot. Because the foot's various pedaling motions primarily represent movement at the center of the sole, it can also drive the corresponding calf to perform the same movement. Therefore, the corresponding location on the foot can be either the center of the sole or the corresponding calf.

[0081] Then, each inertial sensor 520 can collect inertial data of the worn foot when performing a specified pedaling action in real time. The inertial data may include but is not limited to relevant directional posture data and angular velocity information of the foot when performing various pedaling actions. Then, the pedaling motion data of the user's foot can include the inertial data collected by each inertial sensor 520 worn on the associated parts of each foot. Then, the pedaling motion data is transmitted to the electronic device 510, so that the electronic device 510 can obtain the pedaling motion data of each foot. Then, the electronic device 510 can determine the pedaling state of at least two target parts of the foot based on the pedaling motion data of the foot.

[0082] Among them, as for the number of inertial sensors, assuming that this application mainly focuses on the stepping events of one of the left and right feet, then one inertial sensor can be configured and worn on the relevant part of the foot that is focused on this time, so as to collect inertial data when the foot performs various stepping actions as the stepping motion data of the foot.

[0083] Assuming that this application mainly focuses on the pedaling events of two feet consisting of the left and right feet, two inertial sensors can be configured and worn on the relevant parts of the left and right feet respectively, so that the inertial data of the left and right feet when performing various pedaling actions can be collected respectively by the two inertial sensors as the pedaling motion data of the left and right feet.

[0084] It should be noted that, in addition to wearing inertial sensors 520 on relevant parts of the user's foot, electronic components with certain computing capabilities can also be worn to independently analyze the stepping events of multiple target parts of the user's foot and / or the overall stepping events of the foot. Therefore, it can be seen that the electronic device 510 and the inertial sensor 520 in this application can not be a one-to-many relationship, but a one-to-one relationship.

[0085] Moreover, for the stepping recognition model in this application, the stepping recognition model may include a skeleton network (backbone network) and multiple head networks (head networks), and the total number of head networks is equal to the total number of target parts of the feet, and the target parts corresponding to each head network are different.

[0086] On the other hand, the present application can also set at least two stepping recognition models, each stepping recognition model can include a skeleton network with the same structure and at least one head network, and the total number of head networks in each stepping recognition model is equal to the total number of target parts of the foot, and the target parts corresponding to each head network are different.

[0087] In both of the aforementioned pedaling recognition models, the skeleton network can be composed of multiple basic blocks (basic_block) with the same structure, which are used to predict the temporal fusion features of the user's foot pedaling motion data. These temporal fusion features are then input into each head network connected to the skeleton network. This allows each head network in each pedaling recognition model to obtain the temporal fusion features of the user's foot pedaling motion data. Each head network can then be used to predict the pedaling state of a target foot part based on the temporal fusion features.

[0088] In other words, by setting the total number of head networks within each stepping recognition model to be equal to the total number of target parts of the foot, all head networks within each stepping recognition model can be mapped one-to-one to each target part of the foot, allowing one head network to be used to predict the stepping state of one target part. Then, using all head networks within each stepping recognition model, the stepping state of each target part of the foot can be predicted.

[0089] Furthermore, each basic block within the backbone network of each stepping recognition model can include sequentially connected fully connected layers (FC), a layer normalization layer (LayerNorm), and a recurrent neural network (RNN). The head network in each stepping recognition model can be a multilayer perceptron (MLP).

[0090] For example, if the feet in this application are composed of two feet, each consisting of a left and a right foot, and the target parts of the feet are the toes and heels, then the application ultimately needs to output four pieces of information through the stepping recognition model: the stepping status of the left toe, the stepping status of the left heel, the stepping status of the right toe, and the stepping status of the right heel. Therefore, the total number of head networks in each stepping recognition model can be four.

[0091] Therefore, depending on the number of stepping recognition models, the number of head networks in each stepping recognition model constructed in this application may be different. Figures 6a, 6b, 6c, and 6d are four different exemplary model structure diagrams of the stepping recognition model in this application.

[0092] In the present application, after obtaining the pedaling motion data of the user's foot, the pedaling motion data can be directly input into the pedaling recognition model constructed above. Then, the pedaling motion data of the foot is sequentially subjected to corresponding feature fusion processing through multiple basic blocks (basic_block) with the same structure in the skeleton network (backbone network) of the pedaling recognition model, thereby obtaining corresponding time series fusion features. Then, each head network (head network) in the pedaling recognition model continues to perform corresponding feature mapping processing on the time series fusion features to determine the pedaling state of each target part of the foot.

[0093] When there are at least two pedaling recognition models, the present application can input the pedaling motion data of the user's feet into at least two pre-built pedaling recognition models respectively, and output the pedaling status of at least two target parts of the foot. Thus, for the pedaling motion data input into each pedaling recognition model, the pedaling motion data of the foot can be sequentially subjected to corresponding feature fusion processing through multiple basic blocks (basic_block) with the same structure in the skeleton network (backbone network) of each pedaling recognition model, thereby obtaining corresponding temporal fusion features. Then, the temporal fusion features are further subjected to corresponding feature mapping processing through each head network (head network) in each pedaling recognition model to determine the pedaling status of each target part of the foot.

[0094] In some feasible methods, in order to ensure the recognition accuracy of the pedaling state of each target part of the foot, the present application outputs the pedaling state of each target part of the foot through a pedaling recognition model, which can be specifically: preprocessing the pedaling motion data of the user's foot to obtain the pedaling feature vector of the foot; inputting the pedaling feature vector into the pre-built pedaling recognition model, and outputting the pedaling state of at least two target parts of the foot.

[0095] That is, after obtaining the pedaling motion data of the user's foot, since the pedaling motion data has multiple structures, it is not convenient for the pedaling recognition model to efficiently process it. Therefore, the present application can pre-process the pedaling motion data of the user's foot and transform it into a one-dimensional feature vector, which is used as the foot pedaling feature vector in the present application.

[0096] For example, when the user's foot performs various pedaling movements, the present application can collect the user's foot pedaling motion data in real time, indicating that the user's foot pedaling motion data is a set of time series data, which can be expressed as: X = {x1, x2, ..., x T}, represents the T-frame inertial data collected by each inertial sensor worn on the foot. Among them, the stepping motion data x of each frame t It may include but is not limited to: the rotation matrix corresponding to the direction and posture information collected by each inertial sensor worn on the user's feet, and the angular velocity information.

[0097] Then, for each frame of pedaling motion data x t , this application can convert the rotation matrix corresponding to the direction and attitude information collected by each inertial sensor into a corresponding one-dimensional vector. Moreover, using the pedaling motion data x of the previous frame t-1 The inverse of the rotation matrix in is multiplied by the stepping motion data x of the frame t The rotation matrix in can be used to obtain the corresponding angular velocity information and convert it into another one-dimensional vector. Then, the above two one-dimensional vectors and the angular velocity information collected by each inertial sensor are combined to obtain the stepping motion data x of the frame. t The corresponding stepping feature vector f t .

[0098] After determining the foot's stepping feature vector, the present application can input the stepping feature vector into a pre-built stepping recognition model. Then, the stepping feature vector is processed accordingly through the skeleton network and each head network in the stepping recognition model to predict the stepping state of each target part of the foot.

[0099] In one or more embodiments of the present application, to ensure that the pedaling recognition model accurately identifies the pedaling state of each target foot part, the present application requires pre-specifying a large number of users' feet performing various pedaling actions to obtain a large amount of training data. A corresponding loss function can then be set within the pedaling recognition model to calculate the loss between the predicted pedaling state values ​​of each target foot part after inputting the large amount of training data and the true pedaling state values ​​represented by the sample labels. This allows the pedaling recognition model to be continuously updated until the loss of the pedaling recognition model converges.

[0100] Among them, the loss function in the stepping recognition model can be a binary cross entropy loss (BCE loss for short).

[0101] As for the training data of the pedaling recognition model, the present application may pre-build a training data acquisition system to obtain the corresponding training data through the training data acquisition system.

[0102] As shown in FIG. 7 , the training data collection system may include a computer device 710 , at least one inertial sensor 720 communicatively connected to the computer device 710 , at least two rigid bodies 730 of different shapes, and an optical motion capture device 740 .

[0103] Specifically, the number of inertial sensors 720 is equal to the number of the user's feet, and each inertial sensor 720 can be worn at a corresponding location on the user's foot. Each inertial sensor 720 can then be used to collect inertial data when the foot performs a specified pedaling motion and transmit the inertial data as a corresponding training sample to the computer device 710.

[0104] The number of rigid bodies 730 is equal to the total number of target parts of the user's feet, and each rigid body 730 is fixed one-to-one at each target part of the user's feet, so that rigid bodies of different shapes are used to intuitively represent different target parts, so that the optical motion capture device 740 can easily distinguish the actual stepping state of each target part by capturing rigid body images of different shapes of each target part.

[0105] The optical motion capture device 740 can include a surround camera array to capture omnidirectional images of various shapes of rigid bodies worn by various target parts of the user's foot as they perform various pedaling motions from various angles, thereby conveniently distinguishing the true pedaling state of each target part. The optical motion capture device 740 can then use the surround camera array to capture a rigid body image of each target part of the foot as the foot performs a specified pedaling motion, thereby determining the true pedaling state value of each target part, and transmitting the true pedaling state value of each target part as the corresponding sample label to the computer device 710.

[0106] It is understandable that the computer device 710 in the training data acquisition system and the electronic device 510 in the stepping recognition system can be the same device or different devices, such as a head-mounted display device, and this application does not limit this.

[0107] Therefore, during the training data acquisition stage, according to the stepping events of different feet that the stepping recognition model terminal focuses on, a large number of users can be required to fix rigid bodies 730 of different shapes on each target part of the foot that is of focus, and wear corresponding inertial sensors 720 at the associated parts of the foot that are of focus.

[0108] Then, when the user's foot performs various designated pedaling movements, the various inertial sensors 720 worn at the associated parts of the user's foot can collect inertial data of the worn foot when performing the designated pedaling movement in real time. The inertial data may include, but is not limited to, relevant directional posture data and angular velocity information of the foot when performing various pedaling movements. Then, the present application can use the inertial data collected by each inertial sensor 720 as historical pedaling motion data of the corresponding foot on which the inertial sensor 720 is worn, and as a corresponding training sample. Furthermore, the above-mentioned training sample is transmitted to the computer device 710 together with the timestamp represented by the collection time of the training sample.

[0109] Moreover, at the same time point, the optical motion capture device 740 can use a surround camera array to capture a rigid body of different shapes fixed to each target part of the foot when the corresponding foot on which each inertial sensor 720 is worn performs a specified stepping action, thereby obtaining a rigid body image corresponding to each target part. Then, the optical motion capture device 740 determines the height of the rigid body from the ground in the rigid body image by performing corresponding feature positioning on the rigid body image of each target part, thereby determining whether each target part is in contact with the ground, and thus obtaining the true value of the stepping state of each target part. Furthermore, the true value of the stepping state of each target part can be used as the corresponding sample label and transmitted to the computer device 710 together with the timestamp represented by the shooting time of the rigid body image.

[0110] Then, the computer device 710 can obtain a large number of training samples and sample labels. By combining the training samples and sample labels under the same timestamp, the corresponding training data can be obtained, thereby accurately training the stepping recognition model in this application.

[0111] FIG8 is a block diagram of a foot stepping recognition device provided by an embodiment of the present application. As shown in FIG8 , the foot stepping recognition device 800 may include:

[0112] a pedaling state determination module 810 for determining the pedaling state of at least two target parts of the user's foot based on the pedaling motion data of the user's foot, wherein the target parts are parts of the user's foot that can contact the ground when the foot lands;

[0113] The foot stepping determination module 820 is used to determine a stepping event of each target part according to the stepping state change of the target part.

[0114] In some implementations, the pedaling state determination module 810 may be specifically configured to:

[0115] The pedaling motion data of the user's foot is input into a pre-built pedaling recognition model, and the pedaling states of at least two target parts of the foot are output.

[0116] In some implementations, the pedaling state determination module 810 may be specifically configured to:

[0117] Preprocessing the pedaling motion data of the user's foot to obtain a pedaling feature vector of the foot;

[0118] The stepping feature vector is input into a pre-built stepping recognition model, and the stepping status of at least two target parts of the foot is output.

[0119] In some implementations, the stepping recognition model includes a skeleton network and multiple head networks, and the total number of the head networks is equal to the total number of the target parts of the foot, and the target parts corresponding to each of the head networks are different.

[0120] In some implementations, there are at least two stepping recognition models, each of which includes a skeleton network and at least one head network of the same structure, and the total number of head networks in each of the stepping recognition models is equal to the total number of the target parts of the foot, and each head network corresponds to a different target part;

[0121] Accordingly, the pedaling state determination module 810 may be specifically configured to:

[0122] The pedaling motion data of the user's foot are respectively input into at least two pre-built pedaling recognition models, and the pedaling states of at least two target parts of the foot are output.

[0123] In some implementations, the training data of the stepping recognition model is obtained through a pre-built training data acquisition system; the training data acquisition system includes a computer device, at least one inertial sensor in communication with the computer device, at least two rigid bodies of different shapes, and an optical motion capture device; wherein,

[0124] The number of the inertial sensors is equal to the number of the feet, and the inertial sensors are worn at relevant parts of the feet, for collecting inertial data when the feet perform a specified stepping action, and transmitting the inertial data as corresponding training samples to the computer device;

[0125] The number of the rigid bodies is equal to the total number of the target parts of the foot, and each rigid body is fixed to each target part of the foot in a one-to-one correspondence;

[0126] The optical motion capture device includes a surround camera array for acquiring a rigid body image of each target part of the foot when the foot performs a specified stepping action, thereby determining a true value of the stepping state of the target part, and transmitting the true value of the stepping state of each target part as a corresponding sample label to the computer device;

[0127] The computer device combines the training samples and sample labels at the same timestamp to obtain corresponding training data.

[0128] In some implementations, the foot stepping motion data is obtained through a pre-built stepping recognition system; the stepping recognition system includes an electronic device and at least one inertial sensor in communication with the electronic device; wherein,

[0129] The number of the inertial sensors is equal to the number of the feet, and the inertial sensors are worn at relevant parts of the feet, for collecting inertial data when the feet perform a specified stepping action, and transmitting the inertial data as stepping motion data of the feet to the electronic device;

[0130] The electronic device is used to determine the pedaling status of at least two target parts of the foot according to the pedaling motion data of the foot.

[0131] In some implementations, the stepping state includes a stepping state and a lifting state. The foot stepping determination module 820 may be specifically configured to:

[0132] For each target part of the foot, if the target part changes from a grounded state to a lifted state, a lift event of the target part is generated;

[0133] If the target part changes from the lifted state to the stepped-on state, a stepped-on event of the target part is generated.

[0134] In some implementations, the foot stepping recognition device 800 may further include:

[0135] The whole foot stepping recognition module is used to determine the whole foot stepping event according to the stepping events of at least two target parts of the foot.

[0136] In some implementations, the whole foot stepping recognition module can be specifically used to:

[0137] Merging the stepping events of each target part of the foot to obtain a stepping event combination of the foot;

[0138] According to the preset stepping logic of the foot, the stepping event combination is filtered to obtain the overall stepping event of the foot.

[0139] In the embodiment of the present application, for at least two target areas of the ground that can be contacted by the user's foot when landing, the pedaling state of each target area can be determined. Then, based on the pedaling state changes of each target area, a pedaling event for that target area is determined, thereby accurately identifying pedaling events for at least two target areas of the foot. This supports accurate recognition of various pedaling actions performed by the foot through different target areas, ensuring diversity in foot pedaling recognition.

[0140] It should be understood that the device embodiment and the method embodiment in the present application may correspond to each other, and similar descriptions may refer to the method embodiment in the present application. To avoid repetition, they will not be described here.

[0141] Specifically, the device 800 shown in Figure 8 can execute any method embodiment provided in this application, and the aforementioned and other operations and / or functions of each module in the device 800 shown in Figure 8 are respectively for implementing the corresponding processes of the above-mentioned method embodiments. For the sake of brevity, they will not be repeated here.

[0142] The above-mentioned method embodiment of the embodiment of the present application is described from the perspective of the functional module in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in the form of hardware, can be implemented by instructions in the form of software, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above-mentioned method embodiment in conjunction with its hardware.

[0143] FIG9 is a schematic block diagram of an electronic device provided in an embodiment of the present application.

[0144] As shown in FIG9 , the electronic device 900 may include:

[0145] The memory 910 and the processor 920 are configured to store computer programs and transmit the program code to the processor 920. In other words, the processor 920 can call and run the computer program from the memory 910 to implement the method in the embodiment of the present application.

[0146] For example, the processor 920 may be configured to execute the above method embodiments according to instructions in the computer program.

[0147] In some embodiments of the present application, the processor 920 may include but is not limited to:

[0148] General-purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0149] In some embodiments of the present application, the memory 910 includes but is not limited to:

[0150] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0151] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 910 and executed by the processor 920 to implement the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program on the electronic device 900.

[0152] As shown in FIG9 , the electronic device may further include:

[0153] The transceiver 930 may be connected to the processor 920 or the memory 910 .

[0154] The processor 920 may control the transceiver 930 to communicate with other devices. Specifically, the processor 920 may send information or data to other devices or receive information or data sent by other devices. The transceiver 930 may include a transmitter and a receiver. The transceiver 930 may further include one or more antennas.

[0155] It should be understood that the various components in the electronic device 900 are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0156] The present application also provides a computer storage medium having a computer program stored thereon, which enables the computer to perform the method of the above method embodiment when the computer program is executed by the computer.

[0157] An embodiment of the present application further provides a computer program product comprising a computer program / instruction, which, when executed by a computer, enables the computer to perform the method of the above method embodiment.

[0158] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0159] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for identifying footstepping, comprising: determining the pedaling status of at least two target parts of the user's foot based on the pedaling motion data of the user's foot, wherein the target parts are parts of the user's foot that can contact the ground when the foot lands; According to the change of the trampling state of each target part, the trampling event of the target part is determined.

2. The method according to claim 1, wherein determining the pedaling states of at least two target parts of the user's foot according to the pedaling motion data of the user's foot comprises: The pedaling motion data of the user's foot is input into a pre-built pedaling recognition model, and the pedaling states of at least two target parts of the foot are output.

3. The method according to claim 2, wherein inputting the pedaling motion data of the user's foot into a pre-built pedaling recognition model and outputting the pedaling states of at least two target parts of the foot comprises: Preprocessing the pedaling motion data of the user's foot to obtain a pedaling feature vector of the foot; The stepping feature vector is input into a pre-built stepping recognition model, and the stepping status of at least two target parts of the foot is output.

4. The method according to claim 2, wherein the stepping recognition model includes a skeleton network and multiple head networks, and the total number of the head networks is equal to the total number of the target parts of the feet, and the target parts corresponding to each of the head networks are different.

5. The method according to claim 2, wherein there are at least two stepping recognition models, each of which includes a skeleton network and at least one head network of the same structure, and the total number of head networks in each stepping recognition model is equal to the total number of each target part of the foot, and each head network corresponds to a different target part; Accordingly, the stepping motion data of the user's foot is input into a pre-built stepping recognition model to output the stepping status of at least two target parts of the foot, including: The pedaling motion data of the user's foot are respectively input into at least two pre-built pedaling recognition models, and the pedaling states of at least two target parts of the foot are output.

6. The method according to claim 2, wherein the training data of the stepping recognition model is obtained through a pre-built training data acquisition system; the training data acquisition system includes a computer device, at least one inertial sensor in communication with the computer device, at least two rigid bodies of different shapes, and an optical motion capture device; wherein, The number of the inertial sensors is equal to the number of the feet, and the inertial sensors are worn at relevant parts of the feet, for collecting inertial data when the feet perform a specified stepping action, and transmitting the inertial data as corresponding training samples to the computer device; The number of the rigid bodies is equal to the total number of the target parts of the foot, and each rigid body is fixed to each target part of the foot in a one-to-one correspondence; The optical motion capture device includes a surround camera array for acquiring a rigid body image of each target part of the foot when the foot performs a specified stepping action, thereby determining a true value of the stepping state of the target part, and transmitting the true value of the stepping state of each target part as a corresponding sample label to the computer device; The computer device combines the training samples and sample labels at the same timestamp to obtain corresponding training data.

7. The method according to claim 2, wherein the pedaling motion data of the foot is obtained through a pre-built pedaling recognition system; the pedaling recognition system comprises an electronic device and at least one inertial sensor in communication with the electronic device; wherein, The number of the inertial sensors is equal to the number of the feet, and the inertial sensors are worn at relevant parts of the feet, for collecting inertial data when the feet perform a specified stepping action, and transmitting the inertial data as stepping motion data of the feet to the electronic device; The electronic device is used to determine the pedaling status of at least two target parts of the foot according to the pedaling motion data of the foot.

8. The method according to claim 1, wherein the stepping state includes a stepping state and a lifting state, and determining the stepping event of each target part according to the stepping state change of the target part comprises: For each target part of the foot, if the target part changes from a grounded state to a lifted state, a lift event of the target part is generated; If the target part changes from the lifted state to the stepped-on state, a stepped-on event of the target part is generated.

9. The method according to claim 1, further comprising: An overall stepping event of the foot is determined according to the stepping events of at least two target parts of the foot.

10. The method according to claim 9, wherein determining the overall stepping event of the foot according to the stepping events of at least two target parts of the foot comprises: Merging the stampede events of each target part of the foot to obtain a stampede event combination of the foot; According to the preset stepping logic of the foot, the stepping event combination is filtered to obtain the overall stepping event of the foot.

11. A foot stepping recognition device, comprising: a pedaling state determination module, configured to determine the pedaling states of at least two target parts of the user's foot based on the pedaling motion data of the user's foot, wherein the target parts are parts of the user's foot that can contact the ground when the foot lands; The foot stepping determination module is used to determine the stepping event of each target part according to the stepping state change of the target part.

12. An electronic device comprising: processor; as well as a memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the foot stepping recognition method according to any one of claims 1 to 10 by executing the executable instructions.

13. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method for identifying foot stepping according to any one of claims 1 to 10 is implemented.

14. A computer program product comprising a computer program / instruction, which, when the computer program / instruction contained in the computer program product is run on an electronic device, enables the electronic device to execute the foot step recognition method according to any one of claims 1 to 10.