IMU encoder
By pre-training an IMU encoder and using a multimodal dataset and combined loss training of multiple modal encoders, the problem of difficult IMU data annotation is solved, generating quantitative representations suitable for downstream tasks and improving the interpretation and annotation capabilities of IMU data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2025-11-19
- Publication Date
- 2026-05-22
AI Technical Summary
Obtaining large amounts of labeled IMU data is difficult, making it challenging to train a good IMU encoder, especially when building large-scale, high-quality datasets.
A pre-training method is employed, utilizing a multimodal dataset and multiple modal encoders to determine a pre-trained IMU encoder. The encoder is trained by combining self-supervised (SS) loss, multimodal (MM) loss, and nearest neighbor (NN) loss to generate quantitative representations that can be used in downstream tasks.
It has been realized that an efficient IMU encoder can be trained with limited labeled data, which can be effectively applied to tasks such as classifiers, data matching and data analysis, and improves the interpretation and annotation capabilities of IMU data.
Smart Images

Figure CN122072830A_ABST
Abstract
Description
Technical Field
[0001] Various exemplary embodiments relate to the field of computer science, and more particularly to electronic devices, methods, computer-readable storage media, and computer program products for pre-training and using inertial measurement unit (IMU) encoders. Background Technology
[0002] Wearable devices can be equipped with / embedded inertial measurement unit (IMU) sensors, including accelerometers, gyroscopes, and magnetometers, which track human movement, acceleration, and orientation. IMU data provides valuable insights into human physical and emotional behavior, thus playing a crucial role in health monitoring and overall well-being. For example, step count data from IMU sensors has been shown to be one of the most effective indicators of cognitive impairment progression in older adults. This potential has inspired communities to collect large amounts of IMU data in time-series format. However, obtaining large amounts of labeled IMU data remains a major challenge when modeling using machine learning (ML) methods because IMU time series data are inherently difficult to interpret and label, even for experts. Summary of the Invention
[0003] Overall, exemplary embodiments of this disclosure provide a solution for pre-training and using an IMU encoder.
[0004] In a first aspect, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory storing instructions, wherein the instructions, when executed by the at least one processor, cause the electronic device to at least: obtain a pre-trained IMU encoder, wherein the pre-trained IMU encoder is determined based on a multimodal dataset together with multiple modal encoders; determine a model output by inputting a model input to the pre-trained IMU encoder, wherein the model input indicates data obtained from one or more IMU sensors; and provide the model output to a downstream task.
[0005] In a second aspect, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory storing instructions, wherein the instructions, when executed by the at least one processor, cause the electronic device to at least: obtain a multimodal dataset comprising multiple segments, wherein the segments are time-aligned and include IMU data samples and multiple modal data samples having multiple types; and determine a pre-trained IMU encoder based on the multimodal dataset together with multiple modal encoders corresponding to the multiple modalities.
[0006] In a third aspect, a method is provided. This method includes: obtaining a pre-trained IMU encoder, wherein the pre-trained IMU encoder is determined based on a multimodal dataset along with multiple modal encoders; determining a model output by inputting a model input to the pre-trained IMU encoder, wherein the model input indicates data obtained from one or more IMU sensors; and providing the model output to a downstream task.
[0007] In a fourth aspect, a method is provided. This method includes: obtaining a multimodal dataset comprising multiple segments, wherein the segments are time-aligned and include IMU data samples and multiple modal data samples of multiple types; and determining a pre-trained IMU encoder based on the multimodal dataset together with multiple modal encoders corresponding to the multiple modalities.
[0008] In a fifth aspect, an apparatus is provided. The apparatus includes: components for obtaining a pre-trained IMU encoder, wherein the pre-trained IMU encoder is determined based on a multimodal dataset together with multiple modal encoders; components for determining a model output by inputting a model input to the pre-trained IMU encoder, wherein the model input indicates data obtained from one or more IMU sensors; and components for providing the model output to a downstream task.
[0009] In a sixth aspect, an apparatus is provided. The apparatus includes: components for acquiring a multimodal dataset comprising multiple segments, wherein the segments are time-aligned and include IMU data samples and multiple modal data samples having multiple types; and components for determining a pre-trained IMU encoder based on the multimodal dataset together with multiple modal encoders corresponding to the multiple modalities.
[0010] In a seventh aspect, there is an apparatus. The apparatus includes: an acquisition circuit configured to acquire a pre-trained IMU encoder, wherein the pre-trained IMU encoder is determined based on a multimodal dataset together with multiple modal encoders; a determination circuit configured to determine a model output by inputting a model input to the pre-trained IMU encoder, wherein the model input indicates data obtained from one or more IMU sensors; and a providing circuit configured to provide the model output to a downstream task.
[0011] In the eighth aspect, there is an apparatus. The apparatus includes: an acquisition circuit configured to acquire a multimodal dataset comprising multiple segments, wherein the segments are time-aligned and include IMU data samples and multiple modal data samples having multiple types; and a determination circuit configured to determine a pre-trained IMU encoder based on the multimodal dataset together with multiple modal encoders corresponding to the multiple modalities.
[0012] In a ninth aspect, a non-transitory computer-readable medium is provided, comprising program instructions for causing the apparatus to perform at least the methods of the third or fourth aspect.
[0013] In a tenth aspect, a computer program including instructions is provided that, when executed by a device, causes the device to perform at least the method of the third or fourth aspect.
[0014] It should be understood that the summary portion is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0015] Some exemplary embodiments will now be described with reference to the accompanying drawings, in which: Figure 1 An exemplary schematic diagram shows an environment in which some exemplary embodiments of the present disclosure may be implemented; Figure 2 A flowchart illustrating a method for using a pre-trained IMU encoder according to some exemplary embodiments of the present disclosure is shown; Figure 3 A flowchart is shown illustrating a method for pre-training an IMU encoder according to some exemplary embodiments of the present disclosure; Figure 4 An exemplary architecture for pre-training an IMU encoder is shown according to some exemplary embodiments of the present disclosure; Figure 5 An exemplary architecture of an IMU encoder according to some exemplary embodiments of the present disclosure is shown; Figure 6 Exemplary schematic diagrams of nearest neighbor supervision according to some exemplary embodiments of the present disclosure are shown; and Figure 7 A schematic block diagram of an exemplary device that can be used to implement embodiments of the present disclosure is shown.
[0016] Throughout the accompanying drawings, unless otherwise stated, the same or similar reference numerals denote the same or similar elements. Detailed Implementation
[0017] The principles of this disclosure will now be described with reference to some exemplary embodiments. It should be understood that these embodiments are described for illustrative purposes only and to assist those skilled in the art in understanding and implementing this disclosure, without imposing any limitation on the scope of this disclosure. This disclosure described herein can be implemented in various ways other than those described below.
[0018] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0019] References to "an embodiment," "embodiment," "exemplary embodiment," etc., in this disclosure indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment needs to include that particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Moreover, when a particular feature, structure, or characteristic is described in connection with an embodiment, whether explicitly described or not, it is considered that incorporating such a feature, structure, or characteristic into other embodiments is within the knowledge of those skilled in the art.
[0020] It should be understood that although the terms “first” and “second” may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0021] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the terms “comprising,” “including,” “having,” “possessing,” “including,” and / or “containing” as used herein specify the presence of stated features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. As used herein, “at least one of the following: ” and “at least one of ” and similar wording (where the list of two or more elements is connected by “and” or “or”) means at least any one of the elements, or at least any two or more of the elements, or at least all of the elements.
[0022] As used in this application, the term "circuit" may refer to one or more, or all of the following: (a) Hardware circuit implementation only (such as implementation in analog and / or digital circuits only). (b) A combination of hardware circuitry and software, such as (if applicable): (i) A combination of analog and / or digital hardware circuitry and software / firmware, and (ii) Any part of a device (such as a mobile phone or server) that has software (one or more) of hardware processors (including one or more digital signal processors), software, and memory (one or more) working together to enable the device to perform various functions, and (c) One or more hardware circuits and / or one or more processors, such as one or more microprocessors or a portion thereof, that require software (e.g., firmware) to operate (but such software may not exist when it is not required to operate).
[0023] This definition of "circuit" applies to all uses of the term in this application, including in any claim. As another example, as used herein, the term "circuit" also covers only hardware circuitry or a processor (or processors) or a portion thereof and its accompanying software and / or firmware implementation. For example, if applicable to a particular claim element, the term "circuit" also covers baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices.
[0024] As used herein, the term "electronic device" refers to a device, server, or system with computing capabilities. Electronic devices may include, but are not limited to, portable computers, desktop computers, laptop embedded devices (LEE), laptop devices (LME), medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronic devices, devices operating on commercial and / or industrial wireless networks, etc.
[0025] As mentioned above, obtaining a large amount of labeled IMU data presents challenges, which may make it difficult to acquire well-trained IMU encoders.
[0026] A promising solution to the problem of scarce labels is the "train once, adapt many times" approach. This involves initially training an encoder on a large-scale dataset of unlabeled or weakly labeled data. Then, a smaller ML model is trained on top of this task-specific (often frozen) encoder using a relatively small amount of labeled data. While this approach has shown significant success in image, video, audio, and natural language processing, its potential for IMU data remains limited, primarily due to the challenges of constructing large-scale, high-quality datasets.
[0027] Embodiments of this disclosure provide a solution for training and using an IMU encoder. In this solution, a pre-trained IMU encoder is available, and the model output can therefore be determined based on the model input using the pre-trained encoder. The pre-trained IMU encoder is determined based on a multimodal dataset along with multiple modal encoders. Thus, the pre-trained IMU encoder can be used to determine a quantitative representation of the model input, and this representation can be aligned with at least one modal representation.
[0028] In this disclosure, the pre-trained IMU encoder can be determined or generated by a pre-training process, which may also be referred to as a training process or an upstream task. The pre-trained IMU encoder can be further used for one or more downstream tasks, which may include some model training and / or model inference processes. For example, downstream tasks may include, but are not limited to, classifiers, data matching, data analyzers, etc.
[0029] In this disclosure, the term "representation" is used as the output of the encoder. In some examples, representation may also be used interchangeably with any of the following: latent representation, feature, embedding, vector representation, etc., and this disclosure is not limited to this aspect.
[0030] Figure 1 An exemplary schematic diagram of an environment 100 in which some exemplary embodiments of the present disclosure may be implemented is shown. As shown, environment 100 includes a pre-training phase 101 and a usage phase 102.
[0031] Pre-training phase 101 can be a pre-training process during which the IMU encoder 112 is updated based on the pre-training dataset 110. The pre-trained IMU encoder 122 can be generated or determined after pre-training phase 101. Pre-training phase 101 can also be referred to as a training phase for the IMU encoder 112, and this disclosure is not limiting in this respect.
[0032] Use phase 102 can use a pre-trained IMU encoder 122, for example, to determine model output 124 based on model input 120. Furthermore, model output 124 or the pre-trained IMU encoder 122 can be used for downstream tasks. Use phase 102 can also be referred to as an inference phase for the pre-trained IMU encoder 122, and this disclosure is not limiting in this respect.
[0033] It should be noted that the pre-training phase 101 and the usage phase 102 can be performed on the same device or different devices. For example, the pre-training phase 101 can be performed on a first device, while the usage phase 102 can be performed on a second device.
[0034] Figure 2A flowchart of a method 200 for using a pre-trained IMU encoder 122 according to some exemplary embodiments of the present disclosure is shown. For ease of description, it is assumed that method 200 is performed by an electronic device.
[0035] At box 210, the pre-trained IMU encoder 122 is obtained. At box 220, the model output 124 is determined by feeding the model input 120 into the pre-trained IMU encoder 122. At box 230, the model output 124 is provided for downstream tasks.
[0036] In some embodiments, the pre-trained IMU encoder 122 may be generated or determined by different devices, and therefore the electronic device may obtain (or receive or download) the pre-trained IMU encoder 122 from different devices. In some embodiments, the pre-trained IMU encoder 122 may be pre-loaded, generated, or determined by the electronic device, and therefore the electronic device may obtain (or access) the pre-trained IMU encoder 122 stored in the storage of the electronic device. For example, the pre-trained IMU encoder 122 may be determined based on the method 300 described below.
[0037] In some embodiments, the pre-trained IMU encoder 122 may include a stacked recurrent neural network (RNN) comprising at least one of the following: a convolutional layer, a group normalization layer, a max pooling layer, or a gated recurrent unit (GRU) layer. It should be noted that although a stacked RNN is used in the following embodiments, other model architectures may be used for the pre-trained IMU encoder 122; for example, a Transformer may be used for the pre-trained IMU encoder 122, and this disclosure is not limited to this aspect.
[0038] In some embodiments, the pre-trained IMU encoder 122 may have a multimodal head, and the output of the multimodal head is regarded as the model output 124.
[0039] In some embodiments, model input 120 may indicate data obtained from one or more IMU sensors. In some examples, the object may include or may be equipped with one or more IMU sensors. For example, the object may include, but is not limited to, one or more wearable devices, consumer electronics devices, vehicles, buildings, bridges, animals, etc. For ease of description, some embodiments in the following examples are provided with model input 120, which includes IMU data obtained from one or more IMU sensors of at least one wearable device associated with a user. In some examples, the IMU data is in a time-series format. For example, the IMU data includes multivariate IMU time-series data.
[0040] In some embodiments, model output 124 may be a quantitative representation of model input 120. In some examples, the quantitative representation may also be referred to as a latent representation, for example, in the form of a vector.
[0041] In some embodiments, the downstream task may be a model training task or a model inference task, or another task using the model output 124. In some examples, the downstream task may include a classifier for determining the validity of the model output 124. For example, the quality of a pre-trained IMU encoder 122 can be evaluated based on the validity of the model output 124 using a classifier. In some examples, the downstream task may include a matching module for determining the labels of the model input 120. For example, the model input and labels may form sample pairs that can be used to train another model. In some examples, the downstream task may include a training task that uses the model output 124 to determine state information of a user associated with at least one wearable device equipped with one or more IMU sensors. For example, the model input 120 may include IMU data obtained from one or more IMU sensors of at least one wearable device, and the model output 124 may be used to determine state information of a user associated with at least one wearable device (e.g., wearing at least one wearable device). Examples of at least one wearable device include, but are not limited to, wrist-worn devices, smart rings, smart clothing (including sensors), and head-mounted devices such as headsets, earpieces, in-ear headphones, smart glasses, or head-mounted devices such as head-mounted displays (HMDs).
[0042] As discussed, the model output 124 of the pre-trained IMU encoder 122 can be used for a variety of downstream tasks, so determining the pre-trained IMU encoder 122 is crucial.
[0043] Figure 3 A flowchart of a method 300 for pre-training an IMU encoder according to some exemplary embodiments of the present disclosure is shown. Method 300 may be performed by an electronic device, which may be the same as or different from the electronic device performing method 200.
[0044] At box 310, the multimodal dataset is obtained. At box 320, the pre-trained IMU encoder is determined based on the multimodal dataset.
[0045] Alternatively, pre-trained IMU encoders can be stored or provided for use.
[0046] In some embodiments, the multimodal dataset may be referred to as the pre-training dataset 110, and the multimodal dataset may include multiple segments. In some examples, the size of the multimodal dataset may be represented as an integer N, that is, the number of multiple segments in the multimodal dataset is represented as N.
[0047] In some examples, a fragment in a multimodal dataset may include IMU data samples and multiple modal data samples, where different modal data samples may be associated with different types. In some examples, a fragment may include IMU data samples, first modal data samples of a first type, and second modal data samples of a second type. For example, the first type may be text, and the first modal data sample of the first type may refer to a text sample. For example, the second type may be video, and the second modal data sample of the second type may refer to a video sample.
[0048] For example, IMU data samples can be obtained from head-mounted sensors, modal data samples of the first type can be obtained from free-form text annotations, and modal data samples of the second type can be obtained from first-person perspective video. It should be noted that multiple modal data samples may also include another modal data sample of another type (such as audio, image, etc.), and this disclosure is not limited to this aspect.
[0049] In this disclosure, a multimodal dataset can be represented as ,in( ) is any segment in the multimodal dataset (e.g., the i-th segment). In some examples, the segment ( A segment can be called a triplet, which consists of three time-aligned samples, for example, the i-th segment. These refer to time-aligned IMU data samples, video samples, and text samples, respectively.
[0050] To pre-train an IMU encoder, multiple modal encoders corresponding to multiple types can be used. For example, multiple modal encoders can be open-source pre-trained models developed by others. For example, multiple modal encoders can include video encoders and text encoders, which can be pre-trained on web-scale data, for example.
[0051] In this disclosure, a multi-objective representation learning strategy can be constructed, and the multi-objective representation learning strategy combines self-supervised (SS) loss and multimodal (MM) loss to pre-train the IMU encoder. Optionally, in some embodiments, nearest neighbor (NN) loss can be further combined.
[0052] Figure 4 An exemplary architecture 400 for pre-training an IMU encoder according to some exemplary embodiments of the present disclosure is shown. As shown, it uses a loss term consisting of three loss terms. and Multi-objective pre-training.
[0053] It should be noted that, although in Figure 4The nearest neighbor supervision 450 is shown, but this disclosure is not limiting in this respect. For example, in some embodiments, the nearest neighbor supervision 450 may be omitted, and the two items ( and ) is used for pre-training the IMU encoder.
[0054] In some examples, the IMU encoder during pre-training can be represented as A video encoder can be represented as And the text encoder can be represented as For example, an IMU encoder takes IMU data samples as input, a video encoder takes video samples as input, and a text encoder takes text samples as input.
[0055] Figure 5 An exemplary architecture 500 of an IMU encoder according to some exemplary embodiments of the present disclosure is shown. As shown, the backbone 510 of the architecture 500 includes a one-dimensional convolutional neural network (1D CNN) and GRU layers. During pre-training, the IMU encoder has two multilayer perceptron (MLP) heads, which include a multimodal head 502 and a unimodal head 504.
[0056] As an example, the IMU encoder can be a stacked RNN, including convolutional layers, group normalization layers, max pooling layers, and GRU layers, with a total of 1.4 million parameters. The unimodal header 504 is used to determine the unimodal self-supervised loss during pre-training, i.e. Figure 4 shown .
[0057] In some examples, the pre-trained IMU encoder can be used in method 200 after pre-training, where the unimodal head 504 is discarded. For example, the model input 120 can be IMU data 501, and therefore the model output 124 is based on the multimodal head 502. For example, the multimodal head 502 can provide a richer or more generalized latent representation than the unimodal head 504 provides.
[0058] In some examples, the motivation for architecture 500 is its deployment efficiency on devices such as mobile phones or wearable devices, which can be the target platform for collecting IMU data. In some examples, architecture 500 has demonstrated efficient generalization performance when processing ML tasks on IMU data. However, it should be understood that architecture 500 for IMU encoders is for illustrative purposes only and is not intended to be limiting; for example, Transformer can also be used for IMU encoders, and this disclosure does not limit this aspect.
[0059] In some embodiments, for pre-training the IMU encoder, a first loss can be determined based on a first IMU data sample in the multimodal dataset and data augmentation of the first IMU data sample (i.e., ).
[0060] In some examples, the first output can be determined by inputting a first IMU data sample into the IMU encoder. For example, if the first IMU data sample is represented as... Then the first output can be expressed as .
[0061] In some examples, the second output can be determined by inputting data augmentation of a first IMU data sample into the IMU encoder. For example, if the data augmentation of the first IMU data sample is represented as... Then the second output can be expressed as In some examples, a random transformation module can be defined for data augmentation. For example, the random transformation module may include transformations for scaling by a random factor and / or transformations for reversing the time direction. It should be noted that other modules may be used to determine data augmentation for the first IMU data; for example, augmentation features for adding Gaussian noise may be used, and this disclosure is not limited to this aspect.
[0062] In some examples, the first loss can be determined by the following formula, where It is a learnable temperature parameter:
[0063] Based on the first loss, the representation of the IMU data sample is encouraged to be similar to the data augmentation representation of the IMU data sample.
[0064] In some embodiments, in order to pre-train the IMU encoder, a second loss can be determined based on a first IMU data sample from the same segment and corresponding multiple modal data samples (i.e., In the case of a segment containing multiple modal data samples of different types, the second loss can correspond to a combination of multiple sub-losses of different types. For example, different sub-losses can be determined based on different modal data samples of different types.
[0065] As an example, for a fragment ( ), which includes the first IMU data sample First modal data samples of the first type and second modal data samples with second type The first output can be determined by inputting a first IMU data sample into the IMU encoder; for example, the first output can be represented as... The third output can be determined by inputting the first modality data sample into the corresponding encoder (a text encoder in this example). For example, the third output can be represented as... The fourth output can be determined by inputting samples of the second modality data into the corresponding encoder (in this example, a video encoder). For example, the fourth output can be represented as... .
[0066] In some examples, the first sub-loss corresponding to the first type (text) can be expressed as: And the second sub-loss corresponding to the second type (video) can be expressed as In some examples, the second loss can be determined based on the following formula, where It is a learnable temperature parameter:
[0067] Based on the second loss, the IMU encoder can learn semantic features that exist in rich modalities such as text and video.
[0068] In some embodiments, in order to pre-train the IMU encoder, a third loss can be determined based on a first segment in the feature queue and a target feature segment (i.e., In some examples, the feature queue can be represented as... ,in It is a cached representation of the IMU data, video, and text generated from its respective encoder. In some examples, the first segment can be represented as... And the target feature fragment can be represented as ,,in Determined by the following formula: (3) For example, the first segment may include video samples. And it can be done by using video samples The input is fed into the video encoder to output the video representation. For example, target feature fragments. It can indicate the representation of the target segment, which includes the target video, and the target video representation of the target video is... For example, video samples The similarity between the target video and the target video can be based on and Therefore, video representations are used to identify target segments (and thus target feature segments). Since the video encoder is pre-trained on a large dataset, the video representation can be a stable representation. Furthermore, video can capture far more detail than text.
[0069] Figure 6 An exemplary schematic diagram of nearest neighbor supervision 600 according to some exemplary embodiments of the present disclosure is shown. As shown, the first segment 610 may be a query segment and a target feature segment. It can be retrieved from the feature queue.
[0070] In some examples, target feature fragments This can indicate the representation of target segment 620. In some examples, it is based on video samples in the first segment 610 ( The similarity between the target feature segment and the target video 624 in the target segment 620 is used to determine the target feature segment.
[0071] In some examples, video-to-video similarity can be used to determine the target segment 620. For example, if video samples ( The similarity between the target video (624) and the video sample () is higher than that between the target video () and the video sample () The similarity between the target video 624 and any video segment 630 can be used to determine the target segment 620, which includes the target video 624. In some examples, the target segment 620 may be considered as the closest segment, nearest neighbor, most similar segment, etc.
[0072] In some examples, text-to-text similarity and / or IMU-to-IMU similarity can also be used to determine the target segment 620. For example, the similarity between the first segment 610 and the target segment 620 can be determined by some or all of the following: IMU data samples ( The first similarity between the video sample and the target IMU data 622, and the video sample ( The second similarity between the text sample and the target video 624, or the text sample ( The third similarity is calculated between the first segment 610 and the target segment 620. For example, the sum of the first, second, and third similarities can be considered as the similarity between the first segment 610 and the target segment 620. Similarly, the similarity between the first segment 610 and another segment 630 can be determined. If the similarity between the first segment 610 and the target segment 620 is higher than the similarity between the first segment 610 and another segment 630, then the target segment 620 can be identified.
[0073] In some examples, the target feature fragment may include multiple outputs determined by inputting corresponding data from the target fragment 620 into corresponding encoders (e.g., inputting target IMU data into an IMU encoder, inputting target video data into a video encoder, and inputting target text data into a text encoder).
[0074] In some examples, the first segment includes a first IMU data sample. And the first output can be determined as In some examples, the third loss can be determined on the first output and at least one feature in the target feature fragment (i.e., at least one output among multiple outputs of the data in the target fragment).
[0075] In some examples, the third loss can be determined by the following formula, where It is a learnable temperature parameter:
[0076] In some embodiments, the total loss may be a weighted sum of the first loss, the second loss, and the third loss, which may be determined by the following formula: (5) in , and It is the weight, and In some examples, the third loss can be omitted, for example... The total loss, also known as the final loss or multi-objective loss, can be used for pre-training the IMU encoder.
[0077] As mentioned above, learnable temperature parameters Used in formulas (1), (2), and (4), which can be updated during pre-training. In some examples, a learnable temperature parameter is included. The multiple trainable parameters can be updated via gradient descent based on the total loss, the details of which will not be discussed in this disclosure.
[0078] According to some embodiments of this disclosure, multiple learning objectives are incorporated during pre-training. A first loss ensures that the IMU encoder remains invariant to noise (similar to noise introduced by slight variations in sensor position or type). A second loss prompts the IMU encoder to generate representations oriented towards aligned text and video representations, thereby allowing the IMU encoder to learn rich semantic information present in the video and / or text modalities. A third loss uses the closest example from the feature queue as a positive pair, enabling the IMU encoder to leverage natural data similarity for more adaptive contrastive learning.
[0079] Given the promising applications of the synergistic relationship between self-supervised learning and multimodal learning in computer vision and natural language processing, and with the recent public availability of large multimodal datasets such as Ego-Exo4D, which include synchronized video, text, and IMU snippets, diverse information sources can be fully utilized, and the merging of multiple learning objectives can be explored for pre-training IMU encoders. In the case of analyzing IMU data from one or more IMU sensors in wearable devices, ubiquitous and effective health monitoring can be enabled.
[0080] In some exemplary embodiments, an apparatus (e.g., an electronic device) capable of performing method 200 may include components for performing the corresponding steps of method 200. The apparatus may be implemented in any suitable form. For example, the apparatus may be implemented in a circuit or software module.
[0081] In some exemplary embodiments, the apparatus includes: components for obtaining a pre-trained IMU encoder, wherein the pre-trained IMU encoder is determined based on a multimodal dataset together with multiple modal encoders; components for determining a model output by inputting a model input to the pre-trained IMU encoder, wherein the model input indicates data obtained from one or more IMU sensors; and components for providing the model output to a downstream task.
[0082] In some exemplary embodiments, the model input includes IMU data in time-series format, and the model output includes a quantitative representation of the IMU data.
[0083] In some exemplary embodiments, the model output is based on a multimodal head of a pre-trained IMU encoder.
[0084] In some exemplary embodiments, the apparatus (e.g., an electronic device) capable of performing method 300 may include components for performing the corresponding steps of method 300. The apparatus may be implemented in any suitable form. For example, the apparatus may be implemented in a circuit or software module.
[0085] In some exemplary embodiments, the apparatus includes: components for obtaining a multimodal dataset comprising multiple segments, wherein the segments are time-aligned and include IMU data samples and multiple modal data samples having multiple types; and components for determining a pre-trained IMU encoder based on the multimodal dataset together with multiple modal encoders corresponding to the multiple modalities.
[0086] In some exemplary embodiments, the apparatus includes: components for determining a first loss based on a first IMU data sample in a first segment and data augmentation of the first IMU data sample; components for determining a second loss based on the first IMU data sample and corresponding multiple modal data samples in the first segment; and components for determining a pre-trained IMU encoder by training based on the first loss and the second loss.
[0087] In some exemplary embodiments, the apparatus includes: components for determining a first output by inputting a first IMU data sample to an IMU encoder; components for determining a second output by inputting data enhancement of the first IMU data sample to an IMU encoder; and components for determining a first loss based on the first output and the second output.
[0088] In some exemplary embodiments, the apparatus includes: means for determining a first output by inputting a first IMU data sample to an IMU encoder; means for determining a third output by inputting a first modal data sample having a first type to a modal encoder corresponding to the first type, wherein the first modal data sample having the first type and the first IMU data sample are in a first segment; and means for determining a first sub-loss based on the first output and the third output.
[0089] In some exemplary embodiments, the apparatus includes: components for determining a target segment from a feature queue based on the similarity between a first segment and a target segment; components for determining a first output by inputting a first IMU data sample of the first segment to an IMU encoder; components for determining a plurality of outputs by inputting corresponding data from the target segment to a corresponding encoder; and components for determining a third loss based on the first output and at least one of the plurality of outputs.
[0090] In some exemplary embodiments, the data augmentation of the first IMU data sample is based on at least one of the following: scaling by a random factor; or reversing the time direction.
[0091] In some exemplary embodiments, the second loss is determined based on a plurality of sub-losses, wherein different sub-losses among the plurality of sub-losses are based on data samples with different modalities having different types.
[0092] In some exemplary embodiments, the pre-trained IMU encoder is further determined based on a third loss, which is determined based on the first segment and the target segment in the feature queue.
[0093] In some exemplary embodiments, the first segment includes a first IMU data sample and a second modal data sample having a second type, the target segment includes target IMU data and target modal data having a second type, and the similarity between the first segment and the target segment includes the similarity between the second modal data sample and the target modal data.
[0094] In some exemplary embodiments, the pre-trained IMU encoder is determined based on the total loss, which is a weighted sum of the first loss, the second loss, and the third loss.
[0095] As used in the specification and claims, the term "means" can refer to one or more individual elements configured to perform the corresponding recounted functions or functions, or it can refer to several elements performing such functions or functions. Furthermore, the functions recounted in the claims can be performed by the same individual means or a combination of the same means. For example, performing such functions or functions can be initiated in the device by a processor executing instructions stored in the device's memory.
[0096] Figure 7 A schematic block diagram of an exemplary device 700 for implementing embodiments of the present disclosure is shown. For example, the electronic device discussed above can be implemented by device 700. As shown, device 700 includes a central processing unit (CPU) 701 that performs various appropriate actions and processes based on computer program instructions stored in read-only memory (ROM) 702 or computer program instructions loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 stores various programs and data required for the operation of device 700. CPU 701, ROM 702, and RAM 703 are connected to each other via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.
[0097] The following components in device 700 are connected to I / O interface 705: input unit 706, such as a keyboard, mouse, etc.; output unit 707, including various displays and speakers, etc.; storage unit 708, including disks, optical disks, etc.; and communication unit 709, including network interface cards, modems, and wireless communication transceivers, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. It should be understood that, according to this disclosure, output unit 707 can be used to display real-time dynamic changes in customer satisfaction, key factor identification information for groups or individual users participating in satisfaction assessment, optimization strategy information, and strategies for achieving effect evaluation information, etc.
[0098] Processing unit 701 may be executed by one or more processing circuits. Processing unit 701 may be configured to perform various processes and procedures as described above. For example, in some embodiments, the processes described above may be implemented as computer software programs tangibly included in a machine-readable medium (e.g., storage unit 708). In some embodiments, part or all of the computer program may be loaded and / or installed onto device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more steps of the processes described above may be performed.
[0099] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to perform aspects of this disclosure.
[0100] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punch cards or raised structures in recesses on which instructions are recorded), and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0101] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or downloaded via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them for storage in a computer-readable storage medium within the suitable computing / processing device.
[0102] Computer-readable program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and traditional procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet through an Internet service provider). In some embodiments, electronic circuits including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions to personalize the electronic circuits in order to perform aspects of this disclosure by utilizing state information from the computer-readable program instructions.
[0103] This document describes aspects of the disclosure with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0104] These computer-readable program instructions may be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that, when executed via the processing unit of the computer or other programmable data processing apparatus, the instructions create parts for implementing the functions / actions specified in one or more boxes of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other apparatus to function in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of writing having instructions that implement aspects of the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0105] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other apparatus to cause a series of operational steps to be performed on the computer, other programmable apparatus or other apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other apparatus, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0106] Flowcharts and block diagrams illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code comprising one or more executable instructions for implementing one or more specified logical functions. In some alternative embodiments, the functions indicated in the blocks may occur in a non-consecutive order. For example, depending on the functions involved, two consecutive blocks may actually execute substantially simultaneously, or these blocks may sometimes execute in reverse order. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs the specified functions or actions.
[0107] Various embodiments of this disclosure have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to better explain the principles of the embodiments, practical applications of techniques found in the market, or technical improvements, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. An electronic device, comprising: At least one processor; as well as At least one memory storing instructions that, when executed by the at least one processor, cause the electronic device to at least: Obtain a pre-trained inertial measurement unit (IMU) encoder, wherein the pre-trained IMU encoder is determined based on a multimodal dataset and multiple modal encoders; The model input is fed into the pre-trained IMU encoder to determine the model output, wherein the model input indicates data obtained from one or more IMU sensors; as well as The model output is then provided to downstream tasks.
2. The electronic device of claim 1, wherein the model input comprises IMU data in time-series format, and wherein the model output comprises a quantitative representation of the IMU data; and The model output is based on the multimodal head of the pre-trained IMU encoder.
3. The electronic device according to claim 1 or 2, further configured as follows: Obtain the multimodal dataset comprising multiple segments, wherein the segments are time-aligned and include IMU data samples and multiple modal data samples, each having multiple types; and Based on the multimodal dataset, the pre-trained IMU encoder is determined together with the multiple modal encoders corresponding to the multiple modalities.
4. The electronic device according to claim 3, further configured as follows: The first loss is determined based on the first IMU data sample in the first segment and the data augmentation of the first IMU data sample; The second loss is determined based on the first IMU data sample and the corresponding multiple modal data samples in the first segment; as well as The pre-trained IMU encoder is determined by training based on the first loss and the second loss.
5. The electronic device according to claim 4, further configured to: The first output is determined by inputting the first IMU data sample into the IMU encoder; The second output is determined by inputting the data augmentation of the first IMU data sample into the IMU encoder; and The first loss is determined based on the first output and the second output.
6. The electronic device of claim 5, wherein the data enhancement of the first IMU data sample is based on at least one of the following: Scaling by a random factor; or Reverse the time direction.
7. The electronic device of claim 4, wherein the second loss is a combination of a plurality of sub-losses, wherein each sub-loss among the plurality of sub-losses is determined based on corresponding modal data samples of corresponding types.
8. The electronic device according to claim 7, further configured to: The first output is determined by inputting the first IMU data sample into the IMU encoder; A third output is determined by inputting a first modal data sample of a first type into a modal encoder corresponding to the first type, wherein the first modal data sample of the first type and the first IMU data sample are in the first segment; and Based on the first output and the third output, the first sub-loss among the plurality of sub-losses is determined.
9. The electronic device of claim 4, wherein the pre-trained IMU encoder is further determined based on a third loss, wherein the third loss is determined based on the first segment and the target segment in the feature queue.
10. The electronic device according to claim 9, further configured to: The target segment is determined from the feature queue based on the similarity between the first segment and the target segment; The first IMU data sample of the first segment is input into the IMU encoder to determine the first output; The corresponding data in the target segment is input into the corresponding encoder to determine multiple outputs; as well as The third loss is determined based on the first output and at least one of the plurality of outputs.
11. The electronic device of claim 10, wherein the first segment comprises the first IMU data sample and a second modal data sample having a second type, and the target segment comprises target IMU data and target modal data having the second type, and The similarity between the first segment and the target segment includes: The similarity between the second modal data sample and the target modal data.
12. The electronic device of claim 9, wherein the IMU encoder is trained based on a total loss that is a weighted sum of the first loss, the second loss, and the third loss.
13. The electronic device of claim 1, wherein the downstream task comprises at least one of the following: A classifier is used to determine the validity of the model's output. The matching module is used to determine the label of the model input, or The training task uses the model output to determine the state information of a user associated with at least one wearable device equipped with one or more of the IMU sensors.
14. The electronic device of claim 1, wherein during the pre-training phase, the IMU encoder has a single-modal head and a multi-modal head, and after the pre-training phase, the single-modal head is discarded.
15. An electronic device comprising: At least one processor; as well as At least one memory storing instructions that, when executed by the at least one processor, cause the electronic device to at least: Obtain a multimodal dataset comprising multiple segments, wherein the segments are time-aligned and include inertial measurement unit (IMU) data samples and multiple modal data samples of multiple types; as well as Based on the multimodal dataset, a pre-trained IMU encoder is determined together with multiple modal encoders corresponding to multiple modalities.
16. The electronic device according to claim 15, further configured to: The first loss is determined based on the first IMU data sample in the first segment and the data augmentation of the first IMU data sample; Based on the first IMU data sample and the corresponding multiple modal data samples in the first segment, a second loss is determined; as well as The pre-trained IMU encoder is determined by training based on the first loss and the second loss.
17. The electronic device of claim 16, wherein the pre-trained IMU encoder is further determined based on a third loss, wherein the third loss is determined based on the first segment and the target segment in the feature queue; and wherein the electronic device further comprises a component configured to: The target segment is determined from the feature queue based on the similarity between the first segment and the target segment; The first IMU data sample of the first segment is input into the IMU encoder to determine the first output; The corresponding data in the target segment is input into the corresponding encoder to determine multiple outputs; as well as The third loss is determined based on the first output and at least one of the plurality of outputs.
18. The electronic device of claim 17, wherein the first segment comprises the first IMU data sample and a second modal data sample having a second type, and the target segment comprises target IMU data and target modal data having the second type, and The similarity between the first segment and the target segment includes: The similarity between the second modal data sample and the target modal data.
19. The electronic device of claim 17, wherein the pre-trained IMU encoder is determined based on a total loss that is a weighted sum of the first loss, the second loss, and the third loss.
20. A method for communication, comprising: Obtain a multimodal dataset comprising multiple segments, wherein the segments are time-aligned and include inertial measurement unit (IMU) data samples and multiple modal data samples of multiple types; as well as Based on the multimodal dataset, a pre-trained IMU encoder is determined together with multiple modal encoders corresponding to multiple modalities.