Training method of signal separation model and wearing detection method, device and equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-08-11
AI Technical Summary
[0011]由上述实施例可知,本说明书通过获取训练样本,并将训练样本中的惯性信号输入待训练的信号分离模型,以使信号分离模型从惯性信号中分离出运动干扰特征和发声振动特征;进一步地,基于发声振动特征与运动干扰特征之间的第一偏差,以及发声振动特征与语音信号对应的语音特征之间的第二偏差,确定针对信号分离模型的损失值,并根据所述损失值对信号分离模型的模型参数进行调整。
Smart Images

Figure CN122551818A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to one or more embodiments in the field of terminal technology, and in particular to a training method, device, and equipment for a signal separation model and a wearing detection method. Background Technology
[0002] With the continuous development of smart wearable devices such as smart glasses, voice interaction, voice wake-up, and voice-based identity verification functions are widely used in scenarios such as payment authentication, privacy access control, and personalized service invocation.
[0003] Since such devices can typically access users' voice, biometrics, and other sensitive data, if the device can still trigger corresponding functions when not actually being worn, the relevant application data may be accessed, called, or even misused without effective user constraints, thereby affecting user account security, privacy security, and data usage security. Summary of the Invention
[0004] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a training method for a signal separation model is proposed, comprising: Acquire training samples, which include voice signals and inertial signals collected by the smart wearable device within the same time period; The inertial signal is input into the signal separation model to be trained, so that the motion interference features corresponding to the motion signal generated when the user performs body movements and the vocal vibration features corresponding to the vibration signal generated when the user speaks can be separated from the inertial signal through the signal separation model. Based on the first deviation between the vocal vibration feature and the motion interference feature, and the second deviation between the vocal vibration feature and the speech feature corresponding to the speech signal, the loss value for the signal separation model is determined; The model parameters of the signal separation model are adjusted based on the loss value.
[0005] According to a second aspect of one or more embodiments of this specification, a wear detection method is provided, comprising: Acquire the voice signals and inertial signals collected by the smart wearable device within the same time period; The inertial signal is input into a pre-trained signal separation model so that the sound vibration characteristics corresponding to the vibration signal generated when the user speaks can be separated from the inertial signal through the signal separation model. Extract the speech features corresponding to the speech signal, and determine whether the speech features and the vocal vibration features satisfy a preset correspondence; Based on the judgment result, it is determined whether the smart wearable device is in the state of being actually worn by the user.
[0006] According to a third aspect of one or more embodiments of this specification, a training apparatus for a signal separation model is provided, comprising: An acquisition module is used to acquire training samples, which include voice signals and inertial signals collected by the smart wearable device within the same time period; The separation module is used to input the inertial signal into the signal separation model to be trained, so as to separate the motion interference features corresponding to the motion signal generated when the user performs body movements and the vocal vibration features corresponding to the vibration signal generated when the user speaks from the inertial signal through the signal separation model. The determination module is used to determine the loss value for the signal separation model based on a first deviation between the vocal vibration feature and the motion interference feature, and a second deviation between the vocal vibration feature and the speech feature corresponding to the speech signal. An adjustment module is used to adjust the model parameters of the signal separation model based on the loss value.
[0007] According to a fourth aspect of one or more embodiments of this specification, a wear detection device is provided, comprising: The acquisition module is used to acquire the voice signals and inertial signals collected by the smart wearable device within the same time period; The separation module is used to input the inertial signal into a pre-trained signal separation model, so as to separate the vocal vibration features corresponding to the vibration signal generated when the user speaks from the inertial signal through the signal separation model. The judgment module is used to extract the speech features corresponding to the speech signal and determine whether the speech features and the vocal vibration features satisfy a preset correspondence. The detection module is used to determine, based on the judgment result, whether the smart wearable device is in the state of being actually worn by the user.
[0008] According to a fifth aspect of one or more embodiments of this specification, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor implements the steps of the method described above by executing the executable instructions.
[0009] According to a sixth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the method described above.
[0010] According to a seventh aspect of one or more embodiments of this specification, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the steps of the method described above.
[0011] As can be seen from the above embodiments, this specification obtains training samples and inputs the inertial signals in the training samples into the signal separation model to be trained, so that the signal separation model can separate motion interference features and vocal vibration features from the inertial signals; further, based on the first deviation between the vocal vibration features and the motion interference features, and the second deviation between the vocal vibration features and the corresponding speech features of the speech signal, the loss value for the signal separation model is determined, and the model parameters of the signal separation model are adjusted according to the loss value.
[0012] Therefore, the trained signal separation model can retain the vibration features related to vocalization activities while improving its ability to separate motion interference components generated by body movements, and enhance the correspondence between vocal vibration features and speech features. This provides a more accurate model basis for subsequent smart wearable device wearing detection based on the correspondence between speech signals and vocal vibration features, thereby improving the reliability and security of related wearing detection scenarios. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the architecture of a wear detection service system provided in an exemplary embodiment; Figure 2 This is a schematic flowchart illustrating a training method for a signal separation model provided in an exemplary embodiment; Figure 3 This is a flowchart illustrating the training process of a motion feature encoder, as provided in an exemplary embodiment. Figure 4 This is a flowchart illustrating the construction of a sample motion signal, provided in an exemplary embodiment. Figure 5 This is an exemplary embodiment of a flowchart for generating a three-dimensional human motion sequence; Figure 6 This is a schematic diagram of the training process of an action parameter generation model provided in an exemplary embodiment; Figure 7 This is a schematic diagram of feature constraints during the training process of a signal separation model, provided in an exemplary embodiment. Figure 8 This is a schematic flowchart of a wear detection method provided in an exemplary embodiment; Figure 9 This is an exemplary embodiment providing a flowchart of a wear detection process; Figure 10This is a schematic diagram of the structure of a device provided in an exemplary embodiment; Figure 11 This is a block diagram of a training apparatus for a signal separation model provided in an exemplary embodiment; Figure 12 This is a block diagram of a wear detection device provided in an exemplary embodiment. Detailed Implementation
[0014] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0015] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.
[0016] With the continuous development of smart wearable devices such as smart glasses, voice interaction, voice wake-up, and voice-based identity verification functions are widely used in scenarios such as payment authentication, privacy access control, and personalized service invocation. These devices typically collect, store, or access user-related voice information, identity information, and application data during actual use. When a device can trigger voice wake-up, voice control, or voice verification even when not actually worn, the relevant application data may be accessed, accessed, or even misused without effective user constraints, thereby affecting user account security, privacy security, and data usage security. Therefore, in the context of smart wearable devices, verifying the voice itself is usually insufficient to establish a complete security trust chain; further confirmation of whether the device is actually being worn by the user is also necessary.
[0017] Currently, the detection of wearing status for smart wearable devices typically employs capacitive sensors, contact sensors, infrared sensors, pressure sensors, or similar sensing devices to detect the contact state between the device and the user's skin, thereby determining whether the device is being worn. While this approach is relatively straightforward, its judgment relies heavily on physical contact or proximity, making it susceptible to manipulation by external objects. For example, in the context of smart glasses, using conductive materials, thin metal sheets, or other external media to simulate contact can mislead users into believing the device is being worn even when it is not actually being worn. This allows subsequent voice wake-up, voice control, or voice verification processes to be exploited through methods such as recording playback or remote injection, resulting in insufficient overall security.
[0018] To address the aforementioned issues, this specification provides a training method for a signal separation model. During model training, the model first learns to extract vibrational components related to sound production from inertial signals and suppresses motion interference components caused by body movements. This trained signal separation model provides a more stable foundation of sound vibrational signals for subsequent wear detection, enabling the device to better distinguish between genuine wear and non-wearing attack states even when the user is in dynamic usage scenarios such as speaking, turning their head, or walking, thereby improving the reliability of wear detection.
[0019] To facilitate understanding of the technical solutions in this specification, the key concepts involved are explained below: Smart wearable devices: Electronic devices that can be worn by a user and perform voice interaction, identity verification, or other application functions while worn. In this specification, a smart wearable device can be smart glasses or other head-mounted wearable devices, such as smart headphones, head-mounted interactive devices, etc.
[0020] Inertial signal: A timing signal acquired by inertial devices in a smart wearable device. This inertial signal contains at least two components: motion signals generated when the user performs physical actions and vibration signals generated when the user speaks. The inertial device may include an accelerometer, a gyroscope, or an inertial measurement unit (IMU) that includes both an accelerometer and a gyroscope. The inertial signal can be a one-dimensional signal or a multi-dimensional combined signal, such as a three-axis acceleration signal, a three-axis angular velocity signal, or mixed timing data formed by stitching together multiple inertial channels.
[0021] Motion signal: The interference component generated and reflected in the inertial signal when a user performs body movements. These body movements can include walking, turning the head, nodding, raising the head, lowering the head, body swaying, and other macroscopic movements. This signal typically has a large amplitude, long duration, and its variation pattern is related to human movements; it is a major source of noise affecting the accuracy of sound vibration extraction.
[0022] Vibration signal: When a user speaks, the vibration of the vocal organs is transmitted to the smart wearable device through head tissues, bones, or the device contact path, and is reflected in the target vibration component in the inertial signal. This signal is usually relatively weak, but it has a temporal correspondence with the speech activity and is an important basis for subsequent wear detection.
[0023] Voice signal: The acoustic signal collected by the smart wearable device during the user's speech. This voice signal can be acquired by a microphone, bone conduction pickup, array microphone, or other acquisition device. In this specification, the voice signal and the mixed inertial signal can be acquired simultaneously by different sensors, and they correspond to the same time period in the time dimension.
[0024] First bias: an index used to characterize the degree of difference between acoustic vibration characteristics and motion interference characteristics. The larger the first bias, the less overlap and the lower the correlation between acoustic vibration characteristics and motion interference characteristics.
[0025] Second bias: an index used to characterize the degree of difference or correspondence between vocal vibration features and speech features. The smaller the second bias, the higher the consistency between vocal vibration features and speech features.
[0026] Signal separation model: A model used to separate different components from inertial signals. This model can be implemented using neural networks or other learning models that include signal encoding and decoding structures.
[0027] The technical methods described in the various embodiments of this specification will be explained in detail below with reference to the accompanying drawings.
[0028] Figure 1 This is a schematic diagram of the architecture of a wear detection service system provided in an exemplary embodiment. Figure 1 As shown, the system may include a server 11, a network 12, a smart wearable device 13, a mobile phone 14, etc.
[0029] Server 11 can be a physical server containing an independent host, or it can be a virtual server hosted in a host cluster. During operation, server 11 can run server-side programs for a certain application to implement the relevant functions of that application. For example, when server 11 runs a model training service program, it can function as a corresponding model training service platform.
[0030] Wearable devices 13 and mobile phones 14 are just some of the types of electronic devices that users can use. In reality, users can obviously also use electronic devices such as tablets, laptops, and PDAs (Personal Digital Assistants), etc., and one or more embodiments in this specification do not limit this. During operation, the electronic device can run a client-side program of an application to achieve the relevant functions of that application. For example, when the electronic device runs a model training service program, it can act as a client for that model training service. The client application for the aforementioned model training service can be launched and run on the electronic device. This client-side program can be a native application installed on the electronic device, or it can be a mini-program, quick app, or other similar form. Of course, when using web technologies such as HTML5 or similar, the relevant functions can be achieved through a page displayed by a browser. This browser can be a standalone browser application or a browser module embedded in some applications.
[0031] As for the network 12 that enables interaction between electronic devices such as smart wearable devices 13 and mobile phones 14 and the server 11, the communication can be implemented using either wired or wireless networks based on the communication methods supported by the respective electronic devices. This specification does not impose any restrictions on this. For example, smart wearable devices 13 can support both wired and wireless communication, so they can use either wired or wireless networks as needed. Mobile phones 14 typically only support wireless communication, so they can use wireless networks for communication.
[0032] In this specification, the entity executing the training method for the signal separation model can be the aforementioned server, or a cloud training platform, edge computing node, electronic device, or a distributed training system composed of multiple computing nodes.
[0033] The entity executing the wear detection method can be the aforementioned smart wearable device, or a terminal device, edge processing device, or server that is communicatively connected to the smart wearable device. The acquisition of signals and detection of wear status related to wear detection can be performed by the smart wearable device. The smart wearable device can also send the collected voice and inertial signals to a paired mobile phone, tablet, or other terminal device, which will then perform the wear detection. Alternatively, the relevant data can be uploaded to a server, which will process the wear detection data and return the results to the smart wearable device or terminal device.
[0034] For ease of description, the following will use a server as the execution subject to illustrate the training method of the signal separation model; and use a smart wearable device as the execution subject to illustrate the wearing detection method.
[0035] Figure 2 This is a flowchart illustrating a training method for a signal separation model provided in an exemplary embodiment, including the following steps: S200: Acquire training samples, which include voice signals and inertial signals collected by smart wearable devices within the same time period.
[0036] The server can first acquire training samples for training the signal separation model. The speech and inertial signals in these training samples correspond to the same acquisition period in time, thus enabling the speech activity and inertial changes to be aligned. In this specification, "the same period" can refer to a time window acquired through strict synchronous acquisition, or it can refer to data segments that, after timestamp alignment, resampling, or interpolation, correspond to the same time axis.
[0037] In practical applications, training samples can be directly derived from real-world data collection. Specifically, test subjects can wear smart wearable devices to simultaneously collect speech and inertial signals while speaking still, walking, turning their heads, nodding, looking down, or in other action scenarios. To improve sample coverage, the server can obtain collection results with different speech content, different action patterns, different action amplitudes, and different usage postures as training samples.
[0038] Furthermore, the inertial signal in the training samples can also be constructed from the vocalization-related signal and the sample motion signal. For example, vocalization-related inertial data collected from the user in a relatively static or low-interference state can be acquired first, and then this vocalization-related inertial data can be superimposed with the sample motion signal to form an inertial signal containing both vocalization vibration components and motion interference components. This allows for the flexible construction of training samples under different combinations of motion intensity and speaking states.
[0039] Optionally, to make subsequent training more stable, the server can also preprocess the speech and inertial signals after acquiring the training samples. This preprocessing can include time synchronization, normalization, DC component removal, filtering, window segmentation, channel splicing, resampling, or noise suppression.
[0040] For example, the server can first align the speech signal and inertial signal to the same sampling window based on timestamps, and then stitch together the multi-axis inertial channels to form a unified input format. Another example is that the energy envelope or temporal intensity sequence can be extracted from the speech signal first, in order to perform subsequent timing bias calculations.
[0041] S202: Input the inertial signal into the signal separation model to be trained, so as to separate the motion interference features corresponding to the motion signal generated when the user performs body movements and the vocal vibration features corresponding to the vibration signal generated when the user speaks from the inertial signal through the signal separation model.
[0042] After obtaining the training samples, the server can input the inertial signals in the training samples into the signal separation model to be trained, so as to extract motion interference features and acoustic vibration features from the inertial signals through the signal separation model.
[0043] Specifically, the signal separation model may include a comprehensive encoder and a motion feature encoder, and optionally, a decoder. Accordingly, separating motion interference features and acoustic vibration features from an inertial signal may include the following process: First, the inertial signal is encoded using the comprehensive encoder to obtain comprehensive features; simultaneously, the inertial signal is encoded using the motion feature encoder to obtain motion interference features; subsequently, the comprehensive features are separated based on the motion interference features to obtain the acoustic vibration features.
[0044] The integrated encoder extracts the overall information representation from the inertial signal, which includes both motion-related and sound-related features. The motion feature encoder extracts inertial features related to macroscopic human movements, such as walking rhythm, head swaying patterns, and head turning inertial change patterns. After obtaining the integrated features and motion interference features, the server can separate the integrated features based on the motion interference features, weakening the motion interference-related components while retaining the sound-related components, thus obtaining the sound vibration features.
[0045] Optionally, the signal separation model may also include a decoder, which the server can use to decode motion interference features and / or acoustic vibration features to obtain the corresponding reconstructed signal for subsequent signal-level constraints, result verification, or visualization analysis.
[0046] Alternatively, the separation process can be implemented in various ways. For example, it can weaken the parts of the composite features that overlap with motion interference features based on feature cancellation; it can generate suppression weights based on gating mechanisms to suppress the motion components in the composite features; or it can preserve the feature subspace that is relatively independent of the motion interference features based on projection separation. The specific method used can be determined according to the model structure design.
[0047] Furthermore, inertial signals can be one-dimensional or multi-dimensional inputs. When the inertial signal includes triaxial acceleration signals, triaxial angular velocity signals, or a combination of both, the integrated encoder and motion feature encoder can process data from different axes separately, or they can fuse the multi-channel data first and then process it. For example, the triaxial acceleration signals can be directly concatenated and input into the encoder, or sub-encoding structures can be set for different inertial channels, and then feature fusion can be performed.
[0048] Furthermore, since the motion feature encoder is responsible for extracting motion disturbance features, the server can also perform pre-training on the motion feature encoder first.
[0049] Figure 3 This is a flowchart illustrating the training process of a motion feature encoder, as provided in an exemplary embodiment.
[0050] The server can acquire sample motion signals, input them into a motion feature encoder, and obtain the corresponding motion interference features. Then, based on the motion interference features, the decoder or motion decoder... The reconstructed motion signal is then used to construct a reconstruction loss based on the reconstruction deviation between the reconstructed motion signal and the sample motion signal. The model parameters of the motion feature encoder are adjusted with the optimization objective of minimizing this loss value.
[0051] The sample motion signals can originate from real-world data acquisition. For example, test subjects can wear smart wearable devices and perform actions such as walking, turning their heads, nodding, shaking their heads, raising their heads, lowering their heads, or other movements without making any noise, thereby collecting the corresponding sample motion signals. The sample motion signals obtained in this way more closely resemble the hardware characteristics of real devices.
[0052] Furthermore, sample motion signals can be constructed through simulation.
[0053] Figure 4 This is an exemplary embodiment of the flowchart for constructing a sample motion signal.
[0054] The server first acquires a 3D human motion sequence, and then, based on the preset wearing position of the smart wearable device on the human body, simulates the motion signals of the human body movements corresponding to the 3D human motion sequence to obtain sample motion signals generated by the smart wearable device during the execution of human movements. The preset wearing position can be a fixed location set in the simulation environment, such as the preset installation position of smart glasses on the head.
[0055] The server can first obtain action description information; then input the action description information into a pre-trained action parameter generation model to obtain motion sequence parameters corresponding to the action description information; then input the motion sequence parameters into a pre-trained visual action generation model to map the motion sequence parameters to a three-dimensional pose space to obtain a three-dimensional human motion sequence.
[0056] For example, action descriptions can be natural language descriptions such as "stand up from a sitting position and walk five steps," "slowly turn your head to the left and then look straight ahead," or "look down while reading and then look up to speak." The action parameter generation model is used to convert these descriptions into human motion control parameters, while the visual action generation model is used to further output continuous three-dimensional posture change results.
[0057] Optionally, the process of generating a three-dimensional human motion sequence from motion description information can also be completed directly by a single model, that is, the motion description information is directly input into a single generation model and the three-dimensional human motion sequence is output.
[0058] Figure 5 This is an exemplary embodiment of a flowchart for generating a three-dimensional human motion sequence.
[0059] The server can input action description information into the interface call layer; this interface call layer can be a large language model interface, used to parse or transform the action description information. Then, the processed action description information is input into a text encoder to extract text features corresponding to the action description information, and further, these text features are input into a motion decoder, which generates the corresponding 3D human motion sequence.
[0060] Figure 6 This is a schematic diagram of the training process of an action parameter generation model provided in an exemplary embodiment.
[0061] The server can input sample human action sequences into the action encoder to obtain corresponding action feature representations. Then, the action features are represented by the action decoder. Decoding is performed to obtain the reconstructed human action sequence, and a reconstruction loss is constructed based on the reconstruction deviation between the reconstructed human action sequence and the sample human action sequence. On the other hand, the server can also obtain action description information corresponding to the human action sequence of the sample, and input the action description information into the text encoder to obtain the corresponding text feature representation. Subsequently, based on action feature representation Text feature representation The feature deviations between them are used to construct the text alignment loss. Ultimately, the server can base its decision on the reconstruction loss. and text alignment loss The model parameters of the motion parameter generation model are adjusted so that the trained motion parameter generation model can output the corresponding motion sequence parameters based on the motion description information.
[0062] S204: Determine the loss value for the signal separation model based on the first deviation between the vocal vibration features and the motion interference features, and the second deviation between the vocal vibration features and the speech features corresponding to the speech signal.
[0063] After obtaining motion interference features and vocal vibration features, the server can also extract corresponding speech features based on the speech signals in the training samples, and determine the loss value for the signal separation model based on the first deviation between the vocal vibration features and motion interference features, and the second deviation between the vocal vibration features and speech features.
[0064] This step allows the model training to be subject to two types of constraints simultaneously: one type of constraint is used to enhance the separability between motion disturbance features and vocal vibration features, and the other type of constraint is used to enhance the correspondence between vocal vibration features and speech features.
[0065] The first bias reflects the degree of difference between the vocal vibration features and the motion interference features. To ensure that the extracted vocal vibration features contain as few motion interference components as possible, the server can constrain model training by increasing the degree of difference between the vocal vibration features and the motion interference features. The first bias can be determined based on distance metrics, correlation metrics, mutual information, cosine similarity, or other indicators that can reflect the degree of feature difference in the feature space.
[0066] The second deviation is used to reflect the degree of difference or correspondence between the vocal vibration features and the speech features. The second deviation can be determined based on at least one of the following: the deviation between the start and end times of vocalization represented by the speech features and the vocal vibration features; the deviation between the intensity change trend represented by the speech features and the vocal vibration features; and the deviation between the peak occurrence times represented by the speech features and the vocal vibration features. For the above multiple indicators, the server can choose any one as the second deviation, or it can weight and fuse multiple indicators to obtain a comprehensive second deviation.
[0067] In this specification, the loss value is negatively correlated with the first bias and positively correlated with the second bias. The server can combine the first bias and the second bias to construct a joint loss function to achieve the training objective of increasing the first bias and decreasing the second bias.
[0068] For example, speech features can be used to characterize the start and end times of speech activities, intensity change trends, and peak occurrence times, while vocal vibration features can also be used to characterize the start and end times of corresponding vibration activities, intensity change trends, and peak occurrence times.
[0069] The server can determine the time difference between the start and end times of the speech activity represented by the speech features and the start and end times of the vibration activity represented by the vocal vibration features, and use the time difference as part of the second deviation.
[0070] The server can also determine the trend difference between the speech activity intensity change trend represented by the speech features and the vibration activity intensity change trend represented by the vocal vibration features, and use the trend difference as part of the second deviation.
[0071] The server can also determine the peak time offset between the peak occurrence time of speech activity characterized by speech features and the peak occurrence time of vibration activity characterized by vocal vibration features, and use the peak time offset as part of the second deviation.
[0072] S206: Adjust the model parameters of the signal separation model based on the loss value.
[0073] After determining the loss value, the server can adjust the model parameters of the signal separation model based on the loss value, thereby completing the model training.
[0074] Specifically, the server can adjust the model parameters of the signal separation model with the optimization objective of minimizing the loss value. This parameter adjustment can be achieved using gradient descent, stochastic gradient descent, adaptive learning rate optimization, or other commonly used optimization algorithms.
[0075] In the specific training process, the server can employ either batch training or mini-batch iterative training. For example, multiple training samples can be divided into several batches, with forward separation, bias calculation, loss determination, and backpropagation updates performed on each batch. Alternatively, the motion feature encoder can be pre-trained first, and then incorporated into the complete signal separation model for joint training with the integrated encoder and decoder. The motion feature encoder and the subsequent overall separation network can be components of the same overall model, with the pre-training results of the former serving as the basis for the overall training of the latter.
[0076] Optionally, the server can also reconstruct the corresponding specific signal based on the acoustic vibration characteristics and / or motion interference characteristics using a decoder, and determine the loss value according to the degree of difference between the reconstructed specific signal and the actual signal, or use the signal difference as an auxiliary constraint, together with the loss value obtained based on feature deviation, for model optimization. In this way, the separation results can be further constrained at the signal level, enhancing the interpretability of the training process.
[0077] Figure 7This is a schematic diagram of feature constraints during the training process of a signal separation model, provided as an exemplary embodiment.
[0078] Among them, after separating the inertial signal input signal into a model, motion disturbance characteristics can be obtained. and acoustic vibration characteristics After the speech signal is input into the speech encoder, speech features can be obtained. Based on this, the server can determine the motion interference characteristics. With acoustic vibration characteristics The deviation between them determines the motion separation loss used to characterize the degree of separation. Based on the characteristics of sound vibration With speech features The deviation between them determines the micro-alignment loss. Subsequently, based on motion separation loss Alignment loss with micro-motion Determine the target loss value The target loss value can be expressed as:
[0079] in, Preset weights are used to balance motion separation loss. Alignment loss with micro-motion At the target loss value The relative influence of the two is considered so that the model can improve the correspondence between vocal vibration features and speech features while reducing the similarity between vocal vibration features and motion interference features.
[0080] After training is complete, the server can save the adjusted model parameters to form a trained signal separation model, which can then be used by subsequent wear detection methods.
[0081] In some embodiments, considering the limitations typically imposed on smart wearable devices in terms of processor computing power, storage capacity, and power consumption, the server can perform lightweight processing on the trained signal separation model after training to meet the local deployment requirements of smart wearable devices. This lightweight processing may include at least one of the following: model compression, model pruning, parameter quantization, knowledge distillation, low-rank decomposition, or other methods to reduce model computational complexity and storage overhead.
[0082] For example, the server can prune redundant connections, channels, or parameters in the signal separation model based on the importance of parameters, channels, or feature response strength of each network layer, in order to reduce the model size while preserving the ability to separate acoustic vibration features as much as possible. Another example is that the server can convert floating-point parameters into low-bit parameter representations to reduce model storage and inference computation overhead. Yet another example is that the server can use a large-scale trained model as a teacher model and train a lightweight student model through knowledge distillation, allowing the lightweight model to inherit the teacher model's ability to separate motion interference features and acoustic vibration features.
[0083] Furthermore, this specification also provides a wearing detection method applied to the above signal separation model, such as... Figure 8 As shown.
[0084] Figure 8 This is a schematic flowchart of a wear detection method provided in an exemplary embodiment, including the following steps: S800: Acquires voice signals and inertial signals collected by smart wearable devices within the same time period.
[0085] Before performing sensitive operations such as voice wake-up, voice control, or payment authentication, a smart wearable device can first acquire voice and inertial signals collected by the device within the same time period. This "same time period" can be a real-time acquisition window or a corresponding segment within a sliding detection window. The voice and inertial signals can be acquired synchronously by different sensors; for example, the voice signal can be acquired by a microphone, and the inertial signal by an accelerometer or inertial measurement unit.
[0086] In real-time detection scenarios, smart wearable devices can extract the closest voice signal and inertial signal to the current moment as the current detection input. Based on the signal changes near the current moment, they can quickly determine whether the smart wearable device is actually being worn by the user, thus meeting the application requirements with high response latency, such as voice wake-up and payment authentication.
[0087] In continuous monitoring scenarios, smart wearable devices can also acquire voice signals and inertial signals from multiple past time periods, thereby making joint judgments based on historical signal changes over multiple time periods. This reduces the impact of signal fluctuations, instantaneous noise interference, or short-term abnormal movements in a single time period on the detection results and improves the stability of the wear detection results.
[0088] S802: Input the inertial signal into a pre-trained signal separation model, so as to separate the vocal vibration characteristics corresponding to the vibration signal generated when the user speaks from the inertial signal through the signal separation model; wherein, the signal separation model is trained by the aforementioned training method.
[0089] After acquiring voice and inertial signals, the smart wearable device can input the inertial signal into a pre-trained signal separation model. After separation, the smart wearable device can obtain the vocal vibration characteristics, which are used for subsequent wearing status determination.
[0090] In one embodiment, the smart wearable device can run a pre-trained signal separation model locally. This reduces the transmission of raw signals to external systems, lowers the risk of privacy breaches, and improves real-time performance. In another embodiment, the smart wearable device can also send the collected inertial signals to a terminal device or server connected to it, whereby the external device runs the signal separation model and then returns the separation results to the smart wearable device.
[0091] S804: Extract the speech features corresponding to the speech signal, and determine whether the speech features and the vocal vibration features satisfy a preset correspondence.
[0092] After obtaining the vocal vibration characteristics, the smart wearable device can further extract the speech features corresponding to the speech signal and determine whether the speech features and vocal vibration characteristics satisfy a preset correspondence. This preset correspondence can be consistent with the construction logic of the second deviation in the training phase, or it can adopt a feature correspondence judgment method that is consistent with or similar to that in the training phase.
[0093] Optionally, the smart wearable device can make a judgment from at least one of the following dimensions: whether the start and end times of speech represented by the speech features and the vocal vibration features correspond; whether the intensity change trend represented by the speech features and the vocal vibration features correspond; and whether the peak occurrence time represented by the speech features and the vocal vibration features correspond.
[0094] In practical applications, this judgment process can be implemented in various ways. For example, the correspondence between speech features and vocal vibration features can be determined based on the correlation coefficient; the temporal alignment between the two can be calculated based on dynamic time warping; and the overall consistency result can be output based on a weighted average of multiple scores.
[0095] For example, smart wearable devices can determine whether at least one of the following—start and end times, intensity change trends, and peak occurrence times—satisfies a preset correspondence based on the information represented by voice features and vocal vibration features. If the time difference between the start and end times of the voice activity represented by the voice features and the start and end times of the vibration activity represented by the vocal vibration features is less than a preset threshold, then the two are determined to satisfy a preset correspondence in the dimension of start and end times; if the difference between the intensity change trend of the voice activity represented by the voice features and the intensity change trend of the vibration activity represented by the vocal vibration features is less than a preset threshold, then the two are determined to satisfy a preset correspondence in the dimension of intensity change trends; if the time offset between the peak occurrence time of the voice activity represented by the voice features and the peak occurrence time of the vibration activity represented by the vocal vibration features is less than a preset threshold, then the two are determined to satisfy a preset correspondence in the dimension of peak occurrence times.
[0096] In some embodiments, before determining whether the voice features and vocal vibration features satisfy a preset correspondence, the smart wearable device can also determine the current business scenario and obtain a deviation threshold that matches the current business scenario. Then, the smart wearable device can compare the deviation result determined based on the voice features and vocal vibration features with the deviation threshold; when the deviation result is less than the deviation threshold, it is determined that the voice features and vocal vibration features satisfy the preset correspondence; otherwise, it is determined that the voice features and vocal vibration features do not satisfy the preset correspondence.
[0097] S806: Based on the judgment result, determine whether the smart wearable device is actually being worn by the user.
[0098] After completing the correspondence determination, the smart wearable device can determine whether it is actually being worn by the user based on the determination result.
[0099] Figure 9 This is an exemplary embodiment of a wear detection flowchart.
[0100] Among them, after separating the inertial signal input signal collected by the smart wearable device into a model, the acoustic vibration characteristics can be obtained. After the speech signal is input into the speech encoder, speech features can be obtained. Then, the consistency between the vocal vibration characteristics and the voice characteristics is judged to determine whether the two meet the preset correspondence, and based on this, it is determined whether the smart wearable device is actually being worn.
[0101] In one embodiment, if the determination result shows that the voice features and the vocal vibration features satisfy a preset correspondence, then it is determined that the smart wearable device is in the state of being actually worn by the user.
[0102] Accordingly, if the judgment result shows that the voice features and the vocal vibration features do not meet the preset correspondence, it is determined that the smart wearable device is not in the state of being actually worn by the user, or at least the current state does not meet the requirements for bio-level wear detection.
[0103] For example, when recording audio aloud without wearing the device, although the device may collect the voice signal and extract the corresponding voice features, it is usually difficult to extract the vocal vibration features that satisfy the preset correspondence with the voice features from the inertial signal because the device is not actually worn. Furthermore, even if an attacker simulates the wearing state using a head model, bracket, or other carrier, it is still difficult to construct the actual vocal vibration features corresponding to the voice activity, thus making it difficult to pass the correspondence verification.
[0104] After determining the wearing status, smart wearable devices can also execute corresponding subsequent control policies. For example, when it is determined that the device is actually being worn by the user, subsequent payment authentication, privacy access control, or other sensitive functions can be allowed; when it is determined that the device is not actually being worn, related operations can be refused, secondary verification can be triggered, or security logs can be recorded.
[0105] Optionally, when a smart wearable device determines that its current state has failed the wearing detection, it can send a detection failure message to the terminal device or server it communicates with, indicating that the current operation poses a risk of being triggered by improper wearing. Correspondingly, upon receiving this detection failure message, the terminal device or server can output a risk warning message to the user, indicating that the device may not be in an actual wearing state and that the relevant function poses a security risk. This risk warning message can be output through pop-up notifications, message notifications, voice prompts, or other prompting methods.
[0106] Furthermore, considering the potential for misjudgments in practical applications due to environmental noise, device vibration, momentary signal loss, or fluctuations in model judgment, smart wearable devices can also perform further fallback processing based on multiple consecutive test results or the duration of the failed state. Specifically, when the number of consecutive failed tests reaches a preset number, or the duration of the failed wearing test reaches a preset duration, the smart wearable device can send the corresponding abnormal detection information to the terminal device or server to trigger secondary detection processing.
[0107] The secondary detection process can include at least one of the following: the terminal device or server re-analyzes the voice and inertial signals based on a higher-precision detection strategy; the terminal device or server calls a more complex signal separation or discrimination model for verification; a re-acquisition request is sent to the user to re-acquire the voice and inertial signals; or other authentication factors are combined to assist in the judgment of the current state. This fallback mechanism can maintain the security of wear detection while reducing the impact of occasional misjudgments on normal user operation.
[0108] To facilitate understanding, several specific application examples are given below.
[0109] In one example, a user wears smart glasses and initiates payment authentication. The smart glasses simultaneously acquire voice and inertial signals when the user speaks the authentication command. Because the user makes a slight head turn while speaking, the inertial signal includes both vibrational components related to vocalization and motion-related components from the head turn. The smart glasses input the inertial signal into a pre-trained signal separation model to obtain vocal vibration features; simultaneously, it extracts corresponding voice features from the voice signal. Then, the smart glasses determine whether the voice features and vocal vibration features satisfy a preset correspondence in dimensions such as the start and end times of vocalization and the trend of intensity changes. If they satisfy the correlation, it is determined that the device is actually being worn by the user, and the payment authentication process continues.
[0110] In another example, although smart glasses can collect speech signals from external recordings and extract corresponding speech features, since the device is not actually worn, it is usually difficult to extract the vocal vibration features that satisfy a preset correspondence with the speech features from the inertial signals. After signal separation and correspondence determination, the smart glasses can determine that the current speech features and vocal vibration features do not satisfy a preset correspondence, thereby determining that the device is not actually worn by the user and refusing to perform related operations.
[0111] Figure 10 This is a schematic structural diagram of a device provided in an exemplary embodiment. For example... Figure 10As shown, device 1000 mainly consists of a communication interface 1002, a user interface 1004, a processor 1006, and a data storage 1008. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1010. The communication interface 1002 enables device 1000 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1002 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1002 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1002 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1002 may also include multiple physical communication interfaces, such as Wi-Fi interfaces, Bluetooth interfaces, and wide-area wireless interfaces.
[0112] User interface 1004 includes receiving user input and providing output to the user. Therefore, user interface 1004 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1004 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1004 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, device 1000 may support remote access from other devices via communication interface 1002 or another physical interface (not shown). User interface 1004 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1004 can also be configured as a display device for rendering or displaying text fragments.
[0113] Processor 1006 may contain one or more general-purpose processors and / or special-purpose processors.
[0114] Data storage 1008 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1006. Data storage 1008 may include removable and non-removable components.
[0115] Processor 1006 is capable of executing program instructions 1018 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 1008 to perform the various functions described herein. Data storage 1008 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1000, enable device 1000 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1018 by processor 1006 may result in processor 1006 using data 1012.
[0116] For example, program instructions 1018 may include an operating system 1022 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1000 and one or more applications 1020 (e.g., a browser, social media application, or game application). Similarly, data 1012 may include operating system data 1016 and application data 1014. Operating system data 1016 is primarily accessible to the operating system 1022, while application data 1014 is primarily accessible to one or more applications 1020. Application data 1014 may reside in a file system visible or hidden from the user of device 1000.
[0117] Application 1020 can communicate with operating system 1022 through one or more application programming interfaces (APIs). These APIs help application 1020 read and / or write application data 1014, transmit or receive information via communication interface 1002, receive or display information on user interface 1004, etc.
[0118] In some terminology, application 1020 may be simply referred to as "app". Furthermore, application 1020 can be downloaded to device 1000 through one or more online app stores or app markets. However, applications can also be installed on device 1000 in other ways, such as through a web browser or a physical interface on device 1000 (e.g., a USB port).
[0119] Please refer to Figure 11 The training device for the signal separation model can be applied to, for example... Figure 10 The device shown is used to implement the technical solution of this specification. The training apparatus for the signal separation model may include: The acquisition module 1100 is used to acquire training samples, which include voice signals and inertial signals collected by the smart wearable device in the same time period. The separation module 1102 is used to input the inertial signal into the signal separation model to be trained, so as to separate the motion interference features corresponding to the motion signal generated when the user performs body movements and the vocal vibration features corresponding to the vibration signal generated when the user speaks from the inertial signal through the signal separation model. The determination module 1104 is used to determine the loss value for the signal separation model based on the first deviation between the vocal vibration feature and the motion interference feature, and the second deviation between the vocal vibration feature and the speech feature corresponding to the speech signal. The adjustment module 1106 is used to adjust the model parameters of the signal separation model according to the loss value.
[0120] Optionally, the signal separation model includes: a comprehensive encoder and a motion feature encoder; The separation module 1102 is specifically used to encode the inertial signal using the integrated encoder to obtain integrated features; and to encode the inertial signal using the motion feature encoder to obtain motion interference features; and to separate the integrated features based on the motion interference features to obtain the acoustic vibration features.
[0121] Optionally, the loss value is negatively correlated with the first deviation and positively correlated with the second deviation; The adjustment module 1106 is specifically used to adjust the model parameters of the signal separation model with the optimization objective of minimizing the loss value.
[0122] Optionally, the signal separation model includes a motion feature encoder, and the device further includes: The pre-training module 1108 is used to acquire sample motion signals; input the sample motion signals into the motion feature encoder to obtain corresponding motion interference features; and adjust the model parameters of the motion feature encoder according to the reconstruction deviation between the sample motion signals and the motion signals reconstructed based on the motion interference features.
[0123] Optionally, the pre-training module 1108 is specifically used to acquire a three-dimensional human motion sequence; based on the preset wearing position of the smart wearable device on the human body, to simulate the human motion corresponding to the three-dimensional human motion sequence, and to obtain the sample motion signal generated by the smart wearable device during the execution of the human motion.
[0124] Optionally, the pre-training module 1108 is specifically used to: acquire motion description information; input the motion description information into a pre-trained motion parameter generation model to obtain motion sequence parameters corresponding to the motion description information; and input the motion sequence parameters into a pre-trained visual motion generation model to map the motion sequence parameters to a three-dimensional pose space to obtain the three-dimensional human motion sequence.
[0125] Optionally, the second deviation is determined based on at least one of the following: The deviation between the speech features and the vocal vibration features representing the start and end times of vocalization; The deviation between the intensity change trend represented by the speech features and the vocal vibration features; The deviation between the peak occurrence time represented by the speech feature and the vocal vibration feature.
[0126] Please refer to Figure 12 Wearable detection devices can be applied to, for example Figure 9 The device shown implements the technical solution described in this specification. The wear detection device may include: The acquisition module 1200 is used to acquire the voice signals and inertial signals collected by the smart wearable device within the same time period; The separation module 1202 is used to input the inertial signal into a pre-trained signal separation model, so as to separate the motion interference features corresponding to the motion signal generated when the user performs body movements and the vocal vibration features corresponding to the vibration signal generated when the user speaks from the inertial signal through the signal separation model. The judgment module 1204 is used to extract the speech features corresponding to the speech signal and determine whether the speech features and the vocal vibration features satisfy a preset correspondence. The detection module 1206 is used to determine whether the smart wearable device is in the state of being actually worn by the user based on the judgment result.
[0127] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0128] Based on the same concept as the methods described above, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor performs the steps of the method as described in any of the above embodiments by executing the executable instructions.
[0129] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0130] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0131] What those skilled in the art will understand is: In this specification, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.
[0132] In this specification, “a,” “an,” and “the” do not specifically refer to the singular, but may also include the plural.
[0133] In this specification, ordinal numbers such as "first," "second," etc., do not necessarily indicate order; they are often used to distinguish between objects. For example, "first server" and "second server" usually refer to two servers. To differentiate between these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.
[0134] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.
[0135] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.
[0136] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0137] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.
Claims
1. A training method for a signal separation model, comprising: Acquire training samples, which include voice signals and inertial signals collected by the smart wearable device within the same time period; The inertial signal is input into the signal separation model to be trained, so that the motion interference features corresponding to the motion signal generated when the user performs body movements and the vocal vibration features corresponding to the vibration signal generated when the user speaks can be separated from the inertial signal through the signal separation model. Based on the first deviation between the vocal vibration feature and the motion interference feature, and the second deviation between the vocal vibration feature and the speech feature corresponding to the speech signal, the loss value for the signal separation model is determined; The model parameters of the signal separation model are adjusted based on the loss value.
2. The method of claim 1, the signal separation model comprising: Integrated encoder and motion feature encoder; Separating motion interference features corresponding to the motion signals generated when the user performs body movements and vocal vibration features corresponding to the vibration signals generated when the user speaks from the inertial signal includes: The inertial signal is encoded by the integrated encoder to obtain integrated features; and The motion feature encoder encodes the inertial signal to obtain motion interference features; The comprehensive features are separated based on the motion interference features to obtain the acoustic vibration features.
3. The method as described in claim 1, wherein the loss value is negatively correlated with the first deviation and positively correlated with the second deviation; Adjusting the model parameters of the signal separation model based on the loss value includes: The model parameters of the signal separation model are adjusted with the goal of minimizing the loss value.
4. The method of claim 3, wherein the signal separation model includes a motion feature encoder, and the method further includes: Acquire sample motion signals; The sample motion signal is input into the motion feature encoder to obtain the corresponding motion interference features; The model parameters of the motion feature encoder are adjusted based on the reconstruction deviation between the sample motion signal and the motion signal reconstructed based on the motion interference features.
5. The method as described in claim 4, wherein acquiring the sample motion signal specifically includes: Obtain three-dimensional human motion sequences; Based on the preset wearing position of the smart wearable device on the human body, motion signal simulation is performed on the human body movements corresponding to the three-dimensional human body motion sequence to obtain sample motion signals generated by the smart wearable device during the execution of the human body movements.
6. The method as described in claim 5, specifically comprising: obtaining a three-dimensional human motion sequence, including: Obtain action description information; The motion description information is input into a pre-trained motion parameter generation model to obtain motion sequence parameters corresponding to the motion description information; The motion sequence parameters are input into a pre-trained visual motion generation model to map the motion sequence parameters to a three-dimensional pose space, thereby obtaining the three-dimensional human motion sequence.
7. The method of claim 1, wherein the second deviation is determined based on at least one of the following: The deviation between the speech features and the vocal vibration features representing the start and end times of vocalization; The deviation between the intensity change trend represented by the speech features and the vocal vibration features; The deviation between the peak occurrence time represented by the speech feature and the vocal vibration feature.
8. A method for detecting wear, comprising: Acquire the voice signals and inertial signals collected by the smart wearable device within the same time period; The inertial signal is input into a pre-trained signal separation model to separate the vocal vibration characteristics corresponding to the vibration signal generated when the user speaks from the inertial signal through the signal separation model; wherein, the signal separation model is trained by the method described in any one of claims 1-7. Extract the speech features corresponding to the speech signal, and determine whether the speech features and the vocal vibration features satisfy a preset correspondence; Based on the judgment result, it is determined whether the smart wearable device is in the state of being actually worn by the user.
9. A training device for a signal separation model, comprising: An acquisition module is used to acquire training samples, which include voice signals and inertial signals collected by the smart wearable device within the same time period; The separation module is used to input the inertial signal into the signal separation model to be trained, so as to separate the motion interference features corresponding to the motion signal generated when the user performs body movements and the vocal vibration features corresponding to the vibration signal generated when the user speaks from the inertial signal through the signal separation model. The determination module is used to determine the loss value for the signal separation model based on a first deviation between the vocal vibration feature and the motion interference feature, and a second deviation between the vocal vibration feature and the speech feature corresponding to the speech signal. An adjustment module is used to adjust the model parameters of the signal separation model based on the loss value.
10. A wear detection device, comprising: The acquisition module is used to acquire the voice signals and inertial signals collected by the smart wearable device within the same time period; A separation module is used to input the inertial signal into a pre-trained signal separation model, so as to separate the vocal vibration features corresponding to the vibration signal generated when the user speaks from the inertial signal through the signal separation model; wherein, the signal separation model is trained by the method described in any one of claims 1-7. The judgment module is used to extract the speech features corresponding to the speech signal and determine whether the speech features and the vocal vibration features satisfy a preset correspondence. The detection module is used to determine, based on the judgment result, whether the smart wearable device is in the state of being actually worn by the user.
11. An electronic device comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as described in any one of claims 1-8 by executing the executable instructions.
12. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-8.