Automatic wake-up method and device, wearable device, and storage medium
By detecting the user's head characteristic information and motion trajectory, combined with the two-step head action matching mechanism, the intelligent wake-up of the device is achieved, solving the inconvenience and false triggering of the traditional wake-up method, and improving the user experience and device response speed.
Patent Information
- Application Number
- CN202510212964.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Traditional wake-up methods such as key wake-up and single action wake-up have problems such as inconvenient operation and false triggering, which affects the device battery life and user experience.
By detecting the user's head characteristic information and motion trajectory, the user's head movement is determined, and through the two-step head movement matching mechanism, a wake-up prompt information is generated and the device is awakened after the user confirms.
It simplifies the operation process of device wake-up, improves the response speed of the device, enhances the interactive experience between users and devices, and improves the accuracy and reliability of wake-up operations.
Smart Images

Figure CN119718088B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of artificial intelligence technology, and more specifically, relates to an automatic wake-up method and device, a wearable device, and a storage medium. Background Art
[0002] With the increasing popularity of smart wearable devices, how to wake up the device conveniently and accurately has become a key issue. Traditional wake-up methods, such as button wake-up or single action wake-up, have many inconveniences and defects. Button wake-up is inconvenient to operate in some scenarios, for example, it is difficult for the user to reach the button when both hands are occupied. Single action wake-up is very easy to be triggered by mistake due to the user's daily inadvertent actions, causing the device to wake up unexpectedly, which not only affects the device's battery life but also interferes with the user's normal use, thus giving the user a poor experience. Summary of the invention
[0003] The purpose of the present disclosure is to provide an automatic wake-up method and apparatus, a wearable device, and a storage medium to enhance the user experience.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an automatic wake-up method, which is applied to a first device and includes:
[0005] Determine a head movement trajectory of the first user based on the head feature information of the first user;
[0006] determining a first head motion based on a head motion trajectory of the first user;
[0007] In response to a first head movement of the first user matching a first preset movement, generating a wake-up prompt message;
[0008] In response to a second head movement of the first user within the first time period matching a second preset movement, waking up the first device;
[0009] The first device is a device worn by a first user.
[0010] According to a second aspect of an embodiment of the present disclosure, there is provided an automatic wake-up device, which is applied to a first device and includes:
[0011] A motion trajectory generating module, configured to determine a head motion trajectory of the first user based on the head feature information of the first user;
[0012] A head action generating module, configured to determine a first head action based on a head movement trajectory of a first user;
[0013] A wake-up prompt module, configured to generate wake-up prompt information in response to a first head movement of a first user matching a first preset movement;
[0014] A wake-up module, configured to wake up the first device in response to a second head movement of the first user matching a second preset movement within a first time;
[0015] The first device is a device worn by a first user.
[0016] According to a third aspect of an embodiment of the present disclosure, a wearable device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the above-mentioned automatic wake-up method when executing the computer program.
[0017] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the automatic wake-up method described above are implemented.
[0018] The automatic wake-up method and apparatus, wearable device, and storage medium provided by the disclosed embodiments have the following beneficial effects: On the one hand, by detecting the user's specific head movements, the embodiment can provide immediate feedback (wake-up prompt information) when the user expects to wake up the device, and quickly wake up the device after further confirmation by the user (second head movement matching). This not only simplifies the operation process of waking up the device, improves the response speed of the device, but also enhances the interactive experience between the user and the device. On the other hand, the embodiment uses a two-step head movement matching mechanism to avoid the situation where the device is mistakenly woken up because the user accidentally makes a similar preset movement, thereby increasing the accuracy and reliability of the wake-up operation; by dividing the action requirements and time limits into two steps, the user's true intention to actively wake up the device is confirmed in a more hierarchical manner, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 A flowchart of an automatic wake-up method provided by an embodiment of the present disclosure;
[0021] Figure 2 A structural block diagram of an automatic wake-up device provided by an embodiment of the present disclosure;
[0022] Figure 3 A schematic block diagram of a wearable device provided in one embodiment of the present disclosure. DETAILED DESCRIPTION
[0023] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present disclosure. However, it should be clear to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present disclosure with unnecessary details.
[0024] In order to make the purpose, technical solutions and advantages of the present disclosure more clear, specific embodiments will be described below in conjunction with the accompanying drawings.
[0025] Please refer to Figure 1 , Figure 1 A flowchart of an automatic wake-up method provided in an embodiment of the present disclosure is provided. The method is applied to a first device and includes:
[0026] S101: Determine a head movement trajectory of the first user based on head feature information of the first user.
[0027] In this embodiment, the first device is a specific device in the context of the automatic wake-up method, and the first device is a device that can be worn by the user. The first device may have certain intelligent functions and built-in corresponding sensors and other components. In this embodiment, the first device may be smart glasses. It is capable of detecting the user's head movements, and then executing subsequent wake-up prompts and final wake-up operations and other related processes according to the set rules. The first user is an individual who has an interactive relationship with the first device and expects to wake up the above-mentioned first device through head movements.
[0028] In this embodiment, sensors built into the first device (such as an accelerometer and a gyroscope, etc.) can continuously capture motion data of the user's head in three-dimensional space, and these motion data can be recorded in the form of head coordinate information or motion parameters.
[0029] The actual motion trajectory of the first user's head is constructed based on the head coordinate information or the motion parameters.
[0030] For example, the acceleration sensor can sense the changes in acceleration of the head in different directions, from which the speed change of the head movement can be calculated. Through operations such as integration of the speed over time, the change in the head position can be further determined, and then the displacement information of the movement can be obtained, that is, the amplitude of the movement. The gyroscope can accurately detect the angle of head rotation and the speed of angle change, reflecting the angle and frequency characteristics of the head movement.
[0031] S102: Determine a first head action based on a head movement trajectory of the first user.
[0032] In this embodiment, when the user's head starts to rotate, the sensor can record the offset of the head position relative to the initial position at each moment, as well as the angular velocity, angular acceleration and other information of the rotation. These information are integrated to form a complete head movement trajectory curve.
[0033] The actual head motion trajectory curve is compared with the pre-set standard motion trajectory corresponding to the first preset action, so as to determine the corresponding first head action.
[0034] The standard motion trajectory of the first preset action is an ideal trajectory model defined and stored in the device according to a specific action form (such as nodding, shaking or turning the head at a specific angle) during the development or setting process of the first device. It includes key parameters such as the range of motion, speed change, and duration of the action in various dimensions.
[0035] S103: In response to the first head movement of the first user matching the first preset movement, generating a wake-up prompt message.
[0036] In this embodiment, when the first head movement made by the user matches a pre-set first preset movement, the first device may generate a wake-up prompt message. The first preset movement in this embodiment may be a pre-defined movement form such as nodding, shaking head or turning head at a specific angle.
[0037] Exemplarily, the first preset action is to quickly turn the head 30 degrees to the left and hold it for 1 second. When the user makes such a precisely corresponding head movement, the relevant sensors in the device (such as accelerometers, gyroscopes, etc., which can monitor changes in the movement state of the head) detect that this action meets the set conditions. At this time, the corresponding program logic can be triggered to generate a wake-up prompt message.
[0038] In this embodiment, the role of the wake-up prompt information can be understood as giving the user feedback, indicating that the first device has noticed the user's specific action, and can inform the user through sound, vibration, or prompts on a display device connected to it, and then continue with subsequent operations to complete the device wake-up.
[0039] S104: In response to a second head movement of the first user within a first time period matching a second preset movement, waking up the first device; the first device is a device worn by the first user.
[0040] In this embodiment, after the wake-up prompt information is generated, the first user is required to make a second head movement within the first time, and the second head movement must match the second preset movement. Only in this way can the first device be truly woken up.
[0041] The first time is a limited time range, for example, it can be set to 5 seconds, which means that the user needs to complete the second head movement that meets the requirements within these 5 seconds. The second preset action is also a specific head movement form that is set in advance. The second preset action is different from the first preset action and serves to further confirm the user's wake-up intention. For example, if the second preset action is set to nod twice quickly, the user needs to make such an action within the specified first time (such as within 5 seconds). After the first device detects that the action meets the requirements again through the sensor, it executes the wake-up operation to make the first device enter a normal working state from an incompletely activated state such as standby or low power consumption. At this time, the user can start to control the first device normally.
[0042] For example, suppose a user wears smart glasses (i.e., the first device) and goes out for a ride. When the user wants to check the riding route or other information, if the first preset action is to quickly turn the head 30 degrees to the left and hold for 1 second. When the user turns his head to the left during riding, the acceleration sensor and gyroscope in the smart glasses detect that this action matches the first preset action, and then generates a wake-up prompt message. The smart glasses emit a slight prompt sound, and the temples vibrate briefly.
[0043] In the next 5 seconds (the first time), if the second preset action is to nod twice quickly, the user nods twice quickly in time, and the smart glasses detect that the action meets the requirements again, they will be fully awakened from the low-power standby state. At this time, the display screen of the smart glasses lights up, showing the current riding speed, route navigation and other information. The user can control the smart glasses to switch the display content, zoom in and out the map, etc. through voice commands or head movements, which is convenient for obtaining information during riding without manual operation, greatly improving the user experience and riding safety.
[0044] From the above, it can be concluded that, on the one hand, this embodiment can provide immediate feedback (wake-up prompt information) when the user expects to wake up the device by detecting the user's specific head movement, and quickly wake up the device after further confirmation by the user (second head movement matching). It not only simplifies the operation process of waking up the device, improves the response speed of the device, but also enhances the interactive experience between the user and the device. On the other hand, this embodiment can avoid the situation where the device is mistakenly awakened because the user accidentally makes a similar preset action through a two-step head movement matching mechanism, thereby increasing the accuracy and reliability of the wake-up operation; through two steps with different action requirements and time limits, the user's true intention to actively wake up the device is confirmed in a more hierarchical manner, thereby improving the user experience.
[0045] Determining a head movement trajectory of the first user based on head feature information of the first user includes:
[0046] A head movement trajectory of the first user is constructed based on the spatiotemporal information corresponding to the head feature information of the first user.
[0047] In this embodiment, the head feature information of the first user may include key elements such as amplitude, frequency, angle, and speed of the head movement.
[0048] In this embodiment, the spatiotemporal information includes both the time dimension and the space dimension. In this context, the time dimension records the order in which the head movements occur and the duration of each movement stage, etc.; the space dimension involves the position change and rotation direction of the head in three-dimensional space.
[0049] For example, the speed of head movement determines how fast the head moves in space per unit time, the angle reflects the direction and amplitude of the head's rotation in space, the frequency reflects the pattern of repetition of specific movements in time, and the amplitude intuitively shows the range covered by the head movement in space.
[0050] After obtaining a large amount of spatiotemporal information corresponding to head feature information at different times, the first device can integrate this information. Taking time as the main line, the position of the head in space (obtained through relevant calculations of the acceleration sensor), the rotation angle (detected by the gyroscope) and the corresponding speed, frequency and other feature information obtained at each time point are arranged in order. A data set arranged in time order and containing spatial state information is formed, thereby forming a complete trajectory curve that can reflect the actual movement of the first user's head. The trajectory curve (head movement trajectory) can be represented as the first head action of the first user.
[0051] From the above, it can be concluded that this embodiment can accurately capture subtle changes in head movements through detailed head feature information, and the constructed head movement trajectory can fully reflect the overall picture of the movement, reduce misjudgment caused by fuzzy movement recognition, and greatly improve the accuracy of head movement judgment. Different users may have differences when performing the same type of action. This embodiment can form an exclusive movement trajectory based on the individual's unique head feature information. Regardless of the size of the movement, speed, or frequency, it can be effectively identified, adapting to a wide range of user groups and diverse movement styles.
[0052] In one embodiment of the present disclosure, determining a first head motion based on a head motion trajectory of a first user includes:
[0053] Calculating the similarity between the head movement trajectory of the first user and the first preset action;
[0054] The first head action is determined based on the similarity between the head movement trajectory of the first user's head and the first preset action.
[0055] In this embodiment, the similarity between the actual movement trajectory of the first user's head and the first preset action is calculated.
[0056] During the comparison process, this embodiment may use a specific algorithm to calculate the similarity or difference between the actual trajectory and the standard trajectory.
[0057] For example, the similarity between the actual trajectory and the standard trajectory can be calculated based on the vector space model. For each time point or key node of the head position, its coordinates can be represented by a three-dimensional vector. Then, the similarity is measured by calculating the cosine value of the angle between the two vector sets. The closer the cosine value of the angle is to 1, the closer the directions of the two vectors are, that is, the more similar the actual head motion trajectory is to the preset motion trajectory.
[0058] If the similarity between the two reaches a certain threshold, it indicates that the first head movement of the first user successfully matches the first preset movement. At this time, the device can trigger the corresponding program module to generate a wake-up prompt message to inform the user that the device has detected the first step of the effective wake-up movement and can proceed to the next step to finally wake up the device.
[0059] From the above, it can be concluded that this embodiment determines the first head movement based on the head movement trajectory, and uses the built-in sensor of the device to accurately capture the movement data of the head in three-dimensional space, which can comprehensively and meticulously record the head movement and make the action recognition more accurate. Secondly, the actual movement trajectory is compared with the standard movement trajectory of the preset action and the similarity is calculated. For example, the cosine value of the angle is calculated using a vector space model to quantify the degree of action matching and effectively reduce misjudgment. When the similarity reaches the threshold, a wake-up prompt message is generated, which not only allows the user to know that the action has been recognized in time, but also lays the foundation for the subsequent complete wake-up process, enhances the reliability and fluency of the user's interaction with the device, and improves the overall user experience.
[0060] In one embodiment of the present disclosure, determining a first head motion based on a head motion trajectory of a first user includes:
[0061] Determine a head movement trajectory of the first user based on the head feature information of the first user;
[0062] A first head motion is determined based on the head movement trajectory of the first user.
[0063] In this embodiment,
[0064] In one embodiment of the present disclosure, a head movement trajectory of the first user is constructed based on the spatiotemporal information corresponding to the head feature information of the first user.
[0065] In one embodiment of the present disclosure, it also includes:
[0066] Retrieving a third preset action corresponding to the first user based on the identity information of the first user;
[0067] In response to the first user's third head movement matching the third preset movement, waking up the first device.
[0068] In this embodiment, the first device has the ability to identify the identity information of the first user, and the identity information can be obtained in a variety of ways, such as the user registering and logging in when the first device is used for the first time, and entering unique identification information such as user name, voice, fingerprint or facial features; or it may be indirectly obtained by associating and binding with other authenticated devices (such as mobile phones). When the user uses the first device, identity authentication and identification can be performed, that is, the collected information is compared and verified with the user identity database stored locally on the device to determine whether the user is a registered historical user.
[0069] If it is determined that the user is a historical user, the first device can search and retrieve the third preset action pre-set for the user in the corresponding database record based on the specific identity of the user. The third preset action is customized according to the user's own habits, preferences and possible usage scenario requirements when the user uses the device for the first time or in the subsequent personalized settings. For example, some users may be accustomed to shaking their heads three times quickly to wake up the device, while some users may set it to nod slowly and keep it for a certain length of time. The third preset action of each user is unique and is stored in a database indexed by user identity information for accurate retrieval.
[0070] During the use of the first device, after the identity information is confirmed, the user's third head movement can be obtained in real time, and the actual third head movement related features can be compared one by one with the standard parameters required by the retrieved third preset action. If the matching requirements are met in all key indicators, for example, the actual head movement angle, speed, frequency, etc. are all within the reasonable range set by the third preset action, then it is determined that the third head movement matches the third preset action successfully. When the match is successful, the first device is awakened.
[0071] From the above, it can be concluded that this embodiment calls exclusive preset actions based on user identity information, fully respects the differences in personal habits and usage scenario preferences of different users, and realizes personalized customization of the wake-up method. Secondly, this design greatly improves the user experience. Users no longer need to adapt to the unified wake-up standard. They can use their familiar and comfortable head movements to wake up the device, which makes the operation more natural and convenient.
[0072] In one embodiment of the present disclosure, before calling the third preset action corresponding to the first user based on the identity information of the first user, the method further includes:
[0073] The identity information of the first user is determined based on a degree of matching between a first audio feature of the first user and a preset audio feature.
[0074] In this embodiment, when the user uses the first device for the first time, the first device can turn on the audio collection function and guide the user to record a specific audio content, such as asking the user to say a few specific sentences, such as "Hello device, I am so-and-so". In this process, the built-in audio sensor (such as a microphone, etc.) of the first device can collect relevant acoustic parameters of the audio, and these parameters together constitute the first audio feature of the first user. For example, audio features cover many aspects such as the pitch, the characteristics of the timbre, the spectrum distribution of the audio, the speaking speed, and the rhythm of the voice. Subsequently, the first device can organize and encode these collected first audio features, and store them in a local database as preset audio features, and at the same time associate and bind them with the identity information of the user to lay the foundation for subsequent identity recognition.
[0075] During use, when it is necessary to call the corresponding preset action based on the identity information of the first user, the first device first collects the audio information emitted by the current user through the audio sensor, and also extracts the corresponding audio features (first audio features). Then, these real-time audio features are compared with the preset audio features previously stored in the database to calculate the degree of match between the two. If the degree of match reaches a threshold standard pre-set by the device, it means that the current user is the first user who previously entered the audio features, thereby determining the identity information of the user, and then based on this identity, the corresponding third preset action can be called to start the subsequent device wake-up process based on personalized needs.
[0076] From the above, it can be concluded that this embodiment provides a convenient and relatively unique identity recognition method for personalized use of the device and automatic wake-up function by determining identity information based on audio feature matching, which can play a good role in some scenarios where it is inconvenient to verify identity through other methods (such as fingerprints, facial recognition, etc.).
[0077] In one embodiment of the present disclosure, before determining the identity information of the first user based on the matching degree between the first audio feature of the first user and the preset audio feature, the method further includes:
[0078] Segmenting the first audio of the first user to obtain a plurality of first audio segments;
[0079] Extracting features from the plurality of first audio segments to obtain a plurality of first feature segments;
[0080] Performing dimensionality reduction processing on multiple feature segments to obtain multiple target feature segments;
[0081] Multiple target feature segments are concatenated to obtain a first audio feature.
[0082] In this embodiment, when the user uses the first device for audio-related settings or identity registration for the first time, the first audio received by the microphone is transmitted to the processing unit. Since a complete audio may contain a variety of different voice elements, pauses, and background noise changes, in order to more accurately extract effective information, the first audio needs to be segmented. A specific audio segmentation algorithm is used, such as an endpoint detection algorithm based on short-time energy and zero-crossing rate, to divide the first audio into multiple relatively independent first audio segments with similar acoustic characteristics by analyzing the energy changes of the audio signal in a short period of time and the number of times the signal crosses the zero value.
[0083] In this embodiment, for each first audio segment obtained by segmentation, the Mel-frequency cepstral coefficient (MFCC) of the audio segment can be extracted. The Mel-frequency cepstral coefficient can simulate the human ear's perception characteristics of sound frequency and effectively characterize the spectral characteristics of the audio; at the same time, the linear prediction coding coefficient (LPC) is extracted. The linear prediction coding coefficient can reflect the resonance characteristics of the vocal tract and plays an important role in identifying the timbre and other characteristics of the speech. Through feature extraction, a corresponding first feature segment is generated for each first audio segment. The first feature segment can include key elements such as the frequency distribution of the audio, the position of the resonance peak, and the pitch of the speech.
[0084] In this embodiment, since the extracted multiple first feature segments may have high dimensions and contain a lot of redundant information, if these high-dimensional features are directly used for identity matching calculation, a lot of computing resources will be consumed and matching efficiency and accuracy may be reduced. Therefore, this embodiment can use dimensionality reduction technology.
[0085] For example, by finding the main components in the data, multiple first feature segments are projected from the first dimension to the feature space of the second dimension. The first dimension is larger than the second dimension. The main features that contribute more to the data variance are retained, and the secondary redundant features are discarded, thereby obtaining multiple target feature segments.
[0086] In this embodiment, multiple target feature segments after dimensionality reduction processing need to be recombined into a complete first audio feature that can be used for identity recognition. Multiple target feature segments are spliced in the original order of the audio segments to form a first audio feature in the form of a comprehensive audio feature vector. The first audio feature represents a unique acoustic identifier of the user's audio. The first audio feature is matched with the preset audio feature, and the first device can determine the identity information of the first user based on the match, and then retrieve the corresponding personalized settings based on the identity information.
[0087] From the above, it can be concluded that this embodiment helps to focus on local audio features through audio segmentation, reduce the overall audio complexity, and improve the accuracy of subsequent processing. Extracting multiple feature segments can characterize audio from different dimensions, making the audio representation more comprehensive and accurate. Dimensionality reduction processing can significantly reduce redundant data, save computing resources and time costs, and make the identity recognition process more efficient. The first audio feature formed by splicing is used as a unique acoustic identifier, which is matched with the preset audio feature to determine the identity information, thereby enhancing the reliability and uniqueness of identity recognition, laying a solid foundation for the realization of subsequent personalized functions (such as personalized wake-up actions), and improving user experience and device intelligence.
[0088] In one embodiment of the present disclosure, determining the identity information of the first user based on the matching degree between the first audio feature of the first user and the preset audio feature includes:
[0089] Calculating the similarity between the first audio feature and the preset audio feature to obtain a first matching value;
[0090] In response to the first matching value being greater than the first threshold, a transition probability and an emission probability of the first audio feature sequence are calculated, and identity information of the first user is determined based on the transition probability and the emission probability of the first audio feature sequence.
[0091] In this embodiment, a cosine similarity algorithm may be used to calculate the similarity between the first audio feature and the preset audio feature.
[0092] Exemplarily, the first audio feature and the preset audio feature are regarded as two vectors, and the cosine value of the angle between the two vectors is calculated. The cosine value represents their similarity, and the obtained value is the first matching value. If this value is closer to 1, it indicates that the two are more similar in feature performance; the closer to 0, the greater the difference. Through such calculation, the similarity between the currently collected audio and the pre-stored audio features representing the user identity is preliminarily measured.
[0093] When the calculated first matching value is greater than the first threshold, it means that from the perspective of simple feature similarity alone, the current audio is more likely to come from the corresponding first user, but in order to more accurately determine the identity and avoid misjudgment, further in-depth analysis is required.
[0094] At this time, the transition probability and emission probability of the first audio feature sequence may be calculated.
[0095] In this embodiment, the first audio feature sequence can be regarded as a plurality of hidden states converted in time sequence, and these hidden states can correspond to different speech units (such as phonemes, syllables, etc.) or speech states (such as pronunciation start, duration, end state, etc.). The transition probability reflects the possibility of transitioning from one hidden state to another hidden state.
[0096] For example, in a speech, the probability of transitioning from a hidden state corresponding to a certain phoneme to a hidden state corresponding to the next phoneme. Due to differences in pronunciation habits, speaking speed, intonation, etc., different users have different transition probabilities between these hidden states in their speech. By analyzing the transition probability of the first audio feature sequence, more detailed speech pattern characteristics can be mined and compared with the pre-stored audio features corresponding to the correct identity in this regard.
[0097] In this embodiment, the emission probability reflects the probability of generating the actually observed audio features in a certain hidden state. In other words, given a hidden state (such as the pronunciation state of a specific vowel), the probability of the specific audio features currently collected (such as the corresponding frequency, amplitude, etc.) appearing. Each user's unique timbre, pronunciation method and other factors will result in different emission probabilities in the same hidden state. Analyzing the emission probability of the first audio feature sequence can further determine whether it matches the user represented by the preset audio feature from the perspective of the association between the audio feature and the hidden state.
[0098] For example, suppose that the user registers and records the voice "Hello, Xiaojing" when using the smart glasses for the first time, and the device extracts and stores the corresponding preset audio features. When using it later, the user says this sentence again. The smart glasses first process the real-time audio to obtain the first audio feature, and use the cosine similarity algorithm to calculate the similarity with the preset audio feature to obtain a first matching value of 0.8, which is greater than the set first threshold of 0.7. Then the transition probability and emission probability of the first audio feature sequence are further calculated. For example, the transition probability of the conversion of the phoneme "good" to "mirror" in the voice is analyzed, and it is found that the user is accustomed to fast continuous reading, and this transition probability is high, which is consistent with the pre-stored rules. Looking at the emission probability when the phoneme "mirror" is pronounced, due to the user's unique timbre and pronunciation method, the emission probability corresponding to the characteristics such as the audio frequency and amplitude generated by it is also similar to the pre-stored data of the user. Based on the analysis results of the transition probability and the emission probability, it is determined that this audio comes from the registered user, thereby completing the accurate determination of the identity information, laying the foundation for the activation of subsequent personalized functions (such as exclusive wake-up action response).
[0099] From the above, it can be concluded that this embodiment obtains the first matching value by calculating the similarity between the first audio feature and the preset audio feature, which can quickly perform preliminary screening, exclude audio that is obviously inconsistent, and improve the efficiency of identity recognition. Secondly, when the first matching value is greater than the threshold, the transfer probability and the emission probability are further calculated to deeply explore the inherent laws of the audio feature sequence. The differences in pronunciation habits of different users are reflected in the transfer probability, and factors such as unique timbre are presented in the emission probability, which makes identity recognition more accurate, effectively reduces the misjudgment rate, enhances the reliability of the device for user identity authentication, and ensures the security and accuracy of subsequent identity-based personalized function applications.
[0100] Corresponding to the automatic wake-up method in the above embodiment, Figure 2 This is a structural block diagram of an automatic wake-up device provided by an embodiment of the present disclosure. For ease of description, only the parts related to the embodiment of the present disclosure are shown. Figure 2 The automatic awakening device 20 includes: a motion trajectory generating module 21, a head movement generating module 22, a awakening prompting module 23, and a awakening module 24.
[0101] The motion trajectory generating module 21 is used to determine the head motion trajectory of the first user based on the head feature information of the first user;
[0102] A head action generating module 22, configured to determine a first head action based on a head movement trajectory of a first user;
[0103] A wake-up prompt module 23, configured to generate a wake-up prompt message in response to a first head movement of a first user matching a first preset movement;
[0104] A wake-up module 24, configured to wake up the first device in response to a second head movement of the first user matching a second preset movement within the first time;
[0105] The first device is a device worn by a first user.
[0106] In one embodiment of the present disclosure, the motion trajectory generating module 21 is specifically used for:
[0107] A head movement trajectory of the first user is constructed based on the spatiotemporal information corresponding to the head feature information of the first user.
[0108] In one embodiment of the present disclosure, the head action generation module 22 is specifically used for:
[0109] Calculating the similarity between the head movement trajectory of the first user and the first preset action;
[0110] The first head action is determined based on the similarity between the head movement trajectory of the first user's head and the first preset action.
[0111] In one embodiment of the present disclosure, the automatic wake-up device 20 further includes: an identity recognition module;
[0112] The identity recognition module is specifically used for:
[0113] Retrieving a third preset action corresponding to the first user based on the identity information of the first user;
[0114] In response to the first user's third head movement matching the third preset movement, waking up the first device.
[0115] In one embodiment of the present disclosure, the identity recognition module is further used to:
[0116] The identity information of the first user is determined based on a degree of matching between a first audio feature of the first user and a preset audio feature.
[0117] In one embodiment of the present disclosure, the identity recognition module is specifically used for:
[0118] Segmenting the first audio of the first user to obtain a plurality of first audio segments;
[0119] Extracting features from the plurality of first audio segments to obtain a plurality of first feature segments;
[0120] Performing dimensionality reduction processing on multiple feature segments to obtain multiple target feature segments;
[0121] Multiple target feature segments are concatenated to obtain a first audio feature.
[0122] In one embodiment of the present disclosure, the identity recognition module is specifically used for:
[0123] Calculating the similarity between the first audio feature and the preset audio feature to obtain a first matching value;
[0124] In response to the first matching value being greater than the first threshold, a transition probability and an emission probability of the first audio feature sequence are calculated, and identity information of the first user is determined based on the transition probability and the emission probability of the first audio feature sequence.
[0125] See also Figure 3 , Figure 3 A schematic block diagram of a wearable device provided by an embodiment of the present disclosure. Figure 3 The wearable device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303 and one or more memories 304. The processors 301, input devices 302, output devices 303 and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of each module in the above-mentioned device embodiments, such as Figure 2 The functions of modules 21 to 22 are shown.
[0126] It should be understood that in the embodiment of the present disclosure, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0127] The input device 302 may include a touch panel, a fingerprint collection sensor (for collecting the user's fingerprint information and fingerprint direction information), a microphone, etc., and the output device 303 may include a display (LCD, etc.), a speaker, etc.
[0128] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0129] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present disclosure can execute the implementation methods described in the first and second embodiments of the automatic wake-up method provided in the embodiments of the present disclosure, and can also execute the implementation methods of the wearable device described in the embodiments of the present disclosure, which will not be repeated here.
[0130] In another embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by the processor, all or part of the processes in the above-mentioned embodiment method are implemented, and the computer program can also be completed by instructing the relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0131] The computer-readable storage medium may be an internal storage unit of the wearable device of any of the foregoing embodiments, such as a hard disk or memory of the wearable device. The computer-readable storage medium may also be an external storage device of the wearable device, such as a plug-in hard disk equipped on the wearable device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Furthermore, the computer-readable storage medium may also include both an internal storage unit of the wearable device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the wearable device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0132] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.
[0133] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the wearable device and unit described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0134] In the several embodiments provided in the present application, it should be understood that the disclosed wearable devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or it can be an electrical, mechanical or other form of connection.
[0135] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present disclosure.
[0136] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0137] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present disclosure, and these modifications or replacements should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. An automatic wake-up method, applied to a first device, characterized in that: include: Determine a head movement trajectory of the first user based on the head feature information of the first user; determining a first head motion based on a head motion trajectory of the first user; In response to a first head movement of the first user matching a first preset movement, generating a wake-up prompt message; In response to a second head movement of the first user within the first time period matching a second preset movement, waking up the first device; The second preset action is different from the first preset action; Controlling the first device through voice commands or head movements; The first device is a device worn by a first user.
2. The automatic wake-up method according to claim 1, characterized in that: The determining the head movement trajectory of the first user based on the head feature information of the first user includes: A head movement trajectory of the first user is constructed based on the spatiotemporal information corresponding to the head feature information of the first user.
3. The automatic wake-up method according to claim 2, characterized in that: The determining the first head action based on the head movement trajectory of the first user includes: Calculating the similarity between the head movement trajectory of the first user and the first preset action; The first head action is determined based on the similarity between the head movement trajectory of the first user's head and the first preset action.
4. The automatic wake-up method according to claim 1, characterized in that: Also includes: Retrieving a third preset action corresponding to the first user based on the identity information of the first user; In response to the first user's third head movement matching the third preset movement, waking up the first device.
5. The automatic wake-up method according to claim 4, characterized in that: Before calling the third preset action corresponding to the first user based on the identity information of the first user, the method further includes: The identity information of the first user is determined based on a degree of matching between a first audio feature of the first user and a preset audio feature.
6. The automatic wake-up method according to claim 5, characterized in that: Before determining the identity information of the first user based on the matching degree between the first audio feature of the first user and the preset audio feature, the method further includes: Segmenting the first audio of the first user to obtain a plurality of first audio segments; Extracting features from the plurality of first audio segments to obtain a plurality of first feature segments; Performing dimensionality reduction processing on the multiple feature segments to obtain multiple target feature segments; Multiple target feature segments are concatenated to obtain a first audio feature.
7. The automatic wake-up method according to claim 5, characterized in that: The determining the identity information of the first user based on the matching degree between the first audio feature of the first user and the preset audio feature includes: Calculating the similarity between the first audio feature and the preset audio feature to obtain a first matching value; In response to the first matching value being greater than a first threshold, a transition probability and an emission probability of a first audio feature sequence are calculated, and identity information of the first user is determined based on the transition probability and the emission probability of the first audio feature sequence.
8. An automatic wake-up device, applied to a first device, characterized in that: include: A motion trajectory generating module, configured to determine a head motion trajectory of the first user based on the head feature information of the first user; A head action generating module, configured to determine a first head action based on a head movement trajectory of a first user; A wake-up prompt module, configured to generate wake-up prompt information in response to a first head movement of a first user matching a first preset movement; A wake-up module, configured to wake up the first device in response to a second head movement of the first user matching a second preset movement within a first time; The second preset action is different from the first preset action; Controlling the first device through voice commands or head movements; The first device is a device worn by a first user.
9. A wearable device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Application awakening method and device, computer equipment and storage medium
CN119376798A
Wakeup method, apparatus and device based on lip reading, and computer readable medium
US20190228212A1