A smart watch-based voice data processing method and system

By constructing a voice perturbation feature vector and jointly determining the operating status on a smartwatch, and dynamically adjusting the voice processing path, the power consumption and stability issues caused by invalid voice input in smartwatches are solved, achieving more efficient resource utilization and stable operation.

CN122369435APending Publication Date: 2026-07-10ZHOUHAI INTELLIGENT (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHOUHAI INTELLIGENT (SHENZHEN) CO LTD
Filing Date
2026-04-01
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing smartwatch voice processing technology is prone to being misjudged as valid voice input due to environmental noise, meaningless speech, or occasional sounds, leading to increased system power consumption, processor load fluctuations, and unstable operating status. It lacks an effective voice input modeling and judgment mechanism, and cannot avoid the waste of resources and frequent switching of operating status caused by invalid voice triggers.

Method used

By combining speech disturbance modeling with operational status determination, the energy change, temporal continuity, and rhythm change features of the speech signal are extracted to construct a speech disturbance feature vector, generate a disturbance state representation, and combine it with the operational status for joint determination, dynamically adjusting the speech processing path and avoiding invalid triggering.

Benefits of technology

It effectively reduces power consumption and system load fluctuations caused by invalid voice triggers, improves the operational stability and battery life of smartwatches, and reduces frequent system status switching caused by invalid voice triggers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369435A_ABST
    Figure CN122369435A_ABST
Patent Text Reader

Abstract

This invention discloses a voice data processing method and system based on a smartwatch, comprising the following steps: acquiring continuous voice signals and performing preprocessing; performing segmented processing to extract energy change features, temporal continuity features, and rhythmic change features of the voice signals, constructing a voice disturbance feature vector; performing temporal aggregation and linear mapping on the voice disturbance feature vector set, and generating disturbance level identifiers through hierarchical quantization, and generating a voice disturbance state representation by combining the voice disturbance intensity score; constructing a set of operating status states and determining the current operating status state; jointly determining the voice disturbance state representation and the current operating status state; performing hold or migration operations on the current operating status state; and dynamically determining the voice data processing path. This invention, based on a voice disturbance modeling and operating status joint determination method, achieves dynamic control of the voice processing path, possessing advantages such as reduced power consumption, improved stability, and avoidance of invalid triggering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice data processing, and in particular to a voice data processing method and system based on a smartwatch. Background Technology

[0002] With the popularization of wearable devices such as smartwatches, voice interaction has gradually become one of the common human-computer interaction methods. The voice processing technology in existing smartwatches is usually based on voice content recognition. After detecting voice input, it directly triggers voice wake-up, command recognition or interactive response process. Such solutions generally rely on the existence of voice or whether the wake-up conditions are met to determine whether to enter the voice processing state. Voice is regarded as an information carrier that needs to be recognized and understood in the system, and its processing logic is function trigger-oriented.

[0003] However, in actual operation, smartwatches operate in environments with limited energy consumption and computing resources. Environmental noise, meaningless speech, or occasional sounds are easily misjudged as valid voice input, frequently triggering voice processing or interaction preparation processes. This leads to increased system power consumption, processor load fluctuations, and unstable operating status. Existing technologies generally lack mechanisms for modeling and judging voice input from the perspective of system operating status and stability. It is difficult to distinguish between voice input that has a real impact on system operation and voice input that has no real value to the system. As a result, it is impossible to effectively avoid the problem of resource waste and frequent switching of operating status caused by invalid voice triggers. Summary of the Invention

[0004] One objective of this invention is to propose a voice data processing method and system based on a smartwatch. This invention is based on a voice disturbance modeling and operational status joint determination method to achieve dynamic control of the voice processing path, which has the advantages of reducing power consumption, improving stability and avoiding invalid triggering.

[0005] A voice data processing method based on a smartwatch according to an embodiment of the present invention includes the following steps: During the operation of the smartwatch, continuous voice signals are collected through the smartwatch's voice acquisition component to generate raw voice data, and preprocessing is performed to generate a standardized voice data sequence. The standardized speech data sequence is segmented and processed. Within each time window, the energy change features, temporal continuity features, and rhythm change features of the speech signal are extracted to construct a speech perturbation feature vector and form a set of speech perturbation feature vectors. Temporal aggregation and linear mapping are performed on the speech disturbance feature vector set to generate speech disturbance intensity scores, and level quantization is used to generate disturbance level labels. The speech disturbance state representation is generated by combining the speech disturbance intensity scores. A set of operational status states is pre-built on the smartwatch, and the current operational status state is determined based on the current operational parameters of the smartwatch; The voice disturbance state representation is jointly determined with the current operational status, and an operational status transition determination result is generated based on the degree of influence of the voice disturbance state representation on the stability of the current operational status. Based on the operational status migration determination result, perform a maintain or migrate operation on the current operational status to obtain the updated operational status. Based on the updated operational status, the processing path for voice data is dynamically determined, and voice data processing is completed.

[0006] Optionally, the preprocessing includes time synchronization, outlier removal, and amplitude normalization.

[0007] Optionally, the generation of the speech perturbation feature vector set specifically includes: The standardized speech data sequence is segmented according to a preset time window, and divided into multiple consecutive time window speech data segments. Energy change features are extracted for speech data segments in adjacent time windows. Energy statistics are performed on the speech amplitude within each speech data segment in a time window to obtain the corresponding window energy value. The change in window energy value of speech data segments in adjacent time windows is then calculated to generate energy change features. Temporal continuity feature extraction is performed on the speech data segments within the time window. Speech activity indication results are generated based on the window energy value and a preset window energy value threshold. The continuous satisfaction of the speech activity indication results in multiple adjacent speech data segments within the time window is statistically analyzed to obtain the continuous duration of the speech activity, which is used as the temporal continuity feature. The rhythm change features are extracted from the speech data segments within the time window. The rhythm period is calculated for the speech amplitude sequence within each time window speech data segment to obtain the corresponding rhythm period parameters. The change in the rhythm period parameters of adjacent time window speech data segments is calculated to generate rhythm change features. The energy change features, temporal continuity features, and rhythm change features corresponding to the speech data segments within the same time window are combined to construct speech perturbation feature vectors for the corresponding time window. The speech perturbation feature vectors are then aggregated according to the order of the time windows to form a set of speech perturbation feature vectors.

[0008] Optionally, the generation of the voice perturbation state representation specifically includes: The speech perturbation feature vector set is processed by temporal aggregation according to the time window index order to generate a temporal aggregation result vector corresponding to each time window. For each time window, a linear mapping process is performed on the time-series aggregation result vector to generate the corresponding speech disturbance intensity score. For each time window, the speech disturbance intensity score is subjected to grade quantization processing to generate the corresponding disturbance level label; The speech disturbance intensity score and disturbance level identifier corresponding to the same time window are combined to generate a speech disturbance state representation corresponding to that time window.

[0009] Optionally, the disturbance level identifier is used to represent the discrete identifier of the disturbance level range to which the speech disturbance intensity score belongs. Its value distinguishes different levels of speech disturbance states, including no disturbance state, low disturbance state, medium disturbance state and high disturbance state.

[0010] Optionally, determining the current operational status specifically includes: A set of operational status states is pre-built on the smartwatch, including low-power stable status, interactive preparation status and disturbance suppression status, and corresponding operational status judgment rules are configured for each operational status state. Within each time window, the current operating parameters of the smartwatch are obtained and the operating parameters are collected to form the operating parameter set corresponding to the current time window. The operating parameters include processor load status, system power consumption status, voice subsystem working status and interactive module wake-up status. Based on the operational status determination rules, the operational status determination process is performed on the set of operational parameters corresponding to the current time window to generate the operational status determination result corresponding to the current time window; Based on the operational status determination results, the operational status corresponding to the current time window is determined from the operational status status set. When the determination result matches the operational status determination rule corresponding to the low-power stable status, the current operational status is determined to be the low-power stable status. When the determination result matches the operational status determination rule corresponding to the interaction preparation status, the current operational status is determined to be the interaction preparation status. When the determination result matches the operational status determination rule corresponding to the disturbance suppression status, the current operational status is determined to be the disturbance suppression status.

[0011] Optionally, the generation of the operational status transition determination result specifically includes: Obtain the voice disturbance status representation corresponding to the current time window, and obtain the corresponding current operating status. The voice disturbance status representation includes a voice disturbance intensity score and a disturbance level identifier. Based on the current operational status, determine the status stability structure parameters corresponding to the current operational status; Normalize the speech disturbance intensity score, and sequentially perform numerical encoding and normalization on the disturbance level label to generate a set of speech disturbance impact parameters; Based on the set of parameters affected by voice disturbance, dynamic modulation processing is performed on the situation stability structure parameters corresponding to the current operating situation to generate the situation stability structure parameters after the disturbance. Based on the status stabilization structure parameters after the disturbance, the operational status transition judgment result is generated.

[0012] Optionally, the generation of the updated operational status specifically includes: Obtain the operation status migration judgment result and the current operation status corresponding to the current time window, and convert the operation status migration judgment result into a migration indication result; When the migration indication result indicates that the current operational status should be maintained, the updated operational status will be determined to be consistent with the current operational status, and the updated operational status will be output as the operational status corresponding to the current time window. When the migration indication result indicates that an operational status migration needs to be performed, obtain the disturbance level identifier corresponding to the current time window; Based on the current operational status and disturbance level indicators, determine the target operational status; The target's operational status is determined as the updated operational status, and the updated operational status is output as the operational status corresponding to the current time window.

[0013] Optionally, the dynamic determination of the voice data processing path specifically includes: Obtain the updated operational status corresponding to the current time window, and obtain the standardized voice data sequence corresponding to the current time window; Based on the updated operational status, determine the speech data processing method corresponding to the current standardized speech data sequence; When the speech data processing method indicates that speech processing should not be performed, suppression processing is performed on the standardized speech data sequence; When the voice data processing method indicates that the interaction preparation state has been entered, the wake-up state of the interaction module is set to on, and the standardized voice data sequence is cached in the interaction preparation buffer. When the speech data processing method indicates that speech processing should be performed, the speech processing procedure is executed on the standardized speech data sequence. The suppression processing result, interaction preparation result, or speech processing result is used as the speech data processing output corresponding to the current time window to complete the speech data processing.

[0014] A voice data processing system based on a smartwatch according to an embodiment of the present invention includes: The voice acquisition and preprocessing module is used to acquire voice signals during the operation of the smartwatch and preprocess the acquired voice signals to generate standardized voice data sequences. The speech perturbation feature construction module is used to extract speech perturbation-related features from standardized speech data sequences and construct a set of speech perturbation feature vectors. The voice disturbance state generation module is used to generate a voice disturbance state representation; The operational status management module is used to build and maintain the operational status set of the smartwatch and determine the current operational status. The joint situation determination module is used to jointly determine the voice disturbance state representation and the current operational situation state, and generate the operational situation transition determination result. The situation update module is used to maintain or migrate and update the operational situation status based on the operational situation migration judgment results. The voice data processing module is used to determine the processing method for voice data based on the updated operational status and to complete the corresponding voice data processing.

[0015] The beneficial effects of this invention are: This invention transforms voice input from a traditional "command trigger object" into a "source of operational disturbance." Without relying on voice content recognition, it quantifies and models the impact of voice input on the operating state of a smartwatch, achieving a deep integration of voice data processing and system operation status management. By extracting energy change features, temporal continuity features, and rhythm change features to construct a voice disturbance feature vector, and further generating a voice disturbance state representation, this invention can accurately characterize the degree of impact of different voice inputs on system operational stability in both time and intensity dimensions, thereby avoiding misjudgments caused by triggering the processing flow solely based on the presence or absence of voice.

[0016] Based on this, the present invention introduces an operational status set and an operational status transition determination mechanism, which jointly determines the voice disturbance status representation with the current operating parameters of the smartwatch. This makes it so that whether voice input triggers processing is no longer an isolated decision, but is jointly constrained by multiple operating conditions such as system power consumption status, processor load status, voice subsystem working status, and interactive module wake-up status. By precisely controlling the maintenance or transition of the operational status, invalid voice processing requests can be effectively suppressed under conditions of limited system resources or unstable operation, reducing unnecessary power consumption and system load fluctuations, and improving the overall stability of the smartwatch.

[0017] Furthermore, this invention dynamically determines the voice data processing path, enabling differentiated processing strategies for voice data under different operating conditions. In a low-power, stable state, it avoids entering the voice processing flow; in an interactive preparation state, it completes interactive preparation in advance; and in a disturbance suppression state, it restricts or delays voice processing. This ensures a reasonable allocation of resources while guaranteeing user experience. Compared to existing technologies, this invention reduces frequent system state switching caused by invalid voice triggers, improving the robustness and battery life of smartwatches in complex voice environments. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a voice data processing method based on a smartwatch proposed in this invention; Figure 2 This is a schematic diagram illustrating the operational status determination process of a voice data processing method based on a smartwatch proposed in this invention. Figure 3 This is a schematic diagram illustrating the process of determining and updating the operational status of a voice data processing method based on a smartwatch, as proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figures 1-3 A voice data processing method based on a smartwatch includes the following steps: During the operation of the smartwatch, continuous voice signals are collected through the smartwatch's voice acquisition component to generate raw voice data, and preprocessing is performed to generate a standardized voice data sequence. The standardized speech data sequence is segmented and processed. Within each time window, the energy change features, temporal continuity features, and rhythm change features of the speech signal are extracted to construct a speech perturbation feature vector and form a set of speech perturbation feature vectors. Temporal aggregation and linear mapping are performed on the speech disturbance feature vector set to generate speech disturbance intensity scores, and level quantization is used to generate disturbance level labels. The speech disturbance state representation is generated by combining the speech disturbance intensity scores. A set of operational status states is pre-built on the smartwatch, and the current operational status state is determined based on the current operational parameters of the smartwatch; The voice disturbance state representation is jointly determined with the current operational status, and an operational status transition determination result is generated based on the degree of influence of the voice disturbance state representation on the stability of the current operational status. Based on the operational status migration determination result, perform a maintain or migrate operation on the current operational status to obtain the updated operational status. Based on the updated operational status, the processing path for voice data is dynamically determined, and voice data processing is completed.

[0021] In this embodiment, the preprocessing includes time synchronization, outlier removal, and amplitude normalization.

[0022] In this embodiment, the generation of the speech perturbation feature vector set specifically includes: The standardized speech data sequence is segmented according to a preset time window, and divided into multiple consecutive time window speech data segments. Energy change features are extracted for speech data segments in adjacent time windows. Energy statistics are performed on the speech amplitude within each speech data segment in a time window to obtain the corresponding window energy value. The change in window energy value of speech data segments in adjacent time windows is then calculated to generate energy change features. Temporal continuity feature extraction is performed on the speech data segments within the time window. Speech activity indication results are generated based on the window energy value and a preset window energy value threshold. The continuous satisfaction of the speech activity indication results in multiple adjacent speech data segments within the time window is statistically analyzed to obtain the continuous duration of the speech activity, which is used as the temporal continuity feature. The generation of the temporal continuity feature specifically includes: obtaining the window energy value corresponding to each time window speech data segment, and comparing the window energy value with a preset window energy value threshold window by window. When the window energy value is greater than or equal to the preset window energy value threshold, the speech activity indication result of the corresponding time window speech data segment is generated as an active state; when the window energy value is less than the preset window energy value threshold, the speech activity indication result of the corresponding time window speech data segment is generated as an inactive state. The speech activity indication results are arranged according to the time window index order to form a speech activity indication sequence. Then, the continuity of the speech activity indication results of adjacent time window speech data segments in the speech activity indication sequence is statistically analyzed to identify the time window speech data segment intervals that are continuously in an active state, and the number of time window speech data segments contained in each continuous interval is calculated to obtain the corresponding continuous duration of speech activity. The continuous duration of speech activity is used as a temporal continuity feature to characterize the degree of continuous occurrence of speech in the time dimension. The rhythm change features are extracted from the speech data segments within the time window. The rhythm period is calculated for the speech amplitude sequence within each time window speech data segment to obtain the corresponding rhythm period parameters. The change in the rhythm period parameters of adjacent time window speech data segments is calculated to generate rhythm change features. The generation of the rhythm variation features specifically includes: acquiring the speech amplitude sequence within each time window speech data segment, and performing rhythm period calculation on the speech amplitude sequence within the corresponding time window to obtain the rhythm period parameter corresponding to the speech data segment of that time window; calculating the change in the rhythm period parameter corresponding to the speech data segments of adjacent time windows according to the time window index order to obtain the rhythm variation between adjacent time windows; and using the rhythm variation as a rhythm variation feature to characterize the change in speech rhythm. The energy change features, temporal continuity features, and rhythm change features corresponding to the speech data segments within the same time window are combined to construct the speech perturbation feature vector for the corresponding time window. The speech perturbation feature vectors are then aggregated according to the order of the time windows to form a set of speech perturbation feature vectors. In this invention, voice disturbance refers to voice input disturbance factors introduced during the operation of a smartwatch that affect the stability of the system's operating status but do not directly constitute control commands. The time continuity feature is a feature value calculated from the historical window state with the current time window as the statistical endpoint. This value is used as the time continuity feature of the current time window. When there is no preceding time window in the current time window, its time continuity feature value is 0. The continuous duration of the speech activity is obtained by counting the number of speech data segments in the time window that continuously meet the speech activity conditions.

[0023] In this embodiment, the generation of the speech disturbance state representation specifically includes: The speech perturbation feature vector set is subjected to temporal aggregation processing according to the time window index order to generate a temporal aggregation result vector corresponding to each time window. The temporal aggregation processing uses weighted aggregation. A linear mapping process is performed on the time-series aggregation result vector corresponding to each time window to generate a corresponding voice disturbance intensity score. The voice disturbance intensity score is used to quantify the comprehensive impact of voice disturbance on the smartwatch's operating state. For each time window, the speech disturbance intensity score is subjected to grade quantization processing to generate the corresponding disturbance level label; The speech disturbance intensity score and disturbance level identifier corresponding to the same time window are combined to generate a speech disturbance state representation corresponding to that time window.

[0024] In this embodiment, the disturbance level identifier is used to represent the discrete identifier of the disturbance level range to which the speech disturbance intensity score belongs. Its value distinguishes different levels of speech disturbance states, including no disturbance state, low disturbance state, medium disturbance state and high disturbance state.

[0025] In this embodiment, determining the current operational status specifically includes: A set of operational status states is pre-built on the smartwatch, including low-power stable status, interactive preparation status and disturbance suppression status, and corresponding operational status judgment rules are configured for each operational status state. When pre-building a set of operational status states on the smartwatch, a set of operational status state identifiers is defined during the system initialization phase. This set includes a low-power stable status identifier, an interaction preparation status identifier, and a disturbance suppression status identifier. Corresponding operational status judgment rules are configured for the low-power stable status identifier, the interaction preparation status identifier, and the disturbance suppression status identifier, respectively. Each operational status judgment rule corresponds to the processor load state, system power consumption state, voice subsystem working state, and interaction module wake-up state. Each operational status identifier is then associated with its corresponding operational status judgment rule. When configuring corresponding operational status judgment rules for each operational status, firstly, independent status rule identifiers are established for the low-power stable status, interaction preparation status, and disturbance suppression status. Then, using processor load status, system power consumption status, voice subsystem working status, and interaction module wake-up status as rule input parameters, a corresponding parameter judgment value range is set for each status rule identifier, forming an operational status judgment condition set corresponding to that status rule identifier. Each operational status judgment condition set is bound one-to-one with its corresponding status rule identifier to generate corresponding operational status judgment rule entries. The operational status judgment rule entries generated for the low-power stable status, interaction preparation status, and disturbance suppression status are collected and structured for storage, forming an operational status judgment rule set. The low-power stable state refers to the smartwatch operating in a low-resource-occupancy state within the current time window, where the processor load, system power consumption, voice subsystem operation, and interaction module wake-up state all meet the low-power operation requirements, and the system maintains basic monitoring capabilities without entering the voice processing or interaction preparation process. The interaction preparation state refers to the smartwatch having the operating conditions to respond to voice interaction or user operation within the current time window, where the voice subsystem and interaction module are available, and the system resource configuration allows entry into the voice processing or interaction response process. The disturbance suppression state refers to the smartwatch maintaining system stability by limiting or suppressing the voice processing path when it detects voice disturbances affecting system stability within the current time window, and the current system resource or power consumption state does not meet the conditions corresponding to the interaction preparation state. The low-power stable state, interaction preparation state, and disturbance suppression state are clearly distinguished in terms of operational objectives, system resource occupancy, and voice processing behavior. In the low-power stable state, the smartwatch prioritizes maintaining basic operational stability. The processor load and system power consumption are in a low-occupancy range, and both the voice subsystem and the interaction module are inactive, not entering the voice processing or interaction response process. In the interaction preparation state, the smartwatch is in a state where it can respond to user voice or operations. The processor load and system power consumption meet the interaction requirements, and the voice subsystem and interaction module are available, allowing entry into the voice processing or interaction preparation process. In the disturbance suppression state, the smartwatch detects voice disturbances affecting system operational stability, and the current processor load or system power consumption does not meet the conditions for the interaction preparation state. The voice subsystem's operating state or the interaction module's wake-up state is restricted or suppressed, maintaining operational stability by blocking or delaying the voice processing path. Within each time window, the current operating parameters of the smartwatch are obtained and the operating parameters are collected to form the operating parameter set corresponding to the current time window. The operating parameters include processor load status, system power consumption status, voice subsystem working status and interactive module wake-up status. The processor load status refers to the processor resource usage of the smartwatch within the current time window, including the current operating frequency of the processor, the task execution usage ratio, and whether the processor is in a high or low load range. The system power consumption status refers to the energy consumption level of the smartwatch within the current time window, including the instantaneous power consumption level recorded by the power management module, the battery discharge rate, and whether it is in a low power consumption range. The voice subsystem working status refers to the working status of the voice processing-related functional modules in the smartwatch within the current time window, including whether the voice acquisition component is enabled, whether the voice signal processing link is running, and whether the voice processing task is scheduled for execution. The interaction module wake-up status refers to the wake-up status of the display, touch, or feedback modules used for user interaction within the current time window, including whether the interaction module is woken up, whether it is in an active state that can respond to user operations, and whether it is allowed to enter the interaction processing flow. Based on the operational status determination rules, the operational status determination process is performed on the set of operational parameters corresponding to the current time window to generate the operational status determination result corresponding to the current time window; Based on the operational status determination results, the operational status corresponding to the current time window is determined from the operational status status set. When the determination result matches the operational status determination rule corresponding to the low-power stable status, the current operational status is determined to be the low-power stable status. When the determination result matches the operational status determination rule corresponding to the interaction preparation status, the current operational status is determined to be the interaction preparation status. When the determination result matches the operational status determination rule corresponding to the disturbance suppression status, the current operational status is determined to be the disturbance suppression status.

[0026] In this embodiment, the generation of the operational status transition determination result specifically includes: Obtain the voice disturbance status representation corresponding to the current time window, and obtain the corresponding current operating status. The voice disturbance status representation includes a voice disturbance intensity score and a disturbance level identifier. Based on the current operational status, determine the status stability structure parameters corresponding to the current operational status; When determining the situation stability structure parameters corresponding to the current operating situation, the current operating situation state corresponding to the current time window is obtained, and the corresponding situation type identifier is determined according to the current operating situation state; then the stability boundary parameter value corresponding to the situation type identifier is obtained; and the stability boundary parameter value is assigned to the situation stability structure parameters corresponding to the current operating situation. When determining the stability boundary parameter values, firstly, the situation type identifier corresponding to the current operating situation is obtained; then, the stability boundary configuration item bound to the situation type identifier is obtained, which provides the lower and upper bounds of the parameter ranges for the processor load state, system power consumption state, voice subsystem working state, and interactive module wake-up state, respectively; within the current time window, the values ​​of four operating parameters are obtained respectively; then, boundary deviation calculation is performed for each operating parameter. When the operating parameter value falls within the corresponding parameter range, the boundary deviation is set to 0; when the operating parameter value is less than the lower bound of the parameter range, the boundary deviation is set to the difference between the lower bound of the parameter range and the operating parameter value; when the operating parameter value is greater than the upper bound of the parameter range, the boundary deviation is set to the difference between the operating parameter value and the upper bound of the parameter range; the maximum value of the four boundary deviations is taken as the stability boundary parameter value; Normalize the speech disturbance intensity score and perform numerical encoding and normalization on the disturbance level identifier in sequence to generate a speech disturbance impact parameter set. The speech disturbance impact parameter set is used to characterize the disturbance intensity of speech disturbance on the operational status structure within the current time window. Based on the set of parameters affected by voice disturbance, dynamic modulation processing is performed on the situation stability structure parameters corresponding to the current operating situation to generate the situation stability structure parameters after the disturbance. The generation of the situational stability structure parameters after the disturbance specifically includes: obtaining the situational stability structure parameters and voice disturbance impact parameter set corresponding to the current operating situation, and taking a weighted average of the normalized value of the voice disturbance intensity and the normalized value of the disturbance level identifier to obtain the current disturbance modulation amount; obtaining the preset modulation step size value; multiplying the current disturbance modulation amount by the modulation step size value to obtain the boundary contraction amount; subtracting the boundary contraction amount from the situational stability structure parameters to obtain the stability boundary parameter value corresponding to the current time window, and assigning the stability boundary parameter value to 0 when the calculation result is less than 0; and assigning the stability boundary parameter value corresponding to the current time window as the situational stability structure parameter after the disturbance. Based on the status stabilization structure parameters after the disturbance, the operational status transition judgment result is generated; When generating the operational status transition judgment result, the stable structure parameters of the status after the disturbance corresponding to the current time window are obtained; then the current operational status is obtained; next, the stable structure parameters of the status after the disturbance are compared with the stability boundary range corresponding to the current operational status. When the stable structure parameters of the status after the disturbance exceed the stability boundary range, the operational status transition judgment result is "migration"; when the stable structure parameters of the status after the disturbance are within the stability boundary range, the operational status transition judgment result is "maintain". Finally, the operational status transition judgment result is output as the operational status transition judgment result corresponding to the current time window.

[0027] In this embodiment, the generation of the updated operational status specifically includes: Obtain the operational status migration judgment result and the current operational status corresponding to the current time window, and convert the operational status migration judgment result into a migration indication result to indicate whether a migration operation needs to be performed on the current operational status. When the migration indication result indicates that the current operational status should be maintained, the updated operational status will be determined to be consistent with the current operational status, and the updated operational status will be output as the operational status corresponding to the current time window. When the migration indication result indicates that an operational status migration needs to be performed, obtain the disturbance level identifier corresponding to the current time window; Based on the current operational status and disturbance level indicators, determine the target operational status; When determining the target operational status, the current operational status corresponding to the current time window is obtained, along with the disturbance level identifier corresponding to the same time window. Then, based on the current operational status, the migration rule corresponding to the current operational status is located. Next, within the migration rule, a uniquely matching target status entry is determined based on the value of the disturbance level identifier. This target status entry explicitly records the target operational status corresponding to the disturbance level identifier. The target operational status is read from the target status entry as the determination result. Finally, the target operational status is output as the target operational status corresponding to the current time window. When locating the migration rule corresponding to the current operational status, the status identifier of the current operational status is obtained; then, the operational status identifiers associated with each migration rule are matched one by one; when the operational status identifier recorded in a certain migration rule is consistent with the status identifier of the current operational status, the migration rule is determined to be the migration rule corresponding to the current operational status. The migration rule is a deterministic rule record, which is used to uniquely determine the corresponding target operational status given the current operational status and the disturbance level identifier. When setting migration rules, the set of operational status states and the set of disturbance level identifiers are obtained. Then, the set of operational status states is used as the first-dimensional index and the set of disturbance level identifiers is used as the second-dimensional index to construct a rule structure for the operational status migration mapping relationship. Next, for each operational status state, each disturbance level identifier is traversed, and a target operational status state is uniquely assigned to each combination of operational status state and disturbance level identifier. Then, the operational status state, disturbance level identifier, and corresponding target operational status state are written into the rule structure in a one-to-one correspondence to form multiple migration rules. All migration rules are collected and stored to form a complete operational status migration mapping relationship. The target's operational status is determined as the updated operational status, and the updated operational status is output as the operational status corresponding to the current time window.

[0028] In this embodiment, the dynamic determination of the voice data processing path specifically includes: Obtain the updated operational status corresponding to the current time window, and obtain the standardized voice data sequence corresponding to the current time window; Based on the updated operational status, determine the speech data processing method corresponding to the current standardized speech data sequence; When determining the speech data processing method corresponding to the current standardized speech data sequence, the updated operational status corresponding to the current time window is obtained; it is determined whether the updated operational status is a low-power stable status, an interaction preparation status, or a disturbance suppression status; when the updated operational status is a low-power stable status, it is determined that no speech processing will be performed on the current standardized speech data sequence; when the updated operational status is an interaction preparation status, it is determined that interaction preparation will be performed on the current standardized speech data sequence; when the updated operational status is a disturbance suppression status, it is determined that speech processing will be performed on the current standardized speech data sequence. When the speech data processing mode indicates that speech processing is not to be performed, suppression processing is performed on the standardized speech data sequence. The suppression processing includes preventing the standardized speech data sequence from entering the subsequent speech processing flow. When the voice data processing method indicates that the interaction preparation state has been entered, the wake-up state of the interaction module is set to on, and the standardized voice data sequence is cached in the interaction preparation buffer. When the speech data processing method indicates that speech processing should be performed, the speech processing procedure is executed on the standardized speech data sequence. The suppression processing result, interaction preparation result, or speech processing result is used as the speech data processing output corresponding to the current time window to complete the speech data processing.

[0029] A voice data processing system based on a smartwatch, comprising: The voice acquisition and preprocessing module is used to acquire voice signals during the operation of the smartwatch and preprocess the acquired voice signals to generate standardized voice data sequences. The speech perturbation feature construction module is used to extract speech perturbation-related features from standardized speech data sequences and construct a set of speech perturbation feature vectors. The voice disturbance state generation module is used to generate a voice disturbance state representation; The operational status management module is used to build and maintain the operational status set of the smartwatch and determine the current operational status. The joint situation determination module is used to jointly determine the voice disturbance state representation and the current operational situation state, and generate the operational situation transition determination result. The situation update module is used to maintain or migrate and update the operational situation status based on the operational situation migration judgment results. The voice data processing module is used to determine the processing method for voice data based on the updated operational status and to complete the corresponding voice data processing.

[0030] Example 1: To verify the feasibility of the present invention in practice, it was applied to the voice data processing scenario of a smartwatch in a real daily wearing environment. The smartwatch is worn on the user's wrist for extended periods and operates continuously in various states such as commuting, working, resting, and outdoor activities. In actual use, the smartwatch frequently enters a low-power monitoring state. The surrounding environment contains a large amount of non-command voice, environmental noise, and unintentional voice from the user. Under existing technology, these voice inputs are easily recognized as valid voice, thus repeatedly triggering voice processing or interaction preparation processes, leading to increased system power consumption, processor load fluctuations, and frequent switching of operating states, which seriously affects the device's battery life and operational stability. The application scenario of this example is precisely aimed at the above problems, reasonably distinguishing and controlling the impact of voice input on the system's operating state without relying on voice content recognition.

[0031] In this scenario, the smartwatch continuously collects environmental voice signals during operation. After preprocessing the collected voice signals, it forms a standardized voice data sequence. The system does not directly determine whether the voice is a command. Instead, it divides the standardized voice data sequence into continuous time windows and extracts energy change features, temporal continuity features, and rhythm change features from the voice signals within each time window to construct a voice disturbance feature vector. By performing temporal aggregation and linear mapping on the voice disturbance feature vector, a voice disturbance intensity score is generated and further quantified into a disturbance level identifier, forming a voice disturbance state representation. Simultaneously, the smartwatch determines its own operating status based on the current processor load status, system power consumption status, voice subsystem operating status, and interaction module wake-up status. When the system detects that the voice disturbance state representation affects the stability of the current operating status, it does not immediately trigger voice processing. Instead, it generates an operating status transition judgment result through a joint judgment mechanism. Based on the judgment result, it performs a maintain or transition operation on the operating status. In the updated operating status, the system dynamically determines the voice data processing path, choosing to suppress voice processing, enter interaction preparation, or execute voice processing, thereby achieving fine-grained control of voice input.

[0032] In practical applications, by comparing the changes in the operating state of smartwatches over continuous time periods, it can be observed that this invention can effectively reduce the frequent switching of system states caused by meaningless voice input. Within the same usage time, the number of voice processing triggers is significantly reduced in office environments, commuting environments, and daily wear conditions, resulting in a more stable system operation and effective control of power consumption fluctuations. Especially in situations with complex ambient voices and where the user does not actively engage in voice interaction, this invention can continuously maintain a low-power, stable state, avoiding the occupation of system resources by invalid voice input.

[0033] To verify the performance of the present invention, it was compared with the traditional method. The comparison results are shown in Table 1.

[0034] Table 1. Overall Performance Comparison of the Invention and Traditional Voice Triggering Methods Comparison indicators Traditional voice triggering processing methods This invention Daily voice processing trigger count (times / day) 480 165 Percentage of invalid voice triggers (%) 62.5 28.2 Average processor load fluctuation (%) 23.4 9.1 Daily voice-related power consumption percentage (%) 28.6 14.3 Number of operational status transitions (times / day) 310 97 Percentage of time spent in a stable low-power state (%) 41.2 68.9 Interactive response availability (%) 94.6 96.1 Continuous operation stability score (out of 100) 71 88 The number of voice processing triggers per day shows that traditional voice triggering methods are highly sensitive to environmental voice during smartwatch operation, with up to 480 triggers per day. Many of these triggers are not user-initiated interactions. However, after adopting the method of this invention, the number of voice processing triggers per day decreased to 165, a significant reduction. This indicates that the present invention effectively suppresses the direct triggering effect of meaningless voice input on the system through voice disturbance modeling and joint determination of operational status.

[0035] The data on the percentage of invalid voice triggers further confirms this point. In traditional methods, the percentage of invalid voice triggers reaches 62.5%, meaning that more than half of the voice processing behavior does not generate actual interactive value, resulting in a waste of system resources. In contrast, the percentage of invalid voice triggers in this invention is reduced to 28.2%, indicating that voice input has undergone comprehensive screening based on disturbance intensity and operational status before entering the processing path. Only voice that has a real impact on system operation enters the processing flow. This change directly stems from the technical approach of this invention, which does not treat voice as an instruction object but rather as an operational disturbance for quantitative evaluation.

[0036] In terms of system load and power consumption, this invention also shows significant advantages. Under traditional methods, frequent triggering of voice processing leads to processor load fluctuations of up to 23.4%, with power consumption accounting for nearly 30%. However, in the method of this invention, the average processor load fluctuation is reduced to 9.1%, and the proportion of voice-related power consumption is reduced to 14.3%. This indicates that by limiting or suppressing the voice processing path under low-power stable conditions and disturbance suppression conditions, this invention effectively reduces the continuous occupation of system resources by voice processing, keeping the processor and power consumption in a more stable range.

[0037] The number of operational status transitions is a key indicator of system stability. Traditional methods, due to voice input directly triggering the processing flow, result in frequent transitions between different operational statuses, reaching up to 310 transitions per day. This invention, through a situational stability structural parameter modulation and migration judgment mechanism, transforms voice disturbances into impacts on stability boundaries. Migration only occurs when stability conditions are violated, reducing the number of status transitions to 97. Correspondingly, the proportion of low-power stable status maintenance time increases from 41.2% to 68.9%, allowing the system to remain in a stable operating range for extended periods, further demonstrating the significant effectiveness of operational status management.

[0038] In terms of user experience-related metrics, this invention does not sacrifice interactivity. The interaction response availability rate remains at 96.1% under the method of this invention, slightly higher than the 94.6% of the traditional method. This indicates that while suppressing invalid voice, scenarios that truly require voice interaction can still be responded to in a timely manner. The overall continuous operation stability score has increased from 71 to 88, demonstrating a significant enhancement in overall reliability and consistency under long-term operation.

[0039] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A voice data processing method based on a smartwatch, characterized in that, Includes the following steps: During the operation of the smartwatch, continuous voice signals are collected through the smartwatch's voice acquisition component to generate raw voice data, and preprocessing is performed to generate a standardized voice data sequence. The standardized speech data sequence is segmented and processed. Within each time window, the energy change features, temporal continuity features, and rhythm change features of the speech signal are extracted to construct a speech perturbation feature vector and form a set of speech perturbation feature vectors. Temporal aggregation and linear mapping are performed on the speech disturbance feature vector set to generate speech disturbance intensity scores, and level quantization is used to generate disturbance level labels. The speech disturbance state representation is generated by combining the speech disturbance intensity scores. A set of operational status states is pre-built on the smartwatch, and the current operational status state is determined based on the current operational parameters of the smartwatch; The voice disturbance state representation is jointly determined with the current operational status, and an operational status transition determination result is generated based on the degree of influence of the voice disturbance state representation on the stability of the current operational status. Based on the operational status migration determination result, perform a maintain or migrate operation on the current operational status to obtain the updated operational status. Based on the updated operational status, the processing path for voice data is dynamically determined, and voice data processing is completed.

2. The voice data processing method based on a smartwatch according to claim 1, characterized in that, The preprocessing includes time synchronization, outlier removal, and amplitude normalization.

3. The voice data processing method based on a smartwatch according to claim 1, characterized in that, The generation of the speech perturbation feature vector set specifically includes: The standardized speech data sequence is segmented according to a preset time window, and divided into multiple consecutive time window speech data segments. Energy change features are extracted for speech data segments in adjacent time windows. Energy statistics are performed on the speech amplitude within each speech data segment in a time window to obtain the corresponding window energy value. The change in the window energy value of speech data segments in adjacent time windows is then calculated to generate energy change features. Temporal continuity feature extraction is performed on the speech data segments within the time window. Speech activity indication results are generated based on the window energy value and a preset window energy value threshold. The continuous satisfaction of the speech activity indication results in multiple adjacent speech data segments within the time window is statistically analyzed to obtain the continuous duration of the speech activity, which is used as the temporal continuity feature. Rhythm variation feature extraction is performed on the speech data segments within the time window. Rhythm cycle calculation is performed on the speech amplitude sequence within each time window speech data segment to obtain the corresponding rhythm cycle parameters. The change in rhythm cycle parameters of adjacent time window speech data segments is calculated to generate rhythm variation features. The energy change features, temporal continuity features, and rhythm change features corresponding to the speech data segments within the same time window are combined to construct speech perturbation feature vectors for the corresponding time window. The speech perturbation feature vectors are then aggregated according to the order of the time windows to form a set of speech perturbation feature vectors.

4. The voice data processing method based on a smartwatch according to claim 1, characterized in that, The generation of the speech perturbation state representation specifically includes: The speech perturbation feature vector set is processed by temporal aggregation according to the time window index order to generate a temporal aggregation result vector corresponding to each time window. For each time window, a linear mapping process is performed on the time-series aggregation result vector to generate the corresponding speech disturbance intensity score. For each time window, the speech disturbance intensity score is subjected to grade quantization processing to generate the corresponding disturbance level label; The speech disturbance intensity score and disturbance level identifier corresponding to the same time window are combined to generate a speech disturbance state representation corresponding to that time window.

5. The voice data processing method based on a smartwatch according to claim 4, characterized in that, The disturbance level identifier is used to represent the discrete identifier of the disturbance level range to which the speech disturbance intensity score belongs. Its value distinguishes different levels of speech disturbance states, including no disturbance state, low disturbance state, medium disturbance state and high disturbance state.

6. The voice data processing method based on a smartwatch according to claim 1, characterized in that, Determining the current operational status specifically includes: A set of operational status states is pre-built on the smartwatch, including low-power stable status, interactive preparation status and disturbance suppression status, and corresponding operational status judgment rules are configured for each operational status state. Within each time window, the current operating parameters of the smartwatch are obtained and the operating parameters are collected to form the operating parameter set corresponding to the current time window. The operating parameters include processor load status, system power consumption status, voice subsystem working status and interactive module wake-up status. Based on the operational status determination rules, the operational status determination process is performed on the set of operational parameters corresponding to the current time window to generate the operational status determination result corresponding to the current time window; Based on the operational status determination results, the operational status corresponding to the current time window is determined from the operational status status set. When the determination result matches the operational status determination rule corresponding to the low-power stable status, the current operational status is determined to be the low-power stable status. When the determination result matches the operational status determination rule corresponding to the interaction preparation status, the current operational status is determined to be the interaction preparation status. When the determination result matches the operational status determination rule corresponding to the disturbance suppression status, the current operational status is determined to be the disturbance suppression status.

7. The voice data processing method based on a smartwatch according to claim 1, characterized in that, The generation of the operational status transition determination result specifically includes: Obtain the voice disturbance status representation corresponding to the current time window, and obtain the corresponding current operating status. The voice disturbance status representation includes a voice disturbance intensity score and a disturbance level identifier. Based on the current operational status, determine the status stability structure parameters corresponding to the current operational status; Normalize the speech disturbance intensity score, and sequentially perform numerical encoding and normalization on the disturbance level label to generate a set of speech disturbance impact parameters; Based on the set of parameters affected by voice disturbance, dynamic modulation processing is performed on the situation stability structure parameters corresponding to the current operating situation to generate the situation stability structure parameters after the disturbance. Based on the status stabilization structure parameters after the disturbance, the operational status transition judgment result is generated.

8. The voice data processing method based on a smartwatch according to claim 1, characterized in that, The generation of the updated operational status specifically includes: Obtain the operation status migration judgment result and the current operation status corresponding to the current time window, and convert the operation status migration judgment result into a migration indication result; When the migration indication result indicates that the current operational status should be maintained, the updated operational status will be determined to be consistent with the current operational status, and the updated operational status will be output as the operational status corresponding to the current time window. When the migration indication result indicates that an operational status migration needs to be performed, obtain the disturbance level identifier corresponding to the current time window; Based on the current operational status and disturbance level indicators, determine the target operational status; The target's operational status is determined as the updated operational status, and the updated operational status is output as the operational status corresponding to the current time window.

9. The voice data processing method based on a smartwatch according to claim 1, characterized in that, The dynamic determination of the voice data processing path specifically includes: Obtain the updated operational status corresponding to the current time window, and obtain the standardized voice data sequence corresponding to the current time window; Based on the updated operational status, determine the speech data processing method corresponding to the current standardized speech data sequence; When the speech data processing method indicates that speech processing should not be performed, suppression processing is performed on the standardized speech data sequence; When the voice data processing method indicates that the interaction preparation state has been entered, the wake-up state of the interaction module is set to on, and the standardized voice data sequence is cached in the interaction preparation buffer. When the speech data processing method indicates that speech processing should be performed, the speech processing procedure is executed on the standardized speech data sequence. The suppression processing result, interaction preparation result, or speech processing result is used as the speech data processing output corresponding to the current time window to complete the speech data processing.

10. A voice data processing system based on a smartwatch, executing the voice data processing method based on a smartwatch as described in any one of claims 1 to 9, characterized in that, include: The voice acquisition and preprocessing module is used to acquire voice signals during the operation of the smartwatch and preprocess the acquired voice signals to generate standardized voice data sequences. The speech perturbation feature construction module is used to extract speech perturbation-related features from standardized speech data sequences and construct a set of speech perturbation feature vectors. The voice disturbance state generation module is used to generate a voice disturbance state representation; The operational status management module is used to build and maintain the operational status set of the smartwatch and determine the current operational status. The joint situation determination module is used to jointly determine the voice disturbance state representation and the current operational situation state, and generate the operational situation transition determination result. The situation update module is used to maintain or migrate and update the operational situation status based on the operational situation migration judgment results. The voice data processing module is used to determine the processing method for voice data based on the updated operational status and to complete the corresponding voice data processing.