Dialogue type human-computer interaction method and system applied to smart wearable device

By analyzing emotional fluctuations and text similarity to select the optimal matching speech for noise reduction, the problem of noise interference in dynamic environments for smartwatches is solved, improving the accuracy of semantic recognition and the reliability of conversational human-computer interaction.

CN122369447APending Publication Date: 2026-07-10SHENZHEN XINGUAN PRECISION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN XINGUAN PRECISION TECH CO LTD
Filing Date
2026-04-23
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Smartwatches are affected by complex noises such as wind noise and human voice interference in dynamic environments. Traditional noise reduction algorithms have difficulty effectively separating the target speech from the interference components, resulting in low semantic understanding accuracy and affecting the reliability of conversational human-computer interaction.

Method used

By analyzing the differences in emotional fluctuations and text similarity between the current input speech and historical input speech, and combining the degree of environmental noise interference and environmental differences, the optimal matching speech is selected for denoising. Context matching is performed using two dimensions of emotional fluctuation and text similarity analysis to identify historical input speech related to the context of the current input speech, and an environmental consistency model is constructed to improve the effectiveness of denoising.

Benefits of technology

It significantly improves the semantic recognition accuracy of the current input speech, enhances the coherence and reliability of continuous dialogue-based human-computer interaction, ensures that the historical input speech and the current input speech are in a similar noise field during the denoising process, avoids single index bias, and improves the accuracy of semantic understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369447A_ABST
    Figure CN122369447A_ABST
Patent Text Reader

Abstract

This invention relates to the field of voice interaction technology, specifically to a conversational human-computer interaction method and system applied to smart wearable devices. The method includes: determining the matching degree between the current input voice and each historical input voice based on the differences in emotional fluctuations and text similarity between the current input voice and each historical input voice, thereby obtaining a matched voice; determining the degree of interference from environmental noise on the current input voice and each matched voice; determining the environmental similarity between the current input voice and each matched voice based on the differences in the environment and interference levels of the environments in which the current input voice and each matched voice are located; combining the matching degree to obtain the final matching degree between the current input voice and each matched voice, thereby obtaining the optimal matched voice; and denoising the current input voice based on the optimal matched voice, which can significantly improve the semantic recognition accuracy of the current input voice and enhance the coherence and reliability of continuous conversational human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice interaction technology, and more specifically to a conversational human-computer interaction method and system for smart wearable devices. Background Technology

[0002] Smartwatches, as a common type of wearable smart device, typically connect to users' smartphones via Bluetooth, allowing users to interact via voice through their built-in microphones to perform functions such as checking the weather, managing schedules, and controlling music. However, because smartwatches are worn on the wrist, their microphones are far from the user's mouth, and in actual use, they are often in dynamic environments (such as outdoor walking or sports), making them susceptible to combined noises such as wind noise, human voice interference, and clothing friction. More importantly, these noises are highly time-varying and individualized (for example, the friction sound patterns of the same user differ under different movement states), making it difficult for traditional noise reduction algorithms based on fixed models or general statistical characteristics to effectively separate the target speech from interference components. Furthermore, existing speech recognition systems typically treat each voice input as an independent event, ignoring the potential correlation between the user's speech content, emotional state, and environment in continuous dialogue. This fragmented processing significantly affects the semantic understanding accuracy of the current input speech, thereby impacting the reliability of conversational human-computer interaction. Summary of the Invention

[0003] To address the technical problem of low semantic recognition accuracy in existing noise reduction methods, which affects the reliability of conversational human-computer interaction, the present invention aims to provide a conversational human-computer interaction method and system for smart wearable devices. The specific technical solution adopted is as follows: In a first aspect of the present invention, a conversational human-computer interaction method for use in a smart wearable device is provided, comprising: Determine the differences in emotional fluctuations between the current input speech and each historical input speech, and combine this with text similarity to obtain the degree of matching between the current input speech and each historical input speech; The degree of interference from environmental noise on the current input speech and each matched speech is determined respectively; the matched speech is obtained by filtering from each historical input speech based on the degree of matching. Determine the degree of environmental difference between the current input speech and each matched speech, and combine the difference in interference levels to obtain the environmental similarity between the current input speech and each matched speech; By combining the matching degree and environmental similarity, the final matching degree between the current input speech and each matched speech is obtained; The current input speech is denoised based on the optimal matching speech; the optimal matching speech is selected from each matching speech according to the final matching degree.

[0004] In an exemplary embodiment, the process of obtaining the emotional fluctuation difference includes: Determine the reference semantic units of each current semantic unit in the current input speech in each historical input speech; The emotional fluctuation level of candidate semantic units is determined based on tone changes; the candidate semantic units are any semantic units between the current semantic unit and the reference semantic unit. The difference in emotional fluctuation between the current input speech and each historical input speech is obtained by comparing the emotional fluctuation levels of each current semantic unit in the current input speech with those of the reference semantic units in each historical input speech.

[0005] In an exemplary embodiment, the process of obtaining the reference semantic unit includes: Convert the current semantic unit and each historical semantic unit in the historical input speech into TF-IDF vectors respectively; The vector similarity between the TF-IDF vector of the current semantic unit and the TF-IDF vector of each historical semantic unit in the historical input speech is determined. The historical semantic unit corresponding to the largest vector similarity is used as the reference semantic unit of the current semantic unit in the historical input speech.

[0006] In an exemplary embodiment, the process of obtaining the degree of emotional fluctuation includes: Determine the pitch value difference between adjacent pitches in candidate semantic units; The number of emotional fluctuations in candidate semantic units is obtained from the difference in pitch values; The emotional fluctuation level of a candidate semantic unit is obtained from the maximum pitch difference, the number of emotional fluctuations, and the signal-to-noise ratio of the candidate semantic unit; the emotional fluctuation level is positively correlated with the maximum pitch difference, the number of emotional fluctuations, and the signal-to-noise ratio.

[0007] In an exemplary embodiment, the process of obtaining text similarity cases includes: By fusing the text similarity between each current semantic unit of the current input speech and the reference semantic units in each historical input speech, the text similarity between the current input speech and each historical input speech is obtained.

[0008] In an exemplary embodiment, the process of obtaining the interference level includes: The candidate input speech is decomposed into a main speech component and a reference speech component; the candidate input speech is any input speech between the current input speech and each matched speech. Determine the amplitude difference between the candidate input speech and the main speech component at the same time to obtain the noise interference amplitude at each time. The environmental interference time and duration in the candidate input speech are obtained from the noise interference amplitude. The degree of environmental noise interference on the candidate input speech is obtained based on the number of environmental interference moments in the candidate input speech, the duration of each environmental interference moment, and the maximum noise interference amplitude of each environmental interference moment; the degree of interference is positively correlated with the number of environmental interference moments, the duration of each environmental interference moment, and the maximum noise interference amplitude.

[0009] In an exemplary embodiment, the step of decomposing the candidate input speech into a primary speech component and a reference speech component includes: The candidate input speech is decomposed into multiple speech components; wherein, the main speech component is the speech component with the largest contribution, and the reference speech components are all other speech components except the main speech component.

[0010] In one exemplary embodiment, the process of obtaining the degree of environmental difference includes: Obtain the difference distance between each reference speech component of the current input speech and each reference speech component of the matching speech, determine the minimum difference distance, and obtain the environmental difference representation between each reference speech component of the current input speech and the matching speech. By integrating the environmental differences between the various reference speech components of the current input speech and the matched speech, the degree of environmental difference between the current input speech and the matched speech is obtained.

[0011] In an exemplary embodiment, the denoising of the current input speech based on the optimal matching speech includes: Extract the noise distribution features of the current input speech and the spectral envelope features of the best matching speech; The noise distribution features and spectral envelope features are used as prior constraints and input into a preset denoising algorithm along with the current input speech, and the denoised current input speech is output.

[0012] In a second aspect of the present invention, a conversational human-computer interaction system for smart wearable devices is provided, comprising: a memory and a processor; the memory is connected to the processor; the memory is used to store program instructions; the processor is used to implement the above-described conversational human-computer interaction method for smart wearable devices when the program instructions are executed.

[0013] This invention has the following beneficial effects: It introduces a two-dimensional analysis of emotional fluctuations and text similarity to perform context matching, determining the degree of matching between the current input speech and each historical input speech, identifying historical input speech related to the context of the current input speech, avoiding irrelevant historical input speech (such as different topics or sudden emotional changes) as subsequent references, and improving the accuracy of matching speech acquisition; by determining the degree of interference from environmental noise on the current input speech and matching speech, the impact of noise is quantified, providing a basis for subsequent environmental analysis; based on the environmental differences between the current input speech and each matching speech, and the differences in the degree of interference from environmental noise, the environmental similarity between the current input speech and each matching speech is obtained, constructing an environment-based analysis system. Consistency models (such as whether both are outdoors or both are exercising) can solve the problem of denoising failure caused by the same semantics but different environmental noise characteristics. This ensures that the historical input speech used for denoising is in a similar noise field as the current input speech, thus improving the effectiveness of denoising. By combining environmental similarity and matching degree analysis in multiple dimensions, the final matching degree between the current input speech and each matching speech is obtained, avoiding the bias of a single indicator and obtaining the optimal matching speech that best matches the current input speech. Finally, the optimal matching speech with high correlation and high environmental consistency is used as a priori to guide the noise suppression of the current input speech, which can significantly improve the semantic recognition accuracy of the current input speech and enhance the coherence and reliability of continuous conversational human-computer interaction. Attached Figure Description

[0014] Figure 1 This is a flowchart of a conversational human-computer interaction method applied to a smart wearable device, provided by an embodiment of the present invention; Figure 2 This is a flowchart of the process for obtaining differences in emotional fluctuations provided in one embodiment of the present invention; Figure 3 This is a flowchart illustrating the process of obtaining the level of interference according to an embodiment of the present invention. Detailed Implementation

[0015] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the specific implementation methods, structures, features, and effects of the present invention are described in detail below with reference to the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. All data and information collected in this application have been obtained with full consent.

[0017] This embodiment provides a conversational human-computer interaction method for smart wearable devices, taking a smartwatch as an example. This method filters the voice data that best matches the current input voice from multiple historical input voices of the user, thus assisting in noise reduction of the current input voice. The smartwatch has a built-in microphone, processor, and storage module for voice acquisition and processing. The data processing provided in this embodiment can be executed by the user's smartphone.

[0018] When interacting with a smartwatch, users often use common commands or phrases, which may appear multiple times at different times and in different scenarios. Therefore, this embodiment can analyze the user's previous historical voice inputs to identify the relevant features of repeated commands, thereby removing noise from the current voice input.

[0019] The smartwatch's built-in storage module can save recent historical voice input, such as 7-30 days of historical voice input. This allows the smartwatch to acquire all of the user's voice input, including the current input and multiple historical inputs. To improve the accuracy and reliability of historical voice input selection, the number of days for acquiring historical voice input can be longer, ensuring a larger quantity of historical voice input. This embodiment can perform time-frequency alignment between the current input and each historical input, for example, using a fast dynamic time warping algorithm to achieve millisecond-level time-frequency alignment between the current input and each historical input.

[0020] Since the signal-to-noise ratio of the current input speech and each historical input speech is low (there is a lot of noise), this embodiment can first use a simple bandpass filter to filter out the extremely low / extreme high frequency noise.

[0021] This embodiment uses speech recognition technology to obtain the text content of the current input speech and each historical input speech. Speech recognition technology is a well-known technique, and its specific methods are not described here. Furthermore, this embodiment can also use a forced alignment algorithm to map each character in the text to the time axis of the speech signal, determining the start time of each character. Additionally, this embodiment can use fractional scaling normalization to process the current input speech and each historical input speech, making them have similar scales and distributions for better comparison. Fractional scaling normalization is a well-known technique, and its specific methods are not described here.

[0022] Starting from the moment the smartwatch begins using the voice function, in the initial stage of the solution provided in this embodiment, the number of historical input voices is very small, or even non-existent (for example, there are no other historical input voices before the first historical input voice in the time sequence, or the second historical input voice in the time sequence is preceded by only one historical input voice, resulting in a very small sample size). In this case, this embodiment uses an existing general denoising model (such as spectral subtraction) to denoise the user's current input voice. As the number of user input voices gradually increases, the denoising process provided in this embodiment is switched to. As an example, this embodiment can start executing the denoising process provided in this embodiment from the 11th input voice in the time sequence.

[0023] like Figure 1 As shown, the conversational human-computer interaction method for smart wearable devices provided in this embodiment includes the following steps: Step S1: Determine the differences in emotional fluctuations between the current input speech and each historical input speech, and combine this with text similarity to obtain the degree of matching between the current input speech and each historical input speech; Step S2: Determine the degree of interference from environmental noise on the current input speech and each matched speech; Step S3: Determine the degree of environmental difference between the current input speech and each matched speech, and combine the differences in interference levels to obtain the environmental similarity between the current input speech and each matched speech; Step S4: Combine the matching degree and environmental similarity to obtain the final matching degree between the current input speech and each matched speech; Step S5: Denoise the current input speech based on the optimal matching speech.

[0024] The following detailed explanation of each step, in conjunction with the accompanying drawings, is provided.

[0025] Step S1: Determine the difference in emotional fluctuation between the current input speech and each historical input speech, and combine this with the text similarity to obtain the degree of matching between the current input speech and each historical input speech.

[0026] This step analyzes both the emotional fluctuations represented in the speech and the text information, obtaining the differences in emotional fluctuations between the current input speech and each historical input speech, as well as the text similarity between the current input speech and each historical input speech. Combining these two aspects, the matching degree between the current input speech and each historical input speech is determined. The difference in emotional fluctuations refers to the difference in emotional fluctuations represented by the current input speech and the historical input speech. In an exemplary embodiment, such as... Figure 2 As shown, the following is one process for obtaining differences in emotional fluctuations: Step S11: Determine the reference semantic units in each historical input speech for each current semantic unit in the current input speech.

[0027] This embodiment employs a sentence segmentation algorithm to divide the current input speech and each historical input speech into several semantic units. Each current input speech and each historical input speech is divided into at least one semantic unit. Specifically, each semantic unit obtained from the segmentation of the current input speech is defined as a current semantic unit, and each semantic unit obtained from the segmentation of each historical input speech is defined as a historical semantic unit. It should be understood that each speech unit contains at least two characters.

[0028] For any given semantic unit in the current input speech, this embodiment needs to determine the reference semantic units in each historical input speech. The reference semantic units are those that are more relevant to the current semantic unit. In an exemplary embodiment, for any historical input speech, the current semantic unit and each historical semantic unit in the historical input speech are converted into TF-IDF vectors (specifically, the text corresponding to the semantic unit is converted into a TF-IDF vector). Then, the vector similarity between the TF-IDF vector of the current semantic unit and the TF-IDF vectors of each historical semantic unit in the historical input speech is determined. This vector similarity is specifically cosine similarity. Since the numerical range of cosine similarity is -1 to 1, the cosine similarity is normalized using the formula: (cosine similarity + 1) / 2. All cosine similarities mentioned below are the normalized results. This cosine similarity represents the text similarity between the TF-IDF vector of the current semantic unit and the TF-IDF vector of the historical semantic units. The higher the cosine similarity value, the more similar the TF-IDF vector of the current semantic unit is to the corresponding historical semantic unit. Therefore, the highest cosine similarity is determined from the cosine similarities between the TF-IDF vector of the current semantic unit and the TF-IDF vectors of all historical semantic units in the historical input speech. The historical semantic unit corresponding to this highest cosine similarity is used as the reference semantic unit of the current semantic unit in that historical input speech. This yields the reference semantic units of the current semantic unit in each historical input speech, and consequently, the reference semantic units of each current semantic unit in each historical input speech.

[0029] Step S12: Determine the degree of emotional fluctuation of candidate semantic units based on tone changes.

[0030] Since this embodiment needs to obtain the emotional fluctuation level of each current semantic unit in the current input speech and each reference semantic unit in the historical input speech, any one of the current semantic units in the current input speech and each reference semantic unit in the historical input speech is defined as a candidate semantic unit.

[0031] The process involves retrieving individual characters from the text recognition results of candidate semantic units and mapping them to the corresponding audio segments. For any given character, the corresponding audio segment is divided into short frames (frame length can be set to 20-40 milliseconds). For each audio frame, the fundamental frequency (FFF) is calculated using either time-domain or frequency-domain methods. The FFF is the basic frequency of speech, reflecting its pitch variations. By arranging these frames sequentially, a FFF sequence is obtained. The average FFF of this sequence is then calculated and used as the pitch value of the character. This process yields the pitch values ​​for each character within the candidate semantic units, accurately reflecting the pitch variations during pronunciation.

[0032] Based on the tonal variations between the tonal values ​​of each character in a candidate semantic unit, the degree of emotional fluctuation in the candidate semantic unit is determined. In an exemplary embodiment, the tonal value difference between two adjacent tones in a candidate semantic unit is obtained, i.e., the tonal value difference between two adjacent characters. The tonal value difference can be calculated as follows: calculate the absolute value of the difference between the tonal values ​​of two adjacent characters, then determine the maximum tonal value among the tonal values ​​of each character in the candidate semantic unit, and calculate the ratio of the absolute value of the difference between the tonal values ​​of two adjacent characters to the maximum tonal value. This normalizes the absolute value of the difference between the tonal values ​​of two adjacent characters, and the result is taken as the tonal value difference between two adjacent characters, thus obtaining the tonal value difference between each pair of adjacent tones in the candidate semantic unit. The larger the tonal value difference, the greater the change in tone and emotion between the two adjacent tones, i.e., the greater the emotional fluctuation.

[0033] This embodiment presets a tone difference threshold, which is used to compare the tone difference between any two adjacent characters in the candidate semantic unit, thereby determining the tone difference that is greater than or equal to the tone difference threshold. It should be understood that the tone difference threshold ranges from 0 to 1, and the specific value is set according to actual needs; this embodiment uses 0.2 as an example.

[0034] Determine the pitch difference values ​​in candidate semantic units that are greater than or equal to a pitch difference threshold, obtain the number of pitch difference values ​​in candidate semantic units that are greater than or equal to the pitch difference threshold, and use this number as the emotional fluctuation count in the candidate semantic unit. Simultaneously, obtain the maximum pitch difference value in the candidate semantic unit.

[0035] The more emotional fluctuations a candidate semantic unit exhibits, the stronger the user's emotional fluctuations within that unit, and the higher the degree of emotional fluctuation; the two are positively correlated. Similarly, the greater the difference in the maximum pitch value among candidate semantic units, the more pronounced the user's emotional fluctuations within those units, and the higher the degree of emotional fluctuation; the two are also positively correlated.

[0036] The signal-to-noise ratio (SNR) of candidate semantic units is obtained and normalized using the sigmoid function. All SNR values ​​discussed later are the normalized results. As a confidence parameter for the number of sentiment fluctuations and the difference in maximum pitch value, a lower SNR indicates stronger noise in the candidate semantic unit, requiring a reduction in the weighting of these factors and reliance more on textual or environmental features for matching. Therefore, logically, the degree of sentiment fluctuation in candidate semantic units is positively correlated with the SNR.

[0037] Based on the above logical analysis, the following is a specific method for calculating the degree of emotional fluctuation: ; in, This indicates the degree of emotional fluctuation in candidate semantic units. The signal-to-noise ratio of the candidate semantic unit is represented. This indicates the number of emotional fluctuations within a candidate semantic unit. This indicates the number of pitch value differences among candidate semantic units. This indicates the percentage of emotional fluctuations in candidate semantic units. This represents the maximum pitch difference among candidate semantic units. The calculation is achieved by averaging the percentage of emotional fluctuations in candidate semantic units with the maximum pitch difference. This fusion is then combined with the signal-to-noise ratio of the candidate semantic units as a confidence parameter to calculate the degree of emotional fluctuation.

[0038] In this embodiment, if there is no pitch difference greater than or equal to the pitch difference threshold in the candidate semantic unit, it indicates that the emotional fluctuation level of the candidate semantic unit is low, and the emotional fluctuation level of the candidate semantic unit is directly set to 0.

[0039] Using the above method, the emotional fluctuation level of each semantic unit in the current input speech and each semantic unit in the reference semantic units in each historical input speech is obtained.

[0040] Step S13: Based on the difference in emotional fluctuation between each current semantic unit in the current input speech and each reference semantic unit in the historical input speech, obtain the difference in emotional fluctuation between the current input speech and each historical input speech.

[0041] For any current semantic unit in the current input speech, determine the emotional fluctuation level of that current semantic unit, and the emotional fluctuation level of the reference semantic units in each historical input speech. Then, determine the difference in emotional fluctuation level between the current semantic unit and the reference semantic units in each historical input speech, where the difference in emotional fluctuation level is specifically the absolute value of the difference in emotional fluctuation level. The absolute value of the difference in emotional fluctuation level between the current semantic unit and the reference semantic units in each historical input speech is taken as the emotional fluctuation difference between the current semantic unit and each historical input speech. Then, based on the emotional fluctuation differences between each current semantic unit in the current input speech and each historical input speech, obtain the emotional fluctuation difference between the current input speech and each historical input speech, calculated as follows: ; in, This indicates the difference in emotional fluctuation between the current input speech and the h-th historical input speech. This represents the emotional fluctuation level of the y-th semantic unit in the current input speech. Y represents the emotional fluctuation level of the reference semantic unit in the h-th historical input speech, and Y represents the number of current semantic units in the current input speech.

[0042] Smartwatches typically use short, high-frequency text commands for interaction, such as checking steps, setting an alarm, or checking today's weather. However, actual voice input may contain redundant information, such as "Can you check my steps today?" or become semantically fragmented due to noise interference, such as "Check...steps...". If the similarity calculation is directly performed on the text corresponding to the entire input voice sentence, the redundant information will interfere with the matching accuracy.

[0043] This embodiment needs to determine the text similarity between the current input speech and each historical input speech. In an exemplary embodiment, for any current semantic unit in the current input speech, the text similarity between the current semantic unit and the reference semantic units in each historical input speech is determined. The method for obtaining the text similarity has been described above and will not be repeated here. The text similarity between each current semantic unit of the current input speech and the reference semantic units in each historical input speech is fused to obtain the text similarity between the current input speech and each historical input speech. A specific calculation method for the text similarity is given below: ; in, This indicates the text similarity between the current input speech and the h-th historical input speech. This represents the text similarity between the y-th current semantic unit in the current input speech and the reference semantic unit in the h-th historical input speech.

[0044] It should be understood that in the presence of noise interference, speech recognition may lead to increased semantic errors, affecting the accuracy of commands. For example, if a user expresses "I want to listen to music" in a noisy environment, but it is identified as other content due to noise interference, this embodiment can infer the user's true needs by analyzing the emotional tendencies of historical similar semantic units in order to better understand the user's true intention. This emotional context information helps to provide feedback that is more in line with the user's emotions in the event of misidentification. Relying solely on text similarity matching would include all historical speech with the same semantics indiscriminately in the filtering scope, increasing the amount of processing computation and failing to meet the real-time requirements of voice interaction. Therefore, this embodiment, based on text analysis, further analyzes the differences in emotional fluctuations between the current input speech and historical input speech, and integrates the two aspects of information to obtain the matching degree between the current input speech and each historical input speech. Accordingly, based on the differences in emotional fluctuations between the current input speech and each historical input speech, as well as the text similarity between the current input speech and each historical input speech, the matching degree between the current input speech and each historical input speech is obtained. The smaller the difference in emotional fluctuation between the current input speech and the historical input speech, the more consistent the emotional features, indicating a higher degree of emotional matching between the two. This higher degree of matching between the historical and historical input speech further enhances the accuracy of noise removal for the current input speech. Matching degree is inversely correlated with difference in emotional fluctuation. Conversely, the higher the text similarity between the current and historical input speech, the higher the content matching between the two. This higher degree of matching between the historical and historical input speech further enhances the accuracy of noise removal for the current input speech. Matching degree is positively correlated with text similarity. Based on the above logical analysis, a specific calculation method for matching degree is given below: ; in, This represents the degree of matching between the current input speech and the h-th historical input speech. The above method uses averaging to fuse the two parameters and obtain the matching degree. Thus, the matching degree between the current input speech and each historical input speech is obtained.

[0045] Step S2: Determine the degree of interference from environmental noise on the current input speech and each matched speech.

[0046] After obtaining the matching degree between the current input speech and each historical input speech, the higher the matching degree, the more similar the corresponding historical input speech is to the current input speech, and the more likely it is to be the matching speech of the current input speech. Therefore, matching speech for the current input speech is obtained from each historical input speech. This embodiment presets a matching degree threshold, which is used as the comparison object for the matching degree between the current input speech and each historical input speech, thereby determining whether the matching degree between the current input speech and each historical input speech is high. The numerical range of this matching degree threshold is 0-1, and the specific value is set according to actual needs. This embodiment uses 0.6 as an example. Historical input speech with a matching degree greater than or equal to the preset matching degree threshold is taken as the matching speech of the current input speech, thus obtaining several matching speech samples.

[0047] It should be understood that if the number of matching voices obtained through threshold comparison is too large, in order to reduce the amount of data processing, this embodiment can remove some matching voices. For example, the matching voices of the current input voice can be sorted in descending order of matching degree, and the first preset number (e.g., 5) matching voices with matching degree can be selected as the final matching voices of the current input voice. In addition, if there are no matching voices for the current input voice, it means that there are no voice signals in the historical input voices that are similar to the current input voice. In this case, the subsequent data processing process of this embodiment will terminate, and existing general filtering algorithms will be used to filter noise from the current input voice.

[0048] After obtaining several matching voices of the current input speech, the user's intent can be understood to some extent. However, the current input speech is still mixed with environmental noise (such as wind noise and footsteps). To improve the real-time performance and accuracy of voice interaction, this embodiment also needs to further filter out the historical input speech that best matches the user's current input speech to assist in denoising the current input speech data. First, the degree of interference from environmental noise on the current input speech and each of its matching voices is determined. In an exemplary embodiment, such as Figure 3 As shown, the following is a specific process for obtaining the level of interference: Step S21: Decompose the candidate input speech into main speech components and reference speech components.

[0049] For ease of explanation, we define the candidate input speech as any one of the current input speech and each matched speech. The candidate input speech is decomposed into a primary speech component and reference speech components. The primary speech component is the speech component that best represents the features of the candidate input speech, and the reference speech components are all other speech components besides the primary speech component. It should be understood that there is one primary speech component and at least one reference speech component.

[0050] In an exemplary embodiment, the candidate input speech is decomposed into multiple speech components (decomposed components) using a variational mode decomposition (VMD) algorithm or a singular spectrum analysis (SSA) algorithm. The feature values ​​of each speech component are then obtained, and the speech component with the highest contribution is selected as the primary speech component, also known as the effective speech component. The primary speech component contains the user's speech features, while the reference speech component contains environmental features.

[0051] The dominant speech component is determined based on the energy proportion of each decomposed component or its contribution to the original signal: in VMD, the dominant component is the modal component with the highest energy; in SSA, the dominant speech component is the reconstructed component corresponding to the largest singular value; the reference speech component is the remaining decomposed components excluding the dominant component. Therefore, the dominant speech component corresponding to the largest contribution usually contains the main information of the user input, such as the speech content, and is thus considered to be dominated by effective speech. Conversely, after removing the effective speech components, the remaining components usually contain less variance information, representing background noise, environmental interference, etc., and are thus considered to be dominated by noise.

[0052] Step S22: Determine the amplitude difference between the candidate input speech and the main speech component at the same time, and obtain the noise interference amplitude at each time.

[0053] The candidate input speech and its main speech components are mapped onto the same two-dimensional coordinate system. The horizontal axis of this system represents time, and the vertical axis represents the amplitude value of the speech. For a given moment on the horizontal axis, the amplitude difference between the candidate input speech and its main speech component at that moment is determined, i.e., the absolute value of the amplitude difference. This difference is then normalized, and the normalized result is used as the noise interference amplitude at that moment, thus obtaining the noise interference amplitude at each moment. The normalization method can be as follows: obtain the maximum amplitude value between the candidate input speech and its main speech component, and calculate the ratio of the absolute value of the amplitude difference to the maximum amplitude value. This ratio is the normalized result and is used as the noise interference amplitude.

[0054] Step S23: Obtain the environmental interference time and environmental interference period in the candidate input speech from the noise interference amplitude.

[0055] For the noise interference amplitude of the candidate input speech and its main speech components at various times, the larger the noise interference amplitude, the more the speech signal at the corresponding time is affected by environmental interference, and the more likely the corresponding time is to be considered an environmentally disturbed moment. In an exemplary embodiment, this embodiment presets a noise interference amplitude threshold, the value of which ranges from 0 to 1, and the specific value is set according to actual needs; this embodiment uses 0.3 as an example. The noise interference amplitude of the candidate input speech and its main speech components at various times is compared with the noise interference amplitude threshold. The times corresponding to noise interference amplitudes greater than or equal to the threshold are determined, and these times are identified as environmentally disturbed moments, thus obtaining the environmentally disturbed moments in the candidate input speech, and simultaneously obtaining the number of environmentally disturbed moments in the candidate input speech. The more environmentally disturbed moments there are, the stronger the interference of environmental noise on the candidate input speech; the two are positively correlated.

[0056] An environmental interference period is formed by combining at least two consecutive environmental interference moments in time, thus obtaining several environmental interference periods in the candidate input speech. It should be understood that isolated environmental interference moments are also considered as one environmental interference period. The duration of each environmental interference period is determined, where duration represents the number of moments it contains. The longer the duration of the environmental interference period, the stronger the interference from environmental noise on the candidate input speech; the two are positively correlated.

[0057] For any given period of environmental interference, the maximum noise interference amplitude is determined from the noise interference amplitudes at each time point within that period. This maximum noise interference amplitude is then used as the maximum noise interference amplitude for that period, thus obtaining the maximum noise interference amplitude for each environmental interference period. The larger the maximum noise interference amplitude for each period, the stronger the interference from environmental noise on the candidate input speech; the two are positively correlated.

[0058] It should be understood that if there are no environmental interference moments in the candidate input speech, then the degree of interference from environmental noise is directly set to 0.

[0059] Step S24: Based on the number of environmental interference moments in the candidate input speech, the duration of each environmental interference period, and the maximum noise interference amplitude of each environmental interference period, the degree of interference of the candidate input speech by environmental noise is obtained.

[0060] The degree of environmental noise interference in the candidate input speech is obtained by fusing and analyzing the number of environmental interference moments, the duration of each interference moment, and the maximum noise interference amplitude of each interference moment. Based on the above logical analysis, a specific calculation method for the degree of environmental noise interference in the candidate input speech is given below: ; in, This indicates the degree of interference from environmental noise in the candidate input speech, where u represents the number of environmental interference moments in the candidate input speech, and U represents the total number of moments in the candidate input speech. This represents the percentage of environmental interference moments in the candidate input speech, which is equivalent to normalizing u. This represents the duration of the g-th environmental disturbance period. Let represent the sum of the durations of all environmental interference periods in the candidate input speech. Then, the sum of the durations of all environmental interference periods in the candidate input speech... The sum is 1. This represents the maximum noise interference amplitude during the g-th environmental interference period, where G represents the number of environmental interference periods.

[0061] The essence is to As weights, the maximum noise interference amplitudes for each environmental interference period are summed using a weighted average method. Furthermore, the above interference level calculation formula employs an averaging approach to integrate the two parameters.

[0062] The more environmental interference moments in the candidate input speech, the longer the duration of each interference moment, and the greater the maximum noise interference amplitude during each interference moment, the noisier the environment in which the candidate input speech is located. This indicates that the speech data collected by the microphone will contain more noise, and the candidate input speech will be more severely affected by environmental noise interference. This allows us to determine the degree of environmental noise interference affecting the current input speech and each of the matched speech samples.

[0063] Step S3: Determine the degree of environmental difference between the current input speech and each matched speech, and combine the differences in interference levels to obtain the environmental similarity between the current input speech and each matched speech.

[0064] When the current input speech is significantly affected by ambient noise, such as in crowded or noisy places or on busy streets, this noise can interfere with the clarity and intelligibility of the speech signal, thus affecting the accurate recognition of the user's intent. Next, we determine the degree of environmental difference between the current input speech and each matched speech, and combine this with the difference in the degree of interference from ambient noise between the current input speech and each matched speech to obtain the environmental similarity between the current input speech and each matched speech. Based on this environmental similarity, we can better select the most suitable historical input speech for matching and noise reduction, thereby improving the accuracy of the interaction.

[0065] In an exemplary embodiment, the following is a specific process for obtaining the degree of environmental difference between the current input speech and the environment in which each matched speech is located: Define any reference speech component of the current input speech as a candidate reference speech component, and define any matching speech component as a candidate matching speech component.

[0066] The difference distance between each candidate reference speech component of the current input speech and each reference speech component of the candidate matching speech is obtained. If the difference distance is obtained using cosine similarity, since the lengths of the candidate reference speech components and the reference speech components of the candidate matching speech may differ, the shorter speech components need to be extended to the same length as the longer speech components through interpolation (such as linear interpolation or spline interpolation). Then, the cosine similarity between the two length-aligned speech components is calculated. Since the cosine similarity ranges from -1 to 1 and is inversely correlated with the difference distance, the difference distance can be calculated as (1 - cosine similarity) / 2. Alternatively, in this embodiment, the difference distance is specifically the Dynamic Time Warping (DTW) distance. In this case, length alignment is not required; the DTW distance between the two speech components is directly calculated, and then the DTW distance is normalized (e.g., using minimum-maximum value normalization) to obtain the difference distance.

[0067] The minimum difference distance is determined from the difference distances between each reference speech component of the candidate reference speech component and each reference speech component of the candidate matched speech. This minimum difference distance is used as the environmental difference representation between the candidate reference speech component and the candidate matched speech, thus obtaining the environmental difference representation between each reference speech component of the current input speech and the candidate matched speech. The environmental difference representations of each reference speech component of the current input speech and the candidate matched speech are then fused. Specifically, the average value of the environmental difference representations of each reference speech component of the current input speech and the candidate matched speech is calculated, and the result is used as the degree of environmental difference between the current input speech and the candidate matched speech. This yields the degree of environmental difference between the current input speech and each matched speech component.

[0068] The difference in the degree of interference between the current input speech and each matched speech is obtained. Specifically, the difference in the degree of interference is the absolute value of the difference in the degree of interference.

[0069] The environmental similarity between the current input speech and each matched speech is determined by the degree of environmental difference between their respective environments and the degree of interference from environmental noise. A smaller degree of environmental difference indicates a smaller difference in the environments in which the current input speech and matched speech are located, meaning they are more likely to be in the same environment, resulting in greater environmental similarity and an inverse correlation. Conversely, a smaller difference in the degree of interference from environmental noise between the current input speech and matched speech also indicates greater environmental similarity and an inverse correlation. Therefore, a specific method for calculating environmental similarity is given below: ; in, This indicates the environmental similarity between the current input speech and the p-th matching speech. This indicates the degree of environmental difference between the current input speech and the p-th matching speech. This represents the absolute value of the difference between the degree of interference of the current input speech with environmental noise and the degree of interference of the p-th matched speech with environmental noise, i.e., the difference in the degree of interference.

[0070] Step S4: Combine the matching degree and environmental similarity to obtain the final matching degree between the current input speech and each matched speech.

[0071] Step S3 obtains the environmental similarity between the current input speech and each matched speech. Using the environmental similarity as an adjustment coefficient, the matching degree between the current input speech and each matched speech is adjusted to obtain the final matching degree between the current input speech and each matched speech, as follows: ; in, This indicates the final matching degree between the current input speech and the p-th matching speech. This indicates the degree of matching between the current input speech and the p-th matching speech.

[0072] Step S5: Denoise the current input speech based on the optimal matching speech.

[0073] The higher the final matching degree between the current input speech and the matched speech, the more similar the current input speech and the more likely it is to be used as the historical input speech that best matches the user's current input speech for denoising the current input speech data. In an exemplary embodiment, the final matching degree with the highest value is determined from the final matching degrees between the current input speech and each matched speech, and the matched speech corresponding to the highest final matching degree is taken as the optimal matched speech for the current input speech.

[0074] The optimal matching speech can be understood as a clean or denoised speech signal. The current input speech is denoised based on the optimal matching speech. In an exemplary embodiment, the noise distribution features (such as noise spectral density) of the current input speech and the spectral envelope features of the optimal matching speech are extracted. Then, the noise distribution features and spectral envelope features are used as prior constraints and input along with the current input speech into a preset denoising algorithm (such as Wiener filtering, spectral subtraction, or a deep learning denoising model) to output the denoised current input speech.

[0075] In the processing, this embodiment can also unify the sampling rate of the optimal matching speech and the current input speech to ensure that the frequency range is consistent, and can also perform pre-emphasis (such as using a first-order high-pass filter) to enhance high-frequency components, which is convenient for subsequent feature extraction.

[0076] The current input speech is divided into multiple short-time frames (e.g., frame length is typically 20-40 milliseconds). A Hamming window can be applied to each frame to reduce spectral leakage. A short-time Fourier transform is performed on each frame to obtain the spectrum. The squared amplitude spectrum of each frame is averaged to obtain the initial noise power spectral density. Based on the noise power spectral density of each frame, the noise distribution characteristics, i.e., the noise power spectral density matrix, are output as the noise prior.

[0077] Similarly, a short-time Fourier transform is performed on the optimally matched speech to obtain its spectrum. Then, the spectral envelope is calculated. In this embodiment, a linear predictive coding method can be used. Linear predictive coding analysis (e.g., order 12-16) is applied to each frame to obtain linear prediction coefficients. Then, the spectral envelope is synthesized using the linear prediction coefficients. Finally, the spectral envelope features are output, represented as an envelope matrix, which serves as the spectral envelope prior of the optimally matched speech.

[0078] Since the duration of the optimal matching speech may differ from that of the current input speech, it is necessary to perform time alignment on the features. In this embodiment, a dynamic time warping algorithm can be used to align the spectral envelope features and the frame sequence of the current input speech to ensure that the prior constraints correspond to the current input speech.

[0079] The spectrum, noise prior, and spectral envelope prior obtained from the short-time Fourier transform of the current input speech are input into the denoising algorithm. The following denoising algorithm is illustrated using Wiener filtering or spectral subtraction as examples. Wiener filtering and spectral subtraction are computationally efficient and suitable for real-time processing.

[0080] If Wiener filtering is chosen as the denoising algorithm, the denoising process is as follows: the power spectrum of the best-matched speech is estimated using the spectral envelope prior; for each frame and each frequency point, the Wiener gain is calculated; the gain is applied to the spectrum of the current input speech to achieve spectral enhancement; the phase information of the current input speech is preserved, since Wiener filtering only modifies the amplitude spectrum.

[0081] If the denoising algorithm uses spectral subtraction, the denoising process is as follows: First, use the noise prior as the noise amplitude spectrum estimate. Calculate the enhanced amplitude spectrum. Combine the enhanced amplitude spectrum with the spectral envelope prior. Iteratively adjust the enhanced amplitude spectrum envelope to approximate the spectral envelope prior, for example, using constrained least squares. Finally, perform phase processing, i.e., retain the original phase.

[0082] The denoised spectrum is subjected to an inverse short-time Fourier transform, and combined with the original phase or estimated phase, the signals are superimposed and added to synthesize the denoised time-domain signal, which is the denoised current input speech. Additionally, this embodiment can also apply a de-emphasis filter to reverse the pre-emphasis effect. Optionally, gain normalization or dynamic range compression can also be performed to improve auditory quality.

[0083] This embodiment also provides a conversational human-computer interaction system for smart wearable devices, including: a memory and a processor; the memory is connected to the processor, and the memory is used to store program instructions; the processor is used to implement the steps in the above-described conversational human-computer interaction method embodiment for smart wearable devices when the program instructions are executed.

[0084] In one exemplary embodiment, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the embodiment of the conversational human-computer interaction method applied to a smart wearable device.

[0085] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0086] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A conversational human-computer interaction method applied to smart wearable devices, characterized in that, include: Determine the differences in emotional fluctuations between the current input speech and each historical input speech, and combine this with text similarity to obtain the degree of matching between the current input speech and each historical input speech; The degree of interference from environmental noise on the current input speech and each matched speech is determined respectively; the matched speech is obtained by filtering from each historical input speech based on the degree of matching. Determine the degree of environmental difference between the current input speech and each matched speech, and combine the difference in interference levels to obtain the environmental similarity between the current input speech and each matched speech; By combining the matching degree and environmental similarity, the final matching degree between the current input speech and each matched speech is obtained; The current input speech is denoised based on the optimal matching speech; the optimal matching speech is selected from each matching speech according to the final matching degree.

2. The conversational human-computer interaction method applied to smart wearable devices as described in claim 1, characterized in that, The process of obtaining the differences in emotional fluctuations includes: Determine the reference semantic units of each current semantic unit in the current input speech in each historical input speech; The emotional fluctuation level of candidate semantic units is determined based on tone changes; the candidate semantic units are any semantic units between the current semantic unit and the reference semantic unit. The difference in emotional fluctuation between the current input speech and each historical input speech is obtained by comparing the emotional fluctuation levels of each current semantic unit in the current input speech with those of the reference semantic units in each historical input speech.

3. The conversational human-computer interaction method for smart wearable devices as described in claim 2, characterized in that, The process of obtaining the reference semantic unit includes: Convert the current semantic unit and each historical semantic unit in the historical input speech into TF-IDF vectors respectively; The vector similarity between the TF-IDF vector of the current semantic unit and the TF-IDF vector of each historical semantic unit in the historical input speech is determined. The historical semantic unit corresponding to the largest vector similarity is used as the reference semantic unit of the current semantic unit in the historical input speech.

4. The conversational human-computer interaction method for smart wearable devices as described in claim 2, characterized in that, The process of obtaining the degree of emotional fluctuation includes: Determine the pitch value difference between adjacent pitches in candidate semantic units; The number of emotional fluctuations in candidate semantic units is obtained from the difference in pitch values; The emotional fluctuation level of a candidate semantic unit is obtained from the maximum pitch difference, the number of emotional fluctuations, and the signal-to-noise ratio of the candidate semantic unit; the emotional fluctuation level is positively correlated with the maximum pitch difference, the number of emotional fluctuations, and the signal-to-noise ratio.

5. The conversational human-computer interaction method applied to smart wearable devices as described in claim 2, characterized in that, The process of obtaining the text similarity cases includes: By fusing the text similarity between each current semantic unit of the current input speech and the reference semantic units in each historical input speech, the text similarity between the current input speech and each historical input speech is obtained.

6. The conversational human-computer interaction method applied to smart wearable devices as described in claim 1, characterized in that, The process of obtaining the level of interference includes: The candidate input speech is decomposed into a main speech component and a reference speech component; the candidate input speech is any input speech between the current input speech and each matched speech. Determine the amplitude difference between the candidate input speech and the main speech component at the same time to obtain the noise interference amplitude at each time. The environmental interference time and duration in the candidate input speech are obtained from the noise interference amplitude. The degree of environmental noise interference on the candidate input speech is obtained based on the number of environmental interference moments in the candidate input speech, the duration of each environmental interference moment, and the maximum noise interference amplitude of each environmental interference moment; the degree of interference is positively correlated with the number of environmental interference moments, the duration of each environmental interference moment, and the maximum noise interference amplitude.

7. The conversational human-computer interaction method applied to smart wearable devices as described in claim 6, characterized in that, The process of decomposing candidate input speech into primary speech components and reference speech components includes: The candidate input speech is decomposed into multiple speech components; wherein, the main speech component is the speech component with the largest contribution, and the reference speech components are all other speech components except the main speech component.

8. The conversational human-computer interaction method applied to smart wearable devices as described in claim 6, characterized in that, The process of obtaining the degree of environmental difference includes: Obtain the difference distance between each reference speech component of the current input speech and each reference speech component of the matching speech, determine the minimum difference distance, and obtain the environmental difference representation between each reference speech component of the current input speech and the matching speech. By integrating the environmental differences between the various reference speech components of the current input speech and the matched speech, the degree of environmental difference between the current input speech and the matched speech is obtained.

9. The conversational human-computer interaction method applied to smart wearable devices as described in claim 1, characterized in that, The denoising of the current input speech based on the optimal matching speech includes: Extract the noise distribution features of the current input speech and the spectral envelope features of the best matching speech; The noise distribution features and spectral envelope features are used as prior constraints and input into a preset denoising algorithm along with the current input speech, and the denoised current input speech is output.

10. A conversational human-computer interaction system for smart wearable devices, characterized in that it includes: Memory and processor; The memory is connected to the processor; The memory is used to store program instructions; The processor is configured to implement, when program instructions are executed, the conversational human-computer interaction method applied to a smart wearable device as described in any one of claims 1-9.