A smart home voice interaction testing method and device

By determining voice interaction parameters in smart home scenarios and conducting real-time testing and compensation, the instability of voice interaction testing in complex scenarios is solved, and high recognition rate and natural interaction experience in multiple environments are achieved.

CN120220649BActive Publication Date: 2025-09-02TIANJIN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510370640.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-09-02
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing voice interaction testing methods are difficult to achieve stable and accurate testing results in complex and specialized smart home scenarios, especially in multi-room environments, open and closed spaces with great impact on sound propagation characteristics and echo problems.

Method used

By determining the scene characteristic data of the specific scene to be tested, determining the voice interaction parameters, conducting real-time voice data testing, obtaining test scores, and correcting and compensating the voice interaction parameters based on the comparison results, including the adjustment of identification threshold, noise suppression parameters and voice enhancement coefficient, and optimizing interaction performance based on user distance, behavior and device status data.

Benefits of technology

It improves the stability and accuracy of voice interaction testing, adapts to different user needs and environmental changes, and provides a natural and smooth interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220649B_ABST
    Figure CN120220649B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of voice interaction testing, and discloses a method and device for testing voice interaction in a smart home. The method comprises: determining a specific scenario to be tested, determining voice interaction parameters based on scenario feature data; obtaining a test score based on the voice interaction data in the specific scenario to be tested; comparing the test score with a test score threshold, and determining whether to modify the voice interaction parameters based on the comparison result; after determining the modified voice interaction parameters, determining an interaction influence index between the user and the smart home to be tested based on distance data, behavior data, and device status data; determining whether to compensate for the modified voice interaction parameters based on the interaction influence index; and when it is determined that the modified voice interaction parameters are to be compensated, compensating the modified voice interaction parameters based on a compensation coefficient to obtain compensated voice interaction parameters. The present invention improves the stability and accuracy of voice interaction testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice interaction testing, and in particular to a method and device for testing smart home voice interaction. Background Art

[0002] With the continuous advancement and widespread adoption of smart home technology, voice interaction has increasingly become one of the key ways for users to communicate and interact with smart home devices. Currently, existing voice interaction testing methods perform well in general application scenarios and can well meet basic user needs. However, in specific and complex scenarios, these testing methods have exposed limitations, making it difficult to achieve stable and accurate testing results. Such specialized scenarios include complex home space layouts, diverse device types, multi-room environments, open and closed spaces, personalized space configurations, and other complex and varied scenarios.

[0003] In these specialized scenarios, voice interaction testing faces multiple challenges due to the diversity and complexity of the environments. For example, in a multi-room environment, users may need to interact with devices in different rooms, requiring voice interaction testing to accurately identify and respond to commands from different rooms. Furthermore, in both open and enclosed spaces, sound propagation characteristics and echo issues can also affect the accuracy of voice interaction.

[0004] Therefore, it is necessary to design a smart home voice interaction testing method and device for specialized scene recognition to solve the problems existing in current technology. Summary of the Invention

[0005] In view of this, the present invention proposes a smart home voice interaction testing method and device, aiming to improve the stability and accuracy of voice interaction testing.

[0006] In one aspect, the present invention provides a smart home voice interaction testing method, comprising the following steps:

[0007] S100: Determine a specific scenario to be tested, extract scenario feature data of the specific scenario to be tested, and determine voice interaction parameters according to the scenario feature data;

[0008] S200: Performing a voice interaction test on the smart home to be tested using real-time voice data in the specific scenario to be tested, obtaining voice interaction data, and obtaining a test score based on the voice interaction data;

[0009] S300: Compare the test score with a test score threshold, and determine whether to modify the voice interaction parameter based on the comparison result;

[0010] S400: When it is determined that the voice interaction parameter needs to be corrected, feature extraction is performed on the real-time voice data to obtain a real-time voice feature value; the real-time voice feature value is compared with the historical voice data, a correction coefficient is determined based on the comparison result, and the voice interaction parameter is corrected based on the correction coefficient to obtain a corrected voice interaction parameter;

[0011] S500: After determining the modified voice interaction parameters, collecting distance data between the user and the smart home to be tested, the user's behavior data, and the device status data of the smart home to be tested, and determining an interaction impact index between the user and the smart home to be tested based on the distance data, behavior data, and device status data; and determining whether to compensate for the modified voice interaction parameters based on the interaction impact index;

[0012] S600: When it is determined to compensate the modified voice interaction parameter, a compensation coefficient is determined according to the interaction influence index, and the modified voice interaction parameter is compensated according to the compensation coefficient to obtain a compensated voice interaction parameter.

[0013] Furthermore, when determining the voice interaction parameters according to the scene feature data, it includes:

[0014] The scene feature data includes noise feature data and spatial echo feature data;

[0015] The speech interaction parameters include recognition threshold, noise suppression parameter and speech enhancement coefficient;

[0016] The recognition threshold is obtained by the following formula:

[0017]

[0018] The noise suppression parameter is obtained by the following formula:

[0019]

[0020] The speech enhancement coefficient is obtained by the following formula:

[0021]

[0022] Among them, T represents the recognition threshold; T0 represents the basic recognition threshold; a1 represents the noise recognition weight coefficient; xn represents the noise characteristic value; bn represents the noise standard value; a2 represents the spatial recognition weight coefficient; xr represents the current spatial echo characteristic value; br represents the current spatial echo standard value; N represents the noise suppression parameter; N0 represents the basic noise suppression parameter; b1 represents the noise suppression weight coefficient; b2 represents the spatial suppression weight coefficient; C represents the speech enhancement coefficient; C0 represents the basic speech enhancement coefficient; k1 represents the speech enhancement weight coefficient; k2 represents the spatial enhancement weight coefficient.

[0023] Furthermore, comparing the test score with the test score threshold, and determining whether to modify the voice interaction parameter according to the comparison result, includes:

[0024] The test score is obtained by the following formula:

[0025]

[0026] Among them, S represents the test score; c1 represents the weight coefficient of speech recognition accuracy; Rmax represents the maximum value of speech recognition accuracy; R1 represents the speech recognition accuracy; c2 represents the weight coefficient of response time; T2 represents the response time; Tmin represents the minimum value of response time; c3 represents the weight coefficient of response accuracy; Amax represents the maximum value of response accuracy; A3 represents response accuracy.

[0027] Furthermore, when comparing the test score with the test score threshold and determining whether to modify the voice interaction parameter based on the comparison result, the method further includes:

[0028] When the test score is greater than the test score threshold, determining not to modify the voice interaction parameter;

[0029] When the test score is less than or equal to the test score threshold, it is determined to modify the voice interaction parameter.

[0030] Furthermore, the real-time voice feature value is compared with the historical voice data, a correction coefficient is determined according to the comparison result, and the voice interaction parameter is corrected according to the correction coefficient to obtain the corrected voice interaction parameter, including:

[0031] Calculating the maximum similarity between the real-time speech feature value and the historical speech data, comparing the maximum similarity with a maximum similarity threshold, and determining the correction coefficient based on the comparison result;

[0032] The maximum similarity is obtained by the following formula:

[0033]

[0034] Among them, M max represents the maximum similarity; X i represents the i-th real-time speech feature vector; Y j represents the jth historical speech feature vector; m represents the number of real-time speech feature vectors; n represents the number of historical speech feature vectors.

[0035] Furthermore, comparing the maximum similarity with a maximum similarity threshold and determining the correction coefficient according to the comparison result includes:

[0036] Comparing the maximum similarity with a first maximum similarity threshold and a second maximum similarity threshold, and determining the correction coefficient according to the comparison result; wherein the first maximum similarity threshold is less than the second maximum similarity threshold;

[0037] Setting a correction coefficient interval, wherein the correction coefficient interval includes a first correction coefficient, a second correction coefficient, and a third correction coefficient;

[0038] When the maximum similarity is less than or equal to the first maximum similarity threshold, determining the correction coefficient of the voice interaction parameter as the first correction coefficient, and taking the product of the first correction coefficient and the voice interaction parameter as the corrected voice interaction parameter;

[0039] When the maximum similarity is greater than the first maximum similarity threshold and less than or equal to the second maximum similarity threshold, determining the correction coefficient of the voice interaction parameter to be a second correction coefficient, and using the product of the second correction coefficient and the voice interaction parameter as the corrected voice interaction parameter;

[0040] When the maximum similarity is greater than the second maximum similarity threshold, the correction coefficient of the voice interaction parameter is determined to be a third correction coefficient, and the product value of the third correction coefficient and the voice interaction parameter is used as the corrected voice interaction parameter.

[0041] Furthermore, when determining the interaction impact index between the user and the smart home to be tested based on the distance data, the behavior data, and the device status data, the method includes:

[0042] The interaction impact index is obtained by the following formula:

[0043]

[0044] Among them, I represents the interaction influence index; D represents distance data; B represents behavior data; E represents device status data; α1 represents the distance weight coefficient; α2 represents the behavior weight coefficient; α3 represents the device status weight coefficient; β1 represents the first influence coefficient; β2 represents the second influence coefficient; β3 represents the third influence coefficient.

[0045] Furthermore, when determining whether to compensate the modified voice interaction parameter according to the interaction influence index, it includes:

[0046] Subtracting the interaction impact index from the interaction impact index threshold to obtain an index difference;

[0047] Comparing the index difference with an index difference threshold, and determining whether to compensate the modified voice interaction parameter based on the comparison result;

[0048] When the index difference is greater than or equal to the index difference threshold, determining to compensate the modified voice interaction parameter;

[0049] When the index difference is less than the index difference threshold, it is determined that the modified voice interaction parameter is not compensated.

[0050] Furthermore, determining a compensation coefficient according to the interaction influence index and compensating the modified voice interaction parameter according to the compensation coefficient includes:

[0051] Comparing the index difference threshold with a first index difference threshold and a second index difference threshold, and determining the compensation coefficient according to the comparison result; wherein the first index difference threshold is smaller than the second index difference threshold;

[0052] Setting a compensation coefficient interval, wherein the compensation coefficient interval includes a first compensation coefficient, a second compensation coefficient, and a third compensation coefficient;

[0053] When the exponent difference is less than or equal to the first exponent difference threshold, determining the compensation coefficient of the modified voice interaction parameter to be a first compensation coefficient, and taking the product of the first compensation coefficient and the modified voice interaction parameter as the compensated voice interaction parameter;

[0054] When the exponent difference is greater than the first exponent difference threshold and less than or equal to the second exponent difference threshold, determining the compensation coefficient of the modified voice interaction parameter as the second compensation coefficient, and using the product value of the second compensation coefficient and the modified voice interaction parameter as the compensated voice interaction parameter;

[0055] When the exponent difference is greater than the second exponent difference threshold, the compensation coefficient of the modified voice interaction parameter is determined to be a third compensation coefficient, and the product value of the third compensation coefficient and the modified voice interaction parameter is used as the compensated voice interaction parameter.

[0056] Compared with the prior art, the beneficial effects of the present invention are as follows: the smart home voice interaction test method for specialized scene recognition provided by the present invention can improve the stability and accuracy of the voice interaction test; by determining the specialized scene to be tested, extracting the scene feature data of the specialized scene to be tested, and determining the voice interaction parameters according to the scene feature data, the adaptability and accuracy of the test can be improved, and good performance in various scenes can be guaranteed; in the specialized scene to be tested, real-time voice data is used to perform a voice interaction test on the smart home to be tested, voice interaction data is obtained, and a test score is obtained according to the voice interaction data, which can clearly grasp the test situation; the test score is used to determine the test score. By comparing with the test score threshold and judging whether to correct the voice interaction parameters based on the comparison results, it is helpful to continuously optimize and adjust in actual applications to adapt to the needs of different users and environmental changes; by extracting features from real-time voice data and comparing it with historical voice data, the voice interaction parameters can be further refined to ensure a high recognition rate in a variety of noise environments; by collecting distance data, behavior data and device status data between the user and the smart home, the user interaction experience can be more comprehensively evaluated, so as to make necessary compensation for the voice interaction parameters, better adapt to the specific needs of users and environmental changes, and provide a more natural and smooth interaction experience.

[0057] In another aspect, the present invention also proposes a smart home voice interaction testing device for specialized scene recognition, comprising:

[0058] a determination module configured to determine a specific scenario to be tested, extract scenario feature data of the specific scenario to be tested, and determine voice interaction parameters based on the scenario feature data; and further configured to perform a voice interaction test on the smart home to be tested using real-time voice data in the specific scenario to be tested, obtain voice interaction data, and obtain a test score based on the voice interaction data;

[0059] A first judgment module is configured to compare the test score with a test score threshold, and determine whether to modify the voice interaction parameter according to the comparison result;

[0060] a correction module configured to, when determining to correct the voice interaction parameter, perform feature extraction on the real-time voice data to obtain a real-time voice feature value; compare the real-time voice feature value with historical voice data, determine a correction coefficient based on the comparison result, and correct the voice interaction parameter based on the correction coefficient to obtain a corrected voice interaction parameter;

[0061] A second judgment module is configured to, after determining the modified voice interaction parameter, collect distance data between the user and the smart home to be tested, the user's behavior data, and the device status data of the smart home to be tested, and determine an interaction impact index between the user and the smart home to be tested based on the distance data, behavior data, and device status data; and determine whether to compensate for the modified voice interaction parameter based on the interaction impact index;

[0062] The compensation module is configured to determine a compensation coefficient according to the interaction influence index when determining to compensate the modified voice interaction parameter, and compensate the modified voice interaction parameter according to the compensation coefficient to obtain a compensated voice interaction parameter.

[0063] It is understandable that the above-mentioned smart home voice interaction testing method and device for specialized scene recognition have the same beneficial effects and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0065] Figure 1 A flowchart of a smart home voice interaction testing method for specialized scene recognition provided by an embodiment of the present invention;

[0066] Figure 2 This is a structural diagram of a smart home voice interaction testing device for specialized scene recognition provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0067] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0068] See Figure 1 As shown, in some embodiments of the present application, this embodiment provides a smart home voice interaction testing method for specialized scene recognition, including the following steps:

[0069] S100: Determine a specific scenario to be tested, extract scenario feature data of the specific scenario to be tested, and determine voice interaction parameters according to the scenario feature data;

[0070] S200: Performing a voice interaction test on the smart home to be tested using real-time voice data in the specific scenario to be tested, obtaining voice interaction data, and obtaining a test score based on the voice interaction data;

[0071] S300: Compare the test score with a test score threshold, and determine whether to modify the voice interaction parameter based on the comparison result;

[0072] S400: When it is determined that the voice interaction parameter needs to be corrected, feature extraction is performed on the real-time voice data to obtain a real-time voice feature value; the real-time voice feature value is compared with the historical voice data, a correction coefficient is determined based on the comparison result, and the voice interaction parameter is corrected based on the correction coefficient to obtain a corrected voice interaction parameter;

[0073] S500: After determining the modified voice interaction parameters, collecting distance data between the user and the smart home to be tested, the user's behavior data, and the device status data of the smart home to be tested, and determining an interaction impact index between the user and the smart home to be tested based on the distance data, behavior data, and device status data; and determining whether to compensate for the modified voice interaction parameters based on the interaction impact index;

[0074] S600: When it is determined to compensate the modified voice interaction parameter, a compensation coefficient is determined according to the interaction influence index, and the modified voice interaction parameter is compensated according to the compensation coefficient to obtain a compensated voice interaction parameter.

[0075] In this embodiment, the specialized scenario to be tested is preferably a high-noise environment. For example, in a high-noise environment, it is necessary to be able to distinguish and ignore background noise and accurately recognize the user's voice command.

[0076] It can be understood that the smart home voice interaction test method for specialized scene recognition provided by this embodiment can improve the stability and accuracy of voice interaction testing; by determining the specialized scene to be tested, extracting the scene feature data of the specialized scene to be tested, and determining the voice interaction parameters based on the scene feature data, it can improve the adaptability and accuracy of the test and ensure good performance in a variety of scenes; in the specialized scene to be tested, real-time voice data is used to perform voice interaction testing on the smart home to be tested, voice interaction data is obtained, and test scores are obtained based on the voice interaction data, which can clearly grasp the test situation; the test score is compared with the test score threshold The system compares the real-time voice data with the real-time voice data, and determines whether to modify the voice interaction parameters based on the comparison results. This helps to continuously optimize and adjust in actual applications to adapt to the needs of different users and environmental changes. By extracting features from real-time voice data and comparing it with historical voice data, the voice interaction parameters can be further refined to ensure a high recognition rate in a variety of noise environments. By collecting distance data, behavior data, and device status data between the user and the smart home, the user interaction experience can be more comprehensively evaluated, thereby making necessary compensation for the voice interaction parameters to better adapt to the specific needs of users and environmental changes, thereby providing a more natural and smooth interaction experience.

[0077] Specifically, when determining the voice interaction parameters based on the scene feature data, it includes:

[0078] The scene feature data includes noise feature data and spatial echo feature data;

[0079] The speech interaction parameters include recognition threshold, noise suppression parameter and speech enhancement coefficient;

[0080] The recognition threshold is obtained by the following formula:

[0081]

[0082] The noise suppression parameter is obtained by the following formula:

[0083]

[0084] The speech enhancement coefficient is obtained by the following formula:

[0085]

[0086] Among them, T represents the recognition threshold; T0 represents the basic recognition threshold; a1 represents the noise recognition weight coefficient; xn represents the noise characteristic value; bn represents the noise standard value; a2 represents the spatial recognition weight coefficient; xr represents the current spatial echo characteristic value; br represents the current spatial echo standard value; N represents the noise suppression parameter; N0 represents the basic noise suppression parameter; b1 represents the noise suppression weight coefficient; b2 represents the spatial suppression weight coefficient; C represents the speech enhancement coefficient; C0 represents the basic speech enhancement coefficient; k1 represents the speech enhancement weight coefficient; k2 represents the spatial enhancement weight coefficient.

[0087] In this embodiment, the recognition threshold is a voice interaction parameter used to determine whether to respond to an input voice signal; the noise suppression parameter is used to reduce the impact of background noise on voice recognition; and the voice enhancement coefficient is used to improve voice signal quality, ensuring clear recognition of user commands even in noisy environments. By adjusting these parameters, user voice commands can be more accurately recognized, maintaining good interactive performance even in specialized scenarios with high noise levels.

[0088] In this embodiment, the basic recognition threshold is a pre-set threshold under an ideal environment with no noise or echo. The basic noise suppression parameters and basic speech enhancement coefficients are also set under similar ideal conditions. They serve as a benchmark for adjustment based on scenario characteristics in actual applications. This allows for adaptation to a variety of environments, providing a more stable and reliable voice interaction experience in smart home scenarios.

[0089] In this embodiment, a1 represents the noise recognition weight coefficient, preferably between 0.1 and 0.3. A larger value indicates a greater impact of noise characteristics on the recognition threshold. a2 represents the spatial recognition weight coefficient, preferably between 0.1 and 0.3. A larger value indicates a greater impact of spatial echo characteristics on the recognition threshold. Similarly, b1 and b2 represent the noise suppression weight coefficient and spatial suppression weight coefficient, respectively. Their preferred ranges are also between 0.1 and 0.3 to ensure that the noise suppression parameters and speech enhancement coefficients can be effectively adjusted to adapt to different noise and echo environments. k1 and k2 represent the speech enhancement weight coefficients, respectively. Their preferred ranges are between 0.1 and 0.3 to ensure that the speech enhancement coefficients can be appropriately adjusted based on the actual noise and echo conditions, thereby improving the clarity and intelligibility of the speech signal. By properly selecting and adjusting these weight coefficients, it is possible to flexibly respond to various specialized scenarios, achieving more accurate speech recognition and a better interactive experience.

[0090] Specifically, comparing the test score with the test score threshold, and determining whether to modify the voice interaction parameter based on the comparison result, includes:

[0091] The test score is obtained by the following formula:

[0092]

[0093] Among them, S represents the test score; c1 represents the weight coefficient of speech recognition accuracy; Rmax represents the maximum value of speech recognition accuracy; R1 represents the speech recognition accuracy; c2 represents the weight coefficient of response time; T2 represents the response time; Tmin represents the minimum value of response time; c3 represents the weight coefficient of response accuracy; Amax represents the maximum value of response accuracy; A3 represents response accuracy.

[0094] In this embodiment, c1 represents the weight coefficient of speech recognition accuracy, preferably 0.4, c2 represents the weight coefficient of response time, preferably 0.3, and c3 represents the weight coefficient of response accuracy, preferably 0.3. This weight distribution ensures that speech recognition accuracy accounts for a larger proportion of the test score, while not neglecting the importance of response time and response accuracy.

[0095] It's understandable that determining the test score based on speech recognition accuracy, response time, and response accuracy provides a comprehensive assessment of voice interaction test performance. Speech recognition accuracy reflects the ability to understand user voice commands, response time reflects the speed of command processing, and response accuracy focuses on the correctness of command execution. By combining these three indicators, a comprehensive test score can be obtained, which more accurately reflects overall performance.

[0096] In this embodiment, the test score threshold is set based on historical test data and user feedback. By analyzing historical test data, a reasonable threshold range can be determined to ensure that the test score is authentic and reliable. User feedback is also an important reference factor, which can help adjust the threshold to better reflect users' actual experience and expectations.

[0097] Specifically, when comparing the test score with the test score threshold and determining whether to modify the voice interaction parameter based on the comparison result, the method further includes:

[0098] When the test score is greater than the test score threshold, determining not to modify the voice interaction parameter;

[0099] When the test score is less than or equal to the test score threshold, it is determined to modify the voice interaction parameter.

[0100] It's understandable that by setting a reasonable test score threshold, voice interaction parameters can be flexibly adjusted based on actual test conditions. When the test score exceeds the threshold, it indicates that the current voice interaction parameters are sufficient to meet user needs and environmental requirements and no adjustment is required. Conversely, if the test score is below or equal to the threshold, it indicates that the expected interaction effect cannot be achieved under the current parameter settings, and the voice interaction parameters need to be corrected.

[0101] Specifically, the real-time voice feature value is compared with the historical voice data, a correction coefficient is determined according to the comparison result, and the voice interaction parameter is corrected according to the correction coefficient to obtain the corrected voice interaction parameter, including:

[0102] Calculating the maximum similarity between the real-time speech feature value and the historical speech data, comparing the maximum similarity with a maximum similarity threshold, and determining the correction coefficient based on the comparison result;

[0103] The maximum similarity is obtained by the following formula:

[0104]

[0105] Among them, M max represents the maximum similarity; X i represents the i-th real-time speech feature vector; Y j represents the jth historical speech feature vector; m represents the number of real-time speech feature vectors; n represents the number of historical speech feature vectors.

[0106] As can be seen, the real-time speech feature vector refers to the feature vector extracted from the current speech data, while the historical speech feature vector refers to the feature vector collected and stored in previous speech interaction tests. By calculating the maximum similarity between the real-time speech feature vector and the historical speech feature vector, we can evaluate the degree of match between the current speech data and the historical data.

[0107] Specifically, comparing the maximum similarity with the maximum similarity threshold and determining the correction coefficient according to the comparison result includes:

[0108] Comparing the maximum similarity with a first maximum similarity threshold and a second maximum similarity threshold, and determining the correction coefficient according to the comparison result; wherein the first maximum similarity threshold is less than the second maximum similarity threshold;

[0109] Setting a correction coefficient interval, wherein the correction coefficient interval includes a first correction coefficient, a second correction coefficient, and a third correction coefficient;

[0110] When the maximum similarity is less than or equal to the first maximum similarity threshold, determining the correction coefficient of the voice interaction parameter as the first correction coefficient, and taking the product of the first correction coefficient and the voice interaction parameter as the corrected voice interaction parameter;

[0111] When the maximum similarity is greater than the first maximum similarity threshold and less than or equal to the second maximum similarity threshold, determining the correction coefficient of the voice interaction parameter to be a second correction coefficient, and using the product of the second correction coefficient and the voice interaction parameter as the corrected voice interaction parameter;

[0112] When the maximum similarity is greater than the second maximum similarity threshold, the correction coefficient of the voice interaction parameter is determined to be a third correction coefficient, and the product value of the third correction coefficient and the voice interaction parameter is used as the corrected voice interaction parameter.

[0113] It is understandable that by setting different similarity thresholds and correction coefficient intervals, the voice interaction parameters can be flexibly adjusted according to the degree of match between real-time voice data and historical data. When the similarity between real-time voice data and historical data is low, a larger correction coefficient is used to significantly adjust the voice interaction parameters, thereby improving the adaptability and accuracy of the test. On the contrary, if the similarity is high, it indicates that the current voice interaction parameters are already relatively appropriate. At this time, a smaller correction coefficient is used for fine-tuning to maintain the stability and continuity of the test. In this way, it can ensure that a high-quality voice interaction experience can be provided in a variety of specialized scenarios.

[0114] Specifically, determining the interaction impact index between the user and the smart home to be tested based on the distance data, behavior data, and device status data includes:

[0115] The interaction impact index is obtained by the following formula:

[0116]

[0117] Among them, I represents the interaction influence index; D represents distance data; B represents behavior data; E represents device status data; α1 represents the distance weight coefficient; α2 represents the behavior weight coefficient; α3 represents the device status weight coefficient; β1 represents the first influence coefficient; β2 represents the second influence coefficient; β3 represents the third influence coefficient.

[0118] It will be appreciated that in this embodiment, α1, α2, and α3 represent the weighting coefficients for distance, behavior, and device status data, respectively. Their selection is based on in-depth research into the impact of these factors on the voice interaction process. For example, the weighting coefficient α1 for distance data is preferably 0.3. This is because the distance between the user and the device directly affects the strength and clarity of the voice signal; the closer the distance, the easier it is to accurately capture the voice signal. The weighting coefficient α2 for behavior data is preferably 0.4. This is because user behavior patterns, such as speaking speed and volume, have a significant impact on the accuracy of voice recognition. The weighting coefficient α3 for device status data is preferably 0.3. This is because the operating status of smart home devices, such as whether they are in silent mode or playing music, also affects voice interaction. β1, β2, and β3 represent the first, second, and third influence coefficients related to distance, behavior, and device status data, respectively. Their setting is intended to further refine the calculation of the interaction impact index. For example, β1 can be set to 0.5 to emphasize the importance of distance data in interaction influence; β2 can be set to 0.3 to reflect the moderate influence of behavioral data on interaction; and β3 can be set to 0.2 to indicate the minimal influence of device status data on interaction. By setting these weight coefficients and influence coefficients, we ensure that the interaction influence index accurately reflects the interaction between users and smart home devices, thus providing a scientific basis for adjusting voice interaction parameters. This data-driven adjustment mechanism not only improves the adaptability and personalization of voice interaction testing, but also helps better meet the specific needs of users in actual use.

[0119] Specifically, determining whether to compensate the modified voice interaction parameter according to the interaction influence index includes:

[0120] Subtracting the interaction impact index from the interaction impact index threshold to obtain an index difference;

[0121] Comparing the index difference with an index difference threshold, and determining whether to compensate the modified voice interaction parameter based on the comparison result;

[0122] When the index difference is greater than or equal to the index difference threshold, determining to compensate the modified voice interaction parameter;

[0123] When the index difference is less than the index difference threshold, it is determined that the modified voice interaction parameter is not compensated.

[0124] Specifically, determining a compensation coefficient according to the interaction influence index, and compensating the modified voice interaction parameter according to the compensation coefficient includes:

[0125] Comparing the index difference threshold with a first index difference threshold and a second index difference threshold, and determining the compensation coefficient according to the comparison result; wherein the first index difference threshold is smaller than the second index difference threshold;

[0126] Setting a compensation coefficient interval, wherein the compensation coefficient interval includes a first compensation coefficient, a second compensation coefficient, and a third compensation coefficient;

[0127] When the exponent difference is less than or equal to the first exponent difference threshold, determining the compensation coefficient of the modified voice interaction parameter to be a first compensation coefficient, and taking the product of the first compensation coefficient and the modified voice interaction parameter as the compensated voice interaction parameter;

[0128] When the exponent difference is greater than the first exponent difference threshold and less than or equal to the second exponent difference threshold, determining the compensation coefficient of the modified voice interaction parameter as the second compensation coefficient, and using the product value of the second compensation coefficient and the modified voice interaction parameter as the compensated voice interaction parameter;

[0129] When the exponent difference is greater than the second exponent difference threshold, the compensation coefficient of the modified voice interaction parameter is determined to be a third compensation coefficient, and the product value of the third compensation coefficient and the modified voice interaction parameter is used as the compensated voice interaction parameter.

[0130] It is understandable that by setting different index difference thresholds and compensation coefficient intervals, the modified voice interaction parameters can be flexibly compensated according to the difference between the interaction influence index and the threshold. When the difference between the interaction influence index and the threshold is large, a larger compensation coefficient is used to significantly adjust the correction parameters, thereby improving the adaptability and personalization of the voice interaction. On the contrary, if the difference is small, it means that the current correction parameters are already more appropriate. At this time, a smaller compensation coefficient is used for fine-tuning to maintain the stability and continuity of the voice interaction. In this way, it can ensure that a high-quality voice interaction experience can be provided in a variety of specialized scenarios.

[0131] See Figure 2 As shown, in some embodiments of the present application, this embodiment provides a smart home voice interaction testing device for specialized scene recognition, including:

[0132] a determination module configured to determine a specific scenario to be tested, extract scenario feature data of the specific scenario to be tested, and determine voice interaction parameters based on the scenario feature data; and further configured to perform a voice interaction test on the smart home to be tested using real-time voice data in the specific scenario to be tested, obtain voice interaction data, and obtain a test score based on the voice interaction data;

[0133] A first judgment module is configured to compare the test score with a test score threshold, and determine whether to modify the voice interaction parameter according to the comparison result;

[0134] a correction module configured to, when determining to correct the voice interaction parameter, perform feature extraction on the real-time voice data to obtain a real-time voice feature value; compare the real-time voice feature value with historical voice data, determine a correction coefficient based on the comparison result, and correct the voice interaction parameter based on the correction coefficient to obtain a corrected voice interaction parameter;

[0135] A second judgment module is configured to, after determining the modified voice interaction parameter, collect distance data between the user and the smart home to be tested, the user's behavior data, and the device status data of the smart home to be tested, and determine an interaction impact index between the user and the smart home to be tested based on the distance data, behavior data, and device status data; and determine whether to compensate for the modified voice interaction parameter based on the interaction impact index;

[0136] The compensation module is configured to determine a compensation coefficient according to the interaction influence index when determining to compensate the modified voice interaction parameter, and compensate the modified voice interaction parameter according to the compensation coefficient to obtain a compensated voice interaction parameter.

[0137] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0138] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0139] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0140] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A smart home voice interaction testing method, characterized in that: include: Determining a specialized scenario to be tested, extracting scenario feature data of the specialized scenario to be tested, and determining voice interaction parameters based on the scenario feature data; The scene feature data includes noise feature data and spatial echo feature data; the voice interaction parameters include recognition threshold, noise suppression parameter and voice enhancement coefficient; In the specific scenario to be tested, using real-time voice data to perform a voice interaction test on the smart home to be tested, obtaining voice interaction data, and obtaining a test score based on the voice interaction data; Comparing the test score with a test score threshold, and determining whether to modify the voice interaction parameter based on the comparison result; When it is determined that the voice interaction parameter is to be modified, feature extraction is performed on the real-time voice data to obtain a real-time voice feature value; Comparing the real-time voice feature value with historical voice data, determining a correction coefficient based on the comparison result, and correcting the voice interaction parameter based on the correction coefficient to obtain a corrected voice interaction parameter; After determining the modified voice interaction parameters, collecting distance data between the user and the smart home to be tested, the user's behavior data, and the device status data of the smart home to be tested, and determining an interaction impact index between the user and the smart home to be tested based on the distance data, behavior data, and device status data; and determining whether to compensate for the modified voice interaction parameters based on the interaction impact index; When it is determined that the modified voice interaction parameter is to be compensated, a compensation coefficient is determined according to the interaction influence index, and the modified voice interaction parameter is compensated according to the compensation coefficient to obtain a compensated voice interaction parameter; Determining the interaction impact index between the user and the smart home to be tested based on the distance data, the behavior data, and the device status data includes: The interaction impact index is obtained by the following formula: Among them, I represents the interaction influence index; D represents distance data; B represents behavior data; E represents device status data; α1 represents the distance weight coefficient; α2 represents the behavior weight coefficient; α3 represents the device status weight coefficient; β1 represents the first influence coefficient; β2 represents the second influence coefficient; β3 represents the third influence coefficient.

2. The smart home voice interaction testing method according to claim 1, characterized in that: Determining the voice interaction parameters based on the scene feature data includes: The recognition threshold is obtained by the following formula: The noise suppression parameter is obtained by the following formula: The speech enhancement coefficient is obtained by the following formula: Among them, T represents the recognition threshold; T0 represents the basic recognition threshold; a1 represents the noise recognition weight coefficient; xn represents the noise characteristic value; bn represents the noise standard value; a2 represents the spatial recognition weight coefficient; xr represents the current spatial echo characteristic value; br represents the current spatial echo standard value; N represents the noise suppression parameter; N0 represents the basic noise suppression parameter; b1 represents the noise suppression weight coefficient; b2 represents the spatial suppression weight coefficient; C represents the speech enhancement coefficient; C0 represents the basic speech enhancement coefficient; k1 represents the speech enhancement weight coefficient; k2 represents the spatial enhancement weight coefficient.

3. The smart home voice interaction testing method according to claim 1, characterized in that: Comparing the test score with the test score threshold, and determining whether to modify the voice interaction parameter according to the comparison result, includes: The test score is obtained by the following formula: Among them, S represents the test score; c1 represents the weight coefficient of speech recognition accuracy; Rmax represents the maximum value of speech recognition accuracy; R1 represents the speech recognition accuracy; c2 represents the weight coefficient of response time; T2 represents the response time; Tmin represents the minimum value of response time; c3 represents the weight coefficient of response accuracy; Amax represents the maximum value of response accuracy; A3 represents response accuracy.

4. The smart home voice interaction testing method according to claim 3, characterized in that: When comparing the test score with the test score threshold and determining whether to modify the voice interaction parameter based on the comparison result, the method further includes: When the test score is greater than the test score threshold, determining not to modify the voice interaction parameter; When the test score is less than or equal to the test score threshold, it is determined to modify the voice interaction parameter.

5. The smart home voice interaction testing method according to claim 1, characterized in that: Comparing the real-time voice feature value with historical voice data, determining a correction coefficient based on the comparison result, and correcting the voice interaction parameter based on the correction coefficient to obtain the corrected voice interaction parameter, including: Calculating the maximum similarity between the real-time speech feature value and the historical speech data, comparing the maximum similarity with a maximum similarity threshold, and determining the correction coefficient based on the comparison result; The maximum similarity is obtained by the following formula: Among them, M max represents the maximum similarity; X i represents the i-th real-time speech feature vector; Y j represents the jth historical speech feature vector; m represents the number of real-time speech feature vectors; n represents the number of historical speech feature vectors.

6. The smart home voice interaction testing method according to claim 5, characterized in that: Comparing the maximum similarity with the maximum similarity threshold and determining the correction coefficient according to the comparison result includes: Comparing the maximum similarity with a first maximum similarity threshold and a second maximum similarity threshold, and determining the correction coefficient according to the comparison result; wherein the first maximum similarity threshold is less than the second maximum similarity threshold; Setting a correction coefficient interval, wherein the correction coefficient interval includes a first correction coefficient, a second correction coefficient, and a third correction coefficient; When the maximum similarity is less than or equal to the first maximum similarity threshold, determining the correction coefficient of the voice interaction parameter as the first correction coefficient, and taking the product of the first correction coefficient and the voice interaction parameter as the corrected voice interaction parameter; When the maximum similarity is greater than the first maximum similarity threshold and less than or equal to the second maximum similarity threshold, determining the correction coefficient of the voice interaction parameter to be a second correction coefficient, and using the product of the second correction coefficient and the voice interaction parameter as the corrected voice interaction parameter; When the maximum similarity is greater than the second maximum similarity threshold, the correction coefficient of the voice interaction parameter is determined to be a third correction coefficient, and the product value of the third correction coefficient and the voice interaction parameter is used as the corrected voice interaction parameter.

7. The smart home voice interaction testing method according to claim 1, characterized in that: When determining whether to compensate the modified voice interaction parameter according to the interaction influence index, the method includes: Subtracting the interaction impact index from the interaction impact index threshold to obtain an index difference; Comparing the index difference with an index difference threshold, and determining whether to compensate the modified voice interaction parameter based on the comparison result; When the index difference is greater than or equal to the index difference threshold, determining to compensate the modified voice interaction parameter; When the index difference is less than the index difference threshold, it is determined that the modified voice interaction parameter is not compensated.

8. The smart home voice interaction testing method according to claim 7, characterized in that: Determining a compensation coefficient according to the interaction influence index, and compensating the modified voice interaction parameter according to the compensation coefficient, includes: Comparing the index difference threshold with a first index difference threshold and a second index difference threshold, and determining the compensation coefficient according to the comparison result; wherein the first index difference threshold is smaller than the second index difference threshold; Setting a compensation coefficient interval, wherein the compensation coefficient interval includes a first compensation coefficient, a second compensation coefficient, and a third compensation coefficient; When the exponent difference is less than or equal to the first exponent difference threshold, determining the compensation coefficient of the modified voice interaction parameter to be a first compensation coefficient, and taking the product of the first compensation coefficient and the modified voice interaction parameter as the compensated voice interaction parameter; When the exponent difference is greater than the first exponent difference threshold and less than or equal to the second exponent difference threshold, determining the compensation coefficient of the modified voice interaction parameter as the second compensation coefficient, and using the product value of the second compensation coefficient and the modified voice interaction parameter as the compensated voice interaction parameter; When the exponent difference is greater than the second exponent difference threshold, the compensation coefficient of the modified voice interaction parameter is determined to be a third compensation coefficient, and the product value of the third compensation coefficient and the modified voice interaction parameter is used as the compensated voice interaction parameter.

9. A smart home voice interaction testing device, applied to the smart home voice interaction testing method according to any one of claims 1 to 8, characterized in that: include: a determination module configured to determine a specific scenario to be tested, extract scenario feature data of the specific scenario to be tested, and determine voice interaction parameters based on the scenario feature data; and further configured to perform a voice interaction test on the smart home to be tested using real-time voice data in the specific scenario to be tested, obtain voice interaction data, and obtain a test score based on the voice interaction data; A first judgment module is configured to compare the test score with a test score threshold, and determine whether to modify the voice interaction parameter according to the comparison result; a correction module, configured to, when determining to correct the voice interaction parameter, perform feature extraction on the real-time voice data to obtain a real-time voice feature value; Comparing the real-time voice feature value with historical voice data, determining a correction coefficient based on the comparison result, and correcting the voice interaction parameter based on the correction coefficient to obtain a corrected voice interaction parameter; A second judgment module is configured to, after determining the modified voice interaction parameter, collect distance data between the user and the smart home to be tested, the user's behavior data, and the device status data of the smart home to be tested, and determine an interaction impact index between the user and the smart home to be tested based on the distance data, behavior data, and device status data; and determine whether to compensate for the modified voice interaction parameter based on the interaction impact index; The compensation module is configured to determine a compensation coefficient according to the interaction influence index when determining to compensate the modified voice interaction parameter, and compensate the modified voice interaction parameter according to the compensation coefficient to obtain a compensated voice interaction parameter.

Citation Information

Patent Citations

  • Threshold adaptive adjustment method for off line speech recognition

    CN108550365A

  • Stochastic modeling of user interactions with a detection system

    US9899021B1