Smart home voice interaction test method and device
By identifying specific scenes in smart homes, extracting scene feature data, and adjusting voice interaction parameters, the stability and accuracy of voice interaction testing in complex environments are achieved, and the problem of poor testing results in existing technologies in specific scenarios is solved.
Patent Information
- Application Number
- CN202510370640.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-27
Smart Images

Figure CN120220649A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice interaction testing, and in particular, to a smart home voice interaction testing method and device. Background Art
[0002] With the continuous progress and wide popularization of smart home technologies, voice interaction has increasingly become one of the key ways for users to communicate and interact with smart home devices. Currently, existing voice interaction testing means perform excellently in general application scenarios and can better meet the basic needs of users. However, in specific and complex scenarios, these testing methods expose limitations and are difficult to achieve stable and accurate testing effects. Such specialized scenarios cover complex home space layouts, diverse device type environments, as well as various complex and changeable scenarios such as multi-room environments, open and closed spaces, and personalized space configurations.
[0003] In these specialized scenarios, due to the diversity and complexity of the environment, voice interaction testing faces multiple challenges. For example, in a multi-room environment, users may need to interact with devices in different rooms, which requires voice interaction testing to accurately identify and respond to commands from different rooms. At the same time, in open and closed spaces, problems such as the propagation characteristics of sound and echoes will also affect the accuracy of voice interaction.
[0004] Therefore, it is necessary to design a smart home voice interaction testing method and device for specialized scenario recognition to solve the problems existing in the current technology. Summary of the Invention
[0005] In view of this, the present invention proposes a smart home voice interaction testing method and device, aiming to improve the stability and accuracy of voice interaction testing.
[0006] On the one hand, the present invention proposes a smart home voice interaction testing method, including the following steps:
[0007] S100: Determine the specialized scenario to be tested, extract the scenario feature data of the specialized scenario to be tested, and determine the voice interaction parameters according to the scenario feature data;
[0008] S200: Under the specialized scenario to be tested, use real-time voice data to perform voice interaction testing on the smart home to be tested, obtain voice interaction data, and obtain a test score according to the voice interaction data;
[0009] S300: Compare the test score with a test score threshold, and judge whether to correct the voice interaction parameters according to the comparison result;
[0010] S400: When it is determined to correct the voice interaction parameter, extract features from the real-time voice data to obtain real-time voice feature values; compare the real-time voice feature values with historical voice data, determine a correction coefficient according to the comparison result, and correct the voice interaction parameter according to the correction coefficient to obtain a corrected voice interaction parameter;
[0011] S500: After determining the corrected voice interaction parameter, collect distance data between the user and the smart home to be tested, the user's behavior data, and device status data of the smart home to be tested, and determine an interaction influence index between the user and the smart home to be tested according to the distance data, behavior data, and device status data; determine whether to compensate the corrected voice interaction parameter according to the interaction influence index;
[0012] S600: When it is determined to compensate the corrected voice interaction parameter, determine a compensation coefficient according to the interaction influence index, and compensate the corrected voice interaction parameter according to the compensation coefficient to obtain a compensated voice interaction parameter.
[0013] Further, when determining the voice interaction parameter according to the scene feature data, it includes:
[0014] The scene feature data includes noise feature data and spatial echo feature data;
[0015] The voice interaction parameters include recognition threshold, noise suppression parameter, and voice enhancement coefficient;
[0016] The recognition threshold is obtained by the following formula:
[0017]
[0018] The noise suppression parameter is obtained by the following formula:
[0019]
[0020] The voice enhancement coefficient is obtained by the following formula:
[0021]
[0022] Among them, T represents the recognition threshold; T0 represents the basic recognition threshold; a1 represents the noise recognition weight coefficient; xn represents the noise eigenvalue; bn represents the noise standard value; a2 represents the space recognition weight coefficient; xr represents the current space echo eigenvalue; br represents the current space echo standard value; N represents the noise suppression parameter; N0 represents the basic noise suppression parameter; b1 represents the noise suppression weight coefficient; b2 represents the space suppression weight coefficient; C represents the voice enhancement coefficient; C0 represents the basic voice enhancement coefficient; k1 represents the voice enhancement weight coefficient; k2 represents the space enhancement weight coefficient.
[0023] Further, when comparing the test score with the test score threshold and determining whether to correct the voice interaction parameter according to the comparison result, it includes:
[0024] The test score is obtained by the following formula:
[0025]
[0026] Among them, S represents the test score; c1 represents the weight coefficient of speech recognition accuracy; Rmax represents the maximum value of speech recognition accuracy; R1 represents the speech recognition accuracy; c2 represents the weight coefficient of response time; T2 represents the response time; Tmin represents the minimum value of response time; c3 represents the weight coefficient of response accuracy; Amax represents the maximum value of response accuracy; A3 represents the response accuracy.
[0027] Further, when comparing the test score with the test score threshold and determining whether to correct the voice interaction parameter according to the comparison result, it also includes:
[0028] When the test score is greater than the test score threshold, it is determined not to correct the voice interaction parameter;
[0029] When the test score is less than or equal to the test score threshold, it is determined to correct the voice interaction parameter.
[0030] Further, when comparing the real-time voice eigenvalue with the historical voice data, determining the correction coefficient according to the comparison result, and correcting the voice interaction parameter according to the correction coefficient to obtain the corrected voice interaction parameter, it includes:
[0031] Calculate the maximum similarity between the real-time voice eigenvalue and the historical voice data, compare the maximum similarity with the maximum similarity threshold, and determine the correction coefficient according to the comparison result;
[0032] The maximum similarity is obtained by the following formula:
[0033]
[0034] Among them, M max represents the maximum similarity; X i represents the i-th real-time speech feature vector; Y j represents the j-th historical speech feature vector; m represents the number of real-time speech feature vectors; n represents the number of historical speech feature vectors.
[0035] Furthermore, when comparing the maximum similarity with a maximum similarity threshold and determining the correction coefficient according to the comparison result, it includes:
[0036] Comparing the maximum similarity with a first maximum similarity threshold and a second maximum similarity threshold, and determining the correction coefficient according to the comparison result; wherein, the first maximum similarity threshold is less than the second maximum similarity threshold;
[0037] Set a correction coefficient interval, and the correction coefficient interval includes a first correction coefficient, a second correction coefficient, and a third correction coefficient;
[0038] When the maximum similarity is less than or equal to the first maximum similarity threshold, determine that the correction coefficient of the voice interaction parameter is the first correction coefficient, and use the product value of the first correction coefficient and the voice interaction parameter as the corrected voice interaction parameter;
[0039] When the maximum similarity is greater than the first maximum similarity threshold and less than or equal to the second maximum similarity threshold, determine that the correction coefficient of the voice interaction parameter is the second correction coefficient, and use the product value of the second correction coefficient and the voice interaction parameter as the corrected voice interaction parameter;
[0040] When the maximum similarity is greater than the second maximum similarity threshold, determine that the correction coefficient of the voice interaction parameter is the third correction coefficient, and use the product value of the third correction coefficient and the voice interaction parameter as the corrected voice interaction parameter.
[0041] Furthermore, when determining the interaction influence index between the user and the smart home to be tested according to the distance data, behavior data, and device status data, it includes:
[0042] The interaction influence index is obtained through the following formula:
[0043]
[0044] Among them, I represents the interaction influence index; D represents the distance data; B represents the behavior data; E represents the device status data; α1 represents the distance weight coefficient; α2 represents the behavior weight coefficient; α3 represents the device status weight coefficient; β1 represents the first influence coefficient; β2 represents the second influence coefficient; β3 represents the third influence coefficient.
[0045] Further, when determining whether to compensate the corrected voice interaction parameter according to the interaction influence index, it includes:
[0046] Subtract the interaction influence index from the interaction influence index threshold to obtain an index difference;
[0047] Compare the index difference with the index difference threshold, and determine whether to compensate the corrected voice interaction parameter according to the comparison result;
[0048] When the index difference is greater than or equal to the index difference threshold, it is determined to compensate the corrected voice interaction parameter;
[0049] When the index difference is less than the index difference threshold, it is determined not to compensate the corrected voice interaction parameter.
[0050] Further, when determining a compensation coefficient according to the interaction influence index and compensating the corrected voice interaction parameter according to the compensation coefficient, it includes:
[0051] Compare the index difference threshold with a first index difference threshold and a second index difference threshold, and determine the compensation coefficient according to the comparison result; wherein, the first index difference threshold is less than the second index difference threshold;
[0052] Set a compensation coefficient range, and the compensation coefficient range includes a first compensation coefficient, a second compensation coefficient, and a third compensation coefficient;
[0053] When the index difference is less than or equal to the first index difference threshold, determine that the compensation coefficient of the corrected voice interaction parameter is the first compensation coefficient, and use the product value of the first compensation coefficient and the corrected voice interaction parameter as the compensated voice interaction parameter;
[0054] When the index difference is greater than the first index difference threshold and less than or equal to the second index difference threshold, determine that the compensation coefficient of the corrected voice interaction parameter is the second compensation coefficient, and use the product value of the second compensation coefficient and the corrected voice interaction parameter as the compensated voice interaction parameter;
[0055] When the index difference is greater than the second index difference threshold, determine that the compensation coefficient of the corrected voice interaction parameter is the third compensation coefficient, and use the product value of the third compensation coefficient and the corrected voice interaction parameter as the compensated voice interaction parameter.
[0056] Compared with the prior art, the beneficial effects of the present invention are as follows: The smart home voice interaction test method for specialized scenario recognition provided by the present invention can improve the stability and accuracy of voice interaction tests; by determining the specialized scenario to be tested, extracting the scenario feature data of the specialized scenario to be tested, and determining the voice interaction parameters according to the scenario feature data, the adaptability and accuracy of the test can be improved, ensuring good performance in multiple scenarios; in the specialized scenario to be tested, using real-time voice data to perform voice interaction tests on the smart home to be tested, obtaining voice interaction data, and obtaining test scores according to the voice interaction data, which can clearly understand the test situation; comparing the test scores with the test score threshold, and judging whether to correct the voice interaction parameters according to the comparison result, which helps to continuously optimize and adjust in practical applications to adapt to the needs of different users and environmental changes; through the feature extraction of real-time voice data and the comparison with historical voice data, the voice interaction parameters can be further refined to ensure high recognition rates in multiple noise environments; by collecting distance data, behavior data, and device status data between users and smart homes, the user interaction experience can be evaluated more comprehensively, so as to make necessary compensation for the voice interaction parameters, better adapt to the specific needs and environmental changes of users, and thus provide a more natural and smooth interaction experience.
[0057] On the other hand, the present invention also proposes a smart home voice interaction test device for specialized scenario recognition, including:
[0058] A determination module, configured to determine the specialized scenario to be tested, extract the scenario feature data of the specialized scenario to be tested, and determine the voice interaction parameters according to the scenario feature data; and further configured to, in the specialized scenario to be tested, use real-time voice data to perform voice interaction tests on the smart home to be tested, obtain voice interaction data, and obtain test scores according to the voice interaction data;
[0059] A first judgment module, configured to compare the test scores with the test score threshold, and judge whether to correct the voice interaction parameters according to the comparison result;
[0060] A correction module, configured to, when it is determined to correct the voice interaction parameters, extract features from the real-time voice data to obtain real-time voice feature values; compare the real-time voice feature values with historical voice data, determine a correction coefficient according to the comparison result, and correct the voice interaction parameters according to the correction coefficient to obtain corrected voice interaction parameters;
[0061] A second judgment module, configured to collect distance data between a user and a to-be-tested smart home, behavior data of the user, and device status data of the to-be-tested smart home after determining the corrected voice interaction parameters, and determine an interaction influence index between the user and the to-be-tested smart home according to the distance data, behavior data, and device status data; and determine whether to compensate the corrected voice interaction parameters according to the interaction influence index.
[0062] A compensation module, configured to determine a compensation coefficient according to the interaction influence index and compensate the corrected voice interaction parameters according to the compensation coefficient to obtain compensated voice interaction parameters when it is determined that the corrected voice interaction parameters need to be compensated.
[0063] It can be understood that the above-mentioned smart home voice interaction test method and device for specific scenario recognition have the same beneficial effects, which will not be elaborated here. Description of the Drawings
[0064] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0065] Figure 1 is a flowchart of the smart home voice interaction test method for specific scenario recognition provided by an embodiment of the present invention;
[0066] Figure 2 is a structural diagram of the smart home voice interaction test device for specific scenario recognition provided by an embodiment of the present invention. Detailed Embodiments
[0067] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other. Hereinafter, the present invention will be described in detail with reference to the drawings and in conjunction with the embodiments.
[0068] Refer to Figure 1 As shown, in some embodiments of the present application, this embodiment provides a smart home voice interaction test method for specific scenario recognition, including the following steps:
[0069] S100: Determine the to-be-tested specific scenario, extract the scenario feature data of the to-be-tested specific scenario, and determine the voice interaction parameters according to the scenario feature data;
[0070] S200: In the to-be-tested specific scenario, perform a voice interaction test on the to-be-tested smart home using real-time voice data, obtain voice interaction data, and obtain a test score according to the voice interaction data;
[0071] S300: Compare the test score with a test score threshold, and determine whether to correct the voice interaction parameters according to the comparison result;
[0072] S400: When it is determined to correct the voice interaction parameters, extract features from the real-time voice data to obtain real-time voice feature values; compare the real-time voice feature values with historical voice data, determine a correction coefficient according to the comparison result, and correct the voice interaction parameters according to the correction coefficient to obtain corrected voice interaction parameters;
[0073] S500: After determining the corrected voice interaction parameters, collect the distance data between the user and the to-be-tested smart home, the behavior data of the user, and the device status data of the to-be-tested smart home, and determine the interaction influence index between the user and the to-be-tested smart home according to the distance data, behavior data, and device status data; determine whether to compensate the corrected voice interaction parameters according to the interaction influence index;
[0074] S600: When it is determined to compensate the corrected voice interaction parameters, determine a compensation coefficient according to the interaction influence index, and compensate the corrected voice interaction parameters according to the compensation coefficient to obtain compensated voice interaction parameters.
[0075] In this embodiment, the to-be-tested specific scenario is preferably a high-noise environment. For example, in a high-noise environment, it is necessary to be able to distinguish and ignore background noise and accurately recognize the user's voice commands.
[0076] It can be understood that the smart home voice interaction test method provided in this embodiment for specialized scenario recognition can improve the stability and accuracy of voice interaction tests; by determining the specialized scenario to be tested, extracting the scenario feature data of the specialized scenario to be tested, and determining the voice interaction parameters according to the scenario feature data, the adaptability and accuracy of the test can be improved, ensuring good performance in multiple scenarios; in the specialized scenario to be tested, real-time voice data is used to conduct voice interaction tests on the smart home to be tested, obtaining voice interaction data, and obtaining test scores based on the voice interaction data, which can clearly understand the test situation; comparing the test scores with the test score threshold, and judging whether to correct the voice interaction parameters according to the comparison result, which helps to continuously optimize and adjust in actual applications to adapt to the needs of different users and environmental changes; through feature extraction of real-time voice data and comparison with historical voice data, the voice interaction parameters can be further refined to ensure high recognition rates in multiple noise environments; by collecting distance data, behavior data, and device status data between users and smart homes, the user interaction experience can be evaluated more comprehensively, thereby making necessary compensation for the voice interaction parameters to better adapt to the specific needs and environmental changes of users, and thus providing a more natural and smooth interaction experience.
[0077] Specifically, when determining the voice interaction parameters according to the scenario feature data, it includes:
[0078] The scenario feature data includes noise feature data and spatial echo feature data;
[0079] The voice interaction parameters include recognition threshold, noise suppression parameter, and voice enhancement coefficient;
[0080] The recognition threshold is obtained by the following formula:
[0081]
[0082] The noise suppression parameter is obtained by the following formula:
[0083]
[0084] The voice enhancement coefficient is obtained by the following formula:
[0085]
[0086] Wherein, T represents the recognition threshold; T0 represents the basic recognition threshold; a1 represents the noise recognition weight coefficient; xn represents the noise eigenvalue; bn represents the noise standard value; a2 represents the spatial recognition weight coefficient; xr represents the current spatial echo eigenvalue; br represents the current spatial echo standard value; N represents the noise suppression parameter; N0 represents the basic noise suppression parameter; b1 represents the noise suppression weight coefficient; b2 represents the spatial suppression weight coefficient; C represents the speech enhancement coefficient; C0 represents the basic speech enhancement coefficient; k1 represents the speech enhancement weight coefficient; k2 represents the spatial enhancement weight coefficient.
[0087] In this embodiment, the recognition threshold is a voice interaction parameter used to determine whether to respond to an input voice signal; the noise suppression parameter is used to reduce the impact of background noise on speech recognition; and the speech enhancement coefficient is used to improve the quality of the voice signal to ensure that user instructions can still be clearly recognized in a noisy environment. By adjusting these parameters, user voice instructions can be recognized more accurately, and good interaction performance can be maintained even in highly noisy specific scenarios.
[0088] In this embodiment, the basic recognition threshold refers to the threshold preset in an ideal environment without noise and echo. The basic noise suppression parameter and the basic speech enhancement coefficient are also set under similar ideal conditions, and they serve as a reference point for adjustment according to scene feature data in actual applications. In this way, various different environments can be adapted, thus providing a more stable and reliable voice interaction experience in the smart home scenario.
[0089] In this embodiment, a1 represents the noise recognition weight coefficient, and a1 is preferably between 0.1 and 0.3. The larger its value, the greater the influence of the noise feature on the recognition threshold; a2 represents the spatial recognition weight coefficient, and a2 is preferably between 0.1 and 0.3. The larger its value, the greater the influence of the spatial echo feature on the recognition threshold. Similarly, b1 and b2 represent the noise suppression weight coefficient and the spatial suppression weight coefficient respectively, and their preferred ranges are also between 0.1 and 0.3 to ensure that the noise suppression parameter and the speech enhancement coefficient can be effectively adjusted to adapt to different noise and echo environments. k1 and k2 represent the speech enhancement weight coefficient respectively, and the preferred range is 0.1 to 0.3 to ensure that the speech enhancement coefficient can be appropriately adjusted according to the actual situation of noise and echo, thereby improving the clarity and recognizability of the voice signal. Through the reasonable selection and adjustment of these weight coefficients, various specific scenarios can be flexibly handled to achieve more accurate speech recognition and a better interaction experience.
[0090] Specifically, when comparing the test score with the test score threshold and determining whether to correct the voice interaction parameter according to the comparison result, it includes:
[0091] The test score is obtained by the following formula:
[0092]
[0093] Wherein, S represents the test score; c1 represents the weight coefficient of speech recognition accuracy; Rmax represents the maximum value of speech recognition accuracy; R1 represents the speech recognition accuracy; c2 represents the weight coefficient of response time; T2 represents the response time; Tmin represents the minimum value of response time; c3 represents the weight coefficient of response accuracy; Amax represents the maximum value of response accuracy; A3 represents the response accuracy.
[0094] In this embodiment, c1 represents the weight coefficient of speech recognition accuracy, preferably 0.4, c2 represents the weight coefficient of response time, preferably 0.3, and c3 represents the weight coefficient of response accuracy, preferably 0.3. Such weight allocation ensures that the speech recognition accuracy accounts for a large proportion in the test score, while not ignoring the importance of response time and response accuracy.
[0095] It can be understood that determining the test score according to the speech recognition accuracy, response time, and response accuracy can comprehensively evaluate the performance of the speech interaction test. The speech recognition accuracy reflects the ability to understand the user's speech commands, the response time reflects the speed of processing commands, and the response accuracy focuses on the correctness of executing commands. By integrating these three indicators, a comprehensive test score can be obtained, so as to more accurately reflect the overall performance.
[0096] In this embodiment, the test score threshold is set according to historical test data and user feedback. By analyzing historical test data, a reasonable threshold range can be determined to ensure that the test score is true and reliable. At the same time, user feedback is also an important reference factor, which can help adjust the threshold to make it more in line with the actual experience and expectations of users.
[0097] Specifically, when comparing the test score with the test score threshold and determining whether to correct the speech interaction parameters according to the comparison result, it further includes:
[0098] When the test score is greater than the test score threshold, it is determined not to correct the speech interaction parameters;
[0099] When the test score is less than or equal to the test score threshold, it is determined to correct the speech interaction parameters.
[0100] It can be understood that by setting a reasonable test score threshold, the voice interaction parameters can be flexibly adjusted according to the actual test situation. When the test score exceeds the threshold, it indicates that the current voice interaction parameters can already meet the user's needs and environmental requirements, and no adjustment is required. On the contrary, if the test score is lower than or equal to the threshold, it means that the expected interaction effect cannot be achieved under the current parameter settings, and at this time, the voice interaction parameters need to be corrected.
[0101] Specifically, when comparing the real-time voice feature value with the historical voice data, determining a correction coefficient according to the comparison result, and correcting the voice interaction parameters according to the correction coefficient to obtain corrected voice interaction parameters, it includes:
[0102] Calculating the maximum similarity between the real-time voice feature value and the historical voice data, comparing the maximum similarity with a maximum similarity threshold, and determining the correction coefficient according to the comparison result;
[0103] The maximum similarity is obtained through the following formula:
[0104]
[0105] where, M max represents the maximum similarity; X i represents the i-th real-time voice feature vector; Y j represents the j-th historical voice feature vector; m represents the number of real-time voice feature vectors; n represents the number of historical voice feature vectors.
[0106] It can be seen that the real-time voice feature vector refers to the feature vector extracted from the current voice data, while the historical voice feature vector refers to the feature vector collected and stored in the previous voice interaction tests. By calculating the maximum similarity between the real-time voice feature vector and the historical voice feature vector, the matching degree between the current voice data and the historical data can be evaluated.
[0107] Specifically, when comparing the maximum similarity with the maximum similarity threshold and determining the correction coefficient according to the comparison result, it includes:
[0108] Comparing the maximum similarity with a first maximum similarity threshold and a second maximum similarity threshold, and determining the correction coefficient according to the comparison result; where, the first maximum similarity threshold is less than the second maximum similarity threshold;
[0109] Setting a correction coefficient interval, the correction coefficient interval includes a first correction coefficient, a second correction coefficient, and a third correction coefficient;
[0110] When the maximum similarity is less than or equal to the first maximum similarity threshold, determine that the correction coefficient of the voice interaction parameter is the first correction coefficient, and use the product value of the first correction coefficient and the voice interaction parameter as the corrected voice interaction parameter;
[0111] When the maximum similarity is greater than the first maximum similarity threshold and less than or equal to the second maximum similarity threshold, determine that the correction coefficient of the voice interaction parameter is the second correction coefficient, and use the product value of the second correction coefficient and the voice interaction parameter as the corrected voice interaction parameter;
[0112] When the maximum similarity is greater than the second maximum similarity threshold, determine that the correction coefficient of the voice interaction parameter is the third correction coefficient, and use the product value of the third correction coefficient and the voice interaction parameter as the corrected voice interaction parameter.
[0113] It can be understood that by setting different similarity thresholds and correction coefficient intervals, the voice interaction parameters can be flexibly adjusted according to the matching degree between real-time voice data and historical data. When the similarity between real-time voice data and historical data is low, a larger correction coefficient is adopted to significantly adjust the voice interaction parameters, thereby improving the adaptability and accuracy of the test. On the contrary, if the similarity is high, it indicates that the current voice interaction parameters are already relatively appropriate. At this time, a smaller correction coefficient is used for fine-tuning to maintain the stability and continuity of the test. In this way, it can be ensured that a high-quality voice interaction experience can be provided in a variety of specific scenarios.
[0114] Specifically, when determining the interaction influence index between the user and the smart home to be tested according to the distance data, behavior data, and device status data, it includes:
[0115] The interaction influence index is obtained by the following formula:
[0116]
[0117] Where, I represents the interaction influence index; D represents the distance data; B represents the behavior data; E represents the device status data; α1 represents the distance weight coefficient; α2 represents the behavior weight coefficient; α3 represents the device status weight coefficient; β1 represents the first influence coefficient; β2 represents the second influence coefficient; β3 represents the third influence coefficient.
[0118] It is understandable that in this embodiment, α1, α2, and α3 respectively represent the weight coefficients of distance, behavior, and device status data. Their selection is based on in-depth research on the influence degrees of these factors during the voice interaction process. For example, the weight coefficient α1 of distance data is preferably 0.3 because the distance between the user and the device directly affects the intensity and clarity of the voice signal. The closer the distance, the easier it is to accurately capture the voice signal. The weight coefficient α2 of behavior data is preferably 0.4 because the user's behavior patterns, such as the speaking speed and volume, have a significant impact on the accuracy of voice recognition. The weight coefficient α3 of device status data is preferably 0.3 because the operating status of smart home devices, such as whether it is in the mute mode or whether music is being played, also affects the voice interaction. β1, β2, and β3 respectively represent the first, second, and third influence coefficients related to distance, behavior, and device status data. Their settings are to further refine the calculation of the interaction influence index. For example, β1 can be set to 0.5 to emphasize the importance of distance data in the interaction influence; β2 can be set to 0.3 to reflect the medium degree of influence of behavior data on the interaction; β3 can be set to 0.2 to indicate the relatively small degree of influence of device status data on the interaction. By setting the above weight coefficients and influence coefficients, it can be ensured that the interaction influence index can accurately reflect the interaction situation between the user and the smart home device, thereby providing a scientific basis for the adjustment of voice interaction parameters. This data-driven adjustment mechanism not only improves the adaptability and personalization level of voice interaction testing but also helps to better meet the specific needs of users in actual use.
[0119] Specifically, when determining whether to compensate the corrected voice interaction parameters according to the interaction influence index, it includes:
[0120] Subtract the interaction influence index from the interaction influence index threshold to obtain an index difference;
[0121] Compare the index difference with the index difference threshold, and determine whether to compensate the corrected voice interaction parameters according to the comparison result;
[0122] When the index difference is greater than or equal to the index difference threshold, it is determined to compensate the corrected voice interaction parameters;
[0123] When the index difference is less than the index difference threshold, it is determined not to compensate the corrected voice interaction parameters.
[0124] Specifically, when determining the compensation coefficient according to the interaction influence index and compensating the corrected voice interaction parameters according to the compensation coefficient, it includes:
[0125] Compare the exponential difference threshold with the first exponential difference threshold and the second exponential difference threshold, and determine the compensation coefficient according to the comparison result; wherein, the first exponential difference threshold is less than the second exponential difference threshold;
[0126] Set a compensation coefficient interval, which includes a first compensation coefficient, a second compensation coefficient, and a third compensation coefficient;
[0127] When the exponential difference is less than or equal to the first exponential difference threshold, determine that the compensation coefficient for the corrected voice interaction parameter is the first compensation coefficient, and use the product value of the first compensation coefficient and the corrected voice interaction parameter as the compensated voice interaction parameter;
[0128] When the exponential difference is greater than the first exponential difference threshold and less than or equal to the second exponential difference threshold, determine that the compensation coefficient for the corrected voice interaction parameter is the second compensation coefficient, and use the product value of the second compensation coefficient and the corrected voice interaction parameter as the compensated voice interaction parameter;
[0129] When the exponential difference is greater than the second exponential difference threshold, determine that the compensation coefficient for the corrected voice interaction parameter is the third compensation coefficient, and use the product value of the third compensation coefficient and the corrected voice interaction parameter as the compensated voice interaction parameter.
[0130] It can be understood that by setting different exponential difference thresholds and compensation coefficient intervals, the corrected voice interaction parameters can be flexibly compensated according to the difference between the interaction influence index and the threshold. When the difference between the interaction influence index and the threshold is large, a larger compensation coefficient is adopted to significantly adjust the corrected parameter, thereby improving the adaptability and personalization level of voice interaction. On the contrary, if the difference is small, it indicates that the current corrected parameter is already relatively appropriate. At this time, a smaller compensation coefficient is used for fine-tuning to maintain the stability and continuity of voice interaction. In this way, it can be ensured that a high-quality voice interaction experience can be provided in a variety of specialized scenarios.
[0131] Refer to Figure 2 As shown, in some embodiments of the present application, this embodiment provides a smart home voice interaction test device for specialized scenario recognition, including:
[0132] A determination module, configured to determine a to-be-tested specialized scenario, extract scenario feature data of the to-be-tested specialized scenario, and determine voice interaction parameters according to the scenario feature data; it is also configured to perform a voice interaction test on a to-be-tested smart home using real-time voice data in the to-be-tested specialized scenario, obtain voice interaction data, and obtain a test score according to the voice interaction data;
[0133] The first judgment module is configured to compare the test score with a test score threshold, and determine whether to correct the voice interaction parameter according to the comparison result;
[0134] The correction module is configured to, when it is determined to correct the voice interaction parameter, extract features from the real-time voice data to obtain real-time voice feature values; compare the real-time voice feature values with historical voice data, determine a correction coefficient according to the comparison result, and correct the voice interaction parameter according to the correction coefficient to obtain corrected voice interaction parameters;
[0135] The second judgment module is configured to, after the corrected voice interaction parameter is determined, collect distance data between the user and the smart home to be tested, behavior data of the user, and device status data of the smart home to be tested, and determine an interaction influence index between the user and the smart home to be tested according to the distance data, behavior data, and device status data; determine whether to compensate the corrected voice interaction parameter according to the interaction influence index;
[0136] The compensation module is configured to, when it is determined to compensate the corrected voice interaction parameter, determine a compensation coefficient according to the interaction influence index, and compensate the corrected voice interaction parameter according to the compensation coefficient to obtain compensated voice interaction parameters.
[0137] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0138] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0139] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0140] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A smart home voice interaction testing method, characterized in that: include: Determine a specialized scene to be tested, extract scene feature data of the specialized scene to be tested, and determine voice interaction parameters according to the scene feature data; In the specific scenario to be tested, using real-time voice data to perform a voice interaction test on the smart home to be tested, obtaining voice interaction data, and obtaining a test score according to the voice interaction data; Comparing the test score with the test score threshold, and determining whether to modify the voice interaction parameter according to the comparison result; When it is determined that the voice interaction parameter is to be modified, feature extraction is performed on the real-time voice data to obtain a real-time voice feature value; Comparing the real-time voice feature value with the historical voice data, determining a correction coefficient according to the comparison result, and correcting the voice interaction parameter according to the correction coefficient to obtain a corrected voice interaction parameter; After determining the modified voice interaction parameters, collecting distance data between the user and the smart home to be tested, the user's behavior data and the device status data of the smart home to be tested, and determining the interaction impact index between the user and the smart home to be tested according to the distance data, behavior data and device status data; judging whether to compensate for the modified voice interaction parameters according to the interaction impact index; When it is determined to compensate the modified voice interaction parameter, a compensation coefficient is determined according to the interaction influence index, and the modified voice interaction parameter is compensated according to the compensation coefficient to obtain a compensated voice interaction parameter.
2. The smart home voice interaction testing method according to claim 1, characterized in that: Determining the voice interaction parameters according to the scene feature data includes: The scene feature data includes noise feature data and spatial echo feature data; The speech interaction parameters include recognition threshold, noise suppression parameter and speech enhancement coefficient; The recognition threshold is obtained by the following formula: The noise suppression parameter is obtained by the following formula: The speech enhancement coefficient is obtained by the following formula: Among them, T represents the recognition threshold; T0 represents the basic recognition threshold; a1 represents the noise recognition weight coefficient; xn represents the noise characteristic value; bn represents the noise standard value; a2 represents the spatial recognition weight coefficient; xr represents the current spatial echo characteristic value; br represents the current spatial echo standard value; N represents the noise suppression parameter; N0 represents the basic noise suppression parameter; b1 represents the noise suppression weight coefficient; b2 represents the spatial suppression weight coefficient; C represents the speech enhancement coefficient; C0 represents the basic speech enhancement coefficient; k1 represents the speech enhancement weight coefficient; k2 represents the spatial enhancement weight coefficient.
3. The smart home voice interaction testing method according to claim 1, characterized in that: The test score is compared with the test score threshold, and judging whether to modify the voice interaction parameter according to the comparison result includes: The test score is obtained by the following formula: Among them, S represents the test score; c1 represents the weight coefficient of speech recognition accuracy; Rmax represents the maximum value of speech recognition accuracy; R1 represents the speech recognition accuracy; c2 represents the weight coefficient of response time; T2 represents the response time; Tmin represents the minimum value of response time; c3 represents the weight coefficient of response accuracy; Amax represents the maximum value of response accuracy; A3 represents response accuracy.
4. The smart home voice interaction testing method according to claim 3, characterized in that: When comparing the test score with the test score threshold and determining whether to modify the voice interaction parameter according to the comparison result, the method further includes: When the test score is greater than the test score threshold, determining not to modify the voice interaction parameter; When the test score is less than or equal to the test score threshold, it is determined to modify the voice interaction parameter.
5. The smart home voice interaction testing method according to claim 1, characterized in that: The real-time voice feature value is compared with the historical voice data, a correction coefficient is determined according to the comparison result, and the voice interaction parameter is corrected according to the correction coefficient to obtain the corrected voice interaction parameter, including: Calculating the maximum similarity between the real-time speech feature value and the historical speech data, comparing the maximum similarity with a maximum similarity threshold, and determining the correction coefficient according to the comparison result; The maximum similarity is obtained by the following formula: Among them, M max represents the maximum similarity; X i represents the i-th real-time speech feature vector; Y j represents the jth historical speech feature vector; m represents the number of real-time speech feature vectors; n represents the number of historical speech feature vectors.
6. The smart home voice interaction testing method according to claim 5, characterized in that: The maximum similarity is compared with the maximum similarity threshold, and the correction coefficient is determined according to the comparison result, including: Comparing the maximum similarity with a first maximum similarity threshold and a second maximum similarity threshold, and determining the correction coefficient according to the comparison result; wherein the first maximum similarity threshold is less than the second maximum similarity threshold; Setting a correction coefficient interval, wherein the correction coefficient interval includes a first correction coefficient, a second correction coefficient, and a third correction coefficient; When the maximum similarity is less than or equal to the first maximum similarity threshold, determining the correction coefficient of the voice interaction parameter as the first correction coefficient, and taking the product value of the first correction coefficient and the voice interaction parameter as the corrected voice interaction parameter; When the maximum similarity is greater than the first maximum similarity threshold and less than or equal to the second maximum similarity threshold, determining the correction coefficient of the voice interaction parameter as the second correction coefficient, and taking the product value of the second correction coefficient and the voice interaction parameter as the corrected voice interaction parameter; When the maximum similarity is greater than the second maximum similarity threshold, the correction coefficient of the voice interaction parameter is determined to be a third correction coefficient, and the product value of the third correction coefficient and the voice interaction parameter is used as the corrected voice interaction parameter.
7. The smart home voice interaction testing method according to claim 6, characterized in that: Determining the interaction impact index between the user and the smart home to be tested according to the distance data, the behavior data and the device status data includes: The interaction impact index is obtained by the following formula: Among them, I represents the interaction influence index; D represents distance data; B represents behavior data; E represents device status data; α1 represents the distance weight coefficient; α2 represents the behavior weight coefficient; α3 represents the device status weight coefficient; β1 represents the first influence coefficient; β2 represents the second influence coefficient; β3 represents the third influence coefficient.
8. The smart home voice interaction testing method according to claim 1, characterized in that: When judging whether to compensate the modified voice interaction parameter according to the interaction influence index, it includes: Subtracting the interaction impact index from the interaction impact index threshold to obtain an index difference; Comparing the index difference with the index difference threshold, and determining whether to compensate the modified voice interaction parameter according to the comparison result; When the index difference is greater than or equal to the index difference threshold, determining to compensate the modified voice interaction parameter; When the index difference is less than the index difference threshold, it is determined that the modified voice interaction parameter is not compensated.
9. The smart home voice interaction testing method according to claim 8, characterized in that: Determining a compensation coefficient according to the interaction influence index, and compensating the modified voice interaction parameter according to the compensation coefficient, includes: Comparing the index difference threshold with the first index difference threshold and the second index difference threshold, and determining the compensation coefficient according to the comparison result; wherein the first index difference threshold is smaller than the second index difference threshold; Setting a compensation coefficient interval, wherein the compensation coefficient interval includes a first compensation coefficient, a second compensation coefficient, and a third compensation coefficient; When the exponential difference is less than or equal to the first exponential difference threshold, determining the compensation coefficient of the modified voice interaction parameter as the first compensation coefficient, and taking the product value of the first compensation coefficient and the modified voice interaction parameter as the compensated voice interaction parameter; When the exponent difference is greater than the first exponent difference threshold and less than or equal to the second exponent difference threshold, determining the compensation coefficient of the modified voice interaction parameter as the second compensation coefficient, and taking the product value of the second compensation coefficient and the modified voice interaction parameter as the compensated voice interaction parameter; When the exponent difference is greater than the second exponent difference threshold, the compensation coefficient of the modified voice interaction parameter is determined to be a third compensation coefficient, and the product value of the third compensation coefficient and the modified voice interaction parameter is used as the compensated voice interaction parameter.
10. A smart home voice interaction test device, applied to the smart home voice interaction test method according to any one of claims 1 to 9, characterized in that: include: A determination module is configured to determine a specialized scenario to be tested, extract scenario feature data of the specialized scenario to be tested, and determine a voice interaction parameter according to the scenario feature data; and is also configured to perform a voice interaction test on the smart home to be tested using real-time voice data in the specialized scenario to be tested, obtain voice interaction data, and obtain a test score according to the voice interaction data; A first judgment module is configured to compare the test score with a test score threshold, and determine whether to modify the voice interaction parameter according to the comparison result; A correction module, configured to extract features from the real-time voice data to obtain a real-time voice feature value when it is determined that the voice interaction parameter is to be corrected; Comparing the real-time voice feature value with the historical voice data, determining a correction coefficient according to the comparison result, and correcting the voice interaction parameter according to the correction coefficient to obtain a corrected voice interaction parameter; The second judgment module is configured to collect the distance data between the user and the smart home to be tested, the behavior data of the user and the device status data of the smart home to be tested, and determine the interaction influence index between the user and the smart home to be tested according to the distance data, the behavior data and the device status data after determining the modified voice interaction parameter; and determine whether to compensate the modified voice interaction parameter according to the interaction influence index; The compensation module is configured to determine a compensation coefficient according to the interaction influence index when it is determined to compensate the modified voice interaction parameter, and compensate the modified voice interaction parameter according to the compensation coefficient to obtain a compensated voice interaction parameter.
Citation Information
Patent Citations
Threshold adaptive adjustment method for off line speech recognition
CN108550365A
Environment correction method and device based on voice air conditioner
CN109469969A
Method and system for improving speech recognition accuracy and medium
CN117219058A
Smart home voice interaction test method and device for special scene recognition
CN118658459A
Method, system and chip for adaptively adjusting noise reduction mode
CN118865936A