Tone configuration method, device and equipment for intelligent voice interaction and medium
By recognizing user information and adjusting interaction scenario information, a target timbre is generated, which solves the problem of low intelligence caused by single timbre configuration in existing technologies. This achieves flexible adaptation and high intelligence in intelligent voice interaction, thereby improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-03-27
AI Technical Summary
In existing intelligent voice interaction technologies, voice interaction usually relies on a single, fixed timbre, which cannot provide differentiated services to users, resulting in low intelligence and affecting user experience.
By acquiring voice commands from the user's terminal, recognizing user information, matching candidate timbres, and adjusting timbre parameters according to interaction scenario information, the target timbre is generated. This includes data of various information types such as user information, environmental information, voice terminal device information, and interaction time, enabling the timbre to flexibly adapt to different situations.
It enables different timbre variations for different users and different voice commands, improving the intelligence and adaptability of voice interaction and enhancing the user interaction experience.
Smart Images

Figure CN121747553A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of smart home, in particular to a timbre configuration method and device for smart voice interaction, equipment and medium. BACKGROUND
[0002] Smart voice interaction technology refers to the voice communication and information interaction between people and machines through artificial intelligence algorithms and natural language processing technology. In recent years, with the improvement of computer performance and the accumulation of big data, smart voice interaction technology has made remarkable development. At present, this technology has been widely used in smart speakers, smart homes, intelligent customer service and other fields.
[0003] However, there are still some problems in the actual application of voice interaction. Voice interaction usually only relies on fixed configuration of single timbre to interact with users. In complex actual scenarios, single configuration cannot provide differentiated services for users in the interaction process, which is low in intelligence and affects user experience. SUMMARY
[0004] In view of the above problems, the present application provides a timbre configuration method, device, equipment and medium for smart voice interaction.
[0005] In a first aspect, the present application provides a timbre configuration method for smart voice interaction, comprising:
[0006] Obtaining a voice instruction of a user terminal, and identifying user information for current voice interaction;
[0007] According to the candidate timbre matching set corresponding to the user information, determining a candidate timbre corresponding to the voice instruction, wherein different voice instructions correspond to different candidate timbers in the candidate timbre matching set;
[0008] Obtaining interaction scene information corresponding to the current voice interaction, and adjusting parameters of the candidate timbre according to the interaction scene information to obtain a target timbre, wherein the interaction scene information includes one or more information types of data such as user information, environment information, voice terminal device information and interaction time;
[0009] After obtaining the response content corresponding to the voice instruction, the user terminal uses the target timbre to broadcast the response content.
[0010] In a possible implementation manner, the adjusting of the parameters of the candidate timbre according to the interaction scene information comprises:
[0011] Obtaining a weight list corresponding to each information type in the interaction scene information, wherein the weight list includes secondary weights corresponding to a plurality of influence factors under each information type.
[0012] generate a timbre tuning matrix according to the weight list, wherein each row or column of the timbre tuning matrix represents a different information type, and a single value in a row or column represents an influence factor secondary weight of a corresponding information type;
[0013] adjust at least two parameters of the candidate timbre according to the timbre tuning matrix and the interaction scenario information, the parameters being pitch, volume, frequency response or tone quality.
[0014] In a possible implementation, the adjusting at least two parameters of the candidate timbre according to the timbre tuning matrix and the interaction scenario information comprises:
[0015] obtaining an actual value corresponding to each influence factor according to the interaction scenario information;
[0016] obtaining a suggested value of the at least two parameters corresponding to each influence factor according to the actual value of each influence factor and a preset mapping relationship, wherein the mapping relationship is used to indicate a corresponding relationship between different actual values of each influence factor and suggested values of the at least two parameters;
[0017] adjusting at least two parameters of the candidate timbre according to the suggested values of the at least two parameters and the timbre tuning matrix.
[0018] In a possible implementation, the adjusting at least two parameters of the candidate timbre according to the suggested values of the at least two parameters and the timbre tuning matrix comprises:
[0019] for each parameter, obtaining an intermediate suggested value corresponding to each information type in the interaction scenario information by weighted calculation according to the suggested value of the parameter and the timbre tuning matrix;
[0020] obtaining a primary weight corresponding to each information type, and obtaining an actual updated value of the parameter by weighted calculation according to the intermediate suggested value corresponding to each information type and the primary weight;
[0021] adjusting at least two parameters of the candidate timbre according to the actual updated values corresponding to the at least two parameters.
[0022] In a possible implementation, after the target timbre is used to play the response content to the user end, the method further comprises:
[0023] in the process of the current voice interaction, if a new voice instruction is sent by the user end, obtaining latest interaction scenario information of the current voice interaction;
[0024] confirm whether the latest interaction scene information is same as interaction scene information corresponding to a previous voice instruction;
[0025] If not, at least two parameters corresponding to a previously used target timbre are adjusted according to a timbre deployment matrix and interaction scene information, to obtain a latest target timbre, so that the user end uses the latest target timbre for voice interaction.
[0026] In a possible implementation, the method further includes:
[0027] After the current voice interaction ends, a corresponding conversation record is obtained;
[0028] According to the conversation record, the timbre deployment matrix is updated.
[0029] In a possible implementation, the updating of the timbre deployment matrix according to the conversation record includes:
[0030] According to the conversation record, it is confirmed through semantic analysis whether the user evaluates the voice timbre;
[0031] If yes, the evaluation content of the user on the voice timbre is obtained, and according to the evaluation content, an intention analysis is performed to obtain an intention analysis result, the intention analysis result is used to indicate an impact factor corresponding to the evaluation content and an evaluation attribute of the impact factor to which the evaluation content is directed, the attribute is positive or negative;
[0032] According to the intention analysis result, the secondary weight of the corresponding impact factor is adjusted to update the timbre deployment matrix, and the adjustment direction of the secondary weight of the impact factor is positively correlated with the corresponding evaluation attribute.
[0033] In a second aspect, the present application provides a timbre configuration device for intelligent voice interaction, the device includes:
[0034] An acquisition module is configured to acquire a voice instruction of a user end and identify user information for current voice interaction;
[0035] A selection module is configured to determine a candidate timbre corresponding to the voice instruction from a candidate timbre matching set corresponding to the user information, wherein different voice instructions correspond to different candidate timbres in the candidate timbre matching set;
[0036] A processing module is configured to acquire interaction scene information corresponding to current voice interaction, and adjust parameters of the candidate timbre according to the interaction scene information to obtain a target timbre, wherein the interaction scene information includes one or more information types of data such as user information, environment information, voice terminal device information and interaction time.
[0037] The playing module is configured to play the response content corresponding to the voice instruction using the target voice tone.
[0038] In a third aspect, a computer readable storage medium is provided, which includes a stored program. When the program is executed, the method in any one of the first aspect is performed.
[0039] In a fourth aspect, an electronic device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the method in any one of the first aspect by using the computer program.
[0040] The voice tone configuration method, device, and equipment for intelligent voice interaction provided in the present application can obtain a voice instruction of a user terminal, determine user information of a current voice interaction, determine a candidate voice tone corresponding to the voice instruction according to a candidate voice tone matching set corresponding to the user information, obtain interaction scene information corresponding to the current voice interaction, and adjust parameters of the candidate voice tone according to the interaction scene information to obtain a target voice tone. Different users can have corresponding candidate voice tone matching sets, and the corresponding voice tone can be different when the user issues voice instruction content. After the candidate voice tone is determined, the interaction scene information corresponding to the current voice interaction is obtained, and the parameters of the candidate voice tone are adjusted according to the interaction scene information to obtain the target voice tone. This method realizes the voice tone change corresponding to different users and different voice instructions, can flexibly adapt to the needs of users in different situations, and can adjust the voice tone parameters according to actual conditions to adapt to various interaction scenes. The voice interaction process is highly intelligent and has strong self-adaptation capability, and the user interaction experience is improved. BRIEF DESCRIPTION OF DRAWINGS
[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application together with the specification.
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without any creative effort.
[0043] Figure 1 is a flowchart of a voice tone configuration method for intelligent voice interaction according to an embodiment of the present application Figure 1 ;
[0044] Figure 2 is a flowchart of a voice tone configuration method for intelligent voice interaction according to an embodiment of the present application Figure 2 ;
[0045] Figure 3 is a flowchart of a timbre configuration method of intelligent voice interaction according to an embodiment of the application Figure 3
[0046] Figure 4 is a structural diagram of a timbre configuration device of intelligent voice interaction according to an embodiment of the application
[0047] Figure 5 is a hardware diagram of an electronic device according to an embodiment of the application. DETAILED DESCRIPTION
[0048] In order to make the personnel in the art better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.
[0049] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0050] To address the problems in existing technologies, this application provides a method for configuring the timbre of intelligent voice interaction. The method involves acquiring voice commands from the user's terminal to determine the user information for the current voice interaction. Based on the candidate timbre matching set corresponding to the user information, a candidate timbre corresponding to the voice command is determined. Specifically, different users can have corresponding candidate timbre matching sets, and the corresponding timbre may differ depending on the content of the user's voice command. After determining the candidate timbre, the interaction scenario information corresponding to the current voice interaction is acquired, and the parameters of the candidate timbre are adjusted according to the interaction scenario information to obtain the target timbre. The interaction scenario information may include one or more types of information data, such as user information, environmental information, voice terminal device information, and interaction time. After acquiring the response content corresponding to the voice command, the response content is played to the user's terminal using the target timbre. This method enables timbre changes for different users and different voice commands, flexibly adapting to the user's needs in different situations. Simultaneously, adjusting the timbre parameters according to the actual situation makes the voice output more adaptable, capable of adapting to various interaction scenarios to enhance communication effectiveness. The voice interaction process is highly intelligent and has strong adaptive capabilities, improving the user interaction experience.
[0051] The aforementioned method for configuring the tone of intelligent voice interaction can be widely applied to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and smart residential ecosystems. The user-end devices in this method are not limited to PCs, mobile phones, tablets, smart air conditioners, smart range hoods, smart refrigerators, smart ovens, smart stoves, smart washing machines, smart water heaters, smart washing equipment, smart dishwashers, smart projectors, smart TVs, smart clothes racks, smart curtains, smart audio-visual systems, smart sockets, smart speakers, smart speakers, smart ventilation systems, smart kitchen and bathroom equipment, smart bathroom fixtures, smart robot vacuums, smart window cleaning robots, smart mopping robots, smart air purifiers, smart steam ovens, smart microwave ovens, smart water heaters, smart air purifiers, smart water dispensers, and smart door locks.
[0052] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Where there is no conflict between the embodiments, the following embodiments and features can be combined with each other. Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0053] Figure 1 This is a flowchart illustrating the tone configuration method for intelligent voice interaction according to an embodiment of this application. Figure 1 .like Figure 1 As shown, the method includes:
[0054] S101, acquire a voice instruction of a user terminal, and identify user information for current voice interaction.
[0055] In this step, in the process of receiving the voice instruction from the user equipment, not only the voice content of the user is included, but also the identification information of the user identity. The user equipment can be a microphone or other input device to acquire the voice input of the user. The user information can be acquired in various ways, such as voice recognition (identifying the user identity through voiceprint), account login information, or device binding information, etc. The user identity is identified by analyzing the voiceprint characteristics (such as pitch, tone, speed, etc.) of the user. The voiceprint is a biometric identification method, which can be matched with the information in the user database; in some cases, the device is already bound to the user account, such as a mobile phone or a smart speaker bound to a specific user account, and the user identity can also be directly acquired through such binding information; in addition, the current user can also be inferred according to the previous interaction history (such as previous voice commands or use habits) of the user.
[0056] S102, determine a candidate voice color corresponding to the voice instruction from a candidate voice color matching set corresponding to the user information.
[0057] Among the candidate voice color matching set, different voice instructions correspond to different candidate voice colors.
[0058] In this step, based on the user information acquired in step S101, a candidate voice color matching set corresponding to the user is obtained. This list can be generated according to the user's personalized data, preferences and historical use records. For example, the user usually prefers a softer voice color for daily communication, and may prefer a more serious voice color when reminding or informing important information. At the same time, this list can also be some pre-set voice colors selected by the user for each interaction content, as well as the user's own voice color or other users' voice color. When the user wants to set his own voice color as a candidate voice color, through the voice color configuration process, the user's training sample is obtained, and the sample is input into the voice color generation model based on machine learning or deep learning, to obtain a candidate voice color identical to the user's voice for the user to choose.
[0059] For example, different voice instructions correspond to different candidate voice colors, which can also be analyzed by natural language processing (NLP) technology to identify the core content of the instruction. The type of voice instruction (such as command type, question type, entertainment type, etc.) determines the selection of the voice color. For example, if the user asks about the weather, a friendly and clear voice color may be selected; if the user issues an emergency instruction, a serious and efficient voice color may be selected. If the user specifies different voice colors for different types of instructions, the voice color used is determined according to the user's pre-selection.
[0060] S103, obtain the interaction scene information corresponding to the current voice interaction, and adjust the parameters of the candidate vocal color according to the interaction scene information to obtain a target vocal color.
[0061] The interaction scene information includes one or more information types of data such as user information, environment information, voice terminal device information, and interaction time.
[0062] In this step, the obtained interaction scene information is multi-dimensional, which can include user information: including personalized data of the user, such as age, gender, emotional state (which can be obtained through voice emotion analysis), etc. Environment information: such as background noise, light intensity, etc. For example, the user may enhance the volume and clarity of the vocal color in a noisy environment. Device information: the type of device for voice interaction is also an important dimension. Different devices (such as smart speakers, mobile phones, car systems, etc.) have different audio output capabilities, so the adaptive adjustment of the vocal color needs to consider the characteristics of the device. Interaction time: the interaction time period also affects the selection of vocal color, for example, a softer and quieter vocal color may be automatically selected at night, while a more lively vocal color may be selected during the day.
[0063] According to the obtained scene information, the specific parameters of the candidate vocal color can be dynamically adjusted. These parameters include but are not limited to pitch, volume, tone, speech rate, etc. When the interaction scene information contains the single-sided information exemplified above, the relevant adaptive adjustment can be made according to the environmental situation of the information feedback. For example, in a noisy environment, the volume may be automatically increased and the clarity of the tone may be improved; in a quiet environment, the volume may be lowered to make the vocal color softer. When the interaction scene information contains the multi-aspect information exemplified above, a comprehensive analysis can be made, and the multi-dimensional condition linkage adjustment can be made. For example, when the user interacts through a smart speaker device in a quiet environment, and it is currently nighttime, the vocal color can be comprehensively optimized in volume, clarity, frequency range, and emotion, so as to adapt to the current situational needs.
[0064] S104, after obtaining the response content corresponding to the voice instruction, the user terminal uses the target vocal color to broadcast the response content.
[0065] In this step, after generating the response content, the target vocal color determined in step S103 is applied to the response content. For example, the voice can be synthesized according to the parameter settings (such as pitch, speech rate, volume, etc.) of the target vocal color through voice synthesis technology.
[0066] For example, during the broadcasting process, the user's feedback can also be monitored in real time, such as whether the user interrupts or modifies the voice instruction. If it is detected that the user is not satisfied with the vocal color or the environment has changed, the vocal color parameters can be dynamically adjusted and the synthesized voice can be regenerated.
[0067] The intelligent voice interaction timbre configuration method provided by the embodiments of the present application comprises the following steps: obtaining a voice instruction of a user end, determining user information of a current voice interaction; determining a candidate timbre corresponding to the voice instruction from a candidate timbre matching set corresponding to the user information. Specifically, different users can have corresponding candidate timbre matching sets, and the corresponding timbre can be different when the user issues voice instructions with different contents. After the candidate timbre is determined, interaction scene information corresponding to the current voice interaction is obtained, and the parameters of the candidate timbre are adjusted according to the interaction scene information to obtain a target timbre. This method realizes the change of the timbre corresponding to different users and different voice instructions, and can flexibly adapt to the needs of users in different situations. At the same time, the timbre parameters are adjusted according to the actual situation, which can adapt to various interaction scenes. The voice interaction process is highly intelligent and has strong self-adaptation ability, and the user interaction experience is improved.
[0068] Figure 2 is a flowchart of an intelligent voice interaction timbre configuration method according to an embodiment of the present application Figure 2 . As shown in Figure 3 , the method comprises the following steps:
[0069] S201, obtaining a weight list corresponding to each information type in the interaction scene information.
[0070] The weight list comprises secondary weights corresponding to a plurality of influence factors under each information type.
[0071] S202, generating a timbre deployment matrix according to the weight list.
[0072] Each row or column of the timbre deployment matrix represents a different information type, and a single value in the row or column represents the secondary weight of an influence factor corresponding to the information type.
[0073] It should be noted that these weights are used to measure the influence degree of different information types on timbre deployment. The information type refers to user characteristics (such as age, gender), environmental conditions (such as noise, quietness), device information (such as device performance), time period, etc. The weight list contains primary weights (overall weights of different information types) and secondary weights (secondary weights corresponding to the subdivided influence factors under each information type). For example, user characteristics can be further subdivided into age, gender, etc., and environmental information can be subdivided into noise level, indoor and outdoor environment, etc. The role of each information type in timbre deployment can be determined through the weight list.
[0074] The rows or columns of the timbre tuning matrix represent different information types (such as user characteristics, environmental information, device information, time period, etc.). The values in the matrix represent the secondary weights of the subdivided influence factors under each information type. For example, a row represents user characteristics, and the columns are the corresponding secondary weights of the user's age, gender, etc. Each numerical value in the matrix represents the relative weight of the influence factor in the timbre tuning. Through the matrix, a structured way can be provided to integrate all information types and their influence factors into the timbre tuning process, so as to more accurately adjust the timbre parameters.
[0075] S203, obtaining an actual value corresponding to each influence factor according to the interaction scene information.
[0076] In this step, the actual value corresponding to each influence factor is obtained from the current interaction scene, which will affect the generation of the final timbre parameters. For example, the age information in the user characteristics can be obtained through the user's registration data, the noise level in the environmental information can be detected in real time through the microphone, the speaker performance in the device information can be read through the device configuration file, and the time information can come from the system time. These actual values reflect the specific conditions of the current scene and are an important basis for subsequent mapping of timbre parameters.
[0077] S204, obtaining a suggested value of the at least two parameters corresponding to each influence factor according to the actual value of each influence factor and a preset mapping relationship.
[0078] The mapping relationship is used to indicate the corresponding relationship between different actual values of each influence factor and the suggested values of the at least two parameters.
[0079] In the timbre tuning process, based on the actual data of the influence factors in different dimensions and the mapping relationship, a set of corresponding optimized timbre parameters can be provided for each influence factor in each dimension. That is, each influence factor in each dimension is regarded as an input variable, and each influence factor in each dimension corresponds to a set of specific timbre parameter suggestions. Among them, for each influence factor in each dimension, the suggested timbre parameters can be generated through a pre-trained model or rules.
[0080] S205, adjusting the at least two parameters of the candidate timbre according to the suggested values of the at least two parameters and the timbre tuning matrix.
[0081] The parameters are pitch, volume, frequency response, or tone quality.
[0082] In this step, the weights of each information type and its influencing factors are integrated through the tone tuning matrix to determine how to adjust the tone parameters according to the recommended values of multiple information sources. The weights in the matrix determine the priority of each factor. For example, the volume may be affected by user information, device information, and environmental information at the same time, and through the tone tuning matrix, these factors can be integrated to finally determine the specific adjustment value of the volume. Similarly, tone, frequency response, and sound quality can also be adjusted comprehensively through similar methods. According to the tone tuning matrix and the recommended value of each influencing factor, the final parameters of the candidate tone are calculated comprehensively to generate a suitable tone for announcement.
[0083] For example, adjusting the at least two parameters of the candidate tone according to the recommended values of the at least two parameters and the tone tuning matrix comprises:
[0084] For each parameter, according to the recommended value of the parameter and the tone tuning matrix, the intermediate recommended value of the parameter corresponding to each information type in the interaction scene information is calculated by weighting;
[0085] Obtain the first-level weight corresponding to each information type, and calculate the actual update value of the parameter by weighting according to the intermediate recommended value and the first-level weight corresponding to each information type;
[0086] Adjust the at least two parameters of the candidate tone according to the actual update values corresponding to the at least two parameters.
[0087] In order to better explain this embodiment, for example, it is assumed that when analyzing the tone, four dimensions, i.e. information types (user characteristics, environment, device, and time) are analyzed at the same time.
[0088] The first-level weights of the four dimensions are: user characteristic weight (W_user): 0.4; environmental weight (W_environment): 0.2; device weight (W_device): 0.3; and time weight (W_time): 0.1. It is assumed that the frequency response is mainly processed, and the specific frequency band distribution is: high frequency, medium frequency, and low frequency.
[0089] In order to simplify the example, the process of calculating and obtaining the intermediate recommended value is omitted. At this time, each dimension will have its corresponding intermediate recommended value. For example:
[0090] 1. User characteristic recommended frequency response:
[0091] High frequency: 0.3
[0092] Medium frequency: 0.4
[0093] Low frequency: 0.7
[0094] 2. Environmental dimension recommended frequency response:
[0095] High frequency: 0.5
[0096] Mid frequency: 0.3
[0097] Low frequency: 0.5
[0098] 3. Frequency response of the equipment dimension suggestion:
[0099] High frequency: 0.4
[0100] Mid frequency: 0.5
[0101] Low frequency: 0.6
[0102] 4. Frequency response of the time dimension suggestion:
[0103] High frequency: 0.2
[0104] Mid frequency: 0.3
[0105] Low frequency: 0.6
[0106] For each timbre parameter, the final timbre parameter is determined by multiplying the weight of each dimension with its corresponding timbre parameter value and then summing the results. The calculation formula is:
[0107] Timbre parameter value = W_user x P_user + W_environment x P_environment + W_equipment x P_equipment + W_time x P_time, wherein P represents the specific timbre parameter value under each dimension.
[0108] High frequency parameter value = (0.4 x 0.3) + (0.2 x 0.5) + (0.3 x 0.4) + (0.1 x 0.2) = 0.36
[0109] Mid frequency parameter value = (0.4 x 0.4) + (0.2 x 0.3) + (0.3 x 0.5) + (0.1 x 0.3) = 0.40
[0110] Low frequency parameter value = (0.4 x 0.7) + (0.2 x 0.5) + (0.3 x 0.6) + (0.1 x 0.6) = 0.62
[0111] Through the above calculation, the timbre parameters of the frequency response obtained are: high frequency: 0.36; mid frequency: 0.40; and low frequency: 0.62.
[0112] When multiple timbre parameters are analyzed and confirmed, each timbre parameter is obtained as described above.
[0113] For example, after the target timbre plays the response content to the user end, it further includes:
[0114] In the current voice interaction process, if a user terminal sends a new voice instruction, the latest interaction scene information of the current voice interaction is acquired;
[0115] It is determined whether the latest interaction scene information is same as the interaction scene information corresponding to the previous voice instruction;
[0116] If not, at least two parameters corresponding to the target timbre used last time are adjusted according to the timbre adjustment matrix and the interaction scene information, to obtain the latest target timbre, so that the user terminal uses the latest target timbre for voice interaction.
[0117] It should be noted that in continuous voice interaction, if a user issues a new voice instruction (such as a new request or command), the latest interaction scene information at this time can be collected in real time. That is, the scene information related to the current environment, device, time, etc. is reacquired, to ensure that the timbre can adapt to the current interaction demand. By comparing the latest acquired interaction scene information and the interaction scene information recorded at the time of the previous voice instruction, it is checked whether a change has occurred. If the interaction scene information is same, it indicates that the current environment and condition have not changed, and the timbre parameter does not need to be adjusted; if not, it indicates that the scene has changed, and the timbre may need to be adjusted accordingly. For example, if it is detected that the interaction scene has changed (for example, the noise increases, or the user changes a device), the timbre parameter is updated according to the new scene information and the timbre adjustment matrix.
[0118] The timbre configuration method for intelligent voice interaction provided in the embodiments of the present application dynamically adjusts the timbre parameter by using the weight list and the timbre adjustment matrix, in combination with the actual value of each dimension influence factor and the mapping relationship, to ensure that the timbre can adapt to the complex and changeable interaction scene in the timbre interaction process, and to provide the most suitable timbre experience for the user.
[0119] Figure 3 is a flowchart of the timbre configuration method for intelligent voice interaction according to the embodiments of the present application Figure 4 . As shown in Figure 4 , the method comprises:
[0120] S301, after the current voice interaction ends, corresponding conversation records are acquired.
[0121] S302, according to the conversation records, it is determined by semantic analysis whether the user evaluates the voice timbre.
[0122] It should be noted that the natural language processing (NLP) technology can be used to perform semantic analysis on the conversation records, to determine whether the user has evaluated the timbre. For example, the user can say sentences such as “this sound is too sharp” or “the volume is a little small”, and the semantic analysis can determine that these sentences are related to the timbre.
[0123] In order to enable the user to provide more evaluation information, a related inquiry on timbre evaluation can also be inserted in the voice interaction process to obtain the user's related answers, so that more specific and clear evaluation data can be obtained.
[0124] S303, if not, no operation is performed.
[0125] S304, if yes, the evaluation content of the user on the voice timbre is obtained, and the intention analysis result is obtained according to the evaluation content.
[0126] The intention analysis result is used to indicate the influence factor corresponding to the evaluation content and the evaluation attribute of the influence factor for the evaluation content, and the attribute is positive or negative.
[0127] It should be noted that if the semantic analysis determines that the user has indeed evaluated the timbre, the system then further analyzes the user's evaluation content through intention analysis. Intention analysis is to identify the actual intention or expectation contained in the user's evaluation, and to clarify the user's specific demand for timbre adjustment. For example, the intention analysis can be determined by extracting the user's specific evaluation content, such as "the sound is too harsh" indicating that the user has a negative evaluation on the pitch, and "the volume is not loud enough" indicating a negative evaluation on the volume. The result of the intention analysis can include the influence factor: for example, the pitch, volume, and sound quality are the influence factors that the user is concerned about; and the evaluation attribute: to determine whether the user's evaluation is positive (such as satisfaction, praise) or negative (such as dissatisfaction, criticism), so as to determine the subsequent adjustment direction.
[0128] S305, according to the intention analysis result, the secondary weight of the corresponding influence factor is adjusted to update the timbre deployment matrix.
[0129] The adjustment direction of the secondary weight of the influence factor is positively correlated with the corresponding evaluation attribute.
[0130] The adjustment direction of the secondary weight is based on the evaluation attribute of the user. For example, a positive evaluation may indicate that no significant adjustment is needed, while a negative evaluation requires increasing the adjustment weight of the parameter. For example, if the user's evaluation of a certain parameter is negative, the secondary weight will be increased, so that the parameter will have a greater proportion in future timbre deployment, prompting the system to pay more attention to the adjustment of the parameter. If the user's evaluation is positive, the secondary weight will be reduced or remain unchanged, indicating that no much adjustment is needed.
[0131] The voice color configuration method for intelligent voice interaction provided in the embodiments of the present application can understand user feedback on voice color through semantic analysis and intent analysis, and convert the feedback into adjustment of the weight of an impact factor in a voice color deployment matrix. By continuously adjusting and optimizing the weight, a voice color setting that is more in line with user expectations is generated, so that voice interaction is more personalized and accurate.
[0132] Corresponding to the above method, the embodiments of the present application also provide a voice color configuration device for intelligent voice interaction, Figure 5 is a structural schematic diagram of a voice color configuration device for intelligent voice interaction according to an embodiment of the present application. As shown in Figure 5 , the device comprises:
[0133] The acquisition module 401 is configured to acquire a voice instruction of a user terminal and identify user information for current voice interaction.
[0134] The selection module 402 is configured to determine a candidate voice color corresponding to the voice instruction from a candidate voice color matching set corresponding to the user information, wherein different voice instructions correspond to different candidate voice colors in the candidate voice color matching set.
[0135] The processing module 403 is configured to acquire interaction scene information corresponding to the current voice interaction, and adjust parameters of the candidate voice color according to the interaction scene information to obtain a target voice color, wherein the interaction scene information comprises one or more information types of data such as user information, environment information, voice terminal device information, and interaction time.
[0136] The playing module 404 is configured to play back the response content corresponding to the voice instruction using the target voice color by the user terminal after the response content is acquired.
[0137] Figure 5 is a hardware schematic diagram of an electronic device according to an embodiment of the present application. As shown in , the electronic device 50 provided in the embodiment comprises at least one processor 501 and a memory 502. The device 50 further comprises a communication component 503. The processor 501, the memory 502, and the communication component 503 are connected through a bus 504.
[0138] In the specific implementation process, the at least one processor 501 executes computer execution instructions stored in the memory 502, so that the at least one processor 501 executes the above method.
[0139] The specific implementation process of the processor 501 can refer to the above method embodiments, which have similar implementation principles and technical effects, and will not be described here in detail.
[0140] In the above In the illustrated embodiment, it is to be understood that the processor can be a central processing unit (CPU), but can also be other general purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), etc. The general purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the disclosed method can be directly embodied as hardware processor execution or a combination of hardware and software modules in the processor.
[0141] The memory can include a random access memory (RAM) and can also include a non-volatile memory (NVM), such as at least one disk memory.
[0142] The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0143] The present application also provides a computer readable storage medium, the computer readable storage medium stores computer execution instructions / computer programs, when the processor executes the computer execution instructions / computer programs, the method as described above is realized.
[0144] The above readable storage medium can be realized by any type of volatile or non-volatile storage device or their combination, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0145] An example readable storage medium is coupled to the processor such that the processor can read information from the readable storage medium and can write information to the readable storage medium. Of course, the readable storage medium can also be a part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.
[0146] The division of the units is only a logical function division, and in actual implementation, another division manner can be used, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0147] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, can be located in one place or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0148] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0149] If the functions are realized in the form of software function units and sold or used as independent products, the functions can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0150] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware. The foregoing program can be stored in a computer readable storage medium. The program executes to perform the steps of the above-mentioned method embodiments; and the foregoing storage medium includes various storage media that can store program codes, such as ROM, RAM, magnetic disk or optical disk.
[0151] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments, and the above embodiments are only used to illustrate the technical solutions of the present application, not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for configuring the timbre of intelligent voice interaction, characterized in that, include: Obtain voice commands from the user and identify the user information for the current voice interaction; Based on the candidate timbre matching set corresponding to the user information, a candidate timbre corresponding to the voice command is determined, wherein different voice commands correspond to different candidate timbres in the candidate timbre matching set; The interaction scenario information corresponding to the current voice interaction is obtained, and the parameters of the candidate timbre are adjusted according to the interaction scenario information to obtain the target timbre. The interaction scenario information includes one or more information types of data, such as user information, environment information, voice terminal device information, and interaction time. After obtaining the response content corresponding to the voice command, the user terminal uses the target timbre to broadcast the response content.
2. The method according to claim 1, characterized in that, The step of adjusting the parameters of the candidate timbre according to the interaction scenario information includes: Obtain a weight list corresponding to each information type in the interactive scenario information, wherein the weight list includes secondary weights corresponding to multiple influencing factors under each information type; Based on the weight list, a tone matching matrix is generated, wherein each row or column of the tone matching matrix represents a different information type, and a single value in a row or column represents the secondary weight of the influence factor for the corresponding information type. Based on the tone matching matrix and the interaction scene information, at least two parameters of the candidate tone are adjusted, wherein the parameters are pitch, volume, frequency response or tone quality.
3. The method according to claim 2, characterized in that, The step of adjusting at least two parameters of the candidate timbre based on the timbre matching matrix and the interaction scene information includes: Based on the interaction scenario information, obtain the actual value corresponding to each influencing factor; Based on the actual value of each impact factor and the preset mapping relationship, the suggested values of the at least two parameters corresponding to each impact factor are obtained, wherein the mapping relationship is used to indicate the correspondence between different actual values of each impact factor and the suggested values of the at least two parameters; Based on the suggested values of the at least two parameters and the tone matching matrix, at least two parameters of the candidate tone are adjusted.
4. The method according to claim 3, characterized in that, The step of adjusting at least two parameters of the candidate timbre based on the suggested values of the at least two parameters and the timbre matching matrix includes: For each parameter, based on the suggested value of the parameter and the tone matching matrix, the intermediate suggested value of the parameter for each information type in the interactive scene information is obtained by weighted calculation; Obtain the primary weight corresponding to each information type, and calculate the actual updated value of the parameter by weighting the intermediate suggestion value and the primary weight corresponding to each information type. Based on the actual updated values corresponding to the at least two parameters, adjust at least two parameters of the candidate timbre.
5. The method according to claim 4, characterized in that, After the user terminal uses the target timbre to broadcast the response content, it also includes: During the current voice interaction, if the user sends a new voice command, the latest interaction scenario information of the current voice interaction will be obtained. Confirm whether the latest interaction scenario information is the same as the interaction scenario information corresponding to the previous voice command; If they are different, then based on the tone matching matrix and the interaction scenario information, at least two parameters corresponding to the previously used target tone are adjusted to obtain the latest target tone, so that the user terminal can use the latest target tone for voice interaction.
6. The method according to claim 5, characterized in that, The method further includes: After the current voice interaction ends, retrieve the corresponding conversation record; The tone matching matrix is updated based on the session records.
7. The method according to claim 6, characterized in that, The step of updating the tone matching matrix based on the session records includes: Based on the conversation records, semantic analysis is used to determine whether the user has evaluated the voice tone. If so, the user's evaluation of the voice timbre is obtained, and intent analysis is performed based on the evaluation content to obtain intent analysis results. The intent analysis results are used to indicate the influence factor corresponding to the evaluation content and the evaluation attribute of the evaluation content for the influence factor. The attribute is positive or negative. The secondary weights of the corresponding influencing factors are adjusted based on the intent analysis results to update the tone matching matrix. The adjustment direction of the secondary weights of the influencing factors is positively correlated with the corresponding evaluation attributes.
8. A voice interaction timbre configuration device, characterized in that, The device includes: The acquisition module is used to acquire voice commands from the user's end and identify user information for the current voice interaction. The selection module is used to determine the candidate timbre corresponding to the voice command based on the candidate timbre matching set corresponding to the user information, wherein different voice commands correspond to different candidate timbres in the candidate timbre matching set; The processing module is used to acquire the interaction scenario information corresponding to the current voice interaction, and adjust the parameters of the candidate timbre according to the interaction scenario information to obtain the target timbre. The interaction scenario information includes one or more information types of data, such as user information, environment information, voice terminal device information, and interaction time. The playback module is used to enable the user terminal to play the response content using the target timbre after obtaining the response content corresponding to the voice command.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 7.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 7 through the computer program.