A voice response method
By analyzing the environment and extracting features of user voice information, the intelligent outbound call robot can recognize valid voices and provide voice responses based on the response status, solving the shortcomings of existing technologies in identifying conversation environments and user status, and improving response accuracy and user experience.
Patent Information
- Application Number
- CN202210251736.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-03-15
AI Technical Summary
Existing intelligent outbound call robots are unable to effectively identify the conversation environment and user status during the voice response process, resulting in problems such as talking to themselves and interrupting users before they have finished speaking, affecting the accuracy of responses and user experience.
By obtaining the user's voice information, performing voice environment analysis, determining the effective voice and extracting voiceprint features, combining semantic and compensation features to determine the user's response status, and making voice responses based on the status.
It improves the answering accuracy and user experience of the intelligent outbound call robot, avoids noise interference in the conversation environment, and achieves more accurate voice response.
Smart Images

Figure CN114724587B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of speech recognition technology, and in particular to a method for speech response. Background Art
[0002] At present, with the rapid development of the market, intelligent outbound call robots are widely used in the field of voice response.
[0003] However, in actual application, the intelligent outbound call robots in the existing technology often cannot effectively identify the dialogue environment and the user's status, resulting in problems such as the intelligent outbound call robots talking to themselves and interrupting the user before they finish speaking, which seriously affects the intelligent outbound call robot's answering accuracy and the user experience.
[0004] Therefore, how to improve the intelligence of intelligent outbound call robots so that they can make special responses based on the user's conversation status is an urgent problem to be solved. Summary of the Invention
[0005] This specification provides a method and device for voice response to partially solve the above-mentioned problems existing in the prior art.
[0006] This manual adopts the following technical solutions:
[0007] This specification provides a voice response method, including:
[0008] Obtaining voice information collected during the user's voice response process;
[0009] Performing voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process;
[0010] If the voice information is determined to be valid, determining the user's response status during the voice response process based on the voice information;
[0011] According to the response status, a voice response is given to the user by a response robot during the voice response process.
[0012] Optionally, performing voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process specifically includes:
[0013] The voice information collected for the first time during the voice response process is used as the target voice;
[0014] Extracting voiceprint features from the target speech as the voiceprint features of the user;
[0015] For each voice message collected during the voice response process other than the target voice, determining whether the voiceprint feature of the voice message matches the voiceprint feature of the user;
[0016] If so, it is determined that the voice information is a valid voice uttered by the user during the voice response process.
[0017] Optionally, performing voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process specifically includes:
[0018] Inputting the voice information into a preset validity recognition model to obtain a validity recognition result of the validity recognition model for the voice information;
[0019] According to the validity recognition result, it is determined whether the voice information is a valid voice uttered by the user during the voice response process.
[0020] Optionally, determining, according to the voice information, a response status of the user in the voice response process specifically includes:
[0021] Extracting semantic features from the voice message, and extracting compensation features from the voice response process, the compensation features including at least one of: the number of words contained in the voice message, the duration of the response voice played by the response robot when the user sends the voice message, the completion ratio of the response voice played by the response robot when the user sends the voice message, and the playback status of the response robot when the user sends the voice message, the playback status being used to indicate whether the response robot is playing the response voice when the user sends the voice message;
[0022] The semantic features and the compensation features are input into a preset cooperation prediction model to determine the cooperation of the user when sending the voice information, and based on the cooperation, the response status of the user in the voice response process is determined.
[0023] Optionally, according to the response state, providing a voice response to the user by a response robot during the voice response process specifically includes:
[0024] If it is determined based on the response status that the user had abnormal emotions when sending the voice message, determining emotional information used to represent that the user had abnormal emotions when sending the voice message;
[0025] The emotional information and the voice information are input into a preset response model to determine a response voice that maintains the voice response process, and the response voice is played by the response robot.
[0026] Optionally, according to the response state, providing a voice response to the user by a response robot during the voice response process specifically includes:
[0027] If it is determined according to the response status that the user was in a normal mood when sending the voice message, judging according to the voice message whether the voice message is a complete voice message that the user needs to send;
[0028] If the voice information is not the complete voice that the user needs to utter, stop generating the voice response, and after monitoring the user to utter the complete voice, give the user a voice response through the answering robot based on the complete voice.
[0029] Optionally, according to the response state, providing a voice response to the user by a response robot during the voice response process specifically includes:
[0030] If it is determined according to the response status that the user was in a normal mood when sending the voice message, determining whether the actual intention of the user can be identified through the voice message;
[0031] If it is determined that the actual intention of the user is not recognized through the voice information, the answering robot plays a voice response to clarify the actual intention of the user to the user.
[0032] Optionally, the method further includes:
[0033] If it is determined that the voice information is not a valid voice, no response is given to the voice information, and after the valid voice uttered by the user during the voice response process is re-identified, a voice response is given to the user through the response robot.
[0034] This specification provides a voice response device, including:
[0035] An acquisition module is used to acquire the voice information collected during the voice response process of the user;
[0036] An analysis module, configured to perform voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process;
[0037] a determination module for determining, when determining that the voice information is valid voice, a response status of the user in the voice response process based on the voice information;
[0038] The answering module is used to give a voice answer to the user through an answering robot during the voice answering process according to the answering status.
[0039] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned voice response method is implemented.
[0040] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned voice response method when executing the program.
[0041] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0042] The voice response method provided in this specification first obtains the voice information collected during the voice response process for the user, and performs a voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process. If the voice information is determined to be a valid voice, the user's response status during the voice response process is determined based on the voice information, and then, based on the response status, a voice response is given to the user through the response robot during the voice response process.
[0043] It can be seen from the above method that by performing voice environment analysis on the voice information collected during the voice response process for the user, effective voice for voice response can be extracted from the voice response process, thereby avoiding the influence of the ambient sound of the conversation environment. Furthermore, based on the effective voice, the user's response status is determined, and then based on the user's response status, the voice response to the user by the answering robot during the voice response process is determined, thereby improving the response accuracy of the intelligent outbound call robot and the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0045] Figure 1 A flowchart of a voice response method provided in this specification;
[0046] Figure 2 This is a schematic diagram of a user's abnormal emotional response provided in this manual;
[0047] Figure 3 A schematic diagram of a voice response device provided in this manual;
[0048] Figure 4 This manual provides a corresponding Figure 1 Schematic diagram of the electronic device. DETAILED DESCRIPTION
[0049] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0050] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0051] Figure 1 This is a flow chart of a voice response method provided in this specification, which includes the following steps:
[0052] S101: Acquire voice information collected during the voice response process of the user.
[0053] In this specification, a voice response robot equipped with a voice response method can analyze the user's voice environment and response status during a voice response with a user, and provide a voice response based on the user's response status. Previously, a voice response robot equipped with a voice response method needed to first collect the user's voice information during the voice response process, and then analyze the user's voice environment and response status based on the user's voice information.
[0054] The voice response process may refer to the process in which the user communicates with the answering robot when performing services based on the answering robot. For example, while providing services to users, the telecom operator platform needs to continuously collect user feedback and optimize its own services based on user feedback to make the services more in line with the actual needs of users, thereby improving the user experience. The way to collect user feedback can be for the answering robot to play a voice requesting the user to evaluate or make suggestions on the service, and collect the voice information sent by the user containing the user's opinion, and obtain the user's opinion by analyzing the collected voice information sent by the user. This process of the call can be regarded as a voice response process.
[0055] In this specification, the execution entity used to implement the voice response method can refer to a server equipped with an answering robot, or it can refer to a designated device such as a desktop computer, laptop computer, etc. For the sake of convenience of description, the voice response method provided in this specification is explained below using the server as the execution entity as an example.
[0056] S102: Perform voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process.
[0057] During the process of voice response with the user, considering that there may be noise in the environment where the user is during the voice response, in order to avoid the impact of such noise on the voice response, the server can perform a voice environment analysis on the voice information collected during the voice response with the user to determine whether the current voice information is valid voice.
[0058] Among them, the noise that may exist in the environment where the user is during the voice answering process may include: environmental noise, signal current sound generated by the call, speech echo, chatting of people around, and voice played by the intelligent voice assistant when the person answering the call is smart.
[0059] In actual applications, when a user answers a call, he or she often first utters voice commands such as "hello" or "goodbye" before making subsequent calls. In this process, the first voice command uttered by the user can, to a certain extent, represent the user's voiceprint characteristics.
[0060] Therefore, the server can use the voice information collected for the first time during the voice response process with the user as the target voice, extract the voiceprint features in the target voice as the user's voiceprint features, and then determine whether the voiceprint features of each voice information collected during the user's voice response process except the target voice match the user's voiceprint features. If so, the voice information is determined to be a valid voice uttered by the user during the voice response process. Otherwise, it is determined that the current voice information is not a valid voice.
[0061] Of course, the server can also send the voice information collected during the voice response process with the user to a preset validity recognition model, and use the validity recognition model to determine whether the collected voice information is valid voice. The validity recognition model can adopt the bidirectional encoder representation from transformers (BERT) model, twin network model, etc.
[0062] In practical applications, the validity recognition model needs to be trained in advance before it can be deployed in the server. When training the above-mentioned validity recognition model in the server, the voice information collected from the user during the voice response process needs to be input into the validity recognition model to obtain the probability value of the voice output by the validity recognition model as a valid voice. Then, the validity recognition model is trained with the optimization goal of minimizing the deviation between the probability value of the voice output by the validity recognition model as a valid voice and the actual result of the sample.
[0063] S103: If the voice information is determined to be valid voice, determine the response status of the user in the voice response process according to the voice information.
[0064] After determining that the voice information collected during the voice response process for the user is valid voice, the server can extract semantic features from the collected voice information, and then determine the user's cooperation level when sending the voice information based on the extracted semantic features, and determine the user's response status based on the cooperation level.
[0065] In actual applications, considering that in the process of voice response, in addition to the semantic features in the voice information that can reflect the user's response status, information such as the time when the user sends the voice information and the status of the response robot playing the response voice when the user sends the voice information can also reflect the user's response status to a certain extent. Therefore, in order to further improve the accuracy of the user's response status determined by the server, the server can also extract compensation features from the voice response process, and then determine the user's cooperation degree when sending the voice information based on the extracted semantic features and compensation features, and determine the user's response status based on the cooperation degree.
[0066] Among them, the compensation features include: the number of words contained in the voice message, the duration of the response voice that the answering robot has played when the user sends the voice message, the playback completion ratio of the response voice that the answering robot has played when the user sends the voice message, and at least one of the playback status of the answering robot when the user sends the voice message. The playback status is used to indicate whether the answering robot is playing the response voice when the user sends the voice message.
[0067] In the above content, the duration of the response voice that the response robot has played when the user sends a voice message can refer to: the total duration of all the response voices that the response robot has played from the start time of the entire voice response process to the start time of the voice message currently sent by the user, or it can refer to the duration of the part of the voice response that the response robot is currently playing when the user currently sends a voice message.
[0068] The completion ratio of the response voice played by the answering robot when the user sends a voice message may refer to the ratio of the response voice played by the answering robot without being interrupted by the user during the entire response process to the total response voice played by the answering robot during the entire response process.
[0069] Among them, the number of words contained in the voice message can reflect the user's cooperation with the voice response to a certain extent. For example: after the answering robot plays the voice content "Do you think there is any room for improvement in our service?", the number of words contained in the voice message sent by the user is relatively large. To a certain extent, it can be reflected that the user is more interested in the entire voice response process, and the higher the user's cooperation.
[0070] In this specification, the duration of the voice response played by the robot when the user sends a voice message can be positively correlated with the aforementioned level of cooperation. For example, the longer the total duration of the voice response played by the robot from the start time of the entire voice response process to the start time of the user's current voice message, the higher the user's level of cooperation with the entire voice response. For another example, if the duration of the portion of the voice response currently played by the robot when the user sends a voice message is shorter, it can indicate a lower level of user interest and a lower level of user cooperation.
[0071] The above-mentioned playback completion ratio may be positively correlated with the degree of cooperation. That is, if the playback completion ratio of the response voice played by the response robot is low when the user sends a voice message, it may reflect that the user has abnormal emotions towards the entire voice response, and it can be determined that the user's cooperation is low.
[0072] In this specification, the server can extract semantic features from the collected voice information and compensatory features from the voice response process and send them to a preset cooperation prediction model, and determine the user's cooperation when sending the voice information through the preset cooperation prediction model, wherein the preset cooperation prediction model can adopt a back propagation (BP) neural network model, etc.
[0073] Among them, when the above-mentioned cooperation prediction model is trained in the server, it is necessary to input the collected voice information of the user during the voice response process into the cooperation prediction model to obtain the user's cooperation output by the cooperation prediction model, and then train the cooperation prediction model with the optimization goal of minimizing the deviation between the user's cooperation output by the cooperation prediction model and the actual cooperation of the sample.
[0074] Of course, the server can also determine the user's level of cooperation based on pre-set rules (for example, whether the voice information collected during the user's voice response contains keywords related to abnormal emotions) and the user's compensation characteristics. Abnormal emotions refer to, for example, the user's negative emotions and willingness to complain. For example, if a user utters a voice message with the text "I'm so annoyed!" during a voice response, the server can extract the word "annoyed" from the voice message and determine that the user has abnormal emotions. Furthermore, the server can determine the user's level of cooperation based on the keywords related to abnormal emotions extracted from the voice message.
[0075] S104: According to the response status, during the voice response process, a response robot performs a voice response to the user.
[0076] After determining the user's response status during the voice response process, the server can determine whether the user has abnormal emotions based on the user's response status, such as Figure 2 shown.
[0077] Figure 2 This is a schematic diagram of a user's abnormal emotional response provided in this manual;
[0078] Combine Figure 2 If the server determines that the user's cooperation level in the response state when sending the voice message is lower than the preset threshold, it is determined that the user has abnormal emotions when sending the voice message. If the server determines that the user's cooperation level in the response state when sending the voice message is higher than the preset threshold, it is determined that the user has normal emotions when sending the voice message.
[0079] Furthermore, if the server determines, based on the response status, that the user was in an abnormal emotion when sending the voice message, the server may also determine emotional information that characterizes the user's abnormal emotion when sending the voice message, and input the emotional information and voice information into a preset response model to determine a response voice that maintains the voice response process, and then play the response voice through a response robot installed in the server. For example, if a user sends a voice message with the text content "How annoying!" during the voice response process, and the server determines that the user's corresponding cooperation level when sending the voice message is below a preset threshold, it determines that the user was in an abnormal emotion, and inputs the emotional information and voice information into a preset response model. The response model then generates a voice response, for example, explaining the purpose of the call to the user to gain their understanding, thereby maintaining the voice response process with the user. It is worth noting that the response voice generated in the response model can be a voice segment, which is played by the response robot, or it can be a text message corresponding to the response voice, which is generated and played by the response robot based on the text message corresponding to the response voice.
[0080] If the server determines, based on the response status, that the user was in a normal mood when sending the voice message, it can determine, based on the semantic features extracted from the voice message and the contextual features extracted by the server based on the user's historical voice messages, whether the voice message is the complete voice that the user needs to send. A complete voice means that the content of the voice message currently expressed by the user is complete. For example, "I want to handle business A" is the complete voice that the user actually wants to express. If the voice message sent by the user is: I want to handle it, the server determines that the voice message does not contain complete content, that is, the voice message sent by the user is not the complete voice that the user needs to express.
[0081] If the voice message is not the complete voice message that the user intended, the server stops generating the voice response. After detecting that the user has uttered the complete voice message, the server uses the answering robot to provide a voice response to the user based on the complete voice message. Stopping the generation of the voice response means that the server temporarily stops providing the voice response for a period of time, waits for the next voice message from the user, and then re-checks whether the voice message uttered by the user is a complete voice message. It does not mean that the entire voice response process is terminated.
[0082] If the voice information is a complete voice that the user needs to make, the voice information is input into the preset response model, and the corresponding response voice is generated through the preset response model, and then the response voice is played through the response robot set in the server.
[0083] Furthermore, if the server determines, based on the response status, that the user was in a normal mood when sending the voice message, the server may also determine, based on the acquired voice message, whether the user's actual intention can be identified through at least one of the semantic features and contextual features corresponding to the voice message. If it is determined that the user's actual intention cannot be identified through the voice message, the voice message may be input into a preset response model, and a voice response clarifying the user's actual intention may be generated by the response model. The voice response clarifying the user's actual intention may then be played by the response robot.
[0084] For example, when a user sends a voice message containing two ambiguous meanings, A and B, you can ask the user, "Do you want to do intention A or intention B?" After determining the actual intention the user wants to express, the system generates a corresponding response voice based on the user's voice information and a preset response model, and then provides a voice response.
[0085] It is worth noting that the server can not only generate the corresponding response voice through the preset response model, but also generate the corresponding response voice using the preset speech template, that is, the server fills the corresponding content in the preset speech template to generate the corresponding response voice.
[0086] It can be seen from the above method that by performing voice environment analysis on the voice information collected during the voice response process for the user, it can be determined that the collected voice information is valid voice, thereby avoiding the influence of the dialogue environment that the answering robot may be affected by during the response process. Furthermore, based on the valid voice, the user's response status is determined, and based on the user's response status, the voice response to the user by the answering robot during the voice response process is determined. While enabling the answering robot to make special responses according to the user's response status, the answering accuracy of the answering robot and the user experience are improved.
[0087] In this specification, the server performs voice environment analysis on the voice information collected during the user's voice response process, determines the user's cooperation and response status corresponding to the voice information based on the voice information, and generates the corresponding voice response, all of which can be performed in one model. It can be understood that after obtaining the voice information in the voice response process, the voice information can be directly input into the model to determine the corresponding response voice through the model.
[0088] It should be noted that all actions of acquiring signals, information or data in this application are carried out in compliance with the data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0089] The above is a voice response method provided in one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding voice response device, such as Figure 3 shown.
[0090] Figure 3 A schematic diagram of a voice response device provided in this specification includes:
[0091] The acquisition module 301 is used to acquire the voice information collected during the voice response process of the user;
[0092] An analysis module 302 is configured to perform voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process;
[0093] A determination module 303 is configured to determine, when determining that the voice information is valid voice, a response status of the user in the voice response process based on the voice information;
[0094] The answering module 304 is used to give a voice answer to the user through the answering robot during the voice answering process according to the answering status.
[0095] Optionally, the analysis module 302 is specifically used to take the voice information collected for the first time during the voice response process as the target voice, extract the voiceprint features in the target voice as the voiceprint features of the user; for each voice information collected except the target voice during the voice response process, determine whether the voiceprint information of the voice information matches the voiceprint information of the user; if so, determine that the voice information is a valid voice uttered by the user during the voice response process.
[0096] Optionally, the analysis module 302 is specifically configured to input the voice information into a preset validity recognition model to obtain a validity recognition result of the validity recognition model for the voice information;
[0097] According to the validity recognition result, it is determined whether the voice information is a valid voice uttered by the user during the voice response process.
[0098] Optionally, the determination module 303 is specifically used to extract semantic features from the voice information and to extract compensation features from the voice response process, the compensation features including: the number of words contained in the voice information, the duration of the response voice played by the response robot when the user sends the voice information, the playback completion ratio of the response voice played by the response robot when the user sends the voice information, and at least one of the playback status of the response robot when the user sends the voice information, the playback status is used to characterize whether the response robot is playing the response voice when the user sends the voice information; the semantic features and the compensation features are input into a preset cooperation prediction model to determine the cooperation of the user when sending the voice information, and based on the cooperation, determine the response status of the user in the voice response process.
[0099] Optionally, the response module 304 is specifically used to, if it is determined based on the response status that the user has abnormal emotions when sending the voice message, determine emotional information used to characterize the abnormal emotions of the user when sending the voice message; input the emotional information and the voice information into a preset response model to determine the response voice that maintains the voice response process, and play the response voice through the response robot.
[0100] Optionally, the response module 304 is specifically used to, if it is determined according to the response status that the user is in a normal mood when sending the voice message, determine based on the voice information whether the voice information is the complete voice that the user needs to send; if the voice information is not the complete voice that the user needs to send, stop generating the voice response, and after monitoring the user to send the complete voice, give a voice response to the user through the response robot based on the complete voice.
[0101] Optionally, the response module 304 is specifically used to, if it is determined based on the response status that the user is in a normal mood when sending the voice message, determine whether the user's actual intention can be identified through the voice message; if it is determined that the user's actual intention cannot be identified through the voice message, play a voice response through the response robot to clarify the user's actual intention to the user.
[0102] Optionally, the analysis module 302 is specifically used to, if it is determined that the voice information is not a valid voice, not respond to the voice information, and after re-identifying the valid voice uttered by the user during the voice response process, provide a voice response to the user through the response robot.
[0103] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 A voice response method is provided.
[0104] This manual also provides Figure 4 The one shown corresponds to Figure 1 Schematic diagram of the electronic equipment. Figure 4 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0105] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a Hardware Description Language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0106] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.
[0107] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0108] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0109] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0110] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0111] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0112] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0113] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0114] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0115] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0116] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0117] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0118] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0119] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0120] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A voice response method, characterized in that: include: Obtaining voice information collected during the user's voice response process; Performing voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process; If it is determined that the voice message is a valid voice, semantic features are extracted from the voice message, and compensation features are extracted from the voice response process, the compensation features including: the number of words contained in the voice message, the duration of the response voice played by the response robot when the user sends the voice message, the completion ratio of the response voice played by the response robot when the user sends the voice message, and at least one of the playback status of the response robot when the user sends the voice message, the playback status being used to characterize whether the response robot is playing the response voice when the user sends the voice message, the semantic features and the compensation features are input into a preset cooperation prediction model to determine the cooperation degree of the user when sending the voice message, and the response status of the user in the voice response process is determined based on the cooperation degree; If it is determined, based on the response state, that the user has abnormal emotions when sending the voice message, determining emotional information used to characterize the abnormal emotions of the user when sending the voice message, inputting the emotional information and the voice information into a preset response model to determine a response voice that maintains the voice response process, and playing the response voice through the response robot; If it is determined based on the response status that the user is in a normal mood when sending the voice message, judge based on the voice information whether the voice information is the complete voice that the user needs to send. If the voice information is not the complete voice that the user needs to send, stop generating the voice response, and after monitoring the user to send the complete voice, give a voice response to the user through the answering robot based on the complete voice.
2. The method according to claim 1, wherein Performing voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process specifically includes: The voice information collected for the first time during the voice response process is used as the target voice; Extracting voiceprint features from the target speech as the voiceprint features of the user; For each voice message collected during the voice response process other than the target voice, determining whether the voiceprint feature of each voice message matches the voiceprint feature of the user; If so, it is determined that each voice message is a valid voice uttered by the user during the voice response process.
3. The method according to claim 1, wherein Performing voice environment analysis on the voice information to determine whether the voice information is a valid voice uttered by the user during the voice response process specifically includes: Inputting the voice information into a preset validity recognition model to obtain a validity recognition result of the validity recognition model for the voice information; According to the validity recognition result, it is determined whether the voice information is a valid voice uttered by the user during the voice response process.
4. The method according to claim 1, wherein The method further comprises: If it is determined that the voice information is not a valid voice, no response is given to the voice information, and after the valid voice uttered by the user during the voice response process is re-identified, a voice response is given to the user through the response robot.
5. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Voice interaction processing method, device, computer device and storage medium
CN109215646A
Voice processing method and device, computer equipment and storage medium
CN112992147A