Voice response method, device, apparatus and storage medium

By allowing users to select the target voice response technology and conduct multiple rounds of interaction, the problem of low communication efficiency caused by the fixed voice response technology in the existing technology is solved, and more efficient user communication and experience improvement are achieved.

CN116074441BActive Publication Date: 2025-10-10CHINA UNITED NETWORK COMM GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111293584.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-10-10
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

The voice response technology in existing intelligent answering devices is fixed and users cannot flexibly choose, resulting in low communication efficiency between calling and called users and poor user experience.

Method used

Provides a voice response method that allows users to pre-select the target voice response technology, and generates voice interaction text through multiple rounds of voice interaction and pushes it to the called user terminal.

Benefits of technology

It improves the communication efficiency between calling and called users, enhances user experience, and reduces the need for repeated calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116074441B_ABST
    Figure CN116074441B_ABST
Patent Text Reader

Abstract

The application provides a voice response method, device, equipment and storage medium. The method comprises the following steps: receiving a call transfer request of a calling user terminal, wherein the call transfer request comprises identification information of a called user; determining a target voice response technology preselected by the called user according to the identification information of the called user; calling the target voice response technology to perform multi-round voice interaction with the calling user terminal, and generating voice interaction text; and pushing the voice interaction text to a called user terminal, so that the user can flexibly select the voice response technology. The called user can determine specific matters communicated by the voice response device and the calling user according to the voice interaction text, the calling user and the called user do not need to dial a phone again to communicate, the communication efficiency of the calling user and the called user can be effectively improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a voice response method, device, equipment and storage medium. Background Art

[0002] The Smart Answering service is a value-added voice / SMS service for when the called party is unable to make a call. The called party uses the corresponding Smart Answering device to notify the calling party of the reason for the inconvenience by voice or SMS, thereby politely declining the call and achieving harmonious communication.

[0003] However, the types of voice response technologies currently included in smart answering devices are pre-defined, and users cannot flexibly select the type of voice response technology. Furthermore, if a caller's call is not answered, the current smart answering service only provides a reminder that it is inconvenient to answer the call. This results in the caller or called party having to dial back to discuss specific matters based on the reminder data, resulting in low communication efficiency and a poor user experience. Summary of the Invention

[0004] The present invention provides a voice response method, device, equipment and storage medium to solve the technical problem that the calling user or the called user needs to dial the phone again to communicate specific matters according to reminder data, the communication efficiency between the calling user and the called user is low, and the user experience is poor.

[0005] In a first aspect, the present invention provides a voice response method, comprising:

[0006] receiving a call forwarding request from a calling user terminal, wherein the call forwarding request includes identification information of the called user;

[0007] determining, based on the identification information of the called user, a target voice response technology preselected by the called user;

[0008] Invoking the target voice response technology to perform multiple rounds of voice interaction with the calling user terminal and generating voice interaction text;

[0009] Push the voice interaction text to the called user terminal.

[0010] In a second aspect, the present invention provides a voice response device, comprising:

[0011] A receiving module, configured to receive a call forwarding request from a calling user terminal, wherein the call forwarding request includes identification information of the called user;

[0012] A determination module, configured to determine a target voice response technology preselected by the called user based on the identification information of the called user;

[0013] a voice interaction module, configured to invoke the target voice response technology to perform multi-round voice interaction with the calling user terminal and generate voice interaction text;

[0014] a pushing module, configured to push the voice interaction text to the called user terminal.

[0015] In a third aspect, the present application provides an electronic device, comprising: at least one processor; and

[0016] a memory in communication connection with the at least one processor; wherein,

[0017] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect.

[0018] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by a processor to implement the method of the first aspect.

[0019] The voice response method, device, equipment and storage medium provided by the embodiments of the present application can receive a call transfer request of a calling user terminal, the call transfer request comprising identification information of a called user; determine a target voice response technology pre-selected by the called user according to the identification information of the called user; invoke the target voice response technology to perform multi-round voice interaction with the calling user terminal and generate voice interaction text; and push the voice interaction text to the called user terminal. Since the user can pre-select a target voice response technology from multiple types of voice response technologies, the user can flexibly select a voice response technology. The target voice response technology can be invoked to perform multi-round voice interaction with the calling user terminal, and after the interaction is completed, voice interaction text can be generated and pushed to the called user terminal. The called user can determine specific matters communicated between the voice response device and the calling user according to the voice interaction text, and the calling user and the called user do not need to make a call again to communicate, which can effectively improve the communication efficiency of the calling user and the called user and improve the user experience. BRIEF DESCRIPTION OF DRAWINGS

[0020] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, further serve to explain the principles of the present disclosure.

[0021] Figure 1 A network architecture diagram of a voice response method provided by an embodiment of the present application;

[0022] Figure 2 A flow chart of a voice response method according to an embodiment of the present application;

[0023] Figure 3 A flow chart of a voice response method according to another embodiment of the present application;

[0024] Figure 4 A flow chart of a voice response method according to yet another embodiment of the present application;

[0025] Figure 5 A flow chart of a voice response method according to still another embodiment of the present application;

[0026] Figure 6 A flow chart of a voice response method according to yet another embodiment of the present application;

[0027] Figure 7 A flow chart of a voice response method according to still another embodiment of the present application;

[0028] Figure 8 A signaling interaction flow chart of a voice response method according to yet another embodiment of the present application;

[0029] Figure 9 A structural schematic diagram of a voice response apparatus according to an embodiment of the present application;

[0030] Figure 10 A block diagram of an electronic device according to an embodiment of the present application.

[0031] The specific embodiments of the present disclosure have been shown through the above-described drawings, and will be described in more detail hereinafter. The drawings and the written description are not intended to restrict the scope of the present disclosure by any means, but to illustrate the concept of the present disclosure to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0032] The exemplary embodiments will be described in detail herein below with reference to the accompanying drawings. In the following description, the same drawings refer to the same elements or similar elements throughout the different drawings. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure, as detailed in the appended claims.

[0033] First, the prior art related to the embodiments of the present application is described and analyzed in detail.

[0034] In the prior art, when smart answering devices provide smart answering services, since the smart answering devices are developed by a single manufacturer, the voice answering technology included in the smart answering devices is also developed by that manufacturer. In other words, the voice answering technology is fixed. However, since there are many manufacturers currently providing smart answering devices, each manufacturer's voice answering technology has its own advantages. Current smart answering devices do not allow users to flexibly select the voice answering technology that best suits them. Therefore, it is difficult to find the voice answering technology that best suits them. Furthermore, the smart answering services provided in the prior art only provide a reminder function when the caller's call is not answered. For example, if the caller's call is unavailable, only a fixed reminder tone is played, such as "The user is in a meeting and cannot answer the call. I will notify him to call you back as soon as possible." After the called user learns that the caller has called, the calling user or the called user needs to dial back to discuss the details based on the reminder data. This reduces communication efficiency between the calling and called users and results in a poor user experience.

[0035] Therefore, in the face of the technical problem in the prior art that the current intelligent response device cannot allow users to flexibly select the voice response technology independently, the inventor discovered through creative research that in order to enable users to flexibly select the voice response technology independently, a voice response technology configuration function can be provided to users, and various types of voice response technology details information can be pushed to users. The user can select a type of voice response technology as the target voice response technology based on the voice response technology details information. When the user is subsequently the called user and the calling user terminal requests a call but the call is not answered, and the calling user terminal performs call forwarding, the target voice response technology pre-selected by the called user can be obtained based on the called user's identification information, and the target voice response technology can be called to perform voice interaction with the calling user terminal. This enables users to flexibly select the voice response technology.

[0036] In the face of the technical problems of low communication efficiency and poor user experience between the calling user and the called user in the existing technology, the inventors discovered through creative research that various voice response technologies can provide not only a reminder function to the calling user that it is inconvenient to answer the phone, but also multiple rounds of voice interaction with the calling user. Therefore, the target voice response technology can be called to conduct multiple rounds of voice interaction with the calling user terminal. And after the interaction is completed, a voice interaction text can be generated and pushed to the called user terminal. The called user can determine the specific matters to be communicated with the intelligent answering system and the calling user based on the voice interaction text. The calling user and the called user do not need to make a phone call again to communicate, which can effectively improve the communication efficiency between the calling user and the called user and improve the user experience.

[0037] Therefore, based on the above creative findings, the inventors have proposed the technical solutions of the embodiments of the present invention. The network structure of the voice response method provided by the embodiments of the present invention is introduced below.

[0038] Figure 1 A network architecture diagram of a voice response method provided by an embodiment of the present invention, such as Figure 1 As shown, the network architecture corresponding to the voice response method provided in this embodiment includes a calling user terminal 11, a called user terminal 12, a core network device 13, a universal access platform (UAP) 14, electronic devices, various voice response technology servers, and a service platform 19. Among them, voice response technologies may include natural language processing technology (NLP), speech recognition technology (ASR), and text-to-speech synthesis technology (TTS). Various types of NLP can be installed in a server 16. Similarly, various types of ASR can be installed in a server 17. Various types of TTS can be installed in a server 18. A voice response device is installed in the electronic device. The voice response device can be an interactive voice response device (IVR) 15. Specifically, the calling user terminal 11, the core network device 13, the UAP 14, and the IVR 15 are sequentially connected in communication. The IVR 15 is connected in communication with a server 17 loaded with various types of NLPs. The UAP 14 is connected to various ASR servers 17 and various TTS servers 18. The IVR 15 is also connected to a service platform 19, which is connected to the called user terminal 12.

[0039] The following specific embodiments describe in detail the technical solutions of the present invention and how the technical solutions of this application solve the above-mentioned technical problems. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The following embodiments of the present invention are described in conjunction with the accompanying drawings.

[0040] Figure 2 A flowchart of a voice response method provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown, the voice response method provided in this embodiment is performed by a voice response device. Specifically, it can be an IVR. The IVR is located in an electronic device. The voice response method provided in this embodiment includes the following steps:

[0041] Step 201: Receive a call forwarding request from a calling user terminal, where the call forwarding request includes identification information of the called user.

[0042] In the embodiment, the calling user can be set to be transferred before the called user terminal enables the voice response device. Or the specific user such as the express user or the take-out user can be set to be transferred.

[0043] In the embodiment, the call transfer condition of the calling user can also be set before the called user terminal enables the voice response device. For example, the call transfer condition of the calling user can be that the called user does not answer the phone for more than a preset time, or the called user actively hangs up the phone, or the calling user terminal makes a call within a preset time period. The preset time period can be, for example, a rest time period, a meeting time period, etc.

[0044] Specifically, in the embodiment, when the calling user terminal dials the called user terminal and meets the preset call transfer condition of the calling user, the calling user terminal sends a call transfer request to the IVR through the core network device and the UAP in sequence. After receiving the call transfer request, the IVR analyzes the call transfer request and obtains the identification information of the called user and the identification information of the calling user.

[0045] The identification information of the called user can be an international mobile subscriber identity (IMSI) corresponding to the called user or a mobile phone number, etc. Similarly, the identification information of the calling user can be an IMSI or a mobile phone number of the calling user, etc.

[0046] In step 202, the target voice response technology preselected by the called user is determined according to the identification information of the called user.

[0047] In the embodiment, the data corresponding to each user using the voice response device can be stored in a local preset storage space or in a service platform.

[0048] The user using the voice response device can preselect a corresponding voice response technology as a target voice response technology from a plurality of types of voice response technologies before using the voice response function of the voice response device, and the identification information of the user and the identification information of the target voice response technology can be associated and stored. Therefore, after the identification information of the called user is determined, the voice response technology associated and stored with the identification information of the called user is obtained as the target voice response technology by accessing the local preset storage space or the service platform.

[0049] The target voice response technology preselected by the called user can include a target natural language processing technology, a target speech recognition technology, and a target speech recognition technology.

[0050] Step 203: Invoke the target voice response technology to perform multiple rounds of voice interaction with the calling user terminal and generate voice interaction text.

[0051] In this embodiment, the target voice response technology is first invoked to determine the opening remarks for the call with the calling user. A play command is then sent to the UAP, which plays the opening remarks to the calling user's terminal. The target voice response technology is subsequently invoked to conduct multiple rounds of voice interaction with the calling user's terminal. During each round of semantic interaction, the target voice recognition technology is invoked to perform voice recognition on the calling user's voice data, generating calling user text data. The target natural language processing technology is invoked to perform semantic recognition on the calling user's text data, and matching response data is determined based on the semantic recognition results. The response data is then processed through speech synthesis using the target speech synthesis technology to generate voice-type response data. A play command is then sent to the UAP, which plays the response data to the calling user's terminal.

[0052] In this embodiment, if it is determined that the voice interaction text generation conditions are met, the calling user text data and response data of each round of voice interaction are obtained, and the texts of multiple rounds of voice interaction are spliced ​​together as voice interaction text.

[0053] Step 204: Push the voice interaction text to the called user terminal.

[0054] In this embodiment, it is determined whether the call stop condition is met. If the call stop condition is met, the voice interaction text is pushed to the called user terminal. The push format is not limited, such as SMS, instant messaging, message, email, etc.

[0055] The voice response method provided in this embodiment receives a call forwarding request from a calling user terminal, the call forwarding request including the called user's identification information; determines a target voice response technology pre-selected by the called user based on the called user's identification information; invokes the target voice response technology to conduct multiple rounds of voice interaction with the calling user terminal, generating a voice interaction text; and pushes the voice interaction text to the called user terminal. Because the user can pre-select a target voice response technology from a variety of voice response technologies, users can flexibly select a voice response technology. Furthermore, the target voice response technology can be invoked to conduct multiple rounds of voice interaction with the calling user terminal. After the interaction is complete, a voice interaction text is generated and pushed to the called user terminal. The called user can determine the specific matters to be communicated between the voice response device and the calling user based on the voice interaction text, eliminating the need for the calling and called users to make another call. This effectively improves communication efficiency between the calling and called users and enhances the user experience.

[0056] Example 2

[0057] Figure 3 A flowchart of a voice response method provided in another embodiment of the present invention is shown in FIG. Figure 3 As shown, the voice response method provided in this embodiment is based on the voice response method provided in Example 1, and further includes other steps before step 202. The voice response method provided in this embodiment includes the following steps:

[0058] Step 301: Receive a voice response technology configuration request sent by a called user terminal.

[0059] In this embodiment, the voice response device can be accessed by the called user terminal in the form of a client or webpage. After the called user logs in to the client or webpage using their username, a voice response technology configuration component can be displayed on the operation interface. The called user triggers a voice response technology configuration request through the voice response technology configuration component. Upon receiving the voice response technology configuration request, the called user terminal then sends the voice response technology configuration request to the voice response device.

[0060] Step 302: Obtain various types of voice response technology details information according to the voice response technology configuration request and send the information to the called user terminal; the voice response technology details information includes identification information of the voice response technology.

[0061] In this embodiment, each type of voice response technology is developed by a single manufacturer. Therefore, each type of voice response technology corresponds to a specific brand. The voice response device may pre-store detailed information for each type of voice response technology. The detailed information for each type of voice response technology may include identification information for the voice response technology, brand information, response range of the voice response technology, and advantages. The identification information for the voice response technology may include a serial number for the voice response technology or a manufacturer code.

[0062] Specifically, in this embodiment, after receiving a voice response technology configuration request, the voice response device obtains detailed information about various types of voice response technologies based on the voice response technology configuration request, and then sends the detailed information about various types of voice response technologies to the called user terminal. The called user terminal then displays the detailed information about various types of voice response technologies, allowing the called user to view the detailed information about each type of voice response technology and providing a basis for the called user to select a target voice response technology.

[0063] Step 303: Receive the identification information of the target voice response technology selected by the called user sent by the called user terminal, and send the identification information of the called user and the identification information of the target voice response technology to the service platform so that the service platform associates and stores the identification information of the called user and the identification information of the target voice response technology.

[0064] In this embodiment, the called user can determine the target voice response technology that meets his or her needs based on the detailed information of each type of voice response technology. The target voice response technology can be selected through the selection component in the client or web page corresponding to the voice response device. After the called user terminal receives the target voice response technology selected by the called user, it sends the identification information of the target voice response technology selected by the called user to the voice response device. The voice response device sends the identification information of the called user and the identification information of the target voice response technology to the service platform. The service platform associates and stores the identification information of the called user and the identification information of the target voice response technology, and can select the voice response technology associated with the identification information of the called user as the target voice response technology when the called user uses the voice response device.

[0065] It should be noted that if the called user has a poor experience when using the voice response device, the corresponding target voice response technology can be changed through the voice response device client or the web page. Specifically, after the called user triggers the voice response technology change request, the voice response device obtains multiple types of voice response technology detail information according to the voice response technology change request and sends it to the called user terminal, and the called user terminal displays the multiple types of voice response technology detail information. The called user views each type of voice response technology detail information, which provides a basis for the called user to change the target voice response technology. After the called user terminal selects the changed target voice response technology, the voice response device receives the identification information of the changed target voice response technology sent by the called user terminal, and sends the identification information of the called user and the identification information of the changed target voice response technology to the service platform, so that the service platform associates and stores the identification information of the called user and the identification information of the changed target voice response technology.

[0066] The voice response method provided in this embodiment receives a voice response technology configuration request sent by the called user terminal before determining the voice response technology pre-selected by the called user based on the identification information of the called user; obtains multiple types of voice response technology detail information based on the voice response technology configuration request and sends it to the called user terminal; the voice response technology detail information includes the identification information of the voice response technology; receives the identification information of the target voice response technology selected by the called user sent by the called user terminal, and sends the identification information of the called user and the identification information of the target voice response technology to the service platform, so that the service platform associates and stores the identification information of the called user and the identification information of the target voice response technology. The called user can select the target voice response technology based on the multiple types of voice response technology detail information, which can make the selected target voice response technology more suitable for the called user, and the identification information of the called user and the identification information of the target voice response technology are associated and stored in the service platform, which can effectively manage the relationship between the user and the voice response technology.

[0067] Example 3

[0068] Figure 4 A flowchart of a voice response method provided by another embodiment of the present invention is shown in FIG. Figure 4 As shown, the voice response method provided in this embodiment is based on the voice response method provided in Example 2, and further refines step 102. The voice response method provided in this embodiment includes the following steps:

[0069] Step 401: Send a voice response technology determination request to the service platform. The voice response technology determination request includes the identification information of the called user, so that the service platform can determine the identification information of the associated stored voice response technology based on the identification information of the called user.

[0070] In this embodiment, the target voice response technology corresponding to each user is stored in the service platform. Therefore, when the voice response device determines the target voice response technology preselected by the called user based on the called user's identification information, it sends a voice response technology determination request to the service platform. The voice response technology determination request carries the called user's identification information. After obtaining the called user's identification information, the service platform queries the stored data to obtain the identification information of the voice response technology associated with the called user's identification information.

[0071] Step 402: Receive identification information of the associated stored voice response technology sent by the service platform.

[0072] Step 403: Determine the target voice response technology based on the identification information of the associated stored voice response technology.

[0073] In this embodiment, the voice response device receives the identification information of the voice response technology stored in association with the identification information of the called user. Since the identification information of the voice response technology stored in association in the service platform is the identification information of the voice response technology selected independently by the called user, the associated stored voice response technology is determined as the target voice response technology.

[0074] The voice response method provided in this embodiment, when determining the target voice response technology pre-selected by the called user based on the identification information of the called user, sends a voice response technology determination request to the service platform, the voice response technology determination request includes the identification information of the called user, so that the service platform determines the identification information of the associated stored voice response technology based on the identification information of the called user; receives the identification information of the associated stored voice response technology sent by the service platform; and determines the target voice response technology based on the identification information of the associated stored voice response technology. By associating and storing the identification information of the called user and the identification information of the target voice response technology through the service platform, the association relationship between the user and the voice response technology can be effectively managed. And by sending a voice response technology determination request to the service platform, and carrying the identification information of the called user in the voice response technology determination request, the service platform can quickly obtain the target voice response technology based on the identification information of the called user.

[0075] Example 4

[0076] Figure 5 A flowchart of a voice response method provided in yet another embodiment of the present invention is shown in FIG. Figure 5 As shown, the voice response method provided in this embodiment is based on the voice response method provided in any of the above embodiments, and further refines step 103. The target voice response technology includes: target voice recognition technology, natural language processing technology, and target voice synthesis technology. The following steps are performed when the target voice response technology is invoked for each round of voice interaction with the calling user terminal:

[0077] Step 501: receiving calling user voice data sent by a calling user terminal.

[0078] In this embodiment, the primary user terminal sends the calling user's voice data to the voice response device through the UAP. The voice response device receives the calling user's voice data through the UAP.

[0079] It can be understood that the main user voice data is of voice type.

[0080] Step 502: Invoke the target speech recognition technology to perform speech recognition on the voice data of the calling user to obtain text data of the calling user.

[0081] As an optional implementation, in this embodiment, before step 502, the following solution is also included:

[0082] An engine startup request corresponding to the target speech recognition technology is sent to the UAP, where the engine startup request includes identification information of the target speech recognition technology, so that the UAP loads the target grammar file corresponding to the identification information of the target speech recognition technology according to the engine startup request, and controls the startup of the target speech recognition engine according to the target grammar file.

[0083] The UAP locally stores grammar files corresponding to various speech recognition technologies. These files are used by the corresponding speech recognition technologies to recognize speech. Each grammar file includes at least one recognition strategy, and the speech recognition technology uses this strategy to recognize speech and generate corresponding text data.

[0084] Specifically, since the target speech recognition technology is not in the voice response device but in a preset server, it is necessary to start the target speech recognition engine corresponding to the target speech recognition technology. First, the voice response device sends an engine start request corresponding to the target speech recognition technology to the UAP. The UAP parses the engine start request and obtains the identification information of the target speech recognition technology. The UAP can then load the target grammar file corresponding to the identification information of the target speech recognition technology from the local storage area, and then control the start of the target speech recognition engine in the server.

[0085] Accordingly, step 502 specifically includes the following solutions:

[0086] The target speech recognition engine is called to perform speech recognition on the calling user's speech data to obtain the calling user's text data.

[0087] Specifically, in this embodiment, the caller's voice data is transmitted to the server corresponding to the target speech recognition engine via the Real-time Transport Protocol (RTP). The server corresponding to the target speech recognition engine invokes the target speech recognition engine to obtain the caller's voice data, then uses the corresponding target grammar file to perform speech recognition on the caller's voice data and outputs text data corresponding to the caller's voice data.

[0088] Step 503: Invoke the target natural language processing technology to perform semantic recognition on the calling user's text data, and determine matching response data based on the semantic recognition result.

[0089] In this embodiment, the target natural language processing technology may be installed in a preset server, and the preset server also stores response data corresponding to each semantic intent.

[0090] Specifically, the voice response device communicates with a server equipped with the target natural language processing technology, accessing it via the HTTP protocol. The target natural language processing technology then performs semantic recognition on the caller's text data, obtaining a semantic recognition result. This semantic recognition result includes the semantic intent. Based on the semantic intent, the pre-set server retrieves matching response data from pre-stored response data and sends it to the voice response device.

[0091] Among them, the response data corresponding to each semantic intent stored in the preset server can be voice type data or text type data.

[0092] Step 504: If the response data is determined to be voice data, a play instruction is sent to the universal access platform UAP, where the play instruction includes the response data, so that the UAP plays the response data to the calling user terminal.

[0093] In this embodiment, the voice response device determines the type of response data. If the response data is determined to be voice data, the response data can be sent to the UAP via a play command without using target speech synthesis technology. The UAP then plays the response data to the calling user terminal.

[0094] Step 505: If it is determined that the response data is text type data, the target speech synthesis technology is called to perform speech synthesis processing on the response data to generate voice type response data, and a playback instruction is sent to the universal access platform UAP. The playback instruction includes the response data, so that the UAP plays the response data to the calling user terminal.

[0095] In this embodiment, if the response data is determined to be text type data, in order to play the response data to the calling user terminal, the target speech synthesis technology is called to perform speech synthesis processing on the text type data.

[0096] Specifically, when calling the target speech synthesis technology to perform speech synthesis processing on the response data, the target speech synthesis engine corresponding to the target speech synthesis technology may be started and called to perform speech synthesis processing on the response data.

[0097] The voice response method provided in this embodiment, when invoking a target voice response technology to conduct multiple rounds of voice interaction with a calling user terminal, performs the following operations during each round of voice interaction: receives calling user voice data sent by the calling user terminal; invokes the target voice recognition technology to perform voice recognition on the calling user voice data to obtain calling user text data; invokes the target natural language processing technology to perform semantic recognition on the calling user text data and determines matching response data based on the semantic recognition results; if the response data is determined to be voice data, sends a play command to the universal access platform (UAP), including the response data, so that the UAP plays the response data to the calling user terminal. If the response data is determined to be text data, invokes the target voice synthesis technology to perform voice synthesis processing on the response data to generate voice-type response data. Furthermore, sends a play command to the universal access platform (UAP), so that the UAP plays the response data to the calling user terminal. This ensures that the calling user can hear a response that matches the calling user's voice each time the calling user speaks. This effectively improves user experience and communication efficiency.

[0098] Example 5

[0099] Figure 6 A flowchart of a voice response method provided by another embodiment of the present invention is shown as follows: Figure 6 As shown, the voice response method provided in this embodiment is based on the voice response method provided in the fourth embodiment, and further includes other steps after step 502. The voice response method provided in this embodiment includes the following steps:

[0100] Step 601: Control the target speech recognition engine to use the improved MRCP protocol to feed back speech recognition results; the speech recognition results include identification information of the target speech recognition technology and speech recognition status information.

[0101] In this embodiment, the voice response device controls the target voice recognition engine to call the improved MRCP protocol to first feed back the voice recognition result to the UAP, and the UAP then feeds back the voice recognition result to the voice response device.

[0102] Among them, the improved MRCP protocol is an improvement on the original MRCP protocol. The speech recognition results fed back by calling the improved MRCP protocol include the identification information of the target speech recognition technology and the speech recognition situation information. Among them, the identification information of the target speech recognition technology can be represented by the serial number of the target voice response technology, or by the corresponding manufacturer code. The speech recognition situation information can be represented by two fields, namely the cause code (causeCode) and the cause name (causeName). Among them, the speech recognition situation information can be represented as shown in Table 1:

[0103] Table 1: Speech recognition information diagram

[0104]

[0105] It should be noted that the speech recognition result fed back using the improved MRCP protocol may also include recognized text data of the calling user and the calling user's voice data and the storage path.

[0106] For example, the message format used by the speech recognition result is as follows:

[0107] <nlresult>causeCode@causeName@caller text data@identification information of target speech recognition technology@caller speech data and stored path< / nlresult> .

[0108] Among them, in two <nlresult>The fields between are the speech recognition results. @ is the separator of each field.

[0109] The causeCode and causeName fields must be non-null. The calling user's text data may be null. The target speech recognition technology's identification information must not be null. The calling user's voice data and its storage path may be null.

[0110] Step 602: If it is determined based on the speech recognition situation information that the target speech recognition engine replacement condition is met, a candidate speech recognition engine is obtained and called to perform speech recognition on the calling user's speech data.

[0111] In this embodiment, after receiving the voice recognition situation information, the voice response device parses the two fields causeCode and causeName, and then determines whether the target voice recognition engine replacement conditions are met. As shown in Table 1, the situations with sequence numbers 2 and 4 meet the target voice recognition engine replacement conditions. If it is determined that the target voice recognition engine replacement conditions are met, the voice response device sends a candidate voice recognition technology acquisition request to the business platform. The business platform obtains the identification information of the candidate voice recognition technology and sends it to the voice response device. The voice response device sends an engine startup request corresponding to the candidate voice recognition technology to the UAP. The engine startup request includes the identification information of the candidate voice recognition technology, so that the UAP loads the candidate grammar file corresponding to the identification information of the candidate voice recognition technology according to the engine startup request, and controls the startup of the candidate voice recognition engine according to the candidate grammar file, and then calls the candidate voice recognition engine to perform voice recognition on the calling user's voice data.

[0112] When determining the identification information of candidate voice recognition technologies, the service platform may pre-set the candidate voice recognition technologies and thereby obtain the identification information of the candidate voice recognition technologies. Alternatively, when the called user independently selects a target voice recognition technology, the called user may also select the corresponding candidate voice recognition technology and associate the identification information of the candidate voice recognition technology with the identification information of the called user and store it in the service platform, thereby allowing the service platform to obtain the identification information of the candidate voice recognition technology based on the identification information of the called user.

[0113] The voice response method provided in this embodiment, after invoking a target voice recognition engine to perform voice recognition on the caller's voice data to obtain the caller's text data, further includes: controlling the target voice recognition engine to feedback a voice recognition result using an improved MRCP protocol; the voice recognition result includes identification information of the target voice recognition technology and voice recognition status information; if it is determined based on the voice recognition status information that a target voice recognition engine replacement condition is met, obtaining a candidate voice recognition engine and invoking the candidate voice recognition engine to perform voice recognition on the caller's voice data. The improved MRCP protocol can be used to feedback the voice recognition result with identification information of the target voice recognition technology, and thus, if the target voice recognition engine cannot recognize the caller's voice data, the voice recognition technology and engine can be replaced in a timely manner to ensure that the caller's voice data can be successfully recognized.

[0114] Example 6

[0115] Figure 7 A flowchart of a voice response method provided in yet another embodiment of the present invention is shown in FIG. Figure 7 As shown, the voice response method provided in this embodiment further includes other steps based on the voice response method provided in any of the above embodiments. The voice response method provided in this embodiment includes the following steps:

[0116] Step 701: If it is determined that the current load of the target voice response technology exceeds a preset load, a voice response technology update request is sent to the called user terminal.

[0117] In this embodiment, the load of each target voice response technology can be periodically monitored. If, after determining the target voice response technology pre-selected by the called user based on the called user's identification information, the current load corresponding to the target voice response technology exceeds the preset load, this indicates that a large number of called users are using the target voice response technology. If the called user continues to use the target voice response technology for subsequent rounds of voice interaction, lag will occur. Therefore, to ensure the quality of the current call, a voice response technology update request is sent to the called user terminal, causing the called user terminal to display the voice response technology update request.

[0118] The voice response technology update request may include identification information of at least one candidate voice response technology, wherein the current load corresponding to the candidate voice response technology is low and does not reach a preset load.

[0119] Step 702: If a voice response technology update response is received from the called user terminal, the identification information of the candidate voice response technology in the voice response technology update response is obtained, and the candidate voice response technology is called to perform multiple rounds of voice interaction with the calling user terminal.

[0120] In this embodiment, the called user can select a candidate voice response technology. The called user terminal then generates a voice response technology update response and sends it to the voice response device. After receiving the voice response technology update response from the called user terminal, the voice response device parses the response, obtains the identification information of the candidate voice response technology, and then invokes the candidate voice response technology to conduct multiple rounds of voice interaction with the calling user terminal. The specific interaction process is similar to the process of invoking the target voice response technology for multiple rounds of voice interaction with the calling user terminal and will not be further described here.

[0121] It should be noted that if the called user does not perform any operation within the preset time or chooses to reject the update operation, the target voice response technology will continue to be called to perform multiple rounds of voice interaction with the calling user terminal.

[0122] The voice response method provided in this embodiment sends a voice response technology update request to the called user terminal if it is determined that the current load of the target voice response technology exceeds the preset load; if a voice response technology update response is received from the called user terminal, the identification information of the candidate voice response technology in the voice response technology update response is obtained, and the candidate voice response technology is called to perform multiple rounds of voice interaction with the calling user terminal. In order to ensure that each voice response technology can perform multiple rounds of voice interaction with the calling user terminal within the load capacity. After the current load of the target voice response technology exceeds the preset load, by updating the target voice response technology to the candidate voice response technology, the smoothness of the voice interaction with the calling user can be guaranteed and the jamming phenomenon can be avoided as much as possible.

[0123] As an optional implementation, based on any of the above embodiments, step 204 specifically includes pushing the voice interaction text to the called user terminal in a preset form if it is monitored that the call stop condition is met.

[0124] The preset form is any one or more of the following forms:

[0125] SMS form, instant messaging form, message form, email form.

[0126] In this embodiment, if it is detected that the calling user terminal hangs up or does not receive the calling user's voice data for a long time, it is determined that the call termination condition is met, and the voice interaction text is pushed to the called user terminal in at least one preset form, so that the calling user terminal can view the voice interaction text in time.

[0127] Example 7

[0128] Figure 8 A signaling interaction flow chart of a voice response method provided in another embodiment of the present invention is shown in FIG. Figure 8 As shown, the voice response method provided in this embodiment is implemented by a voice response system. The voice interaction method is described using an opening statement and a round of voice interaction with the called user terminal as an example. The voice response method provided in this embodiment includes the following steps:

[0129] Step 801: The calling user terminal sends a call forwarding request to the UAP via the core network device.

[0130] Among them, after the calling user terminal sends the call forwarding request to the core network device, the core network device sends the call forwarding request to the UAP through the traffic routing.

[0131] In step 802, the UAP sends a call forwarding request to the IVR.

[0132] The call forwarding request includes identification information of the called user.

[0133] Step 803: The IVR sends a voice response technology determination request to the service platform.

[0134] Step 804: The service platform determines the identification information of the associated stored voice response technology based on the identification information of the called user.

[0135] In this embodiment, the voice response technology determination request carries the called user's identification information. The service platform determines the identification information of the associated stored voice response technology based on the called user's identification information, and uses the associated stored voice response technology identification information as the target voice response technology identification information.

[0136] Step 805: The service platform sends identification information of the target voice response technology to the IVR.

[0137] Step 806: The IVR sends an opening statement request to the NLP server according to the identification information of the target voice response technology.

[0138] In this embodiment, the opening remarks request includes identification information of the target voice response technology. The NLP server determines the target natural language processing technology based on the identification information of the target voice response technology, and uses the target natural language processing technology to obtain a matching opening remarks.

[0139] Step 807: The NLP server sends a notification of the opening remarks to the IVR.

[0140] The opening remarks announcement may include opening remarks data, which may be voice data or text data.

[0141] Step 808: The IVR sends a notification to play the opening remarks to the UAP.

[0142] In the opening announcement, the announcement parameters can be configured. For example, the announcement parameters can include whether the called user is allowed to be interrupted by the calling user.

[0143] In step 809, the UAP plays the opening announcement to the calling user terminal.

[0144] In this embodiment, if the opening announcement is voice type data, the UAP directly plays the opening announcement to the calling user terminal. If the opening announcement data is text type data, the UAP sends the text type data to the target speech synthesis technology server, performs speech synthesis processing on the opening announcement data, and returns the voice type opening announcement data to the UAP, which plays the opening announcement to the calling user terminal.

[0145] In step 810, the IVR sends an engine start request corresponding to the target speech recognition technology to the UAP.

[0146] In step 811, the UAP loads the target grammar file corresponding to the identification information of the target speech recognition technology according to the engine start request.

[0147] In step 812, the UAP sends a start instruction to the target speech recognition engine server to control the target speech recognition engine to start.

[0148] In step 813, the target speech recognition engine server sends a start result to the UAP.

[0149] If the start result is successful, step 814 is performed, otherwise step 811 is re-executed to start the target speech recognition engine again.

[0150] In step 814, the calling user terminal sends the calling user voice data to the IVR through the UAP.

[0151] In step 815, the IVR sends the calling user voice data to the target speech recognition engine in the target speech recognition engine server.

[0152] In step 816, the target speech recognition engine performs speech recognition on the calling user voice data to obtain calling user text data.

[0153] In step 817, the target speech recognition engine feeds back the speech recognition result to the UAP using the improved MRCP protocol.

[0154] In step 818, the UAP feeds back the speech recognition result to the IVR.

[0155] In step 819, the IVR sends the calling user text data to the NLP server.

[0156] Step 820: The NLP server determines matching response data based on the calling user's text data.

[0157] Step 821: The NLP server sends response data to the IVR.

[0158] In step 822, if the IVR determines that the response data is text type data, the IVR sends the response data to the target speech synthesis engine in the target speech synthesis engine server.

[0159] Step 823: The target speech synthesis engine performs speech synthesis processing on the response data to generate speech-type response data.

[0160] Step 824: The IVR receives the response data.

[0161] Step 825: The IVR sends a playback instruction to the UAP.

[0162] In step 826, the UAP plays the response data to the calling user terminal.

[0163] In this embodiment, the implementation of steps 814 to 826 is similar to the implementation of the corresponding steps above, and will not be described in detail here.

[0164] Example 8

[0165] Figure 9 A schematic diagram of the structure of a voice response device provided by an embodiment of the present invention is shown in FIG. Figure 9 As shown, the voice response device 900 provided in this embodiment is located in an electronic device and includes: a receiving module 901 , a determining module 902 , a voice interaction module 903 , and a pushing module 904 .

[0166] Receiving module 901 is configured to receive a call forwarding request from a calling user terminal, which includes the identification information of the called user. Determining module 902 is configured to determine the target voice response technology pre-selected by the called user based on the identification information of the called user. Voice interaction module 903 is configured to invoke the target voice response technology to conduct multiple rounds of voice interaction with the calling user terminal and generate voice interaction text. Push module 904 is configured to push the voice interaction text to the called user terminal.

[0167] The voice response device provided in this embodiment can execute the method for preventing the container truck from lifting provided in the above embodiment 1. The specific implementation method and principle are similar and will not be described in detail.

[0168] Optionally, the voice response device provided in this embodiment further includes: an acquisition module and a sending module.

[0169] Accordingly, receiving module 901 is further configured to receive a voice response technology configuration request sent by a called user terminal. The acquiring module is configured to acquire detailed information about various types of voice response technologies based on the voice response technology configuration request and transmit it to the called user terminal; the detailed information about the voice response technologies includes identification information about the voice response technologies. Receiving module 901 is further configured to receive identification information about the target voice response technology selected by the called user, sent by the called user terminal. The transmitting module is configured to transmit the identification information of the called user and the identification information of the target voice response technology to the service platform, so that the service platform associates and stores the identification information of the called user and the identification information of the target voice response technology.

[0170] Optionally, the determining module 902 is specifically configured to:

[0171] A voice response technology determination request is sent to the service platform, wherein the voice response technology determination request includes the identification information of the called user, so that the service platform determines the identification information of the associated stored voice response technology based on the identification information of the called user; the identification information of the associated stored voice response technology is received from the service platform; and the target voice response technology is determined based on the identification information of the associated stored voice response technology.

[0172] Optionally, the target voice response technology includes: target voice recognition technology and target natural language processing technology.

[0173] The voice interaction module 903 is specifically used to:

[0174] When calling the target voice response technology to conduct each round of voice interaction with the calling user terminal, the following operations are performed: receiving the calling user voice data sent by the calling user terminal; calling the target voice recognition technology to perform voice recognition on the calling user voice data to obtain the calling user text data; calling the target natural language processing technology to perform semantic recognition on the calling user text data, and determining the matching response data based on the semantic recognition result; if it is determined that the response data is voice type data, sending a playback instruction to the universal access platform UAP, and the playback instruction includes the response data, so that the UAP plays the response data to the calling user terminal.

[0175] Optionally, the target voice response technology further includes a target voice synthesis technology;

[0176] Correspondingly, the voice interaction module 903 is further configured to, if it is determined that the response data is text type data, call the target voice synthesis technology to perform voice synthesis processing on the response data before sending the playback instruction to the called user terminal to generate voice type response data.

[0177] Optionally, the sending module is further configured to:

[0178] Sending an engine start request corresponding to the target speech recognition technology to the UAP, where the engine start request includes identification information of the target speech recognition technology, so that the UAP loads a target grammar file corresponding to the identification information of the target speech recognition technology according to the engine start request, and controls the start of the target speech recognition engine according to the target grammar file;

[0179] Optionally, when the voice interaction module 903 calls the target voice recognition technology to perform voice recognition on the calling user's voice data to obtain the calling user's text data, it is specifically configured to:

[0180] The target speech recognition engine is called to perform speech recognition on the calling user's speech data to obtain the calling user's text data.

[0181] Optionally, the voice interaction module 903 is further configured to:

[0182] The target speech recognition engine is controlled to use the improved MRCP protocol to feedback speech recognition results; the speech recognition results include identification information of the target speech recognition technology and speech recognition status information; if it is determined based on the speech recognition status information that the target speech recognition engine replacement conditions are met, a candidate speech recognition engine is obtained and called to perform speech recognition on the calling user's voice data.

[0183] Optionally, the sending module is further configured to send a voice response technology update request to the called user terminal if it is determined that the current load of the target voice response technology exceeds a preset load; and the acquiring module is further configured to, upon receiving a voice response technology update response from the called user terminal, acquire identification information of the candidate voice response technology in the voice response technology update response. The voice interaction module is further configured to invoke the candidate voice response technology to conduct multiple rounds of voice interaction with the calling user terminal.

[0184] Optionally, the push module 904 is specifically configured to push the voice interaction text to the called user terminal in a preset form if it is detected that the call stop condition is met; the preset form is any one or more of the following forms: SMS form, instant messaging form, message form, and mailbox form.

[0185] The voice response method provided in this embodiment can implement the voice response method provided in any one of the above-mentioned embodiments 2 to 7. The specific implementation methods and principles are similar and will not be described in detail.

[0186] Embodiment 9

[0187] Figure 10 A block diagram of an electronic device provided by an embodiment of the present invention, such as Figure 10 As shown, the electronic device 1000 provided in this embodiment includes: at least one processor 1002; a memory 1001 and a transceiver 1003;

[0188] Among them, the memory 1001, the processor 1002 and the transceiver 1003 are circuit-connected.

[0189] The memory 1001 stores instructions that can be executed by at least one processor 1002, and the transceiver 1003 is used to send and receive data with a user terminal.

[0190] The instructions are executed by at least one processor to enable the at least one processor to execute the voice response method provided by any one of the embodiments.

[0191] The relevant instructions can be understood by referring to the relevant descriptions and effects corresponding to the steps of the voice response method provided in any embodiment, and will not be elaborated here.

[0192] An embodiment of the present invention further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the voice response method provided in any one of the embodiments.

[0193] An embodiment of the present invention further provides a computer program product, including a computer program, which is executed by a processor to implement the voice response method provided in any one of the above embodiments.

[0194] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.

[0195] It should be further noted that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0196] It should be understood that the above-described device embodiments are merely illustrative, and the device of the present application may also be implemented in other ways. For example, the division of units / modules in the above-described embodiments is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0197] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present application may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.

[0198] If the integrated unit / module is implemented in hardware, the hardware may be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the artificial intelligence processor may be any appropriate hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0199] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0200] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.< / nlresult>

Claims

1. A voice response method, characterized in that: include: receiving a call forwarding request from a calling user terminal, wherein the call forwarding request includes identification information of the called user; determining, based on the identification information of the called user, a target voice response technology preselected by the called user; wherein the target voice response technology includes a target natural language processing technology, a target voice recognition technology, and a target voice synthesis technology, and the target natural language processing technology, the target voice recognition technology, and the target voice synthesis technology are provided by different manufacturers; Invoking the target voice response technology to perform multiple rounds of voice interaction with the calling user terminal and generating voice interaction text; Pushing the voice interaction text to the called user terminal; If it is determined that the current load of the target voice response technology exceeds the preset load, a voice response technology update request is sent to the called user terminal; If a voice response technology update response is received from the called user terminal, identification information of the candidate voice response technology in the voice response technology update response is obtained, and the candidate voice response technology is called to perform multiple rounds of voice interaction with the calling user terminal.

2. The method according to claim 1, characterized in that Before determining the voice response technology pre-selected by the called user according to the identification information of the called user, the method further includes: receiving a voice response technology configuration request sent by a called user terminal; Acquire multiple types of voice response technology details information according to the voice response technology configuration request and send the information to the called user terminal; the voice response technology details information includes identification information of the voice response technology; Receive the identification information of the target voice response technology selected by the called user sent by the called user terminal, and send the identification information of the called user and the identification information of the target voice response technology to the service platform, so that the service platform associates and stores the identification information of the called user and the identification information of the target voice response technology.

3. The method according to claim 2, characterized in that The determining, based on the identification information of the called user, a target voice response technology preselected by the called user includes: Sending a voice response technology determination request to the service platform, wherein the voice response technology determination request includes identification information of the called user, so that the service platform determines identification information of the associated stored voice response technology according to the identification information of the called user; Receiving identification information of the associated stored voice response technology sent by the service platform; The target voice response technology is determined according to the identification information of the associated stored voice response technology.

4. The method according to claim 1, wherein The target voice response technology includes: target voice recognition technology and target natural language processing technology; The calling of the target voice response technology to perform multiple rounds of voice interaction with the calling user terminal includes: The following operations are performed when calling the target voice response technology to perform each round of voice interaction with the calling user terminal: Receiving calling user voice data sent by the calling user terminal; Invoking the target speech recognition technology to perform speech recognition on the calling user voice data to obtain calling user text data; Invoking the target natural language processing technology to perform semantic recognition on the calling user text data, and determining matching response data based on the semantic recognition result; If it is determined that the response data is voice type data, a playback instruction is sent to the universal access platform UAP, where the playback instruction includes the response data, so that the UAP plays the response data to the calling user terminal.

5. The method according to claim 4, characterized in that The target speech response technology also includes a target speech synthesis technology; If it is determined that the response data is text type data, before sending the play instruction to the UAP, the method further includes: The target speech synthesis technology is called to perform speech synthesis processing on the response data to generate speech type response data.

6. The method according to claim 4, characterized in that Before calling the target speech recognition technology to perform speech recognition on the calling user's speech data to obtain the calling user's text data, the method further includes: Sending an engine start request corresponding to the target speech recognition technology to the UAP, the engine start request including identification information of the target speech recognition technology, so that the UAP loads a target grammar file corresponding to the identification information of the target speech recognition technology according to the engine start request, and controls the start of the target speech recognition engine according to the target grammar file; The calling of the target speech recognition technology to perform speech recognition on the calling user's speech data to obtain the calling user's text data includes: The target speech recognition engine is called to perform speech recognition on the calling user voice data to obtain calling user text data.

7. The method according to claim 6, characterized in that After calling the target speech recognition engine to perform speech recognition on the calling user's speech data to obtain the calling user's text data, the method further includes: Controlling the target speech recognition engine to use the improved MRCP protocol to feed back speech recognition results; the speech recognition results include identification information of the target speech recognition technology and speech recognition status information; If it is determined according to the voice recognition situation information that the target voice recognition engine replacement condition is met, a candidate voice recognition engine is obtained, and the candidate voice recognition engine is called to perform voice recognition on the calling user voice data.

8. The method according to any one of claims 1 to 7, characterized in that The step of pushing the voice interaction text to the called user terminal includes: If the call stop condition is detected, the voice interaction text is pushed to the called user terminal in a preset form; The preset form is any one or more of the following forms: SMS form, instant messaging form, message form, email form.

9. A voice response device, characterized in that: include: A receiving module, configured to receive a call forwarding request from a calling user terminal, wherein the call forwarding request includes identification information of the called user; a determination module, configured to determine a target voice response technology preselected by the called user based on the identification information of the called user; wherein the target voice response technology includes a target natural language processing technology, a target voice recognition technology, and a target voice synthesis technology, and the target natural language processing technology, the target voice recognition technology, and the target voice synthesis technology are provided by different manufacturers; A voice interaction module, configured to invoke the target voice response technology to perform multiple rounds of voice interaction with the calling user terminal and generate voice interaction text; A push module, configured to push the voice interaction text to the called user terminal; a sending module configured to send a voice response technology update request to the called user terminal if it is determined that the current load of the target voice response technology exceeds a preset load; an acquisition module, configured to acquire identification information of candidate voice response technologies in a voice response technology update response sent by a called user terminal upon receiving the voice response technology update response; The voice interaction module is further configured to call the candidate voice response technology to perform multiple rounds of voice interaction with the calling user terminal.

10. An electronic device, characterized in that: include: at least one processor; Memory and transceivers; The memory, the processor and the transceiver circuit are interconnected; The memory stores instructions executable by the at least one processor, and the transceiver is configured to transmit and receive data with a user terminal; The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Incoming call intelligent response method and system based on intelligent terminal

    CN105592196A

  • Intelligent telephone answering method and system

    CN111294471A