Voice interaction method and voice interaction system

By integrating speech recognition, natural language processing, and speech synthesis services into the SIP communication system, the problem of lack of unified scheduling in existing intelligent voice systems is solved, achieving an efficient, stable voice interaction experience and rapid response.

CN121814745APending Publication Date: 2026-04-07北京云迹科技股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing intelligent voice systems lack unified scheduling and resource integration, resulting in unstable voice interaction experiences and low efficiency.

Method used

By integrating the SIP communication system with intelligent voice processing services, a high degree of integration between voice communication and intelligent voice processing is achieved. A three-stage voice interaction mechanism is adopted, including speech recognition (ASR), natural language processing (NLP), and text-to-speech (TTS) which are centrally scheduled and processed in the SIP communication system.

Benefits of technology

It significantly improves the efficiency and response speed of voice interaction, reduces the complexity of system deployment, has good controllability and maintainability, supports multiple language selection and exception handling, and ensures the continuity and stability of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121814745A_ABST
    Figure CN121814745A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction method and a voice interaction system. The system is constructed and integrated with at least one intelligent voice processing module based on an SIP communication protocol. The method comprises the following steps: an access layer accesses a voice interaction request initiated by an SIP (Session Initiation Protocol) communication client, and transmits the voice interaction request to a corresponding target service instance in a service layer; the target SIP service in the target service instance receives the voice interaction request, establishes a session channel with the SIP communication client, and controls a data stream in the voice interaction request to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy; each target intelligent voice processing service receives the data stream sent by the target SIP service, carries out intelligent voice processing, and returns the processed data stream to the target SIP service; and the target SIP service receives feedback voice obtained after circulation is completed according to the preset strategy, and returns the feedback voice to the SIP communication client through the session channel. Therefore, high integration of voice communication and intelligent voice processing can be realized, and the system deployment complexity is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a voice interaction method and a voice interaction system. Background Technology

[0002] With the rapid development of intelligent voice processing technology, voice-driven human-computer interaction has been widely used in various scenarios such as intelligent customer service, voice assistants, remote office, and smart healthcare. However, most existing intelligent voice systems adopt a decentralized deployment approach for intelligent voice processing modules, lacking unified scheduling and resource integration of each link, making it difficult to achieve an efficient and stable voice interaction experience. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a voice interaction method and a voice interaction system. By integrating the SIP communication system with intelligent voice processing services, a high degree of integration between voice communication and intelligent voice processing is achieved, effectively reducing the complexity of system deployment, improving the efficiency and response speed of voice interaction, and possessing good controllability, maintainability and intelligence.

[0004] This application provides a voice interaction method applied to a voice interaction system. The voice interaction system is built based on the SIP communication protocol and integrates at least one intelligent voice processing module. The voice interaction system includes an access layer and a service layer. At least one service instance is deployed in the service layer. Each service instance includes a SIP service and at least one intelligent voice processing service. The voice interaction method includes: The access layer receives voice interaction requests initiated by the SIP communication client and transmits the voice interaction requests to the corresponding target service instance in the service layer. The target SIP service in the target service instance receives the voice interaction request, establishes a session channel with the SIP communication client, and controls the data flow in the voice interaction request to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy. Each target intelligent voice processing service receives the data stream sent by the target SIP service, performs intelligent voice processing on the data stream, and returns the processed data stream to the target SIP service; The target SIP service receives the feedback voice obtained after the process is completed according to the preset strategy, and returns the feedback voice to the SIP communication client through the session channel.

[0005] Furthermore, the intelligent voice processing service includes: a speech recognition service, a dialogue analysis service, and a speech synthesis service; the target SIP service controls the data flow in the voice interaction request to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy, including: The target SIP service sends the voice stream in the voice interaction request to the speech recognition service and receives the recognized text returned by the speech recognition service; The target SIP service sends the identified text to the dialogue analysis service and receives the text analysis results returned by the dialogue analysis service; The target SIP service sends the text analysis results to the speech synthesis service and receives the feedback speech returned by the speech synthesis service.

[0006] Furthermore, after receiving the text analysis results returned by the dialogue analysis service, the method further includes: If the text analysis results indicate that at least one of the following is true: the user's intent cannot be recognized, a preset keyword is detected, or an abnormal situation is detected, the voice interaction request will be transferred to a human agent.

[0007] Furthermore, the voice interaction system also includes a management layer; the management layer is pre-configured with SIP account groups and configuration parameters corresponding to each customer with voice interaction permissions; the voice interaction method further includes: The target SIP service parses the SIP account included in the voice interaction request and sends the SIP account to the management layer; The management layer queries the target configuration parameters of the target customer corresponding to the SIP account and returns the target configuration parameters to the target SIP service. The target SIP service processes the voice interaction request based on the target configuration parameters.

[0008] Furthermore, the target SIP service processes the voice interaction request based on the target configuration parameters, including: The target SIP service determines whether the target customer has configured multi-language selection based on the target configuration parameters. If configured, a multilingual selection guide voice is returned to the SIP communication client through the session channel; The system receives language parameters returned by the SIP communication client and controls the processing of the data stream by the at least one target intelligent voice processing service according to the language parameters. Furthermore, the target SIP service processes the voice interaction request based on the target configuration parameters, and further includes: The target SIP service determines whether the target customer has enabled the intelligent voice processing service based on the target configuration parameters. If enabled, the data flow in the voice interaction request is controlled to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy; If not enabled, the voice interaction request will be transferred to a human agent according to the transfer method in the target configuration parameters.

[0009] Furthermore, based on the SIP communication protocol, the access layer supports various voice communication resources, including: supporting traditional telephone signals to access via SIP trunk through the voice gateway of the IAD device, supporting SIP phone access, supporting VoIP resources to access via SIP TRUNK, and supporting EI gateway device access.

[0010] This application also provides a voice interaction system, which is built on the SIP communication protocol and integrates at least one intelligent voice processing module; the voice interaction system includes an access layer and a service layer; at least one service instance is deployed in the service layer; each service instance includes a SIP service and at least one intelligent voice processing service; The access layer is used to access voice interaction requests initiated by SIP communication clients and transmit the voice interaction requests to the corresponding target service instance in the service layer. The target SIP service in the target service instance is used to receive the voice interaction request, establish a session channel with the SIP communication client, and control the data flow in the voice interaction request to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy. Each target intelligent voice processing service is used to receive the data stream sent by the target SIP service, perform intelligent voice processing on the data stream, and return the processed data stream to the target SIP service; The target SIP service is used to receive feedback voice after the process is completed according to the preset strategy, and to return the feedback voice to the SIP communication client through the session channel.

[0011] This application also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the voice interaction method described above are performed.

[0012] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the voice interaction method described above.

[0013] This application provides a voice interaction method and system that introduces a phased intelligent voice processing service into the SIP communication architecture. It integrates multiple functional modules required for traditional intelligent voice interaction into the communication system scheduling process, achieving a high degree of integration between voice communication and intelligent processing. This breaks the structural limitations of dispersed deployment and loose coupling of voice processing modules, significantly improving the system's integration and collaborative efficiency. It also avoids the burden of redundant development, interface adaptation, and operation and maintenance caused by multiple systems being connected in series, effectively reducing the overall system complexity and deployment cost.

[0014] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart of a voice interaction method provided in an embodiment of this application is shown; Figure 2 This paper shows a schematic diagram of the structure of a voice interaction system provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.

[0018] Research has shown that with the rapid development of intelligent voice processing technology, voice-driven human-computer interaction has been widely used in various scenarios such as intelligent customer service, voice assistants, remote work, and smart healthcare. However, most existing intelligent voice systems adopt a decentralized deployment approach for intelligent voice processing modules, lacking unified scheduling and resource integration across different stages, making it difficult to achieve an efficient and stable voice interaction experience.

[0019] Based on this, the embodiments of this application provide a voice interaction method to achieve a high degree of integration between voice communication and intelligent voice processing, effectively reduce system deployment complexity, improve voice interaction efficiency and response speed, and have good controllability, maintainability and intelligence level, realizing an end-to-end, highly integrated intelligent voice interaction process.

[0020] Please see Figure 1 , Figure 1 A flowchart illustrating a voice interaction method provided in an embodiment of this application. Please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram of the structure of a voice interaction system provided in an embodiment of this application. The voice interaction system is built on the SIP communication protocol and integrates at least one intelligent voice processing module. The voice interaction system includes an access layer and a service layer; the service layer deploys at least one service instance; each service instance includes a SIP service and at least one intelligent voice processing service (MRCP service). In specific implementations, the voice interaction system can be implemented based on the Freeswitch softswitch platform.

[0021] like Figure 1 As shown in the embodiments of this application, the voice interaction method includes: S101. The access layer receives the voice interaction request initiated by the SIP communication client and transmits the voice interaction request to the corresponding target service instance in the service layer.

[0022] Here, based on the SIP communication protocol, the access layer supports a variety of voice communication resources, including: supporting traditional telephone signals to access via SIP trunk through voice gateways of IAD devices (such as FXO, FXS), supporting SIP phone access, supporting VoIP resources (such as IP phones, cloud communication platforms, etc.) to access via SIP trunk, and supporting EI gateway device access.

[0023] S102. The target SIP service in the target service instance receives the voice interaction request, establishes a session channel with the SIP communication client, and controls the data flow in the voice interaction request to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy.

[0024] Here, the SIP server acts as the unified scheduling and control center of the voice interaction system, responsible for: session establishment and release; audio stream forwarding and transmission; calling logic of intelligent voice processing services at each stage; and fault tolerance mechanisms such as state transition control and timeout retry.

[0025] It is worth noting that the voice interaction system in this embodiment may not have a load balancing layer configured as a unified entry point for the gateway layer, in order to further reduce the complexity of the system. In this case, the corresponding target service instance is determined based on the IP address accessed in the voice interaction request. Simultaneously, the accessed SIP communication client can set the server addresses of the primary and backup services, and automatically switch to the backup service when the primary service request fails.

[0026] S103. Each target intelligent voice processing service receives the data stream sent by the target SIP service, performs intelligent voice processing on the data stream, and returns the processed data stream to the target SIP service.

[0027] In one possible implementation, the intelligent voice processing service includes: a voice recognition service, a dialogue analysis service, and a voice synthesis service. Therefore, the three-stage voice processing mechanism, divided into phases, includes: Phase 1 (ASR): After receiving the user's voice, the target SIP service sends the voice stream in the voice interaction request to the ASR service via the MRCP (Media Resource Control Protocol). The ASR service performs real-time speech recognition and converts it into corresponding text information. The target SIP service receives the recognized text returned by the speech recognition service.

[0028] The second stage (NLP dialogue analysis): The target SIP service sends the identified text to the locally deployed NLP dialogue analysis service. The NLP performs intent recognition, context understanding, and dialogue state management on the identified text output by ASR to generate text analysis results of structured response content or action instructions; the target SIP service receives the text analysis results returned by the dialogue analysis service.

[0029] The third stage (TTS speech synthesis): The text analysis results processed by NLP are converted into speech responses by the TTS module, generating feedback speech files or audio streams; the target SIP service receives the feedback speech returned by the speech synthesis service and returns it to the user through the SIP relay channel, completing one full interaction round.

[0030] Alternatively, some of the intelligent voice processing modules (such as NLP services related to text processing) can be deployed outside the system. The SIP service can call these intelligent voice processing modules through interfaces, enabling the intelligent voice processing modules to be shared by multiple systems.

[0031] Furthermore, after receiving the text analysis results returned by the dialogue analysis service, the method further includes: If the text analysis results indicate that at least one of the following is true: the user's intent cannot be recognized, a preset keyword is detected, or an abnormal situation is detected, the voice interaction request will be transferred to a human agent.

[0032] When the NLP module fails to accurately identify the user's intent, or detects special keywords / abnormal situations, the system can automatically transfer the current call to a human customer service agent according to preset rules, ensuring the continuity of interaction and the closed loop of service.

[0033] Furthermore, the voice interaction system also includes a management layer; the management layer is pre-configured with SIP account groups and configuration parameters corresponding to each customer with voice interaction permissions.

[0034] Because the system architecture supports the integration of multiple voice platforms, third-party AI services, standard voice gateways, and TTS engines, it boasts excellent scalability and platform compatibility. Therefore, it can flexibly adapt to the voice service needs of enterprises in different industries, such as banking, government, healthcare, telecommunications, and hotels.

[0035] Furthermore, taking hotels as an example, the configuration parameters include information management for each hotel, SIP account management, call transfer method configuration, historical call detail records, data statistics, and an AI master switch for controlling the status of intelligent voice processing services. In addition, the voice interaction system also includes a data layer, which contains a database shared by all layers within the system.

[0036] The voice interaction method also includes: S201. The target SIP service parses the SIP account included in the voice interaction request and sends the SIP account to the management layer. S202. The management layer queries the target configuration parameters of the target customer corresponding to the SIP account and returns the target configuration parameters to the target SIP service. S203. The target SIP service processes the voice interaction request based on the target configuration parameters.

[0037] In one possible implementation, step S203 may include: The target SIP service determines whether the target customer has configured multi-language selection based on the target configuration parameters; if configured, it returns multi-language selection guidance voice to the SIP communication client through the session channel; it receives the language parameters returned by the SIP communication client and controls the processing of the data stream by the at least one target intelligent voice processing service according to the language parameters. The language parameters can include language, speaker style, dialect, and speech rate. SIP communication can configure the language of the ASR service and the speaker for the TTS service based on these language parameters.

[0038] In another possible implementation, step S203 may include: The target SIP service determines whether the target customer has enabled the intelligent voice processing service based on the target configuration parameters. If enabled, the data flow in the voice interaction request is controlled to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy. If not enabled, the voice interaction request is transferred to a human agent according to the transfer method in the target configuration parameters.

[0039] In this way, when the intelligent voice service is not enabled, the call can be automatically transferred to a human agent for further processing, ensuring the continuity and reliability of the interactive service, which is suitable for enterprise-level intelligent customer service systems.

[0040] S104. The target SIP service receives the feedback voice obtained after the process is completed according to the preset strategy, and returns the feedback voice to the SIP communication client through the session channel.

[0041] The voice interaction method provided in this application has the following beneficial effects: 1. A three-stage voice interaction mechanism based on a SIP server.

[0042] The three-stage voice interaction processing flow includes an orderly linkage and control mechanism for three stages: Automatic Speech Recognition (ASR), Natural Language Processing (NLP), and Text-to-Speech (TTS). By highly integrating traditional speech recognition, semantic analysis, and speech synthesis modules into a unified SIP communication system, it significantly improves the fluency of interaction and system processing efficiency. A clear logical link is established between ASR, NLP, and TTS, enabling independent scheduling, flexible configuration, and scalable replacement of processing tasks at each stage. This modular design allows the voice interaction flow to be customized and optimized according to actual needs, providing efficient adaptability to different business scenarios.

[0043] 2. Deep integration of SIP communication system and intelligent voice processing module.

[0044] The deep integration of the SIP communication protocol with intelligent voice processing modules (ASR, NLP, TTS), especially the centralized scheduling and processing of voice signals at the SIP server level, effectively solves the problem of the separation between voice communication and voice processing modules in traditional systems, and improves the system's collaborative efficiency, response speed and scalability.

[0045] 3. Intelligent transfer and exception handling mechanism.

[0046] The technical solution of this application also includes intelligent transfer and exception handling mechanisms. When the NLP module cannot accurately identify the user's intent or intelligent voice processing is not enabled, the system can automatically transfer the voice call to a human customer service representative. This enables seamless switching between the intelligent voice processing system and human service, ensuring the continuity of interactive services and the stability of the user experience.

[0047] 4. An open architecture based on the Freeswitch platform.

[0048] The communication framework is built using the open-source SIP server platform Freeswitch, and integrated with external MRCP, NLP, and TTS services, employing open protocols and interfaces (such as MRCP and RESTful APIs) for modular integration. This architecture provides excellent flexibility and scalability, enabling the system to easily connect to different voice platforms, AI services, and hardware devices, supporting future technology iterations and various application scenarios.

[0049] 5. Support and compatibility with multiple access methods.

[0050] The system supports connection to external voice gateways through various access methods such as FXO, FXS devices, SIP trunks, and SIP TRUNKs, ensuring stable operation in both traditional circuit-switched (PSTN) and modern IP communication environments, demonstrating its broad applicability in terms of communication interface adaptability and compatibility.

[0051] Based on the same inventive concept, this application also provides a voice interaction system, which is built on the SIP communication protocol and integrates at least one intelligent voice processing module; the voice interaction system includes an access layer and a service layer; at least one service instance is deployed in the service layer; each service instance includes a SIP service and at least one intelligent voice processing service; The access layer is used to access voice interaction requests initiated by SIP communication clients and transmit the voice interaction requests to the corresponding target service instance in the service layer. The target SIP service in the target service instance is used to receive the voice interaction request, establish a session channel with the SIP communication client, and control the data flow in the voice interaction request to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy. Each target intelligent voice processing service is used to receive the data stream sent by the target SIP service, perform intelligent voice processing on the data stream, and return the processed data stream to the target SIP service; The target SIP service is used to receive feedback voice after the process is completed according to the preset strategy, and to return the feedback voice to the SIP communication client through the session channel.

[0052] Furthermore, the intelligent voice processing service includes: a speech recognition service, a dialogue analysis service, and a speech synthesis service; when the target SIP service controls the data flow in the voice interaction request to flow through at least one target intelligent voice processing service in the target service instance according to a preset strategy, the target SIP service is used to: The voice stream in the voice interaction request is sent to the voice recognition service, and the recognized text returned by the voice recognition service is received. The identified text is sent to the dialogue analysis service, and the text analysis results returned by the dialogue analysis service are received. The text analysis results are sent to the speech synthesis service, and the feedback speech returned by the speech synthesis service is received.

[0053] Furthermore, after receiving the text analysis results returned by the dialogue analysis service, the target SIP service is also used to: If the text analysis results indicate that at least one of the following is true: the user's intent cannot be recognized, a preset keyword is detected, or an abnormal situation is detected, the voice interaction request will be transferred to a human agent.

[0054] Furthermore, the voice interaction system also includes a management layer; the management layer is pre-configured with SIP account groups and configuration parameters corresponding to each customer with voice interaction permissions; The target SIP service is also used to parse the SIP account included in the voice interaction request and send the SIP account to the management layer; The management layer is used to query the target configuration parameters of the target customer corresponding to the SIP account, and return the target configuration parameters to the target SIP service. The target SIP service is also used to process the voice interaction request based on the target configuration parameters.

[0055] Furthermore, when the target SIP service processes the voice interaction request based on the target configuration parameters, the target SIP service is used to: Based on the target configuration parameters, determine whether the target customer has configured multi-language selection; If configured, a multilingual selection guide voice is returned to the SIP communication client through the session channel; The system receives language parameters returned by the SIP communication client and controls the processing of the data stream by the at least one target intelligent voice processing service according to the language parameters.

[0056] Furthermore, when the target SIP service processes the voice interaction request based on the target configuration parameters, the target SIP service is also used to: Based on the target configuration parameters, determine whether the target customer has enabled the intelligent voice processing service; If enabled, the data flow in the voice interaction request is controlled to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy; If not enabled, the voice interaction request will be transferred to a human agent according to the transfer method in the target configuration parameters.

[0057] Furthermore, based on the SIP communication protocol, the access layer supports various voice communication resources, including: supporting traditional telephone signals to access via SIP trunk through the voice gateway of the IAD device, supporting SIP phone access, supporting VoIP resources to access via SIP TRUNK, and supporting EI gateway device access.

[0058] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 3 As shown, the electronic device 300 includes a processor 310, a memory 320, and a bus 330.

[0059] The memory 320 stores machine-readable instructions executable by the processor 310. When the electronic device 300 is running, the processor 310 and the memory 320 communicate via the bus 330. When the machine-readable instructions are executed by the processor 310, they can perform the operations described above. Figure 1 The steps of the voice interaction method in the illustrated method embodiment can be found in the method embodiment for specific implementation, and will not be repeated here.

[0060] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 1 The steps of the voice interaction method in the illustrated method embodiment can be found in the method embodiment for specific implementation, and will not be repeated here.

[0061] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0062] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0063] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0064] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0065] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0066] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice interaction method, characterized in that, The method is applied to a voice interaction system; the voice interaction system is built on the SIP communication protocol and integrates at least one intelligent voice processing module; the voice interaction system includes an access layer and a service layer; at least one service instance is deployed in the service layer; each service instance includes a SIP service and at least one intelligent voice processing service. The voice interaction method includes: The access layer receives voice interaction requests initiated by the SIP communication client and transmits the voice interaction requests to the corresponding target service instance in the service layer. The target SIP service in the target service instance receives the voice interaction request, establishes a session channel with the SIP communication client, and controls the data flow in the voice interaction request to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy. Each target intelligent voice processing service receives the data stream sent by the target SIP service, performs intelligent voice processing on the data stream, and returns the processed data stream to the target SIP service; The target SIP service receives the feedback voice obtained after the process is completed according to the preset strategy, and returns the feedback voice to the SIP communication client through the session channel.

2. The method according to claim 1, characterized in that, The intelligent voice processing service includes: voice recognition service, dialogue analysis service, and voice synthesis service; the target SIP service controls the data flow in the voice interaction request to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy, including: The target SIP service sends the voice stream in the voice interaction request to the speech recognition service and receives the recognized text returned by the speech recognition service; The target SIP service sends the identified text to the dialogue analysis service and receives the text analysis results returned by the dialogue analysis service; The target SIP service sends the text analysis results to the speech synthesis service and receives the feedback speech returned by the speech synthesis service.

3. The method according to claim 2, characterized in that, After receiving the text analysis results returned by the dialogue analysis service, the method further includes: If the text analysis results indicate that at least one of the following is true: the user's intent cannot be recognized, a preset keyword is detected, or an abnormal situation is detected, the voice interaction request will be transferred to a human agent.

4. The method according to claim 1, characterized in that, The voice interaction system also includes a management layer; the management layer is pre-configured with SIP account groups and configuration parameters corresponding to each customer with voice interaction permissions; The voice interaction method further includes: The target SIP service parses the SIP account included in the voice interaction request and sends the SIP account to the management layer; The management layer queries the target configuration parameters of the target customer corresponding to the SIP account and returns the target configuration parameters to the target SIP service. The target SIP service processes the voice interaction request based on the target configuration parameters.

5. The method according to claim 4, characterized in that, The target SIP service processes the voice interaction request based on the target configuration parameters, including: The target SIP service determines whether the target customer has configured multi-language selection based on the target configuration parameters. If configured, a multilingual selection guide voice is returned to the SIP communication client through the session channel; The system receives language parameters returned by the SIP communication client and controls the processing of the data stream by the at least one target intelligent voice processing service according to the language parameters.

6. The method according to claim 4, characterized in that, The target SIP service processes the voice interaction request based on the target configuration parameters, and further includes: The target SIP service determines whether the target customer has enabled the intelligent voice processing service based on the target configuration parameters. If enabled, the data flow in the voice interaction request is controlled to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy; If not enabled, the voice interaction request will be transferred to a human agent according to the transfer method in the target configuration parameters.

7. The method according to claim 1, characterized in that, Based on the SIP communication protocol, the access layer supports various voice communication resources, including: supporting traditional telephone signals to access via SIP trunk through the voice gateway of the IAD device, supporting SIP phone access, supporting VoIP resources to access via SIP TRUNK, and supporting EI gateway device access.

8. A voice interaction system, characterized in that, The voice interaction system is built on the SIP communication protocol and integrates at least one intelligent voice processing module; the voice interaction system includes an access layer and a service layer; at least one service instance is deployed in the service layer; each service instance includes a SIP service and at least one intelligent voice processing service. The access layer is used to access voice interaction requests initiated by SIP communication clients and transmit the voice interaction requests to the corresponding target service instance in the service layer. The target SIP service in the target service instance is used to receive the voice interaction request, establish a session channel with the SIP communication client, and control the data flow in the voice interaction request to flow in at least one target intelligent voice processing service in the target service instance according to a preset strategy. Each target intelligent voice processing service is used to receive the data stream sent by the target SIP service, perform intelligent voice processing on the data stream, and return the processed data stream to the target SIP service; The target SIP service is used to receive feedback voice after the process is completed according to the preset strategy, and to return the feedback voice to the SIP communication client through the session channel.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the machine-readable instructions are executed by the processor to perform the steps of the voice interaction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the voice interaction method as described in any one of claims 1 to 7.