Service Status Monitoring Method, Device, Equipment and Storage Medium
By analyzing incremental logs and counter monitoring methods, the problem of inability to timely detect abnormal communication between media resource servers and intelligent voice devices in the prior art is solved, and timely monitoring and early warning of service status is achieved, and the stability of the system and customer experience are improved.
Patent Information
- Application Number
- CN202211203405.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-09-29
AI Technical Summary
The prior art cannot timely monitor the service status of media resource servers and smart voice devices, resulting in the inability to issue warnings in time when communication is abnormal, affecting customer experience and system stability.
By analyzing the incremental log of the interaction process between the media resource server and the intelligent voice device, using the handshake round counter and early warning threshold to perform early warning processing in a timely manner, monitoring the text-to-speech failure counter and automatic voice recognition failure rate, and promptly discovering and warning communication abnormalities.
It realizes timely monitoring of the interaction status of media resource servers and intelligent voice devices, prevents communication abnormalities, and improves system stability and customer experience.
Smart Images

Figure CN115604155B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data analysis, and in particular, to a service status monitoring method, device, equipment, and storage medium. Background Art
[0002] In recent years, customer service robot systems built based on artificial intelligence speech recognition (ASR), text-to-speech (TTS), and natural language processing technologies have gradually emerged. Such intelligent customer service systems have advantages such as automation, intelligence, batch processing, and standardization, greatly liberating human resources and reducing costs for enterprise after-sales, marketing, etc. Building such a system requires four major modules: a media resource server, intelligent voice devices, NLP, and application systems. Among them, the media resource server and intelligent voice products act as the "mouth" and "ears" of the customer service robot. The media resource server receives the speech synthesized by the intelligent voice product and plays it to the customer through subsequent modules to respond to the customer; the media resource server transmits the customer's replied speech and sends it to the intelligent voice product for recognition as text to receive customer feedback. The media resource server and intelligent voice products are at the forefront of customer interaction, and whether they can interact stably and normally affects the customer's intuitive experience. However, there are many hidden dangers in their interaction: the interaction process is very complex, and there are many factors that can be affected; once an interaction anomaly occurs, it is very difficult to detect it in a timely manner.
[0003] In the prior art, if the traditional log sampling monitoring method of a general application system is used to monitor the service status of the media resource server and intelligent voice devices, there are at least the following technical problems: a warning of communication anomalies cannot be issued in a timely manner, making it impossible for administrators and application systems to intervene in a timely manner to prevent the occurrence of communication anomalies. Summary of the Invention
[0004] This application provides a service status monitoring method, device, equipment, and storage medium, which solves the problem that the prior art cannot monitor the service status of the media resource server and intelligent voice devices in a timely manner and issue a warning.
[0005] In a first aspect, this application provides a service status monitoring method, which is applied to a monitoring device and includes:
[0006] Obtain the incremental log of the interaction process between the media resource server and the intelligent voice device;
[0007] Initialize the preset handshake round counter to 0, where the handshake round counter is used to record the handshake rounds in the interaction process between the media resource server and the intelligent voice device;
[0008] Parse each incremental log. If a handshake request message is parsed, increment the handshake round counter by 1. If a handshake request message is not parsed, set the handshake round counter to 0;
[0009] If the count of the handshake round counter is greater than or equal to the warning threshold, enter the warning process, and enter the sleep for a preset duration after the warning process is completed;
[0010] After completing the sleep for the preset duration, re - execute the step of parsing each incremental log.
[0011] In a possible design, the service status monitoring method further includes:
[0012] If the count of the handshake round counter is less than the warning threshold, continue to determine whether the count of the handshake round counter is greater than or equal to the alarm threshold, where the alarm threshold is greater than the warning threshold;
[0013] If the count of the handshake round counter is greater than or equal to the alarm threshold, enter the emergency process, and enter the sleep for a preset duration after the emergency process is completed; if the count of the handshake round counter is less than the alarm threshold, enter the sleep for a preset duration;
[0014] After completing the sleep for the preset duration, re - execute the step of parsing each incremental log.
[0015] In a possible design, after obtaining the incremental log of the interaction process between the media resource server and the intelligent voice device, it further includes:
[0016] Initialize the preset text - to - speech failure counter to 0, where the text - to - speech failure counter is used to record the number of failures of text - to - speech requests during the interaction process between the media resource server and the intelligent voice device;
[0017] Parse each incremental log to determine the request message type of the media resource control protocol;
[0018] If the request message type of the media resource control protocol is a text - to - speech request, determine whether there is a channel identifier in the cache for the text - to - speech request;
[0019] If there is a channel identifier in the cache for the text - to - speech request, increment the text - to - speech failure counter by 1; if there is no channel identifier in the cache for the text - to - speech request, cache the channel identifier;
[0020] If the return message of the channel identifier is parsed in each subsequent incremental log, delete the channel identifier from the cache and enter the sleep for a preset duration;
[0021] After completing the sleep for the preset duration, re - execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0022] After obtaining the incremental log of the interaction process between the media resource server and the intelligent voice device, the following steps are further included:
[0023] If the return message of the channel identifier is not parsed in each incremental log, increment the text-to-speech failure counter by 1, and continue to determine whether the text-to-speech failure count is greater than the warning threshold;
[0024] If the text-to-speech failure count is greater than the warning threshold, enter the warning process, and enter a sleep state for a preset duration after the warning process is completed;
[0025] If the text-to-speech failure count is less than or equal to the warning threshold, continue to determine whether the text-to-speech failure count is greater than the alarm threshold, where the alarm threshold is greater than the warning threshold;
[0026] If the text-to-speech failure count is greater than the alarm threshold, enter the emergency process, and enter a sleep state for a preset duration after the emergency process is completed; if the text-to-speech failure count is less than or equal to the alarm threshold, enter a sleep state for a preset duration;
[0027] After completing the sleep state for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0028] In a possible design, after parsing each incremental log to determine the request message type of the media resource control protocol, the following steps are further included:
[0029] Initialize both the preset automatic speech recognition request counter and the automatic speech recognition failure counter to 0. The automatic speech recognition request counter is used to record the number of automatic speech recognition requests sent during the interaction between the media resource server and the intelligent voice device, and the automatic speech recognition failure counter is used to record the number of failures of automatic speech recognition requests during the interaction between the media resource server and the intelligent voice device;
[0030] If the request message type of the media resource control protocol is an automatic speech recognition request, increment the automatic speech recognition request counter by 1; if the channel identifier exists in the cache for the automatic speech recognition request, increment the automatic speech recognition failure counter by 1; if the channel identifier does not exist in the cache for the automatic speech recognition request, cache the channel identifier;
[0031] If the return message of the channel identifier is parsed in each incremental log, set the automatic speech recognition request counter and the automatic speech recognition failure counter to 0, delete the channel identifier in the cache, and enter a sleep state for a preset duration;
[0032] After completing the sleep for the preset duration, the step of parsing each incremental log again to determine the request message type of the media resource control protocol is performed again.
[0033] In a possible design, after parsing each incremental log to determine the request message type of the media resource control protocol, the following is further included:
[0034] If the return message of the channel identifier is not parsed in each incremental log, the automatic speech recognition failure counter is incremented by 1, and it continues to determine whether the automatic speech recognition failure rate reaches a preset threshold, where the automatic speech recognition failure rate is the ratio of the number of automatic speech recognition failures to the number of automatic speech recognition requests;
[0035] If the automatic speech recognition failure rate reaches the preset threshold, emergency processing is entered, and after the emergency processing is completed, sleep for the preset duration is entered; if the automatic speech recognition failure rate does not reach the preset threshold, sleep for the preset duration is entered;
[0036] After completing the sleep for the preset duration, the step of parsing each incremental log again to determine the request message type of the media resource control protocol is performed again.
[0037] In a possible design, the service status monitoring method further includes:
[0038] Obtain the incremental log and incremental voice file during the interaction between the media resource server and the intelligent voice device;
[0039] Initialize the preset text-to-speech request counter to 0, where the text-to-speech request count is used to record the number of text-to-speech requests during the interaction between the media resource server and the intelligent voice device;
[0040] Parse each incremental log to determine the request message type of the media resource control protocol;
[0041] If the request message type of the media resource control protocol is a text-to-speech request, continue to determine whether the channel identifier exists in the cache for the text-to-speech request;
[0042] If the channel identifier exists in the cache for the text-to-speech request, increment the text-to-speech failure counter by 1;
[0043] If the channel identifier does not exist in the cache for the text-to-speech request, configure the corresponding relationship between the synthesized text for text-to-speech and the size value of the voice packet generated by the synthesized text for text-to-speech;
[0044] Parse the text for speech synthesis in this text-to-speech request, query the size value of the corresponding speech package, and save the mapping relationship between the channel identifier of this text-to-speech request and the size value of the speech package;
[0045] If the return message of the channel identifier is parsed in each incremental log, increment the text-to-speech request counter by 1;
[0046] Parse the incremental speech file, use the size value of the speech package corresponding to the channel identifier to traverse the size of the incremental speech file, and calculate the deviation rate, where the deviation rate is: the absolute value of the difference between the size value of the speech package corresponding to the text for speech synthesis and the size value of the incremental speech file divided by the size value of the speech package corresponding to the text for speech synthesis;
[0047] If there is no incremental speech file with a deviation rate less than a specific threshold, enter emergency processing, and enter a sleep for a preset duration after the emergency processing is completed;
[0048] After completing the sleep for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0049] In a possible design, the service status monitoring method further includes:
[0050] If there is an incremental speech file with a deviation rate less than a specific threshold, continue to determine whether the number of incremental speech files is less than the number of text-to-speech requests. If the number of incremental speech files is less than the number of text-to-speech requests, enter emergency processing, and enter a sleep for a preset duration after the emergency processing is completed; if the number of incremental speech files is greater than or equal to the number of text-to-speech requests, enter a sleep for a preset duration;
[0051] After completing the sleep for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0052] In a possible design, the obtaining the incremental log and the incremental speech file in the interaction process between the media resource server and the intelligent voice device includes:
[0053] Read the incremental log and the incremental speech file in the interaction process between the media resource server and the intelligent voice device from the network storage.
[0054] In a second aspect, the present application provides a service status monitoring device, including:
[0055] An obtaining module, configured to obtain an incremental log in the interaction process between a media resource server and an intelligent voice device;
[0056] A processing module is used to initialize a preset handshake round counter to 0, where the handshake round counter is used to record the number of handshake rounds in the interaction process between the media resource server and the intelligent voice device;
[0057] The processing module is further used to parse each incremental log. If a handshake request message is parsed, the handshake round counter is incremented by 1. If no handshake request message is parsed, the handshake round counter is set to 0;
[0058] The processing module is further used to enter warning processing if the count of the handshake round counter is greater than or equal to a warning threshold, and enter a sleep for a preset duration after the warning processing is completed;
[0059] The processing module is further used to continue to determine whether the count of the handshake round counter is greater than or equal to an alarm threshold if the count of the handshake round counter is less than the warning threshold, where the alarm threshold is greater than the warning threshold;
[0060] The processing module is further used to enter emergency processing if the count of the handshake round counter is greater than or equal to the alarm threshold, and enter a sleep for a preset duration after the emergency processing is completed; if the count of the handshake round counter is less than the alarm threshold, enter a sleep for a preset duration;
[0061] The processing module is further used to re-execute the step of parsing each incremental log after completing the sleep for the preset duration.
[0062] In a possible design, the service status monitoring device further includes:
[0063] A processing module is used to initialize a preset text-to-speech failure counter to 0, where the text-to-speech failure counter is used to record the number of failures of text-to-speech requests in the interaction process between the media resource server and the intelligent voice device;
[0064] The processing module is further used to parse each incremental log to determine the request message type of the media resource control protocol;
[0065] The processing module is further used to determine whether a channel identifier exists in the cache for a text-to-speech request if the request message type of the media resource control protocol is a text-to-speech request;
[0066] The processing module is further used to increment the text-to-speech failure counter by 1 if the channel identifier exists in the cache for the text-to-speech request; if the channel identifier does not exist in the cache for the text-to-speech request, cache the channel identifier;
[0067] The processing module is further configured to, if a return message of the channel identifier is parsed from subsequent incremental logs, delete the channel identifier from the cache and enter a sleep state for a preset duration;
[0068] The processing module is further configured to, if a return message of the channel identifier is not parsed from each incremental log, increment the text-to-speech failure counter by 1 and continue to determine whether the text-to-speech failure count is greater than a warning threshold;
[0069] The processing module is further configured to, if the text-to-speech failure count is greater than the warning threshold, enter warning processing and enter a sleep state for a preset duration after the warning processing is completed;
[0070] The processing module is further configured to, if the text-to-speech failure count is less than or equal to the warning threshold, continue to determine whether the text-to-speech failure count is greater than an alarm threshold, where the alarm threshold is greater than the warning threshold;
[0071] The processing module is further configured to, if the text-to-speech failure count is greater than the alarm threshold, enter emergency processing and enter a sleep state for a preset duration after the emergency processing is completed; if the text-to-speech failure count is less than or equal to the alarm threshold, enter a sleep state for a preset duration;
[0072] The processing module is further configured to, after completing the sleep state for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0073] In a possible design, the service status monitoring device further includes:
[0074] A processing module, configured to initialize both a preset automatic speech recognition request counter and an automatic speech recognition failure counter to 0, where the automatic speech recognition request counter is used to record the number of automatic speech recognition requests sent during the interaction between the media resource server and the intelligent voice device, and the automatic speech recognition failure counter is used to record the number of failures of the automatic speech recognition requests during the interaction between the media resource server and the intelligent voice device;
[0075] The processing module is further configured to, if the request message type of the media resource control protocol is an automatic speech recognition request, increment the automatic speech recognition request counter by 1; if a channel identifier exists in the cache for the automatic speech recognition request, increment the automatic speech recognition failure counter by 1; if a channel identifier does not exist in the cache for the automatic speech recognition request, cache the channel identifier;
[0076] The processing module is further configured to, if a return message of the channel identifier is parsed from each incremental log, set both the automatic speech recognition request counter and the automatic speech recognition failure counter to 0, delete the channel identifier in the cache, and enter a sleep state for a preset duration;
[0077] The processing module is further configured to, if the return message of the channel identifier is not parsed in each incremental log, increment the automatic speech recognition failure counter by 1, and continue to determine whether the automatic speech recognition failure rate reaches a preset threshold, where the automatic speech recognition failure rate is the ratio of the number of automatic speech recognition failures to the number of automatic speech recognition requests;
[0078] The processing module is further configured to, if the automatic speech recognition failure rate reaches the preset threshold, enter emergency processing, and enter a sleep state for a preset duration after the emergency processing is completed; if the automatic speech recognition failure rate does not reach the preset threshold, enter a sleep state for a preset duration;
[0079] The processing module is further configured to, after completing the sleep state for the preset duration, re - execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0080] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0081] The memory stores computer - executable instructions;
[0082] The processor executes the computer - executable instructions stored in the memory to implement the method as described in the first aspect.
[0083] In a fourth aspect, the present application provides a computer - readable storage medium, where computer - executable instructions are stored in the computer - readable storage medium, and when the computer - executable instructions are executed by a processor, they are used to implement the method as described in the first aspect.
[0084] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method as described in the first aspect.
[0085] The service status monitoring method, device, equipment, and storage medium provided by the present application parse the incremental logs in the interaction process between the media resource server and the intelligent voice device to obtain the handshake rounds of the handshake request message, and perform early warning processing in a timely manner based on the handshake rounds, so that the administrator and the application system can intervene in a timely manner to prevent communication anomalies from occurring. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0087] Figure 1 It is an application scenario diagram of the service status monitoring method applicable to the embodiments of the present application;
[0088] Figure 2 The flowchart of the service status monitoring method provided by the embodiment of the present application Figure 1 ;
[0089] Figure 3 The flowchart of the service status monitoring method provided by the embodiment of the present application Figure 2 ;
[0090] Figure 4 The flowchart of the service status monitoring method provided by the embodiment of the present application Figure 3 ;
[0091] Figure 5 Schematic diagrams of normal handshake interaction and abnormal handshake interaction between the media resource server and the intelligent voice device during the handshake phase in the embodiment of the present application;
[0092] Figure 6 Schematic diagrams of normal interaction and abnormal interaction between the media resource server and the intelligent voice device during the MRCP communication phase in the embodiment of the present application;
[0093] Figure 7 Schematic diagram of the structure of the service status monitoring device provided by the embodiment of the present application;
[0094] Figure 8 Schematic diagram of the structure of the electronic device provided by the embodiment of the present application.
[0095] Through the above-mentioned drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0096] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0097] The customer service robot system built based on automatic speech recognition (ASR), text-to-speech (TTS), and natural language processing technology (NLP) has the advantages of automation, intelligence, batch processing, and standardization. Building such a system requires four major modules: a media resource server, intelligent voice devices, NLP, and an application system. Among them, the media resource server and intelligent voice devices act as the "mouth" and "ear" of the customer service robot. The media resource server receives the voice synthesized by the intelligent voice product and plays it to the customer through the subsequent module to respond to the customer; the media resource server transmits the customer's replied voice to the intelligent voice device for recognition as text to receive customer feedback. Thus, it can be seen that the media resource server and intelligent voice devices are at the forefront of customer interaction. Whether they can interact stably and normally affects the customer's intuitive experience. However, there are many hidden dangers in their interaction: First, the interaction process is very complex and there are many factors that can affect it. The interaction between the two is divided into two stages - the handshake stage and the voice interaction stage. In the handshake stage, the Session Initiation Protocol (SIP) and Session Description Protocol (SDP) are used to negotiate the channel ID, UDP communication port, voice interaction rules, etc.; in the voice interaction stage, the Media Resource Control Protocol (MRCP) and Real-Time Transport Protocol (RTP) are used to control the interaction process and transmit voice packets, etc. It involves both Transmission Control Protocol (TCP) communication and User Datagram Protocol (UDP) communication, and the two communications also need to cooperate. This kind of interaction is very complex and closely linked. Many factors (such as network conditions, port abnormalities, etc.) can easily cause a destructive butterfly effect; Second, once an interaction anomaly occurs, it is very difficult to detect it in time. Because the customer service robot system is one-way and outbound in a patterned way without human participation. Once a communication anomaly occurs, it will cause many problems such as the call being hung up all the time after being connected, the call being hung up immediately after being connected, no sound on the phone, no response after the customer replies, etc. Customers can discover it in time, but there is no feedback channel; or even if there is a feedback channel, the customer's intention to feedback problems is not high, which causes the problem to continue to ferment. Coupled with the superimposed factor of the large volume of outbound calls of the customer service robot, the scope of influence is huge, which has an adverse impact on the company's image and even causes customer complaints.
[0098] In the prior art, if the traditional log sampling monitoring method of a general application system is used to monitor the service status of the media resource server and intelligent voice devices, there will be many limitations, such as the keyword matching rules are not fully applicable, the problem determination time windows are different, the recognition rules for transaction response timeouts are single, and the communication status cannot be monitored.
[0099] In view of the above technical problems, the present application proposes the following technical concept: by utilizing the interaction characteristics between the media resource server and the intelligent voice device, parsing the incremental log generated during the interaction between the media resource server and the intelligent voice device to obtain the handshake round of the handshake request message; and performing alarm processing or early warning processing in a timely manner based on the handshake round, so that the administrator and the application system can intervene in a timely manner to prevent communication anomalies from occurring.
[0100] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0101] Figure 1 FIG. is an application scenario diagram of the service status monitoring method applicable to the embodiments of the present application. As Figure 1 shown, this application scenario includes: an intelligent voice device 101, a media resource server 102, a network storage device 103, and a monitoring device 104. Among them, a large number of logs and voice files are generated during the interaction between the intelligent voice device 101 and the media resource server 102, and the generated large number of logs and voice files are stored in the network storage device 103. The monitoring device 104 can read the logs and voice files stored in the monitoring device 104, and monitor the interaction status between the intelligent voice device 101 and the media resource server 102 by parsing the logs and / or voice files.
[0102] Based on Figure 1 the application scenario shown, the embodiments of the present application provide a service status monitoring method. The execution subject of this method can be Figure 1 the monitoring device shown in Figure 2 FIG. is the flow of the service status monitoring method provided by the embodiments of the present application Figure 1 , and this embodiment provides monitoring of the service status in the handshake stage between the intelligent voice device 101 and the media resource server 102. As Figure 2 shown, this service status monitoring method includes:
[0103] S201. Obtain the incremental log during the interaction between the media resource server and the intelligent voice device.
[0104] In this embodiment, the media resource server is used to receive requests from the application system, send a voice synthesis request to the intelligent voice device, and then play the synthesized voice for the customer to listen to; receive the customer's voice, send a voice recognition request to the intelligent voice device, and then return the voice recognition result to the application system. The intelligent voice device is used to respond to the voice recognition request and voice synthesis request of the media resource server.
[0105] S202. Initialize the preset handshake round counter to 0, where the handshake round counter is used to record the handshake rounds during the interaction between the media resource server and the intelligent voice device.
[0106] In this embodiment, the service status monitoring method is used to monitor the service status during the handshake phase. Figure 5 This is a schematic diagram of the normal handshake interaction and abnormal handshake interaction between the media resource server and the intelligent voice device during the handshake phase in the embodiment of the present application. The normal handshake interaction and abnormal handshake interaction are as Figure 5 shown:
[0107] Normal handshake interaction:
[0108] Step 1 - A - 2: The media resource server sends an INVITE message, and the intelligent voice device returns Trying and OK messages;
[0109] Step B - 3 - C: The media resource server sends an ACK message to successfully end this handshake.
[0110] Abnormal handshake interaction:
[0111] Step 1 - A1: The media resource server sends an INVITE message, and the intelligent voice device does not return a response message;
[0112] Step 2 - E1: The media resource server sends a BYE message to end this handshake unsuccessfully.
[0113] S203. Parse each incremental log. If a handshake request message is parsed, increment the handshake round counter by 1. If no handshake request message is parsed, set the handshake round counter to 0.
[0114] In this embodiment, the handshake round counter is used to record the handshake rounds during the interaction between the media resource server and the intelligent voice device.
[0115] S204. If the count of the handshake round counter is greater than or equal to the warning threshold, enter the warning processing, and enter a sleep for a preset duration after the warning processing is completed.
[0116] In this embodiment, the warning threshold can be set as needed, and the embodiment of the present application does not impose any restrictions on this. For example, the warning threshold is 1, 2, or 3 times.
[0117] The execution entity of the warning processing is the monitoring device. The warning processing method can be sending an alarm text message, an email, etc. The preset duration of sleep can be set to sleep for 30s. The configuration of the sleep time is the handshake round interval time between the media resource server and the intelligent voice device, which is convenient for monitoring each handshake in a timely manner.
[0118] S205. After completing the sleep for the preset duration, resume the step of parsing each incremental log.
[0119] In this embodiment, after completing the monitoring of one handshake round, continue with the monitoring of the next handshake round.
[0120] In summary, the service status monitoring method provided in this embodiment parses the incremental logs during the interaction between the media resource server and the intelligent voice device to obtain the handshake rounds of the handshake request messages, and performs early warning processing in a timely manner based on the handshake rounds, enabling the administrator and the application system to intervene in a timely manner and prevent communication anomalies from occurring.
[0121] In the embodiment of the present application, Figure 1 Based on the provided embodiment, after S204, the following steps are further included, which explain the situation where the count of the handshake round counter is less than the early warning threshold, including:
[0122] S206. If the count of the handshake round counter is less than the early warning threshold, continue to determine whether the count of the handshake round counter is greater than or equal to the alarm threshold.
[0123] In this embodiment, the alarm threshold can be set as needed, and the embodiments of the present application do not impose any restrictions on this. For example, the alarm threshold is 2, 3, or 4 times, and the alarm threshold is greater than the early warning threshold.
[0124] S207. If the count of the handshake round counter is greater than or equal to the alarm threshold, enter the emergency handling, and enter the sleep for the preset duration after the emergency handling is completed; if the count of the handshake round counter is less than the alarm threshold, enter the sleep for the preset duration.
[0125] After completing the sleep for the preset duration, resume Figure 1 the step of parsing each incremental log in the provided embodiment.
[0126] In this embodiment, the execution entity of the emergency handling is the monitoring device, and the ways of emergency handling can be sending alarm text messages, emails, etc. The preset duration of sleep can be set to 30s of sleep, and the configuration of the sleep time is the handshake round interval time between the media resource server and the intelligent voice device, which is convenient for monitoring each round of handshake in a timely manner.
[0127] In summary, the service status monitoring method provided in this embodiment processes the situation where the count of the handshake round counter is less than the early warning threshold, so as to be able to perform emergency handling in a timely manner, enabling the administrator and the application system to intervene in a timely manner and prevent communication anomalies from occurring.
[0128] Figure 3 This is the process of the service status monitoring method provided in the embodiment of the present application Figure 2 . AsFigure 3 As shown, the monitoring method for the intelligent voice service status monitors the service status during the communication stage between the intelligent voice device 101 and the media resource server 102 using the Media Resource Control Protocol (MRCP). The MRCP interaction uses the TCP communication protocol and includes three types of interactions: Automatic Speech Recognition (ASR), Text To Speech (TTS), and heartbeat. The method includes the following steps:
[0129] S301. Obtain the incremental log of the interaction process between the media resource server and the intelligent voice device.
[0130] In this embodiment, the service status monitoring method is used for the service status during the MRCP communication stage. Figure 6 This is a schematic diagram of the normal and abnormal interactions between the media resource server and the intelligent voice device during the MRCP communication stage in the embodiment of the present application. The normal and abnormal interactions are as Figure 6 shown:
[0131] Normal TTS interaction:
[0132] Step 1 - D - 2: The media resource server sends a SPEAK message, and the intelligent voice device returns an IN-PROGRESS message.
[0133] Step E - 3: The intelligent voice device returns a SPEAK-COMPLETE message, successfully ending a TTS interaction.
[0134] Abnormal TTS interaction:
[0135] Step 1 - D1 - 2: The media resource server sends a SPEAK message, and the intelligent voice device returns a 40*COMPLETE message.
[0136] Normal ASR interaction:
[0137] Step 4 - F - 5: The media resource server sends a RECOGNIZE message, and the intelligent voice device returns an IN-PROGRESS message and a START-OF-INPUT.
[0138] Step G - 6: The intelligent voice device returns a RECOGNIZE-COMPLETE message with a Completion-cause of 000success, successfully ending a TTS interaction.
[0139] Abnormal ASR interaction:
[0140] Step 4—F1—5: The media resource server sends a RECOGNIZE message, and the intelligent voice device returns an IN-PROGRESS message and a START-OF-INPUT message.
[0141] Step G1—6: The intelligent voice device returns a RECOGNIZE-COMPLETE message, but the Completion-cause is not 000, and a TTS interaction ends in failure.
[0142] S302. Analyze each incremental log.
[0143] In this embodiment, the incremental log is the incremental log of the interaction process between the media resource server and the intelligent voice device.
[0144] S303. Determine the request message type of the media resource control protocol;
[0145] In this embodiment, only two types of request messages of the media resource control protocol are discussed: the text-to-speech request and the automatic speech recognition request.
[0146] S304. If the request message type of the media resource control protocol is a text-to-speech request, determine whether there is a channel identifier in the cache for the text-to-speech request. If not, execute S305; if so, execute S306.
[0147] In this embodiment, there is a channel identifier (Channel-Identifier) field in the messages of the three interactions of automatic speech recognition, text-to-speech, and heartbeat, which can be regarded as the unique ID for each ASR or TTS path.
[0148] S305. Cache the channel identifier.
[0149] In this embodiment, the channel identifier is cached in the monitoring device.
[0150] S306. Increment the text-to-speech failure counter by 1.
[0151] In this embodiment, the preset text-to-speech failure counter is first initialized to 0, where the text-to-speech failure counter is used to record the number of failures of text-to-speech requests during the interaction process between the media resource server and the intelligent voice device.
[0152] S307. Determine whether a return message of the channel identifier is parsed in the subsequent incremental log. If not, execute S308; if so, execute S314.
[0153] S308. Increment the text-to-speech failure counter by 1.
[0154] In this embodiment, the function of the text-to-speech failure counter is the same as that in step S306.
[0155] S309. Determine whether the number of text-to-speech failures is greater than the warning threshold. If not, execute S310; if so, execute S311.
[0156] In this embodiment, the warning threshold can be set as needed, and the embodiments of this application do not impose any restrictions on this. For example, the warning threshold is 1, 2, or 3 times.
[0157] S310. Determine whether the number of text-to-speech failures is greater than the alarm threshold. If not, execute S313; if so, execute S312.
[0158] In this embodiment, the alarm threshold can be set as needed, and the embodiments of this application do not impose any restrictions on this. For example, the warning threshold is 2, 3, or 4 times, and the alarm threshold is greater than the warning threshold.
[0159] In this embodiment, the warning threshold in step S309 and the alarm threshold in S310 are both very small. Because once there is a problem with the text-to-speech request and the user cannot hear the voice, it will seriously affect the customer experience and business process, so the sensitivity requirements for monitoring are very high.
[0160] S311. Enter the warning processing.
[0161] In this embodiment, the execution entity of the warning processing is the monitoring device, and the warning processing method can be sending alarm text messages, emails, etc.
[0162] S312. Enter the emergency processing.
[0163] In this embodiment, the execution entity of the emergency processing is the monitoring device, and the emergency processing method can be sending alarm text messages, emails, etc.
[0164] S313. Enter the sleep for a preset duration; after step S313 ends, execute step S301 again.
[0165] In this embodiment, the preset duration of sleep is inversely proportional to the number of concurrent calls of the media resource server. The greater the concurrency, the more text-to-speech requests it means, and the higher the monitoring frequency should be, and the shorter the sleep time.
[0166] S314. Delete the channel identifier from the cache; after step S314 ends, execute step S313.
[0167] S315. If the request message type of the media resource control protocol is an automatic speech recognition request, increment the automatic speech recognition request counter by 1.
[0168] In this embodiment, the automatic speech recognition request counter is used to record the number of automatic speech recognition requests sent during the interaction between the media resource server and the intelligent voice device.
[0169] S316. Determine whether a channel identifier exists in the cache; if not, execute S317; if so, execute S318.
[0170] In this embodiment, the concept of the channel identifier is the same as that in step S304.
[0171] S317. Cache the channel identifier.
[0172] In this embodiment, the channel identifier is cached in the monitoring device.
[0173] S318. Increment the automatic speech recognition failure counter by 1.
[0174] In this embodiment, the automatic speech recognition failure counter is used to record the number of failures of automatic speech recognition requests during the interaction between the media resource server and the intelligent voice device.
[0175] S319. Determine whether a return message of the channel identifier is parsed in each incremental log; if not, execute S320; if so, execute S322.
[0176] S320. Increment the automatic speech recognition failure counter by 1.
[0177] S321. Determine whether the automatic speech recognition failure rate reaches a preset threshold; if so, execute S311; if not, execute S313.
[0178] In this embodiment, the preset threshold is automatic speech recognition failure count / automatic speech recognition request count = 100%.
[0179] According to practical experience, for automatic speech recognition interactions, due to signal problems, low user voice, unclear pronunciation, dialects, etc., some automatic speech recognition can successfully return, but the recognition result is empty, and the keyword of the return message is no - input - timeout. However, this cannot be completely regarded as an automatic speech recognition failure of the intelligent voice product. Only when the automatic speech recognition failure count / automatic speech recognition request count = 100% can it be determined that an automatic speech recognition failure has occurred. Therefore, in addition to the automatic speech recognition failure counter, an automatic speech recognition request counter also needs to be set.
[0180] S322. Set the automatic speech recognition request counter and the automatic speech recognition failure counter to 0.
[0181] In this embodiment, if it is detected that the automatic speech recognition successfully returns a message, it is determined that the automatic speech recognition is fault-free. Therefore, the automatic speech recognition request counter needs to be cleared to prevent historical data from affecting the calculation result of the automatic speech recognition failure rate.
[0182] S323. Delete the channel identifier in the cache. After step S323 ends, execute step S313.
[0183] In this embodiment, after entering the sleep state for a preset duration, the step of obtaining the incremental log of the interaction process between the media resource server and the intelligent voice device is executed again.
[0184] In summary, the service status monitoring method provided in this embodiment determines the request message type of the media resource control protocol according to the incremental logs of the automatic speech recognition interaction and the text-to-speech interaction phases in the MRCP communication phase between the media resource server and the intelligent voice device, discovers the anomalies in the interaction process between the media resource server and the intelligent voice device, and issues warnings and early warnings of communication anomalies in a timely manner according to the number of text-to-speech failures and the magnitude of the automatic speech recognition failure rate, enabling the administrator and the application system to intervene in a timely manner and prevent the occurrence of communication anomalies.
[0185] Figure 4 is the flow of the service status monitoring method provided in the embodiment of the present application Figure 3 , this embodiment provides the monitoring of the service status in the Real-time Transport Protocol (RTP) communication phase between the intelligent voice device 101 and the media resource server 102, such as Figure 4 shown, this service status monitoring method includes the following steps:
[0186] S401. Obtain the incremental log and incremental voice files of the interaction process between the media resource server and the intelligent voice device.
[0187] In this embodiment, the RTP communication phase is accompanied by the MRCP communication phase. The MRCP controls the interaction processes of ASR, TTS, and heartbeat, and the RTP controls the transmission of voice files. During TTS communication, the intelligent voice device returns the synthesized TTS voice file. Since the UDP protocol is used for transmission, there may be problems such as the normal interaction of TTS MRCP messages, but the media resource server does not receive the voice file or receives an incomplete voice file. Therefore, in this stage, the focus is on monitoring the voice files received by the media resource server.
[0188] In this embodiment, the incremental log and incremental voice files of the interaction process between the media resource server and the intelligent voice device are read from the Network Attached Storage (NAS).
[0189] A network storage device is a special dedicated data storage server that provides file sharing functionality. During the interaction between a media resource server and an intelligent voice product, a large number of logs and voice files are stored on the network storage device.
[0190] S402. Initialize the preset text-to-speech request counter to 0.
[0191] In this embodiment, the text-to-speech request count is used to record the number of text-to-speech requests during the interaction between the media resource server and the intelligent voice device. This step is only preset when the monitoring program starts.
[0192] S403. Parse each incremental log.
[0193] S404. Determine whether the request message type of the media resource control protocol is a text-to-speech request. If so, execute S405.
[0194] S405. Determine whether a channel identifier exists in the cache for the text-to-speech request. If not, execute S406. If so, execute S407.
[0195] In this embodiment, the concept of the channel identifier is the same as that in step S304.
[0196] S406. Configure the correspondence between the synthesized text for text-to-speech and the size value of the voice packet generated by the synthesized text; parse the synthesized text in this text-to-speech request, query the corresponding size value of the voice packet, and save the mapping relationship between the channel identifier and the size value of the voice packet for this text-to-speech request.
[0197] In this embodiment, in the customer service robot system, except for words such as identity confirmation and balance inquiry that vary from person to person, most of the robot words are pre-configured. Therefore, under normal interaction conditions, the size of the voice packet generated by the synthesized text is a fixed value, which may have slight fluctuations due to network conditions. Therefore, the correspondence between the synthesized text for text-to-speech and the size value of the voice packet generated by the synthesized text can be configured in the monitoring device. By querying the size of the voice packet, the mapping relationship between the channel identifier and the size value of the voice packet for this text-to-speech request is saved, which is convenient for subsequent query and comparison of the voice packet size.
[0198] S407. Increment the text-to-speech failure counter by 1.
[0199] In this embodiment, the text-to-speech failure counter is used to record the number of failed text-to-speech requests during the interaction between the media resource server and the intelligent voice device.
[0200] S408. Determine whether a return message of the channel identifier is found in each incremental log. If not, execute S409 to S414. If so, execute S415.
[0201] Among them, the steps of S409 are the same as those of S308, the steps of S410 are the same as those of S309, the steps of S411 are the same as those of S310, and the steps of S412 are the same as those of S312; the steps of S413 are the same as those of S311, and the steps of S414 are the same as those of S3131. They will not be elaborated here.
[0202] S415. Increment the text-to-speech request counter by 1.
[0203] In this embodiment, the text-to-speech request count is used to record the number of text-to-speech requests during the interaction between the media resource server and the intelligent voice device.
[0204] S416. Parse the incremental voice file.
[0205] In this embodiment, the voice file is a voice file generated during the interaction between the media resource server and the intelligent voice device.
[0206] S417. Use the size value of the voice packet corresponding to the synthesized speech to traverse the size of the incremental voice file and calculate the deviation rate.
[0207] In this embodiment, the deviation rate is: the ratio of the absolute value of the difference between the size value of the voice packet corresponding to the synthesized speech and the size value of the incremental voice file to the size value of the voice packet corresponding to the synthesized speech.
[0208] S418. Determine whether there is an incremental voice file with a deviation rate less than a specific threshold. If so, execute S419. If not, execute S413.
[0209] In this embodiment, the specific threshold can be set as needed, and this application embodiment does not impose any restrictions on it. For example, the specific threshold is 4%, 5% or 6%.
[0210] If there is a file with a deviation rate less than the specific threshold, it is considered that the voice packet file corresponding to the speech successfully transmits with little loss; otherwise, it is considered that there is a problem with the voice packet transmission.
[0211] S419. Determine whether the number of incremental voice files is less than the number of text-to-speech requests. If so, execute S413. If not, execute S414.
[0212] In this embodiment, after parsing all incremental logs, obtain the number of incremental voice files and compare it with the number of text-to-speech requests. If the number of incremental voice files is less than the number of text-to-speech requests, it means that some voice packets for text-to-speech are not received and emergency processing is required.
[0213] In summary, for the service status monitoring method provided in this embodiment, based on the incremental logs and voice files during the text-to-speech interaction stage in the RTP communication stage between the media resource server and the intelligent voice device, the voice packets received by the media resource server are monitored to promptly detect anomalies during the interaction between the media resource server and the intelligent voice device. A warning and early warning of communication anomalies are promptly issued based on the deviation rate between the size value of the voice packet and the size value of the incremental voice file, enabling the administrator and the application system to intervene in a timely manner and prevent the occurrence of communication anomalies.
[0214] Figure 7 It is a schematic structural diagram of the service status monitoring device provided in an embodiment of the present application. As Figure 7 shown, the service status monitoring device 70 includes: an acquisition module 701 and a processing module 702.
[0215] Among them, the acquisition module 701 is used to acquire the incremental logs during the interaction between the media resource server and the intelligent voice device.
[0216] The processing module 702 is used to initialize the preset handshake round counter to 0, where the handshake round counter is used to record the number of handshake rounds during the interaction between the media resource server and the intelligent voice device.
[0217] The processing module 702 is further used to parse each incremental log. If a handshake request message is parsed, the handshake round counter is incremented by 1. If no handshake request message is parsed, the handshake round counter is set to 0.
[0218] The processing module 702 is further used to enter the early warning process if the count of the handshake round counter is greater than or equal to the early warning threshold, and enter a sleep state for a preset duration after the early warning process is completed.
[0219] The processing module 702 is further used to re-execute the step of parsing each incremental log after completing the sleep state for the preset duration.
[0220] In some embodiments, the processing module 702 is further specifically used for:
[0221] If the count of the handshake round counter is less than the early warning threshold, continue to determine whether the count of the handshake round counter is greater than or equal to the alarm threshold, where the alarm threshold is greater than the early warning threshold;
[0222] If the count of the handshake round counter is greater than or equal to the alarm threshold, enter the emergency handling process, and enter a sleep state for a preset duration after the emergency handling process is completed; if the count of the handshake round counter is less than the alarm threshold, enter a sleep state for a preset duration;
[0223] After completing the sleep state for the preset duration, re-execute the step of parsing each incremental log.
[0224] In some embodiments, the processing module 702 is specifically configured to:
[0225] Initialize a preset text-to-speech failure counter to 0, where the text-to-speech failure counter is used to record the number of failures of text-to-speech requests during the interaction between the media resource server and the intelligent voice device.
[0226] Parse each incremental log to determine the request message type of the media resource control protocol.
[0227] If the request message type of the media resource control protocol is a text-to-speech request, determine whether a channel identifier exists in the cache for the text-to-speech request.
[0228] If a channel identifier exists in the cache for the text-to-speech request, increment the text-to-speech failure counter by 1; if a channel identifier does not exist in the cache for the text-to-speech request, cache the channel identifier.
[0229] If a return message of the channel identifier is parsed in subsequent incremental logs, delete the channel identifier from the cache and enter a sleep state for a preset duration.
[0230] If a return message of the channel identifier is not parsed in each incremental log, increment the text-to-speech failure counter by 1 and continue to determine whether the text-to-speech failure count is greater than the warning threshold.
[0231] If the text-to-speech failure count is greater than the warning threshold, enter warning processing and enter a sleep state for a preset duration after the warning processing is completed.
[0232] After completing the sleep state for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0233] In some embodiments, the processing module 702 is further specifically configured to:
[0234] If a return message of the channel identifier is not parsed in each incremental log, increment the text-to-speech failure counter by 1 and continue to determine whether the text-to-speech failure count is greater than the warning threshold;
[0235] If the text-to-speech failure count is greater than the warning threshold, enter warning processing and enter a sleep state for a preset duration after the warning processing is completed;
[0236] If the text-to-speech failure count is less than or equal to the warning threshold, continue to determine whether the text-to-speech failure count is greater than the alarm threshold, where the alarm threshold is greater than the warning threshold;
[0237] If the number of text-to-speech failures is greater than the alarm threshold, enter emergency handling and enter a sleep state for a preset duration after the emergency handling is completed; if the number of text-to-speech failures is less than or equal to the alarm threshold, enter a sleep state for a preset duration;
[0238] After completing the sleep state for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0239] In some embodiments, the processing module 702 is specifically configured to:
[0240] Initialize both the preset automatic speech recognition request counter and the automatic speech recognition failure counter to 0, where the automatic speech recognition request counter is used to record the number of automatic speech recognition requests sent during the interaction between the media resource server and the intelligent voice device, and the automatic speech recognition failure counter is used to record the number of failures of the automatic speech recognition requests during the interaction between the media resource server and the intelligent voice device;
[0241] If the request message type of the media resource control protocol is an automatic speech recognition request, increment the automatic speech recognition request counter by 1; if the channel identifier exists in the cache for the automatic speech recognition request, increment the automatic speech recognition failure counter by 1; if the channel identifier does not exist in the cache for the automatic speech recognition request, cache the channel identifier;
[0242] If a return message of the channel identifier is parsed in each incremental log, set both the automatic speech recognition request counter and the automatic speech recognition failure counter to 0, delete the channel identifier in the cache, and enter a sleep state for a preset duration;
[0243] After completing the sleep state for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0244] In some embodiments, the acquisition module 702 is also specifically configured to:
[0245] If a return message of the channel identifier is not parsed in each incremental log, increment the automatic speech recognition failure counter by 1, and continue to determine whether the automatic speech recognition failure rate reaches a preset threshold, where the automatic speech recognition failure rate is the ratio of the number of failures of the automatic speech recognition requests to the number of automatic speech recognition requests;
[0246] If the automatic speech recognition failure rate reaches the preset threshold, enter emergency handling and enter a sleep state for a preset duration after the emergency handling is completed; if the automatic speech recognition failure rate does not reach the preset threshold, enter a sleep state for a preset duration;
[0247] After completing the sleep for the preset duration, the step of parsing each incremental log again to determine the request message type of the media resource control protocol is performed again.
[0248] In some embodiments, the obtaining module 702 is specifically configured to: obtain the incremental log and the incremental voice file during the interaction between the media resource server and the intelligent voice device;
[0249] The processing module 702 is specifically configured to:
[0250] Initialize the preset text-to-speech request counter to 0, where the text-to-speech request count is used to record the number of text-to-speech requests during the interaction between the media resource server and the intelligent voice device;
[0251] Parse each incremental log to determine the request message type of the media resource control protocol;
[0252] If the request message type of the media resource control protocol is a text-to-speech request, continue to determine whether there is a channel identifier in the cache for the text-to-speech request;
[0253] If there is a channel identifier in the cache for the text-to-speech request, increment the text-to-speech failure counter by 1;
[0254] If there is no channel identifier in the cache for the text-to-speech request, configure the correspondence between the synthesized text in the text-to-speech request and the size value of the voice packet generated by the synthesized text;
[0255] Parse the synthesized text in the current text-to-speech request, query the corresponding size value of the voice packet, and save the mapping relationship between the channel identifier and the size value of the voice packet for the current text-to-speech request;
[0256] If a return message of the channel identifier is parsed in each incremental log, increment the text-to-speech request counter by 1;
[0257] Parse the incremental voice file, use the size value of the voice packet corresponding to the channel identifier to traverse the size of the incremental voice file, and calculate the deviation rate, where the deviation rate is: the ratio of the absolute value of the difference between the size value of the voice packet corresponding to the synthesized text and the size value of the incremental voice file to the size value of the voice packet corresponding to the synthesized text;
[0258] If there is no incremental voice file with a deviation rate less than a specific threshold, enter emergency processing, and enter sleep for the preset duration after the emergency processing is completed;
[0259] After completing the sleep for the preset duration, the step of parsing each incremental log again to determine the request message type of the media resource control protocol is performed again.
[0260] In some embodiments, the processing module 702 is further specifically configured to:
[0261] If there is an incremental voice file with a deviation rate less than a specific threshold, continue to determine whether the number of incremental voice files is less than the number of requests for text-to-speech. If the number of incremental voice files is less than the number of requests for text-to-speech, enter emergency processing, and enter a sleep state for a preset duration after the emergency processing is completed; if the number of incremental voice files is greater than or equal to the number of requests for text-to-speech, enter a sleep state for a preset duration;
[0262] After completing the sleep state for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
[0263] In some embodiments, the acquisition module 702 is specifically configured to: read the incremental log and incremental voice files of the interaction process between the media resource server and the intelligent voice device from the network memory.
[0264] The service status monitoring device provided by the embodiments of the present application can be used to execute the technical solutions of the service status monitoring method in the above embodiments, and its implementation principle and technical effects are similar, and will not be described in detail here.
[0265] Figure 8 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 8 shown, the electronic device 80 in this embodiment includes: a processor 801 and a memory 802; where
[0266] The memory 802 is used to store computer execution instructions;
[0267] The processor 801 is used to execute the computer execution instructions stored in the memory to implement each step executed by the receiving device in the above embodiments. For specific reference, see the relevant descriptions in the foregoing method embodiments.
[0268] Optionally, the memory 802 can be either independent or integrated with the processor 801.
[0269] When the memory 802 is independently provided, the electronic control device further includes a bus 803 for connecting the memory 802 and the processor 801.
[0270] The embodiments of the present disclosure also provide a computer-readable storage medium, in which computer execution instructions are stored. When the processor executes the computer execution instructions, the service status monitoring method as described above is implemented.
[0271] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.
[0272] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0273] In addition, each functional module in various embodiments of the present disclosure can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0274] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present application.
[0275] It should be understood that the above processor can be a Central Processing Unit (CPU for short), and can also be other general-purpose processors, Digital Signal Processors (DSP for short), Application Specific Integrated Circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of hardware and software modules in the processor.
[0276] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disc, etc.
[0277] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0278] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0279] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0280] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, a magnetic disk, or an optical disc.
[0281] The embodiments of this application also provide a computer program product. The computer program product includes a computer program, which is stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When at least one processor executes the computer program, the technical solutions of the monitoring method for the intelligent voice service state in the above embodiments can be implemented.
[0282] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the following claims.
[0283] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A service status monitoring method, characterized in that, Applied to a monitoring device, including: Obtain the incremental log of the interaction process between the media resource server and the intelligent voice device; Initialize the preset handshake round counter to 0, where the handshake round counter is used to record the handshake rounds in the interaction process between the media resource server and the intelligent voice device; Parse each incremental log. If a handshake request message is parsed, increment the handshake round counter by 1. If no handshake request message is parsed, set the handshake round counter to 0. If the count of the handshake round counter is greater than or equal to the warning threshold, enter warning processing and enter a sleep for a preset duration after the warning processing is completed; After completing the sleep for the preset duration, re-execute the step of parsing each incremental log; After obtaining the incremental log of the interaction process between the media resource server and the intelligent voice device, it further includes: Initialize the preset text-to-speech failure counter to 0, where the text-to-speech failure counter is used to record the number of failures of text-to-speech requests in the interaction process between the media resource server and the intelligent voice device; Parse each incremental log to determine the request message type of the media resource control protocol; If the request message type of the media resource control protocol is a text-to-speech request, determine whether a channel identifier exists in the cache for the text-to-speech request; If the text-to-speech request has a channel identifier in the cache, increment the text-to-speech failure counter by 1. If the text-to-speech request does not have a channel identifier in the cache, cache the channel identifier; If a return message of the channel identifier is parsed in subsequent incremental logs, delete the channel identifier from the cache and enter a sleep for a preset duration; After completing the sleep for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
2. The method according to claim 1, wherein It further includes: If the count of the handshake round counter is less than the warning threshold, continue to determine whether the count of the handshake round counter is greater than or equal to the alarm threshold, where the alarm threshold is less than the warning threshold; If the count of the handshake round counter is greater than or equal to the alarm threshold, enter emergency processing and enter a sleep for a preset duration after the emergency processing is completed; If the count of the handshake round counter is less than the alarm threshold, enter a sleep for a preset duration; After completing the sleep for the preset duration, re-execute the step of parsing each incremental log.
3. The method according to claim 1, wherein It further includes: If no return message of the channel identifier is parsed in each incremental log, increment the text-to-speech failure counter by 1 and continue to determine whether the text-to-speech failure count is greater than the warning threshold; If the text-to-speech failure count is greater than the warning threshold, enter warning processing and enter a sleep for a preset duration after the warning processing is completed; If the text-to-speech failure count is less than or equal to the warning threshold, continue to determine whether the text-to-speech failure count is greater than the alarm threshold, where the alarm threshold is less than the warning threshold; If the text-to-speech failure count is greater than the alarm threshold, enter emergency processing and enter a sleep for a preset duration after the emergency processing is completed; If the number of text-to-speech failures is less than or equal to the alarm threshold, enter a sleep state for a preset duration; After completing the sleep state for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
4. The method according to claim 3, wherein After parsing each incremental log to determine the request message type of the media resource control protocol, it further includes: Initialize both the preset automatic speech recognition request counter and the automatic speech recognition failure counter to 0, where the automatic speech recognition request counter is used to record the number of automatic speech recognition requests sent during the interaction between the media resource server and the intelligent voice device, and the automatic speech recognition failure counter is used to record the number of failures of automatic speech recognition requests during the interaction between the media resource server and the intelligent voice device; If the request message type of the media resource control protocol is an automatic speech recognition request, increment the automatic speech recognition request counter by 1; if the automatic speech recognition request has a channel identifier in the cache, increment the automatic speech recognition failure counter by 1; if the automatic speech recognition request does not have a channel identifier in the cache, cache the channel identifier; If the return message of the channel identifier is parsed in each incremental log, set both the automatic speech recognition request counter and the automatic speech recognition failure counter to 0, delete the channel identifier in the cache, and enter a sleep state for a preset duration; After completing the sleep state for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
5. The method according to claim 4, characterized in that, It further includes: If the return message of the channel identifier is not parsed in each incremental log, increment the automatic speech recognition failure counter by 1, and continue to determine whether the automatic speech recognition failure rate reaches a preset threshold, where the automatic speech recognition failure rate is the ratio of the number of failures of automatic speech recognition requests to the number of automatic speech recognition requests; If the automatic speech recognition failure rate reaches the preset threshold, enter emergency processing, and enter a sleep state for a preset duration after the emergency processing is completed; If the automatic speech recognition failure rate does not reach the preset threshold, enter a sleep state for a preset duration; After completing the sleep state for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
6. The method according to claim 1, wherein It further includes: Obtain the incremental log and the incremental voice file during the interaction between the media resource server and the intelligent voice device; Initialize the preset text-to-speech request counter to 0, where the text-to-speech request counter is used to record the number of text-to-speech requests during the interaction between the media resource server and the intelligent voice device; Parse each incremental log to determine the request message type of the media resource control protocol; If the request message type of the media resource control protocol is a text-to-speech request, continue to determine whether the text-to-speech request has a channel identifier in the cache; If the text-to-speech request has the channel identifier in the cache, increment the text-to-speech failure counter by 1; If the channel identifier does not exist in the cache for the text-to-speech request, configure the correspondence between the synthesized speech text in the text-to-speech request and the size value of the speech packet generated by the synthesized speech text; Parse the synthesized speech text in the current text-to-speech request, query the size value of the corresponding speech packet, and save the mapping relationship between the channel identifier of the current text-to-speech request and the size value of the speech packet; If the return message of the channel identifier is parsed in each incremental log, increment the text-to-speech request counter by 1; Parse the incremental speech file, use the size value of the speech packet corresponding to the channel identifier to traverse the size of the incremental speech file, and calculate the deviation rate, where the deviation rate is: the ratio of the absolute value of the difference between the size value of the speech packet corresponding to the synthesized speech text and the size value of the incremental speech file to the size value of the speech packet corresponding to the synthesized speech text; If there is no incremental speech file with a deviation rate less than a specific threshold, enter emergency processing and enter a sleep for a preset duration after the emergency processing is completed; After completing the sleep for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
7. The method according to claim 6, wherein Further included: If there is an incremental speech file with a deviation rate less than a specific threshold, continue to determine whether the number of incremental speech files is less than the number of text-to-speech requests. If the number of incremental speech files is less than the number of text-to-speech requests, enter emergency processing and enter a sleep for a preset duration after the emergency processing is completed; if the number of incremental speech files is greater than or equal to the number of text-to-speech requests, enter a sleep for a preset duration; After completing the sleep for the preset duration, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
8. The method according to claim 6, characterized in that, The obtaining of the incremental log and the incremental speech file in the interaction process between the media resource server and the intelligent voice device includes: Read the incremental log and the incremental speech file in the interaction process between the media resource server and the intelligent voice device from the network storage.
9. A service status monitoring device, characterized in that, Included: An obtaining module for obtaining an incremental log in the interaction process between the media resource server and the intelligent voice device; A processing module for initializing a preset handshake round counter to 0, where the handshake round counter is used to record the handshake rounds in the interaction process between the media resource server and the intelligent voice device; The processing module is further used to parse each incremental log. If a handshake request message is parsed, increment the handshake round counter by 1. If no handshake request message is parsed, set the handshake round counter to 0; The processing module is further used to enter warning processing if the count of the handshake round counter is greater than or equal to the warning threshold, and enter a sleep for a preset duration after the warning processing is completed; The processing module is further used to re-execute the step of parsing each incremental log after completing the sleep for the preset duration; The processing module is further used for: Initializing a preset text-to-speech failure counter to 0, where the text-to-speech failure counter is used to record the number of failures of text-to-speech requests in the interaction process between the media resource server and the intelligent voice device; Parse each incremental log to determine the request message type of the media resource control protocol; If the request message type of the media resource control protocol is a text-to-speech request, determine whether a channel identifier exists in the cache for the text-to-speech request; If the text-to-speech request has a channel identifier in the cache, increment the text-to-speech failure counter by 1; if the text-to-speech request does not have a channel identifier in the cache, cache the channel identifier; If a return message for the channel identifier is parsed in a subsequent incremental log, delete the channel identifier from the cache and enter a sleep for a preset duration; After completing the preset duration of sleep, re-execute the step of parsing each incremental log to determine the request message type of the media resource control protocol.
10. An electronic device, characterized in that, Comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1-8.
12. A computer program product, characterized in that, Comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
IP telephone fault alarming method and apparatus based on SIP
CN101605075A
Server and method and device for triggering fusing
CN112527544A