Voice quality evaluation method and device and readable storage medium
By identifying and evaluating effective RTP packets and call end signaling in voice calls, and combining time intervals to perform voice quality evaluation, the problem of inaccurate evaluation caused by insufficient RTP packet relationship in the prior art is solved, and a more accurate and reliable voice quality evaluation is achieved.
Patent Information
- Application Number
- CN202510138168.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The existing voice quality evaluation method only evaluates voice quality based on the relationship between RTP packets, and does not judge the validity of RTP packets and the call end process, resulting in inaccurate evaluation results.
By obtaining the SIP signaling and RTP packet data corresponding to the target call sheet, the first and last packets of the RTP packet are identified, and all RTP packets located between the first packet and the last packet are identified as valid RTP packets. Voice quality evaluation is performed based on the time interval between the valid RTP packets and the time interval between the tail packet and the end-of-talk identifier.
It significantly improves the accuracy and reliability of voice quality evaluation, makes the evaluation results closer to the actual call experience, and solves the inaccurate evaluation caused by existing methods ignoring the effectiveness of RTP packets and the judgment of call process.
Smart Images

Figure CN119997055A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a voice quality assessment method, device and readable storage medium. Background Art
[0002] With the popularization of 5G networks, people have higher expectations for more stable call quality. Therefore, the optimization of voice quality assessment technology is particularly important.
[0003] In the process of research and practice of the prior art, the inventor found that the existing voice quality assessment method only considers the relationship between RTP (Real-time Transport Protocol) packets as the basis for judging whether there is a voice problem, and does not judge the validity of the RTP packet and the call end process. For the start stage of a voice call, there are terminals in the existing network that support the early media (P-Early-Media) function. When the caller starts a call and receives a 180Ring message, the caller starts to send uplink RTP messages, which may include audio or video content, such as public service advertisements in video ringback tone services, etc., to keep the connection active or transmit media content in advance. These early media messages are sent before the called party answers the call, and even if the called party does not answer the call, the calling party will continue to send a certain number of RTP packets as part of the keep-alive mechanism, resulting in inaccurate voice call quality assessment. For the end stage of a voice call, the receipt of a 200OK (BYE) message is generally used as a sign of the end process of the voice call. The judgment of the end of the call by not sending an RTP packet cannot accurately reflect the actual end time of the voice call, resulting in inaccurate voice call quality assessment. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a voice quality assessment method, device and readable storage medium in view of the above-mentioned shortcomings of the prior art, so as to solve the problem that the existing voice quality assessment method only assesses the voice quality based on the relationship between RTP packets, but ignores the validity of the RTP packets and the process judgment at the start and end stages of the call, thereby causing inaccurate assessment results.
[0005] In a first aspect, the present invention provides a method for evaluating speech quality, the method comprising:
[0006] Acquire the Session Initiation Protocol (SIP) signaling corresponding to the target call record and the Real-time Transport Protocol (RTP) packet data corresponding to the SIP signaling;
[0007] Identify the first packet and the last packet of the RTP packet according to the SIP signaling and the RTP packet data;
[0008] Identify the first packet, the last packet, and all RTP packets between the first packet and the last packet in the RTP packet data as valid RTP packets;
[0009] The voice quality of the target call record is evaluated according to the time interval between the valid RTP packets and the time interval between the tail packet and the voice call end identifier to obtain a voice quality evaluation result.
[0010] Further, the acquiring of the Session Initiation Protocol SIP signaling corresponding to the target call record and the Real-time Transport Protocol RTP packet data corresponding to the SIP signaling specifically includes:
[0011] Collect SIP signaling and RTP packet data generated by users during voice calls;
[0012] By matching the collected SIP signaling and RTP packet data according to timestamp information and payload type, a corresponding relationship between the SIP signaling and the RTP packet is obtained;
[0013] The SIP signaling corresponding to the target call record and the RTP packet data corresponding to the SIP signaling are obtained according to the corresponding relationship between the SIP signaling and the RTP packet.
[0014] Furthermore, the collected SIP signaling includes a timestamp and session description protocol SDP information, and the collected RTP packet data includes a sequence number, a timestamp, a synchronization source SSRC, and a payload type.
[0015] Further, the identifying the first packet and the last packet of the RTP packet according to the SIP signaling and the RTP packet data specifically includes:
[0016] Determine the time point of the first message in the SIP signaling as the time point of the voice connection;
[0017] Identify the first RTP packet after the voice connection time point in the RTP packet data as the first packet;
[0018] Determine the time point of the second message in the SIP signaling as the time point of the voice call end;
[0019] The last RTP packet before the voice end time point in the RTP packet data is identified as the tail packet.
[0020] Further, the voice quality of the target call record is evaluated according to the time interval between the valid RTP packets and the time interval between the tail packet and the voice call end identifier to obtain a voice quality evaluation result, which specifically includes:
[0021] Calculate a first time interval between any two consecutive RTP packets in the valid RTP packets;
[0022] Calculating a second time interval between the tail packet and the voice call end marker;
[0023] The voice quality of the target call record is evaluated according to the first time interval and the second time interval to obtain a voice quality evaluation result.
[0024] Further, the evaluating the voice quality of the target call record according to the first time interval and the second time interval to obtain the voice quality evaluation result specifically includes:
[0025] If any one of the first time intervals is greater than a preset first threshold or the second time interval is greater than a preset second threshold, the evaluation result of the voice quality is determined to be poor voice quality.
[0026] Furthermore, the first threshold is 500 ms, and the second threshold is 1000 ms.
[0027] In a second aspect, the present invention provides a speech quality assessment device, the device comprising:
[0028] A call record data acquisition module, used to acquire the session initiation protocol SIP signaling corresponding to the target call record and the real-time transport protocol RTP packet data corresponding to the SIP signaling;
[0029] A head and tail packet identification module, connected to the call bill data acquisition module, for identifying the head and tail packets of the RTP packet according to the SIP signaling and the RTP packet data;
[0030] A valid RTP packet identification module, connected to the first and last packet identification modules, for identifying the first packet, the last packet and all RTP packets between the first packet and the last packet in the RTP packet data as valid RTP packets;
[0031] The voice quality assessment module is connected to the valid RTP packet identification module and is used to evaluate the voice quality of the target call record according to the time interval between the valid RTP packets and the time interval between the tail packet and the voice call end mark to obtain a voice quality assessment result.
[0032] In a third aspect, the present invention provides a speech quality assessment device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the speech quality assessment method described in the first aspect.
[0033] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the speech quality assessment method described in the first aspect is implemented.
[0034] The present invention provides a voice quality assessment method, device and readable storage medium. First, the session initiation protocol SIP signaling corresponding to the target call record and the real-time transport protocol RTP packet data corresponding to the SIP signaling are obtained; then the first packet and the last packet of the RTP packet are identified according to the SIP signaling and the RTP packet data; then the first packet, the last packet and all RTP packets between the first packet and the last packet in the RTP packet data are identified as valid RTP packets; finally, the voice quality of the target call record is evaluated according to the time interval between the valid RTP packets and the time interval between the last packet and the voice call end mark, and the voice quality assessment result is obtained. The present invention ensures that the voice quality assessment is performed only based on the valid RTP packet data by giving priority to judging the validity of the RTP packet. At the same time, a comprehensive analysis is performed in combination with the time interval of the valid RTP packet and the time interval between the last packet and the voice call end mark, which significantly improves the accuracy and reliability of the voice quality assessment, making the assessment result closer to the actual call experience. The present invention solves the problem that the existing voice quality assessment method only assesses the voice quality based on the relationship between RTP packets, but ignores the validity of the RTP packets and the process judgment at the start and end of the call, thus causing inaccurate assessment results. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a flow chart of a method for evaluating speech quality according to Embodiment 1 of the present invention;
[0036] Figure 2 A flowchart of RTP first packet identification according to an embodiment of the present invention;
[0037] Figure 3 A flowchart of RTP tail packet identification according to an embodiment of the present invention;
[0038] Figure 4 This is a schematic diagram of the time points of the RTP first and last packets and the 200OK (BYE) message according to an embodiment of the present invention;
[0039] Figure 5 A flowchart of another method for evaluating speech quality according to an embodiment of the present invention;
[0040] Figure 6 This is a structural diagram of a speech quality assessment device according to Embodiment 2 of the present invention;
[0041] Figure 7 This is a structural diagram of a speech quality assessment device according to Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0042] In order to enable those skilled in the art to better understand the technical solution of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0043] It should be understood that the specific embodiments and drawings described herein are only used to explain the present invention rather than to limit the present invention.
[0044] It can be understood that, in the absence of conflict, the various embodiments of the present invention and the various features in the embodiments can be combined with each other.
[0045] It can be understood that, for the convenience of description, the drawings of the present invention only show the parts related to the present invention, while the parts irrelevant to the present invention are not shown in the drawings.
[0046] It can be understood that each unit and module involved in the embodiments of the present invention may correspond to only one physical structure, or may be composed of multiple physical structures, or multiple units and modules may be integrated into one physical structure.
[0047] It can be understood that the terms "first", "second", etc. in the embodiments of the present invention are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.
[0048] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in an order different from that marked in the drawings.
[0049] It is understood that the flowcharts and block diagrams of the present invention illustrate the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the various embodiments of the present invention. Each box in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified functions. Moreover, each box or combination of boxes in the block diagram and flowchart may be implemented by a hardware-based system that implements the specified functions, or may be implemented by a combination of hardware and computer instructions.
[0050] It can be understood that the units and modules involved in the embodiments of the present invention can be implemented by software or hardware. For example, the units and modules can be located in a processor.
[0051] Embodiment 1:
[0052] This embodiment provides a method for evaluating speech quality. Figure 1 As shown, the method includes:
[0053] Step S101: Acquire the SIP (Session Initiation Protocol) signaling corresponding to the target call record and the RTP packet data corresponding to the SIP signaling.
[0054] In this embodiment, the SIP signaling corresponding to the target call record and the RTP packet data corresponding to the SIP signaling can be obtained based on the pre-collected and stored SIP signaling and RTP packet data.
[0055] Optionally, the acquiring of the Session Initiation Protocol SIP signaling corresponding to the target call record and the Real-time Transport Protocol RTP packet data corresponding to the SIP signaling specifically includes:
[0056] Collect SIP signaling and RTP packet data generated by users during voice calls;
[0057] By matching the collected SIP signaling and RTP packet data according to timestamp information and payload type, a corresponding relationship between the SIP signaling and the RTP packet is obtained;
[0058] The SIP signaling corresponding to the target call record and the RTP packet data corresponding to the SIP signaling are obtained according to the corresponding relationship between the SIP signaling and the RTP packet.
[0059] In this embodiment, the SIP signaling and RTP packet data generated by the user during the voice call can be collected through the core network, wherein the collected SIP signaling includes timestamp and SDP (Session Description Protocol) information, and the collected RTP packet data includes sequence number, timestamp, SSRC (Synchronization Source) and payload type. Among them, SDP is used to associate SIP signaling with RTP packet data, and SSRC is used as a unique identifier for the user call.
[0060] In this embodiment, by analyzing the timestamp and SDP information carried in the SIP signaling, and matching and associating the information such as the Payload Type in the SDP information with the timestamp and Payload Type value carried in the RTP packet data, the correspondence between the SIP signaling and the RTP packet is obtained, and the SIP signaling and RTP packet information of each call record is obtained, wherein the target call refers to any one of the call records.
[0061] Step S102: Identify the first packet and the last packet of the RTP packet according to the SIP signaling and the RTP packet data.
[0062] In this embodiment, in order to obtain a valid RTP packet, the first packet and the last packet of the RTP packet are first identified according to the SIP signaling and the RTP packet data.
[0063] Optionally, the identifying the first packet and the last packet of the RTP packet according to the SIP signaling and the RTP packet data specifically includes:
[0064] Determine the time point of the first message in the SIP signaling as the time point of the voice connection;
[0065] Identify the first RTP packet after the voice connection time point in the RTP packet data as the first packet;
[0066] Determine the time point of the second message in the SIP signaling as the time point of the voice call end;
[0067] The last RTP packet before the voice end time point in the RTP packet data is identified as the tail packet.
[0068] In this embodiment, taking the uplink behavior on the calling side as an example, the ACK message sent by the SBC (Session Border Controller) on the calling side after receiving the 200OK (invite) message sent by the calling terminal after the called party picks up the phone is used as the first message indicating the voice connection identifier. The time point is the time point of the voice connection. After the voice is connected, the first RTP packet sent by the calling terminal received by the SBC on the calling side is used as the first packet.
[0069] In this embodiment, taking the calling side uplink as an example, the first 200OK (BYE) message received by the calling side SBC is used as the second message indicating the end of the voice call. The time point is the time point when the voice ends. After the voice call ends, the last RTP packet sent by the calling terminal received by the calling side SBC is used as the tail packet.
[0070] Step S103: Identify the first packet, the last packet, and all RTP packets between the first packet and the last packet in the RTP packet data as valid RTP packets;
[0071] Step S104: Evaluate the voice quality of the target call record according to the time interval between the valid RTP packets and the time interval between the tail packet and the voice call end mark to obtain a voice quality evaluation result.
[0072] In this embodiment, in order to improve the accuracy and reliability of voice quality assessment, voice quality assessment is performed by identifying valid RTP packet data (including the first packet, the last packet, and all RTP packets between the first packet and the last packet), and then combining the time interval between valid RTP packets and the time interval between the last packet and the voice call end mark. This method makes the assessment result closer to the actual call experience.
[0073] Optionally, the evaluating the voice quality of the target call record according to the time interval between the valid RTP packets and the time interval between the tail packet and the voice call end identifier to obtain the voice quality evaluation result specifically includes:
[0074] Calculate a first time interval between any two consecutive RTP packets in the valid RTP packets;
[0075] Calculating a second time interval between the tail packet and the voice call end marker;
[0076] The voice quality of the target call record is evaluated according to the first time interval and the second time interval to obtain a voice quality evaluation result.
[0077] In this embodiment, the time when the calling side SBC receives the nth RTP packet is T(n), the time when the RTP tail packet is T(m), and the time when the first 200OK(BYE) message is T(m+1), then the time interval TR between RTPs and the time interval TB between the tail packet and the first 200OK(BYE) message are:
[0078] TR=T(n)-T(n-1)
[0079] TB=T(m+1)-T(m)
[0080] The time interval between RTP packets and voice call end signaling is used as the basis for voice quality assessment.
[0081] Optionally, the evaluating the voice quality of the target call record according to the first time interval and the second time interval to obtain a voice quality evaluation result specifically includes:
[0082] If any one of the first time intervals is greater than a preset first threshold or the second time interval is greater than a preset second threshold, the evaluation result of the voice quality is determined to be poor voice quality.
[0083] In this embodiment, since the call ending stage involves resource release and other processes, the process time is generally 500ms, so the time interval between the tail packet and the voice call ending mark is increased by 500ms compared to the time interval between RTP packets, so the first threshold is preferably 500ms, and the second threshold is preferably 1000ms. That is, if TR is greater than 500ms or TB is greater than 1000ms, it is considered that the voice quality is poor.
[0084] Through statistical data analysis, the addition of voice head and tail packet and signaling judgment rules has increased the accuracy of voice single-talk intermittent statistics from 92.66% to 96.35%, an increase of 3.69pp. The accuracy of voice quality assessment has been significantly improved compared to the case without adding head and tail packet and signaling judgment rules.
[0085] It should be noted that the voice quality assessment method provided in the embodiment of the present invention, by combining the voice call signaling process, constructs a comprehensive assessment model of voice call SIP signaling and RTP packets, gives priority to judging the validity of RTP packets, and thus performs voice quality assessment based on valid RTP packet data in combination with voice SIP signaling, thereby improving the accuracy of voice perception assessment judgment.
[0086] In a specific embodiment, the voice quality assessment method is applied to a voice quality assessment system. The system outputs a voice perception assessment result by evaluating and analyzing the association between RTP packets and voice call signaling for abnormal problems such as intermittent voice and single-talk. The system specifically includes a data acquisition module, a data storage module, a call list association analysis module, an RTP first packet identification module, an RTP tail packet identification module and an evaluation module. The modules are described as follows:
[0087] (1) Data collection module, which collects SIP signaling and RTP packet data corresponding to the voice generated by the user during the voice call through the core network, wherein the SIP signaling includes but is not limited to timestamp, SDP information, etc., and the RTP packet data includes but is not limited to packet sequence number, timestamp, SSRC (synchronous source), Payload Type, etc. Among them, SDP is used to associate SIP signaling with RTP packet data, and SSRC is used as a unique identifier for the user call.
[0088] (2) A data storage module, which is used to store the collected SIP signaling and RTP information related data.
[0089] (3) Call record association analysis module: This module analyzes the timestamp and SDP information carried in the SIP signaling, and matches and associates the information such as Payload Type in the SDP information with the timestamp and PayloadType value carried in the RTP packet data to obtain the correspondence between the SIP signaling and the RTP packet, and obtains the SIP signaling and RTP packet information of each call record.
[0090] (4) The RTP first packet identification module further confirms the first packet of the RTP packet by checking the SIP signaling and RTP packet information of the call record. Taking the calling side uplink as an example, Figure 2 The flowchart of RTP first packet identification is shown, and the ACK message sent by the calling side SBC after receiving the 200OK (invite) message sent by the calling terminal after the called party picks up the phone is used as the voice connection identifier. This time point is the time point of voice connection. After the voice is connected, the first RTP packet sent by the calling terminal received by the calling side SBC is taken as the first packet.
[0091] (5) The RTP tail packet identification module further confirms the tail packet of the RTP packet by checking the SIP signaling and RTP packet information of the call record. Taking the calling side uplink as an example, Figure 3 The flowchart of RTP tail packet identification is shown, with the first 200OK (BYE) message received by the calling side SBC as the end mark of the voice call. This time point is the time point when the voice ends. After the voice call ends, the last RTP packet sent by the calling terminal received by the calling side SBC is taken as the tail packet.
[0092] (6) An evaluation module evaluates the voice quality according to certain rules based on the obtained voice SIP signaling and RTP packet header and tail packet information, and outputs the evaluation result.
[0093] Take the calling side behavior as an example. Figure 4 The figure shows the time point diagram of the RTP first and last packets and the 200OK (BYE) message. The time when the calling side SBC receives the nth RTP packet is T(n), the time when the RTP last packet is T(m), and the time when the first 200OK (BYE) message is T(m+1). Then, the time interval TR between RTPs and the time interval TB between the last packet and the first 200OK (BYE) message are:
[0094] TR=T(n)-T(n-1)
[0095] TB=T(m+1)-T(m)
[0096] The time interval between RTP packets and the voice call end signaling is used as the basis for voice quality assessment. Since the call end stage involves resource release and other processes, the process time is generally 500ms. Therefore, the time interval between the tail packet and the first 200OK (BYE) message is increased by 500ms compared to the time interval between RTP packets. If TR is greater than 500ms or TB is greater than 1000ms, it is considered that the voice quality is poor.
[0097] Through statistical data analysis, the addition of voice head and tail packet and signaling judgment rules has increased the accuracy of voice single-talk intermittent statistics from 92.66% to 96.35%, an increase of 3.69pp. The accuracy of voice quality assessment has been significantly improved compared to the case without adding head and tail packet and signaling judgment rules.
[0098] It should be noted that, in order to solve the problem of inaccurate evaluation in the existing voice quality evaluation method, the present invention proposes a voice evaluation scheme based on RTP packets and voice call signaling. The scheme takes into account the signaling process of voice connection and hang-up, and judges the voice quality through SIP signaling and RTP packets, thereby achieving a more accurate and comprehensive evaluation.
[0099] Based on the above speech quality assessment system, such as Figure 5 As shown, the corresponding voice quality assessment method includes the following steps:
[0100] 1. The data collection module collects voice call SIP signaling and RTP packet data every day;
[0101] 2. Save the collected SIP signaling and RTP packet data to the data storage server;
[0102] 3. Complete the association between SIP signaling and RTP packet data to obtain different call records;
[0103] 4. Confirm the first RTP packet according to the associated call sheet data (i.e. confirm the first RTP packet according to the call sheet SIP signaling and RTP packet data time information);
[0104] 5. Confirm the tail of the RTP packet according to the associated call sheet data (i.e. confirm the RTP tail packet according to the call sheet SIP signaling and RTP packet data time information);
[0105] 6. Evaluate the time interval between RTP packets and the time interval between the tail packet and the voice call end marker, and output the evaluation results.
[0106] It should be noted that the voice quality assessment method provided by the embodiment of the present invention combines the voice RTP packet with the voice call signaling process to more accurately judge abnormal phenomena such as the discontinuity of the first and last voice packets. Compared with the traditional delay comparison between RTP packets, it is more in line with the user's perception of abnormal phenomena such as call discontinuity.
[0107] The voice quality evaluation method provided by the embodiment of the present invention first obtains the session initiation protocol SIP signaling corresponding to the target call record and the real-time transport protocol RTP packet data corresponding to the SIP signaling; then identifies the first packet and the last packet of the RTP packet according to the SIP signaling and the RTP packet data; then identifies the first packet, the last packet and all RTP packets between the first packet and the last packet in the RTP packet data as valid RTP packets; finally, the voice quality of the target call record is evaluated according to the time interval between the valid RTP packets and the time interval between the last packet and the voice call end mark, and the voice quality evaluation result is obtained. The present invention ensures that the voice quality evaluation is performed only based on the valid RTP packet data by giving priority to judging the validity of the RTP packet. At the same time, a comprehensive analysis is performed in combination with the time interval of the valid RTP packet and the time interval between the last packet and the voice call end mark, which significantly improves the accuracy and reliability of the voice quality evaluation, making the evaluation result closer to the actual call experience. The present invention solves the problem that the existing voice quality assessment method only assesses the voice quality based on the relationship between RTP packets, but ignores the validity of the RTP packets and the process judgment at the start and end of the call, thus causing inaccurate assessment results.
[0108] Embodiment 2:
[0109] like Figure 6 As shown, this embodiment provides a speech quality assessment device, which is used to perform the above-mentioned speech quality assessment method, including:
[0110] A call record data acquisition module 11 is used to acquire the session initiation protocol SIP signaling corresponding to the target call record and the real-time transport protocol RTP packet data corresponding to the SIP signaling;
[0111] A head and tail packet identification module 12, connected to the call bill data acquisition module 11, is used to identify the head and tail packets of the RTP packet according to the SIP signaling and the RTP packet data;
[0112] A valid RTP packet identification module 13, connected to the first and last packet identification module 12, is used to identify the first packet, the last packet and all RTP packets between the first packet and the last packet in the RTP packet data as valid RTP packets;
[0113] The voice quality assessment module 14 is connected to the valid RTP packet identification module 13, and is used to evaluate the voice quality of the target call record according to the time interval between the valid RTP packets and the time interval between the tail packet and the voice call end mark to obtain a voice quality assessment result.
[0114] Optionally, the call record data acquisition module 11 includes:
[0115] A data collection unit is used to collect SIP signaling and RTP packet data generated by the user during a voice call;
[0116] A data matching unit, used for matching the collected SIP signaling and RTP packet data according to timestamp information and payload type to obtain a corresponding relationship between the SIP signaling and the RTP packet;
[0117] The call record data acquisition unit is used to acquire the SIP signaling corresponding to the target call record and the RTP packet data corresponding to the SIP signaling according to the corresponding relationship between the SIP signaling and the RTP packet.
[0118] Optionally, the collected SIP signaling includes a timestamp and session description protocol SDP information, and the collected RTP packet data includes a sequence number, a timestamp, a synchronization source SSRC, and a payload type.
[0119] Optionally, the first and last packet identification module 12 includes:
[0120] A connection time point determining unit, used to determine the time point of the first message in the SIP signaling as a voice connection indicator as the time point of the voice connection;
[0121] A first packet identification unit, used for identifying the first RTP packet after the voice connection time point in the RTP packet data as the first packet;
[0122] An end time point determination unit, used to determine the time point of the second message in the SIP signaling as an indication of the end of the voice call as the time point of the voice end;
[0123] The tail packet identification unit is used to identify the last RTP packet before the voice end time point in the RTP packet data as the tail packet.
[0124] Optionally, the voice quality assessment module 14 includes:
[0125] A first calculation unit, configured to calculate a first time interval between any two consecutive RTP packets in the valid RTP packets;
[0126] A third calculation unit, used to calculate a second time interval between the tail packet and the voice call end mark;
[0127] An evaluation and judgment unit is used to evaluate the voice quality of the target call record according to the first time interval and the second time interval to obtain an evaluation result of the voice quality.
[0128] Optionally, the evaluation and judgment unit is specifically used to:
[0129] If any one of the first time intervals is greater than a preset first threshold or the second time interval is greater than a preset second threshold, the evaluation result of the voice quality is determined to be poor voice quality.
[0130] Optionally, the first threshold is 500 ms and the second threshold is 1000 ms.
[0131] Embodiment 3:
[0132] refer to Figure 7 This embodiment provides a speech quality assessment device, including a memory 21 and a processor 22. The memory 21 stores a computer program, and the processor 22 is configured to run the computer program to execute the speech quality assessment method in Example 1.
[0133] The memory 21 is connected to the processor 22. The memory 21 may be a flash memory, a read-only memory or other memory. The processor 22 may be a central processing unit or a single-chip microcomputer.
[0134] Embodiment 4:
[0135] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the speech quality assessment method in the above-mentioned embodiment 1 is implemented.
[0136] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules or other data). Computer-readable storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory), flash memory or other memory technology, CD-ROM (Compact Disc Read-Only Memory), digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0137] In summary, the voice quality evaluation method, device and readable storage medium provided by the embodiment of the present invention first obtain the session initiation protocol SIP signaling corresponding to the target call record and the real-time transport protocol RTP packet data corresponding to the SIP signaling; then identify the first packet and the last packet of the RTP packet according to the SIP signaling and the RTP packet data; then identify the first packet, the last packet and all RTP packets between the first packet and the last packet in the RTP packet data as valid RTP packets; finally, evaluate the voice quality of the target call record according to the time interval between the valid RTP packets and the time interval between the last packet and the voice call end mark, and obtain the voice quality evaluation result. The present invention ensures that the voice quality evaluation is performed only based on the valid RTP packet data by giving priority to judging the validity of the RTP packet. At the same time, the comprehensive analysis is combined with the time interval of the valid RTP packet and the time interval between the last packet and the voice call end mark, which significantly improves the accuracy and reliability of the voice quality evaluation, making the evaluation result closer to the actual call experience. The present invention solves the problem that the existing voice quality assessment method only assesses the voice quality based on the relationship between RTP packets, but ignores the validity of the RTP packets and the process judgment at the start and end of the call, thus causing inaccurate assessment results.
[0138] It is to be understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of the present invention, but the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered to be within the scope of protection of the present invention.
Claims
1. A method for evaluating speech quality, characterized in that: The method comprises: Acquire the Session Initiation Protocol (SIP) signaling corresponding to the target call record and the Real-time Transport Protocol (RTP) packet data corresponding to the SIP signaling; Identify the first packet and the last packet of the RTP packet according to the SIP signaling and the RTP packet data; Identify the first packet, the last packet, and all RTP packets between the first packet and the last packet in the RTP packet data as valid RTP packets; The voice quality of the target call record is evaluated according to the time interval between the valid RTP packets and the time interval between the tail packet and the voice call end identifier to obtain a voice quality evaluation result.
2. The method according to claim 1, characterized in that The obtaining of the session initiation protocol SIP signaling corresponding to the target call record and the real-time transport protocol RTP packet data corresponding to the SIP signaling specifically includes: Collect SIP signaling and RTP packet data generated by users during voice calls; By matching the collected SIP signaling and RTP packet data according to timestamp information and payload type, a corresponding relationship between the SIP signaling and the RTP packet is obtained; The SIP signaling corresponding to the target call record and the RTP packet data corresponding to the SIP signaling are obtained according to the corresponding relationship between the SIP signaling and the RTP packet.
3. The method according to claim 2, characterized in that The collected SIP signaling includes a timestamp and session description protocol SDP information, and the collected RTP packet data includes a sequence number, a timestamp, a synchronization source SSRC, and a payload type.
4. The method according to claim 1, characterized in that The step of identifying the first packet and the last packet of the RTP packet according to the SIP signaling and the RTP packet data specifically includes: Determine the time point of the first message in the SIP signaling as the time point of the voice connection; Identify the first RTP packet after the voice connection time point in the RTP packet data as the first packet; Determine the time point of the second message in the SIP signaling as the time point of the voice call end; The last RTP packet before the voice end time point in the RTP packet data is identified as the tail packet.
5. The method according to claim 1, characterized in that The step of evaluating the voice quality of the target call record according to the time interval between the valid RTP packets and the time interval between the tail packet and the voice call end identifier to obtain a voice quality evaluation result specifically includes: Calculate a first time interval between any two consecutive RTP packets in the valid RTP packets; Calculating a second time interval between the tail packet and the voice call end marker; The voice quality of the target call record is evaluated according to the first time interval and the second time interval to obtain a voice quality evaluation result.
6. The method according to claim 5, characterized in that The step of evaluating the voice quality of the target call record according to the first time interval and the second time interval to obtain a voice quality evaluation result specifically includes: If any one of the first time intervals is greater than a preset first threshold or the second time interval is greater than a preset second threshold, the evaluation result of the voice quality is determined to be poor voice quality.
7. The method according to claim 6, characterized in that The first threshold is 500ms, and the second threshold is 1000ms.
8. A speech quality assessment device, characterized in that: The device comprises: A call record data acquisition module, used to acquire the session initiation protocol SIP signaling corresponding to the target call record and the real-time transport protocol RTP packet data corresponding to the SIP signaling; A head and tail packet identification module, connected to the call bill data acquisition module, for identifying the head and tail packets of the RTP packet according to the SIP signaling and the RTP packet data; A valid RTP packet identification module, connected to the first and last packet identification modules, for identifying the first packet, the last packet and all RTP packets between the first packet and the last packet in the RTP packet data as valid RTP packets; The voice quality assessment module is connected to the valid RTP packet identification module and is used to evaluate the voice quality of the target call record according to the time interval between the valid RTP packets and the time interval between the tail packet and the voice call end mark to obtain a voice quality assessment result.
9. A speech quality assessment device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the speech quality assessment method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the speech quality assessment method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
System for monitoring VOIP voice quality based on SIP protocol and detection method thereof
CN101437032A
Method and device for identifying RTP tail packet loss
CN109587096A
VoLTE voice quality dial test analysis method and system
CN110401965A
RTP packet loss detection method, device and equipment and computer readable storage medium
CN112688824A
Managing early media for communication sessions establishing via the session initiation protocol (SIP)
WO2013138198A1