A method, apparatus, device and storage medium for evaluating voice call quality

By acquiring and processing air interface and base station log information, and using a voice quality assessment model to evaluate voice call quality, the problems of subjective factors and hardware limitations are solved, achieving a more accurate and cost-effective assessment.

CN118098284BActive Publication Date: 2025-11-18RUIJIE NETWORKS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410122955.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-11-18
Estimated Expiration
2044-01-29

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as significant subjective influence, numerous software and hardware limitations, and high costs when assessing voice call quality.

Method used

By acquiring air interface log information and/or base station log information during voice calls, and using a voice quality assessment model to segment and fuse this log information, segment logs are obtained and scored, thereby determining the quality of voice calls and reducing reliance on user subjective awareness and hardware devices.

Benefits of technology

It enables more accurate and effective evaluation of voice call quality, reduces evaluation costs, minimizes the impact of time factors, and avoids reliance on dedicated equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118098284B_ABST
    Figure CN118098284B_ABST
Patent Text Reader

Abstract

The application provides a method, device, equipment and storage medium for evaluating voice call quality, comprising: acquiring air interface log information and / or base station log information collected in a voice call process for a voice receiver; the air interface log information records parameter information representing communication quality at an air interface corresponding to the voice receiver in the voice call process; the base station log information records parameter information representing communication quality at a base station corresponding to the voice receiver in the voice call process; dividing the air interface log information and / or the base station log information according to time slices to obtain a plurality of segment logs; inputting the plurality of segment logs into a voice quality evaluation model to obtain voice call quality of the voice receiver in the voice call process. The scheme can accurately and effectively determine the quality of the voice call environment, and makes the voice quality score not affected by the subjective consciousness of the user; without using a voice call quality evaluation instrument, the cost and hardware limitations are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a method, apparatus, device and storage medium for evaluating the quality of voice calls. Background Technology

[0002] Traditional methods for evaluating voice call quality include subjective testing and using specially customized terminals combined with voice call quality assessment instruments. Subjective testing involves surveying and quantifying user behavior in receiving and perceiving voice call quality. Different users subjectively compare the original standard voice with the degraded sound transmitted over a wireless network, assigning a voice call quality score. However, this method is influenced by user subjectivity. Using specially customized terminals combined with voice call quality assessment instruments and corresponding drive-testing software changes the subjective evaluation method, reducing the influence of user subjectivity. However, due to numerous software and hardware limitations and high license fees, the scenarios for voice call quality assessment are limited.

[0003] Therefore, there is a need for a method that can efficiently evaluate voice call quality by reducing the influence of subjective factors, minimizing software and hardware limitations, and lowering evaluation costs. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for evaluating voice call quality, which can reduce subjective factors, reduce software and hardware limitations, and reduce evaluation costs to efficiently evaluate voice call quality.

[0005] Firstly, embodiments of this application provide a method for evaluating voice call quality. This method can be executed by a device for evaluating voice call quality, which can be a terminal device or a module for a terminal device, or a server or a module for a server. This application does not limit the executing entity of the method. The method includes: acquiring air interface log information and / or base station log information collected during a voice call for the voice receiver; the air interface log information records parameter information characterizing communication quality at the air interface corresponding to the voice receiver during the voice call; the base station log information records parameter information characterizing communication quality at the base station corresponding to the voice receiver during the voice call; dividing the air interface log information and / or base station log information into time slices to obtain multiple segment logs; inputting the multiple segment logs into a voice quality evaluation model to obtain the voice call quality of the voice receiver during the voice call; wherein, the voice quality evaluation model is obtained by training on sample segment logs and the voice quality scores corresponding to the sample segment logs.

[0006] The above solution has several advantages. First, it uses air interface log information and / or base station log information, rather than core network data, making data acquisition easier. Second, it enables accurate and effective determination of voice quality scores through air interface log information and / or base station log information, thereby accurately and effectively determining the quality of the current voice call environment. Third, it uses a voice quality assessment model to determine the voice quality score, making the score unaffected by the user's subjective perception. Furthermore, it eliminates the need for voice call quality assessment instruments, saving costs and hardware limitations.

[0007] In one possible implementation, multiple segment logs are input into a speech quality assessment model to obtain a speech quality score corresponding to each segment log; based on the speech quality score corresponding to each segment log, the speech call quality of the speech receiver during the speech call is determined.

[0008] The above scheme determines the voice call quality of the voice receiver during the voice call by using voice quality scores corresponding to multiple segment logs. This reduces the impact of time factors and makes the assessed voice call quality more accurate and effective.

[0009] In one possible implementation, the air interface log information and / or base station log information are segmented according to the duration corresponding to the scoring cycle of the voice quality assessment instrument to obtain multiple log segments.

[0010] The above scheme can obtain multiple accurate and effective segment logs, which are then input into the speech quality assessment model to obtain multiple accurate and effective speech quality scores, thereby accurately and effectively determining the quality of the current speech call environment.

[0011] In one possible implementation, the air interface log information is fragmented to obtain multiple first fragment logs; the base station log information is fragmented to obtain multiple second fragment logs; for the first fragment logs and second fragment logs belonging to the same time period, the parameter information corresponding to the same parameters in the first fragment logs and the second fragment logs is fused to obtain the fused parameter information, thereby obtaining the fragment logs corresponding to the time period.

[0012] The above solution integrates information with the same parameters, which can reduce data redundancy and resource waste, making the corresponding log segments more concise and effective.

[0013] In one possible implementation, the parameters included in the air interface log information and / or base station log information include one or more of the following: voice signal strength, signal-to-noise ratio, voice coding method, or transmission delay.

[0014] The above scheme inputs the above-mentioned multiple parameter information into the voice quality assessment model, which can obtain an accurate and effective voice quality score, thereby accurately and effectively determining the quality of the current voice call environment.

[0015] In one possible implementation, if the voice call quality is less than a quality threshold, the air interface log information and / or base station log information are reported to the management platform; and / or, the air interface log information and / or base station log information are input into the voice quality detection model to determine the factors affecting the voice call quality of the voice receiver.

[0016] The above scheme helps to identify the factors affecting the quality of voice calls for the recipient, and then to address these factors.

[0017] In one possible implementation, the following steps are taken: First, a sample voice sender and at least one sample voice receiver are identified. Second, the received voice recorded by at least one sample voice receiver during a sample call is acquired, along with sample air interface log information and / or sample base station log information corresponding to that receiver during the call. Third, the recorded received voice is sent to a voice quality assessment instrument, which determines the voice quality score for each of the at least one sample voice receivers. Fourth, the sample air interface log information and / or sample base station log information corresponding to each receiver are segmented according to the scoring period of the voice quality assessment instrument to obtain sample segment logs. Fifth, the sample segment logs and their corresponding voice quality scores are used as training samples to train a voice quality assessment model.

[0018] The above scheme enables the trained speech quality assessment model to accurately and effectively determine the speech quality score, thereby achieving an accurate and effective determination of the quality of the current speech call environment.

[0019] Secondly, embodiments of this application provide an apparatus for evaluating voice call quality, comprising: an acquisition unit and a processing unit. The acquisition unit is used to acquire air interface log information and / or base station log information collected during a voice call for the voice receiver; the air interface log information records parameter information characterizing communication quality at the air interface corresponding to the voice receiver during the voice call; the base station log information records parameter information characterizing communication quality at the base station corresponding to the voice receiver during the voice call; the processing unit is used to divide the air interface log information and / or base station log information into time slices to obtain multiple segment logs; input the multiple segment logs into a voice quality evaluation model to obtain the voice call quality of the voice receiver during the voice call; wherein, the voice quality evaluation model is obtained by training on sample segment logs and the voice quality scores corresponding to the sample segment logs.

[0020] In one possible implementation, a processing unit is used to input multiple segment logs into a speech quality assessment model to obtain a speech quality score corresponding to each segment log; and to determine the speech call quality of the speech receiver during the speech call based on the speech quality score corresponding to each segment log.

[0021] In one possible implementation, the processing unit is used to segment the air interface log information and / or base station log information according to the duration corresponding to the scoring cycle of the voice quality assessment instrument, so as to obtain multiple segment logs.

[0022] In one possible implementation, a processing unit is used to segment the air interface log information to obtain multiple first segment logs; segment the base station log information to obtain multiple second segment logs; and for the first segment logs and second segment logs belonging to the same time period, fuse the parameter information corresponding to the same parameters in the first segment logs and the second segment logs to obtain the fused parameter information, thereby obtaining the segment logs corresponding to the time period.

[0023] In one possible implementation, the parameters included in the air interface log information and / or base station log information include one or more of the following: voice signal strength, signal-to-noise ratio, voice coding method, or transmission delay.

[0024] In one possible implementation, the processing unit is configured to report air interface log information and / or base station log information to the management platform if the voice call quality is less than a quality threshold; and / or input the air interface log information and / or base station log information into the voice quality detection model to determine the factors affecting the voice call quality of the voice receiver.

[0025] In one possible implementation, a processing unit is used to determine the sample voice sender and at least one sample voice receiver; an acquisition unit is used to acquire the received voice recorded by at least one sample voice receiver during the sample call, and to collect the sample air interface log information and / or sample base station log information corresponding to at least one sample voice receiver during the sample call; the processing unit is used to send the recorded received voice to a voice quality assessment instrument, which determines the voice quality score of at least one sample voice receiver; the sample air interface log information and / or sample base station log information corresponding to each sample voice receiver are segmented according to the scoring period of the voice quality assessment instrument to obtain sample segment logs; and the sample segment logs and the corresponding voice quality scores are used as training samples to train a voice quality assessment model.

[0026] Thirdly, embodiments of this application also provide a computing device, including:

[0027] Memory, used to store program instructions;

[0028] The processor is used to call program instructions stored in memory and execute any method that implements the first aspect described above according to the obtained program instructions.

[0029] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-readable instructions, which, when read and executed by a computer, implement any of the methods described in the first aspect.

[0030] Fifthly, embodiments of this application provide a computer program product, including a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform any of the methods described in the first aspect. Attached Figure Description

[0031] Figure 1 A system architecture diagram for evaluating voice call quality is provided in the embodiments of this application;

[0032] Figure 2 A flowchart illustrating a method for evaluating voice call quality provided in an embodiment of this application;

[0033] Figure 3 A flowchart illustrating a method for training a speech quality assessment model provided in an embodiment of this application;

[0034] Figure 4 A network architecture diagram for collecting training samples using MOS boxes is provided in an embodiment of this application;

[0035] Figure 5 A flowchart illustrating a method for collecting training samples provided in an embodiment of this application;

[0036] Figure 6 A flowchart illustrating an evaluation process for a calling terminal and a called terminal provided in an embodiment of this application;

[0037] Figure 7 A schematic diagram of a neural network structure for a multilayer perceptron (MLP) provided in an embodiment of this application;

[0038] Figure 8 A schematic diagram of a device for evaluating voice call quality provided in an embodiment of this application;

[0039] Figure 9 This is a schematic diagram of a device for evaluating voice call quality provided in an embodiment of this application. Detailed Implementation

[0040] Figure 1The system architecture diagram for evaluating voice call quality provided in this application embodiment includes multiple terminals, drive test software, and a base station. The multiple terminals are used for voice calls, the drive test software is used to collect air interface log information and determine the voice initiator and voice receiver, and the base station is used to maintain the voice call and collect base station log information.

[0041] Figure 2 This is a flowchart illustrating a method for evaluating voice call quality according to an embodiment of this application. The method can be executed by a device for evaluating voice call quality, which can be a terminal device or a module for a terminal device, or a server or a module for a server. This application does not limit the entity executing this method.

[0042] The method includes the following steps:

[0043] Step 201: Obtain air interface log information and / or base station log information collected during the voice call for the voice receiver.

[0044] Among them, the air interface log information records the parameter information that represents the communication quality at the air interface corresponding to the voice receiver during the voice call; the base station log information records the parameter information that represents the communication quality at the base station corresponding to the voice receiver during the voice call.

[0045] In one possible implementation, the air interface log information can also be the drive test log information.

[0046] In one possible implementation, the air interface log information and / or base station log information include one or more of the following parameters: voice signal strength, signal-to-noise ratio, voice coding method, or transmission delay. This scheme inputs these multiple parameters into a voice quality assessment model to obtain an accurate and effective voice quality score, thereby accurately and effectively determining the quality of the current voice call environment.

[0047] Step 202: Divide the air interface log information and / or base station log information into time slices to obtain multiple log segments.

[0048] In one possible implementation, the air interface log information and / or base station log information are segmented according to the duration corresponding to the scoring period of the voice quality assessment instrument, resulting in multiple segment logs. The voice quality assessment instrument can be a Mean Opinion Score (MOS) box, or other types of voice quality assessment instruments; this application is not limited to this. In a communication network configured with a voice quality assessment instrument, the instrument periodically evaluates the voice call quality during a voice call. For example, the scoring period of the voice quality assessment instrument corresponds to 8 seconds, but other durations are also possible; this application is not limited to this. This scheme can obtain multiple accurate and effective segment logs, which are then input into a voice quality assessment model to obtain multiple accurate and effective voice quality scores, thereby accurately and effectively determining the quality of the current voice call environment.

[0049] In one possible implementation, the air interface log information is fragmented to obtain multiple first-segment logs; the base station log information is fragmented to obtain multiple second-segment logs; for the first-segment and second-segment logs belonging to the same time period, the parameter information corresponding to the same parameters in the first-segment and second-segment logs is fused to obtain the fused parameter information, thereby obtaining the segment log corresponding to the time period. This scheme, by fusing the same parameter information, can reduce data redundancy and resource waste, making the corresponding segment logs more concise and effective.

[0050] Step 203: Input multiple segment logs into the voice quality assessment model to obtain the voice call quality of the voice receiver during the voice call.

[0051] The speech quality assessment model is trained using sample segment logs and their corresponding speech quality scores.

[0052] One possible implementation involves inputting multiple log segments into a speech quality assessment model to obtain a speech quality score for each segment. Based on these scores, the speech quality of the recipient during the call is determined. This approach, by using speech quality scores from multiple log segments to determine the recipient's speech quality, reduces the impact of time factors, resulting in a more accurate and effective speech quality assessment.

[0053] One possible implementation method is to determine the voice call quality of the voice receiver during the voice call based on the average of the voice quality scores corresponding to multiple segment logs.

[0054] In another possible implementation, the voice call quality of the voice receiver during the voice call is determined based on the weights of the voice quality scores corresponding to multiple segment logs.

[0055] In one possible implementation, the speech quality assessment model is trained using one of the following models: Multilayer Perceptron (MLP), Decision Tree, Support Vector Machine, Gaussian, K-Nearest Neighbor, Hidden Markov Model, Conditional Random Field, etc. It can also be trained by combining multiple of the above models. This application does not limit the type of speech quality assessment model.

[0056] In one possible implementation, if the voice call quality is below a quality threshold, air interface log information and / or base station log information are reported to a management platform for analysis by administrators to determine the factors affecting the voice call quality of the voice receiver; and / or, the air interface log information and / or base station log information are input into a voice quality detection model, and the factors affecting the voice call quality of the voice receiver are determined based on the output of the voice quality detection model. The voice quality detection model is trained on sample segment logs and the corresponding factors affecting voice call quality. This application does not limit the type of voice quality detection model. This scheme helps to determine the factors affecting the voice call quality of the voice receiver and then process these factors.

[0057] The above solution has several advantages. First, it uses air interface log information and / or base station log information, rather than core network data, making data acquisition easier. Second, it enables accurate and effective determination of voice quality scores through air interface log information and / or base station log information, thereby accurately and effectively determining the quality of the current voice call environment. Third, it uses a voice quality assessment model to determine the voice quality score, making the score unaffected by the user's subjective perception. Furthermore, it eliminates the need for voice call quality assessment instruments, saving costs and hardware limitations.

[0058] In steps 201 to 203 above, since the two parties in a voice call are generally engaged in a dialogue, the voice quality assessment model can be used to evaluate the voice call quality of both the voice initiator and the voice receiver. For example, the air interface log information and / or base station log information of the voice initiator are input into the voice quality assessment model to determine the voice call quality of the voice initiator; the air interface log information and / or base station log information of the voice receiver are input into the voice quality assessment model to determine the voice call quality of the voice receiver; alternatively, the air interface log information and / or base station log information of both the voice initiator and the voice receiver can be input into the voice quality assessment model to determine the voice call quality of the voice initiator; the air interface log information and / or base station log information of both the voice initiator and the voice receiver can be input into the voice quality assessment model to determine the voice call quality of the voice receiver; furthermore, the air interface log information and / or base station log information of both the voice initiator and the voice receiver can be input into the voice quality assessment model simultaneously to determine the voice call quality of both the voice initiator and the voice receiver.

[0059] In one possible implementation, in step 203 above, the speech quality assessment model is trained using sample segment logs and their corresponding speech quality scores; wherein the method for training the speech quality assessment model is as follows: Figure 3 As shown, the method includes the following steps:

[0060] Step 301: Determine the sender of the sample voice and at least one receiver of the sample voice.

[0061] Step 302: Obtain the received voice recorded by at least one sample voice receiver during the sample call, and collect the sample air interface log information and / or sample base station log information corresponding to at least one sample voice receiver during the sample call.

[0062] Step 303: The recorded received voice is sent to the voice quality assessment instrument, which determines the voice quality score of at least one sample voice receiver.

[0063] Step 304: Divide the sample air interface log information and / or sample base station log information corresponding to each sample voice receiver into segments according to the scoring cycle of the voice quality assessment instrument to obtain sample segment logs.

[0064] Step 305: Use the sample segment logs and the corresponding speech quality scores as training samples to train the speech quality assessment model.

[0065] The above scheme enables the trained speech quality assessment model to accurately and effectively determine the speech quality score, thereby achieving an accurate and effective determination of the quality of the current speech call environment.

[0066] The following section uses the MOS box as an example to explain in detail how to obtain a speech quality assessment model by training it on sample segment logs and the corresponding speech quality scores.

[0067] One possible implementation involves a network architecture that uses MOS boxes to collect training samples, such as... Figure 4 As shown, where, Figure 4 The system includes two terminals, test phone 1 and test phone 2, as well as drive test software and a MOS box. The drive test software is used to control the terminals to make and receive calls and collect drive test logs (air interface logs). The MOS box is used to control the call content and to evaluate the voice call quality. The base station is used to collect base station log information, which includes wireless environment information reported by the terminals and user plane data during the call.

[0068] One possible implementation method is to use MOS boxes to collect training samples, such as... Figure 5 As shown, the method includes the following steps:

[0069] Step 501: The MOS box determines the original audio of the call content.

[0070] In one possible implementation, before the MOS box determines the raw audio of the call content, the drive test software identifies the calling terminal and the called terminal. The calling terminal can be referred to as the voice initiator, and the called terminal can be referred to as the voice receiver. The MOS box then sends the raw audio of the call content to the calling terminal identified by the drive test software.

[0071] In one possible implementation, the road test software first determines that test mobile phone 1 is the calling terminal and test mobile phone 2 is the called terminal; then, it determines that test mobile phone 1 is the called terminal and test mobile phone 2 is the calling terminal in the second instance.

[0072] One possible implementation involves selecting audio from an existing speech evaluation corpus as the original audio of the call content. This results in more standardized original audio, enabling accurate and effective comparison between the original audio and the recorded audio, thus making the speech quality score determined by the MOS box more accurate.

[0073] Step 502: The MOS box sends the original audio of the call to the calling terminal.

[0074] In one possible implementation, the MOS box plays a standard audio file to the microphone 1 (MIC1) port of the test mobile phone 1 via line output 1 (LINE OUT1). The function of LINE OUT is to output the analog signal processed by the sound card to audio devices such as speakers through the line output interface.

[0075] In one possible implementation, the calling terminal sends the call content to the called terminal via a microphone, and the evaluation process for both the calling and called terminals is as follows: Figure 6 As shown. Within one cycle, the called terminal records audio starting at second 0 for 12 seconds; starting at second 12, it plays 8 seconds of second original audio, which is then recorded by the calling terminal; the calling terminal plays 8 seconds of first original audio starting at second 2, which is then recorded by the called terminal; audio recording begins again at second 8 and continues for 12 seconds. Of course, the evaluation process for the calling and called terminals can also be in other ways, and this application does not limit this to any particular method.

[0076] In one possible implementation, the called terminal records the received call content, obtains the recorded audio, and sends the recorded audio to the MOS box.

[0077] Step 503: The MOS box receives the recorded audio sent by the called terminal.

[0078] In one possible implementation, the test phone 2 receives the audio signal and plays it back to the MOS box's line input 2 (LINE IN2) via the speaker 2 port. Here, LINE IN refers to the audio line input.

[0079] In one possible implementation, test phone 2 and test phone 1 perform caller and receiver tests. The MOS box plays audio to the MIC2 port of test phone 2 via LINEOUT2. After receiving the audio, test phone 1 plays standard audio to the LINE IN1 port of the MOS box via SPEAKER1. The program compares the sample file and calculates the score as the voice value of test phone 1.

[0080] Step 504: The MOS box compares the original audio with the recorded audio to determine the voice quality score.

[0081] One possible implementation method uses the MOS (Mean Offset Syndrome) scoring criteria, as shown in Table 1 below. Table 1 shows that a higher MOS value indicates higher voice call quality. Generally, a MOS value of 4 or higher is considered good voice call quality. If the MOS value is below 3.5, most calls will not be of satisfactory quality.

[0082] Table 1

[0083] audio level MOS value Evaluation criteria excellent 4.0-5.0 Very good, I can hear it clearly; the latency is low, and the communication is smooth. good 3.5-4.0 Slightly worse, but still clear; low latency, but communication is not smooth and there is some background noise. middle 3.0-3.5 It's okay, but the sound isn't very clear; there's a slight delay, but communication is possible. Difference 1.5-3.0 Barely audible; significant delay, requiring multiple repetitions for communication. inferior 0-1.5 Extremely poor quality, incomprehensible; significant latency, hindering communication.

[0084] In one possible implementation, the base station collects caller and called party terminal logs during the call. Drive test logs can be collected from terminal logs, while voice-related logs can be extracted via packet capture to broaden the evaluation scope and reduce the cost of voice quality assessment. During training data collection, a MOS instrument and a test terminal or a test terminal with a built-in sound card are needed to collect MOS data and obtain voice quality values.

[0085] In one possible implementation, the air interface log information includes one or more of the following: SS-RSRP, SS-SINR, PDSCH DMRS SINR, MCSDL, MCS UL, CQI, PDSCH bler, PUSCH bler, PUSCH Pathloss, Jitter, RTP Loss Rate, RTP Delay, POLQA Score SWB.Among them, SS-RSRP (Synchronization Signal Reference Signal Received Power) refers to the received power of the synchronization reference signal; SS-SINR (Synchronization Signal Signal-to-Interference plus Noise Ratio) refers to the signal-to-noise ratio of the synchronization signal; PDSCH DMRS SINR (Physical Downlink Shared Channel Demodulation Reference Signal Signal-to-Interference plus Noise Ratio) refers to the signal-to-noise ratio of the physical downlink shared channel reference signal; MCSDL (Modulation and Coding Scheme Downlink) refers to the downlink modulation and coding scheme; MCS UL (Modulation and Coding Scheme Uplink) refers to the uplink modulation and coding scheme; CQI (Channel Quality Indicator) refers to the channel quality indicator; PDSCH BLER (Physical Downlink Shared Channel Block Error Rate) refers to the physical downlink shared channel block error rate; PUSCH BLER (Physical Uplink Shared Channel Block Error Rate) refers to the physical uplink shared channel block error rate; and PUSCH Pathloss (Physical Uplink Shared Channel Path) refers to the physical uplink shared channel block error rate. Loss refers to the physical uplink shared channel path loss; Jitter refers to jitter; RTP Loss Rate (Real-time Transport Protocol Loss Rate) refers to the packet loss rate of the real-time transport protocol; RTP Delay (Real-time Transport Protocol Delay) refers to the delay of the real-time transport protocol; POLQA Score SWB (Perceptual Objective Listening Quality Analysis Score Super Wideband) refers to the broadband voice multimode audio quality.

[0086] In one possible implementation, the base station log information includes terminal-side log information and cell-level log information. The terminal-side log information includes one or more of the following: Sounding Reference Signal Received Power (SRS / RSRP) and Physical Uplink Shared Channel Signal-to-Interference plus Noise Ratio (PUSCH / SINR). SRS (Sounding Reference Signal) refers to the uplink reference signal, while RSRP (Reference Signal Received Power) represents the power of the received signal. SRS / RSRP indicates the power of the received uplink reference signal in a 5G network. This parameter is commonly used to evaluate the strength of the uplink signal received by the device and is an important indicator for evaluating signal coverage and signal quality. PUSCH (Physical Uplink Shared Channel) refers to the physical uplink shared channel, while SINR (Signal-to-Interference plus Noise Ratio) represents the signal-to-interference plus noise ratio. PUSCH SINR represents the signal-to-interference-plus-noise ratio (SNR) of the physical uplink shared channel in a 5G network. This parameter is used to evaluate the quality of the uplink channel, where signal quality has a significant impact on the reliability and speed of data transmission.

[0087] Cell-level log information includes one or more of the following: Download Physical Resource Block Average Utilization (DL_PRB_AVG_UTIL), Upload Physical Resource Block Average Utilization (UL_PRB_AVG_UTIL), and Voice over New Radio Number (VONR num). Here, PRB (Physical Resource Block) refers to the physical resource block, which is the basic unit used to describe radio resource allocation in Long Term Evolution (LTE) systems. It contains one or more consecutive subframes in the time domain and one or more consecutive carrier frequencies in the frequency domain, used to carry user data and control information. The allocation and configuration of PRBs have a significant impact on the performance and resource utilization of the LTE system. DL (Download) means downloading, UL (Upload) means uploading, AVG_UTIL (Average Utilization) means average utilization, VoNR (Voice over New Radio) means voice over new radio, and 5G VoNR means providing high-definition voice and video services on a standalone 5G network.

[0088] In one possible implementation, this application employs a multilayer perceptron (MLP) to train the speech quality assessment model. The perceptron is a single-neuron model, a precursor to larger neural networks. The power of neural networks lies in their ability to learn representations in training data and how to relate them to the output variables desired for prediction. Mathematically, they can learn any mapping function and have proven to be a general approximation algorithm.

[0089] In one possible implementation, this application employs a neural network structure of a multilayer perceptron (MLP), such as... Figure 7 As shown, it includes one input layer, three hidden layers, and one output layer. Of course, the neural network structure of a multilayer perceptron (MLP) can also be of other types, and this application does not limit it.

[0090] In one possible implementation, during training, the method for determining whether the MLP model has converged is as follows: if the mean squared error between the predicted and actual voice call quality values ​​is below a threshold, then training the prediction model is stopped, and a converged prediction model is obtained. Specifically, the smaller the mean squared error, the higher the similarity between the predicted and actual voice call quality values, the closer the predicted voice call quality value is to the actual voice call quality value, and the more ideal the training result of the prediction model. When the mean squared error between the predicted and actual voice call quality values ​​is less than the threshold, the prediction model is determined to have converged, and a converged prediction model is obtained.

[0091] In another possible implementation, the method for determining whether the MLP model has converged is as follows: input the training set into the initial prediction model and train the initial prediction model to obtain a prediction model under training; input the validation set data into the prediction model under training to obtain the predicted voice call quality value; compare the predicted voice call quality value with the actual voice call quality value corresponding to the data in the validation set; if it is determined that the mean square error between the predicted voice call quality value and the actual voice call quality value is less than or equal to a threshold, then stop training the prediction model under training to obtain a prediction model that has been trained to convergence; if it is determined that the mean square error between the predicted voice call quality value and the actual voice call quality value is greater than a threshold, then train the prediction model under training until it is determined that the sum of squared residuals between the predicted voice call quality value and the actual voice call quality value is greater than a threshold.

[0092] Based on the same technological concept Figure 8 An exemplary illustration is provided by an embodiment of this application for evaluating voice call quality. For example... Figure 8 As shown, it includes: an acquisition unit 801 and a processing unit 802. The acquisition unit 801 is used to acquire air interface log information and / or base station log information collected during a voice call for the voice receiver; the air interface log information records parameter information representing communication quality at the air interface corresponding to the voice receiver during the voice call; the base station log information records parameter information representing communication quality at the base station corresponding to the voice receiver during the voice call; the processing unit 802 is used to divide the air interface log information and / or base station log information into time slices to obtain multiple segment logs; inputting the multiple segment logs into a voice quality evaluation model to obtain the voice call quality of the voice receiver during the voice call; wherein, the voice quality evaluation model is obtained by training on sample segment logs and the corresponding voice quality scores of the sample segment logs.

[0093] In one possible implementation, the processing unit 802 is used to input multiple segment logs into the speech quality assessment model to obtain a speech quality score corresponding to each segment log; and to determine the speech call quality of the speech receiver during the speech call based on the speech quality score corresponding to each segment log.

[0094] In one possible implementation, the processing unit 802 is used to segment the air interface log information and / or base station log information according to the duration corresponding to the scoring cycle of the voice quality assessment instrument, so as to obtain multiple segment logs.

[0095] In one possible implementation, the processing unit 802 is used to fragment the air interface log information to obtain multiple first fragment logs; fragment the base station log information to obtain multiple second fragment logs; and for the first fragment logs and second fragment logs belonging to the same time period, fuse the parameter information corresponding to the same parameters in the first fragment logs and the second fragment logs to obtain the fused parameter information, thereby obtaining the fragment logs corresponding to the time period.

[0096] In one possible implementation, the parameters included in the air interface log information and / or base station log information include one or more of the following: voice signal strength, signal-to-noise ratio, voice coding method, or transmission delay.

[0097] In one possible implementation, the processing unit 802 is configured to report air interface log information and / or base station log information to the management platform if the voice call quality is less than a quality threshold; and / or input the air interface log information and / or base station log information into the voice quality detection model to determine the factors affecting the voice call quality of the voice receiver.

[0098] In one possible implementation, processing unit 802 is used to determine the sample voice sender and at least one sample voice receiver; acquisition unit 801 is used to acquire the received voice recorded by at least one sample voice receiver during the sample call, and to collect the sample air interface log information and / or sample base station log information corresponding to at least one sample voice receiver during the sample call; processing unit 802 is used to send the recorded received voice to a voice quality assessment instrument, which determines the voice quality score of at least one sample voice receiver; the sample air interface log information and / or sample base station log information corresponding to each sample voice receiver are segmented according to the scoring period of the voice quality assessment instrument to obtain sample segment logs; the sample segment logs and the corresponding voice quality scores are used as training samples to train a voice quality assessment model.

[0099] Based on the same technical concept, embodiments of this application provide an apparatus 900 for evaluating voice call quality, which may be, for example, a computing device. Figure 9 As shown, an apparatus 900 for evaluating voice call quality includes at least one processor 901 and a memory 902 connected to the at least one processor. In this embodiment, the specific connection medium between the processor 901 and the memory 902 is not limited. Figure 9 Taking the connection between processor 901 and memory 902 via a bus as an example, the bus can be divided into address bus, data bus, control bus, etc.

[0100] In this embodiment of the application, the memory 902 stores instructions that can be executed by at least one processor 901. By executing the instructions stored in the memory 902, at least one processor 901 can perform the above-described method for evaluating voice call quality.

[0101] The processor 901 is a control center for a device 900 that evaluates voice call quality. It can connect to various parts of a computer device via various interfaces and lines, and performs resource settings by running or executing instructions stored in the memory 902 and accessing data stored in the memory 902. Optionally, the processor 901 may include one or more determining units. The processor 901 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 901. In some embodiments, the processor 901 and the memory 902 may be implemented on the same chip; in some embodiments, they may be implemented on separate chips.

[0102] The processor 901 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0103] Memory 902, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 902 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 902 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 902 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0104] This application also provides a computer-readable storage medium storing a computer-executable program for causing a computer to perform a method for evaluating voice call quality as listed in any of the above embodiments.

[0105] This application provides a computer program product, including a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform a method for evaluating voice call quality as listed in any of the above methods.

[0106] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0107] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0108] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0109] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0110] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for evaluating voice call quality, characterized in that, include: Acquire air interface log information and / or base station log information collected during a voice call for the voice receiver; The air interface log information records parameter information that characterizes the communication quality at the air interface corresponding to the voice receiver during the voice call. The base station log information records parameter information that characterizes the communication quality at the base station corresponding to the voice receiver during the voice call. The air interface log information and / or the base station log information are divided into time slices to obtain multiple log segments. The multiple segment logs are input into the voice quality assessment model to obtain the voice call quality of the voice receiver during the voice call; wherein, the voice quality assessment model is obtained by training on sample segment logs and the voice quality scores corresponding to the sample segment logs.

2. The method as described in claim 1, characterized in that, The step of inputting the multiple log segments into the voice quality assessment model to obtain the voice call quality of the voice receiver during the voice call includes: The multiple log segments are input into the speech quality assessment model to obtain the speech quality score corresponding to each log segment. The voice call quality of the voice receiver during the voice call is determined based on the voice quality score corresponding to each segment log.

3. The method as described in claim 1, characterized in that, The air interface log information and / or the base station log information are divided into multiple log segments according to time slices, including: According to the duration corresponding to the scoring cycle of the voice quality assessment instrument, the air interface log information and / or the base station log information are segmented to obtain multiple log segments.

4. The method as described in claim 3, characterized in that, The process of fragmenting the air interface log information and / or the base station log information to obtain multiple log fragments includes: The air interface log information is fragmented to obtain multiple first fragment logs; The base station log information is fragmented to obtain multiple second-segment logs; For the first and second log segments belonging to the same time period, the parameter information corresponding to the same parameters in the first and second log segments is merged to obtain the merged parameter information, thereby obtaining the log segment corresponding to the time period.

5. The method according to any one of claims 1 to 4, characterized in that, The parameters included in the air interface log information and / or the base station log information include one or more of the following: voice signal strength, signal-to-noise ratio, voice coding method, or transmission delay.

6. The method as described in claim 1, characterized in that, After obtaining the voice call quality of the voice receiver during the voice call, the method further includes: If the voice call quality is less than the quality threshold, the air interface log information and / or the base station log information are reported to the management platform; and / or, the air interface log information and / or the base station log information are input into the voice quality detection model to determine the factors affecting the voice call quality of the voice receiver.

7. The method as described in claim 1, characterized in that, The speech quality assessment model is trained on sample segment logs and their corresponding speech quality scores, including: Identify the sender of the sample voice and at least one receiver of the sample voice; Acquire received voice recordings from at least one sample voice receiver during a sample call, and collect sample air interface log information and / or sample base station log information corresponding to the at least one sample voice receiver during the sample call. The recorded received voice is sent to a voice quality assessment instrument, which then determines the voice quality score of each of the at least one sample voice receiver. The sample air interface log information and / or sample base station log information corresponding to each sample voice receiver are segmented according to the scoring period of the voice quality assessment instrument to obtain sample segment logs; the sample segment logs and the corresponding voice quality scores are used as training samples to train the voice quality assessment model.

8. An apparatus for evaluating the quality of voice calls, characterized in that, Includes an acquisition unit and a processing unit: The acquisition unit is used to acquire air interface log information and / or base station log information collected during a voice call for the voice receiver; the air interface log information records parameter information characterizing the communication quality at the air interface corresponding to the voice receiver during the voice call. The base station log information records parameter information that characterizes the communication quality at the base station corresponding to the voice receiver during the voice call. The processing unit is configured to divide the air interface log information and / or the base station log information into time slices to obtain multiple segment logs; input the multiple segment logs into a voice quality assessment model to obtain the voice call quality of the voice receiver during the voice call; wherein, the voice quality assessment model is obtained by training on sample segment logs and the voice quality scores corresponding to the sample segment logs.

9. A computing device, characterized in that, include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the method as described in any one of claims 1 to 7 according to the obtained program instructions.

10. A computer-readable storage medium, characterized in that, Includes computer-readable instructions that, when read and executed by a computer, cause the method as described in any one of claims 1 to 7 to be implemented.

Citation Information

Patent Citations

  • Voice quality evaluation method, device and system

    CN112509603A

  • Alternating current management device and method

    CN115004297A