A method, apparatus, device and storage medium for voice quality assessment

Through the method of evaluating the voice quality of mobile networks, and through segmentation and model calculation, the problem of testers in the prior art needs to carry special equipment, achieving faster and more efficient voice quality evaluation.

CN115273899BActive Publication Date: 2025-05-30CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210858447.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-20
Publication Date
2025-05-30
Estimated Expiration
2042-07-20

AI Technical Summary

Technical Problem

In the prior art, testers need to carry special equipment when evaluating the voice quality of mobile networks, which is inefficient and costly.

Method used

Voice quality is evaluated by segmenting the voice call data of the target area within the preset time period based on the target period, sample data of multiple time slices are generated, and these data are input into the target model to calculate the average opinion score MOS value of each sample data.

Benefits of technology

This method does not require testers to carry special equipment, and can quickly determine the voice quality of the area to be evaluated, reduce the evaluation cost, and improve the evaluation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273899B_ABST
    Figure CN115273899B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and storage medium for voice quality assessment, relating to the field of communication technologies, and is used to improve the efficiency of assessing the voice quality of a mobile network. The method includes: performing segmentation processing on the data of voice calls corresponding to a target area within a preset time period based on a target period to obtain a plurality of sample data corresponding to a plurality of time slices, where one time slice corresponds to one sample data; inputting each of the plurality of sample data into a target model to obtain the MOS value corresponding to each sample data; the target model is a model constructed according to historical measurement data corresponding to the target area in a target historical period; according to the MOS value corresponding to each sample data, obtaining the MOS value corresponding to a target object, where the target object is at least one of the following: a voice call corresponding to the target area, a grid, a sub-area, and the target area. The present application is applied to the scenario of assessing the voice quality of a mobile network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to a voice quality assessment method, apparatus, device and storage medium. Background Art

[0002] With the evolution of communication technology, the 4th Generation Mobile Communication Technology (4G) and the 5th Generation Mobile Communication Technology (5G) network standards have become the main market players. User voice calls are mainly based on IMS-based voice services (Voice over LTE, VoLTE), and 5G voice calls (Voice over New Radio, VoNR) are gradually being commercialized. Voice quality is the focus of operators when evaluating network quality, as it directly affects the user's call experience.

[0003] At present, the method of evaluating voice quality is to evaluate voice quality by scoring based on the user's subjective auditory experience; or to evaluate voice quality by comparing the corpus sent by the voice sending end with the corpus received by the voice receiving end through an algorithm to obtain the Mean Opinion Score (MOS). Specifically, the operator conducts urban road tests or fixed-point tests in key locations for voice quality within a cycle. When conducting voice MOS value evaluation, it is necessary to carry a dedicated MOS test device. The voice sending end and the voice receiving end use fixed corpus. After the voice receiving end receives the corpus sent by the voice sending end, it compares the received corpus with the locally stored original corpus (i.e., the corpus sent by the sending end) to obtain the voice quality evaluation result.

[0004] In the above method, when performing a voice quality evaluation test, it is necessary to use a dedicated MOS test device in the area to be evaluated to test the voice quality of the area to be evaluated. In this case, the tester needs to carry the MOS test device with him / her at all times to test the voice quality of the area to be evaluated, which is extremely inconvenient for the tester, and the cost of performing voice quality evaluation using the MOS test device is high. Therefore, the current efficiency of evaluating the voice quality of mobile networks is relatively poor. Summary of the invention

[0005] The present application provides a voice quality assessment method, apparatus, device and storage medium for quickly determining the voice quality of an area to be assessed, reducing the cost of voice quality assessment, and improving the efficiency of assessing the voice quality of a mobile network.

[0006] To achieve the above object, the present application adopts the following technical solutions:

[0007] In a first aspect, a method for evaluating voice quality is provided. The method includes: based on a target period, performing segmentation processing on data of a voice call corresponding to a target area within a preset time period to obtain a plurality of sample data corresponding to a plurality of time slices; one time slice corresponds to one sample data, and the preset time period is the time period before the current moment. The data of the voice call includes: Trace data, signaling, and detailed records of traffic XDR voice bills; inputting each of the plurality of sample data into a target model to obtain a mean opinion score (MOS) value corresponding to each sample data; the target model is a model constructed based on historical voice quality measurement data corresponding to the target area in a target historical period; according to the MOS value corresponding to each sample data, obtaining the MOS value corresponding to a target object, and the MOS value corresponding to the target object is used to evaluate the voice quality of the target object. The target object is at least one of the following: a voice call corresponding to the target area, a grid, a sub-area, and the target area. The grid and the sub-area are obtained by dividing the target area.

[0008] In a possible implementation manner, inputting each of the plurality of sample data into the target model to obtain a mean opinion score (MOS) value corresponding to each sample data includes: determining a target performance indicator corresponding to each sample data; the target performance indicator includes at least one of the following: measurement report (MR) data, the number of radio resource control (RRC) reconstructions, and the voice call duration staying in the target network; inputting the target performance indicator corresponding to each sample data into the target model to obtain the MOS value corresponding to each sample data.

[0009] In a possible implementation manner, before performing segmentation processing on data of a voice call corresponding to a target area within a preset time period based on a target period to obtain a plurality of sample data corresponding to a plurality of time slices, the method further includes: obtaining historical voice quality measurement data corresponding to the target area in a target historical period, and performing segmentation processing on the historical voice quality measurement data based on the target period to obtain a plurality of training data corresponding to a plurality of time slices; one time slice corresponds to one training data; determining a target performance indicator and a historical MOS value corresponding to each training data, and training the target model based on the target performance indicator and the historical MOS value corresponding to each training data.

[0010] In a possible implementation, the voice call corresponding to the target area includes multiple voice calls, and each voice call corresponds to at least one sample data. The target object includes: the voice call corresponding to the target area; obtaining the MOS value corresponding to the target object according to the MOS value corresponding to each sample data includes: for any voice call, determining the MOS value corresponding to the voice call based on the MOS value corresponding to each sample data in the at least one sample data corresponding to the voice call and the number of the at least one sample data.

[0011] In a possible implementation, the target object includes: a grid; the method further includes: rasterizing the target area to obtain a plurality of grids and determining the position information of each grid; determining the target position information corresponding to each voice call from the data of the voice calls, and determining the grid corresponding to each voice call according to the target position information corresponding to each voice call; each grid includes at least one voice call; obtaining the MOS value corresponding to the target object according to the MOS value corresponding to each sample data includes: for any grid, determining the MOS value corresponding to the grid based on the MOS value corresponding to each voice call in the at least one voice call included in the grid and the first total number of the sample data corresponding to the grid.

[0012] In a possible implementation, the target area includes a plurality of sub-areas, and each sub-area includes at least one grid of the plurality of grids. The target object includes: a sub-area; obtaining the MOS value corresponding to the target object according to the MOS value corresponding to each sample data includes: for any sub-area, determining the MOS value corresponding to the sub-area based on the MOS value corresponding to each grid in the at least one grid included in the sub-area and the first weight value corresponding to each grid in the at least one grid included in the sub-area, where the first weight value is the ratio between the first total number of the sample data corresponding to each grid and the second total number of the sample data corresponding to the sub-area.

[0013] In a possible implementation, the target object includes: the target area; obtaining the MOS value corresponding to the target object according to the MOS value corresponding to each sample data includes: determining the MOS value corresponding to the target area based on the MOS value corresponding to each grid in the plurality of grids included in the target area and the second weight value corresponding to each grid in the plurality of grids included in the target area, where the second weight value is the ratio between the first total number of the sample data corresponding to each grid and the third total number of the sample data corresponding to the target area.

[0014] In a possible implementation, the target network includes a first network and a second network. The data of the voice call includes the data of the voice call based on the first network and the data of the voice call based on the second network. The voice call duration staying in the target network includes: the voice call duration staying in the first network and the voice call duration staying in the second network; wherein, the data type of the data of the voice call based on the first network is different from the data type of the data of the voice call based on the second network.

[0015] In a second aspect, a voice quality evaluation device is provided. The voice quality evaluation device includes: a processing unit; the processing unit is configured to perform a segmentation process on the data of the voice call corresponding to the target area within a preset time period based on a target period, so as to obtain a plurality of sample data corresponding to a plurality of time slices; one time slice corresponds to one sample data, and the preset time period is the time period before the current moment. The data of the voice call includes: Trace data, signaling, and XDR voice call records of traffic details; the processing unit is configured to input each sample data among the plurality of sample data into a target model to obtain the mean opinion score (MOS) value corresponding to each sample data; the target model is a model constructed according to the historical voice quality measurement data corresponding to the target area in a target historical time period; the processing unit is configured to obtain the MOS value corresponding to the target object according to the MOS value corresponding to each sample data, and the MOS value corresponding to the target object is used to evaluate the voice quality of the target object. The target object is at least one of the following: the voice call corresponding to the target area, a grid, a sub-area, and the target area. The grid and the sub-area are obtained by dividing the target area.

[0016] In a possible implementation, the processing unit is configured to determine the target performance index corresponding to each sample data; the target performance index includes at least one of the following: measurement report (MR) data, the number of radio resource control (RRC) reconstructions, and the voice call duration staying in the target network; the processing unit is configured to input the target performance index corresponding to each sample data into the target model to obtain the MOS value corresponding to each sample data.

[0017] In a possible implementation, the voice quality evaluation device further includes: an acquisition unit; the acquisition unit is configured to acquire the historical voice quality measurement data corresponding to the target area in a target historical time period; the processing unit is configured to perform a segmentation process on the historical voice quality measurement data based on a target period to obtain a plurality of training data corresponding to a plurality of time slices; one time slice corresponds to one training data; the processing unit is configured to determine the target performance index and the historical MOS value corresponding to each training data, and train the target model based on the target performance index and the historical MOS value corresponding to each training data.

[0018] In a possible implementation, the voice call corresponding to the target area includes multiple voice calls, and each voice call corresponds to at least one sample data. The target object includes: the voice call corresponding to the target area; a processing unit, configured to, for any voice call, determine the MOS value corresponding to the voice call based on the MOS value corresponding to each sample data in the at least one sample data corresponding to the voice call and the quantity of the at least one sample data.

[0019] In a possible implementation, the target object includes: a grid; a processing unit, configured to rasterize the target area to obtain multiple grids and determine the position information of each grid; a processing unit, configured to determine the target position information corresponding to each voice call from the data of the voice calls and determine the grid corresponding to each voice call according to the target position information corresponding to each voice call; each grid includes at least one voice call; a processing unit, configured to, for any grid, determine the MOS value corresponding to the grid based on the MOS value corresponding to each voice call in the at least one voice call included in the grid and the first total quantity of the sample data corresponding to the grid.

[0020] In a possible implementation, the target area includes multiple sub-areas, and each sub-area includes at least one grid among the multiple grids. The target object includes: a sub-area; a processing unit, configured to, for any sub-area, determine the MOS value corresponding to the sub-area based on the MOS value corresponding to each grid in the at least one grid included in the sub-area and the first weight value corresponding to each grid in the at least one grid included in the sub-area, where the first weight value is the ratio between the first total quantity of the sample data corresponding to each grid and the second total quantity of the sample data corresponding to the sub-area.

[0021] In a possible implementation, the target object includes: the target area; a processing unit, configured to determine the MOS value corresponding to the target area based on the MOS value corresponding to each grid in the multiple grids included in the target area and the second weight value corresponding to each grid in the multiple grids included in the target area, where the second weight value is the ratio between the first total quantity of the sample data corresponding to each grid and the third total quantity of the sample data corresponding to the target area.

[0022] In a possible implementation, the target network includes a first network and a second network. The data of the voice call corresponding to the target area includes the data of the voice call based on the first network and the data of the voice call based on the second network. The voice call duration staying in the target network includes: the voice call duration staying in the first network and the voice call duration staying in the second network; wherein, the data type of the data of the voice call based on the first network is different from the data type of the data of the voice call based on the second network.

[0023] In a third aspect, an electronic device includes: a processor and a memory; wherein, the memory is used to store one or more programs, and the one or more programs include computer execution instructions. When the electronic device runs, the processor executes the computer execution instructions stored in the memory, so that the electronic device executes a voice quality evaluation method as described in the first aspect.

[0024] In a fourth aspect, a computer-readable storage medium storing one or more programs is provided. The one or more programs include instructions that, when executed by a computer, cause the computer to execute a voice quality evaluation method as described in the first aspect.

[0025] The present application provides a voice quality evaluation method, apparatus, device, and storage medium, which are applied to the scenario of evaluating the voice quality of a mobile network. When evaluating the voice quality, data of a voice call corresponding to a target area within a preset time period before the current moment can be sliced based on a target period to obtain a plurality of sample data corresponding to a plurality of time slices. Thus, each sample data among the plurality of sample data can be further input into a target model to obtain a MOS value corresponding to each sample data, and based on the MOS value corresponding to each sample data, a MOS value for evaluating the voice quality of the voice call corresponding to the target area can be obtained, or a MOS value for evaluating the voice quality of a grid, sub-region, and target area corresponding to the target area can be obtained. Through the above method, there is no need for a tester to always carry a MOS test device to test the voice quality of the area to be evaluated. Instead, the MOS value of the area to be evaluated is determined through historical voice call data and a model, so as to indicate the voice quality of the area to be evaluated through the MOS value of the area to be evaluated. The voice quality of the area to be evaluated can be determined more quickly, the cost of voice quality evaluation can be reduced, and the efficiency of evaluating the voice quality of a mobile network can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic structural diagram of a voice quality evaluation system provided by an embodiment of the present application;

[0027] Figure 2 It is a schematic flow diagram of a voice quality evaluation method provided by an embodiment of the present application Figure 1 ;

[0028] Figure 3 It is a schematic flow diagram of a voice quality evaluation method provided by an embodiment of the present application Figure 2 ;

[0029] Figure 4 It is a schematic flow diagram of a voice quality evaluation method provided by an embodiment of the present application Figure 3 ;

[0030] Figure 5A flow chart of a method for evaluating speech quality provided in an embodiment of the present application Figure 4 ;

[0031] Figure 6 A flow chart of a method for evaluating speech quality provided in an embodiment of the present application Figure 5 ;

[0032] Figure 7 A flow chart of a method for evaluating speech quality provided in an embodiment of the present application Figure 6 ;

[0033] Figure 8 A flow chart of a method for evaluating speech quality provided in an embodiment of the present application Figure 7 ;

[0034] Figure 9 A schematic diagram of the structure of a speech quality assessment device provided in an embodiment of the present application;

[0035] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0037] In the description of this application, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, "at least one" and "plurality" refer to two or more. The words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not limit them to be different.

[0038] At present, the main method for evaluating voice quality is that the tester carries the MOS test equipment specially provided by the instrument manufacturer to conduct urban road tests or fixed-point tests in key places, and uses fixed predictions at the sending and receiving ends. After receiving, the receiving end compares the received predictions with the original corpus saved locally, and periodically scores to achieve the evaluation of the MOS value of voice quality. However, the above method requires carrying special MOS test equipment, which is extremely inconvenient for the tester, and the cost of using MOS test equipment to evaluate voice quality is high.

[0039] A speech quality assessment method provided in an embodiment of the present application can be applied to a speech quality assessment system. Figure 1A structural schematic diagram of the voice quality assessment system is shown. As Figure 1 shown, the voice quality assessment system 20 includes: an electronic device 21 and a server 22. The electronic device 21 is connected to the server 22. The connection between the electronic device 21 and the server 22 can be wired or wireless, and the embodiments of the present application do not limit this.

[0040] The voice quality assessment system 20 can be used in the Internet of Things. The voice quality assessment system 20 may include multiple hardware components such as multiple central processing units (CPUs), multiple memories, and a storage device storing multiple operating systems.

[0041] The electronic device 21 can be used in the Internet of Things to provide data processing services for users, interact with the server 22, and implement the data processing services required by users.

[0042] The server 22 can be used in the Internet of Things, connected to the electronic device 21, and is the server corresponding to the data processed by the electronic device 21. That is, the server 22 is the server of the communication operator, used to provide services such as data information transmission for users. For example, it provides the data information required for operation and processing for the electronic device 21 so that the electronic device 21 can provide data processing services for users.

[0043] It should be noted that the electronic device 21 and the server 22 can be independent devices or integrated into the same device, and the present application does not make specific limitations on this.

[0044] When the electronic device 21 and the server 22 are integrated into the same device, the communication method between the electronic device 21 and the server 22 is the communication between internal modules of the device. In this case, the communication process between the two is the same as the "communication process between the electronic device 21 and the server 22 when they are independent of each other".

[0045] In the following embodiments provided by the present application, the present application takes the electronic device 21 and the server 22 being independently set as an example for description.

[0046] Next, a voice quality assessment method provided by the embodiments of the present application will be described with reference to the accompanying drawings.

[0047] As Figure 2 shown, a voice quality assessment method provided by the embodiments of the present application includes S201 - S203:

[0048] S201. Based on the target period, perform segmentation processing on the data of the voice call corresponding to the target area within the preset time period to obtain multiple sample data corresponding to multiple time slices.

[0049] Among them, one time slice corresponds to one sample data, the preset time period is the time period before the current moment, and the data of the voice call includes: Trace data, signaling, and traffic detail record (XDR) voice call records.

[0050] It can be understood that in a voice quality assessment method proposed in an embodiment of this application, before performing voice quality assessment, it is necessary to pre-determine the target area for voice quality assessment and obtain the data of voice calls within the preset time period before the current moment in the target area.

[0051] It should be noted that the target area can be understood as an administrative area in terms of geographical location (such as the geographical area corresponding to City A, the geographical area corresponding to City B, etc.), or the target area can be a geographical range determined artificially.

[0052] In one implementation, the data of voice calls in the target area includes the data of all voice calls of all users in the target area within the preset time period.

[0053] It can be understood that the preset time period can be a preset duration before the current moment, such as 1 hour, 6 hours, 12 hours, etc. before the current moment.

[0054] Optionally, the Trace data in the data of voice calls can be collected by the operation and maintenance personnel through subscription. The Trace data can be radio side Trace data (i.e., data transmitted in a wireless form); the XDR data can be collected by using a deep packet inspection (DPI) signaling collection system.

[0055] It can be understood that by slicing the data of voice calls through the target period, multiple sample data are obtained. Each sample data corresponds to the data of a target period duration in the data of voice calls. All the sample data constitutes the complete data of voice calls. The target period can be 3 seconds, 6 seconds, 9 seconds, etc.

[0056] It should be noted that when slicing the data of voice calls in chronological order, when the time length of the remaining data of voice calls is less than the target period, the data of voice calls corresponding to the remaining time length is used as one sample data, or the data of voice calls corresponding to the remaining time length is discarded (i.e., regarded as invalid data).

[0057] It can be understood that each sample data contains some of the above-mentioned Trace data and XDR voice call record data.

[0058] S202. Input each of the multiple sample data into the target model to obtain the Mean Opinion Score (MOS) value corresponding to each sample data.

[0059] Among them, the target model is a model constructed based on the historical voice quality measurement data corresponding to the target area in the target historical time period.

[0060] It should be noted that the target model is a neural network model, which is a pre-trained model using historical data. By inputting each sample data into the target model, the MOS value corresponding to each sample data can be predicted.

[0061] Optionally, multiple sample data can be input into the target model simultaneously to obtain the MOS value corresponding to each sample data; or, each of the multiple sample data can be input into the target model sequentially to obtain the MOS value corresponding to each sample data.

[0062] It should be noted that the MOS value refers to the Mean Opinion Value, which is used to evaluate the voice quality.

[0063] S203. Obtain the MOS value corresponding to the target object according to the MOS value corresponding to each sample data.

[0064] Among them, the MOS value corresponding to the target object is used to evaluate the voice quality of the target object. The target object is at least one of the following: the voice call corresponding to the target area, the grid, the sub-region, the target area. The grid and the sub-region are obtained by dividing the target area.

[0065] Optionally, the voice call corresponding to the above target area is the voice call made in the target area during the preset time period; the above grid is the grid obtained after rasterizing the target area; the above sub-region is the sub-region included in the target area.

[0066] It can be understood that after obtaining the MOS value corresponding to each sample data, the MOS value corresponding to any voice call, any grid, any sub-region included in the target area, or the MOS value of the target area can be determined respectively according to the MOS value corresponding to each sample data.

[0067] An embodiment of the present application provides a method for evaluating voice quality, which is applied to a scenario of evaluating the voice quality of a mobile network. When evaluating the voice quality, data of a voice call corresponding to a target area within a preset time period before the current moment can be sliced based on a target period to obtain a plurality of sample data corresponding to a plurality of time slices. Thus, each sample data among the plurality of sample data can be further input into a target model to obtain a MOS value corresponding to each sample data, and based on the MOS value corresponding to each sample data, a MOS value for evaluating the voice quality of the voice call corresponding to the target area can be obtained, or a MOS value for evaluating the voice quality of a grid, a sub-region, or the target area corresponding to the target area can be obtained. Through the above method, it is not necessary for a tester to carry a MOS test device at all times to test the voice quality of the area to be evaluated. Instead, the MOS value of the area to be evaluated is determined through historical voice call data and a model, so as to indicate the voice quality of the area to be evaluated through the MOS value of the area to be evaluated. The voice quality of the area to be evaluated can be determined more quickly, the cost of voice quality evaluation can be reduced, and the efficiency of evaluating the voice quality of a mobile network can be improved.

[0068] In one design, as Figure 3 shown, in a method for evaluating voice quality provided by an embodiment of the present application, the above step S202 may specifically include S2021-S2022:

[0069] S2021. Determine the target performance indicators corresponding to each sample data.

[0070] Among them, the target performance indicators include at least one of the following: Measurement Report (MR) data, the number of Radio Resource Control (RRC) reconstructions (RRC Reest Num), and the voice call duration staying in the target network.

[0071] It should be noted that the target performance indicators are part of the data within the sample data, and other data, such as user information and location information, are also included in the sample data.

[0072] It can be understood that the number of RRC reconstructions refers to the number of radio link reconstructions occurring within the target period.

[0073] Optionally, after slicing the data of the voice call to obtain a plurality of sample data, each sample data can be parsed, so as to obtain the target performance indicators corresponding to each sample data from each sample data.

[0074] In one design, the target network may include a first network and a second network, and the data of the voice call corresponding to the target area may include the data of the voice call based on the first network and the data of the voice call based on the second network.

[0075] Moreover, the data type of the data of the voice call based on the first network is different from the data type of the data of the voice call based on the second network.

[0076] It can be understood that the voice call duration staying in the target network includes: the voice call duration staying in the first network and the voice call duration staying in the second network.

[0077] As a possible implementation, the first network may be a 4G network, and the second network may be a 5G network.

[0078] Table 1

[0079] Associated field of the first network Associated field of the second network IMSI IMSI ECGI NCGI MME Group ID AMF Region ID MME Code AMF Set ID MME UE S1AP ID AMF UE NGAP ID TIME TIME

[0080] Optionally, data fields of the first network data and the second network data can be associated respectively. As shown in Table 1, Table 1 is a relationship table for associating the first network data and the second network data fields. And through the International Mobile Subscriber Identity (IMSI) and the timestamp (TIME) of the terminal, the statistical information of the user's single call time is obtained, including the duration information of the voice calls staying in the first network and the second network. And according to the duration information of the voice calls staying in the first network and the second network within the user's single call time, the duration information of the voice calls staying in the first network and the second network in each sample data is calculated.

[0081] Among them, the associated fields of the first network include IMSI, E-UTRAN Cell Global Identifier (ECGI), Mobility Management Entity function (MME) Group ID (MMEGI), MME Code (MMEC), the unique identifier of the User Equipment (UE) on the S1 interface of the MME side (MME UE S1 Application Protocol ID, MME UE S1AP ID), and TIME; the associated fields of the second network include IMSI, NR Cell Global Identifier (NCGI), Authentication Management Function (AMF) Region ID, AMF Set ID, the identifier that identifies the UE in the AMF at the N2 reference point (AMF UE NG Application Protocol ID, AMF UE NG AP ID), and TIME.

[0082] It should be noted that the information indicated by the IMSI field in the first network is the same as the information indicated by the IMSI field in the second network; the information indicated by the ECGI field in the first network is the same as the information indicated by the NCGI field in the second network; the information indicated by the MME Group ID field in the first network is the same as the information indicated by the AMF Region ID field in the second network; the information indicated by the MME Code field in the first network is the same as the information indicated by the AMF Set ID field in the second network; the information indicated by the MME UE S1AP ID field in the first network is the same as the information indicated by the AMF UE NG AP ID field in the second network; the information indicated by the TIME field in the first network is the same as the information indicated by the TIME field in the second network.

[0083] S2022. Input the target performance metric corresponding to each sample data into the target model to obtain the MOS value corresponding to each sample data.

[0084] In one implementation, the data for the voice call is the data obtained by the test terminal, which is a device supporting voice calls over the second network (e.g., 5G voice call (Voice over New Radio, VoNR)). When there is coverage of the second network and the VoNR of the second network base station is enabled, the test terminal is a 5G mobile network, and the voice service resides on VoNR. Otherwise, the test terminal falls back to the first network, and the voice service resides on the first network for voice calls (e.g., Voice over LTE (VoLTE) based on IMS).

[0085] Table 2

[0086]

[0087]

[0088] Specifically, as shown in Table 2, in the first network, the target performance indicators include MR data, the number of RRC reconstructions (RRCReest Num), and the VoLTE dwell duration (VoLTE Time Dwell) within the target period (i.e., the duration of the time slice); in the second network, the target performance indicators include MR data, the number of RRC reconstructions (RRC Reest Num), and the VoNR dwell duration (VoNR Time Dwell) within the target period.

[0089] Among them, the MR data in the first network includes: the average reference signal receiving power (ReferenceSignal Receiving Power, RSRP), the average reference signal receiving quality (Reference Signal ReceivingQuality, RSRQ), the average uplink signal to interference plus noise ratio (Signal to Interference plus NoiseRatio, SINR), and the average UE transmit power headroom PHR; the MR data in the second network includes: the average synchronization reference signal receiving power (Synchronization Signal Reference Signal Received Power, SS-RSRP), the average synchronization reference signal receiving quality (Synchronization Signal Reference Signal ReceivingQuality, SS-RSRQ), and the average synchronization signal to interference plus noise ratio (Synchronization Signal Signal toInterference plus Noise Ratio, SS-SINR).

[0090] It should be noted that the performance indicators of the sample data are not limited to the listed data types, and testers can also add other performance indicators according to actual needs.

[0091] In one design, as Figure 4 shown, in a voice quality assessment method provided by an embodiment of the present application, before the above step S202, the target model can be obtained through the following S101-S102:

[0092] S101. Obtain historical voice quality measurement data corresponding to the target area in the target historical time period, and perform segmentation processing on the historical voice measurement data based on the target period to obtain a plurality of training data corresponding to a plurality of time slices.

[0093] Among them, one time slice corresponds to one training data, and the historical voice quality measurement data is data obtained by operation and maintenance personnel measuring the historical voice call quality in the target area through a MOS device during the target historical time period. The data type of the historical voice quality measurement data is the same as the data type of the sample data.

[0094] As a possible implementation manner, the historical voice quality measurement data may include historical voice quality measurement data corresponding to the first network and historical voice quality measurement data corresponding to the second network. Specifically, the historical voice quality measurement data may include voice quality measurement data of the first network and voice quality measurement data of the second network.

[0095] It should be noted that when segmenting the historical voice quality measurement data in chronological order, when the time length of the remaining voice call data is less than the target period, the historical voice quality measurement data corresponding to the remaining time length is used as one training data, or the historical voice quality measurement data corresponding to the remaining time length is discarded (i.e., regarded as invalid data).

[0096] S102. Determine the target performance indicator and the historical MOS value corresponding to each training data, and train the target model based on the target performance indicator and the historical MOS value corresponding to each training data.

[0097] Optionally, after segmenting the historical voice quality measurement data to obtain a plurality of training data, each training data can be parsed to obtain the target performance indicator and the historical MOS value corresponding to each training data from each training data.

[0098] It can be understood that the target model is trained according to the correspondence between the performance indicator and the historical MOS value corresponding to each sample data. The trained target model can predict the MOS value data corresponding to the input performance indicator data through the input performance indicator data.

[0099] Further, by inputting the target performance index and the historical MOS value corresponding to each training data obtained into the initial model, a target model is trained.

[0100] It can be understood that an initial model can be selected, the target performance index and the historical MOS value are input into the initial model, and the model is trained by continuously inputting the target performance index and the historical MOS value to adjust the parameters of the model, so as to obtain a prediction model indicating the corresponding relationship between the target performance index and the MOS value, that is, the target model.

[0101] In one design, the voice call corresponding to the target area includes multiple voice calls, each voice call includes at least one sample data, and the target object includes: the voice call corresponding to the target area; as Figure 5 As shown, in a voice quality assessment method provided by an embodiment of the present application, the above step S203 may specifically include S301:

[0102] S301. For any voice call, based on the MOS value corresponding to each sample data in at least one sample data corresponding to the voice call, and the number of at least one sample data, determine the MOS value corresponding to the voice call.

[0103] It can be understood that any voice call refers to the data of a complete voice call.

[0104] Optionally, the average MOS value corresponding to at least one sample data included in any voice call (that is, all sample data included in the voice call) can be calculated, and the average MOS value can be used as the MOS value corresponding to the voice call.

[0105] Optionally, the MOS values corresponding to at least one sample data included in any voice call can also be sorted in ascending order, and the median can be used as the MOS value corresponding to the voice call.

[0106] Specifically, as shown in Formula 1, taking the average MOS value as an example, the MOS value corresponding to the voice call can be obtained by calculating the average value of the MOS values corresponding to at least one sample data included in the voice call.

[0107]

[0108] Among them, is the MOS value of the mth voice call of user n, is the MOS value of the cth sample data in the data of the mth voice call of user n, and H is the number of sample data of this voice call.

[0109] An embodiment of the present application provides a method for evaluating voice quality. When evaluating voice quality, for any voice call among multiple voice calls, the MOS value for evaluating the voice quality of this voice call can be calculated based on the MOS value corresponding to each sample data in at least one sample data included in this voice call. Thus, the MOS value corresponding to each voice call among multiple voice calls is determined through the MOS value corresponding to each sample data, so as to indicate the voice quality of a single voice call through the MOS value of each voice call.

[0110] In one design, the target object includes: grids; as Figure 6 shown, in a method for evaluating voice quality provided by an embodiment of the present application, the following S401 - S403 may further be included, and step S203 above may specifically include S403:

[0111] S401. Perform grid processing on the target area to obtain a plurality of grids, and determine the position information of each grid.

[0112] Optionally, the grid size accuracy can be selected as L*L (L is the side length of the square grid), for example, the grid size accuracy can be selected as a square with a side length of 10 meters; or, the accuracy of the grid size can also be selected as H*L (H and L are the length and width of the rectangular grid respectively), for example, the accuracy of the grid size can be selected as a rectangle with a length of 15 and a width of 10.

[0113] It can be understood that the grid is a geographical grid, that is, by dividing the geographical information of the target area, a plurality of grid data containing different geographical position information is obtained.

[0114] S402. Determine the target position information corresponding to each voice call from the data of the voice call, and determine the grid corresponding to each voice call according to the target position information corresponding to each voice call.

[0115] Among them, each grid includes at least one voice call.

[0116] Optionally, the position information corresponding to at least one sample data corresponding to each voice call can be determined, so as to determine the target position information corresponding to each voice call.

[0117] It should be noted that since the user is in a moving state during a voice call, a single voice call can correspond to multiple position information (that is, the target position information is multiple position information), so a single voice call can correspond to multiple grids.

[0118] In a specific implementation, by determining the longitude and latitude information of at least one sample data corresponding to each voice call, the grid position information corresponding to the voice call is judged. Specifically, by determining the Minimization of Drive Tests (MDT) in the Trace data of the central time point of the time slice corresponding to the sample data, the longitude and latitude information corresponding to the sample data is obtained, so as to judge the grid position information corresponding to each voice call through the longitude and latitude information of each sample data.

[0119] S403. For any grid, based on the MOS value corresponding to each voice call in at least one voice call included in the grid, and the first total number of sample data corresponding to the grid, determine the MOS value corresponding to the grid.

[0120] Optionally, the average MOS value of the MOS values corresponding to each voice call in at least one voice call included in the grid (i.e., all voice calls included in the grid) can be calculated, and this average MOS value can be used as the MOS value corresponding to the grid.

[0121] Optionally, the MOS values corresponding to each voice call in at least one voice call included in the grid can also be sorted in ascending order, and the median can be used as the MOS value corresponding to the grid.

[0122] Specifically, as shown in Formula 2, taking the average MOS value as an example, the MOS value corresponding to the grid can be obtained by calculating the average value of the MOS values corresponding to each voice call in at least one voice call included in the grid.

[0123]

[0124] Specifically, as shown in Formula 2, where j MOS_G is the MOS value of the j-th grid, j is the MOS value of the i-th voice call included in the j-th grid, and D

[0125] is the total number of voice calls included in the j-th grid.

[0126] An embodiment of the present application provides a method for evaluating voice quality. When evaluating voice quality, for any grid among multiple grids, the MOS value for evaluating the voice quality of this grid can be calculated based on the MOS value corresponding to each voice call in at least one voice call included in the grid. Thus, the MOS value corresponding to each grid among the multiple grids is determined through the MOS value corresponding to each voice call, so as to indicate the voice quality within a single grid through the MOS value corresponding to each grid.

[0127] In one design, the target area includes multiple sub-areas, each sub-area includes at least one grid among the multiple grids, and the target object includes: sub-areas; as Figure 7 shown, in a method for evaluating voice quality provided by an embodiment of the present application, step S203 above may specifically include S501:

[0128] S501. For any sub-area, based on the MOS value corresponding to each grid among at least one grid included in the sub-area and the first weight value corresponding to each grid among at least one grid included in the sub-area, determine the MOS value corresponding to the sub-area.

[0129] Among them, the first weight value is the ratio between the first total quantity of sample data corresponding to each grid and the second total quantity of sample data corresponding to the sub-area.

[0130] Optionally, after determining the MOS value corresponding to each grid among the multiple grids included in the target area, the average MOS value of the MOS values corresponding to each grid among at least one grid (i.e., all grids included in the sub-area) included in the sub-area can be further calculated, and this average MOS value is used as the MOS value corresponding to the sub-area.

[0131] Optionally, the MOS values corresponding to each grid among at least one grid included in the sub-area can also be sorted in ascending order, and the median is used as the MOS value corresponding to the sub-area.

[0132] Specifically, as shown in Formula 3, taking the average MOS value as an example, the MOS value corresponding to the sub-area can be obtained by calculating the average value of the MOS values corresponding to each grid among at least one grid included in the sub-area.

[0133]

[0134] Among them, MOS_S e is the MOS value of the e-th sub-area, MOS_G q is the MOS value of the q-th grid within this sub-area, U is the total number of grids contained in this sub-area, and W qAs shown in Formula 4, it is the ratio of the total number of sample data corresponding to the qth grid to the total number of sample data corresponding to the sub-region (ie, the first weight value).

[0135]

[0136] Among them, D q is the total number of sample data corresponding to the qth grid, is the total number of sample data corresponding to the sub-region, and U is the total number of grids contained in the sub-region.

[0137] It can be understood that the MOS value corresponding to each sub-region can be calculated through Formula 3 and Formula 4.

[0138] Optionally, whether the grid belongs to a sub-region can be determined by using the ray method and whether the longitude and latitude of the center of the grid are within the sub-region.

[0139] Specifically, the sub-area may be some specified latitude and longitude boundaries, such as specific areas such as schools and parks within the target area.

[0140] It can be understood that the sub-area is an area divided by the tester according to the actual test needs, based on specific life scenarios or specific functional areas in the target area, and is used to evaluate the voice call quality in a specific scenario.

[0141] It should be noted that if the sub-area is a road line area, such as an urban road, etc., taking into account the particularity of the line area (that is, the length of the area is much greater than the width of the area), the road line area can be specially rasterized through the road GIS layer. The grid width is the variable road width, and the grid length is a fixed length L. The sample data of a single voice call falls into the road grid, and the MOS value of the road line area is evaluated by Formula 3.

[0142] The embodiment of the present application provides a method for evaluating voice quality. When evaluating voice quality, for any sub-region among multiple sub-regions, a MOS value for evaluating the voice quality of the sub-region can be calculated based on the MOS value corresponding to each grid in at least one grid included in the sub-region. Thus, the MOS value corresponding to each sub-region among the multiple sub-regions is determined by the MOS value corresponding to each grid, so that the voice quality in a single sub-region is indicated by the MOS value corresponding to each sub-region.

[0143] In one design, the target object includes: a target area; Figure 8 As shown, in a voice quality assessment method provided in an embodiment of the present application, the above step S203 may specifically include the following S2031:

[0144] S2031. Determine the MOS value corresponding to the target area based on the MOS value corresponding to each grid included in the target area and the second weight value corresponding to each grid included in the target area.

[0145] Among them, the second weight value is the ratio between the first total quantity of the sample data corresponding to each grid and the third total quantity of the sample data corresponding to the target area.

[0146]

[0147] Specifically, as shown in Formula 5, where MOS_Q d is the MOS value of the d-th area (such as the target area), MOS_G s is the MOS value of the s-th grid within this area, A is the total number of grids contained in the target area, W s As shown in Formula 6, it is the ratio between the total quantity of the sample data corresponding to the s-th grid and the total quantity of the sample data corresponding to the target area (i.e., the second weight value).

[0148]

[0149] Among them, D a is the total quantity of the sample data corresponding to the s-th grid, is the total quantity of the sample data corresponding to the target area, and A is the total number of grids contained in the target area.

[0150] In one design, the average MOS value of the MOS values corresponding to multiple sample data included in the target area (i.e., all sample data included in the target area) can be calculated, and this average MOS value can be used as the MOS value corresponding to the target area.

[0151] Optionally, the MOS values corresponding to the multiple sample data included in the target area can also be sorted in ascending order, and the median can be used as the MOS value corresponding to the target area.

[0152] Specifically, as shown in Formula 7, taking the average MOS value as an example, the MOS value corresponding to the target area can be obtained by calculating the average value of the MOS values corresponding to all sample data included in the target area.

[0153]

[0154] Among them, MOS_Q is the MOS value of the target area, D x is the MOS value of the x-th sample data among all sample data included in the target area, is the total MOS value corresponding to the MOS values of all sample data included in the target area, and X is the number of sample data included in the target area.

[0155] The embodiment of the present application provides a voice quality evaluation method, which is applied to the scenario of evaluating the voice quality of a mobile network. Before evaluating the voice quality of the target area, it is necessary to pre-train a neural network model with the historical voice quality measurement data in the target area to obtain a target model. Thus, when evaluating the voice quality of the target area, the data of voice calls corresponding to the preset time period before the current moment in the target area are obtained, and the data of voice calls corresponding to the target area within the preset time period are sliced based on the target period to obtain multiple sample data corresponding to multiple time slices. Thus, according to each sample data among the multiple sample data, the target performance data corresponding to each sample data is determined. And the target performance index corresponding to each sample data is input into the target model to obtain the MOS value corresponding to each sample data. Further, the target area is rasterized to obtain multiple grids. To calculate the MOS value corresponding to a single voice call, the MOS value corresponding to a single grid, the MOS value corresponding to a single sub-area, and the MOS value corresponding to the target area according to the MOS value corresponding to each sample data.

[0156] Specifically, the MOS value corresponding to a single voice call can be calculated according to the MOS values corresponding to at least one sample data included in each voice call; the MOS value of each grid can be calculated according to the MOS values of multiple voice calls corresponding to each grid; and the MOS value of each sub-area can be calculated according to the MOS values of multiple grids included in each sub-area. In addition, the MOS value of the target area can also be calculated through the MOS values corresponding to all the grids included in the target area. Through the above method, there is no need for testers to always carry MOS test equipment to test the voice quality of the area to be evaluated. Instead, the MOS value of the area to be evaluated is determined through the data and model of historical voice calls, so as to indicate the voice quality of the area to be evaluated through the MOS value of the area to be evaluated. The voice quality of the area to be evaluated can be determined more quickly, the cost of voice quality evaluation can be reduced, and the efficiency of evaluating the voice quality of the mobile network can be improved. And, based on the MOS values of the time slices, the MOS values of each voice call, each grid, each sub-area, and the target area can be obtained, and the evaluation range is more flexible, which can meet diversified test requirements.

[0157] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0158] The embodiments of the present application can divide the functional modules of a voice quality evaluation device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. Optionally, the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0159] Figure 9 It is a schematic structural diagram of a voice quality evaluation device provided by the embodiments of the present application. As Figure 9 shown, a voice quality evaluation device 40 is used to improve the efficiency of evaluating the voice quality of a mobile network, for example, used to execute Figure 2 the voice quality evaluation method shown. The voice quality evaluation device 40 includes: a processing unit 401.

[0160] The processing unit 401 is configured to perform slicing processing on the data of the voice call corresponding to the target area within a preset time period based on the target period, so as to obtain a plurality of sample data corresponding to a plurality of time slices; wherein, one time slice corresponds to one sample data, the preset time period is the time period before the current moment, and the data of the voice call includes: Trace data, signaling, and XDR voice call records of traffic details.

[0161] The processing unit 401 is configured to input each sample data in the plurality of sample data into the target model to obtain the mean opinion score MOS value corresponding to each sample data; wherein, the target model is a model constructed according to the historical voice quality measurement data corresponding to the target area in the target historical time period.

[0162] A processing unit 401 is configured to obtain the MOS value corresponding to the target object according to the MOS value corresponding to each sample data, and the MOS value corresponding to the target object is used to evaluate the voice quality of the target object. The target object is at least one of the following: a voice call corresponding to a target area, a grid, a sub-region, and a target area. The grid and the sub-region are obtained by dividing the target area.

[0163] In a possible implementation manner, in a voice quality evaluation device 40 provided in an embodiment of the present application, the processing unit 401 is configured to determine a target performance metric corresponding to each sample data; wherein, the target performance metric includes at least one of the following: measurement report MR data, the number of radio resource control protocol RRC reconstructions, and the voice call duration staying in the target network.

[0164] The processing unit 401 is configured to input the target performance metric corresponding to each sample data into the target model to obtain the MOS value corresponding to each sample data.

[0165] In a possible implementation manner, in a voice quality evaluation device 40 provided in an embodiment of the present application, as Figure 9 shown, it further includes an acquisition unit 402; the acquisition unit 402 is configured to acquire historical voice quality measurement data corresponding to the target area in a target historical time period.

[0166] The processing unit 401 is configured to perform segmentation processing on the historical voice quality measurement data based on a target period to obtain multiple training data corresponding to multiple time slices; wherein, one time slice corresponds to one training data.

[0167] The processing unit 401 is configured to determine the target performance metric and the historical MOS value corresponding to each training data, and train the target model based on the target performance metric and the historical MOS value corresponding to each training data.

[0168] Optionally, the voice call corresponding to the target area includes multiple voice calls, and each voice call corresponds to at least one sample data. The target object includes: the voice call corresponding to the target area. In a voice quality evaluation device 40 provided in an embodiment of the present application, the processing unit 401 is configured to, for any voice call, determine the MOS value corresponding to the voice call based on the MOS value corresponding to each sample data in at least one sample data corresponding to the voice call and the number of at least one sample data.

[0169] Optionally, the target object includes: a grid; in a voice quality evaluation device 40 provided in an embodiment of the present application, the processing unit 401 is configured to perform grid processing on the target area to obtain multiple grids, and determine the position information of each grid.

[0170] The processing unit 401 is configured to determine the target location information corresponding to each voice call from the data of the voice calls, and determine the grid corresponding to each voice call according to the target location information corresponding to each voice call, where each grid includes at least one voice call.

[0171] The processing unit 401 is configured to, for any grid, determine the MOS value corresponding to the grid based on the MOS value corresponding to each voice call in at least one voice call included in the grid and the first total number of the sample data corresponding to the grid.

[0172] Optionally, the target area includes multiple sub-areas, each sub-area includes at least one grid among the multiple grids, and the target object includes: sub-areas. In a voice quality evaluation device 40 provided in an embodiment of the present application, the processing unit 401 is configured to, for any sub-area, determine the MOS value corresponding to the sub-area based on the MOS value corresponding to each grid in at least one grid included in the sub-area and the first weight value corresponding to each grid in at least one grid included in the sub-area, where the first weight value is the ratio between the first total number of the sample data corresponding to each grid and the second total number of the sample data corresponding to the sub-area.

[0173] Optionally, the target object includes: the target area; In a voice quality evaluation device 40 provided in an embodiment of the present application, the processing unit 401 is configured to determine the MOS value corresponding to the target area based on the MOS value corresponding to each grid in the multiple grids included in the target area and the second weight value corresponding to each grid in the multiple grids included in the target area, where the second weight value is the ratio between the first total number of the sample data corresponding to each grid and the third total number of the sample data corresponding to the target area.

[0174] Optionally, in a voice quality evaluation device 40 provided in an embodiment of the present application, the target network includes a first network and a second network, the data of the voice calls corresponding to the target area includes the data of the voice calls based on the first network and the data of the voice calls based on the second network, and the voice call duration staying in the target network includes: the voice call duration staying in the first network and the voice call duration staying in the second network.

[0175] Wherein, the data type of the data of the voice calls based on the first network is different from the data type of the data of the voice calls based on the second network.

[0176] In the case of implementing the functions of the above integrated module in the form of hardware, an embodiment of the present application provides another possible structural schematic diagram of the electronic device involved in the above embodiment. As Figure 10 shown, an electronic device 60 is used to improve the efficiency of evaluating the voice quality of a mobile network, for example, used to execute Figure 2A voice quality assessment method is shown. The electronic device 60 includes a processor 601, a memory 602, and a bus 603. The processor 601 and the memory 602 can be connected through the bus 603.

[0177] The processor 601 is the control center of the communication device and can be a single processor or a collective term for multiple processing elements. For example, the processor 601 can be a general-purpose central processing unit (CPU) or other general-purpose processors. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0178] As an embodiment, the processor 601 can include one or more CPUs, such as Figure 10 the CPUs 0 and CPU 1 shown in

[0179] The memory 602 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0180] As a possible implementation, the memory 602 can exist independently of the processor 601. The memory 602 can be connected to the processor 601 through the bus 603 for storing instructions or program codes. When the processor 601 calls and executes the instructions or program codes stored in the memory 602, it can implement a voice quality assessment method provided by an embodiment of the present application.

[0181] In another possible implementation, the memory 602 can also be integrated with the processor 601.

[0182] The bus 603 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 10 only a thick line is used to represent it in Figure 10 , but it does not mean that there is only one bus or one type of bus.

[0183] It should be noted that Figure 10 the structure shown does not constitute a limitation on the electronic device 60. Except Figure 10 for the components shown, the electronic device 60 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0184] As an example, in combination with Figure 9 , the functions implemented by the processing unit 401 and the acquisition unit 402 in the electronic device are the same as those of the processor 601 in Figure 10 .

[0185] Optionally, as shown in Figure 10 , the electronic device 60 provided in the embodiment of the present application may further include a communication interface 604.

[0186] The communication interface 604 is used to connect to other devices through a communication network. The communication network can be an Ethernet, a wireless access network, a wireless local area network (WLAN), etc. The communication interface 604 may include a receiving unit for receiving data and a transmitting unit for transmitting data.

[0187] In one design, in the electronic device provided in the embodiment of the present application, the communication interface may also be integrated in the processor.

[0188] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit is used as an example. In actual applications, the above functions can be allocated to different functional units as needed, that is, the internal structure of the device is divided into different functional units to complete all or part of the functions described above. The specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0189] The embodiments of the present application also provide a computer-readable storage medium, in which instructions are stored. When a computer executes these instructions, the computer performs each step in the method flow shown in the above method embodiments.

[0190] The embodiments of the present application provide a computer program product containing instructions. When the instructions run on a computer, the computer is enabled to execute a voice quality assessment method in the above method embodiments.

[0191] Among them, the computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk. Random Access Memory (RAM), Read-Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM), registers, hard disks, optical fibers, portable compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above, or any other form of computer-readable storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). In the embodiments of the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.

[0192] Since the electronic devices, computer-readable storage media, and computer program products in the embodiments of the present application can be applied to the above method, the technical effects they can obtain can also refer to the above method embodiments, and the embodiments of the present application will not be elaborated herein.

[0193] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application.

Claims

1. A method for evaluating voice quality, characterized in that, the method includes: Based on a target period, perform segmentation processing on the data of a voice call corresponding to a target area within a preset time period to obtain multiple sample data corresponding to multiple time slices; one time slice corresponds to one sample data, the preset time period is the time period before the current moment, and the data of the voice call includes: Trace data, signaling, and XDR voice call records; Determine the target performance indicators corresponding to each sample data; the target performance indicators include at least one of the following: Measurement Report (MR) data, the number of Radio Resource Control (RRC) reconstructions, and the voice call duration staying in the target network; the target network includes a first network and a second network, and the data of the voice call corresponding to the target area includes the data of the voice call based on the first network and the data of the voice call based on the second network. The voice call duration staying in the target network includes: the voice call duration staying in the first network and the voice call duration staying in the second network; wherein, the data type of the data of the voice call based on the first network is different from the data type of the data of the voice call based on the second network; Input the target performance indicators corresponding to each sample data into a target model to obtain the MOS value corresponding to each sample data; the target model is a model constructed based on the historical voice quality measurement data corresponding to the target area in a target historical period; According to the MOS value corresponding to each sample data, obtain the MOS value corresponding to the target object. The MOS value corresponding to the target object is used to evaluate the voice quality of the target object. The target object is at least one of the following: the voice call corresponding to the target area, a grid, a sub - area, the target area. The grid and the sub - area are obtained by dividing the target area.

2. The method according to claim 1, characterized in that, before the step of performing segmentation processing on the data of the voice call corresponding to the target area within the preset time period based on the target period to obtain multiple sample data corresponding to multiple time slices, the method further includes: Obtain the historical voice quality measurement data corresponding to the target area in a target historical period, and perform segmentation processing on the historical voice quality measurement data based on the target period to obtain multiple training data corresponding to multiple time slices; one time slice corresponds to one training data; Determine the target performance indicators and historical MOS values corresponding to each training data, and train the target model based on the target performance indicators and the historical MOS values corresponding to each training data.

3. The method according to claim 1, characterized in that, the voice call corresponding to the target area includes multiple voice calls, and each voice call corresponds to at least one sample data. The target object includes: the voice call corresponding to the target area; The step of obtaining the MOS value corresponding to the target object according to the MOS value corresponding to each sample data includes: For any voice call, determine the MOS value corresponding to the voice call based on the MOS value corresponding to each sample data in the at least one sample data corresponding to the voice call and the quantity of the at least one sample data.

4. The method according to claim 3, wherein, the target object includes: a grid; and the method further includes: performing rasterization processing on the target area to obtain a plurality of grids, and determining the position information of each grid; determining the target position information corresponding to each voice call from the data of the voice call, and determining the grid corresponding to each voice call according to the target position information corresponding to each voice call; each grid includes at least one voice call; The obtaining the MOS value corresponding to the target object according to the MOS value corresponding to each sample data includes: for any grid, determining the MOS value corresponding to the grid based on the MOS value corresponding to each voice call in the at least one voice call included in the grid and the first total quantity of the sample data corresponding to the grid.

5. The method according to claim 4, wherein, the target area includes a plurality of sub-areas, each sub-area includes at least one grid of the plurality of grids, the target object includes: a sub-area; the obtaining the MOS value corresponding to the target object according to the MOS value corresponding to each sample data includes: for any sub-area, determining the MOS value corresponding to the sub-area based on the MOS value corresponding to each grid in the at least one grid included in the sub-area and the first weight value corresponding to each grid in the at least one grid included in the sub-area, where the first weight value is the ratio between the first total quantity of the sample data corresponding to each grid and the second total quantity of the sample data corresponding to the sub-area.

6. The method according to claim 4, wherein, the target object includes: the target area; the obtaining the MOS value corresponding to the target object according to the MOS value corresponding to each sample data includes: determining the MOS value corresponding to the target area based on the MOS value corresponding to each grid in the plurality of grids included in the target area and the second weight value corresponding to each grid in the plurality of grids included in the target area, where the second weight value is the ratio between the first total quantity of the sample data corresponding to each grid and the third total quantity of the sample data corresponding to the target area.

7. A voice quality evaluation device, wherein, the voice quality evaluation device includes: a processing unit; the processing unit is configured to perform segmentation processing on the data of the voice call corresponding to the target area within a preset time period based on a target period to obtain a plurality of sample data corresponding to a plurality of time slices; one time slice corresponds to one sample data, the preset time period is the time period before the current moment, and the data of the voice call includes: Trace data, signaling, and XDR voice call records. The processing unit determines the target performance metrics corresponding to each sample data; the target performance metrics include at least one of the following: measurement report (MR) data, the number of radio resource control (RRC) reconstructions, and the voice call duration staying in the target network; the target network includes a first network and a second network, and the voice call data corresponding to the target area includes voice call data based on the first network and voice call data based on the second network. The voice call duration staying in the target network includes: the voice call duration staying in the first network and the voice call duration staying in the second network; wherein, the data type of the voice call data based on the first network is different from the data type of the voice call data based on the second network The processing unit is configured to input the target performance metrics corresponding to each sample data into a target model to obtain the MOS value corresponding to each sample data; the target model is a model constructed based on the historical voice quality measurement data corresponding to the target area in a target historical time period; The processing unit is configured to obtain the MOS value corresponding to the target object according to the MOS value corresponding to each sample data. The MOS value corresponding to the target object is used to evaluate the voice quality of the target object. The target object is at least one of the following: the voice call corresponding to the target area, a grid, a sub-area, and the target area. The grid and the sub-area are obtained by dividing the target area.

8. The voice quality evaluation device according to claim 7, wherein, the voice quality evaluation device further includes: an acquisition unit; The acquisition unit is configured to acquire the historical voice quality measurement data corresponding to the target area in a target historical time period; The processing unit is configured to perform a slicing process on the historical voice quality measurement data based on the target period to obtain multiple training data corresponding to multiple time slices; one time slice corresponds to one training data; The processing unit is configured to determine the target performance metrics and the historical MOS value corresponding to each training data, and train the target model based on the target performance metrics and the historical MOS value corresponding to each training data.

9. The voice quality evaluation device according to claim 7, wherein, The voice call corresponding to the target area includes multiple voice calls, and each voice call corresponds to at least one sample data. The target object includes: the voice call corresponding to the target area; The processing unit is configured to, for any voice call, determine the MOS value corresponding to the voice call based on the MOS value corresponding to each sample data in the at least one sample data corresponding to the voice call and the number of the at least one sample data.

10. The voice quality evaluation device according to claim 9, wherein, The target object includes: a grid; The processing unit is configured to perform a grid processing on the target area to obtain multiple grids, and determine the position information of each grid; The processing unit is configured to determine the target location information corresponding to each voice call from the data of the voice call, and determine the grid corresponding to each voice call according to the target location information corresponding to each voice call; each grid includes at least one voice call. The processing unit is configured to, for any grid, determine the MOS value corresponding to the grid based on the MOS value corresponding to each voice call in at least one voice call included in the grid and the first total number of the sample data corresponding to the grid.

11. The voice quality evaluation device according to claim 10, wherein, the target area includes a plurality of sub-areas, each sub-area includes at least one grid among the plurality of grids, and the target object includes: a sub-area; The processing unit is configured to, for any sub-area, determine the MOS value corresponding to the sub-area based on the MOS value corresponding to each grid in at least one grid included in the sub-area and the first weight value corresponding to each grid in at least one grid included in the sub-area, where the first weight value is the ratio between the first total number of the sample data corresponding to each grid and the second total number of the sample data corresponding to the sub-area.

12. The voice quality evaluation device according to claim 10, wherein, the target object includes: the target area; The processing unit is configured to determine the MOS value corresponding to the target area based on the MOS value corresponding to each grid in the plurality of grids included in the target area and the second weight value corresponding to each grid in the plurality of grids included in the target area, where the second weight value is the ratio between the first total number of the sample data corresponding to each grid and the third total number of the sample data corresponding to the target area.

13. An electronic device, wherein, comprises: a processor and a memory; wherein, the memory is used to store one or more programs, and the one or more programs include computer execution instructions. When the electronic device runs, the processor executes the computer execution instructions stored in the memory, so that the electronic device executes a voice quality evaluation method according to any one of claims 1-6.

14. A computer-readable storage medium storing one or more programs, wherein, the one or more programs include instructions that, when executed by a computer, cause the computer to execute a voice quality evaluation method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Voice quality estimation method and device as well as electronic equipment

    CN104581758A

  • Method and system for speech quality perception evaluation based on speech semantic recognition technology

    CN108877839A

  • System, device and method for non-intrusively estimating in real-time the quality of audio communication signals

    WO2022094702A1